Network security monitoring method and system based on distributed nodes
By constructing a multidimensional phase space reconstruction matrix and performing nonlinear dynamic feature analysis, the problem of capturing nonlinear correlation features between nodes in a distributed network is solved, improving the accuracy of attack pattern monitoring and dynamic tracking capabilities, and enabling precise quantification and analysis of abnormal patterns.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 贵州华谊联盛科技有限公司
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to accurately capture the nonlinear dynamic correlation characteristics between nodes in distributed networks, resulting in insufficient ability to capture implicit attack patterns with non-periodic and progressively mutating characteristics, a high false negative rate, and an inability to effectively integrate the temporal dynamics of node communication with the identifying attributes of protocol interactions, thus affecting the tracking and analysis of attack propagation paths.
By acquiring the interaction records and protocol interaction information between nodes, a multi-dimensional phase space reconstruction matrix is constructed, nonlinear dynamic characteristics are analyzed, a set of chaotic feature values is extracted, an attack mode topological expression vector is generated, the abnormal state propagation process is analyzed, and the ability to dynamically track the attack diffusion path is improved.
It enables the accurate capture and quantification of abnormal communication patterns among nodes in distributed networks, improving the accuracy and dynamic analysis capabilities of network security monitoring and ensuring a comprehensive reflection of the attack propagation process.
Smart Images

Figure CN121585476B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data processing and network security, and more specifically, to a network security monitoring method and system based on distributed nodes. Background Technology
[0002] With the development of distributed network technology, analyzing node communication behavior can identify abnormal patterns to prevent network attacks. In distributed networks, nodes interact with each other through various communication protocols, and the communication behavior between nodes exhibits dynamic temporal characteristics and complex correlations. Accurately capturing these characteristics is crucial for effective monitoring. Existing technologies typically simplify node communication behavior into independent or linearly related events, making it difficult to characterize the nonlinear dynamic correlations formed between nodes due to protocol interactions and temporal evolution. This results in insufficient ability to capture implicit attack patterns with non-periodic and progressively changing characteristics. Furthermore, traditional feature representations have limitations, failing to effectively integrate the temporal dynamics of node communication with the identifying attributes of protocol interactions. This makes it difficult to quantify the spatial distribution of abnormal patterns in the network topology, affecting the tracking and analysis of the attack's propagation path from the source node to the target node. Consequently, this leads to a high false negative rate for complex attacks in network security monitoring, making it difficult to meet the monitoring needs of diverse attack patterns and dynamic propagation paths in distributed network environments. Summary of the Invention
[0003] Based on this, the present invention provides a network security monitoring method and system based on distributed nodes.
[0004] According to one aspect of the present invention, a network security monitoring method based on distributed nodes is provided. The method includes: acquiring node interaction records and identification information during protocol interaction processes of each node in a distributed network within a preset communication period to obtain a node communication dataset containing time-series markers, wherein the node communication dataset contains a communication time sequence between nodes and a fixed field sequence during protocol interaction processes; and constructing a multi-dimensional phase space reconstruction matrix containing communication time sequence correlation features and protocol identification correlation features based on the node communication dataset using a time delay embedding method, wherein the row dimension of the multi-dimensional phase space reconstruction matrix corresponds to the occurrence time of node communication events, and the column dimension corresponds to the communication correlation attributes between different nodes. Nonlinear dynamics feature analysis is performed on the multidimensional phase space reconstruction matrix. A chaotic feature set of node communication behavior is extracted by calculating the evolution complexity index of the phase space trajectory. This chaotic feature set characterizes the nonlinear variation characteristics of node communication patterns. A topological association structure of node communication features is constructed through local linear relationships. The chaotic feature set is mapped to a topological space of a preset dimension, generating an attack pattern topological expression vector containing the geometric distribution characteristics of abnormal node communication patterns. Based on this attack pattern topological expression vector, the abnormal state propagation process in the distributed network is analyzed through inter-node energy propagation rules, yielding network security detection results including attack diffusion paths and the energy attenuation distribution of each node.
[0005] According to another aspect of the present invention, a computer system is provided, comprising: a processor; and a memory, wherein the memory stores computer-readable code, which, when executed by the processor, causes the processor to perform the method as described above.
[0006] This invention acquires a node communication dataset containing communication time-series sequences and protocol fixed field sequences. By combining this with a time-delay embedding method to construct a multi-dimensional phase space reconstruction matrix, it transforms the complex communication relationships between nodes in a distributed network into a structured spatial feature matrix. This allows the originally fragmented communication data to form a joint feature expression with temporal dynamics and protocol identification attributes. By performing nonlinear dynamic feature analysis on the multi-dimensional phase space reconstruction matrix to extract a set of chaotic feature values, it can accurately capture the nonlinear and non-periodic changes in node communication patterns, improving the ability to detect gradual communication mutations in implicit attacks and avoiding missed detections due to reliance on fixed thresholds or rule matching. By constructing a topological association structure and mapping the chaotic feature value set to the topological space to generate attack pattern topological expression vectors, it achieves the transformation of abstract nonlinear features into geometrically distributed features. This allows abnormal node communication patterns to form quantifiable distribution relationships in the topological space, facilitating intuitive analysis of the spatial clustering and association characteristics of abnormal patterns, and solving the technical problem that traditional features are difficult to associate with node topological relationships. Based on the attack pattern topology representation vector, the abnormal state propagation process is analyzed through the energy propagation rules between nodes. This can dynamically simulate the evolution path and energy decay law of attacks in distributed networks, improve the dynamic tracking capability of attack diffusion process, and ensure that the generated network security detection results can fully reflect the propagation path of the attack and the degree of node impact, effectively improving the accuracy and dynamic analysis capability of distributed network security monitoring. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of an application scenario provided by the present invention;
[0008] Figure 2 This is a flowchart illustrating a network security monitoring method based on distributed nodes provided by the present invention.
[0009] Figure 3 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0010] To facilitate a clearer understanding of this invention, we will first introduce the application scenarios of the network security monitoring method based on distributed nodes, such as... Figure 1 As shown, the application scenario of this invention includes a computer system 10 and a node cluster. The node cluster may include one or more distributed nodes; the number of distributed nodes will not be limited here. Figure 1As shown, the node cluster may specifically include distributed node 1, distributed node 2, ..., distributed node n; it can be understood that distributed node 1, distributed node 2, distributed node 3, ..., distributed node n can all be connected to the computer system 10 via a network so that each distributed node can interact with the computer system 10 via the network connection.
[0011] It is understood that computer system 10 can refer to a device that executes the network security monitoring method based on distributed nodes provided in the embodiments of the present invention. Computer system 10 can be, for example, a server, a single physical server, or a server cluster or distributed system consisting of at least two physical servers. Distributed nodes can specifically refer to smartphones, tablets, laptops, desktop computers, etc., but are not limited to these. Each distributed node and computer system 10 can be directly or indirectly connected via wired or wireless communication. Furthermore, the number of distributed nodes and computer system 10 can be one or at least two; the present invention does not impose any limitation on this.
[0012] Further, please see Figure 2 This is a flowchart illustrating a network security monitoring method based on distributed nodes provided in an embodiment of the present invention. Figure 2 As shown, this method can be derived from... Figure 1 The network security monitoring method based on distributed nodes is executed by computer system 10, and may include the following steps:
[0013] Step S100: Obtain the node interaction records and identification information in the protocol interaction process of each node in the distributed network within a preset communication period, and obtain a node communication dataset containing time series markers. The node communication dataset contains the communication time sequence between nodes and the fixed field sequence in the protocol interaction process.
[0014] A preset communication period is a pre-defined time range used to define the time period for data collection. Its setting can be flexibly determined based on factors such as network operation patterns and monitoring needs; for example, it could be the length of a workday or a peak business period within a week. Inter-node interaction records are records generated when nodes in a distributed network communicate. They cover the basic elements of communication, such as the initiating node, the receiving node, and the start and end times of the communication. These records reflect the communication activities between nodes. Identification information during protocol interaction is information that serves an identifying function during protocol interaction. The protocol type identifier clarifies the specific protocol used for communication, such as HTTP or TCP. Status codes reflect the status of the protocol interaction; different status codes represent different interaction results, such as success, failure, and not found. Time series markers are time-dimension markers added to the acquired data. Their purpose is to give the data a clear chronological order, which helps in subsequent analysis of the temporal relationships of communication. The node communication dataset is a dataset that integrates the acquired inter-node interaction records and protocol interaction identification information. The communication time sequence between nodes reflects the chronological order and interval of node communication, reflecting the time pattern of communication. The fixed field sequence in the protocol interaction process is a sequence of fixed-format fields in the protocol interaction, which contains key information of the protocol interaction.
[0015] To acquire this data, data acquisition devices can be deployed on various nodes of the distributed network. These devices can monitor the communication interfaces of the nodes in real time, capturing interaction records and protocol information between nodes. Identification information during protocol interactions can be extracted by parsing protocol data packets. Specifically, network packet parsing tools are used to accurately extract protocol type identifiers and status codes from the data packets according to the protocol format specifications. After acquiring the data, time-series markers are added using the system's time recording function, ultimately forming a node communication dataset containing time-series markers.
[0016] Step S200: Based on the node communication dataset, construct a multi-dimensional phase space reconstruction matrix containing communication timing association features and protocol identifier association features using the time delay embedding method. The row dimension of the multi-dimensional phase space reconstruction matrix corresponds to the occurrence time of node communication events, and the column dimension corresponds to the communication association attributes between different nodes.
[0017] Time-delay embedding can transform one-dimensional time-series data into multi-dimensional phase-space data, thereby revealing the potential dynamic structure within the time-series data. Communication temporal correlation features reflect the temporal interrelationships of node communication, such as the regularity of communication time intervals and whether communication is periodic. Protocol identifier correlation features are related to protocol identifier information, reflecting the association between different protocol type identifiers and the changing patterns of status codes. The multi-dimensional phase-space reconstruction matrix is a multi-dimensional matrix structure. Its row dimension represents the specific time of occurrence of node communication events, with each row corresponding to a time point of a communication event; the column dimension corresponds to the communication correlation attributes between different nodes, such as the communication frequency between different nodes and the protocol type used in the communication.
[0018] In one implementation, step S200 may specifically include the following steps S210 to S260:
[0019] Step S210: Parse the fixed field sequence of the protocol interaction process in the node communication data set, extract the protocol type identifier, status code flow sequence and field length variation coefficient, and generate a protocol identifier original tuple containing protocol behavior characteristics. Each element of the protocol identifier original tuple corresponds to the attribute description of a fixed field in the protocol interaction process.
[0020] Parsing the fixed field sequence of the protocol interaction process in the node communication data involves a detailed analysis and interpretation of this sequence according to the protocol's format specifications. The protocol type identifier clarifies the specific type of the protocol; different protocol types have different functions and characteristics. The status code transition sequence is the sequence of status codes changing over time during protocol interaction, recording the state transitions from the start to the end of the interaction and reflecting the execution status of the protocol interaction. The field length variation coefficient is an indicator that measures the degree of change in the length of fixed fields during protocol interaction, reflecting the stability of the protocol data. The protocol identifier raw tuple is a tuple structure containing protocol behavioral characteristics, where each element corresponds to an attribute description of a fixed field during protocol interaction, such as the protocol type identifier or status code. This tuple comprehensively describes the protocol's interaction behavior.
[0021] To parse a fixed-field sequence, regular expression matching can be used. Based on the protocol's format specifications, corresponding regular expressions are written, and their matching capabilities are used to accurately extract information such as the protocol type identifier and status code from the fixed-field sequence. For calculating the coefficient of variation (COP) of field lengths, the length of each fixed field is first calculated, and then the COP is calculated by analyzing the dispersion of these length data. For example, the COP is obtained by comparing the fluctuation range of the length data with the average length. Finally, the extracted protocol type identifier, status code sequence, and COP are combined to form the original protocol identifier tuple.
[0022] Step S220: Divide the communication time sequence in the node communication dataset into hierarchical time windows, segment the continuous communication events into protocol-specific time window units according to the protocol type identifier, calculate the communication frequency density and data packet interval volatility in each protocol-specific time window unit, and generate protocol-differentiated communication time sequence units. The time granularity of the protocol-differentiated communication time sequence units matches the average period of protocol interaction.
[0023] Layered time window partitioning divides a communication time sequence into multiple time windows according to certain rules, with each window containing a certain number of communication events. This aims to analyze the temporal characteristics of communication in greater detail. Protocol-specific time window units are time windows obtained by segmenting continuous communication events based on protocol type identifiers. Each window contains only communication events of the same protocol type, facilitating separate analysis of communication for different protocols. Communication frequency density refers to the number of communication events per unit time within a protocol-specific time window unit, reflecting the level of interactive activity of that protocol during that time period. Packet interval volatility is an indicator that measures the stability of the time interval between packets during protocol interaction, reflecting the fluctuation of the protocol interaction time interval. Protocol-differentiated communication time sequence units are communication time sequence units generated according to the characteristics of different protocols. Their time granularity matches the average period of protocol interaction, thus better reflecting the communication patterns of different protocols.
[0024] In one implementation, step S220 may specifically include the following steps S221 to S226:
[0025] Step S221: Extract all protocol type identifiers from the node communication dataset, count the number of interactions and the duration of each interaction for each protocol type within a preset communication period, calculate the average interaction period for each protocol, and generate a protocol period feature table. The protocol period feature table contains the mapping relationship between protocol types and average interaction periods.
[0026] Extracting all protocol type identifiers from the node communication dataset involves parsing the protocol-related information in the dataset and extracting all protocol type identifiers. The number of interactions refers to the number of communication interactions for each protocol type within a preset communication period, reflecting the frequency of use of that protocol during that time period. The duration of a single interaction refers to the time spent from the start to the end of each protocol interaction, reflecting the time consumption of each interaction. The average interaction period is the average of all interaction durations for each protocol type within the preset communication period, reflecting the temporal pattern of the protocol's interactions. The protocol period feature table is a table containing the mapping relationship between protocol types and average interaction periods. This table allows for quick lookup of the average interaction period for each protocol, providing a basis for subsequent analysis.
[0027] Protocol type identifiers can be extracted using regular expression matching from protocol data packets in the node communication dataset. The number of interactions can be counted by counting communication events for each protocol type, i.e., counting each communication record for each protocol type individually. For calculating the duration of a single interaction, the start and end times of each protocol interaction can be recorded, and the difference between the two can be calculated to obtain the duration of a single interaction. To calculate the average interaction cycle, the total interaction duration for each protocol is divided by the number of interactions to obtain the average interaction cycle for that protocol. Finally, the protocol type and corresponding average interaction cycle for each protocol are recorded in the protocol cycle feature table.
[0028] Step S222: Based on the protocol cycle feature table, set a dedicated time window length for each protocol type. The window length is a preset multiple of the average interaction cycle of the corresponding protocol, so that each window can fully cover at least one protocol interaction process.
[0029] The protocol cycle characteristic table is a table generated in the previous step that contains the mapping relationship between protocol types and average interaction cycles. The dedicated time window length is the length of a time window set individually for each protocol type. The preset multiple is a pre-determined multiple, and its value can be determined based on factors such as the characteristics of the protocol and the required accuracy of the analysis. Setting the dedicated time window length to a preset multiple of the average interaction cycle of the corresponding protocol ensures that each time window can completely cover at least one protocol interaction process. This allows for more accurate analysis of the protocol's communication patterns and avoids the inability to fully observe a protocol interaction due to an excessively short window length.
[0030] Based on the protocol cycle characteristic table, the average interaction cycle of each protocol is found, and then multiplied by a preset multiple to obtain the dedicated time window length. For example, if the preset multiple is an appropriate value, the dedicated time window length will be correspondingly longer for protocols with longer average interaction cycles, and relatively shorter for protocols with shorter average interaction cycles.
[0031] Step S223: Traverse the communication time sequence in the node communication dataset, assign continuous communication events to the dedicated time window unit of the corresponding protocol according to the protocol type identifier, and assign communication events that cross the window boundary to the adjacent window according to the time proportion.
[0032] Traversing the communication time sequence in the node communication dataset involves processing each communication event sequentially according to its occurrence time. Based on the protocol type identifier, consecutive communication events are assigned to the corresponding protocol-specific time window units. This separates communication events from different protocols, facilitating individual analysis of each protocol's communication behavior. For communication events that cross window boundaries—that is, events whose occurrence time spans two adjacent time windows—they are assigned to adjacent windows based on their time proportion. For example, if a communication event lasts a large proportion of its total duration in the first window, then the majority of that event is assigned to the first window; if it lasts a large proportion in the second window, then the majority is assigned to the second window. This approach allows for a more accurate count and analysis of the number and characteristics of communication events within each window.
[0033] For example, allocation can be performed by comparing the start and end times of a communication event with the boundary times of each dedicated time window. For communication events that span multiple windows, the proportion of time they occupy within each window is calculated, and then the relevant data for that communication event is allocated to adjacent windows based on that proportion.
[0034] Step S224: Calculate the total number of communication events within each protocol-specific time window unit, divide by the window length to obtain the communication frequency density, which reflects the activity level of protocol interaction per unit time.
[0035] Calculating the total number of communication events within each protocol-specific time window involves counting all communication events within that window. The window length refers to the time span of each dedicated time window. Dividing the total number of communication events by the window length yields the communication frequency density. Communication frequency density reflects the activity level of protocol interactions per unit of time; a higher frequency density indicates more frequent interactions and higher activity within that time period, while a lower frequency density indicates relatively less interaction and lower activity. To calculate the total number of communication events, each communication record within each dedicated time window can be counted individually. Then, dividing the total count by the window length gives the communication frequency density.
[0036] Step S225: Extract the transmission interval sequence of all data packets within each window, and calculate the relationship between the dispersion of the transmission interval sequence and the average interval as the data packet interval volatility. This volatility reflects the stability of the protocol interaction time interval.
[0037] Extracting the transmission interval sequence of all data packets within each window involves obtaining the transmission time interval data between adjacent data packets from the communication data within each protocol-specific time window unit, forming a sequence. The dispersion of the transmission interval sequence is calculated in relation to the average interval, which is then used as the data packet interval volatility. Dispersion can be measured by analyzing the fluctuation range and variance of the transmission interval data, while the average interval is the average value of the transmission interval sequence. The data packet interval volatility reflects the stability of the protocol interaction time interval; a higher volatility indicates greater fluctuation and poorer stability in the protocol interaction time interval, while a lower volatility indicates more stable time intervals. When extracting the transmission interval sequence, the transmission times of data packets within each window can be sorted, and the time difference between adjacent data packets can be calculated to obtain the transmission interval sequence. Then, the dispersion and average interval of this sequence are analyzed, and the data packet interval volatility is obtained by comparing the relationship between the two.
[0038] Step S226: Arrange the protocol-specific time window units for each protocol in chronological order, and associate each window unit with the corresponding communication frequency density and data packet interval volatility to generate protocol-differentiated communication timing units, whose time granularity is consistent with the average interaction period of the corresponding protocol.
[0039] Arranging the protocol-specific time window units for each protocol in chronological order means sorting the time window units for each protocol according to the order in which communication events occur. Each window unit is associated with a corresponding communication frequency density and packet interval volatility. This binds the previously calculated communication frequency density and packet interval volatility of each window to that window unit, providing a more comprehensive description of the communication characteristics of each window. Protocol-differentiated communication timing units are generated, with the time granularity of these units consistent with the average interaction period of the corresponding protocol. This means that the length of each time unit matches the average interaction period of the protocol, better reflecting the communication patterns of different protocols.
[0040] For example, a data structure can be used to store the dedicated time window unit for each protocol and its corresponding communication frequency density and packet interval volatility, and then these data can be sorted in chronological order to finally generate protocol-differentiated communication timing units.
[0041] Step S230: Determine the dynamic delay parameter of the time delay embedding method based on the protocol status code flow sequence. Increase the delay parameter when the protocol status code changes and decrease the delay parameter when the protocol status code is continuous and stable, thereby generating a delay parameter sequence that dynamically adjusts with the protocol status.
[0042] The protocol status code transition sequence records the changes in status codes over time during protocol interaction, reflecting the execution status and flow of the protocol interaction. The dynamic delay parameter in the time delay embedding method is used to adjust the phase space construction, determining the time interval between adjacent time points. When a protocol status code transitions, it indicates a significant change in the protocol interaction state. Increasing the delay parameter at this time allows for a more comprehensive capture of information before and after the state change, enhancing the capture of temporal correlations during state changes. When the protocol status codes are continuous and stable, it indicates a relatively stable protocol interaction state. Decreasing the delay parameter avoids introducing excessive redundant timing information, improving data processing efficiency. The dynamically adjusted delay parameter sequence is a sequence generated by dynamically adjusting the delay parameters based on the changes in the protocol status codes. Each parameter in this sequence corresponds to a time point of a communication event.
[0043] In one implementation, step S230 may specifically include the following steps S231 to S236:
[0044] Step S231: Parse the status code transition sequence in the original tuple of the protocol identifier, extract the timestamp of the status code transition event and the status code values before and after the transition, and construct a status code transition event list, where each entry contains the transition time, the original status code, and the new status code.
[0045] Parsing the status code transition sequence in the original protocol identifier tuple involves detailed analysis and processing of this sequence. A status code transition event is an event in which the protocol status code changes. Extracting the timestamp of a status code transition event clarifies the specific time the transition occurred, and extracting the status code values before and after the transition reveals the changes in the status code. A status code transition event list is constructed, recording relevant information for each event. Each entry includes the transition time, the original status code, and the new status code, facilitating subsequent analysis and processing of these events.
[0046] When parsing the status code transition sequence, the changes in the status codes can be checked sequentially according to the sequence. When a change in the status code is found, the timestamp of that moment and the status code values before and after the transition are recorded, and this information is added to the status code transition event list.
[0047] Step S232: Calculate the time interval of all status code transition events within the preset communication period, calculate the average transition interval, and set its preset ratio as the basic delay parameter as the initial value of the dynamic delay parameter.
[0048] Calculating the time interval between all status code transition events within a preset communication period involves determining the time difference between adjacent status code transition events. The average transition interval is the average of the time intervals between all status code transition events, reflecting the average frequency of status code transitions. A preset proportion of the average transition interval is set as the basic delay parameter. This preset proportion is a manually determined value. The basic delay parameter obtained in this way serves as the initial value for the dynamic delay parameter, providing a starting point for dynamically adjusting the delay parameter based on the protocol status.
[0049] For example, the timestamp of each status code transition event can be obtained first from the status code transition event list. Then, the difference between adjacent timestamps can be calculated to obtain the time interval of the status code transition events. Next, all time intervals are summed and divided by the number of transition events to obtain the average transition interval. Finally, the average transition interval is multiplied by a preset ratio to obtain the basic delay parameter.
[0050] Step S233: Traverse the status code transition sequence. When a status code transition event is detected, add the relative relationship between the duration of the continuous stable status codes before the transition and the average transition interval to the current delay parameter based on the basic delay parameter, thereby enhancing the capture of the temporal correlation when the state changes.
[0051] Traversing the status code transition sequence involves examining each status code sequentially according to its chronological order. When a status code transition event is detected, it signifies a change in the protocol's state. The duration of consecutive stable status codes before the transition refers to the duration during which the status code remains unchanged before the transition. Adding the relative relationship between the duration of consecutive stable status codes and the average transition interval to the current delay parameter, based on the base delay parameter, aims to more fully capture the temporal correlation before and after a status code transition. Because a transition occurring after a longer duration of consecutive stable status codes may indicate a more significant state transition, increasing the delay parameter can obtain more relevant information.
[0052] For example, the stable duration of each status code can be recorded while traversing the status code transition sequence. When a status code transition is detected, the relative value between this duration and the average transition interval is calculated, and then this relative value is added to the current delay parameter.
[0053] Step S234: When the duration of continuous detection of the same status code without transition exceeds a preset multiple of the average transition interval, the current delay parameter is reduced by a preset ratio of the duration to the average transition interval based on the basic delay parameter to avoid redundant timing information in a stable state.
[0054] When the same status code is detected consecutively without a transition, it indicates that the protocol interaction is in a stable state. If the duration of this stable state exceeds a preset multiple of the average transition interval, it means that the protocol has been in a stable state for a relatively long period of time, and there may be a lot of redundant timing information. By reducing the current delay parameter by a preset ratio of the duration to the average transition interval from the base delay parameter, unnecessary information collection can be reduced, and data processing efficiency can be improved.
[0055] For example, when traversing the status code transition sequence, the duration of consecutive occurrences of the same status code can be recorded. When the duration exceeds a preset multiple of the average transition interval, the ratio of the duration to the average transition interval is calculated, and then this ratio is multiplied by a preset proportion and subtracted from the current delay parameter.
[0056] Step S235: Apply upper and lower limit constraints to the adjusted delay parameters. The lower limit is set to a preset ratio of the basic delay parameters, and the upper limit is set to a preset multiple of the basic delay parameters to prevent the phase space reconstruction from being distorted due to excessively large or small delay parameters.
[0057] The purpose of imposing upper and lower limits on the adjusted delay parameters is to ensure that the delay parameters are within a reasonable range. The lower limit is set as a preset proportion of the base delay parameter, and the upper limit is set as a preset multiple of the base delay parameter. This avoids the delay parameter being too small to capture sufficient timing information, and also prevents the delay parameter being too large to introduce excessive noise and redundant information during phase space reconstruction, thereby ensuring the accuracy and effectiveness of phase space reconstruction.
[0058] For example, the adjusted delay parameter can be set to the lower limit when it is less than the lower limit, and set to the upper limit when it is greater than the upper limit.
[0059] Step S236: Record the dynamic delay parameters corresponding to each communication event in chronological order, and generate a dynamic delay parameter sequence with the same length as the communication time sequence.
[0060] Recording the dynamic delay parameters corresponding to each communication event in chronological order involves sequentially recording the adjusted delay parameters for each communication event according to their order of occurrence. This generates a dynamic delay parameter sequence with the same length as the communication time sequence. This ensures that each communication event has a corresponding delay parameter, providing accurate parameters for subsequently constructing a multidimensional phase space reconstruction matrix using the time delay embedding method. For example, while traversing communication events, the dynamic delay parameters corresponding to each communication event can be added to a sequence in chronological order, ultimately forming the dynamic delay parameter sequence.
[0061] Step S240: Map the protocol-differentiated communication timing units to the phase space using the time delay embedding method, and construct a basic timing correlation matrix by combining the dynamic delay parameter sequence. The row vectors of the matrix correspond to the node communication times under different protocol states, and the column vectors correspond to the communication feature values under different delays.
[0062] Protocol-differentiated communication timing units contain the temporal communication characteristics of different protocols. Mapping them to phase space allows for a more intuitive analysis of the temporal correlations of communication. The dynamic delay parameter sequence, generated in the previous steps, dynamically adjusts the delay parameters according to the protocol state. A basic temporal correlation matrix is constructed by combining the dynamic delay parameter sequence. The row vectors of the matrix correspond to the node communication times under different protocol states, reflecting the time points of communication events; the column vectors correspond to communication characteristic values under different delays. These characteristic values can be communication frequency density, packet interval volatility, etc., thus comprehensively describing the temporal correlation characteristics of node communication. For example, a time delay embedding operation can be performed on the protocol-differentiated communication timing units based on the dynamic delay parameter sequence. Each communication event is processed according to its corresponding delay parameter, converting it into a vector in phase space. These vectors are then combined to form the basic temporal correlation matrix.
[0063] Step S250: Convert the original protocol identifier tuple into a numerical protocol identifier vector, calculate the weight coefficient of each field using the protocol field importance weighting algorithm, and perform weighted expansion on the protocol identifier vector based on the weight coefficient to generate a weighted protocol identifier matrix that matches the column dimensions of the basic time series correlation matrix.
[0064] The original protocol identifier tuple contains various identifier information during protocol interaction. Converting it into a numerical protocol identifier vector facilitates mathematical calculations and analysis. This can be achieved by encoding each element of the original protocol identifier tuple, converting it into a numerical form to form the numerical protocol identifier vector. The protocol field importance weighting algorithm is used to calculate the weight coefficients of protocol fields. Different fields may have different importance in protocol interaction; this algorithm determines the weight of each field. Weighted expansion of the protocol identifier vector based on these weight coefficients involves multiplying each element of the protocol identifier vector by its corresponding weight coefficient and then expanding it so that its column dimensions match the basic temporal correlation matrix. The generated weighted protocol identifier matrix contains the weighted features of the protocol identifier information. Combined with the basic temporal correlation matrix, it can more comprehensively reflect the characteristics of node communication.
[0065] For example, the original protocol identifier tuples can be encoded first, converting them into numerical protocol identifier vectors. Then, a protocol field importance weighting algorithm is used, such as calculating the weight coefficient of each field based on factors like its frequency of occurrence in the protocol and its influence on the protocol interaction results. Finally, the protocol identifier vector is multiplied by the weight coefficients to expand it, generating a weighted protocol identifier matrix.
[0066] Step S260: Concatenate the basic temporal correlation matrix and the weighted protocol identifier matrix along the column dimension, retain the feature dimensions with cumulative contribution rates reaching a preset threshold through incremental singular value decomposition, and generate a multi-dimensional phase space reconstruction matrix with row dimensions corresponding to the occurrence time of node communication events and column dimensions corresponding to the communication correlation attributes between different nodes.
[0067] Concatenating the basic temporal correlation matrix and the weighted protocol identifier matrix along their column dimensions merges the two matrices, integrating their information. The basic temporal correlation matrix contains the temporal correlation features of node communication, while the weighted protocol identifier matrix contains the weighted features of protocol identifiers. The concatenated matrix provides a more comprehensive reflection of node communication. Incremental singular value decomposition (SVD) is a matrix factorization method that can reduce the dimensionality of the concatenated matrix. The cumulative contribution rate refers to the sum of the contribution proportions of each feature dimension to the original matrix information after SVD. The preset threshold is a pre-defined proportion; by retaining feature dimensions with a cumulative contribution rate reaching the preset threshold, some unimportant information in the matrix can be removed, reducing data redundancy while retaining the main feature information. The generated multidimensional phase space reconstruction matrix has rows corresponding to the occurrence times of node communication events and columns corresponding to the communication correlation attributes between different nodes. This matrix can be used for subsequent nonlinear dynamics feature analysis. For example, the basic temporal correlation matrix and the weighted protocol identifier matrix can be concatenated along their column dimensions using a matrix concatenation function. Then, the concatenated matrix is decomposed using the incremental singular value decomposition algorithm, the contribution rate of each feature dimension is calculated, and the contribution rates are accumulated. When the accumulated contribution rate reaches a preset threshold, the feature dimensions are no longer retained, and the multidimensional phase space reconstruction matrix is finally obtained.
[0068] Step S300: Perform nonlinear dynamic characteristic analysis on the multidimensional phase space reconstruction matrix, and extract the chaotic feature set of node communication behavior by calculating the evolution complexity index of the phase space trajectory. The chaotic feature set is used to characterize the nonlinear change characteristics of the node communication mode.
[0069] Nonlinear dynamics feature analysis is a method used to study the dynamic behavior of complex systems. The multidimensional phase space reconstruction matrix contains various characteristic information of node communication, and its nonlinear dynamics feature analysis can reveal the potential patterns in node communication behavior. The phase space trajectory refers to the trajectory formed by the change of node communication states over time in phase space. Calculating the evolutionary complexity index of the phase space trajectory can measure the complexity and variability of the trajectory. The chaotic feature set of node communication behavior is obtained by extracting the evolutionary complexity index of the phase space trajectory. These feature sets can characterize the nonlinear variation characteristics of node communication patterns, such as whether the communication pattern has randomness or periodicity.
[0070] In one implementation, step S300 may specifically include the following steps S310 to S360:
[0071] Step S310: Perform trajectory segmentation preprocessing on the multidimensional phase space reconstruction matrix. Based on the timestamp information of the matrix row dimension, the matrix is divided into multiple state segment matrices according to the timestamp of the protocol status code jump event. Each state segment matrix corresponds to a communication period when the protocol status code is stable.
[0072] The row dimension of the multidimensional phase space reconstruction matrix contains timestamps of node communication events, while the timestamps of protocol status code transition events indicate changes in protocol interaction states. Based on this information, the multidimensional phase space reconstruction matrix undergoes trajectory segmentation preprocessing, dividing the matrix into multiple state segment matrices according to the timestamps of protocol status code transition events. Each state segment matrix corresponds to a communication period where the protocol status code is stable. The purpose of this is to divide the node communication process according to protocol states, facilitating separate analysis of communication behavior under different states.
[0073] In one implementation, step S310 may specifically include the following steps S311 to S316:
[0074] Step S311: Extract the timestamp information of the row dimension of the multidimensional phase space reconstruction matrix, construct a timestamp sequence, and each element in the timestamp sequence corresponds to the time of occurrence of the communication event in a row of the matrix.
[0075] Extracting timestamp information from the row dimensions of the multidimensional phase space reconstruction matrix involves retrieving the timestamp of each communication event from the matrix rows. A timestamp sequence is constructed, arranging these timestamps sequentially according to the matrix rows, so that each element in the timestamp sequence corresponds to the time of the communication event in one row of the matrix. This timestamp sequence provides the temporal basis for subsequent timestamp segmentation of the matrix based on protocol status code transition events.
[0076] For example, matrix operation functions can be used to extract timestamp information from the row dimensions of a multidimensional phase space reconstruction matrix and add it to a sequence to ultimately form a timestamp sequence.
[0077] Step S312: Extract the status code flow sequence from the original protocol identifier tuple, align the status code flow sequence with the timestamp sequence to ensure that each status code corresponds to a unique timestamp, and generate a status code-timestamp lookup table.
[0078] Extracting the status code transition sequence from the original protocol identifier tuple involves extracting the sequence of status codes changing over time during protocol interaction. Time-aligning the status code transition sequence with the timestamp sequence ensures that each status code has a unique timestamp, allowing for accurate determination of the communication time corresponding to each status code. A status code-timestamp lookup table is generated, recording the correspondence between status codes and timestamps, providing accurate information for subsequent timestamp segmentation of the status code transition event matrix.
[0079] For example, by comparing the time information in the status code flow sequence and the timestamp sequence, the status code can be matched with the corresponding timestamp, and a status code-timestamp lookup table can be generated.
[0080] Step S313: Traverse the status code-timestamp lookup table, identify the start and end timestamps of consecutive identical status codes, calculate the difference between the start and end timestamps as the status code stability duration, and generate a status code stability period list. Each entry in the list contains the status code value, start timestamp, and end timestamp.
[0081] Traversing the status code-timestamp lookup table involves sequentially checking each status code and its corresponding timestamp. Identifying the start and end timestamps of consecutive identical status codes reveals the start and end times of the time period in which the status code remains unchanged. The difference between the start and end timestamps is calculated to obtain the status code stability duration, reflecting the length of time the protocol state remains stable under a given status code. A list of stable status code periods is generated, where each entry contains the status code value, start timestamp, and end timestamp. This list clearly demonstrates the stability of the protocol state across different time periods.
[0082] For example, when traversing the status code-timestamp lookup table, the start timestamp of consecutive identical status codes can be recorded. When the status code changes, the end timestamp can be recorded, and the difference can be calculated to obtain the stable duration. The status code value, start timestamp, and end timestamp can then be added to the status code stable period list.
[0083] Step S314: Based on the start and end timestamps in the stable period list of status codes, extract the corresponding row range of the sub-matrix in the multi-dimensional phase space reconstruction matrix. Each sub-matrix is a state segment matrix, and its row dimension corresponds to the time when the communication event occurs within the stable period of the status code.
[0084] Based on the start and end timestamps in the status code stable period list, the corresponding row range is found in the multidimensional phase space reconstruction matrix. Then, a submatrix is extracted from this row range; each submatrix is a state segment matrix. The row dimensions of the state segment matrix correspond to the occurrence times of communication events within the status code stable period. This allows the multidimensional phase space reconstruction matrix to be divided according to the protocol status code stable periods, facilitating the analysis of communication behavior under different states. For example, a matrix truncation function can be used to extract the corresponding row range from the multidimensional phase space reconstruction matrix based on the start and end timestamps in the status code stable period list, thus obtaining the state segment matrix.
[0085] Step S315: Perform row dimension verification on each state segment matrix. If the number of rows in the matrix is less than the minimum row number threshold set according to the protocol interaction characteristics, it is merged with the adjacent state segment matrix. The merged state segment matrix corresponds to two consecutive stable state code periods.
[0086] Performing row dimension verification on each state segment matrix involves checking whether the number of rows in each matrix meets the requirements. The minimum row count threshold, set based on the protocol interaction characteristics, is a predetermined value determined according to the protocol's features and analytical needs. If a state segment matrix has fewer rows than the minimum threshold, it indicates that the matrix contains insufficient communication event information and may not accurately reflect the communication behavior in that state. In this case, it is merged with adjacent state segment matrices. The merged state segment matrix corresponds to two consecutive stable state code periods, thus increasing the number of rows in the matrix and improving the accuracy of the analysis.
[0087] For example, the number of rows in each state segment matrix can be obtained using a matrix row count function and then compared with a minimum row count threshold. If the number of rows is less than the threshold, the matrix is merged with the adjacent state segment matrix using a matrix merge function.
[0088] Step S316: Add a status code label to each status segment matrix. The label value is the status code value of the corresponding stable status code period. Generate a set of status segment matrices containing status code labels. Each matrix corresponds to a stable communication period of the protocol status code.
[0089] Adding a status code label to each state segment matrix means assigning a status code value as a label to each state segment matrix. The label value is the status code value of the corresponding stable period of the status code, which makes it clear which protocol state each state segment matrix corresponds to. A set of state segment matrices containing status code labels is generated. Each matrix in this set corresponds to a communication period of stable protocol status codes, which facilitates subsequent classification and analysis of communication behavior under different protocol states.
[0090] For example, a status code label column can be added to each status segment matrix, and the status code value corresponding to the stable period of the status code can be filled into the column, so as to finally generate a set of status segment matrices containing status code labels.
[0091] Step S320: For each state segment matrix, its row vectors are regarded as trajectory points in phase space, and connected in time order to form a state segment trajectory line. The tangent vector sequence of the trajectory line is calculated, and the trajectory bending intensity is determined by the rate of change of the angle between adjacent tangent vectors, thus generating a trajectory bending intensity sequence.
[0092] For each state segment matrix, its row vectors are treated as trajectory points in phase space, since the row vectors of the state segment matrix correspond to the states of node communication events at different times, and can be considered as points in phase space. These trajectory points are connected in chronological order to form a state segment trajectory line, which reflects the changes in the node communication state during the stable period of the state code. The tangent vector sequence of the trajectory line is calculated; the tangent vector represents the tangent direction of the trajectory line at a certain point. By calculating the tangent vector at each trajectory point, the direction change of the trajectory line can be understood. The trajectory bending intensity is determined by the rate of change of the angle between adjacent tangent vectors. The larger the rate of change of the angle, the greater the curvature of the trajectory line, and the higher the trajectory bending intensity; conversely, the smaller the rate of change of the angle, the lower the trajectory bending intensity. A trajectory bending intensity sequence is generated, which records the bending intensity of the trajectory line at different positions, providing a basis for subsequent segmentation and analysis of the trajectory line. For example, the tangent vectors corresponding to the row vectors of the state segment matrix can be calculated using vector calculation methods, and then the rate of change of the angle between adjacent tangent vectors can be calculated. These rates of change are recorded sequentially to form the trajectory bending intensity sequence.
[0093] Step S330: Based on the peak points of the trajectory bending intensity sequence, the trajectory line of the state segment is segmented twice to obtain multiple micro-trajectory segments. The length, mean curvature and endpoint distance of each micro-trajectory segment are calculated to generate a geometric feature set of micro-trajectory segments.
[0094] In one implementation, step S330 may specifically include the following steps S331 to S336:
[0095] Step S331: Smooth the trajectory bending intensity sequence by eliminating high-frequency noise based on an improved moving average filter, preserving the overall trend of the sequence, and generating a smooth bending intensity sequence.
[0096] Smoothing the trajectory curvature intensity sequence removes high-frequency noise, making the sequence smoother and facilitating accurate peak point identification. An improved moving average filter, an enhancement of the traditional moving average filter, more effectively eliminates high-frequency noise based on the sequence's characteristics while preserving its overall trend. This generates a smoothed curvature intensity sequence that, while removing noise, still reflects the overall variation in trajectory curvature intensity.
[0097] For example, an improved moving average filter can be used to filter the trajectory bending intensity sequence. This filter performs a weighted average of each point in the sequence and its neighbors, adjusting the weights according to an improved rule to obtain a smooth bending intensity sequence.
[0098] Step S332: Set the bending strength threshold to a preset ratio of the maximum value of the smooth bending strength sequence, identify all peak points in the sequence that exceed the threshold, record the time index of the peak points, and generate a peak point index list.
[0099] A bending intensity threshold is set as a preset proportion of the maximum value of the smooth bending intensity sequence. This preset proportion is a manually determined value. The bending intensity threshold determined in this way can serve as a standard for judging peak points. All peak points exceeding the threshold in the sequence are identified. These peak points represent locations with high trajectory bending intensity and are crucial for dividing micro-trajectory segments. The time index of each peak point is recorded, i.e., the position of each peak point in the sequence, generating a peak point index list. This list provides accurate positional information for subsequent segmentation of the state segment trajectory line based on the peak points.
[0100] For example, the maximum value of the smooth bending intensity sequence can be calculated first, and then multiplied by a preset ratio to obtain the bending intensity threshold. The smooth bending intensity sequence is then traversed to find points that exceed the threshold and their time indices are recorded, ultimately generating a list of peak point indices.
[0101] Step S333: On the state segment trajectory line, divide the trajectory line into multiple micro-trajectory segments according to the time index in the peak point index list. The starting point of each micro-trajectory segment is the previous peak point, and the ending point is the current peak point.
[0102] On the trajectory line of the state segment, the trajectory line is divided according to the time index in the peak point index list. The starting point of each micro-trajectory segment is the previous peak point, and the ending point is the current peak point. This division of micro-trajectory segments can highlight the parts of the trajectory line with greater curvature, which is convenient for the analysis of the local features of the trajectory line.
[0103] For example, the dividing point on the state segment trajectory line can be determined based on the time index in the peak point index list, thus dividing the trajectory line into multiple micro-trajectory segments.
[0104] Step S334: Calculate the sum of the distances between all adjacent points on the trajectory line of each micro-trajectory segment as the trajectory length, with the unit consistent with the phase space dimension unit.
[0105] The trajectory length is calculated by summing the distances between all adjacent points on the trajectory line of each micro-trajectory segment. This involves accumulating the distances between adjacent points on the trajectory line of each micro-trajectory segment. The unit is consistent with the unit of phase space dimension, ensuring consistency between the measurement of trajectory length and the dimension of phase space, facilitating subsequent analysis and comparison. For example, a distance calculation function can be used to calculate the distances between adjacent points on the trajectory line of a micro-trajectory segment, and then these distances can be summed to obtain the trajectory length.
[0106] Step S335: Calculate the curvature values of all points on the micro-trajectory segment, and take the arithmetic mean as the mean curvature value of the micro-trajectory segment, which reflects the overall curvature of the trajectory segment.
[0107] The curvature values of all points on the micro-track segment are calculated, where each curvature value represents the degree of bending of the trajectory line at a given point. The arithmetic mean of these values is taken as the mean curvature of the micro-track segment, thus providing a measure of the overall curvature of the segment. For example, the curvature value of each point on the micro-track segment can be calculated using a curvature calculation method, and then all curvature values are summed and divided by the number of points to obtain the mean curvature.
[0108] Step S336: Calculate the straight-line distance between the start and end points of each micro-trajectory segment, generate the endpoint distance parameter, and combine the trajectory length and mean curvature to generate a set of geometric features of the micro-trajectory segment containing geometric features.
[0109] The straight-line distance between the start and end points of each micro-trajectory segment is calculated; this distance, called the endpoint distance parameter, reflects the spatial distance between the start and end points of the micro-trajectory segment. Combining the trajectory length and mean curvature, these three parameters are used to generate a geometric feature set for the micro-trajectory segments, which comprehensively describes the geometric characteristics of each micro-trajectory segment, providing rich information for subsequent analysis of nonlinear changes in node communication behavior. For example, the straight-line distance between the start and end points of a micro-trajectory segment can be calculated using a distance calculation function to obtain the endpoint distance parameter. Then, the trajectory length, mean curvature, and endpoint distance parameter are combined to generate the geometric feature set of the micro-trajectory segments.
[0110] Step S340: Perform fractal interpolation on each micro-trajectory segment and calculate the long-term correlation coefficient of the interpolated trajectory. The long-term correlation coefficient reflects the self-similarity characteristics of the micro-trajectory segment in the time dimension.
[0111] Fractal interpolation is a method for interpolating and fitting data. Applying fractal interpolation to each micro-trajectory segment can make the trajectory smoother and more continuous, while also uncovering potential fractal structures within the trajectory. The long-term correlation coefficient of the interpolated trajectory is then calculated. The long-term correlation coefficient is an indicator that measures the correlation of time series data over a long time scale. Calculating this coefficient reflects the self-similarity of the micro-trajectory segments over time. A high long-term correlation coefficient indicates that the micro-trajectory segments have similar structures and trends across different time scales; conversely, a low long-term correlation coefficient indicates that the changes in the micro-trajectory segments are relatively random and lack self-similarity. For example, a fractal interpolation algorithm can be used to interpolate each micro-trajectory segment, and then correlation analysis methods can be used to calculate the long-term correlation coefficient of the interpolated trajectory.
[0112] Step S350: Calculate the multi-scale Lyapunov exponent spectrum of the trajectory line of each state segment. This is done by dividing the phase space into grids of different scales, calculating the local Lyapunov exponent in each grid, and taking the median of the exponents of each grid to form the exponent spectrum. The exponent spectrum reflects the trajectory separation characteristics at different spatial scales.
[0113] In one implementation, step S350 may specifically include the following steps S351 to S356:
[0114] Step S351: Determine the boundary range of the phase space where the state segment trajectory line is located, calculate the maximum and minimum values of the trajectory points in each dimension, and generate the phase space bounding box. The length of each dimension of the bounding box is a preset multiple of the trajectory point range in the corresponding dimension.
[0115] Determining the boundary range of the phase space containing the state segment trajectory line requires finding the maximum and minimum values of the trajectory points in each dimension. The maximum and minimum values of the trajectory points in each dimension are calculated, and the value range for each dimension is obtained by comparing the values of the trajectory points in each dimension. A phase space bounding box is generated; the bounding box is a multi-dimensional rectangle, and the length of each dimension is a preset multiple of the corresponding trajectory point range. The preset multiple is a pre-determined proportional value, which appropriately expands the boundary range to ensure that the trajectory line is completely contained within the bounding box. For example, maximum and minimum value calculation functions can be used to calculate the maximum and minimum values of the state segment trajectory line in each dimension. Then, the difference between the maximum and minimum values is multiplied by the preset multiple to obtain the length of each dimension of the bounding box, ultimately generating the phase space bounding box.
[0116] Step S352: Generate a set of grid scale parameters that decrease proportionally. The initial scale is a preset ratio of the minimum dimension length of the bounding box, and subsequent scales are set in a proportionally decreasing manner to generate multiple scale parameters covering the spatial range from global to local.
[0117] Generating a set of proportionally decreasing mesh scale parameters is used to divide the phase space into meshes of different scales to cover different spatial ranges from global to local. The initial scale is a preset ratio of the minimum dimension length of the bounding box. This preset ratio is a manually determined value, ensuring that the mesh covers the entire phase space. Subsequent scales are set in a proportionally decreasing manner, meaning each subsequent scale is a fixed proportion of the previous scale. This allows for the generation of multiple scale parameters of different sizes for dividing the mesh into different scales. For example, the minimum dimension length of the bounding box can be calculated first, and then multiplied by the preset ratio to obtain the initial scale. Then, following the proportionally decreasing rule, subsequent scale parameters are calculated sequentially, ultimately generating a set of proportionally decreasing mesh scale parameters.
[0118] Step S353: For each scale parameter, divide the phase space bounding box into cubic grids of equal size, count the number of trajectory points contained in each grid, filter out sparse grids with too few trajectory points, and retain the effective grids.
[0119] For each scale parameter, the phase space bounding box is divided into cubic meshes of equal size. This involves determining the mesh size based on each scale parameter and dividing the phase space bounding box into multiple cubic meshes of this size. The number of trajectory points within each mesh is counted to understand the distribution of trajectory points within each mesh. Sparse meshes with too few trajectory points are filtered out because these meshes contain limited information and may not accurately reflect the characteristics of the trajectory. Valid meshes, which contain a larger number of trajectory points and provide more valuable information, are retained. For example, a mesh division function can be used to divide the phase space bounding box into cubic meshes according to each scale parameter. Then, a counting function is used to count the number of trajectory points within each mesh. Based on a preset trajectory point count threshold, meshes with too few points are filtered out, ultimately retaining the valid meshes.
[0120] Step S354: For each valid grid, select the trajectory point within the grid as the initial reference point, search for its neighboring trajectory points in subsequent time steps, calculate the rate of change of the distance between the reference point and the neighboring points over time, and obtain the local Lyapunov exponent through linear fitting.
[0121] For each valid grid, a trajectory point within the grid is selected as an initial reference point, representing the initial state of the trajectory within that grid. Neighboring trajectory points are searched for in subsequent time steps. Neighboring trajectory points are those that are close to the reference point in subsequent time steps; these neighboring points reflect the temporal evolution of the reference point. The rate of change of the distance between the reference point and its neighbors over time is calculated; this rate of change reflects the separation speed of the trajectories in that region. A local Lyapunov exponent is obtained through linear fitting. Linear fitting is a method used to find linear relationships between data. By linearly fitting the rate of change of distance over time, the local Lyapunov exponent can be obtained, which represents the separation characteristics of the trajectories within that grid.
[0122] For example, a trajectory point can be randomly selected within each effective grid as an initial reference point, and then neighboring trajectory points can be searched in subsequent time steps. The distance between the reference point and neighboring points is calculated, and the change in distance over time is recorded. A linear fitting algorithm is used to fit these changing data to obtain the local Lyapunov exponent.
[0123] Step S355: Calculate the median of the local Lyapunov exponents for all effective grids at each scale parameter to obtain the representative exponent value corresponding to that scale. Multiple scale parameters correspond to multiple representative exponent values.
[0124] The median of the local Lyapunov exponents for all valid grids at each scale parameter is calculated. The median is a statistic that avoids the influence of individual outliers and better represents the middle level of a set of data. The median yields a representative exponent value for that scale, reflecting the overall separation characteristics of the trajectories at that scale. Multiple scale parameters correspond to multiple representative exponent values, which constitute part of the multi-scale Lyapunov exponent spectrum. For example, the median calculation function can be used to calculate the local Lyapunov exponents for all valid grids at each scale parameter, obtaining the representative exponent value for that scale.
[0125] Step S356: Arrange multiple representative index values in descending order of scale parameter to generate a multi-scale Lyapunov index spectrum, where the index values at different positions reflect the separation characteristics of node communication trajectories at the corresponding spatial scale.
[0126] Arranging multiple representative index values in descending order of scale parameter allows the index spectrum to be displayed in a spatial scale order from global to local. A multi-scale Lyapunov index spectrum is generated, where index values at different locations reflect the separation characteristics of node communication trajectories at corresponding spatial scales. Analyzing the multi-scale Lyapunov index spectrum reveals the stability and chaotic characteristics of node communication behavior at different spatial scales.
[0127] For example, a sorting function can be used to arrange multiple representative index values in descending order of scale parameter, ultimately generating a multi-scale Lyapunov index spectrum.
[0128] Step S360: Integrate the long-term correlation coefficients of each state segment, the geometric feature set of the micro-trajectory segment, and the multi-scale Lyapunov exponent spectrum, and generate a set of chaotic feature values of node communication behavior according to the protocol state code classification. This set contains nonlinear change indicators under different protocol states.
[0129] Integrating the long-term correlation coefficients, micro-trajectory segment geometric feature sets, and multi-scale Lyapunov exponent spectra of each state segment involves summarizing and merging the feature information obtained in the previous steps. The long-term correlation coefficients reflect the self-similarity of micro-trajectory segments in the time dimension; the micro-trajectory segment geometric feature sets describe the geometric characteristics of micro-trajectory segments; and the multi-scale Lyapunov exponent spectra reflect the separation characteristics of node communication trajectories at different spatial scales. Integrating this information allows for a more comprehensive description of the nonlinear changes in node communication behavior. Generating a chaotic feature value set of node communication behavior by classifying it according to protocol status codes involves classifying the integrated feature information based on the protocol status codes, ensuring that each protocol state has a corresponding chaotic feature value set. This set contains nonlinear change indicators under different protocol states, providing a basis for subsequent analysis of abnormal node communication patterns.
[0130] For example, the long-term correlation coefficients of each state segment, the geometric feature set of the micro-trajectory segment, and the multi-scale Lyapunov index spectrum can be merged, and then classified according to the state code labels of the state segment matrix to finally generate a set of chaotic feature values of node communication behavior.
[0131] Step S400: Construct a topological association structure of node communication features through local linear relationships, map the chaotic feature value set to a topological space of preset dimensions, and generate an attack pattern topological expression vector containing the geometric distribution features of abnormal node communication patterns.
[0132] A topological association structure of node communication features is constructed through local linear relationships. Local linear relationships refer to the linear relationships between node communication features within a local scope. This relationship can be used to construct the connection structure between nodes, reflecting the association between node communication features. A chaotic feature value set is mapped to a topological space of a preset dimension. The preset dimension is a pre-determined dimension value. The feature information in the chaotic feature value set is represented in the topological space through mapping operations. An attack pattern topological representation vector containing the geometric distribution characteristics of abnormal node communication patterns is generated. This vector can describe the geometric distribution characteristics of abnormal node communication patterns in the topological space, such as the distribution range and shape of the abnormal patterns, providing intuitive information for identifying and analyzing network attacks.
[0133] In one implementation, step S400 may specifically include the following steps S410 to S460:
[0134] Step S410: Construct a node communication feature matrix based on the chaotic feature value set. The row dimension of the matrix corresponds to the node identifier in the distributed network, the column dimension corresponds to different chaotic feature value types, and the matrix element values reflect the chaotic feature attributes of the corresponding node.
[0135] Constructing a node communication feature matrix based on a chaotic eigenvalue set involves organizing the feature information from the chaotic eigenvalue set into a matrix form. The row dimensions of the matrix correspond to node identifiers in the distributed network, clearly defining the position of each node within the matrix. The column dimensions correspond to different chaotic eigenvalue types, such as long-term correlation coefficients and micro-trajectory segment geometric features; different columns distinguish different types of features. The matrix element values reflect the chaotic feature attributes of the corresponding nodes, i.e., the specific feature values of each node under different chaotic feature types. For example, a matrix construction function can be used to arrange the feature information in the chaotic eigenvalue set according to node identifiers and chaotic eigenvalue types to generate a node communication feature matrix.
[0136] Step S420: Calculate the similarity of chaotic feature values between each node and other nodes, and generate an adjacency matrix containing the similarity measure between nodes. The element values of the adjacency matrix are positively correlated with the feature similarity between nodes.
[0137] Calculating the similarity of chaotic feature values between each node and other nodes measures the degree of similarity between nodes in terms of chaotic features. Methods such as Euclidean distance and cosine similarity can be used to calculate the similarity between nodes. An adjacency matrix containing the similarity measure between nodes is generated. This adjacency matrix is a square matrix where the element values are positively correlated with the feature similarity between nodes; that is, the higher the similarity between nodes, the larger the value of the corresponding element in the adjacency matrix; conversely, the lower the similarity, the smaller the element value.
[0138] In one implementation, step S420 may specifically include the following steps S421 to S425:
[0139] Step S421: Analyze the feature types in the chaotic feature value set and determine the feature dimensions used to calculate node similarity. The feature dimensions cover various chaotic feature values that reflect the nonlinear characteristics of node communication.
[0140] Analyzing the feature types within a chaotic eigenvalue set involves classifying and identifying different features. Determining the feature dimensions used to calculate node similarity requires covering various chaotic eigenvalues that reflect the nonlinear characteristics of node communication, such as long-term correlation coefficients and the geometric features of micro-trajectory segments. Choosing appropriate feature dimensions allows for more accurate calculation of node similarity.
[0141] For example, different feature types can be identified by analyzing the set of chaotic feature values, and the feature dimensions used to calculate node similarity can be determined based on the analysis results.
[0142] Step S422: Standardize the chaotic feature set, map different types of chaotic feature values to the same numerical range, eliminate the dimensional differences between feature values, and generate a standardized chaotic feature matrix. The row vectors of the standardized chaotic feature matrix correspond to the chaotic feature expressions of different nodes.
[0143] Standardizing the chaotic feature set is necessary because different types of chaotic features may have different dimensions and numerical ranges, which can affect the accuracy of similarity calculations. Mapping different types of chaotic features to the same numerical interval, for example, mapping all features to the [0, 1] interval, eliminates the dimensional differences between features. A standardized chaotic feature matrix is generated, where the row vectors correspond to the chaotic feature representations of different nodes. After standardization, the chaotic features of each node are represented within the same numerical interval, facilitating similarity calculations. For example, standardization algorithms such as min-max standardization and Z-score standardization can be used to process the chaotic feature set, mapping different types of chaotic features to the same numerical interval, ultimately generating a standardized chaotic feature matrix.
[0144] Step S423: Set the nearest neighbor quantity parameter. The value of the nearest neighbor quantity parameter is determined based on the number of nodes in the distributed network to ensure that each node has a sufficient number of neighbor nodes to build local topology relationships.
[0145] A nearest neighbor count parameter is set to determine the number of nearest neighbors for each node. The value of this parameter is determined based on the number of nodes in the distributed network. Generally, the more nodes there are, the larger the nearest neighbor count parameter can be; conversely, the fewer nodes there are, the smaller the parameter should be. Ensuring that each node has a sufficient number of neighbors is crucial for building local topology relationships. A sufficient number of neighbors more accurately reflects the local associations between nodes, providing a foundation for constructing a topological association structure that reflects node communication characteristics.
[0146] For example, the value of the nearest neighbor parameter can be determined based on the number of nodes in the distributed network, combined with experience and analytical requirements.
[0147] Step S424: For each node, calculate the distance between its normalized chaotic feature row vector and the normalized chaotic feature row vectors of all other nodes, and generate a distance matrix. The element values of the distance matrix reflect the degree of difference in chaotic features between nodes.
[0148] For each node, calculate the distance between its normalized chaotic feature row vector and the normalized chaotic feature row vectors of all other nodes. This distance can be calculated using methods such as Euclidean distance or Manhattan distance. Generate a distance matrix, which is a square matrix whose element values reflect the degree of difference in chaotic features between nodes. The larger the distance value, the greater the difference in chaotic features between nodes; the smaller the distance value, the more similar the chaotic features between nodes.
[0149] For example, a distance calculation function can be used to calculate the distance between the row vector of each node in the normalized chaotic feature matrix and the row vectors of all other nodes, and the distance values can be filled into the corresponding positions in the distance matrix to finally generate the distance matrix.
[0150] Step S425: Sort the row vectors of the distance matrix corresponding to each node, select the first few nodes with the smallest distance values as the nearest neighbors of that node, set the similarity value of the nearest neighbors to the reciprocal of the distance value, and set the similarity value of the non-nearest neighbors to zero, and generate an adjacency matrix containing the similarity measure between nodes.
[0151] The distance matrix row vectors for each node are sorted, meaning the distance values between each node and other nodes are arranged in ascending order. The nodes with the smallest distance values are selected as the node's nearest neighbors; the number of these first few nodes is determined by a previously set parameter for the number of nearest neighbors. The similarity value of the nearest neighbors is set to the reciprocal of the distance value, because a smaller distance value indicates greater similarity between nodes, and taking the reciprocal ensures a positive correlation between similarity and distance. The similarity value of non-nearest neighbors is set to zero, which highlights the connections between nearest neighbors and simplifies the structure of the adjacency matrix. An adjacency matrix containing similarity metrics between nodes is generated, reflecting the local relationships between nodes.
[0152] For example, a sorting function can be used to sort each row vector of the distance matrix, and the nodes with the smallest distance values can be selected as the nearest neighbors. Then, the reciprocal of the distance values of the nearest neighbors is taken as the similarity value, and the similarity values of the non-nearest neighbors are set to zero, thus generating the adjacency matrix.
[0153] Step S430: Construct a topological association structure of node communication features based on the adjacency matrix, represent nodes as vertices of the graph, represent the similarity measure between nodes as edge weights of the graph, and generate a weighted undirected graph containing node connection relationships.
[0154] A topological association structure based on node communication characteristics is constructed using an adjacency matrix. The adjacency matrix contains similarity metrics between nodes, which can be used to build connections between them. Nodes are represented as vertices in the graph, with each node represented by a single point. The similarity metric between nodes is represented by edge weights; higher edge weights indicate greater similarity and stronger connections between nodes. A weighted undirected graph containing node connections is generated. This graph visually displays the topological association structure between nodes, providing a foundation for subsequent community partitioning and analysis. For example, a graph construction function can be used, with nodes as vertices and the elements of the adjacency matrix as edge weights, to generate a weighted undirected graph.
[0155] Step S440: Perform community partitioning on the weighted undirected graph. Determine the optimal number of communities based on the eigenvalue decomposition results of the graph's Laplacian matrix. Divide the topological association structure into node communities with similar communication characteristics. Each node community corresponds to a node communication mode.
[0156] In one implementation, step S440 may specifically include the following steps S441 to S446:
[0157] Step S441: Construct the degree matrix of the weighted undirected graph. The degree matrix is a diagonal matrix, and the element values on the diagonal are equal to the degree values of the corresponding nodes. The degree value of a node is the sum of the edge weights connecting that node to other nodes.
[0158] Construct the degree matrix of the weighted undirected graph. The degree matrix is a diagonal matrix, and the values of its diagonal elements have corresponding meanings. The degree value of a node is the sum of the edge weights connecting that node to other nodes. The degree value reflects how closely the node is connected to other nodes in the graph. Fill the diagonal of the degree matrix with the degree value of each node to form the degree matrix.
[0159] For example, the degree value of each node in a weighted undirected graph can be calculated using a degree value calculation function, and then the degree value can be filled into the corresponding position in a diagonal matrix to generate a degree matrix.
[0160] Step S442: Calculate the Laplacian matrix of the graph based on the degree matrix and adjacency matrix, and generate a symmetric normalized Laplacian matrix through normalization. The symmetric normalized Laplacian matrix can eliminate the influence of node degree differences on the clustering results.
[0161] The Laplacian matrix of a graph is calculated based on the degree matrix and adjacency matrix. The Laplacian matrix is a matrix used to describe the structure of a graph and the relationships between nodes, and it can be obtained through operations on the degree matrix and adjacency matrix. A symmetric normalized Laplacian matrix is generated through normalization. Normalization is performed to eliminate the influence of node degree differences on the clustering results. The symmetric normalized Laplacian matrix has better mathematical properties and can more accurately reflect the community structure of the graph.
[0162] For example, the Laplacian matrix of a graph can be calculated using the Laplacian matrix calculation formula, based on the degree matrix and adjacency matrix. Then, a normalization algorithm is used to process the Laplacian matrix to generate a symmetric normalized Laplacian matrix.
[0163] Step S443: Perform eigenvalue decomposition on the symmetric normalized Laplacian matrix to extract the eigenvalues and corresponding eigenvectors of the matrix. Sort the eigenvalues in ascending order to generate an eigenvalue sequence and an eigenvector matrix.
[0164] Eigenvalue decomposition is performed on the symmetric normalized Laplacian matrix. Eigenvalue decomposition is a method that decomposes a matrix into eigenvalues and eigenvectors. The eigenvalues and corresponding eigenvectors are extracted; eigenvalues reflect some inherent properties of the matrix, while eigenvectors correspond to the eigenvalues. The eigenvalues are sorted in ascending order to generate an eigenvalue sequence and an eigenvector matrix. This facilitates the analysis of the eigenvalue distribution and provides a basis for determining the optimal number of communities.
[0165] For example, an eigenvalue decomposition algorithm can be used to decompose a symmetric normalized Laplacian matrix to extract eigenvalues and eigenvectors. Then, a sorting function is used to sort the eigenvalues, generating an eigenvalue sequence and an eigenvector matrix.
[0166] Step S444: Analyze the distribution characteristics of the eigenvalue sequence, determine the optimal number of communities using the eigenvalue gap method, select the position in the eigenvalue sequence where a significant jump occurs as the basis for judging the number of communities, and the eigenvalue gap reflects the degree of separation of different community structures.
[0167] Analyze the distribution characteristics of the eigenvalue sequence, observing the magnitude changes and distribution patterns of the eigenvalues. Determine the optimal number of communities using the eigenvalue gap method, a method that determines the number of communities based on the jumps in eigenvalues within the sequence. Positions showing significant jumps in the eigenvalue sequence are selected as the basis for judging the number of communities. The eigenvalue gap reflects the degree of separation between different community structures; the larger the gap, the more obvious the differences between different communities, and the more reasonable the community division.
[0168] For example, the feature value sequence can be visualized and analyzed to observe the trend of feature value changes, find the positions where significant jumps occur, and take the number of communities corresponding to the positions as the optimal number of communities.
[0169] Step S445: Extract the feature vectors corresponding to the smallest non-zero feature values of the number of the top optimal communities from the feature vector matrix, construct the feature vector subspace, and perform clustering processing on the sample points in the feature vector subspace using the mean clustering algorithm to generate the node community partitioning results.
[0170] The eigenvectors corresponding to the smallest non-zero eigenvalues of the top-optimal communities are extracted from the eigenvector matrix. These eigenvectors contain information about the community structure of the graph. A eigenvector subspace is constructed by combining these eigenvectors, where sample points correspond to nodes in the graph. The sample points in the eigenvector subspace are clustered using the mean clustering algorithm, a commonly used clustering method that can divide sample points into different categories. The node community partitioning results are generated, which divide the nodes in the graph into different communities, where nodes within each community share similar characteristics.
[0171] For example, a matrix extraction function can be used to extract the eigenvectors corresponding to the smallest non-zero eigenvalues of the number of the top optimal communities from the eigenvector matrix, constructing an eigenvector subspace. Then, the mean clustering algorithm is used to cluster the sample points in the eigenvector subspace to generate the node community partitioning results.
[0172] Step S446: Analyze the chaotic feature value distribution characteristics of each node community, calculate the mean and variance parameters of the chaotic feature values of nodes within the community, and generate feature descriptors that reflect the community communication pattern. Each node community corresponds to a node communication pattern with similar chaotic features.
[0173] Analyze the distribution characteristics of chaotic eigenvalues in each node community to understand the distribution of chaotic features among nodes within the community. Calculate the mean and variance parameters of the chaotic eigenvalues of nodes within the community. The mean reflects the average level of chaotic features among nodes in the community, while the variance reflects the dispersion of chaotic features. Generate feature descriptors that reflect the communication patterns of the communities. These descriptors contain information such as the mean and variance of chaotic features among nodes in the community. Each node community corresponds to a node communication pattern with similar chaotic features. The feature descriptors can be used to distinguish and identify different communication patterns.
[0174] For example, statistical analysis functions can be used to calculate the mean and variance parameters of chaotic features of nodes within each node community, and these parameters can be combined into feature descriptors to ultimately generate a set of feature descriptors that reflect the community communication pattern.
[0175] Step S450: Construct a topological space of a preset dimension. Map the topological structure of the weighted undirected graph to the topological space using a multidimensional scaling analysis method, preserve the relative distance relationship between nodes, and generate the coordinate position matrix of the nodes in the topological space.
[0176] A topological space with a predefined dimension is constructed. This predefined dimension is a value determined manually based on the analytical needs and data characteristics. Multidimensional scaling analysis is used to map the topological structure of the weighted undirected graph to this topological space. Multidimensional scaling analysis is a method used to map high-dimensional data to a low-dimensional space, preserving the relative distance relationships between data points. A matrix of node coordinate positions in the topological space is generated. This matrix records the coordinate position of each node in the topological space, visually displaying the relative positional relationships between nodes and facilitating the analysis of the geometric distribution of abnormal node communication patterns.
[0177] For example, a topology space of a preset dimension can be constructed using a topology space construction function. Then, a multidimensional scaling analysis algorithm is used to map the topology of the weighted undirected graph onto the topology space, obtaining the coordinate positions of the nodes in the topology space. These coordinate positions are then filled into a matrix to generate a coordinate position matrix.
[0178] Step S460: Extract the geometric distribution features of each node community in the coordinate position matrix, calculate the center coordinates, distribution radius and shape parameters of the community, and generate an attack mode topology expression vector to describe the spatial distribution characteristics of abnormal node communication patterns.
[0179] The geometric distribution characteristics of each node community are extracted from the coordinate position matrix. This matrix records the coordinate position of each node in the topological space, allowing analysis of the geometric distribution of node communities. The center coordinates of each community are calculated, representing its central location and reflecting its overall positional characteristics. The distribution radius measures the size of the community's distribution range, and shape parameters describe its shape, such as circular or elliptical. An attack pattern topology representation vector is generated to describe the spatial distribution characteristics of abnormal node communication patterns. This vector contains information such as the community's center coordinates, distribution radius, and shape parameters, comprehensively describing the distribution characteristics of abnormal node communication patterns in the topological space, providing crucial information for identifying and analyzing network attacks.
[0180] For example, geometric calculation methods can be used to calculate the center coordinates, distribution radius, and shape parameters of each node community in the coordinate position matrix. These parameters are then combined into a vector to generate an attack pattern topology representation vector.
[0181] Step S500: Based on the attack pattern topology expression vector, analyze the abnormal state propagation process in the distributed network through the energy propagation rules between nodes to obtain network security detection results that include the attack diffusion path and the energy attenuation distribution of each node.
[0182] Based on attack pattern topology representation vectors, which contain information such as the spatial distribution characteristics of abnormal node communication patterns, potential attack patterns within the network can be identified. The propagation process of abnormal states in a distributed network is analyzed using inter-node energy propagation rules. These rules define the methods and rules for the propagation of abnormal states between nodes, such as the direction and coefficient of energy propagation. Analyzing the propagation process of abnormal states can simulate the spread of attacks within the network. Network security detection results, including attack propagation paths and energy attenuation distributions at each node, are obtained. The attack propagation path clearly defines the route of attack propagation from the originating node to other nodes, while the energy attenuation distribution reflects the attenuation of attack energy at each node during propagation. This information is crucial for the timely detection and prevention of network attacks.
[0183] In one implementation, step S500 may specifically include the following steps S510-S560:
[0184] Step S510: Convert the attack mode topology representation vector into an initial abnormal energy distribution vector. The dimension of the initial abnormal energy distribution vector is consistent with the number of nodes in the distributed network, and the vector element values reflect the initial energy of the abnormal state of the corresponding node.
[0185] The attack pattern topology representation vector is transformed into an initial anomaly energy distribution vector. This vector contains information such as the geometric distribution characteristics of the node communication anomaly patterns. A specific transformation method converts this information into the initial anomaly energy for each node. The dimension of the initial anomaly energy distribution vector is consistent with the number of nodes in the distributed network, ensuring that each node has a corresponding initial energy value. The vector element values reflect the initial anomaly energy of the corresponding node; a higher energy value indicates a more severe anomaly. For example, a vector transformation function can be used to assign an initial anomaly energy value to each node based on the information in the attack pattern topology representation vector, generating the initial anomaly energy distribution vector.
[0186] Step S520: Construct a set of energy propagation rules between nodes, determine the energy propagation path based on the topological association structure of node communication characteristics, the propagation coefficient on the energy propagation path is negatively correlated with the topological distance between nodes, and the propagation coefficient of non-directly connected nodes is set to zero.
[0187] A set of energy propagation rules is constructed between nodes, containing various rules for the propagation of anomalous energy between nodes, such as propagation direction and propagation coefficient. Energy propagation paths are determined based on the topological association structure of node communication characteristics. This topological association structure reflects the connection relationships between nodes, and the path for anomalous energy to propagate from one node to another can be determined based on this structure. The propagation coefficient on the energy propagation path is negatively correlated with the topological distance between nodes; the greater the topological distance, the smaller the propagation coefficient, indicating greater energy loss during propagation. The propagation coefficient for non-directly connected nodes is set to zero, meaning that anomalous energy will not propagate between non-directly connected nodes. For example, the connection relationships and topological distances between nodes can be determined based on the topological association structure of node communication characteristics. Then, a propagation coefficient is set for each energy propagation path based on the topological distance. The setting of the propagation coefficient can be achieved through a mapping function that maps the topological distance to the corresponding propagation coefficient, satisfying the negative correlation between the propagation coefficient and the topological distance. For non-directly connected nodes, their propagation coefficients are set to zero. Finally, these propagation rules are compiled into a set for subsequent anomalous state propagation simulation.
[0188] Step S530: Based on the initial abnormal energy distribution vector and the set of energy propagation rules, the dynamic evolution simulation of abnormal state propagation is carried out through time step iterative update. In each iteration, the energy transfer amount is calculated according to the current energy value of the node and the propagation coefficient of the adjacent nodes, and the dynamic evolution sequence of node abnormal energy at different times is generated.
[0189] A dynamic evolution simulation of anomalous state propagation is performed based on an initial anomalous energy distribution vector and a set of energy propagation rules. The initial anomalous energy distribution vector provides the initial anomalous energy value for each node, while the set of energy propagation rules defines the rules for anomalous energy propagation between nodes. The simulation is conducted through an iterative update method with a time step, where a manually set time interval is used. Within each time step, the energy transfer amount is calculated based on the node's current energy value and the propagation coefficients of neighboring nodes. The energy transfer amount represents the magnitude of anomalous energy transferred from a node to its neighboring nodes within a time step. After each iteration, the anomalous energy value of each node is updated. This iterative process is repeated continuously to simulate the dynamic propagation of anomalous states in a distributed network. A dynamic evolution sequence of node anomalous energy at different times is generated. This sequence records the anomalous energy value of each node at different times, and analysis of this sequence reveals the propagation process and trend of anomalous states in the network.
[0190] In one implementation, step S530 may specifically include the following steps S531 to S536:
[0191] Step S531: Set the time step parameter based on the time resolution of node communication events, determine the total number of iterations required to complete one complete anomaly propagation simulation, and initialize the node anomaly energy distribution at the current moment as the initial anomaly energy distribution vector.
[0192] The time step parameter is set based on the temporal resolution of node communication events. The temporal resolution of node communication events reflects the temporal precision of node communication data recording. Setting the time step appropriately based on this resolution ensures that the time step accurately reflects the propagation process of the abnormal state without being overly precise and causing excessive computational load. The total number of iterations required to complete one full anomaly propagation simulation is determined. The determination of the total number of iterations needs to comprehensively consider factors such as the network size and the propagation speed of the abnormal state to ensure that the simulation can completely represent the propagation process of the abnormal state in the network. The node anomaly energy distribution at the current moment is initialized as the initial anomaly energy distribution vector. This initial anomaly energy distribution vector serves as the starting point of the simulation, providing a foundation for subsequent iterative updates.
[0193] For example, a suitable time step parameter can be determined based on the temporal resolution of node communication events, combined with experience and analytical requirements. Then, based on the network characteristics and the expected propagation of abnormal states, the total number of iterations required to complete one full simulation is estimated. Finally, the initial abnormal energy distribution vector is assigned to the node abnormal energy distribution at the current moment, completing the initialization operation.
[0194] Step S532: Perform an adjacency energy collection operation for each node. Based on the propagation path information in the energy propagation rule set, extract the abnormal energy values of the adjacent nodes at the previous iteration time. Combine the propagation coefficients of the corresponding propagation paths to calculate the weighted summation result and obtain the initial energy input value of the current node.
[0195] Performing an adjacency energy collection operation on each node involves collecting the anomalous energy transferred to it from its neighboring nodes. Based on the propagation path information in the energy propagation rule set, the neighboring nodes of each node and the propagation paths between them are determined. The anomalous energy values of the neighboring nodes at the previous iteration are extracted; these values reflect the anomalous state of the neighboring nodes at the previous time step. A weighted sum is calculated by combining the propagation coefficients of the corresponding propagation paths. The propagation coefficients represent the degree of energy attenuation during propagation. The weighted sum yields the total anomalous energy received by the current node from its neighboring nodes, i.e., the initial energy input value of the current node. For example, for each node, its neighboring nodes and corresponding propagation paths can be found in the energy propagation rule set. Then, the anomalous energy values of the neighboring nodes at the previous iteration are extracted from the node's anomalous energy dynamic evolution sequence. The anomalous energy value of each neighboring node is multiplied by its corresponding propagation coefficient, and these products are summed to obtain the initial energy input value of the current node.
[0196] Step S533: Calculate the energy transfer amount between nodes based on the preset energy transfer formula. The transfer amount is the result of the correlation between the initial energy input value and the propagation coefficient. At the same time, the upper limit constraint of the node processing capacity threshold on the transfer amount is considered to generate the preliminary energy transfer value.
[0197] The energy transfer amount between nodes is calculated based on a pre-defined energy transfer formula. This formula specifies how to calculate the energy transfer amount based on the initial energy input value and the propagation coefficient. This formula can be designed according to the characteristics of the network and the laws governing the propagation of abnormal states. The transfer amount is taken as the result of a calculation involving the initial energy input value and the propagation coefficient, such as their product or other reasonable calculation methods. Simultaneously, a node processing capacity threshold is considered as an upper limit constraint on the transfer amount. This threshold is a pre-set value representing the maximum abnormal energy transfer amount that a node can handle within a time step. If the calculated energy transfer amount exceeds the node processing capacity threshold, the transfer amount is set as the threshold, generating a preliminary energy transfer value.
[0198] For example, a preset energy transfer formula can be used to calculate an energy transfer amount by substituting the initial energy input value and the propagation coefficient into the formula. This transfer amount is then compared with a node processing capacity threshold. If it exceeds the threshold, the transfer amount is set to the threshold, and a preliminary energy transfer value is obtained.
[0199] Step S534: Perform energy attenuation correction processing on the preliminary energy transfer value. Calculate the attenuation factor based on the node's energy attenuation coefficient and propagation time step. Multiply the preliminary energy transfer value by the attenuation factor to obtain the corrected actual energy transfer amount.
[0200] Energy attenuation correction is applied to the initial energy transfer value because energy decays with increasing propagation distance and time during anomalous energy propagation. An attenuation factor is calculated based on the node's energy attenuation coefficient and the propagation time step. The energy attenuation coefficient reflects the degree of energy loss during a node's propagation of anomalous energy, while the propagation time step represents the time interval for energy propagation. These two parameters allow for the calculation of the energy attenuation within a given time step, yielding the attenuation factor. Multiplying the initial energy transfer value by the attenuation factor yields the corrected actual energy transfer amount. This actual energy transfer amount more accurately reflects the magnitude of anomalous energy transferred from one node to its neighboring nodes, considering energy attenuation. For example, the attenuation factor can be calculated using an attenuation calculation function using the energy attenuation coefficient and the propagation time step. Then, the initial energy transfer value is multiplied by the attenuation factor to obtain the corrected actual energy transfer amount.
[0201] Step S535: For edge nodes in the distributed network, a neighborhood energy compensation mechanism is used to handle boundary effects. By calculating the average energy transfer between the edge node and all its neighboring nodes, the actual energy transfer of the edge node is compensated and corrected to ensure the continuity of energy propagation in the boundary region.
[0202] For edge nodes in a distributed network, which are nodes located at the edge of the network topology with a relatively small number of neighboring nodes, discontinuous energy propagation may occur, known as the boundary effect. A neighborhood energy compensation mechanism is employed to address this boundary effect. The core idea of this mechanism is to compensate for and correct the actual energy transfer of the edge node by calculating the average energy transfer between the edge node and all its neighboring nodes. This involves averaging the energy transfer received by the edge node from its neighbors to obtain an average energy transfer value. This average energy transfer value is then used to adjust the actual energy transfer of the edge node, ensuring the continuity of energy propagation in the boundary region and allowing abnormal states to propagate more smoothly throughout the network. For example, for each edge node, the sum of its energy transfer with all its neighboring nodes can be calculated, and then divided by the number of neighboring nodes to obtain the average energy transfer value. Based on the relationship between the average energy transfer value and the actual energy transfer value of the edge node, the actual energy transfer value can be compensated and corrected; for example, the difference between the two can be added to the actual energy transfer value.
[0203] Step S536: Accumulate the actual energy transfer amount of all nodes at the current iteration time into the abnormal energy value of the corresponding node, generate the node abnormal energy distribution vector at the current time, repeat the steps of adjacent energy collection to boundary compensation until the preset total number of iterations is reached, and generate a node abnormal energy dynamic evolution sequence containing time series relationships.
[0204] The process involves accumulating the actual energy transfer amounts of all nodes at the current iteration time into the corresponding abnormal energy value. This means adding the actual abnormal energy transfer amount received by each node in the current iteration to its original abnormal energy value, thus updating the node's abnormal energy state. A node abnormal energy distribution vector is generated at the current time, recording the abnormal energy values of all nodes at that moment. The steps of adjacency energy collection to boundary compensation are repeated until the preset total number of iterations is reached. In each iteration, the abnormal energy values of the nodes are continuously updated, simulating the propagation process of abnormal states in the network. Finally, a dynamic evolution sequence of node abnormal energy containing time-series relationships is generated. This sequence records the abnormal energy value of each node at different times. Analyzing this sequence provides a clear understanding of the propagation process and trend of abnormal states in the distributed network.
[0205] For example, a vector addition function can be used to add the actual energy transfer amount of all nodes at the current iteration time to the abnormal energy value of the corresponding node, generating the node abnormal energy distribution vector at the current time. Then, it is checked whether the preset total number of iterations has been reached. If not, the next iteration is performed, and the steps of adjacent energy collection, energy transfer amount calculation, energy decay correction, boundary compensation, etc. are repeated until the total number of iterations is reached, finally generating the dynamic evolution sequence of node abnormal energy.
[0206] Step S540: Analyze the temporal distribution characteristics of the dynamic evolution sequence of abnormal node energy, extract the set of nodes whose energy values exceed the preset abnormal threshold and their corresponding timestamp information, and generate an abnormal node activation time series. The abnormal node activation time series reflects the spread process of the attack in the network.
[0207] Analyzing the temporal distribution characteristics of the dynamic evolution sequence of abnormal node energy aims to observe how the abnormal energy value of a node changes over time and understand the temporal pattern of abnormal energy propagation in the network. A preset anomaly threshold is a pre-defined value; when the abnormal energy value of a node exceeds this threshold, it indicates that the node is in an abnormal state and may have been attacked. The process involves extracting the set of nodes whose energy values exceed the preset anomaly threshold and their corresponding timestamps, identifying nodes whose abnormal energy values exceed the threshold at different times, and recording the timestamps of when these nodes were activated (i.e., their energy values exceeded the threshold). An abnormal node activation time series is generated, recording the chronological order in which each abnormal node is activated, reflecting the spread of the attack in the network. Analyzing this series reveals how the attack propagates from one node to other nodes. For example, the dynamic evolution sequence of abnormal node energy can be traversed, and for each node's abnormal energy value sequence, it can be checked whether it exceeds the preset anomaly threshold. If it exceeds the threshold, the node's identifier and corresponding timestamp are recorded. These records are then arranged in chronological order by timestamp to generate the abnormal node activation time series.
[0208] Step S550: Based on the activation time series of abnormal nodes and the energy propagation path, reconstruct the attack diffusion path through the temporal relationship of node activation and path preference information, determine the sequential relationship of the attack source node, intermediate propagation nodes and target nodes, and generate an attack diffusion path descriptor containing node identifiers and timestamps.
[0209] Attack propagation paths are reconstructed based on the activation time series of anomalous nodes and energy propagation paths. The activation time series records the chronological order in which anomalous nodes are activated, while the energy propagation path defines the route of anomalous energy propagation between nodes. By analyzing the temporal relationship of node activations and path preference information, the temporal relationship of node activations helps determine the direction of attack propagation, and the path preference information reflects which paths the anomalous energy prefers to propagate through. Using this information, the attack propagation path is reconstructed to identify the specific path from the source node, through intermediate propagation nodes, to the target node. The sequential relationship between the source node, intermediate propagation nodes, and target node is determined, clarifying the propagation order of the attack in the network. An attack propagation path descriptor containing node identifiers and timestamps is generated. This descriptor records the identifier of each node on the attack propagation path and the timestamp of the attack arriving at that node, providing a detailed understanding of the attack propagation process in the network.
[0210] For example, the earliest activated node can be identified from the abnormal node activation time series as the attack source node. Then, based on the energy propagation path and the chronological order of node activation, the sequence of intermediate propagation nodes and the target node is determined step by step. For each node, its identifier and activation timestamp are recorded, ultimately generating an attack propagation path descriptor.
[0211] Step S560: Calculate the energy attenuation parameters of each node during the attack propagation process, determine the energy attenuation coefficient based on the energy propagation distance and propagation time, generate a node energy attenuation distribution matrix, the matrix element values reflect the degree of attenuation of attack energy when passing through the node, and combine the attack propagation path descriptor to obtain the network security detection result containing the attack propagation path and the energy attenuation distribution of each node.
[0212] The energy attenuation parameters of each node during the attack propagation process are calculated. These parameters reflect the attenuation of attack energy as it passes through each node. The energy attenuation coefficient is determined based on the energy propagation distance and propagation time. The energy propagation distance refers to the number of nodes or path length the attack energy passes through from the source node to the target node, and the propagation time refers to the time it takes for the attack to propagate from the source node to the target node. The longer the energy propagation distance and propagation time, the larger the energy attenuation coefficient, indicating greater energy loss during propagation. A node energy attenuation distribution matrix is generated, where the rows and columns correspond to nodes in the network, and the matrix element values reflect the degree of attenuation of attack energy as it passes through the nodes. Combined with the attack propagation path descriptor, a network security detection result containing the attack propagation path and the energy attenuation distribution at each node is obtained. This result includes both the attack propagation path information in the network and the attenuation of attack energy at each node, providing network security administrators with comprehensive information to take appropriate preventative measures.
[0213] For example, for each node on the attack propagation path, its energy attenuation coefficient can be determined using an energy attenuation coefficient calculation function based on its energy propagation distance and propagation time from the attack source node. The energy attenuation coefficient of each node is then filled into the corresponding position in the node energy attenuation distribution matrix. Finally, the attack propagation path descriptor and the node energy attenuation distribution matrix are combined to obtain a network security detection result containing the attack propagation path and the energy attenuation distribution of each node.
[0214] The various algorithms involved in the above embodiments, such as Euclidean distance algorithm, cosine distance algorithm, clustering algorithm, time delay embedding method, etc., can all be learned from relevant content in the prior art, and will not be elaborated on here. In addition, when implementing the solution of the present invention, those skilled in the art can supplement the undisclosed details based on common knowledge in the art. For example, based on common knowledge in the art, normalization can be used to eliminate dimensional conflicts before feature fusion, interpolation can be used to eliminate dimensional differences, thresholds can be reasonably set in combination with historical data, experience or business scenario requirements, the model can be trained based on a general model training method, the number of layers in the model structure can be set according to actual needs, activation functions can be selected, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes.
[0215] Please see details. Figure 3 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. Figure 3 As shown, the computer system 1000 described above may include: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer system 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 3 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program. Figure 3 In the computer system 1000 shown, the network interface 1004 provides network communication functions; the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the methods provided in the above embodiments.
[0216] It should be understood that the computer system 1000 described in the embodiments of the present invention can execute the foregoing text. Figure 2 The implementation principle and beneficial effects of the network security monitoring method based on distributed nodes described in the corresponding embodiments will not be repeated here.
Claims
1. A network security monitoring method based on distributed nodes, characterized in that, The method includes: Obtain the node interaction records and identification information of each node in the distributed network within a preset communication period to obtain a node communication dataset containing time series markers. The node communication dataset includes the communication time sequence between nodes and the fixed field sequence in the protocol interaction process. Based on the node communication dataset, a multidimensional phase space reconstruction matrix containing communication timing association features and protocol identifier association features is constructed using the time delay embedding method. The row dimension of the multidimensional phase space reconstruction matrix corresponds to the occurrence time of node communication events, and the column dimension corresponds to the communication association attributes between different nodes. Nonlinear dynamics feature analysis is performed on the multidimensional phase space reconstruction matrix. The chaotic feature set of node communication behavior is extracted by calculating the evolution complexity index of the phase space trajectory. The chaotic feature set is used to characterize the nonlinear variation characteristics of the node communication mode. The topological association structure of node communication features is constructed by local linear relationships, and the chaotic feature value set is mapped to a topological space of a preset dimension to generate an attack mode topological expression vector containing the geometric distribution features of abnormal node communication patterns. Based on the attack pattern topology representation vector, the abnormal state propagation process in the distributed network is analyzed through the energy propagation rules between nodes, and network security detection results including attack propagation paths and energy attenuation distribution of each node are obtained.
2. The method according to claim 1, characterized in that, The construction of a multi-dimensional phase space reconstruction matrix containing communication timing correlation features and protocol identifier correlation features based on the node communication dataset, using a time delay embedding method, includes: The fixed field sequence of the protocol interaction process in the node communication data is parsed, and the protocol type identifier, status code transition sequence and field length variation coefficient are extracted to generate a protocol identifier raw tuple containing protocol behavior features. Each element of the protocol identifier raw tuple corresponds to the attribute description of a fixed field in the protocol interaction process. The communication time sequence in the node communication dataset is divided into hierarchical time windows. Continuous communication events are divided into protocol-specific time window units according to the protocol type identifier. The communication frequency density and data packet interval volatility within each protocol-specific time window unit are calculated to generate protocol-differentiated communication time sequence units. The time granularity of the protocol-differentiated communication time sequence units is matched with the average period of protocol interaction. The dynamic delay parameter of the time delay embedding method is determined based on the protocol status code transition sequence. When the protocol status code changes, the delay parameter is increased, and when the protocol status code is continuous and stable, the delay parameter is decreased, generating a delay parameter sequence that dynamically adjusts with the protocol status. The protocol-differentiated communication timing units are mapped to the phase space by the time delay embedding method, and a basic timing correlation matrix is constructed by combining the delay parameter sequence. The row vectors of the matrix correspond to the node communication times under different protocol states, and the column vectors correspond to the communication feature values under different delays. The original protocol identifier tuple is converted into a numerical protocol identifier vector. The weight coefficient of each field is calculated by the protocol field importance weighting algorithm. The protocol identifier vector is then weighted and expanded based on the weight coefficient to generate a weighted protocol identifier matrix that matches the column dimensions of the basic time series correlation matrix. The basic temporal correlation matrix and the weighted protocol identifier matrix are concatenated along the column dimension. The feature dimensions with cumulative contribution rates reaching a preset threshold are retained through incremental singular value decomposition. This generates a multi-dimensional phase space reconstruction matrix in which the row dimension corresponds to the occurrence time of node communication events and the column dimension corresponds to the communication correlation attributes between different nodes.
3. The method according to claim 2, characterized in that, The process of hierarchically dividing the communication time sequence in the node communication dataset into time windows, segmenting continuous communication events into protocol-specific time window units according to protocol type identifiers, calculating the communication frequency density and data packet interval volatility within each protocol-specific time window unit, and generating protocol-differentiated communication time sequence units includes: Extract all protocol type identifiers from the node communication dataset, count the number of interactions and the duration of each interaction for each protocol type within a preset communication period, calculate the average interaction period for each protocol, and generate a protocol period feature table. The protocol period feature table contains the mapping relationship between protocol types and average interaction periods. Based on the protocol cycle feature table, a dedicated time window length is set for each protocol type. The window length is a preset multiple of the average interaction cycle of the corresponding protocol, so that each window can completely cover at least one protocol interaction process. Traverse the communication time sequence in the node communication dataset, and assign continuous communication events to the dedicated time window unit of the corresponding protocol according to the protocol type identifier. For communication events that cross the window boundary, assign them to the adjacent window according to the time proportion. Calculate the total number of communication events within each protocol-specific time window unit, divide by the window length to obtain the communication frequency density, which reflects the activity level of protocol interactions per unit time. Extract the transmission interval sequence of all data packets within each window, and calculate the relationship between the dispersion of the transmission interval sequence and the average interval as the data packet interval volatility. The protocol-specific time window units for each protocol are arranged in chronological order, and each window unit is associated with the corresponding communication frequency density and data packet interval volatility to generate protocol-differentiated communication timing units, whose time granularity is consistent with the average interaction cycle of the corresponding protocol.
4. The method according to claim 2, characterized in that, The dynamic delay parameter determined by the protocol status code transition sequence-based time delay embedding method increases the delay parameter when a protocol status code transition occurs and decreases the delay parameter when the protocol status code is continuous and stable, generating a delay parameter sequence that dynamically adjusts with the protocol status, including: Parse the status code transition sequence in the original tuple of the protocol identifier, extract the timestamp of the status code transition event and the status code values before and after the transition, and construct a status code transition event list, where each entry contains the transition time, the original status code, and the new status code; The time interval of all status code transition events within a preset communication period is statistically analyzed, the average transition interval is calculated, and its preset proportion is set as the basic delay parameter, which serves as the initial value of the dynamic delay parameter. Traverse the status code transition sequence. When a status code transition event is detected, increase the current delay parameter by the ratio of the duration of the continuous stable status codes before the transition to the average transition interval, thereby enhancing the capture of the temporal correlation during state changes. When the duration of continuous detection of the same status code without transition exceeds a preset multiple of the average transition interval, the current delay parameter is reduced by a preset ratio of the duration to the average transition interval based on the base delay parameter. The adjusted delay parameter is subject to upper and lower limit constraints. The lower limit is set as a preset ratio of the base delay parameter, and the upper limit is set as a preset multiple of the base delay parameter. Record the dynamic delay parameters corresponding to each communication event in chronological order to generate a dynamic delay parameter sequence with the same length as the communication time sequence.
5. The method according to claim 2, characterized in that, The nonlinear dynamics feature analysis of the multidimensional phase space reconstruction matrix, and the extraction of a set of chaotic feature values of node communication behavior by calculating the evolution complexity index of the phase space trajectory, includes: The multidimensional phase space reconstruction matrix is preprocessed by trajectory segmentation. Based on the timestamp information of the matrix row dimension, the matrix is divided into multiple state segment matrices according to the timestamp of the protocol status code jump event. Each state segment matrix corresponds to a communication period with stable protocol status codes. For each state segment matrix, its row vectors are regarded as trajectory points in phase space, and they are connected in time order to form a state segment trajectory line. The tangent vector sequence of the trajectory line is calculated, and the trajectory bending strength is determined by the rate of change of the angle between adjacent tangent vectors, thus generating a trajectory bending strength sequence. Based on the peak points of the trajectory bending intensity sequence, the trajectory line of the state segment is segmented twice to obtain multiple micro-trajectory segments. The length, mean curvature and endpoint distance of each micro-trajectory segment are calculated to generate a geometric feature set of micro-trajectory segments. Fractal interpolation is performed on each micro-trajectory segment, and the long-term correlation coefficient of the interpolated trajectory is calculated. The long-term correlation coefficient reflects the self-similarity of the micro-trajectory segment in the time dimension. The multi-scale Lyapunov exponent spectrum of the trajectory line in each state segment is calculated by dividing the phase space into grids of different scales, calculating the local Lyapunov exponent in each grid, and taking the median of the exponents of each grid to form the exponent spectrum. The exponent spectrum reflects the trajectory separation characteristics at different spatial scales. By integrating the long-term correlation coefficients of each state segment, the geometric feature set of the micro-trajectory segment, and the multi-scale Lyapunov exponent spectrum, a set of chaotic feature values of node communication behavior is generated according to the protocol state code classification.
6. The method according to claim 5, characterized in that, The trajectory segmentation preprocessing of the multidimensional phase space reconstruction matrix involves dividing the matrix into multiple state segment matrices based on the timestamp information of the matrix row dimensions and the timestamps of protocol status code transition events, including: Extract the timestamp information of the row dimension of the multidimensional phase space reconstruction matrix to construct a timestamp sequence, wherein each element in the timestamp sequence corresponds to the time of occurrence of the communication event in a row of the matrix; Extract the status code flow sequence from the original tuple of the protocol identifier, and align the status code flow sequence with the timestamp sequence to ensure that each status code corresponds to a unique timestamp, thereby generating a status code-timestamp lookup table. Traverse the status code-timestamp lookup table, identify the start and end timestamps of consecutive identical status codes, calculate the difference between the start and end timestamps as the status code stability duration, and generate a status code stability period list. Each entry in the list contains the status code value, start timestamp, and end timestamp. Based on the start and end timestamps in the status code stable period list, a sub-matrix corresponding to the row range is extracted from the multi-dimensional phase space reconstruction matrix. Each sub-matrix is a status segment matrix, and its row dimension corresponds to the time when the communication event occurs within the stable period of the status code. Each state segment matrix is checked for row dimension. If the number of rows in the matrix is less than the minimum row number threshold set according to the protocol interaction characteristics, it is merged with the adjacent state segment matrix. The merged state segment matrix corresponds to two consecutive stable state code periods. Add a status code label to each status segment matrix. The label value is the status code value of the corresponding stable status code period. Generate a set of status segment matrices containing status code labels. Each matrix corresponds to a stable communication period of the protocol status code.
7. The method according to claim 5, characterized in that, The calculation of the multi-scale Lyapunov exponent spectrum for each state segment trajectory line involves dividing the phase space into grids of different scales, calculating the local Lyapunov exponent within each grid, and then taking the median of the exponents from each grid to form the exponent spectrum, including: Determine the boundary range of the phase space where the trajectory line of the state segment is located, calculate the maximum and minimum values of the trajectory points in each dimension, and generate a phase space bounding box. The length of each dimension of the bounding box is a preset multiple of the trajectory point range in the corresponding dimension. Generate a set of grid scale parameters that decrease proportionally. The initial scale is a preset ratio of the minimum dimension length of the bounding box, and subsequent scales are set in a proportionally decreasing manner to generate multiple scale parameters covering the spatial range from global to local. For each scale parameter, the phase space bounding box is divided into cubic grids of equal size. The number of trajectory points contained in each grid is counted. Sparse grids with too few trajectory points are filtered out, and effective grids are retained. For each valid grid, a trajectory point within the grid is selected as the initial reference point. The neighboring trajectory points in subsequent time steps are searched, and the rate of change of the distance between the reference point and the neighboring points over time is calculated. The local Lyapunov exponent is obtained through linear fitting. The median of the local Lyapunov exponents of all effective grids under each scale parameter is calculated to obtain the representative exponent value corresponding to that scale. Multiple scale parameters correspond to multiple representative exponent values. Multiple representative index values are arranged in descending order of scale parameter to generate a multi-scale Lyapunov index spectrum, where the index values at different positions reflect the separation characteristics of node communication trajectories at the corresponding spatial scale.
8. The method according to claim 5, characterized in that, The trajectory line of the state segment is further segmented based on the peak points of the trajectory curvature intensity sequence to obtain multiple micro-trajectory segments. The length, mean curvature, and endpoint distance of each micro-trajectory segment are calculated to generate a geometric feature set of micro-trajectory segments, including: The trajectory bending intensity sequence is smoothed to generate a smooth bending intensity sequence; Set the bending strength threshold to a preset proportion of the maximum value of the smooth bending strength sequence, identify all peak points in the sequence that exceed the threshold, record the time index of the peak points, and generate a peak point index list; On the state segment trajectory line, the trajectory line is divided into multiple micro-trajectory segments according to the time index in the peak point index list. The starting point of each micro-trajectory segment is the previous peak point, and the ending point is the current peak point. The length of the trajectory is calculated as the sum of the distances between all adjacent points on the trajectory line of each micro-trajectory segment, with the unit consistent with the unit of phase space dimension. Calculate the curvature values of all points on the micro-trajectory segment, and take the arithmetic mean as the mean curvature of the micro-trajectory segment, which reflects the overall curvature of the trajectory segment; Calculate the straight-line distance between the start and end points of each micro-trajectory segment, generate the endpoint distance parameter, and combine the trajectory length and mean curvature to generate a set of geometric features of the micro-trajectory segment containing geometric features.
9. The method according to claim 1, characterized in that, The process of constructing a topological association structure for node communication features through local linear relationships maps the chaotic feature value set to a topological space of a preset dimension, generating an attack pattern topological expression vector containing the geometric distribution features of abnormal node communication patterns, including: Based on the chaotic feature value set, a node communication feature matrix is constructed. The row dimension of the matrix corresponds to the node identifier in the distributed network, the column dimension corresponds to different chaotic feature value types, and the matrix element values reflect the chaotic feature attributes of the corresponding node. Calculate the similarity of chaotic feature values of each node with other nodes, and generate an adjacency matrix containing a similarity measure between nodes. The element values of the adjacency matrix are positively correlated with the feature similarity between nodes. Based on the adjacency matrix, a topological association structure of node communication features is constructed, where nodes are represented as vertices of the graph, and the similarity measure between nodes is represented as edge weights of the graph, generating a weighted undirected graph containing node connection relationships. The weighted undirected graph is subjected to community partitioning. The optimal number of communities is determined based on the eigenvalue decomposition of the graph's Laplacian matrix. The topological association structure is divided into node communities with similar communication characteristics, and each node community corresponds to a node communication mode. Construct a topological space of a preset dimension, and map the topological structure of the weighted undirected graph to the topological space using a multidimensional scaling analysis method, preserving the relative distance relationship between nodes, and generating a coordinate position matrix of the nodes in the topological space; Extract the geometric distribution features of each node community in the coordinate position matrix, calculate the center coordinates, distribution radius and shape parameters of the community, and generate an attack mode topology expression vector to describe the spatial distribution characteristics of abnormal node communication patterns.
10. A computer system, characterized in that, include: processor; And a memory, wherein the memory stores computer-readable code that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Network security situation prediction method based on improved BPNN (back propagation neural network)
CN106453293A
Network defense capability verification method and system based on intrusion attack simulation
CN120090868A