A data analysis and management method and system based on marine network security

By building a dynamic topology model and anomaly detection technology, reconstructing the attack path, and generating the optimal protection response plan, the problems of dynamic node changes and threat concealment in marine networks are solved, and efficient threat identification and protection response are achieved.

CN120602232BActive Publication Date: 2025-10-14SHANDONG PROVINCIAL INST OF LAND & SPACE DATA & REMOTE SENSING TECH (SHANDONG PROVINCIAL SEA AREA DYNAMIC SURVEILLANCE & MONITORING CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511093594.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-14
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

The security protection of marine networks faces problems such as dynamic changes in nodes, high threat concealment, complex attack propagation paths, and limited resources. Existing technologies are difficult to adapt to marine scenarios, resulting in delayed threat identification, difficulty in tracing attack sources, and inefficient protection responses.

Method used

By collecting multi-source marine network equipment traffic and log data in real time, constructing dynamic heterogeneous data streams, extracting spatiotemporal correlation feature tensors, performing behavioral topology modeling and anomaly detection, reconstructing attack paths, combining multi-dimensional security assessment functions to make protection response decisions, and using a hybrid multi-objective optimization algorithm to generate the optimal protection plan.

Benefits of technology

It improves environmental adaptability, increases the accuracy of anomaly identification and the efficiency of attack chain boundary identification, outputs customized protection solutions, and enhances the security of marine networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602232B_ABST
    Figure CN120602232B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data analysis management method and system based on marine network security, it is related to marine network security technical field, its technical key points include the following steps: real-time acquisition multi-source marine network equipment flow and log data, construct dynamic heterogeneous data flow, and extract space-time correlation feature tensor;Space-time correlation feature tensor is modeled to behavior topology, generates network behavior dynamic topology graph, and calculates topological stability measure;Abnormal behavior detection is carried out based on topological stability measure, and abnormal behavior cluster is identified by multi-scale spectral clustering algorithm, generates high-risk threat area coordinate set;Technical effect is adapted to the characteristics such as marine network equipment movement, communication interval by dynamic topology modeling, and environmental adaptability is improved;Combined with manifold mapping and adaptive clustering, the accuracy of abnormal identification is improved, false alarm rate is reduced, and threat accurate positioning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of marine network security, and in particular to a data analysis and management method and system based on marine network security. Background Art

[0002] With the deep integration of marine economy and information technology, marine networks have become key infrastructure supporting core scenarios such as ocean shipping, seabed resource exploration, and marine environmental monitoring. Its network architecture covers heterogeneous equipment such as mobile ship terminals, underwater sensor nodes, shore-based control centers, and satellite / underwater acoustic communication links. It features dynamic node movement, complex communication environment (such as underwater signal attenuation and satellite transmission delay), and strong business relevance (such as deep coupling of navigation data and ship control).

[0003] However, the security protection of marine networks faces multiple challenges: First, the network topology changes dynamically (such as ships sailing causing frequent access / exit of communication nodes), and traditional static baseline construction methods are difficult to adapt, which can easily lead to "normal behavior misjudgment"; second, threats are highly concealed, and attack behaviors are often mixed in complex communication noise (such as underwater sensor data packet loss and malicious tampering are difficult to distinguish), and existing detection technologies based on fixed feature libraries have a high false negative rate; third, the attack propagation path is complex, and threats can spread across subnets through satellite links (such as spreading from ship terminals to shore-based servers), and there is a lack of effective path tracking and boundary definition methods; fourth, protection resources are limited, and the computing power and bandwidth of ocean-going ships and underwater equipment are limited, making it difficult to support large-scale protection deployment. A precise balance must be achieved between resource consumption and protection effect.

[0004] Existing network security technologies are mostly designed for fixed terrestrial networks and fail to fully consider the unique characteristics of marine scenarios. For example, traditional topology modeling ignores node mobility, resulting in distorted stability measurements; general clustering algorithms are not adapted to the nonlinear distribution of marine data, resulting in inaccurate anomaly detection; attack chain analysis lacks support for cross-subnet propagation; and protection decisions fail to incorporate multi-objective optimization logic under resource constraints. These shortcomings lead to delayed threat identification, difficulty in tracing attacks, and inefficient protection responses in marine networks. A comprehensive security management solution tailored to marine scenarios is urgently needed. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a data analysis and management method and system based on marine network security to solve the problems in the above-mentioned background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a data analysis and management method based on marine network security, comprising the following steps:

[0007] Collect multi-source marine network equipment traffic and log data in real time, build dynamic heterogeneous data streams, and extract spatiotemporal correlation feature tensors;

[0008] Perform behavioral topology modeling on spatiotemporal correlation feature tensors, generate dynamic topological graphs of network behavior, and calculate topological stability measures;

[0009] Abnormal behavior detection is performed based on topological stability measurement, and abnormal behavior clusters are identified through multi-scale spectral clustering algorithm to generate a coordinate set of high-risk threat areas;

[0010] Reconstruct the attack path of the high-risk threat area coordinate set and generate the initial attack chain boundary using the spatiotemporal graph model matching algorithm;

[0011] Construct a multi-dimensional security assessment function, use a hybrid multi-objective optimization algorithm to make evolutionary decisions on the initial attack chain boundary, output the optimal protection response plan and execute it.

[0012] In a preferred embodiment, the behavior topology modeling of the spatiotemporal correlation feature tensor is performed to generate a network behavior dynamic topology graph, and the topology stability measure is calculated, specifically:

[0013] Parse the three-dimensional data of protocol type, packet size, and access frequency in the spatiotemporal correlation feature tensor and construct a dynamic attribute graph;

[0014] Calculate the behavioral correlation matrix between nodes in the attribute graph and introduce persistent homology theory to extract topologically invariant features;

[0015] Fusion of behavioral association matrix and topology invariant features to generate weighted network behavior dynamic topology graph;

[0016] The rate of change of the Betti number of the topological graph on consecutive time slices is calculated as a measure of topological stability.

[0017] In a preferred embodiment, the abnormal behavior detection is performed based on the topological stability measure, and the abnormal behavior clusters are identified by the multi-scale spectral clustering algorithm to generate the high-risk threat area coordinate set, specifically:

[0018] Map the topological stability measure to the Riemannian manifold space and construct the behavioral geodesic distance matrix;

[0019] The spectral clustering algorithm modified by Mahalanobis distance is used to perform eigendecomposition on the behavioral geodesic distance matrix;

[0020] Adaptively cluster feature vectors using a Dirichlet process mixture model to identify clusters of abnormal behaviors exceeding a risk threshold.

[0021] Extract the IP coordinates, port vectors, and timestamps of nodes in the abnormal behavior cluster to generate a coordinate set of high-risk threat areas.

[0022] In a preferred embodiment, the attack path is reconstructed for the high-risk threat area coordinate set, and the initial attack chain boundary is generated using a spatiotemporal graph model matching algorithm, specifically:

[0023] Analyze the spatiotemporal distribution of high-risk threat area coordinate sets and construct an attack causal graph model;

[0024] Use random walk algorithm to simulate threat propagation paths and generate candidate attack chains;

[0025] Calculate the graph structure similarity between the candidate attack chain and the historical attack pattern, and select the path with similarity higher than the preset value as the initial attack chain boundary.

[0026] In a preferred embodiment, the method for analyzing the spatiotemporal distribution of the high-risk threat area coordinate set and constructing the attack causal graph model is specifically as follows:

[0027] Each high-risk threat coordinate point is used as a graph node, and the node attributes include the threat type weight;

[0028] Calculate the transition probability between nodes based on time sequence and protocol correlation;

[0029] The maximum information coefficient method is used to quantify the causal strength between nodes, and edges with strength exceeding the threshold are retained;

[0030] The transition probability and causal strength are integrated to construct a weighted directed attack causal graph.

[0031] In a preferred embodiment, the multi-dimensional security assessment function is constructed, a hybrid multi-objective optimization algorithm is used to make an evolutionary decision on the initial attack chain boundary, and the optimal protection response solution is output and executed, specifically:

[0032] Construct a multi-dimensional security assessment function that includes resource consumption, response time, and risk coverage;

[0033] The initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-Ⅲ algorithm is used for non-dominated sorting;

[0034] The quantum rotating gate mechanism is introduced to perform adaptive crossover mutation and generate the Pareto optimal solution set;

[0035] The optimal protection response plan is selected from the Pareto solution set based on the entropy weight TOPSIS method.

[0036] In a preferred embodiment, the initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-III algorithm is used for non-dominated sorting, specifically:

[0037] K-means++ is used to initialize the reference point set and dynamically adjust the reference point distribution;

[0038] Introducing convolutional neural networks to predict the convergence trend of the solution set and adaptively adjust the crossover probability;

[0039] A crowding operator based on topological structure similarity is designed to optimize the distribution of solution sets.

[0040] Compared with the existing technology, the present invention provides a data analysis and management method and system based on marine network security, which has the following beneficial effects: through dynamic topology modeling to adapt to the characteristics of marine network equipment mobility, communication intermittentness, etc., environmental adaptability is improved; combined with manifold mapping and adaptive clustering, the accuracy of anomaly identification is improved, the missed reporting rate is reduced, and precise threat positioning is achieved; based on causal graphs and path simulation to restore attack propagation, the accuracy of attack chain boundary identification is improved and the efficiency is improved; through multi-objective optimization to output customized protection solutions, the overall effectiveness is improved; the distributed closed-loop design improves deployment flexibility, forms a complete security link, and comprehensively breaks through the constraints of marine network dynamics, threat concealment and limited resources, providing efficient and safe protection for various marine facilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A flowchart of a data analysis and management method based on marine network security according to the present invention;

[0042] Figure 2 A dynamic attribute graph modeling block diagram of the data analysis and management method based on marine network security of the present invention;

[0043] Figure 3 A multi-objective optimization decision-making block diagram of the data analysis and management method based on marine network security of the present invention;

[0044] Figure 4 A block diagram of topological stability measurement calculation for a data analysis and management method based on marine network security according to the present invention;

[0045] Figure 5 This is a flowchart of the awareness module of the data analysis and management system based on marine network security of the present invention;

[0046] Figure 6 This is a flowchart of the attack deduction module of the data analysis and management system based on marine network security of the present invention;

[0047] Figure 7 This is a flowchart of the threat location module of the data analysis and management system based on marine network security of the present invention;

[0048] Figure 8 This is a flowchart of a topology modeling module of a data analysis and management system based on marine network security according to the present invention;

[0049] Figure 9 This is a flow chart of the dynamic perception module of the data analysis and management system based on marine network security of the present invention. DETAILED DESCRIPTION

[0050] In the present invention, unless otherwise specified, directions such as "up" and "down" are generally used with respect to the directions shown in the drawings, or with respect to the vertical, perpendicular or gravity directions; similarly, for ease of understanding and description, "left" and "right" are generally used with respect to the left and right shown in the drawings; "inside" and "outside" refer to the inside and outside relative to the outline of each component itself, but the above-mentioned directions are not used to limit the present invention.

[0051] Example 1, reference Figures 1 to 9 A data analysis and management method based on marine network security specifically includes the following steps:

[0052] Collect multi-source marine network equipment traffic and log data in real time, build dynamic heterogeneous data streams, and extract spatiotemporal correlation feature tensors.

[0053] The behavior topology modeling of the spatiotemporal correlation feature tensor is performed to generate a network behavior dynamic topology graph and calculate the topology stability measure, specifically:

[0054] Parse the three-dimensional data of protocol type, packet size, and access frequency in the spatiotemporal correlation feature tensor and construct a dynamic attribute graph;

[0055] It should be noted that protocol type: distinguishes interaction rules such as TCP / UDP / ICMP. For example, specific protocols are often used for communication between ships and shore-based platforms, and abnormal protocols may indicate threats; data packet size: normal business data has a fixed size range. For example, the data returned by sensors is usually small, and abnormally large packets may contain malicious code; access frequency: normal communication between devices is periodic, such as underwater robots reporting data at regular intervals. Sudden frequency changes may be a precursor to an attack.

[0056] Dynamic Attribute Graph; Nodes: Represent entities in the marine network, such as ship terminals, underwater sensors, shore-based servers, communication ports, etc.; Edges: Represent the interaction relationship between nodes, such as data transmission and command issuance; Attributes: The weight of the edge integrates the above three-dimensional data, such as high frequency + specific protocol interaction has a higher weight; Dynamicity: The graph structure is updated over time (for example, if the movement of the ship causes the communication object to change, the edges in the graph will increase or decrease in real time).

[0057] Calculate the behavioral correlation matrix between nodes in the attribute graph and introduce persistent homology theory to extract topologically invariant features;

[0058] It should be noted that the behavior association matrix quantifies the closeness of interaction between nodes. The matrix element A[i,j] represents the strength of the behavioral association between nodes i and j (calculated based on three-dimensional data. For example, the higher the interaction frequency and the higher the protocol matching degree, the larger the value).

[0059] It reflects the basic relationship of "who frequently interacts with whom" in the network and provides a numerical basis for the topological structure.

[0060] Function: Extract "core features that are resistant to local perturbations" (i.e., topological invariants) from dynamically changing network structures.

[0061] Specific operation: By analyzing the birth and death of network connected branches at different scales (for example, temporary connections that disappear under small disturbances are filtered out, and long-term stable core connections are retained), the "skeleton structure" of the network is obtained.

[0062] Significance in marine scenarios: Marine network equipment (such as ships) is highly mobile and has frequent temporary connections, but core services (such as navigation data transmission) have stable topology. Persistent synchronization can filter out noise and preserve key structures.

[0063] Fusion of behavioral association matrix and topology invariant features to generate weighted network behavior dynamic topology graph;

[0064] It should be noted that the "behavioral association matrix" (quantified interaction strength) is combined with the "topologically invariant features" (core structure) to form the final topological map.

[0065] Preserve the core nodes and connections extracted by persistent homology (to ensure topological stability);

[0066] Use the weight value of the behavior association matrix to mark the importance of the edge (for example, the weight of the key business link is significantly higher than that of the ordinary link);

[0067] Refresh the graph structure over time slices (e.g., every 10 minutes) to reflect temporal changes in network behavior.

[0068] The value of topology maps: They intuitively present the "behavioral patterns" of marine networks. Under normal circumstances, the core structure of the topology map (such as the communication links between shore-based and main vessels) is stable, and the weight distribution conforms to business rules. However, in abnormal situations, structural mutations may occur (such as unauthorized nodes accessing core links).

[0069] The rate of change of the Betti number of the topological graph on consecutive time slices is calculated as a measure of topological stability.

[0070] It should be noted that Betti-0 represents the number of connected branches (e.g., the network is divided into several independent parts); Betti-1 represents the number of rings (e.g., the number of closed communication loops, reflecting the network redundancy capability); in marine networks, the Betti number changes smoothly under normal conditions (e.g., the connected branches are stable and the key loops continue to exist).

[0071] Calculate the ratio of the difference in Betti numbers between consecutive time slices (e.g., t1 to t2, t2 to t3) to the time interval. A larger rate of change indicates a more significant deviation from the normal topology. For example, the sudden appearance of a large number of new connected branches may indicate a malicious node intrusion; a sudden decrease in the ring structure may indicate the severance of a critical link.

[0072] Ultimately, the rate of change is used as a "topological stability measure" to provide a quantitative basis for subsequent anomaly detection (if it exceeds the threshold, it will be marked as suspicious).

[0073] In the second embodiment, abnormal behavior detection is performed based on topological stability measurement, and abnormal behavior clusters are identified through a multi-scale spectral clustering algorithm to generate a high-risk threat area coordinate set. Specifically,

[0074] Map the topological stability measure to the Riemannian manifold space and construct the behavioral geodesic distance matrix;

[0075] It's important to note that topological stability measures (such as the rate of change of the Betti number) reflect the nonlinear dynamics of network structure. Traditional Euclidean space struggles to capture these complex nonlinear relationships (for example, different types of attacks may lead to similar stability changes, but for different underlying reasons). Riemannian manifold space better fits the distribution characteristics of high-dimensional, nonlinear data and is particularly well-suited for describing data with inherent geometric structure, such as network behavior.

[0076] The topological stability measure of each time slice (which may be a multidimensional vector, such as the rate of change of Betti-0 and Betti-1) is regarded as a point on the manifold, and the "geodesic distance" between points (the shortest path on the manifold surface, not the straight-line distance) is calculated using differential geometry methods.

[0077] The matrix element D[i,j] represents the degree of difference in network behavior between the i-th time slice and the j-th time slice. Larger values ​​indicate a more significant divergence in network behavior patterns between the two time slices. Compared to traditional distance calculations (such as Euclidean distance), this matrix more accurately captures behavioral differences in marine networks that are "superficially similar but fundamentally different," such as topological changes caused by normal ship movement versus those caused by malicious intrusions.

[0078] The spectral clustering algorithm modified by Mahalanobis distance is used to perform eigendecomposition on the behavioral geodesic distance matrix;

[0079] It should be noted that spectral clustering, a graph-theoretic clustering method, is effective for data with non-convex distributions and is particularly well-suited for clustering data with complex correlation structures, such as network behavior (traditional algorithms like K-means are prone to local optima when handling such data). Marine network data exhibits significant "feature correlation" and "scale differences" (for example, the range of access frequency values ​​is much larger than the value encoded by protocol type). The Mahalanobis distance eliminates the dimensionality effects of data of different dimensions and considers covariance between features (for example, specific protocol types are often associated with specific packet sizes), making distance calculations more consistent with real-world business logic. The Mahalanobis distance is used to weight the original geodesic distance matrix, highlighting the influence of key features (such as unusual protocol types). Spectral decomposition (calculating the eigenvalues ​​and eigenvectors of the Laplacian matrix) is performed on the modified distance matrix, mapping the high-dimensional distance matrix to a lower-dimensional space while preserving the core structural information of the data and providing more manageable eigenvectors for subsequent clustering.

[0080] Adaptively cluster feature vectors using a Dirichlet process mixture model to identify clusters of abnormal behaviors exceeding a risk threshold.

[0081] It's important to note that traditional clustering algorithms (such as K-means) require a predefined number of clusters. However, the types of anomalies in marine networks are often unknown (e.g., new attack patterns). DPMM uses a nonparametric Bayesian approach to adaptively learn the number of clusters. The number of clusters automatically corresponds to the number of behavioral patterns present in the data, eliminating the need for manual pre-setting. Using the low-dimensional feature vectors derived from spectral decomposition as input, DPMM iteratively calculates the posterior probability that each sample belongs to a different cluster, ultimately clustering time slices with similar behavioral patterns into a single category (e.g., normal communication clusters, suspected scanning clusters, data leakage clusters, etc.). A "risk score" is calculated for each cluster (based on the degree to which the behavior within the cluster deviates from historical norms). Clusters marked as anomalous when the score exceeds a preset risk threshold are dynamically adjusted for different sea areas (e.g., nearshore vs. offshore) and device types (e.g., sensors vs. servers) to avoid misjudgments due to environmental variations.

[0082] Extract the IP coordinates, port vectors, and timestamps of nodes in the abnormal behavior cluster to generate a coordinate set of high-risk threat areas.

[0083] It should be noted that: from the time slice corresponding to the abnormal behavior cluster, the network nodes involved in the abnormal behavior are located and three types of core information are extracted: IP coordinates: network layer location (such as the public network IP of the ship terminal and the intranet IP of the underwater sensor); port vectors: transport layer identification (such as abnormally open port numbers and frequently interacting port combinations, such as port 21 (FTP) and port 3389 (Remote Desktop) being active at the same time may indicate an intrusion); timestamp: the specific time and duration of the abnormal behavior (used to analyze the attack rhythm, such as whether it is a continuous attack).

[0084] The composition of the high-threat area coordinate set: A coordinate set is a multidimensional set, with each element formatted as (IP address, port set, start time, end time, threat confidence), visually displaying the network location, resources involved, time range, and risk level of the threat. For example, {(192.168.1.105,[22,8080],10:23:15,10:45:30,0.92),...} indicates that the IP address generated high-confidence anomalous behavior on ports 22 and 8080 within the specified timeframe.

[0085] The attack path is reconstructed for the high-risk threat area coordinate set, and the initial attack chain boundary is generated using the spatiotemporal graph model matching algorithm, specifically:

[0086] Analyze the spatiotemporal distribution of high-risk threat area coordinate sets and construct an attack causal graph model;

[0087] It should be noted that the "temporal and spatial distribution characteristics" of the high-risk threat area coordinate concentration include: time dimension: the order of the timestamps of each high-risk coordinate point (for example, the anomaly at point A occurs at t1, point B at t2, and t1 <t2可能暗示A到B的传播);空间维度:节点间的网络拓扑关系(如IP地址所属网段、端口服务依赖关系,如数据库服务器(3306端口)通常依赖认证服务器(8080端口));属性维度:威胁类型关联(如端口扫描(异常端口访问)常伴随暴力破解(多次失败登录))。

[0088] The attack causal graph model consists of the following components: Nodes: coordinates of each high-risk threat point (including attributes such as IP address, port number, timestamp, and threat type); Directed Edges: represent the "possibility of causal association" between nodes, with the direction of the arrow reflecting temporal order or propagation direction; Edge Weights: incorporate factors such as time difference, network distance, and threat type similarity to quantify the strength of causal relationships (e.g., nodes with shorter time intervals and closer network distances receive higher weights). Marine scenario adaptation: To address the mobility of ship networks (e.g., dynamic IP address allocation) and communication latency of underwater equipment (possible timestamp deviations), a "spatiotemporal elasticity coefficient" is introduced into the causal graph to correct for biases in causal judgments caused by environmental characteristics.

[0089] Use random walk algorithm to simulate threat propagation paths and generate candidate attack chains;

[0090] It's important to note that the core function of the random walk algorithm is to simulate the potential threat propagation path within the attack causal graph based on the strength of causal relationships (edge ​​weights) between nodes. The algorithm begins with an initial high-risk node (the node that first exhibits an anomaly) and randomly selects the next node based on the edge weight probability. It iterates until no new nodes can be reached, forming a complete path.

[0091] Key parameter design: Walk step limit: Set a maximum number of steps (e.g., 10) based on the scale of the marine network (e.g., a ship's local area network typically contains 10-50 nodes) to avoid redundancy caused by overly long paths; Restart probability: Set a certain probability (e.g., 15%) to return to the starting point from the current node and restart the walk to ensure coverage of multi-branch paths (e.g., an attack may spread to multiple devices simultaneously); Weight decay factor: Reduce the impact of edge weights as the number of steps increases, simulating the energy attenuation of threat propagation (e.g., later nodes are less directly affected by the initial attack).

[0092] Generation of candidate attack chains: Through multiple random walks (e.g., 100), all valid, non-repeating paths are collected to form a set of candidate attack chains. Each candidate chain includes information such as the node sequence, total propagation time, and cumulative threat strength (e.g., [(A, t1) → (B, t2) → (C, t3)], total propagation time 30 minutes, cumulative threat strength 0.85).

[0093] Calculate the graph structure similarity between the candidate attack chain and the historical attack pattern, and select the path with similarity higher than the preset value as the initial attack chain boundary.

[0094] It should be noted that the construction of the historical attack pattern library stores the typical path structure of known marine network attack cases (such as the chain structure of "port scanning → vulnerability exploitation → data theft"), and each pattern is saved in the form of a graph structure (including node type, edge relationship, key step sequence, etc.).

[0095] Graph structure similarity calculation method: adopt a weighted combination of "graph edit distance" and "node attribute matching degree". Graph edit distance: quantifies the minimum operations required to transform the candidate chain graph into the historical pattern graph (such as node addition / deletion, edge direction change). The smaller the value, the more similar the structure. Node attribute matching degree: compares the consistency of threat type, port service and other attributes of the corresponding nodes in the candidate chain and the historical pattern (if both contain the "SSH port (22) brute force cracking" node, the matching degree is high).

[0096] The method for analyzing the spatiotemporal distribution of the high-risk threat area coordinate set and constructing the attack causal graph model is specifically as follows:

[0097] Each high-risk threat coordinate point is used as a graph node, and the node attributes include the threat type weight;

[0098] It should be noted that: Basic identification: Each node corresponds to a specific threat point in the high-risk threat area coordinate set, contains core identification information (IP address, port number, timestamp), and uniquely determines the network location and time when the threat occurred;

[0099] Threat type weight: A core attribute of a node, quantifying the risk level and type characteristics of the threat point. For example, type classification includes common marine network threat types such as port scanning (weight 0.3), brute force cracking (0.5), data leakage (0.8), and remote control (0.9). Weight calculation combines the destructiveness of the threat behavior (e.g., data leakage has a higher weight than port scanning), target importance (e.g., attacking navigation servers has a higher weight than attacking ordinary sensors), and frequency of occurrence (multiple anomalies from the same IP address are weighted cumulatively). Extended attributes include device type (ship terminal / underwater sensor / shore-based server), subnet (e.g., ship A's local area network / ocean communication network), and protocol type (e.g., satellite communication protocol / underwater acoustic communication protocol), providing a multi-dimensional basis for subsequent causal associations.

[0100] In view of the special characteristics of marine equipment, "mobility coefficient" (such as 0.8 for ship nodes and 0.2 for fixed shore-based nodes) and "communication stability" (such as underwater sensors may have an instability coefficient of 0.3 due to signal attenuation) are added to the node attributes to correct subsequent causal strength calculations.

[0101] Calculate the transition probability between nodes based on time sequence and protocol correlation;

[0102] It should be noted that for two nodes u (timestamp t1) and v (timestamp t2), if t2>t1 (v occurs after u), there is a possibility that u will spread to v. The smaller the time interval Δt=t2-t1, the higher the basic value of the transfer probability (for example, the basic value is 0.8 when Δt<5 minutes, and drops to 0.2 when Δt>30 minutes). If t2 <t1,则v向u传播的概率为0(排除时间倒流的因果关系)。

[0103] Analyze the protocol interaction relationship between nodes u and v. For example, if u's abnormal port is 80 (HTTP) and v's abnormal port is 3306 (MySQL), and there is an HTTP to database business call relationship between the two, then the protocol correlation is 0.7; if u uses the ship-specific AIS protocol and v uses the ordinary TCP protocol and there is no business intersection, then the correlation is 0.1.

[0104] The transition probability from node u to v P(u→v) = α × time factor + (1-α) × protocol association, where α is the weight coefficient (usually 0.6, giving priority to time).

[0105] For example: from u (t1=10:00, HTTP protocol) to v (t2=10:03, MySQL protocol), the time factor is 0.8, and the protocol correlation is 0.7, then P(u→v)=0.6×0.8+0.4×0.7=0.76.

[0106] The maximum information coefficient method is used to quantify the causal strength between nodes, and edges with strength exceeding the threshold are retained;

[0107] It should be noted that MIC is an indicator that measures the strength of the nonlinear correlation between two variables (with a value range of 0-1). It can effectively capture the complex causal relationships in network threats (such as the indirect impact of non-direct transmission and the delayed onset of attack effects), overcoming the limitations of traditional linear correlation analysis.

[0108] Using the multi-dimensional attributes of nodes u and v (threat type, device type, subnet information, time difference, etc.) as variables, the MIC values ​​of both nodes are calculated to obtain the basic causal strength. For example, if u and v are both "brute force attacks" and belong to the same ship subnet, the MIC value is 0.8. If u is "port scanning" and v is "data exfiltration," but they belong to different subnets and have no protocol association, the MIC value is 0.2.

[0109] Combining the transition probability and the MIC value, we calculate the final causal strength: Causal Strength S(u→v) = P(u→v) × MIC(u,v). We set a threshold (e.g., 0.5, which can be adjusted dynamically based on network size) and retain only directed edges where S(u→v) exceeds the threshold. We filter out weakly correlated, redundant edges to ensure the simplicity and effectiveness of the graph structure.

[0110] The transition probability and causal strength are integrated to construct a weighted directed attack causal graph.

[0111] It should be noted that the final composition of the graph structure is as follows: Node set: all high-risk threat coordinate points (including attribute information); Directed edge set: screened strong causal correlation edges, whose direction is determined by time sequence (from earlier occurrence nodes to later occurrence nodes); Edge weight: Using the fusion value S(u→v), it reflects both the transfer possibility (P) and the causal correlation (MIC). The higher the weight, the more credible the propagation path.

[0112] Dynamic update mechanism: As new high-risk threat coordinate points are added (such as new anomalies detected in real time), the attack causal graph will dynamically iterate: new nodes are added and their causal strength with existing nodes is calculated; if the weight of the new edge exceeds the threshold, it is added to the graph; and the weight of the existing edge is re-evaluated (for example, the new node may strengthen or weaken the existing causal relationship).

[0113] To address issues such as satellite communication delays and underwater equipment time lags, a "time elastic window" (such as allowing a time error of ±3 minutes) is introduced to avoid misjudgment of causal relationships due to communication delays; for mobile nodes (such as ships), the subnet association is corrected according to their navigation trajectory to ensure the accuracy of cross-regional propagation paths.

[0114] The multi-dimensional security assessment function is constructed, and a hybrid multi-objective optimization algorithm is used to make an evolutionary decision on the initial attack chain boundary, output the optimal protection response plan and execute it. Specifically,

[0115] Construct a multi-dimensional security assessment function that includes resource consumption, response time, and risk coverage;

[0116] It should be noted that this process is divided into four progressive steps, from building an assessment system to outputting the optimal solution, forming a complete decision-making optimization chain: building a multi-dimensional security assessment function that includes resource consumption, response time, and risk coverage;

[0117] Design logic for evaluation dimensions: Aiming at the core demand for marine cybersecurity protection (rapidly covering risks with limited resources), three dimensions are designed that are mutually constrained and require coordinated optimization:

[0118] Resource consumption function (f1): Quantifies the cost of various resources required to implement the protection plan, including:

[0119] Hardware resources: such as the number of rules enabled in the firewall and the computing power utilization rate of the isolation device; communication resources: such as the bandwidth occupied by encrypted transmission (the bandwidth of marine satellite communications is usually limited); labor costs: such as the time cost of manual intervention (the human response delay is high in offshore scenarios); function form: f1=ω1×hardware utilization rate+ω2×bandwidth consumption+ω3×manpower hours (ω is the weight of each resource, which is dynamically adjusted according to the real-time resource shortage).

[0120] Response time function (f2): Measures the execution speed and effectiveness time of protective measures. Key indicators include: decision delay: the time from attack chain identification to solution generation; execution delay: the time it takes for the solution to be delivered to the device and take effect (ship-to-shore communication delay needs to be considered); threat containment time: the time from the execution of protection to the cessation of threat spread; function form: f2 = τ1 × decision delay + τ2 × execution delay + τ3 × containment time (τ is the time weight, and the τ3 weight is significantly increased in emergency threat scenarios).

[0121] Risk coverage function (f3): Evaluates the interception effect of the protection scheme on the attack chain. The core indicators include: key node protection rate: the protection coverage rate of core equipment in the attack chain (such as navigation servers); attack path blocking rate: the proportion of paths that are effectively cut off in the initial attack chain boundary; secondary threat prevention rate: the potential spread risk avoided by protection measures (such as preventing the attack from spreading to other ships); function form: f3=1-(λ1×proportion of unprotected nodes+λ2×proportion of unblocked paths+λ3×secondary risk probability) (the larger the value, the better the coverage effect).

[0122] There are natural contradictions among the three functions (for example, a solution with low resource consumption may have insufficient coverage, and a solution with fast response may consume too many resources), and a balance point needs to be found through multi-objective optimization.

[0123] The initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-Ⅲ algorithm is used for non-dominated sorting;

[0124] It's important to note that the chromosome encoding method converts the protection decision variables (e.g., the protection measures for each node: blocking, monitoring, or encryption) in the initial attack chain boundary into a gene sequence. For example, the gene locus corresponds to each node or path in the attack chain; the allele represents the specific protection action (e.g., 0 = no action, 1 = port blocking, 2 = traffic encryption, 3 = device isolation); and the chromosome length is equal to the number of nodes / paths requiring decision making in the attack chain (e.g., if there are 10 nodes, the chromosome length is 10). For example, the chromosome "1-2-3-0-1" means blocking node 1, encrypting node 2, isolating node 3, not taking action on node 4, and blocking node 5.

[0125] Improved core optimizations of the NSGA-III algorithm: NSGA-III is a classic algorithm for multi-objective optimization, with three improvements for marine scenarios: Dynamic adjustment of reference points: Dynamically adjusts the distribution of reference points based on the risk level of the attack chain (e.g., extremely high risk when involving nuclear-powered ships), prioritizing the optimization of high-risk dimensions; Improved convergence speed: Introduces attack chain topology complexity factors (e.g., the number of path branches), adopts a fast convergence strategy for simple attack chains, and employs a refined search for complex chains (multi-branch, cross-subnet); Enhanced constraint processing: Incorporates marine network-specific constraints (e.g., energy consumption limits for underwater sensors, latency constraints for satellite communications), and prioritizes eliminating solutions that violate hard constraints during sorting.

[0126] The initial chromosome population (i.e., candidate protection schemes) is classified according to "dominance relationship": if scheme A is better than scheme B in all three evaluation dimensions, then A dominates B. Through multiple rounds of comparison, the population is divided into different levels of non-dominated layers (the first layer is the optimal solution candidate, the second layer is the optimal solution candidate), providing direction for subsequent optimization.

[0127] The quantum rotating gate mechanism is introduced to perform adaptive crossover mutation and generate the Pareto optimal solution set;

[0128] It should be noted that: the role of the quantum revolving door mechanism: drawing on the superposition idea of ​​quantum computing, each gene locus represents the choice of protective action in the form of probability (for example, a node has a 30% probability of being blocked and a 50% probability of being monitored), and the probability distribution is adjusted through the quantum revolving door: rotation angle design: dynamic adjustment according to the dominance level of the current solution (the higher the dominance level, the smaller the rotation angle to avoid destroying high-quality solutions); adaptive strategy: increase the mutation probability for dimensions with slow convergence (such as stagnation of risk coverage improvement), and reduce the mutation amplitude for dimensions that have been well optimized (such as resource consumption).

[0129] Crossover and mutation operations are adapted to marine scenarios: Crossover operator: prioritizes retaining the protection genes of key paths in the attack chain (such as the path leading to the shore-based center) to ensure the stability of the core protection strategy; Mutation operator: sets a higher mutation probability for the protection genes of mobile nodes (such as ships) (because changes in the ship's position may cause the original plan to become invalid).

[0130] Generation of a Pareto-optimal solution set: Through multiple iterations (e.g., 50-100 generations), the algorithm converges to a set of "non-dominated solutions." Each solution in the solution set has a different emphasis on the three evaluation dimensions (e.g., Solution X has low resource consumption but medium coverage, while Solution Y has high coverage but slightly slower response). No solution is superior to others in all dimensions, forming an "optimal trade-off set."

[0131] The optimal protection response plan is selected from the Pareto solution set based on the entropy weight TOPSIS method.

[0132] It should be noted that the advantages of the entropy weight TOPSIS method are: it combines the entropy weight method (objective weighting) and the TOPSIS method (approximating the ideal solution), avoids the deviation of subjective weight setting, and is suitable for the decision-making needs of multiple stakeholders in the ocean network (such as shipping companies and maritime departments).

[0133] Specific implementation steps: Standardization processing: standardize the three evaluation index values ​​of each plan in the Pareto solution set (eliminate dimensional differences); entropy weight calculation: calculate the weight according to the degree of dispersion of the index value (for example, the weight is higher if the risk coverage fluctuates greatly); determination of ideal solution and negative ideal solution: the ideal solution is the combination of optimal values ​​of each dimension, and the negative ideal solution is the combination of worst values ​​of each dimension; closeness calculation: the closer each plan is to the ideal solution and the farther it is from the negative ideal solution, the higher the closeness; selection of the optimal plan: select the plan with the highest closeness as the final protection response plan.

[0134] Decision preference adjustment for ocean scenarios: Emergency scenarios (such as attacks that have endangered navigation safety): Increase the weight of response time through entropy weight correction; Resource-constrained scenarios (such as limited resources for ocean-going ships): Increase the weight of resource consumption; Protection of key facilities (such as submarine optical cable nodes): Increase the weight of risk coverage.

[0135] The initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-III algorithm is used for non-dominated sorting, specifically:

[0136] K-means++ is used to initialize the reference point set and dynamically adjust the reference point distribution;

[0137] It should be noted that the limitations of traditional NSGA-III are: Traditional NSGA-III uses fixed reference points (e.g., evenly distributed in the target space). However, the optimization objectives of the marine cyber attack chain (resource consumption, response time, and risk coverage) vary significantly in importance in different scenarios (e.g., wartime scenarios focus more on response time, while daily scenarios focus more on resource consumption). Fixed reference points are difficult to adapt to such dynamic needs.

[0138] Advantages of K-means++ initialization: First, the Pareto optimal solutions of historical optimization cases are clustered (K-means++ algorithm) to identify dense areas of solutions in the target space. Initial reference points are set based on cluster centers, so that the reference points are naturally close to the actual distribution of valuable solutions, reducing ineffective searches. For example, if historical data shows a high proportion of solutions with "low resource consumption + medium coverage" (which meets the daily protection needs of ships), the initial reference points will be more densely distributed in this area.

[0139] Dynamic adjustment mechanism: After each iteration (e.g., 20 generations), the matching degree between the current population and the reference points is calculated. If the quality of the solution in a certain area continues to decline (e.g., convergence of the solution in a high-coverage area stagnates), the number of reference points in that area is automatically increased. The range of reference points is dynamically contracted or expanded (e.g., expanding the distribution of reference points in the resource consumption dimension as the attack chain scales up) based on real-time changes in the attack chain (e.g., adding new threat nodes causing target space offsets). Marine scenario adaptation: Preset reference point adjustment rules are implemented for different sea areas (nearshore / open-sea) and different ship types (commercial / military). For example, in the military ship scenario, the reference point density in the "high coverage" area is always maintained.

[0140] Introducing convolutional neural networks to predict the convergence trend of the solution set and adaptively adjust the crossover probability;

[0141] It is important to note the role and challenges of crossover probability: Crossover probability (PC) controls the probability of exchanging genes between two parent chromosomes in the algorithm. A high PC can easily destroy high-quality genes, while a low PC can slow convergence. Traditional NSGA-III uses a fixed or linearly adjustable PC, which is difficult to adapt to the complex convergence dynamics of attack chain optimization.

[0142] The prediction mechanism of the convolutional neural network (CNN): Input design: The non-dominated layer distribution, objective function value matrix, and attack chain topology characteristics (such as the number of nodes and the number of path branches) of the current population are converted into a multidimensional feature graph as the CNN input; Output target: Predict the convergence indicators of the next generation population (such as the distribution entropy of the solution in the target space and the improvement of the optimal solution); Training data: The model is trained using the population status during the historical optimization process and the subsequent convergence results, so that it can accurately predict the impact of different PCs on convergence.

[0143] Adaptive adjustment strategy: If the CNN predicts "slow convergence" (e.g., a slow decrease in the solution distribution entropy), the Pc is increased (e.g., from 0.6 to 0.8) to enhance population diversity. If the prediction is "large fluctuations in solution quality" (e.g., large alternations in the optimal solution), the Pc is reduced (e.g., from 0.6 to 0.4) to stabilize high-quality genes. Special handling for marine scenarios: When the attack chain includes underwater sensor nodes (whose protection decisions are strongly constrained by energy consumption), the CNN will additionally input the "energy sensitivity" feature, so that the Pc adjustment focuses more on retaining low-energy solutions.

[0144] A crowding operator based on topological structure similarity is designed to optimize the distribution of solution sets.

[0145] It should be noted that the traditional crowding operator has a flaw: the traditional NSGA-III measures crowding by calculating the neighborhood density of the solution in the target space, but ignores the topological correlation of the attack chain protection scheme. Two solutions that are close to each other in the target space may correspond to very different protection topologies (for example, one blocking path A and the other blocking path B), resulting in insufficient actual diversity of the solution set.

[0146] Quantification of topological similarity: Each chromosome (protection scheme) is converted into a "protection topology graph" where nodes represent attack chain nodes and edges represent the synergistic relationships between protection measures (e.g., if nodes u and v are blocked simultaneously, a synergistic edge exists). The similarity between two protection topologies is calculated using the "graph edit distance": the smaller the distance, the more similar the topologies (e.g., if there is only one difference in protection node, the distance is 1).

[0147] Combining the target space distance and topological similarity, a comprehensive crowding index is constructed: crowding = α × target space distance + (1-α) × (1-topological similarity), (α is the weight, usually 0.4, giving priority to ensuring topological diversity).

[0148] Optimization of operators: in non-dominated sorting, solutions with low crowding degree (target space dispersion and large topological difference) are given higher selection priority; avoid a large number of redundant solutions in the population with similar protection logic but different target values, and ensure that the Pareto solution set contains more diverse protection strategies (such as some focusing on node isolation and some focusing on traffic encryption); ocean scene adaptation: for cross-subnet attack chains (such as ship-satellite-land-based), increase the "subnet boundary protection" weight in topological similarity calculation to ensure that the solution set contains enough cross-network collaborative protection schemes.

[0149] Embodiment two, Figures 5 to 9 The system of the data analysis management method based on marine network security according to the present application is given in the present application, comprising:

[0150] The dynamic perception module is used for collecting multi-source network data in real time and constructing a space-time correlation feature tensor;

[0151] The topology modeling module is used for generating a network behavior dynamic topology graph and a stability measure;

[0152] The threat positioning module is used for identifying abnormal behavior clusters and generating a high-risk threat area coordinate set;

[0153] The attack deduction module is used for reconstructing an attack path and generating an initial attack chain boundary;

[0154] The decision optimization module is used for executing multi-objective optimization and outputting an optimal protection response scheme.

[0155] The above formulas are all dimensionless numerical calculations, the formula is obtained by software simulation of a large amount of data to obtain a formula of the nearest real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0156] The above embodiments can be realized wholly or partially by software, hardware, firmware or any other combination. When realized by software, the above embodiments can be realized in the form of a computer program product wholly or partially.

[0157] Those skilled in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0158] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0159] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0160] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A data analysis and management method based on marine network security, characterized in that: The following steps are involved: Collect multi-source marine network equipment traffic and log data in real time, build dynamic heterogeneous data streams, and extract spatiotemporal correlation feature tensors; Perform behavioral topology modeling on spatiotemporal correlation feature tensors, generate dynamic topological graphs of network behavior, and calculate topological stability measures; Specifically: Parse the three-dimensional data of protocol type, packet size, and access frequency in the spatiotemporal correlation feature tensor and construct a dynamic attribute graph; Calculate the behavioral correlation matrix between nodes in the attribute graph and introduce persistent homology theory to extract topologically invariant features; Fusion of behavioral association matrix and topology invariant features to generate weighted network behavior dynamic topology graph; Calculate the rate of change of the Betti number of the topological graph on consecutive time slices as a measure of topological stability; Abnormal behavior detection is performed based on topological stability measurement, and abnormal behavior clusters are identified through multi-scale spectral clustering algorithm to generate a coordinate set of high-risk threat areas; Specifically: Map the topological stability measure to the Riemannian manifold space and construct the behavioral geodesic distance matrix; The spectral clustering algorithm modified by Mahalanobis distance is used to perform eigendecomposition on the behavior geodesic distance matrix to obtain eigenvectors. Adaptively cluster feature vectors using a Dirichlet process mixture model to identify clusters of abnormal behaviors exceeding a risk threshold. Extract the IP coordinates, port vectors, and timestamps of nodes in the abnormal behavior cluster to generate a coordinate set of high-risk threat areas; Reconstruct the attack path of the high-risk threat area coordinate set and generate the initial attack chain boundary using the spatiotemporal graph model matching algorithm; Construct a multi-dimensional security assessment function, use a hybrid multi-objective optimization algorithm to make evolutionary decisions on the initial attack chain boundary, output the optimal protection response plan and execute it.

2. The data analysis and management method based on marine network security according to claim 1 is characterized by: The attack path is reconstructed for the high-risk threat area coordinate set, and the initial attack chain boundary is generated using the spatiotemporal graph model matching algorithm; specifically: Analyze the spatiotemporal distribution of high-risk threat area coordinate sets and construct an attack causal graph model; Use random walk algorithm to simulate threat propagation paths and generate candidate attack chains; Calculate the graph structure similarity between the candidate attack chain and the historical attack pattern, and select the path with similarity higher than the preset value as the initial attack chain boundary.

3. The data analysis and management method based on marine network security according to claim 2 is characterized by: The method for analyzing the spatiotemporal distribution of the high-risk threat area coordinate set and constructing the attack causal graph model is specifically as follows: Each high-risk threat coordinate point is used as a graph node, and the node attributes include the threat type weight; Calculate the transition probability between nodes based on time sequence and protocol correlation; The maximum information coefficient method is used to quantify the causal strength between nodes, and edges with strength exceeding the threshold are retained; The transition probability and causal strength are integrated to construct a weighted directed attack causal graph.

4. The data analysis and management method based on marine network security according to claim 3 is characterized by: The multi-dimensional security assessment function is constructed, and a hybrid multi-objective optimization algorithm is used to make an evolutionary decision on the initial attack chain boundary, output the optimal protection response plan and execute it. Specifically, Construct a multi-dimensional security assessment function that includes resource consumption, response time, and risk coverage; The initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-Ⅲ algorithm is used for non-dominated sorting; The quantum rotating gate mechanism is introduced to perform adaptive crossover mutation and generate the Pareto optimal solution set; The optimal protection response plan is selected from the Pareto optimal solution set based on the entropy weight TOPSIS method.

5. The data analysis and management method based on marine network security according to claim 4 is characterized in that: The initial attack chain boundary is encoded as a chromosome population, and the improved NSGA-III algorithm is used for non-dominated sorting, specifically: K-means++ is used to initialize the reference point set and dynamically adjust the reference point distribution; Introducing convolutional neural networks to predict the convergence trend of the solution set and adaptively adjust the crossover probability; A crowding operator based on topological structure similarity is designed to optimize the distribution of solution sets.

6. A data analysis and management system based on marine network security, used to implement the method described in any one of claims 1 to 5, characterized in that: include: Dynamic perception module, used to collect multi-source network data in real time and construct spatiotemporal correlation feature tensors; Topology modeling module, used to generate dynamic topology diagrams of network behavior and stability measures; Threat location module, used to identify abnormal behavior clusters and generate a set of high-risk threat area coordinates; Attack deduction module, used to reconstruct the attack path and generate the initial attack chain boundary; The decision optimization module is used to perform multi-objective optimization and output the optimal protection response plan.

Citation Information

Patent Citations

  • Computer network security intelligent analysis system and method based on big data

    CN117896137A

  • Intelligent maritime affair security situation awareness system and method

    CN120257200A