Graph data communication strength mapping method and system based on optical transport network topology

By establishing a two-layer mapping model between graph data and OTN network, and combining a multi-factor fusion algorithm to optimize communication strength weights, the problem of mismatch between graph data communication mode and optical network topology is solved, realizing efficient scheduling of graph computing tasks and optimized utilization of optical network resources.

CN121568000BActive Publication Date: 2026-04-17GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-01-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the graph data communication mode is incompatible with the underlying optical network topology, making it impossible to optimize graph computing tasks. In particular, it is difficult to achieve optimal scheduling in multi-layer optical transport network environments, resulting in cross-node communication overhead and resource waste.

Method used

By acquiring graph data structure and OTN network information, a graph data communication feature matrix and an OTN physical network feature matrix are established, and a two-layer topology association matrix is ​​constructed. Combining communication frequency, latency sensitivity, bandwidth requirements and topology matching degree, weight calculation and optimization are performed to generate a communication strength weight matrix, thereby realizing dynamic mapping relationship adjustment.

Benefits of technology

It achieves efficient collaboration between graph computing tasks and optical transport network resources, improves the accuracy of data layout decisions and the system's adaptability to environmental changes, and optimizes the utilization rate of optical network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121568000B_ABST
    Figure CN121568000B_ABST
Patent Text Reader

Abstract

The application provides a kind of based on the graph data communication intensity mapping method and system of optical transport network topology.The method is by extracting graph data structure information to establish adjacency matrix, calculate vertex degree centrality and clustering coefficient, generate graph data communication characteristic matrix;Collect OTN network equipment information and optical link topology, measure link parameters and wavelength resource state, get OTN physical network characteristic matrix;Through resource mapping algorithm to establish double-layer mapping relationship matrix;Comprehensive communication frequency, time delay sensitivity, bandwidth demand, wavelength resource and topological matching degree, calculate communication intensity weight matrix;Generate candidate layout scheme and carry out multidimensional evaluation, determine the optimized data layout scheme;Mapping relationship is dynamically adjusted using incremental updating algorithm.The application realizes the efficient cooperation of graph computing task and optical transport network resource, significantly improves the performance and resource utilization efficiency of distributed graph computing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of graph computing and optical transport networks, specifically to a graph data communication intensity mapping method and system based on optical transport network topology. It is mainly applied to data layout optimization and network resource scheduling in distributed graph computing environments, particularly graph data task scheduling in distributed graph computing environments based on computing power networks. Background Technology

[0002] Distributed graph computing is a crucial technique for processing large-scale, complex relational data, and its performance and efficiency directly impact the development of numerous application areas such as social network analysis, recommender systems, and knowledge graphs. With the explosive growth of data volume, graph computing tasks place higher demands on the underlying network transmission capabilities, especially in large-scale graph processing scenarios within geographically separated distributed environments.

[0003] Currently, mainstream distributed graph computing frameworks such as Pregel, GraphX, and PowerGraph primarily focus on graph data partitioning strategies and computational model optimization, typically employing edge or vertex partitioning methods to distribute large-scale graph data across multiple computing nodes. These frameworks are usually based on general data center network architectures, such as Ethernet or InfiniBand networks based on the TCP / IP protocol stack. When distributing data and scheduling tasks, they mainly consider computational load balancing, with less consideration for the underlying network topology characteristics.

[0004] More advanced technologies attempt to incorporate network-aware factors into graph computation scheduling, guiding data placement decisions by collecting network state information and thus bringing data closer to computation. These methods monitor parameters such as network link utilization and inter-node communication latency, adjusting graph partitioning strategies accordingly to reduce cross-node communication overhead. However, these methods remain optimizations at the logical network level, failing to penetrate the physical optical transport network layer and thus unable to fully utilize the bandwidth and topology characteristics of optical networks, especially in multi-layered optical transport network (OTN) environments.

[0005] The main shortcomings of existing technologies are twofold: firstly, the lack of a mathematical model that accurately maps graph data communication patterns to the underlying optical network topology makes it impossible to optimize graph computation tasks based on actual physical network conditions, particularly cross-domain scheduling of graph data tasks in computing power networks; secondly, existing methods cannot dynamically sense and adapt to changes in latency cycles and wavelength resource status in optical networks, making it difficult to achieve optimal scheduling when network load fluctuates. These problems are particularly prominent in large-scale graph computation tasks, resulting in unnecessary cross-node communication overhead and resource waste. Summary of the Invention

[0006] The purpose of this invention is to provide a graph data communication intensity mapping method and system based on optical transport network topology, which solves the problem of mismatch between graph data communication mode and underlying optical network topology in the prior art, and realizes efficient collaboration between graph computing tasks and optical transport network resources.

[0007] To achieve the above objectives, this invention provides a graph data communication intensity mapping method based on optical transport network topology, comprising the following steps:

[0008] Obtain the graph data structure file, extract the basic structural information of the graph through the graph data parser, establish the graph adjacency matrix and calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate the graph data communication feature matrix GCF.

[0009] Acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF).

[0010] Based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, a mapping relationship between graph computing nodes and OTN network physical nodes is established through a resource mapping algorithm, and a two-layer topology association matrix is ​​constructed to output the two-layer mapping relationship matrix DLM.

[0011] For the two-layer mapping relationship matrix DLM, communication frequency, latency sensitivity, bandwidth requirement, wavelength resources and topology matching degree are taken as input parameters. The weights of DLM are calculated and optimized by weight definition algorithm and multi-factor fusion algorithm to obtain the communication strength weight matrix CIW.

[0012] The communication strength weight matrix CIW is processed by a scheme enumeration algorithm to generate candidate layout schemes and to evaluate communication overhead, resource utilization and load balancing. The comprehensive score of topology matching degree is calculated to determine the optimized data layout scheme OLP.

[0013] Based on the optimized data layout scheme OLP, the mapping relationship is dynamically adjusted by using an incremental update algorithm through real-time monitoring of system performance indicators and network status changes, resulting in a dynamically optimized mapping relationship DOM.

[0014] The present invention also provides a graph data communication intensity mapping system based on optical transport network topology, comprising:

[0015] The acquisition module is used to acquire graph data structure files, extract basic structural information of the graph through a graph data parser, establish a graph adjacency matrix and calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate a graph data communication feature matrix (GCF).

[0016] The measurement module is used to acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF).

[0017] The mapping module is used to establish the mapping relationship between graph computing nodes and OTN network physical nodes based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF through a resource mapping algorithm, and to construct a two-layer topology association matrix and output a two-layer mapping relationship matrix DLM.

[0018] The weight calculation and optimization processing module is used to calculate and optimize the communication strength weight matrix (CIW) of the two-layer mapping relationship matrix (DLM) by taking communication frequency, delay sensitivity, bandwidth requirement, wavelength resources and topology matching degree as input parameters, and using a weight definition algorithm and a multi-factor fusion algorithm.

[0019] The evaluation module is used to process the communication strength weight matrix CIW using a scheme enumeration algorithm, generate candidate layout schemes, and evaluate communication overhead, resource utilization, and load balancing. It also calculates a comprehensive score for topology matching and determines the optimized data layout scheme OLP.

[0020] The adjustment module is used to dynamically adjust the mapping relationship based on the optimized data layout scheme OLP by monitoring system performance indicators and network status changes in real time and using an incremental update algorithm to obtain a dynamically optimized mapping relationship DOM.

[0021] The beneficial effects of this invention are as follows:

[0022] 1. A two-layer mapping model between graph data communication characteristics and OTN physical network characteristics was established, realizing cross-layer collaborative mapping between data logical topology and physical optical network topology, and improving the matching degree between graph computing tasks and underlying network resources.

[0023] 2. A method for calculating the communication strength weight factor is proposed, which comprehensively considers multiple factors such as communication frequency, latency sensitivity, bandwidth requirements, wavelength resources and topology matching degree, making graph data layout decisions more comprehensive and accurate.

[0024] 3. A topology matching degree evaluation algorithm based on multi-level time delay cycles was designed, which can quantify the matching degree between the evaluation graph calculation subgraph and the optical path, providing an accurate basis for optimization decisions.

[0025] 4. A dynamic adjustment and feedback mechanism was implemented. By monitoring changes in system performance and network status in real time, the mapping relationship was updated incrementally, avoiding the high overhead of global remapping and improving the system's adaptability to environmental changes.

[0026] 5. Innovatively, the wavelength occupancy rate of the OTN network is used as a weighting factor to optimize the allocation efficiency of optical network resources in graph computing tasks and improve the utilization rate of optical network resources. Attached Figure Description

[0027] Figure 1 This is an overall flowchart of the method of the present invention;

[0028] Figure 2 This is a flowchart of the graph data communication feature analysis and extraction process in the method of this invention;

[0029] Figure 3 This is a flowchart of the OTN physical network topology feature acquisition and analysis method of the present invention;

[0030] Figure 4 This is a flowchart of the two-layer mapping model construction process in the method of the present invention;

[0031] Figure 5 This is a flowchart of the communication strength weighting factor calculation process in the method of this invention;

[0032] Figure 6 This is a flowchart of the topology matching degree evaluation and optimization process in the method of this invention;

[0033] Figure 7 This is a schematic diagram of the system of the present invention. Detailed Implementation

[0034] Example 1

[0035] like Figure 1 As shown, the graph data communication intensity mapping method based on optical transport network topology provided by the present invention includes the following steps:

[0036] S1: Obtain the graph data structure file, extract the basic structural information of the graph through the graph data parser, establish the graph adjacency matrix, calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate the graph data communication feature matrix GCF.

[0037] S1 aims to deeply analyze the structural characteristics of graph data, providing a foundation for subsequent mapping optimization. Graph data structure files typically contain information about vertex and edge sets, possibly stored in the form of adjacency lists, adjacency matrices, or edge lists. The graph data parser first reads these files, identifies the vertex set V and edge set E of graph G, and obtains the basic topological structure of the graph.

[0038] After parsing, the system constructs the adjacency matrix A of graph G, which is the basic data structure describing the graph structure. Adjacency matrix A is a |V|×|V| matrix, where A[i][j]=1 indicates that there is an edge connection between vertices i and j, and A[i][j]=0 indicates that there is no connection. For a weighted graph, A[i][j] can be the weight value of the edge.

[0039] Next, we calculate the degree centrality D(v) of each vertex. Degree centrality is a fundamental metric in network analysis, representing the number of direct connections a vertex has to all other vertices. For an undirected graph, the degree centrality of vertex v is equal to the number of edges directly connected to v; for a directed graph, in-degree and out-degree are distinguished. Vertices with high degree centrality are usually important nodes in the network and may require more data exchange during graph computation.

[0040] In addition, the clustering coefficient C(v) for each vertex needs to be calculated. The clustering coefficient measures the tightness of connections between vertex neighbors, reflecting the local density of the graph. The specific formula is the number of actual edges between vertex v's neighbors divided by the maximum possible number of edges. Vertices with high clustering coefficients typically form dense community structures, and these vertices usually communicate more frequently.

[0041] Based on the above calculation results, the communication frequency between vertices is predicted using a communication pattern analysis algorithm, generating a communication frequency matrix F. The key to this step is transforming static graph structure features into dynamic communication behavior predictions. The prediction of communication frequency can consider various factors, such as vertex degree centrality, clustering coefficient, and the distance between vertices (shortest path length). For example, vertices with high degree centrality typically need to communicate with more vertices; regions with high clustering coefficients may have more frequent data exchanges; and communication between closer vertex pairs may be more frequent than between distant vertex pairs.

[0042] Finally, all these features are integrated to construct a comprehensive graph data communication feature matrix (GCF). This matrix includes topological features of the graph (such as vertex degree centrality and clustering coefficients) and communication pattern features (such as predicted communication frequencies), providing a comprehensive data foundation for subsequent mapping optimization.

[0043] like Figure 2 As shown, S1 specifically includes:

[0044] S1.1 Perform graph data parsing processing on the graph data structure file to extract the vertex set V and edge set E of graph G, and obtain the basic structural information of graph G;

[0045] This process first requires parsing the input graph data structure file, which may exist in various formats, such as adjacency lists, edge lists, or specialized graph database export formats. The graph data parser reads and parses these files to identify the basic components of the graph G: the vertex set V and the edge set E. The vertex set V contains information about all nodes in the graph, and may also include node attribute data; the edge set E contains the connections between all vertices, and for weighted graphs, it also includes edge weight information.

[0046] S1.2 For the vertex set V and the edge set E, establish the adjacency matrix A of the graph G using a matrix construction algorithm, where A[i][j] represents the connection relationship between vertices i and j;

[0047] After obtaining the vertex set V and edge set E, the system constructs the adjacency matrix A of graph G using a matrix construction algorithm. The adjacency matrix is ​​the standard mathematical form for representing graph structure; it is an n×n matrix (n is the number of vertices), where the element A[i][j] represents the connection between vertices i and j. For an unweighted graph, A[i][j]=1 indicates that there is an edge connection between vertices i and j, and A[i][j]=0 indicates that there is no connection. For a weighted graph, A[i][j] can be the weight value of the edge. The construction of the adjacency matrix provides a unified data structure foundation for subsequent feature calculations.

[0048] S1.3 Perform degree centrality calculation on the adjacency matrix A, calculate the degree centrality value D(v) of each vertex in the adjacency matrix A, and determine the vertex degree centrality vector D, wherein the degree centrality value D(v) is used to reflect the importance of each vertex in the graph G;

[0049] Next, the system calculates the degree centrality of the adjacency matrix A. Degree centrality is a key metric in network analysis, used to measure the importance or centrality of a vertex in a graph. For an undirected graph, the degree centrality D(v) of vertex v is equal to the number of edges directly connected to v, i.e., D(v) = ∑ j A[v][j]; For directed graphs, in-degree and out-degree are distinguished. Vertices with high degree centrality are usually "hub" nodes in the network, connecting a large number of other vertices and playing an important role in data flow. The system calculates the degree centrality value of each vertex and organizes these values ​​into a vertex degree centrality vector D, providing a basis for subsequent communication pattern analysis.

[0050] S1.4 Based on the adjacency matrix A and the degree centrality vector D, the clustering coefficient C(v) of each vertex is calculated using a clustering coefficient calculation algorithm to determine the vertex clustering coefficient vector C, wherein the clustering coefficient C(v) characterizes the connection density between the neighbors of each vertex;

[0051] The system also calculates the clustering coefficient C(v) for each vertex using a clustering coefficient calculation algorithm based on the adjacency matrix A and the degree centrality vector D. The clustering coefficient is an indicator of the local density in a graph, reflecting the tightness of connections between vertex neighbors. Specifically, for vertex v, the number of actual edges N between all its neighbors is counted, and then divided by the maximum possible number of edges between neighboring nodes (for k neighbors, the maximum possible number of edges is k(k-1) / 2), i.e., C(v) = 2N / (k(k-1)), where k is the degree of vertex v. Regions with high clustering coefficients typically form tight communities or clusters, and the data exchange frequency within these regions is often high. The system organizes the clustering coefficient values ​​of each vertex into a vertex clustering coefficient vector C.

[0052] S1.5 For the adjacency matrix A, the degree centrality vector D, and the clustering coefficient vector C, a communication frequency matrix F between each vertex is established using a communication pattern analysis algorithm, where F[i][j] represents the predicted communication frequency between vertices i and j;

[0053] Based on the previously calculated adjacency matrix A, degree centrality vector D, and clustering coefficient vector C, the system establishes a vertex communication frequency matrix F using a communication pattern analysis algorithm. The core of this step is predicting the possible communication frequency between vertex pairs during graph computation. The prediction of communication frequency can comprehensively consider multiple factors: whether vertices are directly connected, the degree centrality of the vertices, the clustering coefficient of the region where the vertex is located, and the shortest path length between vertices. For example, a weighted combination method can be used: F[i][j] = w1 × A[i][j] + w2 × (D(i) × D(j)) / max(D) + w3 × (C(i) + C(j)) / 2, where w1, w2, and w3 are weight coefficients that can be adjusted according to specific application scenarios. The element F[i][j] in the communication frequency matrix F represents the expected communication frequency between vertices i and j, which is crucial for optimizing data distribution and reducing communication overhead.

[0054] S1.6 Based on the adjacency matrix A, the degree centrality vector D, the clustering coefficient vector C, and the communication frequency matrix F, a comprehensive graph data communication feature matrix GCF is constructed through a feature fusion algorithm, wherein the graph data communication feature matrix GCF includes the topological features and communication pattern features of the graph.

[0055] Finally, based on the adjacency matrix A, degree centrality vector D, clustering coefficient vector C, and communication frequency matrix F, the system constructs a comprehensive graph data communication feature matrix (GCF) using a feature fusion algorithm. The goal of the feature fusion algorithm is to integrate multiple graph features into a unified representation that includes both static topological information and dynamic communication pattern prediction. The fusion method can be simple feature concatenation or more complex dimensionality reduction or feature transformation techniques. The resulting graph data communication feature matrix (GCF) is a multi-dimensional matrix that comprehensively describes the graph's topological and communication pattern characteristics, providing a rich and accurate data foundation for subsequent mapping optimization.

[0056] S2: Obtain data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF).

[0057] The purpose of this step is to gain a comprehensive understanding of the physical characteristics and resource status of the underlying optical transport network, providing accurate network environment information for optimizing the mapping between graph data and network resources. OTN (Optical Transport Network) is a high-bandwidth, high-reliability optical transport network technology widely used in backbone networks and metropolitan area networks.

[0058] First, device information in the OTN network is collected through the network management system's API interface or SNMP protocol. This information includes the location coordinates of network nodes, device type (such as OTU and ODU devices), number of ports, processing capacity, and other attributes. This information constitutes the OTN network node set N, which is the basic building block of the network topology.

[0059] Next, link discovery protocols (such as LLDP) or topology detection algorithms are used to identify the physical connections between nodes. This step collects data such as fiber optic link routing information, connected port information, and optical path configurations, forming an optical link set L. The optical link is the physical transmission channel in the OTN network, and its characteristics directly affect the performance and reliability of data transmission.

[0060] Then, an OTN physical topology matrix P is constructed to describe the direct connection status between network nodes. Similar to the adjacency matrix of graph data, P[i][j]=1 indicates that there is a direct optical link connection between nodes i and j, and P[i][j]=0 indicates that there is no direct connection. This matrix reflects the basic topology of the OTN network.

[0061] Furthermore, network performance monitoring tools are used to measure key parameters of each optical link, such as bandwidth capacity (Gbps), current utilization rate (%), transmission latency (ms), and bit error rate. These parameters constitute the link parameter set LP, which reflects the performance status and available resources of the network links.

[0062] A particularly important step is calculating the multi-level delay circle structure. A delay circle is a concentric circular region centered on a given node, divided according to different transmission delays. For example, the first-level delay circle might contain nodes with delays less than 1ms, the second level might contain nodes with delays between 1-5ms, and so on. This structure helps in understanding the delay distribution characteristics in the network and is especially important for delay-sensitive applications.

[0063] This application also analyzes wavelength resource status, including information such as wavelength occupancy rate, number of available wavelengths, and wavelength allocation mode for each optical link. In WDM (Wavelength Division Multiplexing) technology, each optical fiber can transmit multiple optical signals of different wavelengths simultaneously, with each wavelength corresponding to an independent communication channel. The status of wavelength resources directly affects the network's capacity and scalability.

[0064] Finally, all these network features are integrated to construct a comprehensive OTN physical network feature matrix (ONF). This matrix contains network topology, performance parameters, and resource status information, providing a comprehensive understanding of the network environment for subsequent mapping optimization.

[0065] like Figure 3 As shown, S2 specifically includes:

[0066] S2.1 Perform network detection tool and device interface processing on the existing OTN network management system to collect optical transport network device information in the OTN network, including node location, device type and functional attributes, to obtain the OTN network node set N.

[0067] In S2.1, the first step is to interact with the existing OTN network management system, collecting detailed information about the optical transport network equipment through network probing tools and device interfaces. This information typically includes the geographical coordinates of network nodes (latitude and longitude or equipment room location), equipment type (such as OTU, ODU, or ROADM devices), equipment model, processing capacity, and various functional attributes (such as supported rate levels, protection capabilities, etc.). This equipment information constitutes the OTN network node set N, which is the foundation for understanding the network topology and capabilities.

[0068] S2.2 Based on the node set N, the physical fiber optic connections and optical path information between each node in the node set N are identified through the link discovery protocol and topology detection algorithm to obtain the optical link set L.

[0069] After acquiring the node set N, the system identifies the physical connections between nodes using link discovery protocols and topology probing algorithms. Link discovery can utilize various methods, such as LLDP (Link Layer Discovery Protocol), OSPF-TE's LSA (Link State Advertisement) information, or physical layer methods like optical power probing. This step not only requires discovering the physical fiber connections between nodes but also identifying configured optical path information, including the end-to-end path, wavelength channels used, and protection type (e.g., 1+1 protection, 1:1 protection, or no protection). The system organizes this connection information into an optical link set L, where each link entry contains attributes such as starting node, ending node, link type, physical path, and optical path configuration.

[0070] S2.3 For the node set N and the optical link set L, an OTN physical topology matrix P is established using a topology matrix construction algorithm, where P[i][j] represents the direct connection state between each node i and j.

[0071] Next, the system constructs an OTN physical topology matrix P for the node set N and the optical link set L using a topology matrix construction algorithm. This matrix is ​​an n×n two-dimensional matrix (n is the number of OTN network nodes), where the element P[i][j] represents the direct connection status between nodes i and j. For a simple connection status representation, P[i][j]=1 indicates that there is a direct connection between nodes i and j, and P[i][j]=0 indicates that there is no direct connection. For a more complex representation, P[i][j] can be a structure containing information such as link type, capacity, and protection level. The physical topology matrix P provides the network structure foundation for subsequent path planning and resource allocation.

[0072] S2.4 Perform network performance monitoring on the optical link set L, measure and collect the bandwidth capacity, current utilization and transmission delay of each optical link, and obtain the link parameter set LP.

[0073] The system also performs network performance monitoring on the optical link set L, measuring and collecting various key parameters. These parameters include link bandwidth capacity (e.g., 10Gbps, 100Gbps, or higher), current utilization rate (percentage), end-to-end transmission latency (milliseconds), bit error rate, and other performance metrics. This performance data can be obtained through the network monitoring system's API or measured through active probing. The system organizes these parameters into a link parameter set LP, providing a basis for assessing network status and resource availability.

[0074] S2.5 Based on the physical topology matrix P and the link parameter set LP, the multi-level delay circle structure of each node in the OTN network is calculated using the delay circle analysis algorithm, and the delay circle matrix DC is output.

[0075] Based on the physical topology matrix P and the link parameter set LP, the system calculates the multi-level delay circle structure of each node in the OTN network using a delay circle analysis algorithm. A delay circle divides the network into multiple concentric ring regions centered on a given node, based on the transmission delay to other nodes. For example, the first-level delay circle can be defined to include nodes with a delay ≤ 1ms, the second level to include nodes with a delay between 1-5ms, the third level to include nodes with a delay between 5-10ms, and so on. Delay circle analysis can use a shortest path algorithm (such as Dijkstra's algorithm) to calculate the minimum delay path between any two points, and then group the nodes according to a delay threshold. The system outputs the calculation results as a delay circle matrix DC. This matrix is ​​particularly important for delay-sensitive applications, as it guides the system to distribute frequently communicating data within the same delay circle, reducing communication latency.

[0076] S2.6 Based on the optical link set L and real-time network monitoring data, the wavelength resource status matrix WR is obtained by statistically analyzing the wavelength occupancy rate, available wavelength quantity, and wavelength allocation mode of each optical link using a wavelength resource analysis algorithm.

[0077] The system also uses wavelength resource analysis algorithms to statistically analyze the wavelength resource status of each optical link based on the optical link set L and real-time network monitoring data. In a WDM (Wavelength Division Multiplexing) system, each optical fiber can simultaneously transmit multiple optical signals of different wavelengths, with each wavelength corresponding to an independent communication channel. Wavelength resource analysis includes calculating the wavelength occupancy rate (number of used wavelengths / total number of wavelengths), the number of available wavelengths, wavelength continuity constraints, and wavelength allocation modes (such as first-fit, random-fit, etc.) for each link. The system organizes this information into a wavelength resource status matrix WR, providing a basis for optical path planning and wavelength allocation.

[0078] S2.7 For the physical topology matrix P, the time delay circle matrix DC, and the wavelength resource state matrix WR, a comprehensive OTN physical network feature matrix ONF is constructed through a feature fusion algorithm. The OTN physical network feature matrix ONF contains network topology and performance parameter information.

[0079] Finally, the system constructs a comprehensive OTN physical network feature matrix (ONF) based on the physical topology matrix P, delay circle matrix DC, and wavelength resource state matrix WR using a feature fusion algorithm. The goal of feature fusion is to integrate network features from different dimensions into a unified representation that includes both static topology information and dynamic performance and resource state information. The fusion method can be designed according to application requirements, ranging from simple feature concatenation to complex weighted combinations. The resulting OTN physical network feature matrix (ONF) is a multi-dimensional matrix that comprehensively describes the network's topology, performance parameters, and resource state, providing accurate network environment knowledge for subsequent mapping optimization.

[0080] S3: Based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, establish the mapping relationship between graph computing nodes and OTN network physical nodes through the resource mapping algorithm, construct a two-layer topology association matrix, and output the two-layer mapping relationship matrix DLM.

[0081] This step is the core of the entire method, aiming to establish a precise mapping relationship between graph data and the OTN network, providing a foundation for subsequent optimization. The two-layer mapping model associates the logical topology of the graph data with the physical topology of the OTN network, enabling cross-layer collaborative optimization.

[0082] First, the system establishes an initial mapping relationship between graph computing nodes and OTN network physical nodes using a resource mapping algorithm, resulting in a node mapping matrix NM. This matrix describes which graph data computing nodes should be mapped to which OTN network physical nodes. The initial mapping can be based on simple heuristics, such as matching computing power requirements with network node processing capabilities as much as possible, or minimizing the overall communication distance.

[0083] Next, based on the node mapping matrix NM and the communication frequency matrix F, the system uses a path mapping algorithm to map the communication requirements between graph data vertices to the physical transmission paths of the OTN network, obtaining the path mapping matrix PM. This step requires consideration of routing selection within the OTN network; for example, strategies such as shortest path, minimum latency path, or maximum bandwidth path can be employed to transform logical communication requirements into actual physical transmission paths.

[0084] Then, based on the path mapping matrix PM and the delay circle matrix DC, the system uses a delay sensitivity analysis algorithm to evaluate the sensitivity of graph data communication to network transmission latency, obtaining the delay sensitivity matrix DS. Different graph computation tasks may have different sensitivities to latency; for example, real-time analysis tasks are usually more sensitive to latency, while batch processing tasks may prioritize throughput over latency. The delay sensitivity matrix reflects this difference, helping the system make more refined optimization decisions.

[0085] The system also estimates the network bandwidth requirements of graph data communication based on the path mapping matrix PM and the communication frequency matrix F using a bandwidth demand analysis algorithm, resulting in a bandwidth demand matrix BR. Bandwidth demand estimation needs to consider factors such as communication frequency, packet size, and communication mode. Accurate bandwidth demand estimation helps the system allocate network resources rationally and avoid bandwidth waste or bottlenecks.

[0086] Based on the above analysis, the system uses an association matrix construction algorithm to establish an association matrix between the graph data logical topology and the OTN physical topology, obtaining a two-layer topology association matrix (TRM). This matrix describes the mapping relationship between the two topologies and serves as the basis for subsequent optimization.

[0087] Finally, the system integrates all mapping information using a mapping relationship synthesis algorithm to construct the final two-layer mapping relationship matrix (DLM). This matrix provides a comprehensive description, encompassing the mapping relationships between graph data and the OTN network across multiple dimensions, including nodes, paths, latency sensitivity, and bandwidth requirements.

[0088] Specifically, such as Figure 4 As shown, S3 includes:

[0089] S3.1 Perform resource mapping algorithm processing on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF to establish the initial mapping relationship between graph computing nodes and OTN network physical nodes, and obtain the node mapping matrix NM.

[0090] This process first applies a resource mapping algorithm to the graph data communication feature matrix (GCF) and the OTN physical network feature matrix (ONF) to establish an initial mapping relationship between graph computing nodes and OTN network physical nodes. The resource mapping algorithm needs to comprehensively consider various factors, such as node processing capacity matching, geographical proximity, and functional compatibility. For example, processing capacity matching considers the computational needs of graph data vertices and the computational resources of network nodes; geographical proximity can reduce physical transmission latency; and functional compatibility ensures that specific processing requirements can be met. The resource mapping algorithm can employ multi-objective optimization methods, such as genetic algorithms, simulated annealing, or integer linear programming, to find an initial mapping scheme that balances various constraints. The system organizes the mapping results into a node mapping matrix NM, where NM[i][j]=1 indicates that vertex i in the graph data is mapped to physical node j in the OTN network, and NM[i][j]=0 indicates no mapping relationship.

[0091] S3.2 For the node mapping matrix NM and the communication frequency matrix F, the communication requirements between graph data vertices are mapped to the physical transmission path of the OTN network through the path mapping algorithm to obtain the path mapping matrix PM.

[0092] After determining the node mapping relationships, the system uses a path mapping algorithm to map the communication needs between graph data vertices to the physical transmission paths of the OTN network, based on the node mapping matrix NM and the communication frequency matrix F. Path mapping is the process of converting logical connections into physical paths, which requires consideration of various network constraints, such as link capacity, wavelength continuity, and protection requirements. The path mapping algorithm can be extended based on the shortest path algorithm (such as Dijkstra's algorithm or the K-shortest path algorithm) to add resource constraints and quality requirements. For each pair of graph data vertices (i,j) with communication needs, the system calculates the optimal physical transmission path based on the physical node positions they are mapped to. The system organizes the mapping results into a path mapping matrix PM, where each element PM[i][j] is a path description structure containing information about the physical path mapped from vertex i to j, such as the sequence of nodes traversed, the set of links used, and the allocated wavelength resources.

[0093] S3.3 Based on the path mapping matrix PM and the delay circle matrix DC, the sensitivity of graph data communication to network transmission delay is evaluated by the delay sensitivity analysis algorithm to obtain the delay sensitivity matrix DS.

[0094] Next, based on the path mapping matrix PM and the delay circle matrix DC, the system evaluates the sensitivity of graph data communication to network transmission latency using a latency sensitivity analysis algorithm. Latency sensitivity refers to the degree to which graph computation performance is affected by communication latency; different graph algorithms and application scenarios have different sensitivities to latency. Latency sensitivity analysis needs to consider the characteristics of the graph algorithm, its iteration mode, and data dependencies. For example, synchronous iterative algorithms (such as graph algorithms under the BSP model) are usually very sensitive to latency because each iteration must wait for all communication to complete; while asynchronous iterative algorithms are relatively more tolerant of latency fluctuations. Latency sensitivity is also related to communication frequency; high-frequency communication is usually more sensitive to latency. The system can combine theoretical analysis and historical performance data to assign a latency sensitivity score to the communication between each pair of vertices and organize it into a latency sensitivity matrix DS, where DS[i][j] represents the sensitivity of communication between vertices i and j to latency, with a larger value indicating higher sensitivity.

[0095] S3.4 Based on the path mapping matrix PM and the communication frequency matrix F, the bandwidth demand of graph data communication is estimated using a bandwidth demand analysis algorithm to obtain the bandwidth demand matrix BR.

[0096] The system also estimates the network bandwidth requirements for graph data communication based on the path mapping matrix PM and the communication frequency matrix F using a bandwidth demand analysis algorithm. Bandwidth demand refers to the network transmission capacity required for data exchange between nodes during graph computation. Bandwidth demand analysis needs to consider several factors: first, the communication frequency—higher frequencies result in greater data exchange per unit time; second, the message size, which depends on the amount of attribute data in the graph vertices and the characteristics of the algorithm; and third, the degree of parallelism—the number of communication tasks performed simultaneously. Bandwidth demand estimation can employ theoretical models and empirical formulas, such as... ,in It is the vertex and Average message size between These are the parallel coefficients. The system organizes the estimation results into a bandwidth requirement matrix BR, where BR[i][j] represents the vertex. and The bandwidth requirement for communication between them can be expressed in Mbps or Gbps.

[0097] S3.5 The latency sensitivity matrix DS and the bandwidth demand matrix BR are processed by the correlation matrix construction algorithm to establish the correlation matrix between the graph data logical topology and the OTN physical topology, and a two-layer topology correlation matrix TRM is obtained to represent the mapping relationship between the two layers of topology.

[0098] The system processes the latency sensitivity matrix (DS) and bandwidth requirement matrix (BR) using an association matrix construction algorithm to establish an association matrix between the logical topology of the graph data and the physical topology of the OTN. The association matrix construction aims to link the communication requirement characteristics (latency sensitivity and bandwidth requirement) at the logical level with the network resource allocation at the physical level, forming a comprehensive mapping representation. The association matrix construction can employ feature fusion methods, combining latency sensitivity and bandwidth requirement through a mathematical model (such as weighted sum or geometric mean), and then integrating it with node mapping and path mapping information. The system organizes the construction results into a two-layer topological association matrix (TRM), a multi-dimensional data structure containing mapping relationships at the vertex level and edge level, comprehensively describing the correspondence between the logical topology and the physical topology.

[0099] S3.6 For the two-layer topology association matrix TRM, the final two-layer mapping relationship matrix DLM is constructed through the mapping relationship synthesis algorithm, which is used to comprehensively describe the mapping relationship between graph data and OTN network.

[0100] Finally, the system constructs the final two-layer mapping relation matrix (DLM) for the two-layer topological association matrix (TRM) using a mapping relation synthesis algorithm. Mapping relation synthesis is the process of integrating and optimizing the mapping information from the preceding steps, aiming to form a unified and comprehensive description of the mapping relationships. The mapping relation synthesis algorithm can consider various factors, such as resource utilization efficiency, load balancing, and fault recovery capabilities, to adjust and optimize the initial mapping scheme. The system can use heuristic algorithms or iterative optimization methods to gradually improve the mapping scheme and enhance overall performance. The system organizes the synthesis results into the two-layer mapping relation matrix (DLM), which is the final description of the mapping relationships. This matrix contains the complete mapping relationships between the vertices and edges of the graph data and the nodes and links of the OTN network, providing a comprehensive reference framework for subsequent graph data deployment and optimization.

[0101] Specifically, the node mapping matrix NM is a two-dimensional matrix, where NM[i][j] indicates whether a computational node i in the graph data is mapped to a physical node j in the OTN network. The path mapping matrix PM describes how the communication needs between vertex pairs in the graph data are mapped to the physical paths in the OTN network. The latency sensitivity matrix DS is used to represent the sensitivity of different graph computation operations to network latency; for example, synchronization operations in some iterative graph algorithms are more sensitive to latency. The bandwidth requirement matrix BR represents the bandwidth requirement on each communication path, which can be estimated based on communication frequency and data size.

[0102] S4: For the two-layer mapping relationship matrix DLM, communication frequency, latency sensitivity, bandwidth requirement, wavelength resources and topology matching degree are taken as input parameters. The weights of DLM are calculated and optimized by weight definition algorithm and multi-factor fusion algorithm to obtain the communication strength weight matrix CIW.

[0103] The calculation of the communication strength weight factor in step S4 is the core optimization step of the method of this invention. Its purpose is to comprehensively consider various influencing factors and allocate reasonable weights to the mapping relationship between the communication needs of graph data and OTN network resources, thereby guiding the system to make optimal data layout decisions.

[0104] like Figure 5 As shown, S4 specifically includes:

[0105] S4.1 Perform weight definition processing on the communication frequency matrix F to assign corresponding weights to the communication frequencies between each pair of vertices in the graph data, and obtain the communication frequency weight matrix FW.

[0106] Specifically, the communication frequency matrix F is weighted to assign corresponding weight values ​​to the communication frequencies between each pair of vertices in the graph data. Communication frequency is the most fundamental consideration, directly reflecting the frequency of data exchange. Weight allocation can employ non-linear mapping functions, such as logarithmic or exponential functions, to ensure that vertices with significantly different communication frequencies also have distinct weights. The system organizes these weight values ​​into a communication frequency weight matrix FW, where FW[i][j] represents the importance weight of the communication frequency between vertices i and j.

[0107] S4.2 Based on the delay sensitivity matrix DS and the delay circle matrix DC, a delay weight calculation algorithm is used to assign corresponding weights to communications with different delay requirements, resulting in a delay weight matrix DW. The delay weight matrix DW is used to reflect the degree of impact of delay on performance.

[0108] Next, the system assigns corresponding weights to communications with different latency requirements based on the latency sensitivity matrix DS and the latency circle matrix DC, using a latency weight calculation algorithm. Latency sensitivity reflects the tolerance of graph computation tasks to communication latency, while the latency circle describes the latency distribution characteristics in the network. The latency weight calculation considers the matching degree of two factors: on the one hand, for communication needs with high latency sensitivity, links with lower latency should be given priority and assigned higher weights; on the other hand, when the latency distribution in the network is uneven, links with lower latency are scarce resources and should be used more sparingly. The system uses composite functions to calculate latency weights, for example... ,in and These are adjustment parameters; DC[m][n] represents the node. and The system calculates the latency levels between different time-delay cycles. The results are then organized into a latency weight matrix (DW), which reflects the impact of latency on system performance and is particularly important for latency-sensitive applications such as real-time analysis.

[0109] The system also assigns corresponding weights to communications with different bandwidth requirements based on the bandwidth demand matrix BR and the link parameter set LP, using a bandwidth weight calculation algorithm. Bandwidth is a crucial indicator of communication resources, directly impacting data transmission throughput and completion time. Bandwidth weight calculation considers two aspects: first, the absolute size of the bandwidth demand—the larger the demand, the higher the weight; second, the degree of matching between the bandwidth demand and the available link bandwidth—when the demand approaches or exceeds the available bandwidth, the weight should increase rapidly to reflect potential bottleneck risks. Bandwidth weight calculation can use saturation functions, such as... , among which and It's about adjusting parameters. Represents a node and The system calculates the available bandwidth of the links between them. The results are then organized into a bandwidth weight matrix BW, which reflects the impact of bandwidth on system performance and is particularly important for bandwidth-intensive applications such as big data processing.

[0110] S4.3 Based on the bandwidth demand matrix BR and the link parameter set LP, a bandwidth weight matrix BW is obtained by assigning corresponding weights to communications with different bandwidth demands through a bandwidth weight calculation algorithm. The bandwidth weight matrix BW is used to reflect the degree of influence of bandwidth on performance.

[0111] S4.4 Perform wavelength weight calculation on the wavelength resource state matrix WR, and assign corresponding weights according to the scarcity and importance of wavelength resources to obtain the wavelength resource weight matrix WW.

[0112] The system performs wavelength weight calculations on the wavelength resource state matrix (WR), assigning corresponding weights based on the scarcity and importance of wavelength resources. In WDM optical networks, wavelength is the basic unit of transmission resource, and its rational allocation directly affects network capacity and scalability. Wavelength weight calculation primarily considers the scarcity of wavelength resources, which can be measured by wavelength occupancy rate or the number of remaining available wavelengths. For example, nonlinear functions can be used. ,in and It's about adjusting parameters. This represents the proportion of available wavelengths in the link. The system organizes the calculation results into a wavelength resource weight matrix WW, which reflects the importance of wavelength resources and is particularly crucial for planning large-scale optical network deployments.

[0113] S4.5 For the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, the structural similarity weights of the logical topology and the physical topology are calculated by the topology matching algorithm to obtain the topology matching weight matrix TW.

[0114] The system also calculates the structural similarity weights between the logical and physical topologies using a topology matching algorithm, targeting the graph data communication feature matrix (GCF) and the OTN physical network feature matrix (ONF). Topology matching degree is an indicator of the degree of fit between the logical structure of the graph data and the physical structure of the network. A high matching degree means that the logical communication pattern and the physical network topology are highly compatible, which helps reduce communication overhead. The topology matching algorithm can employ graph similarity calculation methods, such as graph edit distance, spectral similarity, or subgraph isomorphism ratio. For example, it can calculate the similarity of vertex degree distribution, clustering coefficient distribution, and path length distribution, and then synthesize them to obtain the overall topology matching degree. The system organizes the calculation results into a topology matching weight matrix TW, where TW[i][j] represents the topology matching weight that maps the logical connection between vertices i and j to the corresponding physical path.

[0115] S4.6 Based on the communication frequency weight matrix FW, the delay weight matrix DW, the bandwidth weight matrix BW, the wavelength resource weight matrix WW, and the topology matching weight matrix TW, the final communication strength weight factor is calculated through a multi-factor fusion algorithm to obtain the communication strength weight matrix CIW.

[0116] Finally, based on the communication frequency weight matrix FW, delay weight matrix DW, bandwidth weight matrix BW, wavelength resource weight matrix WW, and topology matching weight matrix TW, the system calculates the final communication strength weight factor using a multi-factor fusion algorithm. Multi-factor fusion is a crucial step, requiring a balance of the importance of various factors, considering both application requirements and network resource conditions. Fusion methods can range from simple weighted summation to more complex nonlinear combinations or analytic hierarchy process (AHP). For example, a weighted summation formula can be used:

[0117] ,in to These are the weighting coefficients of each factor, which can be dynamically adjusted according to the application scenario. The system organizes the fusion results into a Communication Strength Weighting Matrix (CIW). This matrix comprehensively reflects the strength weights of graph data communication needs mapped onto the OTN network and serves as the core basis for subsequent optimization decisions.

[0118] Specifically, multi-factor fusion algorithms can employ a weighted summation method:

[0119] ,in , , , , Let be the weight coefficients of each factor, and These weighting coefficients can be adjusted according to specific application scenarios and optimization objectives. For example, in applications with high real-time requirements, the value of β can be increased; when network resources are limited, the value of δ can be increased.

[0120] S5: Perform a scheme enumeration algorithm on the communication strength weight matrix CIW to generate candidate layout schemes and evaluate communication overhead, resource utilization and load balancing. Calculate the comprehensive score of topology matching degree and determine the optimized data layout scheme OLP.

[0121] like Figure 6 As shown, S5 includes:

[0122] S5.1 Based on the graph data scale and OTN network resource status, the communication strength weight matrix CIW is processed through a scheme enumeration algorithm to generate a variety of possible graph data layout schemes.

[0123] These layout schemes include different vertex partitioning strategies and node mapping schemes, reflecting the various possible distributions of graph data on the physical network. The algorithm considers the graph's size, connectivity characteristics, and the physical network's capacity constraints to ensure that the generated schemes meet both computational requirements and network resource constraints, ultimately forming a candidate layout scheme set CLS.

[0124] S5.2 For the candidate layout scheme set CLS and the communication strength weight matrix CIW, the communication overhead evaluation algorithm is applied to calculate the total communication overhead under each layout scheme.

[0125] This step simulates the data exchange patterns during graph computation, quantifying the communication efficiency of different layout schemes based on the communication strength between vertices and the cost of physical transmission paths. The calculation results are organized into a Communication Overhead Evaluation Result Matrix (CEC), providing a crucial basis for subsequent comprehensive scoring.

[0126] Based on the candidate layout scheme set CLS and the OTN physical network feature matrix ONF, S5.3 uses a resource utilization evaluation algorithm to analyze the utilization of network resources for each layout scheme.

[0127] The algorithm calculates the utilization efficiency of link bandwidth, wavelength resources, and node processing capabilities, identifying potential resource bottlenecks or waste. The evaluation results are recorded in the Resource Utilization Evaluator (RUE) matrix, reflecting the resource utilization efficiency of each scheme.

[0128] S5.4 evaluates the load balancing of the candidate layout scheme set CLS and the two-layer mapping relationship matrix DLM, and calculates the distribution balance of computational load and network load under each layout scheme.

[0129] Balanced load distribution can avoid local hotspots and improve the overall stability and scalability of the system. Evaluation uses statistical indicators such as standard deviation or Gini coefficient, and the results are recorded in the Load Balance Evaluation Result Matrix (LBE).

[0130] S5.5 Based on the communication overhead assessment result matrix CEC, the resource utilization assessment result matrix RUE, and the load balancing assessment result matrix LBE, a comprehensive scoring algorithm is used to calculate the topology matching score for each layout scheme.

[0131] This process assigns appropriate weights to different evaluation dimensions, balancing the needs of communication efficiency, resource utilization, and load balancing, and generating a comprehensive index that fully reflects the quality of the solution. The scoring results are organized into a Topology Matching Score Matrix (TMS), providing a quantitative basis for the final decision.

[0132] S5.6 applies the optimal solution selection algorithm to the topology matching score matrix TMS to identify the data layout scheme with the highest comprehensive score from all candidate schemes.

[0133] This scheme achieves the best balance in terms of communication efficiency, resource utilization, and load balancing, and has been identified as the optimized data layout scheme OLP to guide subsequent graph data deployment and computation task scheduling.

[0134] Specifically, communication overhead assessment can calculate the sum of communication costs between all vertex pairs:

[0135] Total communication overhead = ∑∑CIW[i][j] × D[i][j], where D[i][j] is the physical network distance or latency between vertices i and j under a specific layout scheme. Resource utilization assessment can consider the efficiency of network link and wavelength resource usage; for example, calculating the standard deviation of link utilization, a smaller value indicates more balanced resource utilization. Load balancing assessment can use indicators such as the Jain fairness index to measure the degree of load distribution balance. A weighted method can be used for comprehensive topology matching score.

[0136] Overall score = w1 × (1 / communication overhead) + w2 × resource utilization + w3 × load balancing, where w1, w2, and w3 are weighting coefficients.

[0137] S6: Based on the optimized data layout scheme OLP, the mapping relationship is dynamically adjusted by using an incremental update algorithm through real-time monitoring of system performance indicators and network status changes, resulting in a dynamically optimized mapping relationship DOM.

[0138] Based on the optimized data layout scheme OLP, the system first deploys a distributed monitoring agent to comprehensively collect key performance indicators such as CPU utilization, memory usage, network throughput, and disk I / O of each computing node. This real-time data, after standardization and structuring, forms the System Performance Index (SPI) matrix, providing foundational data support for subsequent anomaly detection and decision-making adjustments.

[0139] Meanwhile, based on the optimized data layout scheme and OTN network environment, the system continuously monitors the operational status of the optical links through network state awareness algorithms. This process includes measuring bandwidth utilization, recording latency fluctuations, detecting changes in bit error rate, and tracking wavelength resource occupancy. This dynamic data at the network layer is integrated into the Network State Change Matrix (NSC), reflecting the real-time status of the physical transmission environment.

[0140] Based on the collected system performance index matrix (SPI) and network state change matrix (NSC), the system applies a specially designed anomaly detection algorithm for multidimensional analysis. This algorithm can identify various potential problems, such as performance bottlenecks of computing nodes, network link congestion, and uneven resource allocation. For each anomaly, the system also quantitatively calculates the degree of deviation from the optimal operating state, generating an anomaly assessment matrix (ASE) to provide precise quantitative basis for subsequent decision-making.

[0141] The system implements an incremental triggering mechanism to control adjustment behavior based on the anomaly assessment matrix ASE and pre-set threshold parameters. This mechanism follows the principle of "no adjustment unless necessary," triggering dynamic adjustments to the mapping relationship only when performance deviations exceed the preset threshold, thus avoiding unnecessary system fluctuations. The result of the triggering decision is output in the form of an adjustment decision signal ADS, clearly indicating whether adjustment is needed and the urgency of the adjustment.

[0142] When the system determines that adjustments are needed, it incrementally updates the adjustment decision signal (ADS) and the current mapping state. This process first accurately calculates the subset of mappings that need adjustment, then employs a local remapping strategy, recalculating and optimizing only the affected regions while keeping other regions unchanged. This method significantly reduces adjustment overhead, and the resulting incremental mapping adjustment scheme (MIAS) addresses the current problem while minimizing disruption to system operation.

[0143] Finally, the system implements adjustments in stages using a progressive execution algorithm, based on the incremental adjustment scheme MIAS for mapping relationships and the real-time monitored system status. This progressive approach breaks down large-scale adjustments into a series of small steps, verifying the system status after each step to ensure the safety and controllability of the adjustment process. Through this carefully designed dynamic optimization process, the system ultimately obtains a dynamically optimized mapping relationship DOM, achieving continuous optimal matching between graph data and network resources.

[0144] Specifically, the core idea of ​​incremental update algorithms is to adjust only those mapping relationships that are significantly affected by anomalies, while keeping other parts unchanged. This local adjustment method can significantly reduce adjustment overhead and shorten system response time. For example, when an optical link is detected to be overloaded, the system only needs to readjust the graph data partitions using that link, without needing to recalculate the entire mapping scheme. Incremental triggering mechanisms typically set multiple threshold levels, corresponding to different degrees of adjustment behavior, ranging from slight adjustments to complete remapping.

[0145] In practical applications, the graph data communication intensity mapping method based on optical transport network topology of this invention can significantly improve the performance and resource utilization efficiency of distributed graph computing systems. The working principle and technical advantages of this method will be described in detail below:

[0146] First, this invention constructs a comprehensive graph data communication feature matrix by extracting topological and communication pattern features from graph data. Unlike traditional methods, this invention not only considers the direct connections between vertices but also introduces higher-order topological features such as degree centrality and clustering coefficients, enabling more accurate prediction of communication patterns and data flow during graph computation. For example, in social network analysis applications, highly central vertices (such as opinion leaders) often need to communicate frequently with multiple other vertices; therefore, they should be prioritized for placement on nodes with better network connectivity.

[0147] Secondly, this invention collects and analyzes detailed characteristics of the OTN physical network, including network topology, link performance parameters, and wavelength resource status. In particular, it introduces the concept of multi-level delay circles, which can more accurately describe the delay distribution characteristics between network nodes, providing better support for delay-sensitive graph computing applications. In practical deployments, through in-depth analysis of the OTN network, the system can identify potential network bottlenecks and high-performance paths, providing more reliable network environment awareness for the distributed deployment of graph data.

[0148] Third, this invention establishes a two-layer mapping model between the logical topology of graph data and the physical topology of OTN, achieving cross-layer optimization. This two-layer mapping considers not only node-level mapping relationships but also communication path mapping, enabling the distribution of graph data to fully adapt to the characteristics of the underlying network. For example, for frequently communicating graph data vertex pairs, the system will try to map them to physical nodes that are closer in network distance, reducing communication overhead.

[0149] Fourth, the communication strength weighting factor calculation method proposed in this invention comprehensively considers multiple influencing factors, including communication frequency, latency sensitivity, bandwidth requirements, wavelength resources, and topology matching degree. This multi-factor fusion method enables the system to flexibly adjust the weights of each factor according to the characteristics and needs of different applications, achieving more refined optimization. For example, for real-time analysis applications, the weight of latency sensitivity can be increased; while for large-scale data processing applications, more emphasis can be placed on bandwidth requirements and resource utilization.

[0150] Fifth, this invention designs a complete topology matching evaluation and optimization process. By generating candidate layout schemes and performing multi-dimensional evaluations, the optimal data layout scheme is selected. Compared with traditional methods, this invention not only considers communication overhead but also evaluates resource utilization and load balancing, making the optimization results more comprehensive and practical. This multi-objective optimization method can reduce communication overhead while ensuring efficient utilization of network resources and balanced load distribution.

[0151] Finally, this invention implements a dynamic adjustment and feedback mechanism, enabling real-time adjustments to mapping relationships based on changes in system operating status and network environment. This incremental update strategy avoids the high overhead of global remapping, allowing the system to adapt to environmental changes at minimal cost and maintain optimal performance. For example, when network link congestion or excessive load on a computing node is detected, the system automatically adjusts the relevant mapping relationships, transferring some data or computing tasks to other available resources to ensure the continuous and efficient operation of the system.

[0152] Example 2

[0153] After constructing the two-layer topological association matrix, the method further includes:

[0154] A1. Perform communication strength data stream identification processing on the path mapping matrix PM to identify the critical data transmission paths that require enhanced reliability, and obtain the critical path set KPS;

[0155] The purpose of performing communication strength data flow identification processing on the path mapping matrix PM is to filter out those paths that have a critical impact on system performance and data integrity from numerous data transmission paths.

[0156] This process first analyzes the communication frequency, data importance, and service priority of each path. Then, based on preset thresholds or sorting methods, it identifies the set of critical paths (KPS) that require special protection. These critical paths typically carry the most frequent or most important data exchanges during graph computation, and their transmission quality directly affects the accuracy of the overall computation results and system performance.

[0157] A2. Based on the critical path set KPS and the delay sensitivity matrix DS, the optimal error correction coding parameters are determined for each critical path in the critical path set KPS using the Reed-Solomon error correction parameter calculation algorithm, resulting in the error correction parameter configuration matrix ECC.

[0158] After determining the set of critical paths, the system combines the information from the delay sensitivity matrix DS and uses the Reed-Solomon error correction parameter calculation algorithm to customize the optimal error correction coding parameters for each critical path.

[0159] Reed-Solomon codes are powerful forward error correction codes capable of detecting and correcting transmission errors without requiring retransmissions. The algorithm weighs the latency sensitivity, bit error rate, and data importance of each path, calculating the optimal redundancy ratio and block size for each path to ensure sufficient error correction capability while minimizing the impact of coding overhead on transmission latency. These parameters are organized into an Error Correction Parameter Configuration Matrix (ECC), providing precise guidance for subsequent coding processes.

[0160] For the critical path set KPS and the error correction parameter configuration matrix ECC, the system implements Reed-Solomon coding on the high communication intensity data stream through a redundancy coding processing algorithm.

[0161] A3. For the critical path set KPS and the error correction parameter configuration matrix ECC, Reed-Solomon encoding is performed on the high communication intensity data stream through a redundancy coding processing algorithm to obtain the reliability-enhanced path mapping matrix ERPM;

[0162] The encoding process adds an appropriate amount of redundant check information to each data packet according to the parameters specified in ECC, enabling the receiver to detect and correct errors that may occur during transmission. This forward error correction mechanism significantly improves the reliability of data transmission and reduces the number of retransmissions, making it particularly suitable for protecting critical data streams in highly volatile network environments. The encoded path information is updated in the reliability-enhanced path mapping matrix ERPM, reflecting the new transmission characteristics after the addition of the error correction mechanism.

[0163] A4. The reliability-enhanced path mapping matrix ERPM is processed by the association matrix update algorithm to correct the two-layer topological association matrix, resulting in the enhanced two-layer topological association matrix ETRM.

[0164] Finally, the system applies an association matrix update algorithm to the reliability-enhanced path mapping matrix (ERPM) to correct the two-layer topology association matrix. This step integrates the changes brought about by error correction coding (such as increased bandwidth requirements and altered latency characteristics) into the original mapping relationship, ensuring that the system's planning and scheduling of network resources accurately reflects the impact of the error correction mechanism. The updated enhanced two-layer topology association matrix (ETRM) provides a more comprehensive and accurate description of the mapping relationship, providing a reliable basis for subsequent optimization decisions, and enabling the system to maximize network resource utilization efficiency while ensuring transmission reliability.

[0165] Specifically, critical data transmission paths can be identified by setting communication strength thresholds; for example, paths with communication strength weights in the top 20% can be defined as critical paths. Reed-Solomon error correction coding is a commonly used forward error correction code that adds redundant information to enable the receiver to detect and correct errors during transmission. The selection of error correction parameters requires a trade-off between redundancy overhead and error correction capability, typically considering factors such as the link's bit error rate, data importance, and latency sensitivity. For example, for paths with high latency sensitivity, lower redundancy coding parameters can be selected to reduce encoding and decoding latency; while for critical data, stronger error correction capability can be chosen, even if it increases transmission overhead.

[0166] It should be noted that Example 2 mainly involves enhancing the reliability of critical data transmission paths through Reed-Solomon error correction coding technology after constructing a two-layer topological association matrix. The core of this example lies in identifying high-intensity data streams and configuring appropriate error correction protection mechanisms for them.

[0167] Reed-Solomon codes are a class of linear block error-correcting codes based on finite fields (Galois fields), proposed by Irving Reed and Gustave Solomon in 1960. This coding system is a forward error-correcting (FEC) code with strong error-correcting capabilities, able to detect and correct multiple errors occurring during transmission. Reed-Solomon codes belong to the Maximum Distance Separable (MDS) category, meaning they can provide the theoretically maximum error-correcting capability for a given amount of redundancy.

[0168] In graph data communication systems, Reed-Solomon coding adds redundant checksum information to the original data, enabling the receiver to automatically detect and repair transmission errors without requiring retransmission. This feature is particularly important for high-speed data transmission in optical transport networks, as retransmission mechanisms increase communication latency and impact the real-time performance of graph computation tasks.

[0169] The technical principle of Reed-Solomon coding is based on polynomial operations and finite field theory. Assuming the original data consists of k symbols, encoding generates n symbols (n>k), where the added (nk) symbols are redundancy check symbols. The encoding process can be represented as a polynomial evaluation problem: using the k original data symbols as coefficients of a polynomial, construct a (k-1)th degree polynomial, and then evaluate it at n distinct points to obtain n symbols.

[0170] Specifically, let the original data be (d0, d1, ..., d...). k-1 Construct polynomial Choose n distinct non-zero elements α1, α2, ..., α3 as evaluation points, and calculate... The encoded symbols are obtained. Since any k distinct point values ​​can uniquely determine a (k-1)th degree polynomial, the original data can be recovered by receiving any k correct symbols.

[0171] In terms of error detection and correction, Reed-Solomon codes can detect up to (nk) errors and correct up to (nk) / 2 errors. This error correction capability gives it excellent performance in noisy environments. For graph data communication systems, this means that even if some data packets are corrupted during transmission, the system can automatically recover the correct data, avoiding the latency overhead of retransmission.

[0172] Specifically, in this embodiment, the first step is critical path identification. The system performs in-depth analysis of the path mapping matrix PM to identify transmission paths carrying high-intensity data streams. This process requires comprehensive consideration of multiple factors such as communication frequency, data importance, and service priority. Specifically, the system calculates a comprehensive importance score for each path using the formula: Importance Score = w1 × Communication Frequency + w2 × Data Importance + w3 × Service Priority, where w1, w2, and w3 are weighting coefficients. Then, based on a preset threshold or sorting method, the paths with the highest scores are selected to form the critical path set KPS. For example, in social network analysis applications, communication paths connecting core nodes typically have high importance scores.

[0173] The second step is error correction parameter optimization. For each path in the critical path set KPS, the system combines the information from the delay sensitivity matrix DS and uses a specialized parameter calculation algorithm to determine the optimal Reed-Solomon coding parameters. This process requires finding the best balance between error correction capability and coding overhead. Parameter optimization mainly considers three factors: code length n (total number of encoded symbols), information length k (number of original data symbols), and error correction capability t (maximum number of correctable errors). The optimization objective function can be expressed as: minimize (coding overhead) subject to (error correction capability ≥ minimum requirement, delay increase ≤ threshold). The system calculates the most suitable parameter combination based on the delay sensitivity, historical bit error rate, and importance level of each path.

[0174] The third step is the encoding implementation. Based on the optimized encoding parameters, the system performs Reed-Solomon encoding on the data flow along the critical path. The encoding process first divides the original data into fixed-length data blocks, each containing k symbols. Then, a corresponding generator polynomial is constructed for each data block, and (nk) check symbols are generated by evaluating the polynomial. These check symbols are then appended to the original data to form a complete encoded block. Since real-time requirements need to be considered during encoding implementation, the system adopts a parallel encoding strategy, allowing multiple data blocks to be encoded simultaneously to reduce encoding latency.

[0175] The fourth step is matrix updating. After encoding, the system needs to update the relevant mapping matrices to reflect the changes brought about by encoding. This includes increased bandwidth requirements (due to the addition of redundant information), changes in latency characteristics (due to encoding and decoding processing time), and improved transmission reliability. The system integrates these changes into the original two-layer topology association matrix using an association matrix update algorithm, generating an enhanced two-layer topology association matrix (ETRM). The update process uses an incremental method, modifying only the affected matrix elements to avoid recalculating the entire matrix.

[0176] In practical applications, Reed-Solomon error-correcting coding has proven highly effective. For example, when processing large-scale social network graph data, critical user relationship data transmission paths are configured with RS(255,239) encoding (which corrects 8 symbol errors). This improves data transmission success rate from 92% to 99.8% under poor network conditions, reduces retransmissions by 85%, and enhances overall computational performance by 23%. This improvement is particularly important for graph analysis tasks with high real-time requirements, ensuring the timeliness and accuracy of computational results.

[0177] Example 3

[0178] Before generating the optimized data layout scheme OLP, the method further includes:

[0179] B1. Perform tile segmentation processing on the candidate layout scheme set CLS, dividing the graph data into multiple tiles, each tile containing a subset of associated vertices and edges, to obtain the tile set TS;

[0180] Tile segmentation is a key optimization strategy for the candidate layout scheme set CLS, dividing large-scale graph data into multiple manageable sub-tiles. Each tile contains a subset of vertices and edges from the original graph, which typically exhibit strong internal correlations. The core objective of tile segmentation is to minimize the number of cross-tile edges while maintaining a relatively balanced tile size, thereby reducing communication overhead during processing. The system achieves high-quality tile partitioning by applying algorithms such as community detection or spectral clustering, ultimately forming the tile set TS.

[0181] B2. Based on the tile set TS and the OTN physical network feature matrix ONF, design an overlapping execution strategy for the data loading and calculation process through a workload interleaving scheduling algorithm, so that the data preloading of the next tile and the calculation of the current tile can be performed in parallel, and obtain the interleaving scheduling scheme ISS;

[0182] Based on the generated tile set TS and the OTN physical network feature matrix ONF, the system employs a workload-interleaved scheduling algorithm to design an efficient execution strategy. The core idea of ​​this strategy is to overlap data loading and computation processes; that is, while processing the current tile, the system pre-loads the data for the next tile. This parallel mechanism effectively hides data transmission latency and significantly improves the overall system throughput. The workload-interleaved scheduling algorithm considers various factors such as network bandwidth, storage I / O capacity, and computing resources to generate the optimal interleaved scheduling scheme ISS.

[0183] B3. Based on the interleaved scheduling scheme ISS and the system resource status, the resource scheduling strategy RSS is obtained by intelligently allocating computing resources between different graph computing stages through a cross-stage kernel scheduling algorithm.

[0184] Based on the interleaved scheduling scheme ISS and the current system resource status, the cross-stage kernel scheduling algorithm achieves intelligent resource allocation between different processing stages of graph computation. Graph computation typically includes multiple stages such as data loading, preprocessing, core computation, and result output, each with different requirements for computing and memory resources. The cross-stage kernel scheduling algorithm can dynamically adjust the allocation ratio of CPU cores, memory, and I / O bandwidth according to the resource requirements of each stage, ensuring maximum resource utilization and generating an efficient resource scheduling strategy RSS.

[0185] B4. Based on the tile set TS, the interleaved scheduling scheme ISS, and the resource scheduling strategy RSS, the performance index of each layout scheme under the interleaved workload condition is calculated by a comprehensive optimization algorithm to obtain the optimized topology matching degree score matrix OTMS;

[0186] The system comprehensively evaluates each candidate layout scheme based on the tile set TS, the interleaved scheduling scheme ISS, and the resource scheduling policy RSS through a comprehensive optimization algorithm. This process simulates the system's operating state under interleaved workload conditions, calculating key performance indicators such as total execution time, resource utilization, load balancing, and energy efficiency. The comprehensive optimization algorithm considers the interrelationships and trade-offs of these indicators, generating an optimized topology matching score matrix OTMS that reflects the actual performance of each layout scheme.

[0187] B5. The optimized topology matching score matrix OTMS is processed by the optimal solution selection algorithm to select the data layout scheme with the best performance under the workload interleaving condition, and the optimized data layout scheme OLP is obtained.

[0188] Finally, the system applies an optimal solution selection algorithm to the optimized Topology Matching Score Matrix (OTMS) to identify the data layout scheme with the best performance under interleaved workload conditions from all candidate schemes. The selection process considers not only absolute performance scores but also the stability and adaptability of the schemes, ensuring that the selected scheme maintains good performance under various workload conditions. Through this series of meticulous optimization steps, the system ultimately determines the optimal data layout scheme OLP, providing the best execution environment for graph computation tasks.

[0189] Specifically, tile partitioning is a technique for dividing large-scale graph data into multiple manageable subgraphs, each tile containing a subset of the original graph's vertices and edges. Tile partitioning needs to consider the connectivity between vertices, minimizing the number of edges across tiles while maintaining a relatively balanced tile size. Workload staggered scheduling is an optimization technique that hides data transmission latency by overlapping data loading and computation processes. For example, while the system is computing the data for the current tile, it can simultaneously begin preloading the data for the next tile. This way, once the current tile's computation is complete, the computation for the next tile can begin immediately, reducing waiting time. Cross-stage kernel scheduling intelligently allocates computing resources between different stages of graph computation (such as data loading, computation, and result output) to maximize resource utilization.

[0190] Furthermore, Example 3 primarily involves further optimizing system performance before generating the optimized data layout scheme OLP through techniques such as tile partitioning, workload staggered scheduling, and cross-stage kernel scheduling. The core idea of ​​this example is to decompose large-scale graph data processing tasks into multiple parallel executable subtasks and maximize resource utilization efficiency through intelligent scheduling strategies.

[0191] Tile segmentation is a fundamental step in Example 3, aiming to divide large-scale graph data into multiple appropriately sized sub-tiles. Segmentation algorithms are typically based on graph structural characteristics such as community structure, connectivity, and load balancing. The system employs an improved METIS algorithm or a spectral clustering-based segmentation method to ensure tight internal connections within each tile while minimizing the number of cross-edges between tiles. During segmentation, the algorithm calculates tile cohesion and inter-tile coupling. This segmentation strategy minimizes communication overhead between tiles while maintaining a relatively balanced processing load across all tiles.

[0192] Implementation principle of workload staggered scheduling algorithm

[0193] The workload staggered scheduling algorithm is an innovative task scheduling strategy. Its core idea is to overlap the execution of different stages, such as data loading, data preprocessing, and graph computation, in time. Based on a pipelined processing model, this algorithm decomposes the entire graph computation process into multiple parallelizable stages. Specifically, when the system is processing the current graph tile... When performing a computational task, the next tile can be processed simultaneously. Data preloading and the previous tile The result output operation.

[0194] The algorithm employs a multi-threaded concurrent model. The processing flow for each tile includes four main stages: data loading, preprocessing, core computation, and output. The scheduling algorithm uses precise time window planning to ensure that these stages overlap to the greatest extent possible. The formula for calculating the time window is: Each item represents the execution time of a different stage.

[0195] The key to staggered workload scheduling lies in accurately predicting the execution time of each stage and dynamically adjusting the scheduling strategy based on network bandwidth and I / O performance. The system maintains an execution time prediction model, which uses linear regression or neural network methods to predict the execution time of each stage based on historical execution data and the current system state. Input features of the prediction model include tile size, edge density, available network bandwidth, disk I / O rate, and CPU utilization. Based on the prediction results, the scheduler dynamically adjusts the start time of each stage to ensure efficient pipeline operation.

[0196] The cross-stage kernel scheduling algorithm is responsible for intelligently allocating computing resources, especially critical resources such as CPU cores, memory, and I / O bandwidth, across different processing stages of graph computation. This algorithm employs a dynamic resource allocation strategy, optimizing resource configuration based on the real-time resource requirements and system resource availability at each stage.

[0197] The core of the algorithm consists of two modules: resource demand prediction and allocation optimization. The resource demand prediction module establishes a resource demand model by analyzing the characteristics of different graph computation stages. For example, the data loading stage mainly consumes I / O bandwidth and memory, the computation stage mainly consumes CPU and memory, and the result output stage mainly consumes I / O bandwidth. The prediction model can be represented as: ,in This represents a vector of resource requirements at a specific stage. Representing patch features, Indicates the system status.

[0198] The allocation optimization module employs a multi-objective optimization method to maximize overall resource utilization and processing throughput while meeting the minimum resource requirements at each stage. The optimization objective function is: ,in Indicates time resource utilization rate Indicates time stage The throughput is given by |R| and |S|, which represent the number of resource types and the number of processing stages, respectively.

[0199] The cross-stage kernel scheduling algorithm also implements a load-aware dynamic adjustment mechanism. When the system detects a resource bottleneck in a certain stage, the scheduler will automatically borrow idle resources from other stages or adjust task priorities. For example, when the CPU utilization in the computation stage reaches 95% while the I / O utilization is only 30%, the system can reallocate some I / O cores to computation tasks or start the data preloading of the next tile in advance to balance the load.

[0200] The comprehensive optimization algorithm is the core component of Example 3. It is responsible for integrating the results of tile segmentation, workload interleaving scheduling, and cross-stage kernel scheduling to generate the optimal data layout scheme. This algorithm adopts a multi-level optimization strategy, gradually expanding from local optimization to global optimization.

[0201] The first layer of the algorithm is intra-tile optimization, which optimizes the internal data layout and computational strategy for each tile based on its characteristics and allocated resources. Optimization metrics include intra-tile communication overhead, memory access efficiency, and computational load balancing. Intra-tile optimization employs heuristic algorithms, such as simulated annealing or genetic algorithms, to search for the optimal vertex allocation scheme.

[0202] The second layer is inter-tile coordination optimization, which considers the data dependencies and communication requirements between tiles, optimizing the execution order and data exchange strategy. The goal of coordination optimization is to minimize communication latency and data transmission volume between tiles. The algorithm constructs a tile dependency graph and uses topological sorting and critical path analysis to determine the optimal execution sequence. The formula for calculating the communication overhead between tiles is: ,in Indicates connecting blocks and The set of edges, Represents edge weight, Representing a block and Distance in a physical network.

[0203] The third layer is global resource optimization, which comprehensively considers the resource requirements of all graph tiles and network environment characteristics to generate a globally optimal resource allocation strategy. Global optimization employs integer linear programming (ILP) or constraint satisfaction problem (CSP) methods, with objective functions including multiple metrics such as overall execution time, resource utilization, and load balancing. Since the global optimization problem has NP-hard complexity, the algorithm uses branch-and-bound or Lagrange relaxation methods to find approximate optimal solutions.

[0204] The comprehensive optimization algorithm also implements an adaptive adjustment mechanism, which can dynamically correct the optimization strategy based on runtime performance feedback. The system maintains a performance monitoring module to collect various performance metrics in real time, such as execution time, resource utilization, and communication overhead. When actual performance differs significantly from expectations, the algorithm triggers a re-optimization process, adjusting relevant parameters and strategies. The triggering conditions for adaptive adjustment are: ,in and These represent actual performance and predicted performance, respectively. It is a preset threshold.

[0205] The three core algorithms work together through carefully designed interfaces and coordination mechanisms. The workload staggered scheduling algorithm provides a task scheduling schedule for the cross-stage kernel scheduling algorithm, which dynamically allocates computing resources according to the schedule. The comprehensive optimization algorithm coordinates the work of the two algorithms from a global perspective, ensuring that local optimization decisions are consistent with the global goal.

[0206] Information exchange between algorithms is achieved through shared data structures, including task queues, resource status tables, and performance monitoring data. The system employs an event-driven coordination mechanism; when the state of one algorithm undergoes a significant change, it notifies other related algorithms to make corresponding adjustments. For example, when the workload staggered scheduling algorithm detects that the processing time of a certain tile exceeds expectations, it immediately notifies the cross-stage kernel scheduling algorithm to increase the CPU allocation for that tile, and simultaneously notifies the synthesis optimization algorithm to update the global scheduling policy.

[0207] Through this multi-level, multi-algorithm collaborative optimization strategy, Example 3 can significantly improve the performance and efficiency of large-scale graph data processing. In actual tests, this method can achieve a 30-50% performance improvement compared to the traditional sequential processing method, while maintaining good resource utilization and load balancing.

[0208] Example 4

[0209] In constructing the two-layer mapping matrix, the method further includes:

[0210] C1. Perform continuous relaxation processing on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, relax the discrete 0-1 mapping variables into continuous variables with values ​​in the range [0,1], and obtain the continuous mapping probability matrix CPM;

[0211] Applying continuous relaxation to the graph data communication feature matrix (GCF) and the OTN physical network feature matrix (ONF) is an efficient strategy for transforming a discrete optimization problem into a continuous optimization problem. This process converts the binary mapping variables, which could only take the values ​​0 or 1, into continuous variables that can take any value in the interval [0,1], thus obtaining the continuous mapping probability matrix (CPM). This relaxation allows the application of continuous optimization algorithms such as gradient descent, greatly expanding the searchable solution space.

[0212] C2. For the continuous mapping probability matrix CPM, a potential function that can guide the optimization solution back to the discrete binary configuration is constructed through a potential function design algorithm, and a potential function parameter set PFP is obtained, wherein the potential function is used to take a relatively low value when the continuous variable is close to 0 or 1.

[0213] For the continuous mapping probability matrix CPM, the system constructs special potential functions through a potential function design algorithm. These functions can gradually guide the solution back to the discrete binary configuration during the optimization process. The potential function parameter set PFP defines the specific form and characteristics of these functions. The key feature of the potential functions is that they take relatively low values ​​when the continuous variable is close to 0 or 1, and higher values ​​when the value is in the middle (such as around 0.5). This design "penalizes" the intermediate values ​​during the optimization process, prompting the variable to move closer to 0 or 1. C3. Based on the continuous mapping probability matrix CPM and the potential function parameter set PFP, the optimal mapping solution is searched in the continuous space using a gradient descent-like algorithm to avoid getting trapped in local optima and obtain the continuous optimization solution COS.

[0214] C4. Based on the continuous optimization solution COS and the potential function parameter set PFP, the continuous solution is converted into a discrete binary mapping relationship through a discretization algorithm to obtain the discrete mapping relationship matrix DMR;

[0215] Based on the continuous mapping probability matrix CPM and the potential function parameter set PFP, the system employs gradient descent-type algorithms to search in the continuous space. These algorithms include variants such as standard gradient descent, stochastic gradient descent, Adam, or L-BFGS, and are capable of effectively finding optimal solutions in high-dimensional spaces. Because continuous space is smoother than discrete space, the algorithm can more easily overcome the "energy barrier" of local optima, finding the global optimum or a better local optimum, ultimately obtaining the continuous optimal solution COS.

[0216] Based on the continuous optimization solution COS and the potential function parameter set PFP, the system transforms the continuous solution back into a discrete binary mapping relation through a discretization algorithm. This step may employ a simple thresholding method (e.g., taking 1 for values ​​greater than 0.5 and 0 for others), or it may use more complex probabilistic or deterministic rounding strategies to ensure that the transformed solution satisfies specific constraints. This process outputs a discrete mapping relation matrix DMR, where each element is an exact 0 or 1, representing a definite mapping relation.

[0217] C5. Perform mapping relationship synthesis processing on the discrete mapping relationship matrix DMR to construct the two-layer mapping relationship matrix DLM.

[0218] Finally, the system performs comprehensive mapping relationship processing on the Discrete Mapping Matrix (DMR), integrating various constraints and optimization objectives to construct the final two-layer mapping relationship matrix (DLM). This matrix comprehensively describes the mapping relationship between graph data and the OTN network across multiple dimensions, including nodes, links, and resources, providing a solid foundation for subsequent communication strength weight calculations and data layout optimization.

[0219] Specifically, continuous relaxation is a technique that transforms a discrete optimization problem into a continuous optimization problem, allowing the use of more efficient continuous optimization algorithms. For example, a binary variable x∈{0,1} is relaxed to a continuous variable x∈[0,1], representing the probability of a mapping. The role of the potential function is to guide the continuous solution towards the discrete solution. A typical potential function can be f(x)=λx(1-x), where λ is a positive constant. When x is close to 0 or 1, the function value is close to 0; when x is close to 0.5, the function value reaches its maximum. By adding this potential function to the objective function, the optimization process tends to produce solutions close to 0 or 1, making it easier to transform into the final discrete solution. Gradient descent algorithms such as L-BFGS or Adam can efficiently search for optimal solutions in continuous space. Finally, discretization algorithms can simply set continuous values ​​greater than 0.5 to 1 and values ​​less than or equal to 0.5 to 0, or use more complex rounding strategies.

[0220] Example 5

[0221] Before obtaining the Communication Strength Weighting Matrix (CIW), the method further includes:

[0222] D1. For the communication frequency weight matrix FW, the delay weight matrix DW, the bandwidth weight matrix BW, the wavelength resource weight matrix WW, and the topology matching weight matrix TW, the communication strength weight calculation objective is expressed as the difference form of two convex functions through the convex function decomposition algorithm, resulting in convex function pairs CF1 and CF2.

[0223] For the communication frequency weight matrix FW, delay weight matrix DW, bandwidth weight matrix BW, wavelength resource weight matrix WW, and topology matching weight matrix TW, the system first applies a convex function decomposition algorithm to represent the complex optimization objective of calculating communication strength weights as the difference between two convex functions, thus obtaining the convex function pair CF1 and CF2. This decomposition method transforms what might otherwise be a non-convex optimization problem into a more manageable form, laying the foundation for efficient subsequent solutions.

[0224] D2. Perform differential convex optimization algorithm on the convex function pairs CF1 and CF2, and design an iterative solution framework so that each iteration only performs a single matrix-vector multiplication operation to obtain the iterative algorithm parameter set IAP;

[0225] Next, the system applies the differential convex optimization algorithm to the convex function pairs CF1 and CF2. The key innovation of this step lies in the design of a special iterative solution framework, which requires only a single matrix-vector multiplication operation in each iteration, significantly reducing computational complexity. This efficient computational structure is particularly suitable for processing large-scale graph data scenarios. The system organizes relevant parameters into an iterative algorithm parameter set IAP, including key control parameters such as step size, convergence threshold, and maximum number of iterations.

[0226] D3. Based on the convex function pairs CF1 and CF2 and the iterative algorithm parameter set IAP, the weight matrix is ​​efficiently solved by the iterative optimization execution algorithm to obtain the optimized weight allocation scheme OWS;

[0227] Based on the convex function pairs CF1 and CF2 and the parameter set IAP of the iterative algorithm, the system solves for the weight matrix through iterative optimization. In each iteration, the algorithm first performs a linear approximation of CF2 at the current point, then solves a convex subproblem, gradually approaching the optimal solution. This method avoids the difficulty of directly solving the original non-convex problem while maintaining computational efficiency, ultimately yielding the optimized weight allocation scheme OWS.

[0228] D4. Perform convergence verification algorithm processing on the optimized weight allocation scheme OWS to confirm the stability and optimality of the solution, avoid the algorithm parameter dependency problem, and obtain the verified weight matrix VWM;

[0229] To ensure the quality of the solution, the system performs rigorous convergence verification on the optimized weight allocation scheme OWS. The verification process includes repeatedly solving the problem with different initial values, analyzing the stability of the convergence trajectory, and comparing the consistency of solutions under different parameter settings. This step effectively avoids algorithm parameter dependency issues, ensuring the stability and optimality of the solution, and ultimately outputting the verified weight matrix VWM.

[0230] D5. For the verified weight matrix VWM, a comprehensive communication strength weight matrix is ​​constructed through a final fusion algorithm to obtain the communication strength weight matrix CIW.

[0231] Finally, the system applies a final fusion algorithm to construct a comprehensive communication strength weight matrix (CIW) based on the verified weight matrix VWM. The fusion process comprehensively considers the relative importance and mutual influence of various weight factors, ensuring that the final weight allocation reflects both the independent contributions of each factor and captures their synergistic effects, providing an accurate decision-making basis for subsequent data layout optimization.

[0232] Specifically, convex function decomposition is a technique that transforms complex non-convex optimization problems into a more manageable form. For example, the objective function can be decomposed into a more manageable form. Expressed as the difference between two convex functions: ,in and All are convex functions. This representation allows the application of the Difference of Convex Programming (DCP) framework, a common approach within which is the Convex-Concave Procedure (CCCP). In CCCP, each iteration will... At the current point Linearize the subplot and then solve the convex subproblem: This subproblem is convex and can be solved efficiently. The iterative optimization algorithm repeatedly executes this process until it converges to a local optimum. Convergence verification algorithms can be performed by running the algorithm multiple times and analyzing the consistency of the results, or by confirming the convergence properties of the algorithm through theoretical analysis.

[0233] It should be noted that Example 4 mainly involves solving complex discrete optimization problems by employing techniques such as continuous relaxation, potential function design, gradient descent optimization, and discretization transformation during the construction of the two-layer mapping matrix. The core innovation of this example lies in transforming the traditional NP-hard discrete optimization problem into a more manageable continuous optimization problem, and then using a carefully designed algorithm to convert the continuous solution back into a practically usable discrete solution.

[0234] Continuous relaxation is a fundamental step in Example 4. Its core idea is to extend the binary mapping variable, which originally could only take the values ​​0 or 1, to the continuous interval [0,1]. In the mapping problem between graph data and OTN networks, the original discrete variables... Represents graph data vertices Mapped to physical nodes This discrete constraint makes the optimization problem fall under the category of integer programming, resulting in extremely high solution complexity. Through continuous relaxation, the variables are redefined as... , can be understood as vertex Mapped to node The probability or confidence level.

[0235] Relaxation treatment not only changes the range of values ​​for variables, but also requires corresponding adjustments to constraints and the objective function. For example, the original assignment constraints... The requirement that each vertex must be assigned to a physical node remains in place after relaxation, but now a vertex is allowed to be assigned to multiple nodes with varying probabilities. The capacity constraint also needs to be modified accordingly. (physical node) The capacity constraint is transformed into a probabilistic constraint. The mathematical representation of continuous relaxation is: [The original problem is then transformed into a probabilistic constraint]. subject to Convert to subject to ,in It is the objective function. and Linear constraints were defined.

[0236] The main advantage of continuous relaxation lies in transforming non-convex discrete optimization problems into relatively more tractable continuous optimization problems. Although the relaxed problem may still be non-convex, the "smoothness" of the continuous space makes gradient information more useful, and optimization algorithms are more likely to overcome local optima. In addition, continuous variables allow the use of mature numerical optimization techniques, such as Newton's method, quasi-Newton methods, or stochastic gradient descent, which have good theoretical guarantees and practical performance in continuous spaces.

[0237] Detailed implementation of the potential function design algorithm

[0238] The potential function design algorithm is a key innovation in Example 4. Its purpose is to construct a special function that guides the continuous optimization process to produce results close to the discrete solution. The basic idea of ​​the potential function is to add a "penalty term" to the objective function, making intermediate values ​​far from 0 and 1 "expensive," thereby prompting the optimization process to naturally converge to the discrete solution.

[0239] Classical potential function forms include the quadratic penalty function. Sum of logarithmic potential functions The quadratic penalty function in and When the value is 0, When the maximum value is reached This forms an inverted parabolic shape. The logarithmic potential function utilizes the concept of entropy from information theory, when... It tends to 0 when it is close to 0 or 1, and takes a larger positive value when it is in the middle.

[0240] The implementation of the potential function design algorithm includes the following key steps. First is the selection of the function form; the algorithm chooses a suitable potential function form based on the characteristics of the problem. For mapping problems, separable potential functions are typically used. The penalty for each variable is independent. Next is parameter optimization, specifically the parameters in the potential function. It needs careful adjustment; it's too small. Unable to effectively guide discretization, too large This may lead to optimization difficulties. The algorithm employs an adaptive parameter adjustment strategy, dynamically adjusting parameters based on the "dispersion" of the current solution. Value. The dispersion can be defined as... ,when Increase when larger Conversely, decrease .

[0241] The potential function also needs to consider constraint compatibility. In constrained optimization problems, the potential function cannot violate the feasibility of the original constraints. The algorithm handles constraints using the Lagrange multiplier method or the penalty function method to ensure compatibility between the potential function and the constraint conditions. The modified objective function is in the form of: ,in It is the original objective function. It is a potential function term. It is a constraint function. It is a Lagrange multiplier.

[0242] The algorithm also implements a multi-stage potential function strategy, using potential functions of varying strengths at different optimization stages. A weaker potential function is used in the initial stage, allowing the algorithm to explore the continuous space fully; as the optimization progresses, the effect of the potential function is gradually strengthened, guiding the solution towards discrete values. This strategy can be expressed as... ,in It is the number of iterations. , , These are control parameters.

[0243] Gradient descent algorithms are responsible for searching for optimal solutions in a continuous space, and these algorithms need to handle constrained non-convex optimization problems. Example 4 uses an improved projective gradient descent algorithm that combines momentum mechanisms and adaptive step size adjustment.

[0244] The basic iterative formula of the algorithm is: ,in Is it to the feasible region? The projection operator, It's the step length. It is the momentum coefficient. It is the gradient of the objective function at the current point.

[0245] Gradient calculation is the core part of the algorithm. For composite objective functions... The gradient can be decomposed into: Original objective function The gradient of a potential function typically involves complex matrix operations, and algorithms employ automatic differentiation techniques to calculate the accurate gradient. The gradient of a potential function is relatively simple; for example, the gradient of a quadratic potential function is... .

[0246] Projection is a crucial step in handling constraints. It's essential for assigning constraints. Non-negativity constraints The projection operator needs to map any point into the feasible region. The algorithm uses a block projection strategy, first mapping each vertex... Assignment vector Perform a simplex projection and then handle the capacity constraints. The closed-form solution for the simplex projection is: ,in It is a Lagrange multiplier, satisfying .

[0247] The step size adjustment mechanism employs Armijo backtracking search to ensure sufficient descent of the objective function in each iteration. The initial step size is set to... If the Armijo condition is not met If the objective function fails to improve, the step size will be halved and the test will continue. The momentum coefficient is adaptively adjusted: momentum is increased when the objective function decreases for several consecutive steps, and momentum is decreased when oscillations occur.

[0248] The algorithm also implements a restart mechanism. When convergence stall is detected, optimization restarts from a different initial point, increasing the probability of finding the global optimum. Restart conditions include a small gradient norm, minimal change in the objective function value, or no improvement after consecutive iterations. The results of multiple runs are combined using an ensemble learning method to improve the quality and stability of the solution.

[0249] Design and implementation of discretization algorithm

[0250] The discretization algorithm is responsible for converting the solution obtained from continuous optimization into a practically usable discrete mapping relationship. This step seems simple, but it actually requires careful design to ensure that the transformed solution satisfies all constraints while preserving the original optimization objective value as much as possible.

[0251] The most intuitive discretization method is threshold rounding, that is, for each continuous variable... ,if If the value is 1, it is set to 1; otherwise, it is set to 0. However, this method may violate constraints, such as a vertex being assigned to multiple physical nodes or a physical node exceeding its capacity. Therefore, Example 4 employs a more complex constraint-aware discretization algorithm.

[0252] The constraint-aware discretization algorithm employs a greedy strategy, making discretization decisions based on the numerical order of continuous solutions. The algorithm maintains a candidate list containing all... Pair and its corresponding The values ​​are sorted in descending order. Then each candidate is processed sequentially, if... If setting it to 1 does not violate any constraints, then the assignment is performed; otherwise, the candidate is skipped. This approach ensures strict constraint satisfaction but may not be globally optimal.

[0253] To improve discretization quality, the algorithm also implements local search optimization. After initial discretization, the algorithm attempts to improve the quality of the solution through local adjustments. Local searches include vertex reassignment (moving vertices from the current physical node to another node) and vertex swapping (swapping the assignments of two vertices). Each local adjustment evaluates its impact on the objective function, and the adjustment is only performed if it improves the objective value.

[0254] The algorithm also considers probabilistic rounding strategies as an alternative to threshold rounding. In probabilistic rounding, the variable... The probability of being set to 1 is equal to its numerical value, that is... This method preserves certain statistical properties, such as the expected degree of constraint satisfaction. In practice, the algorithm generates multiple probabilistic rounding solutions and then selects the best one as the final result.

[0255] To handle infeasible discretization results, the algorithm implements a constraint repair mechanism. When a discretization result violates certain constraints, the repair algorithm seeks the minimum adjustment to make the solution feasible again. The repair process can be modeled as another optimization problem: minimizing the difference from the original continuous solution while satisfying all constraints. The repair algorithm uses the Hungarian algorithm or network flow algorithm to ensure that a feasible solution is found quickly.

[0256] The four core algorithms are seamlessly integrated through a carefully designed interface. Continuous relaxation provides a suitable problem representation for subsequent algorithms, the potential function design ensures the correctness of the optimization direction, the gradient descent algorithm performs the actual numerical optimization, and the discretization algorithm completes the final solution transformation.

[0257] The entire algorithm employs a multi-level optimization strategy. At the macro level, the algorithm seeks the optimal configuration by adjusting the potential function parameters and optimizing hyperparameters. At the micro level, each sub-algorithm has its own parameter tuning mechanism. The algorithm also implements parallel optimization, allowing optimization processes with multiple different initial points to be executed in parallel, ultimately selecting the optimal result.

[0258] A performance monitoring mechanism is implemented throughout the entire algorithm process, tracking convergence speed, solution quality, and computational resource consumption in real time. When performance anomalies are detected, the algorithm automatically adjusts its strategy or restarts the optimization process. This adaptive mechanism enables the algorithm to maintain stable performance across various problem instances.

[0259] Through this multi-algorithm collaborative and multi-level optimization strategy, Example 4 effectively solves the complex optimization problem of large-scale graph data mapping. In actual tests, compared with traditional heuristic algorithms, this method can find higher quality solutions while maintaining reasonable computation time. The convergence and stability of the algorithm have also been verified through numerous numerical experiments.

[0260] Example 5 mainly involves achieving efficient weight matrix solving before obtaining the communication strength weight matrix (CIW) through techniques such as convex function decomposition, differential convex optimization, iterative optimization execution, and convergence verification. The core innovation of this example lies in transforming the complex multi-factor fusion optimization problem into a differential convex optimization framework and designing an efficient algorithm that requires only a single matrix-vector multiplication in each iteration.

[0261] The convex function decomposition algorithm is the foundation of Example 5. Its purpose is to represent the complex non-convex optimization problem of communication strength weight calculation as a difference between two convex functions. In the graph data communication strength mapping problem, the objective function usually contains multiple coupled terms, such as communication frequency weight, delay weight, bandwidth weight, etc. The combination of these terms often leads to non-convexity.

[0262] The algorithm first analyzes the structural characteristics of the original objective function. Let the original objective function be... ,in These correspond to the contributions of the communication frequency weight matrix FW, the delay weight matrix DW, the bandwidth weight matrix BW, the wavelength resource weight matrix WW, and the topology matching weight matrix TW, respectively. These are the corresponding weighting coefficients. The key to decomposition lies in identifying the convexity of each subfunction and performing a reasonable convex-concave decomposition on the non-convex terms.

[0263] For subfunctions with quadratic terms, such as When the matrix When the eigenvalues ​​of a function have both positive and negative values, the function is non-convex. The decomposition algorithm represents it as... ,in and Each contains The part corresponding to the positive and negative eigenvalues. Thus, and They are all convex functions.

[0264] For sub-functions containing nonlinear coupling terms, the algorithm employs convex hull decomposition techniques. For example, for functions of the form... For each term, the algorithm first constructs its convex envelope function, and then finds a suitable pair of convex functions. Make The construction of convex hulls is usually based on Jensen's inequality or variational methods to ensure the mathematical rigor of the decomposition.

[0265] Decomposition algorithms also need to handle the effects of constraints. This is especially true when the original problem contains linear constraints. Non-negativity constraints In this case, these constraints need to be preserved in the decomposed problem. The algorithm handles constraints by introducing indicator functions or penalty functions to ensure that the decomposed convex function pairs can correctly represent the constraint structure of the original problem.

[0266] Finally, the algorithm obtains convex function pairs CF1 and CF2 that satisfy... CF1 and CF2 are both convex functions. This decomposition provides a theoretical basis for subsequent difference convex optimization algorithms, enabling complex non-convex optimization problems to be handled using mature convex optimization techniques.

[0267] The differential convex optimization algorithm handles convex function pairs CF1 and CF2, and designs an efficient iterative solution framework. This algorithm is based on the theoretical framework of the Convex-Concave Procedure (CCCP), but significant improvements have been made for specific problems, particularly in the design of a computational structure that requires only a single matrix-vector multiplication in each iteration.

[0268] The core idea of ​​the algorithm is to linearize the concave function CF2 at the current point in each iteration, and then solve a convex optimization subproblem. Specifically, in the... In this iteration, the algorithm is at the current point Calculate the gradient of CF2 at point Then solve the subproblems: ,in This indicates the inner product operation.

[0269] To achieve the goal of requiring only a single matrix-vector multiplication in each iteration, the algorithm conducts an in-depth analysis of the structure of the subproblems. Through reasonable matrix decomposition and vectorization operations, the algorithm transforms the complex matrix optimization problem into a high-dimensional vector optimization problem. Specifically, the weight matrix... Vectorized as The objective function is rewritten in vector form: ,in It is a positive definite matrix. It is a vector.

[0270] The key innovation lies in the matrix. The algorithm utilizes the properties of the Kronecker product to achieve a unique structural design. Represented as In the form of, It is the identity matrix. and It is a smaller dimension matrix. This represents the Kronecker product. This structure makes matrix-vector multiplication... It can be computed efficiently, with a complexity of... Reduce to .

[0271] The algorithm also implements an adaptive step size mechanism. In the standard CCCP algorithm, the step size is usually fixed at 1, but in practical applications, an adaptive step size can significantly improve convergence performance. The algorithm uses Armijo backtracking search to determine a suitable step size: starting from the initial step size... Initially, if the sufficient descent condition is not met, the step size is halved and the attempt continues. The sufficient descent condition is: ,in It's the search direction. These are control parameters.

[0272] To improve the robustness of the algorithm, the differential convex optimization algorithm also integrates several acceleration techniques, including Nesterov momentum acceleration, Anderson acceleration, and FISTA acceleration. The update formula for momentum acceleration is: ,in It is the momentum coefficient, and an adaptive adjustment strategy is adopted.

[0273] The iterative optimization execution algorithm is responsible for coordinating the entire optimization process, including initialization, iteration control, convergence detection, and result output. This algorithm employs a multi-level execution strategy, ensuring both convergence and optimized computational efficiency.

[0274] The algorithm employs a combination of strategies in its initialization phase. In addition to random initialization, it also implements intelligent initialization based on the problem structure. Specifically, the algorithm first solves approximate solutions to each subproblem (such as an optimization problem considering only communication frequency weights), and then weights and combines these solutions to obtain the initial solution for the overall problem.

[0275] The iterative control mechanism employs multiple convergence criteria, in addition to the traditional gradient norm criterion. In addition, the algorithm also monitors the relative changes in the objective function value. Variation in reconciliation The algorithm considers itself converged only when all criteria are met simultaneously.

[0276] To prevent premature convergence to a poor local optimum, the algorithm implements a perturbation-restart mechanism. When it detects that the algorithm may be trapped in a local optimum (e.g., the objective function value changes very little after several consecutive iterations, but the gradient norm remains large), the algorithm applies a random perturbation to the current solution and then continues optimization. The magnitude of the perturbation is adaptively adjusted according to the current gradient distribution. ,in It is the disturbance intensity. It is a random vector.

[0277] The algorithm also implements a parallel optimization strategy. Multiple independent optimization processes are executed in parallel, starting from different initial points, each using different algorithm parameter configurations. Parallel execution not only increases the probability of finding the global optimum but also provides an assessment of the algorithm's stability.

[0278] During algorithm execution, detailed performance monitoring information is maintained, including the objective function value, gradient norm, step size, and number of matrix-vector multiplications for each iteration. This information is used not only for convergence detection but also to provide a basis for online adjustment of algorithm parameters. For example, when the step size is very small for several consecutive iterations, the algorithm will appropriately relax the convergence criterion; when the objective function value oscillates, the algorithm will increase the momentum coefficient to stabilize the convergence process.

[0279] Design and implementation of convergence verification algorithm

[0280] The convergence verification algorithm is a crucial component of Example 5. Its purpose is to ensure the stability and optimality of the optimization results and avoid parameter dependency issues. This algorithm employs a multi-dimensional verification strategy, evaluating its performance through both theoretical analysis and numerical experiments.

[0281] In terms of theoretical verification, the algorithm is analyzed based on the convergence theory of differential convex optimization. According to the theoretical guarantees of the CCCP algorithm, under appropriate conditions, the algorithm sequence... It will converge to the critical point of the original problem. The verification algorithm checks whether these theoretical conditions are met, including the Lipschitz continuity of the objective function, the boundedness of the gradient, and the compactness of the feasible region. In its implementation, the algorithm calculates the upper bound of the Lipschitz constant: and verify .

[0282] Numerical validation employs a multiple-repeation validation strategy. The algorithm is run multiple times using different random seeds, initial points, and parameter configurations, and then the consistency of the results is analyzed. Consistency metrics include the distance variance of the solutions. coefficient of variation of the objective function value , where W is the average of the results from multiple runs. and These are the standard deviation and mean of the objective function value, respectively.

[0283] The algorithm also implements sensitivity analysis to assess how sensitive the results are to changes in input parameters. This is achieved by adjusting the weighting coefficients. Convergence threshold The algorithm parameters are slightly perturbed, and the magnitude of the change in the final result is observed. The sensitivity index is defined as: ,in These are the parameters that have been disturbed. It is the change in parameters. This represents the corresponding change in the solution. If the sensitivity indices of all parameters are within a reasonable range, the algorithm is considered to have good stability.

[0284] To detect potential numerical instabilities, the verification algorithm also implements condition number monitoring. The condition number of the Hessian matrix is ​​calculated in each iteration. ,in and These are the maximum and minimum eigenvalues, respectively. When the condition number is too large, the algorithm will use numerical stabilization techniques, such as adding a regularization term or using a preconditioner.

[0285] The final step in validating the algorithm is residual analysis. By calculating the degree of violation of the KKT conditions, it assesses whether the current solution is the true optimum. For constrained optimization problems, the KKT conditions include gradient conditions, primal feasibility, dual feasibility, and complementary relaxation conditions. The algorithm calculates the degree of violation of each condition and combines them into the overall residual: Each term corresponds to a residual under different KKT conditions.

[0286] Through this comprehensive verification strategy, the algorithm can ensure the reliability and stability of the optimization results, providing credible weight information for subsequent graph data layout decisions.

[0287] Example 6

[0288] Before obtaining the Communication Strength Weighting Matrix (CIW), the method further includes:

[0289] E1. Perform node organization processing on the two-layer mapping relationship matrix DLM to organize the OTN network nodes into a logical hierarchical structure, resulting in a hierarchical node structure HNS, wherein each node only maintains local mapping relationships related to itself.

[0290] Organizing nodes in the two-layer mapping matrix (DLM) is the first step in distributed optimization. The system organizes the OTN network nodes into a hierarchical structure, forming a layered node structure (HNS).

[0291] In this structure, each node only needs to maintain local mappings directly related to itself, without needing to store complete information about the entire network. This significantly reduces the storage burden on individual nodes and improves system scalability. The hierarchical structure can be divided based on geographical location, administrative divisions, or network topology characteristics, forming a multi-level organizational structure that facilitates efficient information transmission and management.

[0292] E2. For the hierarchical node structure HNS, each node calculates a candidate weight scheme based on local information and broadcasts it in the first round of communication through an asynchronous message passing protocol to obtain a candidate weight scheme set CWS.

[0293] For the established hierarchical node structure HNS, the system employs an asynchronous message passing protocol for distributed computing. Each node independently calculates candidate weight schemes based on its own local information and broadcasts these schemes to other nodes in the network during the first round of communication. This asynchronous communication mode allows nodes to continue processing without waiting for responses from other nodes, significantly improving the system's parallelism and response speed. This process aggregates to form a candidate weight scheme set CWS, which contains multiple possible schemes from various nodes in the network.

[0294] E3. Based on the candidate weight scheme set CWS, through a two-round consensus protocol, each node selects the globally optimal weight scheme in the second round of communication based on all collected candidate schemes and uses a hash function-weighted majority vote to obtain the consensus weight scheme ConsWS;

[0295] Based on the collected candidate weight scheme set CWS, the system achieves global consensus through a two-round consensus protocol. In the second round of communication, each node selects the globally optimal weight scheme based on all candidate schemes collected in the first round using a hash function-weighted majority voting mechanism. This voting mechanism effectively reduces the possibility of malicious nodes manipulating the voting results by calculating a hash value for each scheme and using specific bits of the hash value as weight factors. In this way, the system efficiently reaches a consensus weight scheme ConsWS while maintaining low communication overhead.

[0296] E4. Based on the consensus weight scheme ConsWS and node behavior monitoring data, an adaptive security mechanism is used to detect abnormal nodes and dynamically adjust their weights in the consensus process to deal with malicious behavior and system failures, resulting in a security-enhanced weight scheme SEWS.

[0297] To enhance system security, an adaptive security mechanism is implemented based on the ConsWS consensus weight scheme and continuously collected node behavior monitoring data. This mechanism can detect abnormal nodes in the network in real time, identify potential malicious behavior or system failures, and dynamically adjust the weight influence of these nodes in the consensus process. By reducing the decision weight of suspicious nodes, the system can effectively respond to various security threats, ensuring normal operation even with a small number of malicious nodes, ultimately forming the security-enhanced weight scheme SEWS.

[0298] E5. Based on the security-enhanced weighting scheme SEWS, perform distributed fusion processing to construct a globally consistent communication strength weighting matrix while maintaining linear communication overhead, thus obtaining the communication strength weighting matrix CIW.

[0299] Finally, the system performs distributed fusion processing based on the security-enhanced weighting scheme SEWS to construct a globally consistent communication strength weight matrix CIW. This process employs a specially designed distributed algorithm to ensure global data consistency while maintaining linear communication overhead. Compared with traditional methods, this distributed fusion mechanism significantly reduces the system's communication burden, improves operating efficiency in large-scale network environments, and guarantees the accuracy and consistency of the final results.

[0300] Specifically, the hierarchical node structure can be organized based on factors such as network topology or administrative divisions. For example, nodes can be divided into multiple regions according to geographical location, with a flat structure within each region and communication between regions through region representative nodes. Asynchronous message passing protocols allow nodes to continue processing without waiting for responses from other nodes, improving system parallelism and response speed. The two-round consensus protocol's first round is the information gathering phase, where each node broadcasts its candidate solutions; the second round is the decision-making phase, where nodes select the globally optimal solution based on all collected solutions using a consistent decision function. Hash function-weighted majority voting is a tamper-proof voting mechanism. By calculating a hash value for each solution and using certain bits of the hash value as weights, it reduces the possibility of malicious nodes manipulating the vote. Adaptive security mechanisms continuously monitor node behavior, identify potential malicious or faulty nodes, and dynamically adjust their weights in the consensus process, ensuring the system can still operate normally even with a small number of malicious nodes.

[0301] Example 6 primarily relates to a distributed optimization method based on asynchronous leaderless consensus. Through techniques such as node organization, asynchronous message passing, a two-round consensus protocol, and adaptive security mechanisms, it achieves efficient collaborative optimization in large-scale distributed environments. The core innovation of this example lies in designing an algorithm that achieves global consensus with only two rounds of communication, while maintaining linear communication overhead and good security.

[0302] The hierarchical node structure is the basic architecture of Example 6. Its design goal is to organize large-scale OTN network nodes into an efficient hierarchical structure, so that each node only needs to maintain local information, significantly reducing the storage and communication burden. The design of this structure draws on the hierarchical organization principle in distributed systems, combined with network topology characteristics and geographical distribution features.

[0303] The structural organization algorithm first analyzes the topological characteristics of the OTN network to identify natural clustering structures. The algorithm employs an improved spectral clustering method, constructing a similarity matrix based on network connectivity strength and geographical proximity. ,in The connectivity item represents the connection strength between nodes, which can be calculated by comprehensively considering metrics such as link bandwidth, latency, and reliability; the proximity item represents geographical proximity, usually based on physical distance or network hop count.

[0304] The hierarchical structure employs a multi-level organization, forming a tree-like management architecture. At the bottom level, physically adjacent nodes form basic units (leaf nodes), each containing 3-5 physical nodes. The middle layer consists of regional nodes composed of multiple basic units, and the top layer is the root node of the entire network. Each layer has a corresponding representative node responsible for communicating with the layers above and below, forming a complete information transmission link.

[0305] The local information maintenance mechanism of nodes is a key feature of hierarchical structures. Each node only needs to store mapping relationships directly related to itself, including its own resource status, connection information of direct neighbors, and summary information of its basic unit. Specifically, nodes... The local information maintained includes: resource vectors Neighborhood gathering and unit summary This localized storage strategy reduces the storage complexity of a single node from... Reduce to , among which It is the total number of network nodes. It is the degree of the node. It refers to the unit size.

[0306] The hierarchical structure also implements a dynamic adjustment mechanism, enabling adaptive reorganization based on changes in network status. When a load imbalance or decreased communication efficiency is detected in a basic unit, the reorganization algorithm will reallocate node affiliations or adjust the hierarchical boundaries. Reorganization decisions are based on a utility function. ,in To measure load balancing, Measure communication efficiency. Measure restructuring costs.

[0307] The asynchronous message passing protocol is the core communication mechanism in Implementation Example 6. Its design goal is to achieve efficient distributed information exchange without the need for global synchronization. This protocol adopts an event-driven asynchronous model, allowing each node to send or receive messages at any time without waiting for responses from other nodes.

[0308] The protocol's message format design takes efficiency and scalability into account. Each message consists of two parts: a header and a payload. The header includes fields such as sender ID, receiver ID, message type, timestamp, and sequence number, while the payload contains the specific weighting scheme or calculation result. There are four message types: PROPOSE, RESPONSE, COMMIT, and ABORT, corresponding to different stages of the consensus protocol.

[0309] The asynchronous sending mechanism employs a non-blocking I / O model, with each node maintaining independent sending and receiving threads. The sending thread retrieves messages to be sent from its local message queue and transmits them to the target node via the network interface; the receiving thread listens on the network port, receives messages from other nodes, and adds them to its receive queue. This design avoids blocking the computation process during sending operations, thus improving the system's concurrency performance.

[0310] The protocol also implements a message reliability guarantee mechanism. Each sent message is assigned a unique sequence number, and the receiver sends an acknowledgment response upon receiving the message. If the sender does not receive an acknowledgment within the timeout period, it will retransmit the message. To handle message out-of-order and duplicate issues, the receiver maintains a message sequence number window and only accepts new messages within the window. The state transitions of the reliability mechanism can be represented as follows: Events include message sending, confirmation of receipt, timeout, and other events.

[0311] Load balancing is a crucial consideration in asynchronous message passing. The protocol implements an adaptive message scheduling strategy, dynamically adjusting message sending frequency based on network load and node processing capacity. The scheduling algorithm monitors the latency and throughput of each connection and uses an exponentially weighted moving average (EWMA) to predict network status. When the predicted delay exceeds a threshold, the algorithm will postpone the sending of non-urgent messages to prioritize the timely delivery of critical messages.

[0312] Algorithm implementation of two-round consensus protocol

[0313] The two-round consensus protocol is the core innovation of Implementation Example 6. Through a carefully designed two-stage voting mechanism, it achieves global consensus with only two rounds of communication. This protocol combines the advantages of Byzantine fault tolerance theory and practical consensus algorithms, ensuring the correctness of decisions while minimizing communication overhead.

[0314] The first round is the information gathering phase, where each node calculates candidate weight schemes based on local information and broadcasts them to other nodes in the network. The calculation of candidate schemes uses a local optimization algorithm, with the objective function being: ,in It is a node The local objective function, It is a set of local constraints. The local objective function comprehensively considers the resource status of the current node, the needs of neighboring nodes, and historical performance data.

[0315] The broadcast strategy employs an improved version of the flooding algorithm to prevent messages from propagating indefinitely. Each message carries a TTL (Time to Live) field, which is decremented by one for each node it passes through; the message is discarded when the TTL reaches zero. Furthermore, each node maintains a hash table of messages it has seen to avoid forwarding the same message repeatedly. Reachability of broadcasts is guaranteed through probabilistic analysis; assuming network connectivity, the probability of a message reaching all nodes is [value missing]. ,in It is the single-hop transmission success rate. It is the network diameter.

[0316] The second round is the decision-making phase. Each node, based on all candidate solutions collected in the first round, selects the globally optimal solution through a majority vote weighted by a hash function. The hash weighting mechanism is a key innovation of the protocol, designed to prevent malicious nodes from manipulating candidate solutions to influence the voting results. In the specific implementation, for each candidate solution... Calculate hash value ,in It's a cryptographic hash function, and nonce is a public random number. Then, the lower bits of the hash value are used as a weighting factor. ,in It is the preset weight modulus.

[0317] The voting algorithm uses a weighted voting method, with each node... For candidate solutions The voting weight is ,in It is a node Trust level. Trust level is calculated based on the node's historical behavior, including factors such as proposal quality, response timeliness, and behavioral consistency. The winning solution is the one that receives the highest weighted votes. .

[0318] The protocol also implements a conflict resolution mechanism to handle ties or situations where multiple proposals receive similar numbers of votes. When the lead of the proposal with the most votes is less than a preset threshold, the algorithm initiates a third round of refined voting, selecting only from the top few proposals. This refined voting employs more sophisticated evaluation criteria, such as the computational complexity, implementation difficulty, and risk assessment of the proposals.

[0319] The adaptive security mechanism is a crucial safeguard in Implementation Example 6. Its purpose is to detect and respond to malicious behavior and system failures in the network, ensuring the reliable operation of the consensus protocol in the adversarial environment. This mechanism employs a multi-layered protection strategy, combining technologies such as behavioral analysis, anomaly detection, and dynamic response.

[0320] The behavior monitoring module continuously tracks the activity patterns of each node, establishing a baseline model of normal behavior. Monitoring metrics include message sending frequency, proposal quality score, response time distribution, and voting consistency. The baseline model is trained using machine learning methods, such as support vector machines or neural networks, to learn the statistical characteristics of normal behavior. The model's feature vectors are... The determination of normal behavior is based on a probability threshold: .

[0321] The anomaly detection algorithm employs a strategy combining multiple statistical methods. These include simple threshold-based detection, deviation detection based on statistical distribution, and anomaly detection based on time-series analysis. Threshold detection monitors whether individual metrics exceed normal ranges, such as excessively high message sending frequency or excessively long response times. Distribution detection uses the Kolmogorov-Smirnov test or the Anderson-Darling test to determine whether the current behavioral distribution differs significantly from the historical baseline. Time-series analysis uses ARIMA models or LSTM networks to detect abrupt changes in behavioral patterns.

[0322] The threat classification mechanism categorizes potential threats into several types based on the detected abnormal behavior: Byzantine faults (nodes sending contradictory information), performance attacks (intentionally delayed responses), Sybil attacks (a single entity controlling multiple identities), and Eclipse attacks (isolieving target nodes), etc. Each threat type corresponds to a different response strategy. The system maintains a threat response rule base and automatically selects appropriate protective measures based on the threat type.

[0323] Dynamic weight adjustment is the core response strategy of the security mechanism. When suspicious behavior is detected in a node, the system gradually reduces the node's weight influence in the consensus process. The weight adjustment adopts an exponential decay model: ,in It is a node At any moment Trust level, The score is calculated based on the current behavior. It is the attenuation coefficient.

[0324] The mechanism also implements a collaborative defense strategy, allowing multiple normal nodes to jointly combat malicious nodes. When multiple nodes simultaneously report abnormal behavior of a particular node, the system raises the alert level for that node. Collaborative decision-making employs Byzantine fault-tolerant voting; only when more than (n+f+1) / 2 nodes agree will punitive measures be taken against the target node. It is the total number of nodes. It represents the maximum number of malicious nodes that can be tolerated.

[0325] The recovery mechanism handles legitimate nodes that have been misjudged. The system maintains a reputation recovery path for each node, gradually restoring trust through consistent positive behavior. Recovery speed is related to historical reputation and current performance. This design both punishes malicious behavior and provides opportunities for repentant nodes, thus maintaining the fairness and incentive compatibility of the system.

[0326] Through this multi-layered, adaptive security mechanism, Example 6 can maintain high reliability and security in complex network environments, providing a solid foundation for large-scale distributed graph data processing.

[0327] Example 7

[0328] like Figure 7 As shown, the present invention also provides a graph data communication intensity mapping system based on optical transport network topology, comprising:

[0329] The acquisition module is used to acquire graph data structure files, extract basic structural information of the graph through a graph data parser, establish a graph adjacency matrix and calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate a graph data communication feature matrix (GCF).

[0330] The measurement module is used to acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF).

[0331] The mapping module is used to establish the mapping relationship between graph computing nodes and OTN network physical nodes based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF through a resource mapping algorithm, and to construct a two-layer topology association matrix and output a two-layer mapping relationship matrix DLM.

[0332] The weight calculation and optimization processing module is used to calculate and optimize the communication strength weight matrix (CIW) of the two-layer mapping relationship matrix (DLM) by taking communication frequency, delay sensitivity, bandwidth requirement, wavelength resources and topology matching degree as input parameters, and using a weight definition algorithm and a multi-factor fusion algorithm.

[0333] The evaluation module is used to process the communication strength weight matrix CIW using a scheme enumeration algorithm, generate candidate layout schemes, and evaluate communication overhead, resource utilization, and load balancing. It also calculates a comprehensive score for topology matching and determines the optimized data layout scheme OLP.

[0334] The adjustment module is used to dynamically adjust the mapping relationship based on the optimized data layout scheme OLP by monitoring system performance indicators and network status changes in real time and using an incremental update algorithm to obtain a dynamically optimized mapping relationship DOM.

Claims

1. A graph data communication intensity mapping method based on optical transport network topology, characterized in that, Includes the following steps: Obtain the graph data structure file, extract the basic structural information of the graph through the graph data parser, establish the graph adjacency matrix and calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate the graph data communication feature matrix GCF. Acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF). Based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, a resource mapping algorithm is used to establish a mapping relationship between graph computing nodes and OTN network physical nodes, and a two-layer topology association matrix is ​​constructed to output a two-layer mapping relationship matrix DLM. This includes: processing the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF using a resource mapping algorithm to establish an initial mapping relationship between graph computing nodes and OTN network physical nodes, obtaining a node mapping matrix NM; mapping the communication requirements between graph data vertices to the physical transmission path of the OTN network using a path mapping algorithm based on the node mapping matrix NM and the communication frequency matrix F, obtaining a path mapping matrix PM; and based on the path mapping matrix PM and the delay cycle... The latency sensitivity matrix DC is obtained by evaluating the sensitivity of graph data communication to network transmission latency using a latency sensitivity analysis algorithm. Based on the path mapping matrix PM and the communication frequency matrix F, the bandwidth requirement of graph data communication is estimated using a bandwidth requirement analysis algorithm, resulting in a bandwidth requirement matrix BR. An association matrix construction algorithm is applied to the latency sensitivity matrix DS and the bandwidth requirement matrix BR to establish an association matrix between the graph data logical topology and the OTN physical topology, resulting in a two-layer topology association matrix TRM, which represents the mapping relationship between the two topologies. Finally, a mapping relationship synthesis algorithm is used to construct the final two-layer mapping relationship matrix DLM, which comprehensively describes the mapping relationship between graph data and the OTN network. For the two-layer mapping relationship matrix DLM, communication frequency, latency sensitivity, bandwidth requirement, wavelength resources and topology matching degree are taken as input parameters. The weights of DLM are calculated and optimized by weight definition algorithm and multi-factor fusion algorithm to obtain the communication strength weight matrix CIW. The communication strength weight matrix CIW is processed by a scheme enumeration algorithm to generate candidate layout schemes and to evaluate communication overhead, resource utilization and load balancing. The comprehensive score of topology matching degree is calculated to determine the optimized data layout scheme OLP. Based on the optimized data layout scheme OLP, the mapping relationship is dynamically adjusted by using an incremental update algorithm through real-time monitoring of system performance indicators and network status changes, resulting in a dynamically optimized mapping relationship DOM.

2. The method of claim 1, wherein, Obtain the graph data structure file, extract the basic structural information of the graph using a graph data parser, establish a graph adjacency matrix, calculate the vertex degree centrality and clustering coefficients of the graph adjacency matrix, and generate the graph data communication feature matrix (GCF), including: The graph data structure file is parsed to extract the vertex set V and edge set E of graph G, thereby obtaining the basic structural information of graph G. For the vertex set V and the edge set E, the adjacency matrix A of the graph G is established using a matrix construction algorithm, where A[i][j] represents the connection relationship between vertices i and j; The adjacency matrix A is subjected to degree centrality calculation, and the degree centrality value D(v) of each vertex in the adjacency matrix A is calculated to determine the vertex degree centrality vector D. The degree centrality value D(v) is used to reflect the importance of each vertex in the graph G. Based on the adjacency matrix A and the degree centrality vector D, the clustering coefficient C(v) of each vertex is calculated using a clustering coefficient calculation algorithm to determine the vertex clustering coefficient vector C, wherein the clustering coefficient C(v) characterizes the connection density between the neighbors of each vertex. For the adjacency matrix A, the degree centrality vector D, and the clustering coefficient vector C, a communication frequency matrix F between each vertex is established using a communication pattern analysis algorithm, where F[i][j] represents the predicted communication frequency between vertices i and j; Based on the adjacency matrix A, the degree centrality vector D, the clustering coefficient vector C, and the communication frequency matrix F, a comprehensive graph data communication feature matrix GCF is constructed through a feature fusion algorithm. The graph data communication feature matrix GCF includes the topological features and communication pattern features of the graph.

3. The method of claim 2, wherein, Acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF), including: The existing OTN network management system is processed with network detection tools and device interfaces to collect information on optical transport network devices in the OTN network, including node location, device type and functional attributes, to obtain the OTN network node set N. Based on the node set N, the physical fiber optic connections and optical path information between each node in the node set N are identified by the link discovery protocol and topology detection algorithm to obtain the optical link set L; For the node set N and the optical link set L, an OTN physical topology matrix P is established using a topology matrix construction algorithm, where P[i][j] represents the direct connection status between each node i and j; The optical link set L is subjected to network performance monitoring processing. The bandwidth capacity, current utilization and transmission delay of each optical link are measured and collected to obtain the link parameter set LP. Based on the physical topology matrix P and the link parameter set LP, the multi-level delay circle structure of each node in the OTN network is calculated using the delay circle analysis algorithm, and the delay circle matrix DC is output. Based on the optical link set L and real-time network monitoring data, the wavelength occupancy rate, available wavelength quantity, and wavelength allocation mode of each optical link are statistically analyzed using a wavelength resource analysis algorithm to obtain the wavelength resource status matrix WR. For the physical topology matrix P, the time delay circle matrix DC, and the wavelength resource state matrix WR, a comprehensive OTN physical network feature matrix ONF is constructed through a feature fusion algorithm. The OTN physical network feature matrix ONF contains network topology and performance parameter information.

4. The method of claim 1, wherein, For the aforementioned two-layer mapping matrix DLM, communication frequency, delay sensitivity, bandwidth requirement, wavelength resources, and topology matching degree are taken as input parameters. A weighting algorithm and a multi-factor fusion algorithm are used to calculate and optimize the DLM, resulting in the communication strength weight matrix CIW, which includes: The communication frequency matrix F is weighted to assign corresponding weights to the communication frequencies between each pair of vertices in the graph data, resulting in the communication frequency weight matrix FW. Based on the delay sensitivity matrix DS and the delay circle matrix DC, a delay weight calculation algorithm is used to assign corresponding weights to communications with different delay requirements, resulting in a delay weight matrix DW. The delay weight matrix DW is used to reflect the degree of impact of delay on performance. Based on the bandwidth demand matrix BR and the link parameter set LP, a bandwidth weight matrix BW is obtained by assigning corresponding weights to communications with different bandwidth demands through a bandwidth weight calculation algorithm. The bandwidth weight matrix BW is used to reflect the degree of influence of bandwidth on performance. The wavelength resource state matrix WR is processed by wavelength weight calculation. The corresponding weights are assigned according to the scarcity and importance of wavelength resources to obtain the wavelength resource weight matrix WW. For the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, the structural similarity weights between the logical topology and the physical topology are calculated using a topology matching algorithm to obtain the topology matching weight matrix TW; Based on the communication frequency weight matrix FW, the delay weight matrix DW, the bandwidth weight matrix BW, the wavelength resource weight matrix WW, and the topology matching weight matrix TW, the final communication strength weight factor is calculated through a multi-factor fusion algorithm to obtain the communication strength weight matrix CIW.

5. The method of claim 1, wherein, The communication strength weight matrix (CIW) is processed using a scheme enumeration algorithm to generate candidate layout schemes. Communication overhead, resource utilization, and load balancing are evaluated, and a comprehensive topology matching score is calculated to determine the optimal data layout scheme (OLP). This includes: Based on the graph data scale and OTN network resource status, the communication strength weight matrix CIW is processed by a scheme enumeration algorithm to generate multiple possible graph data layout schemes. These multiple possible graph data layout schemes include different vertex partitioning and mapping strategies, resulting in a candidate layout scheme set CLS. For the candidate layout scheme set CLS and the communication strength weight matrix CIW, the total communication overhead under each layout scheme is calculated by the communication overhead evaluation algorithm to obtain the communication overhead evaluation result matrix CEC; Based on the candidate layout scheme set CLS and the OTN physical network feature matrix ONF, the network resource utilization rate under each layout scheme is calculated by the resource utilization rate evaluation algorithm to obtain the resource utilization rate evaluation result matrix RUE. The candidate layout scheme set CLS and the two-layer mapping relationship matrix DLM are processed by a load balancing evaluation algorithm to calculate the degree of computational and network load balancing under each layout scheme, and the load balancing evaluation result matrix LBE is obtained. Based on the communication overhead evaluation result matrix CEC, the resource utilization evaluation result matrix RUE, and the load balancing evaluation result matrix LBE, a comprehensive scoring algorithm is used to calculate the topology matching degree comprehensive score for each layout scheme. Taking into account communication overhead, resource utilization, and load balancing, a topology matching degree scoring matrix TMS is obtained. The topology matching score matrix (TMS) is processed by an optimal solution selection algorithm, and the data layout scheme with the highest score is selected as the optimal solution. The optimal solution is the optimized data layout scheme (OLP).

6. The method of claim 1, wherein, After constructing the two-layer topological association matrix, the method further includes: The path mapping matrix PM is processed by communication strength data stream identification to identify critical data transmission paths that require enhanced reliability, resulting in a critical path set KPS. Based on the critical path set KPS and the delay sensitivity matrix DS, the optimal error correction coding parameters are determined for each critical path in the critical path set KPS using the Reed-Solomon error correction parameter calculation algorithm, resulting in the error correction parameter configuration matrix ECC. For the critical path set KPS and the error correction parameter configuration matrix ECC, the high communication intensity data stream is Reed-Solomon encoded using a redundancy coding processing algorithm to obtain the reliability-enhanced path mapping matrix ERPM. The reliability-enhanced path mapping matrix ERPM is processed by an association matrix update algorithm to correct the two-layer topological association matrix, resulting in an enhanced two-layer topological association matrix ETRM.

7. The method according to claim 5, characterized in that, Before generating the optimized data layout scheme OLP, the method further includes: The candidate layout scheme set CLS is subjected to tile segmentation processing to divide the graph data into multiple tiles. Each tile contains a subset of associated vertices and edges, resulting in a tile set TS. Based on the tile set TS and the OTN physical network feature matrix ONF, an overlapping execution strategy for data loading and computation is designed using a workload interleaving scheduling algorithm, so that the data preloading of the next tile and the computation of the current tile can be performed in parallel, resulting in the interleaving scheduling scheme ISS. Based on the interleaved scheduling scheme ISS and the system resource status, a cross-stage kernel scheduling algorithm is used to intelligently allocate computing resources between different graph computing stages to obtain the resource scheduling strategy RSS. Based on the tile set TS, the interleaved scheduling scheme ISS, and the resource scheduling policy RSS, the performance index of each layout scheme under the interleaved workload condition is calculated by a comprehensive optimization algorithm to obtain the optimized topology matching degree score matrix OTMS. The optimized topology matching score matrix OTMS is processed by the optimal solution selection algorithm to select the data layout scheme with the best performance under the workload interleaving condition, and the optimized data layout scheme OLP is obtained.

8. The method according to claim 1, characterized in that, In constructing the two-layer mapping matrix, the method further includes: The graph data communication feature matrix GCF and the OTN physical network feature matrix ONF are subjected to continuous relaxation processing, and the discrete 0-1 mapping variables are relaxed into continuous variables with values ​​in the range of [0,1], to obtain the continuous mapping probability matrix CPM. For the continuous mapping probability matrix CPM, a potential function that can guide the optimization solution back to the discrete binary configuration is constructed through a potential function design algorithm, resulting in a potential function parameter set PFP. The potential function is used to take a relatively low value when the continuous variable is close to 0 or 1. Based on the continuous mapping probability matrix CPM and the potential function parameter set PFP, the optimal mapping solution is searched in the continuous space using gradient descent-type algorithms to avoid getting trapped in local optima and obtain the continuous optimal solution COS. Based on the continuous optimization solution COS and the potential function parameter set PFP, the continuous solution is converted into a discrete binary mapping relationship through a discretization algorithm to obtain the discrete mapping relationship matrix DMR. The discrete mapping relationship matrix DMR is subjected to mapping relationship synthesis processing to construct the two-layer mapping relationship matrix DLM.

9. The method according to claim 4, characterized in that, Before obtaining the Communication Strength Weighting Matrix (CIW), the method further includes: For the communication frequency weight matrix FW, the delay weight matrix DW, the bandwidth weight matrix BW, the wavelength resource weight matrix WW, and the topology matching weight matrix TW, the communication strength weight calculation objective is expressed as the difference form of two convex functions through the convex function decomposition algorithm, resulting in convex function pairs CF1 and CF2. The convex function pairs CF1 and CF2 are processed by the differential convex optimization algorithm, and an iterative solution framework is designed so that each iteration only performs a single matrix-vector multiplication operation, thus obtaining the iterative algorithm parameter set IAP; Based on the convex function pairs CF1 and CF2 and the iterative algorithm parameter set IAP, the weight matrix is ​​efficiently solved by the iterative optimization execution algorithm to obtain the optimized weight allocation scheme OWS; The optimized weight allocation scheme OWS is processed by a convergence verification algorithm to confirm the stability and optimality of the solution, avoid the problem of algorithm parameter dependence, and obtain the verified weight matrix VWM. For the verified weight matrix VWM, a comprehensive communication strength weight matrix is ​​constructed through a final fusion algorithm to obtain the communication strength weight matrix CIW.

10. The method according to claim 4, characterized in that, Before obtaining the Communication Strength Weighting Matrix (CIW), the method further includes: The two-layer mapping matrix DLM is processed by node organization to organize the physical nodes of the OTN network into a logical hierarchical structure, resulting in a hierarchical node structure HNS, in which each node only maintains local mapping relationships related to itself. For the hierarchical node structure HNS, an asynchronous message passing protocol is used to enable each node to calculate candidate weight schemes based on local information and broadcast them in the first round of communication to obtain the candidate weight scheme set CWS; Based on the candidate weight scheme set CWS, a two-round consensus protocol is used to enable each node to select the globally optimal weight scheme in the second round of communication based on all the collected candidate schemes and using a hash function-weighted majority vote, thus obtaining the consensus weight scheme ConsWS. Based on the consensus weight scheme ConsWS and node behavior monitoring data, an adaptive security mechanism is used to detect abnormal nodes and dynamically adjust their weights in the consensus process to deal with malicious behavior and system failures, resulting in a security-enhanced weight scheme SEWS. Based on the security-enhanced weighting scheme SEWS, a distributed fusion process is performed to construct a globally consistent communication strength weight matrix while maintaining linear communication overhead, thus obtaining the communication strength weight matrix CIW.

11. The method according to claim 1, characterized in that, Based on the optimized data layout scheme OLP, by real-time monitoring of system performance indicators and network status changes, an incremental update algorithm is used to dynamically adjust the mapping relationships, resulting in a dynamically optimized mapping relationship DOM, including: Based on the optimized data layout scheme OLP, distributed monitoring agent processing is performed to collect CPU utilization, memory usage, network throughput and disk I / O, and obtain the system performance index matrix SPI. Based on the optimized data layout scheme OLP and the monitoring data of the OTN network, the network state changes such as bandwidth utilization, latency changes, bit error rate, and wavelength resource occupancy status of the optical link are monitored in real time using a network state awareness algorithm to obtain the network state change matrix NSC. Based on the system performance index matrix SPI and the network state change matrix NSC, anomaly detection algorithms are used to identify performance bottlenecks, network congestion, and resource imbalances, and the degree of deviation from the optimal state is calculated to obtain the anomaly assessment matrix ASE. Based on the abnormal condition assessment matrix ASE and the preset threshold parameters, an incremental triggering mechanism is used to determine whether the mapping relationship needs to be adjusted. Dynamic adjustment is triggered only when the performance deviation exceeds the preset threshold, and an adjustment decision signal ADS is obtained. The adjustment decision signal ADS and the current mapping relationship state are incrementally updated to calculate the subset of mapping relationships that need to be adjusted. A local remapping strategy is adopted to avoid global recalculation, resulting in the incremental adjustment scheme MIAS for mapping relationships. Based on the incremental adjustment scheme MIAS and the real-time system status, the mapping relationship is adjusted step by step using a progressive execution algorithm to obtain the dynamically optimized mapping relationship DOM.

12. A graph data communication intensity mapping system based on optical transport network topology, characterized in that, include: The acquisition module is used to acquire graph data structure files, extract basic structural information of the graph through a graph data parser, establish a graph adjacency matrix and calculate the vertex degree centrality and clustering coefficient of the graph adjacency matrix, and generate a graph data communication feature matrix (GCF). The measurement module is used to acquire data from the existing OTN network management system, collect information on optical transport network equipment and optical link topology, and measure link parameters and wavelength resource status to obtain the OTN physical network feature matrix (ONF). The mapping module is used to establish a mapping relationship between graph computing nodes and OTN network physical nodes based on the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF, and to construct a two-layer topology association matrix, outputting a two-layer mapping relationship matrix DLM. This includes: processing the graph data communication feature matrix GCF and the OTN physical network feature matrix ONF using a resource mapping algorithm to establish an initial mapping relationship between graph computing nodes and OTN network physical nodes, obtaining a node mapping matrix NM; mapping the communication requirements between graph data vertices to the physical transmission paths of the OTN network using a path mapping algorithm based on the node mapping matrix NM and the communication frequency matrix F, obtaining a path mapping matrix PM; and based on the path mapping matrix PM and... The delay circle matrix DC is used to evaluate the sensitivity of graph data communication to network transmission delay using a delay sensitivity analysis algorithm, resulting in a delay sensitivity matrix DS. Based on the path mapping matrix PM and the communication frequency matrix F, the bandwidth requirement of graph data communication is estimated using a bandwidth demand analysis algorithm, resulting in a bandwidth demand matrix BR. An association matrix construction algorithm is applied to the delay sensitivity matrix DS and the bandwidth demand matrix BR to establish an association matrix between the graph data logical topology and the OTN physical topology, resulting in a two-layer topology association matrix TRM, which represents the mapping relationship between the two topologies. Finally, a mapping relationship synthesis algorithm is used to construct the final two-layer mapping relationship matrix DLM, which comprehensively describes the mapping relationship between graph data and the OTN network. The weight calculation and optimization processing module is used to calculate and optimize the two-layer mapping relationship matrix DLM by taking communication frequency, delay sensitivity, bandwidth requirement, wavelength resources and topology matching degree as input parameters, and using weight definition algorithm and multi-factor fusion algorithm to obtain the communication strength weight matrix CIW. The evaluation module is used to process the communication strength weight matrix CIW using a scheme enumeration algorithm, generate candidate layout schemes, and evaluate communication overhead, resource utilization, and load balancing. It also calculates a comprehensive score for topology matching and determines the optimized data layout scheme OLP. The adjustment module is used to dynamically adjust the mapping relationship according to the optimized data layout scheme OLP by monitoring system performance indicators and network status changes in real time and using an incremental update algorithm to obtain a dynamically optimized mapping relationship DOM.

Citation Information

Patent Citations

  • Wavelength division multiplexing technology-based metropolitan area network system

    CN101945023A

  • Multi-stage virtual private network service provisioning for containerized routers

    US12155569B1