Network topology optimization method, electronic device, storage medium, and program product
By acquiring multidimensional data of network nodes, identifying and optimizing links, the problem of low reliability in network topology optimization in existing technologies is solved, and more comprehensive and accurate network topology optimization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies focus on only a single performance metric when optimizing network topology, ignoring multidimensional factors, resulting in low reliability of network topology optimization.
By acquiring the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints of each network node in the network topology, the links to be optimized are accurately identified, and they are comprehensively evaluated and screened to determine the target to perform operations, ultimately completing the optimization of the network topology.
It improves the effectiveness of network topology optimization, avoids the limitations of traditional single-indicator evaluation, ensures more comprehensive and accurate optimization decisions, and avoids misjudgments caused by missing local data.
Smart Images

Figure CN122437804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to network topology optimization methods, electronic devices, storage media, and program products. Background Technology
[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, the performance of the network topology architecture of data centers, as the core infrastructure supporting massive businesses, directly determines the response efficiency, reliability, and user experience of businesses.
[0003] In related technologies, when optimizing network topology, only a single performance indicator such as latency or load is usually considered. Heuristic algorithms (such as minimum spanning tree and shortest path algorithms) are used to optimize link paths, but the combined influence of multiple factors is ignored, resulting in low reliability of network topology optimization. Summary of the Invention
[0004] This application provides a network topology optimization method, electronic device, storage medium, and program product to at least solve the reliability problem of network topology optimization in related technologies.
[0005] Firstly, this application provides a network topology optimization method, including:
[0006] Obtain the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints for each network node in the network topology.
[0007] Based on the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints of each network node, multiple links to be optimized are identified.
[0008] The optimization process is performed on the plurality of links to be optimized and the optimization execution operations corresponding to each link to be optimized to obtain an optimization strategy set corresponding to the network topology. The optimization strategy set includes the target execution operations corresponding to each link to be optimized.
[0009] Based on the target operation corresponding to each link to be optimized, the network topology is optimized.
[0010] Secondly, this application also provides a network topology optimization device, including an acquisition module, a determination module, a first processing module, and a second processing module:
[0011] The acquisition module is used to acquire the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node in the network topology.
[0012] The determining module is used to determine multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators and service layer constraints corresponding to each network node.
[0013] The first processing module is used to perform optimization processing on the plurality of links to be optimized and the optimization execution operations corresponding to each link to be optimized, to obtain an optimization strategy set corresponding to the network topology, wherein the optimization strategy set includes the target execution operations corresponding to each link to be optimized;
[0014] The second processing module is used to perform operations based on the targets corresponding to each link to be optimized, and to optimize the network topology.
[0015] Thirdly, this application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described network topology optimization methods.
[0016] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described network topology optimization methods.
[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described network topology optimization methods.
[0018] The network topology optimization method, electronic device, storage medium, and program products provided in this application can accurately identify multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to each network node in the network topology. Then, the available optimization operations for each link to be optimized are comprehensively evaluated and screened to determine the target execution operation for each link. Finally, the network topology optimization is completed based on the target execution operation. This avoids the limitations of traditional single-indicator evaluation, making optimization decisions more comprehensive and accurate, avoiding misjudgments caused by missing local data, and thus improving the effectiveness of network topology optimization. Attached Figure Description
[0019] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram illustrating the application scenarios provided in the embodiments of this application;
[0021] Figure 2 A flowchart illustrating a network topology optimization method provided in an embodiment of this application;
[0022] Figure 3 A flowchart illustrating another network topology optimization method provided in this application embodiment;
[0023] Figure 4 This is a schematic diagram of the architecture of a link optimization model provided in an embodiment of this application;
[0024] Figure 5 A schematic diagram of an architecture for link alarm switching provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of a network topology optimization device provided in an embodiment of this application;
[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0029] Figure 1 This is a schematic diagram illustrating an application scenario provided in an embodiment of this application. Please refer to [link / reference]. Figure 1 This application scenario can include a data center network 100, which can include multiple network nodes 101, such as servers, switches, and routers. The network topology of the data center network 100 can be formed by these multiple network nodes 101. The performance of the network topology of the data center network 100 directly determines the service response efficiency, reliability, and user experience. By optimizing this network topology, the service response efficiency and reliability of the data center network 100 can be improved, and the user experience can be enhanced.
[0030] In related technologies, when optimizing network topology, only a single performance indicator such as latency or load is usually considered. Heuristic algorithms (such as minimum spanning tree and shortest path algorithms) are used to optimize link paths, but the combined influence of multiple factors is ignored, resulting in low effectiveness of network topology optimization.
[0031] The network topology optimization method provided in this application can accurately identify multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to each network node in the network topology. It then comprehensively evaluates and filters the optional optimization operations for each link to be optimized, determines the target execution operation for each link, and finally completes the network topology optimization based on the target execution operation. This avoids the limitations of traditional single-indicator evaluation, making optimization decisions more comprehensive and accurate, avoiding misjudgments caused by missing local data, and thus improving the effectiveness of network topology optimization.
[0032] Figure 2 This is a flowchart illustrating a network topology optimization method provided in an embodiment of this application. Please refer to [link / reference]. Figure 2 The method may include:
[0033] S201. Obtain the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node in the network topology.
[0034] A three-tiered monitoring system consisting of the physical layer, transport layer, and service layer can be constructed to achieve comprehensive perception of the entire network status and accurate data collection.
[0035] The data acquisition cycle can be 500ms / time, or it can be dynamically adjusted based on the actual situation.
[0036] At the physical layer, device hardware parameters are collected in real time to identify potential hardware faults (such as continuous high temperature or port failure) in advance, thus enabling early warning of hardware anomalies.
[0037] Physical layer hardware parameters refer to the hardware status data of network nodes, including but not limited to the number of available ports, device temperature, fan speed, power status, etc.
[0038] For example, the equipment temperature is 45℃ and the fan speed is 2000RPM.
[0039] In the transport layer, key operational data of the link are collected based on network protocols to identify network congestion and traffic anomalies.
[0040] Transport layer monitoring can capture transport layer traffic metrics based on technologies such as Simple Network Management Protocol (SNMP), Intelligent Platform Management Interface (IPMI), and telemetry. Transport layer traffic metrics refer to network link performance data, which may include, but are not limited to, link load rate, transmission latency, packet loss rate, bit error rate, and number of error frames, thereby accurately identifying micro-burst traffic and link congestion characteristics.
[0041] For example, the link load rate is 80% and the transmission latency is 10ms.
[0042] At the business layer, business operation information can be parsed, and specific requirements of key businesses on preset latency limits, bandwidth requirements, path redundancy levels, and service level agreement (SLA) compliance requirements can be collected, transforming business requirements into quantifiable network optimization constraints.
[0043] Business layer constraints refer to the rigid constraints that services impose on network performance, including but not limited to latency limits, bandwidth requirements, path redundancy levels, and SLA compliance requirements for critical services.
[0044] For example, the latency limit for critical business operations is 50ms, and the bandwidth requirement is 1Gbps.
[0045] S202. Based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node, determine multiple links to be optimized.
[0046] In some embodiments, multiple single-link evaluation values are determined based on the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints corresponding to each network node; based on the single-link evaluation values corresponding to each single link, multiple links to be optimized are determined among the multiple single links.
[0047] A single link is a directed link between two consecutive network nodes in a network topology, and all single links can be identified in the network topology.
[0048] A single link can include a first network node and a second network node, where the first network node is the upstream node of the second network node. In other words, a single link is represented as a directed connection from the first network node to the second network node.
[0049] The single-link evaluation value is used to indicate the optimization needs and instability of the corresponding single link. The larger the evaluation value, the more prominent the performance shortcomings of the link and the higher the optimization priority.
[0050] Based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node, the link load, port availability, link risk value, and latency deviation rate of a single link can be extracted, thereby determining the evaluation value of a single link.
[0051] Optionally, a single link whose evaluation value is greater than a preset threshold can be identified as a link to be optimized.
[0052] In this application, a quantitative evaluation system is constructed by integrating multi-dimensional monitoring data, which can comprehensively cover abnormal performance links and potential risk links, avoid the omissions caused by a single dimension, improve the accuracy and comprehensiveness of the identification of links to be optimized, and lay a reliable foundation for subsequent optimization operations.
[0053] S203. Perform optimization processing on multiple links to be optimized and the optimization execution operations corresponding to each link to be optimized to obtain a set of optimization strategies corresponding to the network topology.
[0054] The optimization strategy set can include the target execution operation corresponding to each link to be optimized. The target execution operation is the operation that meets the constraints and has the best benefit, selected from the optional optimization execution operations.
[0055] Optimization operations are actions that improve link performance or network stability, used to specifically optimize the links to be optimized. Examples include bandwidth expansion, path redundancy configuration, port upgrades, and link traffic migration.
[0056] In some embodiments, multiple optimization execution operations can be determined; for any link to be optimized, the optimization benefit ratio of the link to be optimized after performing each optimization execution operation is determined, so as to obtain multiple optimization benefit ratios corresponding to the link to be optimized; and an optimization strategy set is determined based on the multiple optimization benefit ratios corresponding to each link to be optimized.
[0057] Based on the physical constraints of the network topology, cost budget, and business requirements, multiple optimization operations can be determined for the links to be optimized.
[0058] Optimizing the benefit ratio can be used to quantify the network performance improvement gains obtained per unit cost investment. The higher the benefit ratio, the higher the cost-effectiveness of the operation.
[0059] The optimization strategy set can be selected by combining the optimization benefit ratios of multiple optimization operations corresponding to each link to be optimized.
[0060] It is worth noting that the target operation can be an empty set, meaning that some links to be optimized do not need to be optimized, and the overall network performance can meet the optimization target through the optimization operations of other links.
[0061] In this application, by determining the optimization benefit ratios of multiple optimization operations corresponding to each link to be optimized, and then selecting an optimization strategy, it is possible to achieve cost-controllable and technically feasible precise optimization while taking into account multi-dimensional monitoring data and network performance requirements, thereby significantly improving the executability and implementation effect of the optimization strategy.
[0062] S204. Based on the target execution operation corresponding to each link to be optimized, optimize the network topology.
[0063] In some embodiments, the target execution operation corresponding to each link to be optimized can be verified to obtain the verification result; if the verification result is that the verification is passed, the task to be executed corresponding to the target execution operation of each link to be optimized is executed; each task to be executed is executed, and during the execution process, the constraints are monitored, and if the constraints are violated, an interruption operation is performed.
[0064] Verification processing can include bandwidth expansion operations verification, path redundancy operations, etc.
[0065] The tasks to be executed can be bandwidth expansion tasks, path redundancy tasks, etc., corresponding to the operations performed on each target.
[0066] Constraints can be cost constraints, physical constraints, business constraints, etc.
[0067] Successful task results will be permanently saved, and optimized network topology status data (such as link bandwidth, port availability, device risk values, etc.) will be fed back to the hierarchical monitoring system of the solution in real time, providing the latest data support for the next round of network status assessment and optimization decisions.
[0068] In this application, by constructing a full-process execution mechanism, the technical feasibility and security of the target execution operation can be ensured, and network risks during the optimization process can be effectively avoided. At the same time, through real-time monitoring and circuit breaker mechanisms, illegal operations can be blocked in a timely manner, ensuring that the network topology optimization process is stable and controllable, and ultimately improving the reliability of network topology optimization and the overall network performance stability.
[0069] The network topology optimization method provided in this application can accurately identify multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to each network node in the network topology. It then comprehensively evaluates and filters the optional optimization operations for each link to be optimized, determines the target execution operation for each link, and finally completes the network topology optimization based on the target execution operation. This avoids the limitations of traditional single-indicator evaluation, making optimization decisions more comprehensive and accurate, avoiding misjudgments caused by missing local data, and thus improving the effectiveness of network topology optimization.
[0070] Figure 3This is a flowchart illustrating another network topology optimization method provided in an embodiment of this application. Please refer to... Figure 3 The method may include:
[0071] S301. Obtain the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node in the network topology.
[0072] The execution process of S301 can be found in the execution process of S201, and will not be repeated here.
[0073] S302. Based on the physical layer hardware parameters, transport layer traffic indicators and service layer constraints of each network node, determine the evaluation values of multiple single links corresponding to multiple single links.
[0074] In some embodiments, for any single link, the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of the first network node and the second network node are standardized to obtain a standardized dimensionless feature vector; based on the dimensionless feature vector of the first network node and the dimensionless feature vector of the second network node, the single link evaluation value corresponding to the single link is determined.
[0075] Dimensionless eigenvectors refer to standardized data structures that transform multidimensional heterogeneous indicators into a unified standard through normalization processing. These structures are used as inputs to evaluation models to eliminate the interference of differences in the dimensions of different indicators on the evaluation results.
[0076] Specifically, the evaluation value of a single link can be determined through a preset link evaluation model. For example, the link load rate is normalized to 0.8, and the device risk value is normalized to 0.3.
[0077] Sliding window mid-value filtering can be used to remove impulse interference data. By using the maximum-minimum normalization method, positive indicators (the larger the value, the better the performance, such as port availability) and negative indicators (the larger the value, the worse the performance, such as latency and bit error rate) are converted into dimensionless data in the interval [0,1], forming a standardized evaluation feature vector, which provides a unified input for subsequent evaluation.
[0078] In this application, multidimensional data is unified into dimensionless feature vectors through standardization processing to eliminate the interference of dimensional differences on the evaluation model. At the same time, the evaluation logic is constructed in combination with the business layer constraints to ensure that the optimization decision is based on the global perspective of the network state and avoid the impact of local indicator deviations on the evaluation accuracy.
[0079] In some embodiments, multiple feature indicators corresponding to a single link can be determined based on the dimensionless feature vector of the first network node and the dimensionless feature vector corresponding to the second network node; the indicator weights corresponding to each feature indicator can be obtained; and the sum of the products of each feature indicator and its corresponding indicator weights can be determined as the single link evaluation value corresponding to the single link.
[0080] Multiple metrics can include port availability, link load, link risk value, and latency deviation rate.
[0081] The weight of the port availability metric is the port weight, the weight of the link load metric is the load weight, the weight of the link risk value metric is the risk weight, and the weight of the latency deviation rate metric is the latency weight.
[0082] In this application, the influence of each characteristic indicator on the single-link evaluation value can be adjusted according to the indicator weight corresponding to each characteristic indicator, which can improve the accuracy of determining the single-link evaluation.
[0083] Specifically, the single-link evaluation value can be determined in the following ways:
[0084] Step 1: From the first network node ( ), second network node ( Extract the number of currently available ports from the dimensionless feature vector of ). , ), Maximum number of ports ( Equipment risk value () , Extracting single links from transport layer monitoring data Link load ( ), actual link latency ( Extract the corresponding preset latency from the business layer constraints. ).
[0085] Step 2: Calculate the port availability index based on the minimum number of available ports at both ends of the node. The formula is as follows: The larger this indicator is, the more strained the port resources are.
[0086] The link risk value is taken from the first network node ( ), second network node ( The maximum value of the equipment risk value, that is, To ensure the highest risk of the nodes associated with the coverage link;
[0087] The delay deviation rate is calculated based on the absolute value of the difference between the preset service delay and the actual delay, using the following formula: This reflects the degree to which link latency deviates from business requirements.
[0088] Step 3: Obtain the port weight corresponding to the port availability metric. Load weight corresponding to link load Risk weights corresponding to link risk values And the delay weight corresponding to the delay deviation rate. ;
[0089] Each weight is adaptively calculated using the information entropy method, as shown in the following formula:
[0090]
[0091] in, Indicates the first The weight of each indicator Indicates the first Information entropy of the indicator; This corresponds to four metrics: port availability, link load, device risk value, and latency deviation rate. The weight allocation is ensured to focus on the most unstable metric, and the weights meet the following requirements. .
[0092] Information entropy It can be based on the first The statistical distribution of the indicators across the entire network links is determined using the information entropy formula. It is confirmed that, among them, Let j be the normalized probability of the j-th indicator on the i-th link. The greater the fluctuation of the indicator, Approximately uniform, Approximately small.
[0093] Step 4: The sum of the products of port availability and port weight, link load and load weight, link risk value and risk weight, and latency deviation rate and latency weight is determined as the single-link evaluation value.
[0094] Specifically, it is calculated using the following link evaluation function:
[0095] ,
[0096] in, Indicates a single link from the first network node To the second network node A directed link; Indicates the first network node The number of currently available ports, Indicates the second network node The number of currently available ports; This represents the maximum number of ports for the first network node. This represents the maximum number of ports for the second network node. Indicates a single link The real-time normalized link load; the higher the value, the higher the risk of link congestion. Indicates the first network node The equipment risk value ranges from [0,1], with a higher value indicating a higher risk of equipment failure. Indicates the second network node The equipment risk value ranges from [0,1], with a higher value indicating a higher risk of equipment failure. Indicates the preset delay of the service. This represents the actual latency of the link.
[0097] This is a single-link evaluation value. The larger the value, the higher the instability of the link and the higher the optimization priority.
[0098] The function automatically calculates dynamic weights based on information entropy. The weights are adaptively adjusted in real time to follow the fluctuations of the indicators. There is no need for historical labels or manual parameter tuning. The evaluation results are optimal in real time according to the network status, which solves the problems of inaccurate, rigid and unsuitable evaluation of fixed weight models.
[0099] For example, in the early stages of equipment aging, the volatility of equipment risk indicators increases, and information entropy... Decrease, corresponding weight The system automatically improves its performance, prioritizing the avoidance of potentially faulty links; during peak business periods, the volatility of link load metrics increases, and information entropy... Decrease, corresponding weight Automatically add links, with the system prioritizing optimization of high-load links.
[0100] In this application, the weight allocation can be adaptively adjusted according to the volatility of the indicators, ensuring that the evaluation model focuses on the most unstable indicators at present, avoiding the evaluation bias caused by traditional static weights, making the optimization decision more in line with the actual state of the network, and improving the accuracy and foresight of the global optimization.
[0101] S303. Based on the single-link evaluation value corresponding to each single link, identify multiple links to be optimized among multiple single links.
[0102] In some embodiments, a current stability threshold is determined based on the single-link evaluation value corresponding to each single link; single links whose single-link evaluation values are greater than the current stability threshold are identified as abnormal links; the rate of change of the evaluation value corresponding to each single link within a preset time window is determined based on the single-link evaluation value corresponding to each single link; single links whose rate of change of the evaluation value is greater than the rate of change threshold are identified as risk surge links; abnormal links and risk surge links are identified as first candidate links; a second candidate link is determined based on the first candidate links; the first candidate links and the second candidate links are identified as links to be optimized, so as to obtain multiple links to be optimized.
[0103] An abnormal link is a link whose single-link evaluation value exceeds the current stability threshold, corresponding to a link with long-term performance degradation. For example, if a single link has an evaluation value of 0.95, exceeding the current stability threshold of 0.9, it will be marked as an abnormal link.
[0104] Specifically, the current stable threshold is calculated using a dynamic threshold function, as shown in the following formula:
[0105]
[0106] in, The current stable threshold, This represents the average value of the entire network link evaluation. Indicates the standard deviation of the evaluation. This is the adjustment coefficient. The adjustment coefficient can be set to control the strictness of the screening; a larger value indicates a stricter screening.
[0107] Among them, the adjustment coefficient Typical value ranges and corresponding applicable scenarios can be defined: =1.2-1.5: Suitable for core business scenarios requiring high stability and low redundancy; =1.6-2.0: Suitable for general data center scenarios with high throughput and large traffic.
[0108] If the single-link evaluation value Greater than the current stability threshold If so, the single link is identified as an abnormal link, forming a set of abnormal links. .
[0109] A risk surge link refers to a link whose single-link evaluation value changes at a rate exceeding a critical threshold within a preset time window, corresponding to a link experiencing a short-term, sudden performance degradation. For example, if the evaluation value of a single link suddenly increases from 0.7 to 0.85 within 10 minutes, exceeding the critical threshold of 0.15, it is marked as a risk surge link.
[0110] Specifically, the rate of change of the evaluation value is calculated using a rate of change analysis function, as shown in the following formula:
[0111]
[0112] in, The rate of change of the evaluation value for a single link. Indicates the preset time window. This refers to the change in the single-link evaluation value within the preset time window; if This identifies the risk surge links, forming a set of risk surge links. .
[0113] This represents a critical threshold based on historical data statistics, typically ranging from 0.05 to 0.1 per minute, used to identify sudden risk links that deteriorate rapidly within a short period of time.
[0114] Merging abnormal link sets and risk surge link set This yields the first candidate link set, covering two core links that need optimization: those with long-term performance degradation and those with short-term sudden risks.
[0115] After determining the first candidate link, a node association impact analysis is performed to identify high-risk nodes involved in the first candidate link. The potential failure of such nodes not only affects themselves but also directly impacts all downstream links and services that depend on them. Therefore, the critical downstream links directly connected to the high-risk node are identified as second candidate links and included in the optimization candidate set for unified evaluation to suppress the spread of local risks in the network.
[0116] That is, the second candidate link is the direct downstream critical link of the high-risk node associated with the first candidate link.
[0117] In this application, abnormal links (such as links with long-term high load or high equipment risk) are screened through dynamic thresholds, links with surging risks (such as links with sudden performance degradation) are screened through rate of change analysis, and downstream key links of high-risk nodes are expanded by combining node correlation impact analysis, and finally merged to form a set of links to be optimized. This can avoid the omissions or misjudgments caused by traditional single thresholds, and ensure that optimization decisions simultaneously cover links with long-term performance degradation, short-term sudden risks, and potential correlation risks.
[0118] S304. Determine multiple optimized execution operations.
[0119] The execution process of S304 can be found in the execution process of S203, and will not be repeated here.
[0120] S305. For any link to be optimized, determine the optimization benefit ratio of the link to be optimized after performing each optimization operation, so as to obtain multiple optimization benefit ratios corresponding to the link to be optimized.
[0121] In some embodiments, for any optimization operation, determine the latency reduction of the link to be optimized after performing each optimization operation. Maximum bit error rate reduction value Minimum bandwidth stability improvement value Obtain the latency weight corresponding to the latency reduction amount. Bit error rate weight corresponding to the maximum reduction in bit error rate The bandwidth stability weight corresponding to the minimum bandwidth stability improvement value The sum of the product of latency reduction and latency weight, the negative product of maximum bit error rate reduction and bit error rate weight, and the product of minimum bandwidth stability improvement and bandwidth stability weight is determined as the optimization benefit; the ratio of optimization benefit to the implementation cost corresponding to the optimization operation is determined as the optimization benefit ratio.
[0122] Specifically, for the link to be optimized Each executable optimized execution operation Quantify the overall optimization benefits of its topology network. and implementation costs .
[0123] Optimize revenue By reusing the aforementioned information entropy-based weights ,in The improvement in global metrics by the fusion operation, corresponding to latency, bit error rate, and bandwidth stability, is calculated using the following formula:
[0124]
[0125] in, Indicates the link to be optimized. Execute operation The total latency reduction across the entire network is the amount of latency reduction. Indicates the link to be optimized. Execute operation The maximum reduction in bit error rate across the entire network after that. Indicates link Execute operation The minimum bandwidth stability improvement value for the entire network.
[0126] Implementation costs This can include direct and indirect costs such as the cost of purchasing hardware required for operation, the cost of manpower for configuration and debugging, and the cost of resource occupation.
[0127] Optimize the return ratio It is used to measure the overall performance improvement benefit that can be obtained by a unit cost investment. The higher the benefit ratio, the higher the cost-effectiveness and priority of the operation.
[0128] In this application, three core objective functions—minimizing total latency, minimizing maximum bit error rate, and maximizing minimum bandwidth stability—are incorporated into the optimization benefit evaluation framework. By dynamically weighting information entropy to integrate multi-dimensional performance improvement benefits, and combining implementation costs to calculate the benefit ratio, an optimization scheme set can be generated based on the benefit ratio ranking and constraint verification. This ensures that the optimization scheme maximizes global benefits under cost budget, physical feasibility, and SLA constraints, thereby improving the cost-effectiveness and feasibility of the optimization scheme.
[0129] S306. Determine the set of optimization strategies based on the multiple optimization benefit ratios corresponding to each link to be optimized.
[0130] In some embodiments, a multi-objective optimization model can be constructed based on multiple optimization benefit ratios corresponding to each link to be optimized, combined with preset multi-objective optimization objectives and constraints; based on the multi-objective optimization model, a set of optimization strategies corresponding to multiple links to be optimized can be generated.
[0131] The multi-objective optimization model has the following multi-objectives: minimizing total transmission delay, minimizing maximum bit error rate, and maximizing minimum bandwidth stability score. The constraints are cost budget, physical feasibility, and business requirements.
[0132] The set of optimization strategies refers to the feasible combination of optimization operations that satisfy the constraints, that is, the set of target execution operations corresponding to each link to be optimized.
[0133] For example, a certain set of optimization strategies may include operations such as "expanding bandwidth of link A", "switching path of link B", and "upgrading port of link C".
[0134] Specifically, the multi-objective optimization model can be found below:
[0135] The primary goal is to minimize the total network transmission latency.
[0136]
[0137] in, Indicates the total network transmission delay. Indicates link The actual delay It represents the set of all links in the network.
[0138] The second objective is to improve network transmission reliability, which is achieved by minimizing the maximum bit error rate across all links, i.e.:
[0139]
[0140] in, This represents the maximum bit error rate of the entire network link. Indicates link The bit error rate is E, where E is the set of all links in the entire network.
[0141] The third objective is network stability optimization, maximizing the minimum bandwidth stability score, i.e.:
[0142]
[0143] in, This represents the minimum bandwidth stability score across the entire network. Indicates link The bandwidth stability score is calculated using the following formula:
[0144]
[0145] in, Indicates the current load on the link. This represents the historical average load of the link. This indicates the maximum load on the link; the smaller the load fluctuation, The higher the score, the better, with a value range of [0,1].
[0146] By maximizing the minimum bandwidth stability score across all links, the minimum bandwidth stability of the entire network is improved, avoiding performance fluctuations caused by load volatility and ensuring overall network stability.
[0147] The three objective functions mentioned above together constitute a multidimensional evaluation system for the optimization direction.
[0148] At the same time, the model sets strict constraints to ensure that the optimization solution meets cost budget, physical feasibility and business requirements.
[0149] Cost constraints for ,in, express The implementation cost of the project's execution operations. This is a preset cost ceiling; it requires that the total cumulative cost of all target operations must not exceed the cost ceiling.
[0150] physical constraints for ,in, For nodes Port connection count, Represents the set of all nodes in the entire network. For nodes The physical port limit is set; the actual number of connections for each node must not exceed the maximum capacity of its physical port to ensure the physical feasibility of the solution.
[0151] Business constraints for ,in, Indicates business flow The actual end-to-end path delay, This indicates the upper limit of the SLA latency corresponding to this service flow. This represents the set of critical business flows across the entire network; it requires that the path latency of all critical business flows meet the SLA agreement to ensure the core performance requirements of the business.
[0152] Based on the above multi-objective optimization model, and considering the optimization benefit ratio of each link to be optimized, the optional optimization operations for each link to be optimized are arranged in descending order of benefit ratio to form an optimization operation queue. The execution flow is as follows:
[0153] First, construct an empty solution set S, set the cumulative used cost, and assign the current real-time port status to the node port availability mapping table. Second, retrieve operations sequentially from the top of the queue, and verify the cost constraints, physical port constraints, and service SLA constraints item by item. Only when all three constraints are satisfied is the operation included in solution set S; if any constraint is not satisfied, skip the current operation. After the operation passes verification, update the parameters synchronously: update the cumulative used cost, and then refresh the entire network topology status according to the operation type, adjusting status parameters such as link bandwidth and device risk value.
[0154] Next, continuously cycle through the verification queue operations until the accumulated cost reaches the total cost limit. If the entire optimization operation queue has been traversed, the iteration terminates and a set of feasible optimization solutions is output. Finally, if the solution set S is empty after a complete queue traversal, indicating no feasible optimization solutions, the adjustment coefficient is lowered to expand the range of selectable optimization links. The optimization operation queue is regenerated, and the entire iterative verification process is repeated until a non-empty set of feasible solutions is obtained. Starting from the top of the queue, each operation is sequentially verified to ensure it meets one of the three types of constraints. If all constraints are met, the operation is included in the optimization strategy set, and the network's state parameters are updated synchronously. If constraints are violated, the operation is skipped. After the iteration terminates, a set of all compliant optimization strategies is output.
[0155] In this application, by constructing a multi-objective optimization model that takes into account latency, reliability, and stability, and combining the benefit ratio ranking with triple constraint verification, a globally optimal optimization strategy that is cost-controllable, physically feasible, and business compliant is generated. This avoids the problem of unimplementable solutions caused by single-objective optimization or ignoring constraints, and significantly improves the comprehensiveness and practical implementation effect of network topology optimization.
[0156] S307. Perform verification processing on the target corresponding to each link to be optimized, and obtain the verification result.
[0157] Before performing any optimization operations, complete a full resource and status pre-verification to confirm that the operations are ready for implementation. Specifically, this may include resource compatibility verification, status consistency verification, and business impact verification.
[0158] Resource compatibility verification can include hardware resource verification and cost resource verification.
[0159] Hardware resource verification can check whether the number of available ports and port speed levels at both ends of the link to be optimized meet the operational requirements. For example, bandwidth expansion requires ports to support the target rate, and path redundancy requires spare path ports to be idle; path redundancy / routing adjustment operations additionally verify whether the number of ports required for the spare path meets or does not exceed the service redundancy level requirements.
[0160] Cost resource verification can calculate the cumulative implementation cost of the current target operation and confirm that it has not exceeded the cost limit. and cost constraints echo.
[0161] State consistency verification can check the current operating status of the link to be optimized and its associated nodes (such as whether the port is enabled normally and whether the link is in a non-faulty state), avoiding the execution of optimization operations in abnormal states; and verify whether the operation conflicts with the existing network configuration (such as whether path switching will cause routing loops and whether bandwidth expansion contradicts the existing QoS policy).
[0162] Business impact assessment evaluates the potential impact of the operation on critical business processes and confirms that the operation will not disrupt business operations. Latency exceeds SLA limit (with business constraints) For path redundancy / routing adjustment operations, additionally verify whether the end-to-end latency of the path meets the business SLA requirements; for operations involving traffic migration, verify the carrying capacity of the migration path to avoid causing new link congestion.
[0163] The verification results can be divided into verification passed and verification failed.
[0164] If all verification items meet the requirements, the output verification passes; if any verification fails (such as insufficient port resources or configuration conflicts), the output verification fails, and the reason for failure is marked. Operations that fail the pre-verification are directly removed from the solution set, a complete verification log is retained, and an operation and maintenance alarm is triggered.
[0165] S308. If the verification result is successful, then determine the tasks to be executed corresponding to the target execution operations of each link to be optimized.
[0166] Once the pre-verification is passed, the target operation is executed as an atomic transaction, ensuring that the entire operation either succeeds or is completely rolled back, preventing inconsistencies in network configuration. The target operation is then transformed into standardized, executable tasks to ensure the standardization and consistency of the operation implementation.
[0167] Specifically, independent task units can be generated for each target operation, clearly defining the task identifier and the operation object (such as the link). Operation type (e.g., bandwidth expansion, path redundancy configuration), target parameters (e.g., post-expansion speed, backup path routing information), execution priority (based on benefit ratio). (Settings); For operations with dependencies (such as expanding the port of a node before adding a link), sort the tasks according to the dependency order to avoid failure due to incorrect execution order.
[0168] The parameters of tasks of the same type can be standardized. The task types can include bandwidth expansion tasks, path redundancy tasks, and port upgrade tasks.
[0169] Bandwidth expansion tasks can specify the target port, the speed before and after expansion, and the configuration command template.
[0170] The bandwidth expansion operation follows a four-stage atomic transaction execution process: a. Pre-configuration stage: A temporary configuration command is sent to both ends of the link to create a temporary configuration session. The configuration is not written to the device's operating status at this time. b. Configuration verification stage: The device independently verifies the legality of the configuration. If the verification passes, a success signal is returned; if the verification fails, a failure signal is returned directly. c. Configuration commit stage: After successful verification, a configuration commit command is sent, and the temporary configuration officially takes effect in the device's operating status. Resource parameters such as link capacity and port availability are updated synchronously within the system. d. Configuration rollback stage: If a failure occurs in the pre-configuration or configuration verification stage, a temporary configuration destruction command is immediately sent to restore the device's original configuration. System resource parameters are not changed, operation logs are retained synchronously, and maintenance alarms are triggered.
[0171] Path redundancy tasks can specify the sequence of backup path nodes, route injection commands, and traffic weight allocation ratios; port upgrade tasks can specify the target node, port model, and upgrade process steps (such as hardware replacement, driver installation, and performance testing).
[0172] Finally, all pending tasks can be aggregated into a task queue and synchronized to the network controller; back up the current network topology (such as link bandwidth, routing configuration, and port occupancy). This backup is used to perform a complete rollback operation when atomic transactions fail, reserving a rollback mechanism so that the network can be quickly restored to its original state in the event of an operation failure.
[0173] S309. Execute each pending task and monitor the constraints during execution. If the constraints are violated, perform an interrupt operation.
[0174] Each task to be executed can be performed as an atomic transaction according to the task queue order, ensuring that the operation is either fully completed or fails and rolls back, avoiding network anomalies caused by partial execution.
[0175] During execution, the task progress is fed back in real time, and the network topology status data within the system (such as link bandwidth, port availability, and device risk values) is updated synchronously.
[0176] When monitoring constraints, the implementation cost of executed tasks is accumulated in real time. If it approaches the cost limit... (If achieved) (90%), triggering an alert; if it exceeds Immediately activate the circuit breaker; check the node port connection volume in real time. Ensure that the number of ports does not exceed the maximum number of ports on the node. To avoid port exhaustion; to collect real-time latency data for critical business flows through a hierarchical monitoring system. Ensure that the SLA limit is not exceeded. Simultaneously monitor link error rate and bandwidth stability score to ensure optimization goals are achieved.
[0177] If any constraint is detected to be violated (such as excessive service latency or excessive port connection volume), the circuit breaker mechanism will be triggered immediately to interrupt the execution of all subsequent tasks (to prevent the spread of risk); for tasks that have been executed but not completed, a rollback operation will be initiated to restore the network state to that before the task was executed; the violating task and the reason for the violation will be marked, a high-priority alarm will be triggered, and the operation and maintenance personnel will be notified to intervene and handle the situation.
[0178] After all tasks are completed (or interrupted by circuit breaker), the execution results are summarized (successful task list, failed task list, constraint violation record); the optimized final network topology status data is fed back to the hierarchical monitoring system of the solution in real time, providing the latest data support for the next round of monitoring, evaluation, and optimization loop.
[0179] Figure 4 This is a schematic diagram of the architecture of a link optimization model provided in an embodiment of this application. Please refer to... Figure 4 The process of running the link optimization model begins; physical layer hardware parameters and transport layer traffic indicators of each link-related node in the network topology are collected, and standardized link feature data are formed after sliding window filtering and maximum-minimum normalization, combined with service layer constraints; the evaluation value of each link is calculated based on the link evaluation function.
[0180] The system uses a dynamic threshold function to determine whether the evaluation value of a single link exceeds the threshold, and combines the risk change rate to determine whether it is a link with a surge in risk. If so, the single link is marked as a link to be optimized, and node association impact analysis is performed to expand the downstream critical links of high-risk nodes to the set of links to be optimized. If not, the system returns to the link feature collection stage and continues to monitor the status of all links in the network.
[0181] The multi-objective optimization model aims to minimize total transmission delay, minimize maximum bit error rate, and maximize minimum bandwidth stability score, while simultaneously loading three types of constraints: cost budget, physical feasibility, and business requirements. For each link to be optimized, optional optimization operations are executed, the optimization benefit ratio is calculated, and the operations are sorted in descending order of benefit ratio. Then, the operations are verified one by one to ensure they meet the three types of constraints.
[0182] The output is the combination of target execution operations that satisfy all constraints, i.e. the set of optimization strategies, which ends the operation of this link optimization model. The optimized topology state is fed back to the hierarchical monitoring system to support the next round of closed-loop optimization.
[0183] The network topology optimization method provided in this application can accurately identify multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to each network node in the network topology. It then comprehensively evaluates and filters the optional optimization operations for each link to be optimized, determines the target execution operation for each link, and finally completes the network topology optimization based on the target execution operation. This avoids the limitations of traditional single-indicator evaluation, making optimization decisions more comprehensive and accurate, avoiding misjudgments caused by missing local data, and thus improving the effectiveness of network topology optimization.
[0184] In some embodiments, for link failures during network operation, a self-healing mechanism based on a hierarchical monitoring system can be used to achieve rapid self-healing of link failures, ensuring service continuity and network stability.
[0185] Specifically, when the hierarchical monitoring system detects anomalies such as port outages, persistently excessive latency, or excessively high bit error rates in a link, in order to avoid misjudgment and omission of a single indicator, it can integrate multi-dimensional monitoring data to determine a comprehensive change demand index, distinguish between instantaneous network jitter and real link failures, and flexibly adjust the fault judgment strategy according to the reliability requirements of different services.
[0186] For any slave node To the node directed links The formula for calculating the comprehensive change demand index is:
[0187]
[0188] in, , These are the dynamic weight values calculated based on information entropy in the link evaluation function. This represents the link delay skew rate. This represents the normalized bit error rate of the link. This represents the normalized packet loss rate of the link. This represents the port status factor of the link.
[0189] Data Acquisition Link Under normal operating conditions over the past 30 days Historical data, statistically derived to obtain its mean and standard deviation The security threshold is automatically updated every day at midnight, requiring no manual preset. The security threshold calculation formula is as follows:
[0190]
[0191] in, This is the adjustment coefficient, and its design logic is consistent with that of the adjustment coefficient in the link evaluation function mentioned above.
[0192] The link fault classification and execution logic are divided into three categories:
[0193] a. Forced scene switch: When In case of (port failure), regardless of The value directly triggers a forced link switch, activating the path switching engine.
[0194] b. Fault switching determination: When the link If the abnormal state lasts for more than the preset minimum time window (default 500ms, adjustable), it is determined that the link has an unrecoverable risk of failure, and the path switching engine is automatically activated.
[0195] c. High-risk warning determination: When The link is marked as a high-risk warning link, the monitoring and collection frequency is increased, and the warning is pushed to the observable management module simultaneously without triggering the path switching operation.
[0196] The switching engine calls an optimal avoidance path for the affected critical business from the daily automatically updated backup path library. The optimality is determined based on a multi-dimensional evaluation of path latency, bandwidth margin, equipment risk value, and other factors.
[0197] The specific switching process is as follows: Send probe messages to the new path to verify its connectivity, latency, bandwidth margin, and fault-free status. If the verification fails, a backup path is selected again. By adjusting the hash value or routing weight of the service flow, traffic is migrated to the new path in batches and smoothly to avoid traffic surges causing congestion on the new path. The proportion of traffic migrated in each batch shall not exceed 30%, and the interval between adjacent batches shall not be less than 500ms. After confirming that the new path is running stably for 3 consecutive monitoring cycles (each cycle is 500ms by default, but can be configured) and that the service latency and packet loss rate meet the SLA requirements, the service binding of the original faulty link is removed to ensure that the service has been completely migrated.
[0198] Link self-healing anomaly handling: If there are no available paths in the backup path library, the system triggers a degradation process, first reducing the bandwidth quota for non-critical services, releasing network resources to build temporary transmission paths for critical services, and at the same time sending an emergency alarm to the operation and maintenance personnel.
[0199] After the switchover is complete, the faulty link is marked as isolated, prohibiting new service flows from accessing it, and increasing the internal risk value of its associated devices. This update will serve as input to refresh the parameters of the link evaluation function (such as device risk value weights), driving subsequent optimization decisions to focus more on this high-risk area, thereby achieving closed-loop risk management.
[0200] Figure 5 This is a schematic diagram of a link alarm switching architecture provided in an embodiment of this application. Please refer to... Figure 5 The link alarm switching process begins, with the hierarchical monitoring system triggering link alarms. Alarm scenarios include port downtime, persistent latency exceeding limits, and excessively high bit error rate. A comprehensive change demand index is calculated based on multi-dimensional monitoring data and compared with historical security thresholds to determine if a link switch is necessary. If a switch is required, alternative paths corresponding to the affected critical services are retrieved from the daily updated backup path database. Probe messages are sent to the alternative paths to verify their connectivity, bandwidth capacity, and fault-free status. Test results are categorized as pass or fail. If the test result is pass: a secondary path selection is triggered, returning to the "alternative path retrieval" stage to reselect a path; if the test result is fail: the next step, traffic migration, begins.
[0201] By adjusting the hash value or routing weight of the service flow, traffic is migrated smoothly to alternative paths in batches. This phased traffic migration avoids congestion on new paths caused by sudden traffic surges. After reclaiming old paths and continuously monitoring the stable operation of new paths for a preset period, the service binding of the original faulty link is removed, and it is marked as isolated. The risk values of the associated devices on the faulty link are updated, and the optimized topology status is fed back to the hierarchical monitoring system, refreshing the parameters of the link evaluation function. The link alarm switching process ends, ensuring that services resume normal operation.
[0202] It is worth noting that during operation, a full-domain visual monitoring and control page can be built based on the platform management interface. All layered monitoring data, optimization plan outputs, task execution records, and link self-healing handling records of the platform can be archived and retained through interface screenshots, system operation logs, and backend databases, supporting full-process traceability and verification, and operation and maintenance compliance audit.
[0203] First, the topology monitoring dashboard enables real-time visualization and perception of the entire network status. The topology monitoring dashboard can include at least one of the following sub-panels: a three-level hierarchical monitoring display panel, a link degradation visualization bar chart, and a dynamic weight monitoring panel.
[0204] The three-level hierarchical monitoring display panel can display three core indicators in real time: physical layer device health status, transmission layer link load rate, and service layer SLA compliance status through the hierarchical monitoring system. It can also simultaneously render historical time-series trend curves of the indicators to support the overall network situation analysis.
[0205] The link degradation visualization bar chart is based on the link evaluation function and uses a hierarchical color gradient to indicate the health level of all links in the network. High-risk degraded links are highlighted in red, and normal healthy links are highlighted in green, which can quickly locate potential fault points.
[0206] The dynamic weight monitoring panel can display the dynamic weight values of entropy weight across all dimensions of link evaluation and the time-series fluctuation trend of weight in real time. Combined with the network-wide operation data, it can automatically determine and prompt the current core bottleneck indicators of the network.
[0207] Secondly, by generating structured optimization solution reports, the basis for optimization decisions can be consolidated, including at least one of the following:
[0208] The original indicators and comprehensive evaluation scores of the four evaluation dimensions of each link are displayed one by one to complete the root cause analysis of link degradation; single link adaptation and optimization operations are matched, the expected optimization benefits and implementation costs are marked, and the optimization operations of the whole network are sorted in descending order based on the benefit-cost ratio; the cost budget utilization rate of the optimization plan, the redundancy of node ports, and the predicted results of business SLA compliance rate are summarized to verify the compliance and feasibility of the plan.
[0209] In addition, the full-process execution status monitoring panel can cover the entire lifecycle management of optimized operations, including at least one of the following functional modules:
[0210] Equipped with a progress bar component, it intuitively presents the execution progress and stage execution results of the entire three-stage process, including pre-verification, atomic transaction execution, and post-constraint verification; it archives and optimizes the entire process record of operation issuance time, controlled network devices, issued configuration instructions, stage execution results, and abnormal rollback, enabling traceability of configuration changes; it collects task circuit breaker trigger information, including constraint violation categories, violation parameter details, automated rollback execution ledger, and abnormal fault analysis conclusions, supporting emergency response for operation and maintenance.
[0211] Finally, the self-healing event archive logs can be used to retain all self-healing data based on the self-healing closed-loop mechanism.
[0212] Specifically, archive the comprehensive change requirement index of the faulty link, the safety threshold for fault judgment, the self-healing trigger time, and the root cause location results of the link anomaly; record the progress of the faulty traffic migration in batches, the switchover completion status, the time taken for the entire network path switchover, and the service impact coverage of this fault; retain the operation log of faulty link isolation, unblocking and recovery, archive the details of this self-healing backup path call, and output subsequent special optimization and rectification suggestions based on the link risk level.
[0213] It is worth noting that the embodiments of this application can also be applied to other suitable business scenarios, and this application does not impose specific limitations on the application scenarios.
[0214] For example, this is applicable to data centers where core financial trading systems such as securities and futures are located. This scenario has extremely high requirements for latency jitter, business continuity, and timely fault switching, while having relatively low cost requirements. It includes monitoring layer adaptation, evaluation layer adaptation, self-healing layer adaptation, and optimization layer adaptation.
[0215] Among them, the monitoring layer adaptation will adjust the collection cycle to 50ms / time, focusing on improving the collection frequency of physical layer port status, transmission layer latency, and bit error rate indicators, while enabling packet-by-packet monitoring mode for core transaction business flows.
[0216] When the evaluation layer adapts to calculate weights based on information entropy, a minimum weight limit of 0.4 is set for the latency deviation rate indicator to ensure that the latency indicator is always the core evaluation dimension; the adjustment coefficient of the dynamic degradation threshold is set between 1.2 and 1.5, adopting a more stringent screening standard to identify potential degradation links in advance.
[0217] In the self-healing layer adaptation comprehensive change requirement index, the weight of the port status factor is set to 0.4, the weight of the latency deviation rate is set to 0.35, the time window verification is closed in hard fault scenarios, and a forced switch is directly triggered when the port crashes; the proportion of traffic migration per batch is adjusted to 50%, and the interval between adjacent batches is shortened to 200ms to further accelerate the switchover speed.
[0218] When optimizing the topology, the business SLA constraints are set to the highest priority, the upper limit of cost constraints is relaxed, and the end-to-end latency and path redundancy requirements of the core transaction business are guaranteed first.
[0219] For example, this is applicable to ultra-large-scale intelligent computing centers for artificial intelligence (AI) training, big data computing, etc. This scenario requires high throughput traffic, large node scale, and many links, and has high requirements for network bandwidth utilization and computational efficiency of topology optimization. This scenario can include monitoring layer adaptation, evaluation layer adaptation, optimization layer adaptation, and execution layer adaptation. The four-layer architecture naming is uniformly adopted in the whole text, the independent execution layer naming is deleted, and batch execution control is incorporated into the optimization layer's subordinate control logic.
[0220] Among them, the monitoring layer adaptation adjusts the collection cycle to 1 second / time, focusing on collecting link load rate, bandwidth stability and packet loss rate indicators. It adopts a partitioned collection mode, dividing the monitoring area according to the Performance Optimization Domain (POD) to reduce the collection performance overhead under large-scale topology.
[0221] When the evaluation layer adapts to calculate weights based on information entropy, it sets a minimum weight limit of 0.35 for link load and bandwidth stability indicators; the adjustment coefficient for high-throughput, high-traffic general data center scenarios is set between 1.6 and 2.0 to reduce unnecessary optimization operations and lower the optimization calculation frequency under large-scale topologies; at the same time, it adopts a regional evaluation mode to perform fine-grained evaluation only on links in high-load areas to improve computational efficiency.
[0222] When the optimization layer adapts to multi-objective optimization, bandwidth stability and network-wide load balancing are set as core optimization objectives. During iterative verification, a batch processing mode is adopted, and only multiple optimization operations in the same region are processed in a single iteration to improve solution efficiency. Batch execution control rules are also incorporated: for batch optimization operations, a batch atomic execution mode is adopted, and only operations within the same POD are executed in the same batch to avoid network-wide oscillations caused by cross-regional configuration conflicts.
[0223] The execution layer is adapted for batch optimization operations, and adopts a batch atomic execution mode. The same batch only executes operations within the same POD, avoiding network-wide turbulence caused by cross-region configuration conflicts.
[0224] Figure 6 This is a schematic diagram of a network topology optimization device provided in an embodiment of this application. Please refer to... Figure 6 The network topology optimization device 600 includes an acquisition module 601, a determination module 602, a first processing module 603, and a second processing module 604.
[0225] The acquisition module 601 is used to acquire the physical layer hardware parameters, transport layer traffic indicators, and service layer constraints of each network node in the network topology.
[0226] The determination module 602 is used to determine multiple links to be optimized based on the physical layer hardware parameters, transport layer traffic indicators and service layer constraints of each network node.
[0227] The first processing module 603 is used to perform optimization processing on multiple links to be optimized and the optimization execution operations corresponding to each link to be optimized, to obtain an optimization strategy set corresponding to the network topology, the optimization strategy set including the target execution operations corresponding to each link to be optimized;
[0228] The second processing module 604 is used to perform network topology optimization based on the target execution operation corresponding to each link to be optimized.
[0229] In some embodiments, the determining module 602 is specifically used for:
[0230] Based on the physical layer hardware parameters, transport layer traffic indicators and service layer constraints of each network node, multiple single-link evaluation values are determined for multiple single links. A single link is a link between two consecutive network nodes in the network topology.
[0231] Based on the single-link evaluation value corresponding to each single link, multiple links to be optimized are identified among the multiple single links.
[0232] In some embodiments, a single link includes a first network node and a second network node, wherein the first network node is an upstream node of the second network node; for any single link; the determining module 602 is specifically used for:
[0233] The physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to the first and second network nodes are standardized to obtain standardized dimensionless feature vectors.
[0234] Based on the dimensionless eigenvector of the first network node and the dimensionless eigenvector of the second network node, the single-link evaluation value corresponding to a single link is determined.
[0235] In some embodiments, the determining module 602 is specifically used for:
[0236] Based on the dimensionless feature vector of the first network node and the dimensionless feature vector of the second network node, determine multiple feature indicators corresponding to a single link.
[0237] Obtain the weights of each feature indicator;
[0238] The sum of the products of each feature index and its corresponding index weight is used to determine the single-link evaluation value for a single link.
[0239] In some embodiments, the multiple characteristic indicators include port availability, link load, link risk value, and latency deviation rate; the determination module 602 is specifically used for:
[0240] Based on the dimensionless feature vector of the first network node, determine the current number of available ports, the maximum number of ports, and the device risk value corresponding to the first network node;
[0241] Based on the dimensionless feature vector of the second network node, determine the current number of available ports, the maximum number of ports, and the device risk value corresponding to the second network node;
[0242] Determine the link load and real-time latency corresponding to a single link;
[0243] Based on the number of currently available ports and the maximum number of ports corresponding to the first network node, and the number of currently available ports corresponding to the second network node, determine the port availability index for a single link;
[0244] For any single link, the link risk value of the single link is determined based on the device risk value corresponding to the first network node and the device risk value corresponding to the second network node.
[0245] Obtain the preset service latency for each single link, and determine the latency deviation rate for each single link based on the preset service latency and real-time link latency.
[0246] In some embodiments, the determining module 602 is specifically used for:
[0247] The current stability threshold is determined based on the single-link evaluation value corresponding to each single link;
[0248] A single link whose evaluation value is greater than the current stable threshold is identified as an abnormal link;
[0249] Based on the single-link evaluation value corresponding to each single link, determine the rate of change of the evaluation value of each single link within the preset time window;
[0250] Single links whose rate of change in the assessed value exceeds the rate of change threshold are identified as links with surging risk.
[0251] Abnormal links and links with surging risks were identified as the first candidate links;
[0252] Based on the first candidate link, a second candidate link is determined. The second candidate link is the downstream link of the risk node associated with the first candidate link.
[0253] The first and second candidate links are identified as links to be optimized, resulting in multiple links to be optimized.
[0254] In some embodiments, the first processing module 603 is specifically used for:
[0255] Multiple optimization operations are identified, which are used to optimize the link to be optimized.
[0256] For any link to be optimized, determine the optimization benefit ratio of the link after performing each optimization operation, so as to obtain multiple optimization benefit ratios for the link to be optimized;
[0257] The set of optimization strategies is determined based on the multiple optimization benefit ratios corresponding to each link to be optimized.
[0258] In some embodiments, an optimization operation is performed for any one of the optimization operations; the first processing module 603 is specifically used for:
[0259] Determine the latency reduction, maximum bit error rate reduction, and minimum bandwidth stability improvement of the link to be optimized after performing each optimization operation;
[0260] Obtain the latency weight corresponding to the latency reduction, the bit error rate weight corresponding to the maximum bit error rate reduction, and the bandwidth stability weight corresponding to the minimum bandwidth stability improvement.
[0261] The sum of the products of latency reduction and latency weight, maximum bit error rate reduction and bit error rate weight, and minimum bandwidth stability improvement and bandwidth stability weight is determined as the optimization benefit.
[0262] The ratio of the optimization benefit to the implementation cost corresponding to the optimization execution operation is determined as the optimization benefit ratio.
[0263] In some embodiments, the first processing module 603 is specifically used for:
[0264] A multi-objective optimization model is constructed based on multiple optimization benefit ratios corresponding to each link to be optimized.
[0265] Based on a multi-objective optimization model, a set of optimization strategies corresponding to multiple links to be optimized is generated.
[0266] The multi-objective optimization model has the following multi-objectives: minimizing total latency, minimizing maximum bit error rate, and maximizing minimum bandwidth stability score. The constraints are cost budget, physical feasibility, and business requirements.
[0267] In some embodiments, the second processing module 604 is specifically used for:
[0268] The target operations corresponding to each link to be optimized are verified to obtain the verification results.
[0269] If the verification result is successful, then the corresponding tasks to be executed for the target execution operation of each link to be optimized will be performed.
[0270] Execute each pending task and monitor the constraints during execution. If the constraints are violated, execute an interrupt operation.
[0271] The model deployment device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.
[0272] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7As shown, the electronic device 700 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the electronic device 700 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus.
[0273] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to execute the above-described network topology optimization method embodiment.
[0274] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0275] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0276] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0277] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0278] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the network topology optimization method embodiments described above.
[0279] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0280] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described network topology optimization method embodiments.
[0281] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described network topology optimization method embodiments.
[0282] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0283] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0284] The above provides a detailed description of a network topology optimization method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A network topology optimization method, characterized in that, include: Obtain the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints for each network node in the network topology. Based on the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints of each network node, multiple links to be optimized are identified. The optimization process is performed on the plurality of links to be optimized and the optimization execution operations corresponding to each link to be optimized to obtain an optimization strategy set corresponding to the network topology. The optimization strategy set includes the target execution operations corresponding to each link to be optimized. Based on the target operation corresponding to each link to be optimized, the network topology is optimized.
2. The method according to claim 1, characterized in that, Based on the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints corresponding to each network node, several links to be optimized are identified, including: Based on the physical layer hardware parameters, transport layer traffic indicators and service layer constraints of each network node, multiple single link evaluation values are determined for multiple single links, where a single link is a link between two consecutive network nodes in the network topology. Based on the single-link evaluation value corresponding to each single link, the plurality of links to be optimized are determined among the plurality of single links.
3. The method according to claim 2, characterized in that, The single link includes a first network node and a second network node, where the first network node is the upstream node of the second network node; this applies to any single link. Based on the physical layer hardware parameters, transport layer traffic metrics, and service layer constraints corresponding to each network node, multiple single-link evaluation values are determined for multiple single links, including: The physical layer hardware parameters, transport layer traffic indicators, and service layer constraints corresponding to the first network node and the second network node are standardized to obtain standardized dimensionless feature vectors. The single-link evaluation value corresponding to the single link is determined based on the dimensionless feature vector of the first network node and the dimensionless feature vector of the second network node.
4. The method according to claim 3, characterized in that, Based on the dimensionless feature vector of the first network node and the dimensionless feature vector of the second network node, the single-link evaluation value corresponding to the single link is determined, including: Based on the dimensionless feature vector of the first network node and the dimensionless feature vector of the second network node, determine multiple feature indicators corresponding to the single link; Obtain the weights of each feature indicator; The sum of the products of each feature index and its corresponding index weight is determined as the single-link evaluation value corresponding to the single link.
5. The method according to claim 3, characterized in that, The multiple feature indicators include port availability, link load, link risk value, and latency deviation rate; based on the dimensionless feature vector of the first network node and the dimensionless feature vector corresponding to the second network node, the multiple feature indicators corresponding to the single link are determined, including: Based on the dimensionless feature vector of the first network node, determine the current number of available ports, the maximum number of ports, and the device risk value corresponding to the first network node; Based on the dimensionless feature vector of the second network node, determine the current number of available ports, the maximum number of ports, and the device risk value corresponding to the second network node; Determine the link load and real-time latency corresponding to the single link; The port availability index corresponding to the single link is determined based on the number of currently available ports and the maximum number of ports corresponding to the first network node and the number of currently available ports corresponding to the second network node. For any single link, the link risk value of the single link is determined based on the device risk value corresponding to the first network node and the device risk value corresponding to the second network node; Obtain the preset service latency corresponding to each single link, and determine the latency deviation rate corresponding to each single link based on the preset service latency and real-time link latency.
6. The method according to claim 2, characterized in that, Based on the single-link evaluation value corresponding to each single link, the plurality of links to be optimized are determined from the plurality of single links, including: Based on the single-link evaluation value corresponding to each single link, determine the current stability threshold; A single link whose single-link evaluation value is greater than the current stability threshold is identified as an abnormal link; Based on the single-link evaluation value corresponding to each single link, determine the rate of change of the evaluation value of each single link within a preset time window; A single link whose rate of change of the assessed value is greater than the rate of change threshold is identified as a link with a surge in risk. The abnormal link and the link with the surge in risk were identified as the first candidate links; Based on the first candidate link, a second candidate link is determined, and the second candidate link is the downstream link of the risk node associated with the first candidate link; The first candidate link and the second candidate link are determined as the links to be optimized, so as to obtain the plurality of links to be optimized.
7. The method according to claim 1, characterized in that, The optimization processes are performed on the plurality of links to be optimized and the optimization operations corresponding to each link to be optimized to obtain a set of optimization strategies corresponding to the network topology, including: Multiple optimization execution operations are determined, which are used to optimize the link to be optimized; For any link to be optimized, determine the optimization benefit ratio of the link to be optimized after performing each optimization operation, so as to obtain multiple optimization benefit ratios corresponding to the link to be optimized; The set of optimization strategies is determined based on the multiple optimization benefit ratios corresponding to each link to be optimized.
8. The method according to claim 7, characterized in that, For any given optimization operation, determine the optimization benefit ratio of the link to be optimized after performing each optimization operation, including: Determine the latency reduction, maximum bit error rate reduction, and minimum bandwidth stability improvement of the link to be optimized after performing each optimization operation; Obtain the latency weight corresponding to the latency reduction amount, the bit error rate weight corresponding to the maximum bit error rate reduction value, and the bandwidth stability weight corresponding to the minimum bandwidth stability improvement value; The sum of the product of the latency reduction and the latency weight, the product of the maximum bit error rate reduction and the bit error rate weight, and the product of the minimum bandwidth stability improvement and the bandwidth stability weight is determined as the optimization benefit. The ratio of the optimization benefit to the implementation cost corresponding to the optimization execution operation is determined as the optimization benefit ratio.
9. The method according to claim 7, characterized in that, Based on the multiple optimization benefit ratios corresponding to each link to be optimized, the set of optimization strategies is determined, including: Based on the multiple optimization benefit ratios corresponding to each link to be optimized, a multi-objective optimization model is constructed. Based on the multi-objective optimization model, an optimization strategy set corresponding to the multiple links to be optimized is generated; The multi-objective optimization model has the following multi-objectives: minimizing total latency, minimizing maximum bit error rate, and maximizing minimum bandwidth stability score. The constraints are cost budget, physical feasibility, and business requirements.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-9.