Integrated computing network computing power scheduling method and integrated computing network computing power scheduling system
By constructing joint computing and network capability information and stability scores, an integrated scheduling scheme is generated, which solves the problem of separating computing resources from network planning and achieves precise end-to-end scheduling and self-optimization capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, computing resource planning is separated from network bandwidth and latency planning, lacking a unified end-to-end perspective, resulting in inaccurate scheduling and difficulty in identifying anomalies and eliminating unstable combinations.
By constructing information on the combined computing and network capabilities, setting stability scores, generating integrated scheduling solutions based on business needs, monitoring for anomalies during operation, performing targeted scheduling adjustments, updating stability scores, and eliminating unstable combinations.
It achieves unified assessment of end-to-end capabilities, improves the accuracy and stability of scheduling, reduces the probability of scheduling unstable combinations, and has self-optimization capabilities with self-learning ability.
Smart Images

Figure CN121907932A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computing power network scheduling technology, specifically an integrated computing power network scheduling method and an integrated computing power network scheduling system. Background Technology
[0002] With the development of cloud computing, edge computing, and computing networks, businesses are increasingly demanding higher levels of comprehensive requirements for end-to-end latency, bandwidth assurance, reliability, and computing resource utilization. To meet the needs of different types of businesses, operators typically deploy several computing nodes in multiple regions and connect them through multiple network paths.
[0003] In existing technologies, computing power scheduling and network scheduling are often performed by different systems: One type of system focuses on the allocation of computing resources, such as determining which nodes to deploy business instances based on the processing power, storage capacity, and current load of the computing nodes; Another type of system focuses on the allocation of network resources, such as selecting paths and reserving bandwidth for service traffic based on link bandwidth, latency, and packet loss rate.
[0004] In this type of approach, computing resource planning is typically separated from network bandwidth and latency planning, with only simple interfacing at the policy level. This approach has the following problems: The lack of a unified end-to-end perspective on computing power and network capabilities makes it impossible to evaluate the comprehensive capabilities of business scenarios, computing nodes, and network paths based on actual business operation results.
[0005] The lack of fine-grained anomaly attribution during operation makes it difficult to effectively distinguish between computing power anomalies, network anomalies, and both anomalies simultaneously, resulting in inaccurate scheduling adjustments.
[0006] There is no clear elimination mechanism for long-term unstable business scenarios, computing nodes and network path combinations, which can easily lead to the problem that some combinations are statistically unstable but are still repeatedly scheduled and used. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an integrated computing power scheduling method for computing networks, comprising: Step 1: In at least one preset test service scenario, the test service is scheduled to multiple computing power nodes and transmitted through multiple network paths to obtain basic computing network information and operation indicators of different service scenarios, computing power nodes and network path combinations. The computing network joint capability information is constructed, and an initial value for stability score is set for each service scenario, computing power node and network path combination. Step 2: Receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. Step 3: Determine the candidate set of computing power nodes and the candidate set of network paths based on the target computing network requirements, computing network joint capability information, and stability score; Step 4: Based on the joint computing and network capability information and the preset integrated computing and network scheduling rules, perform joint resource allocation on the candidate computing power node set and the candidate network path set to generate an integrated computing and network computing power scheduling scheme. The integrated computing and network computing power scheduling scheme includes at least the target computing power node and the target network path. Step 5: The integrated computing power scheduling scheme is distributed to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. Step 6: Collect computing network operation status information during the operation of the target business, determine the expected operation indicators based on the computing network joint capability information, compare the computing network operation status information with the expected operation indicators, and determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. Step 7: Perform computing power scheduling adjustment operations according to the anomaly category. Update the computing network joint capability information based on the difference between the computing network operation status information and the expected operation indicators before and after the computing network scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, mark the combination of business scenario, computing power node and network path as a disabled combination. Disabled combinations do not participate in the determination of candidate computing power node set and candidate network path set.
[0008] Furthermore, in at least one preset test service scenario, the test service is scheduled to multiple computing power nodes and transmitted through multiple network paths to obtain basic computing network information and operational indicators of different service scenarios, computing power node and network path combinations, construct computing network joint capability information, and set initial stability scores for each service scenario, computing power node and network path combination, including: The basic information of the computing network includes computing resource information of computing nodes and network resource information of network links. The computing resource information includes computing node identification, processing capacity, storage capacity and current computing load. The network resource information includes link bandwidth limit, link latency, historical packet loss rate and link cost information. The operation indicators include end-to-end latency, throughput, computing power utilization and bandwidth utilization. Under at least one preset test service scenario, the test service is scheduled to multiple computing nodes and transmitted through multiple network paths. The operation indicators of different service scenarios, computing nodes and network path combinations are collected. The correspondence between service scenarios, computing nodes, network paths and operation indicators is stored as computing network joint capability information. An initial value for stability score is set for each service scenario, computing node and network path combination.
[0009] Furthermore, the process of receiving service requests and determining target computing network requirements based on these requests includes the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service, including: Parse the service type, service access location, target end-to-end latency limit, target bandwidth requirement, and service reliability requirement from the service request; determine the computing power type requirement and computing power scale requirement based on the service type and service reliability requirement; combine the service access location, target end-to-end latency limit, and target bandwidth requirement into network-side target constraints; and combine the computing power type requirement, computing power scale requirement, and network-side target constraints into target computing network requirements.
[0010] Furthermore, the determination of the candidate computing power node set and candidate network path set based on target computing network requirements, computing network joint capability information, and stability scores includes: Based on the joint computing and network capability information, computing power nodes that meet the target computing network requirements in terms of computing power type and computing power scale are selected from all computing power nodes to form a candidate computing power node set. For each candidate computing power node in the candidate computing power node set, based on the network-side target constraints, joint computing and network capability information, business scenarios, and the stability score of the combination of computing power nodes and network paths, network paths that meet the target computing network requirements in terms of end-to-end latency and bandwidth and whose stability score is not lower than the first threshold are selected from all network paths to form a candidate network path. The candidate computing power nodes are combined with their respective candidate network paths to form a candidate computing power node set and a candidate network path set.
[0011] Furthermore, based on the joint computing and network capability information and the preset integrated computing and network scheduling rules, the candidate computing power node set and candidate network path set are jointly allocated resources to generate an integrated computing and network computing power scheduling scheme. The integrated computing and network computing power scheduling scheme includes at least target computing power nodes and target network paths, including: According to the preset integrated computing and network scheduling rules, computing power resource quotas and network bandwidth quotas are allocated to each combination of computing power node and network path in the candidate computing power node set and candidate network path set. The preset integrated computing and network scheduling rules comprehensively weigh computing power resource utilization, network resource utilization, stability score and scheduling cost under the premise of meeting the target computing network requirements. Based on the operation indicators and stability scores of each combination recorded in the computing and network joint capability information, an end-to-end capability evaluation is performed on all combinations. The combination with the best comprehensive evaluation result is selected as the target computing power node and target network path, and an integrated computing power scheduling scheme including the target computing power node, target network path and resource quota is formed.
[0012] Furthermore, the integrated computing power scheduling scheme is distributed to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path, including: The system issues computing power scheduling instructions to the computing power orchestration unit. These instructions include the target computing power node identifier, the service instance mirror identifier, the number of service instances, and the computing power resource quota for each service instance. The computing power orchestration unit then creates service instances on the target computing power node and completes the reservation of computing power resources. The system also issues network configuration instructions to the network control unit. These instructions include the target network path identifier, the network link bandwidth reservation value, and the service traffic forwarding identifier. The network control unit then configures the forwarding path and bandwidth reservation in the network devices on the target network path to carry the target service traffic.
[0013] Furthermore, the process of collecting computing network operation status information during the operation of the target service, determining expected operation indicators based on computing network joint capability information, comparing computing network operation status information with expected operation indicators, and determining computing network operation anomalies and their categories include: Periodically collect computing power side operation status information, including processor utilization, storage utilization, and service instance response latency of the target computing power node; periodically collect network side operation status information, including end-to-end latency, link bandwidth utilization, and packet loss rate of the target network path; read operation indicators corresponding to the current business scenario, target computing power node, and target network path from the computing-network joint capability information, determine the expected end-to-end latency, expected computing power utilization, and expected link bandwidth utilization as expected operation indicators, compare the computing network operation status information with the expected operation indicators, and determine computing network operation anomalies and the anomaly categories of computing network operation anomalies.
[0014] Furthermore, the aforementioned anomaly categories include a first anomaly category, a second anomaly category, and a third anomaly category, specifically as follows: The first anomaly category indicates that at least one computing power-side indicator in the computing power-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator, and at least one network-side indicator in the network-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator; the second anomaly category indicates that all computing power-side indicators in the computing power-side operation status information are within the preset deviation range of the corresponding expected operation indicator, and at least one network-side indicator in the network-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator; the third anomaly category indicates that all network-side indicators in the network-side operation status information are within the preset deviation range of the corresponding expected operation indicator, and at least one computing power-side indicator in the computing power-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator.
[0015] Furthermore, the process of performing computing power scheduling adjustment operations based on anomaly categories, updating the computing network joint capability information based on the differences between the computing network operating status information before and after the computing power scheduling adjustment operations and the expected operating indicators, and updating the stability score of the combination of business scenarios, computing power nodes, and network paths, includes: For the first anomaly category, a joint rescheduling operation is performed to migrate the target service to a new target computing power node and switch to a new target network path. Based on the new target computing power node and the new target network path, new operating indicators and new stability scores are recorded in the computing-network joint capability information. For the second abnormal category, while keeping the target computing power node unchanged, adjust the target network path or adjust the network bandwidth quota on the target network path, and correct the operation indicators and stability scores corresponding to the combination of target computing power node and target network path in the computing network joint capability information based on the adjusted computing network operation status information; For the third anomaly category, while keeping the target network path unchanged, the computing power resource quota of the target computing power node is adjusted or the number of service instances is increased within the same computing network scheduling domain. The computing network scheduling domain is used to characterize the set of computing power resources uniformly scheduled by the integrated computing network computing power scheduling method. Based on the adjusted computing network operation status information, the operation indicators and stability scores corresponding to the combination of target computing power node and target network path in the computing network joint capability information are corrected. When the stability score of any service scenario, computing power node and network path combination is lower than the second threshold, the service scenario, computing power node and network path combination is marked as a disabled combination, and the disabled combination is excluded when determining the candidate computing power node set and candidate network path set in step three.
[0016] An integrated computing network power scheduling system, employing the aforementioned integrated computing network power scheduling method, includes: a computing network joint capability information construction module, a target computing network demand determination module, a candidate resource determination module, a scheduling scheme generation module, a scheduling scheme execution module, an operation monitoring and anomaly judgment module, a scheduling adjustment and capability update module, and a data processing module; the computing network joint capability information construction module, the target computing network demand determination module, the candidate resource determination module, the scheduling scheme generation module, the scheduling scheme execution module, the operation monitoring and anomaly judgment module, and the scheduling adjustment and capability update module are respectively connected to the data processing module; The computing network joint capability information construction module is used to acquire basic computing network information and operational indicators of different business scenarios, computing power nodes and network path combinations under at least one preset test business scenario, construct computing network joint capability information, and set initial stability score values for each business scenario, computing power node and network path combination. The target computing network requirement determination module is used to receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. The candidate resource determination module is used to determine the set of candidate computing power nodes and the set of candidate network paths based on the target computing network requirements, computing network joint capability information and stability score; The scheduling scheme generation module is used to perform joint resource allocation on the candidate computing power node set and the candidate network path set based on the computing network joint capability information and the preset integrated computing network scheduling rules, and generate an integrated computing network computing power scheduling scheme. The integrated computing network computing power scheduling scheme includes at least the target computing power node and the target network path. The scheduling scheme execution module is used to distribute the integrated computing network computing power scheduling scheme to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. The aforementioned operation monitoring and anomaly determination module is used to collect computing network operation status information during the operation of the target service, determine expected operation indicators based on computing network joint capability information, and compare computing network operation status information with expected operation indicators to determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. The scheduling adjustment and capability update module is used to perform computing power scheduling adjustment operations according to the anomaly category, update the computing power joint capability information based on the difference between the computing power operation status information and the expected operation indicators before and after the computing power scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, the combination of business scenario, computing power node and network path is marked as a disabled combination, and the disabled combination is excluded when the candidate resource determination module determines the candidate computing power node set and the candidate network path set.
[0017] The beneficial effects of this invention are: During the testing phase, operational metrics mappings are constructed for combinations of business scenarios, computing nodes, and network paths to form computing-network joint capability information, providing a unified end-to-end capability reference for subsequent scheduling and monitoring.
[0018] A stability score is introduced to quantify the performance of the combination over a longer time period. Candidate combinations are screened by a first threshold, and combinations that frequently cause anomalies are eliminated by a second threshold and a combination disabling mechanism, thereby reducing the probability of scheduling to unstable combinations from the root.
[0019] Based on the deviation between expected operating indicators and actual operating status, anomalies are classified into three categories: computing power anomalies, network anomalies, and simultaneous anomalies on both the computing power and network sides. Different scheduling adjustment strategies are implemented for different anomaly categories to enhance the targeting of scheduling adjustments.
[0020] By using the changes in deviations before and after scheduling adjustments to update the network's joint capability information and stability score, the system gains self-learning ability and can automatically optimize the available combination set and scheduling strategy as runtime increases. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating an integrated computing power scheduling method for computing networks. Figure 2 A schematic diagram of the information flow for constructing the computing network joint capability. Detailed Implementation
[0022] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0023] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0024] Example 1: like Figure 1 As shown, an integrated computing power scheduling method for computing networks includes: Step 1: In at least one preset test service scenario, the test service is scheduled to multiple computing power nodes and transmitted through multiple network paths to obtain basic computing network information and operation indicators of different service scenarios, computing power nodes and network path combinations. The computing network joint capability information is constructed, and an initial value for stability score is set for each service scenario, computing power node and network path combination. Step 2: Receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. Step 3: Determine the candidate set of computing power nodes and the candidate set of network paths based on the target computing network requirements, computing network joint capability information, and stability score; Step 4: Based on the joint computing and network capability information and the preset integrated computing and network scheduling rules, perform joint resource allocation on the candidate computing power node set and the candidate network path set to generate an integrated computing and network computing power scheduling scheme. The integrated computing and network computing power scheduling scheme includes at least the target computing power node and the target network path. Step 5: The integrated computing power scheduling scheme is distributed to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. Step 6: Collect computing network operation status information during the operation of the target business, determine the expected operation indicators based on the computing network joint capability information, compare the computing network operation status information with the expected operation indicators, and determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. Step 7: Perform computing power scheduling adjustment operations according to the anomaly category. Update the computing network joint capability information based on the difference between the computing network operation status information and the expected operation indicators before and after the computing network scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, mark the combination of business scenario, computing power node and network path as a disabled combination. Disabled combinations do not participate in the determination of candidate computing power node set and candidate network path set.
[0025] Specifically, the computing network control platform connects to various computing nodes and network devices. Computing nodes can be data center servers, edge nodes, or other devices with computing capabilities. Network devices can be switches, routers, or other data forwarding devices.
[0026] The computing orchestration unit is responsible for creating, destroying, and scaling up / down service instances on specific computing nodes. The network control unit is responsible for issuing path configurations, bandwidth reservations, and routing policies to ensure the transmission of service traffic in the network. A service instance represents a computing unit created and running on a computing node according to a service image. A service instance can be a virtual machine, container, process, or other independently schedulable running entity, and each service instance can independently process data from service requests. The number of service instances represents the number of service instances created for the same service on the target computing node; increasing the number of service instances can improve the service's concurrent processing capabilities and reliability.
[0027] Basic information about the computing network includes: Computing resource information for computing nodes: This includes the computing node identifier, processing capacity, storage capacity, and current computing load. Processing capacity can be described based on the number of processing units, processing frequency, or other levels used to characterize computing power.
[0028] The current computing load can be described by statistically analyzing the proportion of processing capacity and storage capacity allocated to business instances to the total capacity within a fixed time period.
[0029] Network link resource information includes link bandwidth limit, link latency, historical packet loss rate, and link cost information. Link bandwidth limit represents the maximum amount of data a link can transmit per unit time; link latency represents the time required for data to travel from the link ingress to the link egress; historical packet loss rate can be determined by selecting a historical time period and calculating the ratio of lost data packets to the total number of transmitted data packets within that period; link cost information can be pre-configured according to operational strategies and used to evaluate costs during scheduling.
[0030] Basic information about the computing network can be retrieved periodically from the computing power management system and the network management system, or it can be proactively reported by computing power nodes and network devices.
[0031] Preset test scenarios are used to simulate different types of business. Each scenario includes: a business type identifier, such as video processing, online inference, batch processing tasks, etc.; a business request scale, such as the number of requests per second or the amount of data input; and a target end-to-end latency range and a target bandwidth range, used to set expected values during the testing phase.
[0032] For ease of explanation, this specification introduces the concept of similar business scenarios. Similar business scenarios can be understood as business scenarios with the same service type and target end-to-end latency and target bandwidth ranges located within pre-defined adjacent intervals. For example, end-to-end latency requirements and bandwidth requirements can be divided into multiple discrete intervals. When two business scenarios fall into the same service type and adjacent or identical latency and bandwidth intervals, the two business scenarios can be considered similar business scenarios.
[0033] Under at least one preset test service scenario, the computing network control platform performs the following steps: Select several computing nodes, ensuring that these nodes are connected to the service access location via multiple network paths. In a specific test service scenario, the test service is sequentially scheduled to each computing node, and test service instances are run on that computing node.
[0034] For each computing node, multiple different network paths are selected from the service access location to the computing node, and test services are run on each path to ensure that the test services run for a period of time in each combination of service scenario + computing node + network path.
[0035] like Figure 2 As shown, the following performance metrics were collected during the test run: End-to-end latency: For each service request, record the time difference from sending the request from the access location to receiving the service response. The end-to-end latency of multiple requests can be calculated within the test period. Throughput: The number of requests successfully processed within a fixed time window, or the amount of data processed. Computing power usage: The ratio of processing capacity usage and storage capacity usage of computing power nodes during the test period. For example, it can be determined based on the ratio of resources used by business instances to the total resources of computing power nodes. Bandwidth usage: The ratio between traffic usage related to the test service and the link bandwidth limit during the test period.
[0036] For each combination of business scenario, computing node, and network path, the computing network control platform stores the above-mentioned operational indicators along with the combination identifier to form a record. All records together constitute the computing network joint capability information.
[0037] This invention introduces a stability score to represent the stability of the combination of business scenario, computing power node, and network path in long-term operation.
[0038] In one implementation, the stability score can be set as an integer level, such as from level 1 to level 5: level 1 represents the lowest level of stability; level 5 represents the highest level of stability.
[0039] After collecting operational metrics during the testing phase, the initial stability score can be set for each combination according to the following logic: For combinations where the operating indicators remain within the expected range and no obvious operational failures occur during the testing process, the initial value of the stability score is set to the intermediate level, such as level 3. For combinations that slightly exceed preset expectations during testing but can be restored to normal by increasing resources, the initial stability score can be set to a slightly lower level, such as level 2. For combinations that experience multiple serious operational failures during the testing period, they can be excluded from the network joint capability information, or the initial stability score can be set to the lowest level and processed according to the disable strategy in subsequent screening.
[0040] The specific range of values and initial configuration of the stability rating level can be set by the system implementer during deployment based on business needs.
[0041] When a new service request arrives, the computing network control platform performs the following steps to parse the service request: parse the service type, used to select the corresponding service scenario or similar service scenario from the computing network joint capability information; parse the service access location, used to determine the network ingress or edge access point related to that access location; parse the target end-to-end latency limit, as a hard constraint to ensure that the end-to-end latency does not exceed the limit during scheduling; parse the target bandwidth requirement, used to ensure that the available bandwidth provided by the target network path is not lower than the requirement; and parse the service reliability requirements, used to determine the number of service instances, whether multi-path protection is required, etc.
[0042] Based on the above analysis results, the network control platform generates: Computing power type requirement indicates the type of computing power required by the business, such as general computing, graphics computing, or inference computing; computing power scale requirement indicates the scale level of computing power required by the business, such as processing capacity level or number of instances; network-side target constraints indicate network requirements related to the business access location, target end-to-end latency limit, and target bandwidth requirement.
[0043] The computing network control platform combines computing power type requirements, computing power scale requirements, and network-side target constraints into target computing network requirements.
[0044] The computing network control platform selects computing nodes from all computing power nodes that meet the following conditions: The computing power type of the computing power node is consistent with the computing power type requirement; the remaining processing capacity and storage capacity that the computing power node can provide are not less than the computing power scale requirement; in the computing network joint capability information, there is a valid record corresponding to the current business scenario or similar business scenario and the computing power node.
[0045] The computing power nodes that meet the above conditions form a candidate computing power node set.
[0046] For each candidate computing node in the candidate computing node set, the computing network control platform performs the following steps: From all reachable network paths between the service access location and the candidate computing node, select the path that meets the network-side target constraints in terms of link bandwidth and connectivity, such as the link bandwidth limit not being lower than the target bandwidth requirement.
[0047] For each candidate path, search the records in the computing network joint capability information that correspond to the current business scenario or similar business scenario, candidate computing power nodes and the combination of paths, and obtain the operation indicators and stability scores.
[0048] If the stability score has been marked as being below the second threshold, the combination is considered a disabled combination and will not participate in subsequent screening.
[0049] For combinations with a stability score not lower than the first threshold, further check whether the operating indicators meet the target computing network requirements, such as whether the end-to-end latency does not exceed the target end-to-end latency limit.
[0050] For combinations that simultaneously meet the above conditions, the corresponding network path is added to the candidate network path set, and the stability score and operational metrics of the combination are recorded.
[0051] The first threshold can be configured by the system implementer during system deployment to control the minimum stability level for candidates to participate in the screening process.
[0052] The computing network control platform, based on preset integrated computing network scheduling rules, jointly evaluates and allocates resources to the candidate computing power node set and the candidate network path set, generating an integrated computing network computing power scheduling scheme.
[0053] The scheduling rules can be implemented according to the following logic: Check hard constraints: For each candidate computing node + candidate network path combination, check whether the corresponding operational indicators meet the hard constraints, including: end-to-end latency does not exceed the target end-to-end latency limit; available bandwidth is not less than the target bandwidth requirement; computing resources can meet the computing scale requirements under the current load. Combinations that fail the hard constraint check are directly eliminated.
[0054] Calculating combination priority: For combinations that pass the hard constraint check, combination priority can be calculated as follows: First, differentiate between different stability rating levels and group combinations with higher stability ratings into higher priority groups; within the same stability rating level, assign higher priority to combinations with shorter end-to-end latency. When the end-to-end latency difference is within an acceptable range, the utilization rates of computing resources and network resources can be compared, and the combination with more balanced resource utilization and less likely to cause local congestion can be selected first. When the above indicators are still similar, the lower-cost combination can be selected based on link cost information or cross-regional scheduling cost.
[0055] Generate scheduling scheme: Through the above priority comparison, select the combination with the highest overall priority from the candidate combinations as the target computing power node and target network path, and allocate clear computing power resource quotas and network bandwidth quotas to the target computing power node and target network path to form an integrated computing network computing power scheduling scheme.
[0056] Combinatorial priority calculation does not require formula expression and can be achieved through multi-level sorting: first group by stability score, then sort by end-to-end latency within the group, and sort by resource utilization and cost when end-to-end latency is similar.
[0057] The computing network control platform distributes the integrated computing network computing power scheduling scheme to the computing power orchestration unit and the network control unit respectively. It issues computing power scheduling instructions to the computing power orchestration unit, which include: target computing power node identifier, service instance mirror identifier, number of service instances, and computing power resource quota for each service instance.
[0058] The computing power orchestration unit creates service instances on the target computing power nodes according to the computing power scheduling instructions, and reserves the corresponding processing capacity and storage capacity.
[0059] The network configuration command is issued to the network control unit. The network configuration command includes: the target network path identifier, the bandwidth reservation value of each network link on the path, and the forwarding identifier or tag used to identify service traffic.
[0060] The network control unit configures forwarding rules and bandwidth reservations in the network devices on the target network path according to the network configuration instructions, so that the target network path can carry the target service traffic according to the scheduling plan.
[0061] During the operation of the target service, the computing network control platform periodically collects computing network operation status information. The collected information includes: Information on the computing power side's operational status, such as: processor utilization of the target computing power node; storage utilization of the target computing power node; and actual response latency of the business instance.
[0062] Network-side operational status information, such as: end-to-end latency of the target network path; bandwidth utilization of each link on the target network path; packet loss rate of the target network path during the monitoring period.
[0063] To perform anomaly detection, the computing network control platform needs to determine the expected operational indicators for the current business scenario, target computing power nodes, and target network path combination. The expected operational indicators can be determined based on historical records in the computing network joint capability information as follows: Read multiple historical operation records of the combination from the computing network joint capability information.
[0064] For end-to-end latency metrics: The median end-to-end delay among the most recent records can be selected as the expected end-to-end delay. In practice, the most recent records can be sorted in ascending order, and the value at the middle of the sorted sequence can be selected as the expected end-to-end delay.
[0065] Regarding computing power usage and bandwidth usage metrics: You can select the median or typical value from the most recent records as the expected occupancy level, or you can sort the records and select the value at the middle position to determine it.
[0066] For throughput metrics: You can select the throughput value from the most recent records that represents a period of stable business operation as the expected throughput. For example, you can choose the record with the value in the middle of the most recent records as a reference.
[0067] In this way, the computing network control platform determines the expected operating indicators such as expected end-to-end latency, expected computing power consumption, expected bandwidth consumption, and expected throughput for each combination.
[0068] The computing network control platform compares the currently collected computing network operating status information with the expected operating indicators. The comparison logic may include: Set deviation threshold ranges for each operational metric. For example, for end-to-end latency, a fixed time deviation from the expected value can be set as the allowable deviation range; for computing power utilization and bandwidth utilization, a fixed percentage range above and below the expected value can be set as the allowable deviation range; for packet loss rate, a maximum allowable packet loss rate can be set. The deviation threshold ranges can be set by the system implementer during deployment.
[0069] For each indicator, determine whether the current value falls within the corresponding allowable deviation range: if it falls within the range, the indicator is considered to be in a normal state; if it exceeds the range, the indicator is considered to be in an abnormal state.
[0070] Based on the abnormalities in computing power and network indicators, the computing network control platform categorizes computing network operational anomalies into three types: First category of anomalies: at least one indicator on the computing power side is abnormal, and at least one indicator on the network side is abnormal; Second category of anomalies: All indicators on the computing power side are normal, but at least one indicator on the network side is abnormal; The third abnormal category: all network-side indicators are normal, but at least one computing power-side indicator is abnormal.
[0071] If all relevant indicators are within the range of deviation from the threshold, it is considered that there is no abnormal operation of the computing network.
[0072] After detecting an anomaly in the computing network operation, the computing network control platform executes different scheduling adjustment strategies based on the anomaly category: For the first anomaly category (simultaneous anomalies on both the computing power side and the network side): the computing network control platform can reselect a new combination from the candidate computing power node set and the candidate network path set to generate a new scheduling scheme; migrate the target service to the new target computing power node and switch the service traffic to the new target network path; after the new combination takes effect, recollect the operation status information and compare it with the expected operation indicators.
[0073] For the second anomaly category (network-side anomalies only): While keeping the target computing power node unchanged, the computing network control platform can: select a new path with a stability score not lower than the first threshold and whose operation indicators meet the target computing network requirements from other reachable paths between the service access location and the target computing power node, and replace the original target network path; or, without changing the target network path, increase the bandwidth reserved for the target service on that path to alleviate congestion.
[0074] For the third anomaly category (anomalies only on the computing power side): While keeping the target network path unchanged, the computing network control platform can: increase the computing resource quota reserved for the current service instance by the target computing power node; or, add a new service instance for the current service within the same computing network scheduling domain, such as deploying an additional instance on another computing power node and connecting that computing power node to the service access location through the target network path or other paths.
[0075] Before executing the scheduling adjustment operation, the computing network control platform has saved the comparison results between the current combination's operating status information and the expected operating indicators. After the scheduling adjustment operation is executed, the computing network control platform collects new computing network operating status information again and compares it with the same expected operating indicator.
[0076] Based on the deviation changes before and after the scheduling adjustment, the network joint capability information and stability score can be updated according to the following logic: For combinations that continue to carry services after scheduling adjustments: if the new operating status information is closer to the expected operating indicators than before the adjustment, i.e. the deviation is reduced, it can be considered that the operating stability of the combination has improved, and the stability score can be appropriately increased. For example, the score level can be increased by one level as long as the score level does not exceed the maximum level. If, after scheduling adjustments, there are still cases where the deviation from the threshold is significantly exceeded, and similar problems are observed in multiple consecutive monitoring sessions, the combination can be considered unstable in the current business scenario, and the stability rating can be lowered.
[0077] For combinations that no longer carry the current services after scheduling adjustments: if the scheduling adjustment is caused by obvious anomalies in the combination, the stability rating of the combination can be reduced according to the severity of the anomalies.
[0078] For newly introduced combinations (new target computing nodes and target network paths): an initial stability score can be set based on the performance of the initial run. For example, if the performance indicators are significantly better than the target computing network requirements, the stability score can be set to an intermediate level or above. At the same time, the performance indicators of the new combination are written into the computing network joint capability information for subsequent scheduling.
[0079] Stability rating and disabling combination: In one implementation, it can be stipulated that when a combination is repeatedly identified as having the first anomaly category, or when anomalies still occur multiple times even after scheduling adjustments, the stability score of the combination is gradually reduced. When the stability score drops to no higher than the second threshold, the computing network control platform marks the combination as a disabled combination. In the subsequent screening of candidate computing power node sets and candidate network path sets, all records marked as disabled combinations will no longer participate in the screening, thereby eliminating unstable combinations at the source of scheduling.
[0080] The specific value of the second threshold can be configured by the system implementer based on the business's sensitivity to stability.
[0081] Example 2: An integrated computing network power scheduling system, employing the aforementioned integrated computing network power scheduling method, includes: a computing network joint capability information construction module, a target computing network demand determination module, a candidate resource determination module, a scheduling scheme generation module, a scheduling scheme execution module, an operation monitoring and anomaly judgment module, a scheduling adjustment and capability update module, and a data processing module; the computing network joint capability information construction module, the target computing network demand determination module, the candidate resource determination module, the scheduling scheme generation module, the scheduling scheme execution module, the operation monitoring and anomaly judgment module, and the scheduling adjustment and capability update module are respectively connected to the data processing module; The computing network joint capability information construction module is used to acquire basic computing network information and operational indicators of different business scenarios, computing power nodes and network path combinations under at least one preset test business scenario, construct computing network joint capability information, and set initial stability score values for each business scenario, computing power node and network path combination. The target computing network requirement determination module is used to receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. The candidate resource determination module is used to determine the set of candidate computing power nodes and the set of candidate network paths based on the target computing network requirements, computing network joint capability information and stability score; The scheduling scheme generation module is used to perform joint resource allocation on the candidate computing power node set and the candidate network path set based on the computing network joint capability information and the preset integrated computing network scheduling rules, and generate an integrated computing network computing power scheduling scheme. The integrated computing network computing power scheduling scheme includes at least the target computing power node and the target network path. The scheduling scheme execution module is used to distribute the integrated computing network computing power scheduling scheme to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. The aforementioned operation monitoring and anomaly determination module is used to collect computing network operation status information during the operation of the target service, determine expected operation indicators based on computing network joint capability information, and compare computing network operation status information with expected operation indicators to determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. The scheduling adjustment and capability update module is used to perform computing power scheduling adjustment operations according to the anomaly category, update the computing power joint capability information based on the difference between the computing power operation status information and the expected operation indicators before and after the computing power scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, the combination of business scenario, computing power node and network path is marked as a disabled combination, and the disabled combination is excluded when the candidate resource determination module determines the candidate computing power node set and the candidate network path set.
[0082] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. An integrated computing power scheduling method for computing networks, characterized in that, include: Step 1: In at least one preset test service scenario, the test service is scheduled to multiple computing power nodes and transmitted through multiple network paths to obtain basic computing network information and operation indicators of different service scenarios, computing power nodes and network path combinations. The computing network joint capability information is constructed, and an initial value for stability score is set for each service scenario, computing power node and network path combination. Step 2: Receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. Step 3: Determine the candidate set of computing power nodes and the candidate set of network paths based on the target computing network requirements, computing network joint capability information, and stability score; Step 4: Based on the joint computing and network capability information and the preset integrated computing and network scheduling rules, perform joint resource allocation on the candidate computing power node set and the candidate network path set to generate an integrated computing and network computing power scheduling scheme. The integrated computing and network computing power scheduling scheme includes at least the target computing power node and the target network path. Step 5: The integrated computing power scheduling scheme is distributed to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. Step 6: Collect computing network operation status information during the operation of the target business, determine the expected operation indicators based on the computing network joint capability information, compare the computing network operation status information with the expected operation indicators, and determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. Step 7: Perform computing power scheduling adjustment operations according to the anomaly category. Update the computing network joint capability information based on the difference between the computing network operation status information and the expected operation indicators before and after the computing network scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, mark the combination of business scenario, computing power node and network path as a disabled combination. Disabled combinations do not participate in the determination of candidate computing power node set and candidate network path set.
2. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The process involves scheduling test services to multiple computing nodes and transmitting them through multiple network paths under at least one preset test service scenario. This process acquires basic computing network information and operational metrics for different combinations of service scenarios, computing nodes, and network paths. It also constructs joint computing network capability information and sets initial stability scores for each service scenario, computing node, and network path combination, including: The basic information of the computing network includes computing resource information of computing nodes and network resource information of network links. The computing resource information includes computing node identification, processing capacity, storage capacity and current computing load. The network resource information includes link bandwidth limit, link latency, historical packet loss rate and link cost information. The operation indicators include end-to-end latency, throughput, computing power utilization and bandwidth utilization. Under at least one preset test service scenario, the test service is scheduled to multiple computing nodes and transmitted through multiple network paths. The operation indicators of different service scenarios, computing nodes and network path combinations are collected. The correspondence between service scenarios, computing nodes, network paths and operation indicators is stored as computing network joint capability information. An initial value for stability score is set for each service scenario, computing node and network path combination.
3. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The process of receiving service requests and determining target computing network requirements based on these requests includes the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. Parse the service type, service access location, target end-to-end latency limit, target bandwidth requirement, and service reliability requirement from the service request; determine the computing power type requirement and computing power scale requirement based on the service type and service reliability requirement; combine the service access location, target end-to-end latency limit, and target bandwidth requirement into network-side target constraints; and combine the computing power type requirement, computing power scale requirement, and network-side target constraints into target computing network requirements.
4. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The determination of the candidate computing power node set and candidate network path set based on target computing network requirements, computing network joint capability information, and stability scores includes: Based on the joint computing and network capability information, computing power nodes that meet the target computing network requirements in terms of computing power type and computing power scale are selected from all computing power nodes to form a candidate computing power node set. For each candidate computing power node in the candidate computing power node set, based on the network-side target constraints, joint computing and network capability information, business scenarios, and the stability score of the combination of computing power nodes and network paths, network paths that meet the target computing network requirements in terms of end-to-end latency and bandwidth and whose stability score is not lower than the first threshold are selected from all network paths to form a candidate network path. The candidate computing power nodes are combined with their respective candidate network paths to form a candidate computing power node set and a candidate network path set.
5. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The aforementioned method, based on joint computing and network capability information and preset integrated computing and network scheduling rules, performs joint resource allocation on the candidate computing power node set and candidate network path set to generate an integrated computing and network computing power scheduling scheme. This integrated computing and network computing power scheduling scheme includes at least target computing power nodes and target network paths, including: According to the preset integrated computing and network scheduling rules, computing power resource quotas and network bandwidth quotas are allocated to each combination of computing power node and network path in the candidate computing power node set and candidate network path set. The preset integrated computing and network scheduling rules comprehensively weigh computing power resource utilization, network resource utilization, stability score and scheduling cost under the premise of meeting the target computing network requirements. Based on the operation indicators and stability scores of each combination recorded in the computing and network joint capability information, an end-to-end capability evaluation is performed on all combinations. The combination with the best comprehensive evaluation result is selected as the target computing power node and target network path, and an integrated computing power scheduling scheme including the target computing power node, target network path and resource quota is formed.
6. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The integrated computing power scheduling scheme is distributed to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. This includes: The system issues computing power scheduling instructions to the computing power orchestration unit. These instructions include the target computing power node identifier, the service instance mirror identifier, the number of service instances, and the computing power resource quota for each service instance. The computing power orchestration unit then creates service instances on the target computing power node and completes the reservation of computing power resources. The system also issues network configuration instructions to the network control unit. These instructions include the target network path identifier, the network link bandwidth reservation value, and the service traffic forwarding identifier. The network control unit then configures the forwarding path and bandwidth reservation in the network devices on the target network path to carry the target service traffic.
7. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The process of collecting computing network operation status information during the operation of the target service, determining expected operation indicators based on computing network joint capability information, comparing computing network operation status information with expected operation indicators, and determining computing network operation anomalies and their categories include: Periodically collect computing power side operation status information, including processor utilization, storage utilization, and service instance response latency of the target computing power node; periodically collect network side operation status information, including end-to-end latency, link bandwidth utilization, and packet loss rate of the target network path; read operation indicators corresponding to the current business scenario, target computing power node, and target network path from the computing-network joint capability information, determine the expected end-to-end latency, expected computing power utilization, and expected link bandwidth utilization as expected operation indicators, compare the computing network operation status information with the expected operation indicators, and determine computing network operation anomalies and the anomaly categories of computing network operation anomalies.
8. The integrated computing power scheduling method for a computing network according to claim 7, characterized in that, The aforementioned anomaly categories include a first anomaly category, a second anomaly category, and a third anomaly category, specifically as follows: The first anomaly category indicates that at least one computing power-side indicator in the computing power-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator, and at least one network-side indicator in the network-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator; the second anomaly category indicates that all computing power-side indicators in the computing power-side operation status information are within the preset deviation range of the corresponding expected operation indicator, and at least one network-side indicator in the network-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator; the third anomaly category indicates that all network-side indicators in the network-side operation status information are within the preset deviation range of the corresponding expected operation indicator, and at least one computing power-side indicator in the computing power-side operation status information exceeds the preset deviation range of the corresponding expected operation indicator.
9. The integrated computing power scheduling method for a computing network according to claim 1, characterized in that, The aforementioned process of performing computing power scheduling adjustment operations based on anomaly categories, updating computing network joint capability information based on the differences between computing network operating status information and expected operating indicators before and after the computing power scheduling adjustment operations, and updating the stability score of business scenarios, computing power nodes, and network path combinations, includes: For the first anomaly category, a joint rescheduling operation is performed to migrate the target service to a new target computing power node and switch to a new target network path. Based on the new target computing power node and the new target network path, new operating indicators and new stability scores are recorded in the computing-network joint capability information. For the second abnormal category, while keeping the target computing power node unchanged, adjust the target network path or adjust the network bandwidth quota on the target network path, and correct the operation indicators and stability scores corresponding to the combination of target computing power node and target network path in the computing network joint capability information based on the adjusted computing network operation status information; For the third anomaly category, while keeping the target network path unchanged, the computing power resource quota of the target computing power node is adjusted or the number of service instances is increased within the same computing network scheduling domain. The computing network scheduling domain is used to characterize the set of computing power resources uniformly scheduled by the integrated computing network computing power scheduling method. Based on the adjusted computing network operation status information, the operation indicators and stability scores corresponding to the combination of target computing power node and target network path in the computing network joint capability information are corrected. When the stability score of any service scenario, computing power node and network path combination is lower than the second threshold, the service scenario, computing power node and network path combination is marked as a disabled combination, and the disabled combination is excluded when determining the candidate computing power node set and candidate network path set in step three.
10. An integrated computing power scheduling system, characterized in that, An integrated computing network power scheduling method according to any one of claims 1-9 includes: a computing network joint capability information construction module, a target computing network demand determination module, a candidate resource determination module, a scheduling scheme generation module, a scheduling scheme execution module, an operation monitoring and anomaly judgment module, a scheduling adjustment and capability update module, and a data processing module; the computing network joint capability information construction module, the target computing network demand determination module, the candidate resource determination module, the scheduling scheme generation module, the scheduling scheme execution module, the operation monitoring and anomaly judgment module, and the scheduling adjustment and capability update module are respectively connected to the data processing module; The computing network joint capability information construction module is used to acquire basic computing network information and operational indicators of different business scenarios, computing power nodes and network path combinations under at least one preset test business scenario, construct computing network joint capability information, and set initial stability score values for each business scenario, computing power node and network path combination. The target computing network requirement determination module is used to receive service requests and determine the target computing network requirements based on the service requests. The target computing network requirements include the computing power requirements, bandwidth requirements, and end-to-end latency requirements of the target service. The candidate resource determination module is used to determine the set of candidate computing power nodes and the set of candidate network paths based on the target computing network requirements, computing network joint capability information and stability score; The scheduling scheme generation module is used to perform joint resource allocation on the candidate computing power node set and the candidate network path set based on the computing network joint capability information and the preset integrated computing network scheduling rules, and generate an integrated computing network computing power scheduling scheme. The integrated computing network computing power scheduling scheme includes at least the target computing power node and the target network path. The scheduling scheme execution module is used to distribute the integrated computing network computing power scheduling scheme to the computing power orchestration unit and the network control unit. The computing power orchestration unit deploys service instances on the target computing power nodes, and the network control unit completes network configuration on the target network path. The aforementioned operation monitoring and anomaly determination module is used to collect computing network operation status information during the operation of the target service, determine expected operation indicators based on computing network joint capability information, and compare computing network operation status information with expected operation indicators to determine computing network operation anomalies and the anomaly categories of computing network operation anomalies. The scheduling adjustment and capability update module is used to perform computing power scheduling adjustment operations according to the anomaly category, update the computing power joint capability information based on the difference between the computing power operation status information and the expected operation indicators before and after the computing power scheduling adjustment operation, and update the stability score of the combination of business scenario, computing power node and network path. When the stability score is lower than the second threshold, the combination of business scenario, computing power node and network path is marked as a disabled combination, and the disabled combination is excluded when the candidate resource determination module determines the candidate computing power node set and the candidate network path set.