Task scheduling method and system of heterogeneous computing cluster
By constructing a path cost matrix and a health prediction model, the health scores of computing nodes are dynamically monitored, solving the problem of low task scheduling accuracy in harsh environments for heterogeneous intelligent computing clusters, and realizing efficient task migration and a node early warning mechanism.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUOXINYUN (SHANGHAI) INTELLIGENT INFORMATION TECH CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
In embedded environments characterized by high vibration, wide temperature range, and strong electromagnetic interference, existing technologies cannot collect multi-dimensional time-series data through the intelligent platform management bus, resulting in low accuracy of task scheduling in heterogeneous intelligent computing clusters.
By constructing a path cost matrix and a health prediction model, multiple time-series data of computing nodes are collected, node health scores are dynamically monitored, preferred and alternative sets are determined, and task migration is performed when risky nodes appear.
It enables health management of computing nodes, improves the accuracy and reliability of task scheduling, adapts to applications in harsh scenarios, and triggers node alerts to ensure the safe transfer of task content.
Smart Images

Figure CN122132178A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of task scheduling, and in particular to a task scheduling method and system for heterogeneous intelligent computing clusters. Background Technology
[0002] With the development of technology, the demand for deploying high-performance artificial intelligence inference tasks in harsh embedded environments with high vibration, wide temperature range, and strong electromagnetic interference, such as airborne, shipborne, and vehicle-mounted edge computing, is increasing. These tasks (such as real-time target recognition of multiple high-definition videos, complex electromagnetic signal analysis, and path planning) typically require peak computing power exceeding tens of billions of floating-point operations per second (TOPS) and have extreme requirements for task response latency and continuous service reliability. Current technologies often only provide basic "power-on / power-off" or simple instantaneous temperature monitoring, failing to collect multi-dimensional time-series data such as temperature change rate, power consumption fluctuation variance, ECC error accumulation frequency, and PCIe link retransmission rate through the Intelligent Platform Management Bus (IPMB). This affects the health management of each computing node, resulting in low task scheduling accuracy in heterogeneous intelligent computing clusters. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a task scheduling method and system for heterogeneous intelligent computing clusters.
[0004] This invention provides a task scheduling method for heterogeneous intelligent computing clusters, including:
[0005] The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and data input port. Based on the connection relationship, it constructs a path cost matrix. At the same time, it collects normal data of each computing node under standard operating conditions and constructs the corresponding health prediction model.
[0006] The heterogeneous intelligent computing cluster collects multiple time-series data from each computing node through the IPMB bus, determines its characteristics within the time window based on the multiple time-series data, and determines the corresponding health score for each feature and each computing node in the health prediction model.
[0007] Upon receiving a new task, the heterogeneous intelligent computing cluster analyzes the new task and outputs the corresponding data content. Based on this data content, the corresponding computing nodes, and the health score, it determines the preferred set and the candidate set. Based on the preferred set and the candidate set, it determines the best candidate node and executes the new task based on the best candidate node.
[0008] The system dynamically monitors the health score of each computing node, triggers corresponding node alerts based on the health score of each computing node, identifies risk nodes, migrates the tasks corresponding to the risk nodes, and schedules the tasks to new nodes.
[0009] This invention provides a task scheduling system for a heterogeneous intelligent computing cluster, which is applied to the aforementioned task scheduling method for heterogeneous intelligent computing clusters; the task scheduling system for the heterogeneous intelligent computing cluster includes:
[0010] The health prediction module is used to query the corresponding configuration space of the heterogeneous intelligent computing cluster, determine the connection relationship between each computing node and data input port, construct a path cost matrix based on the connection relationship, and collect normal data of each computing node under standard operating conditions to construct the corresponding health prediction model.
[0011] The health score module is used by heterogeneous intelligent computing clusters to collect multiple time-series data from each computing node via the IPMB bus, determine the characteristics of each time-series data within a time window based on multiple time-series data, and determine the corresponding health score of each feature and each computing node in the health prediction model.
[0012] The candidate node module is used to collect new tasks, the heterogeneous intelligent computing cluster parses the new tasks and outputs the corresponding data content, and determines the preferred set and the candidate set based on the data content, the corresponding computing nodes and health scores; the best candidate node is determined based on the preferred set and the candidate set, and the new task is executed based on the best candidate node;
[0013] The new node module is used to dynamically monitor the health score of each computing node, trigger corresponding node alerts based on the health score of each computing node, identify risk nodes, migrate the task content corresponding to the risk node, and schedule the task content to the new node.
[0014] Compared with the prior art, the beneficial effects of the present invention are:
[0015] (1) The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and the data input port. Based on the connection relationship, a path cost matrix is constructed. At the same time, normal data of each computing node under standard working conditions is collected, and a corresponding health prediction model is constructed. The heterogeneous intelligent computing cluster collects multiple time series data from each computing node through the IPMB bus. Based on the multiple time series data, its characteristics within the time window are determined. Each characteristic and each computing node determines the corresponding health score in the health prediction model, realizing the control of the health of each computing node and quantifying it through the health score.
[0016] (2) When a new task is collected, the heterogeneous intelligent computing cluster parses the new task and outputs the corresponding data content. Based on the data content, the corresponding computing nodes and health scores, the preferred set and the alternative set are determined. Based on the preferred set and the alternative set, the best candidate node is determined and the new task is executed based on the best candidate node. The health scores of each computing node are dynamically monitored. Based on the health scores of each computing node, the corresponding node warning is triggered and the risk node is identified. The task content corresponding to the risk node is migrated and the task content is scheduled to the new node. The best candidate node is introduced, and the application of various harsh scenarios is realized. The corresponding node warning is triggered to realize the transfer of the task content corresponding to the risk node and improve the task scheduling accuracy of the heterogeneous intelligent computing cluster. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the task scheduling method for a heterogeneous intelligent computing cluster in an embodiment of the present invention.
[0018] Figure 2 This is a flowchart illustrating step S11 in the task scheduling method of the heterogeneous intelligent computing cluster in this embodiment of the invention.
[0019] Figure 3 This is a flowchart illustrating step S12 in the task scheduling method of the heterogeneous intelligent computing cluster in this embodiment of the invention.
[0020] Figure 4 This is a flowchart illustrating step S13 in the task scheduling method of the heterogeneous intelligent computing cluster in this embodiment of the invention.
[0021] Figure 5 This is a flowchart illustrating step S14 in the task scheduling method of the heterogeneous intelligent computing cluster in this embodiment of the invention.
[0022] Figure 6 This is a schematic diagram of the structural composition of the task scheduling system of the heterogeneous intelligent computing cluster in an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0024] Please see Figures 1 to 6 A task scheduling method for heterogeneous intelligent computing clusters, applied to task scheduling scenarios; the task scheduling method for heterogeneous intelligent computing clusters includes:
[0025] Step S11: The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and the data input port. Based on the connection relationship, a path cost matrix is constructed. At the same time, normal data of each computing node under standard working conditions is collected, and a corresponding health prediction model is constructed.
[0026] Step S12: The heterogeneous intelligent computing cluster collects multiple time-series data from each computing node through the IPMB bus, determines its characteristics within the time window based on the multiple time-series data, and determines the corresponding health score for each feature and each computing node in the health prediction model.
[0027] Step S13: Upon receiving a new task, the heterogeneous intelligent computing cluster parses the new task and outputs the corresponding data content. Based on this data content, the corresponding computing nodes, and the health score, a preferred set and a candidate set are determined. The best candidate node is determined based on the preferred set and the candidate set, and the new task is executed based on the best candidate node.
[0028] Step S14: Dynamically monitor the health score of each computing node, trigger the corresponding node warning based on the health score of each computing node, identify risk nodes, migrate the task content corresponding to the risk node, and schedule the task content to the new node.
[0029] refer to Figure 2 In step S11, the specific steps are as follows:
[0030] S111: Mark heterogeneous intelligent computing clusters. Heterogeneous intelligent computing clusters actively query the configuration space and implement multi-dimensional topology awareness to identify each computing node. Based on the tracing of each computing node, the corresponding data input port is determined, and the corresponding connection relationship is determined along each computing node and the corresponding data input port.
[0031] S112: Collect network delay data from each computing node, determine the corresponding path cost matrix based on the combination of each network delay data and connection relationship, and further dynamically adjust it based on historical traffic data. The path cost matrix dynamically adjusts the path cost according to the task type and system load to match the corresponding scheduling strategy.
[0032] S113: Monitor each computing node in real time, mark the normal data of each computing node under standard operating conditions, train on each normal data, and build the corresponding health prediction model according to the environmental changes during the operation of each computing node.
[0033] In the embodiments of this application, heterogeneous intelligent computing clusters are marked. The heterogeneous intelligent computing clusters actively query the configuration space and implement multi-dimensional topology awareness to identify each computing node. Based on the tracing of each computing node, the corresponding data input port is determined. The corresponding connection relationship is determined along each computing node and the corresponding data input port. This takes into account the overall consideration of each computing node and the corresponding data input port, and ensures the accuracy of the corresponding connection relationship.
[0034] At this point, the target heterogeneous intelligent computing cluster to be queried is determined so that configuration space query and topology awareness can be performed subsequently. It can be marked by the cluster's unique identifier (such as cluster ID) or the cluster's network address. At the same time, the configuration information of the heterogeneous intelligent computing cluster is obtained, including the type, number, and connection relationship of computing nodes. The cluster's configuration information is actively queried through standardized interfaces or protocols, such as the PCIe configuration space access interface or the IPMI protocol.
[0035] Obtain the physical topology of the heterogeneous intelligent computing cluster, including compute nodes, data input ports, and connection links, and evaluate the communication overhead of different paths; parse the configuration information of the PCIe switch to obtain the connection relationship between the switch ports and compute nodes and data input ports; monitor the status of the PCIe links, such as link bandwidth, latency, and error rate, and evaluate the communication quality of different paths; analyze network traffic to obtain the traffic load of different paths, evaluate the congestion level of different paths, identify all compute nodes and data input ports in the heterogeneous intelligent computing cluster, establish the connection relationship between them, and construct a connection relationship graph of compute nodes and data input ports based on the information obtained in the previous steps.
[0036] Specifically, in vehicle scenarios, the Vehicle Identification Number (VIN) or the vehicular network address can be used to identify heterogeneous intelligent computing clusters. Type B clusters can access the PCIe configuration space through the interface to query the PCIe switch configuration information, obtaining information such as the model, quantity, and connection relationships of each NPU card. Type B clusters can parse the PCIe switch configuration information to determine the connection relationship between each NPU card and the 10G Ethernet port, monitor the status of the PCIe link, and evaluate the communication quality of different paths. Type B clusters can construct a connection relationship diagram including components such as NPU cards, 10G Ethernet ports, and PCIe switches, and label the connection relationship and communication overhead of each component. Through proactive querying of the configuration space and multi-dimensional topology awareness, heterogeneous intelligent computing clusters can accurately obtain the physical topology and communication overhead information of the cluster, providing important reference for subsequent task scheduling and resource allocation.
[0037] Furthermore, network latency data from each computing node is collected, and the corresponding path cost matrix is determined based on the combination of each network latency data and connection relationships. This matrix is then dynamically adjusted in conjunction with historical traffic data. The path cost matrix dynamically adjusts the path cost according to the task type and system load to match the corresponding scheduling strategy. This approach takes into account the overall consideration of the combination of each network latency data and connection relationships, ensuring the accuracy of the corresponding path cost matrix.
[0038] At this point, high-performance hardware timers or timestamp mechanisms are used to send test probe packets or periodic beacons to each computing node; the system records the timestamps of data packets being sent from the source port, passing through the backplane and switching network, and arriving at the target computing node, and calculates the round-trip time (RTT) and one-way jitter; in addition, it is also necessary to record the actual throughput and bandwidth utilization under different link layers (such as PCIeGen3 / Gen4x4vsx8).
[0039] The abstract physical connection relationships and measured network delay data are mapped to a mathematical model that the scheduler can compute—the path cost matrix—to quantify the communication cost of each "source port-compute node" path. The cost function Cost(i,j)=f(Delayi,j,Hopi,j,BWi,j) is defined, where i is the data input port, j is the compute node, Delay is the collected delay, Hop is the number of hops (physical distance) through the switch, and BW is the link shared bandwidth. The value of each item in the matrix is calculated so that the larger the cost value, the higher the communication overhead.
[0040] Continuously monitor the bandwidth utilization of each shared link; when a link is detected to be in a high-load state (such as utilization exceeding 80%), dynamically increase the path cost weight passing through that link. This is equivalent to artificially increasing the "resistance" of the congested path when the network is congested, guiding the scheduler to select an idle path.
[0041] Introduce configurable weight coefficients λ; for “I / O intensive” tasks (such as video stream processing), increase the weight λ1 of path cost (communication overhead); for “computation intensive” tasks (such as complex model inference), decrease the communication weight and increase the weight λ2 of node load balancing; the cost function becomes Cost = λ1 × CommCost + λ2 × LoadCost.
[0042] Specifically, assuming the cluster consists of 12 NPU compute nodes interconnected via a high-performance PCIe Switch; LiDAR data stream test: The system simulates high-speed data streams from the LiDAR (such as 10Gbps point cloud data), injected through the 10G Ethernet port 1 on the front panel, and tests the data transmission latency from this port to NPU nodes 1-12; Camera data stream test: Simulates data streams from multiple high-definition cameras, injected through a Gigabit Ethernet port or a dedicated SerDes interface, and tests the latency to other compute nodes.
[0043] When a vehicle is traveling on a bumpy road, the backplane connector may experience a micro-transitional interruption, causing the PCIe link to need to be retrained. At this time, the acquisition module needs to record the entire process time from the link state to L0 (full speed) to Recovery and back to L0. This non-linear delay surge is recorded as special delay data to distinguish between "steady-state delay" and "dynamic interference delay".
[0044] A check of the PCIeSwitch configuration space revealed that 10G Ethernet port 1 is directly connected to the PCIeSwitch's uplink port, while NPU nodes 1-4 are located on the first downlink port of the Switch, and NPU nodes 5-8 are located on the second downlink port. Measurements showed that the average latency from 10G port 1 to NPU node 1 (direct connection or through a few layers of switching) is 5 microseconds, while the average latency to NPU node 12 (through multiple layers of switching or a complex path) is 15 microseconds.
[0045] Construct a matrix C[i][j], and set C[Port1][Node1] = α×1 + β×5μs and C[Port1][Node12] = α×3 + β×15μs (assuming the number of hops is 1 and 3 respectively). The system marks C[Port1][Node1] as a low-cost path and C[Port1][Node12] as a high-cost path. This means that when a task comes from port 1, the scheduling algorithm will prioritize NPU node 1 over node 12 to minimize data transfer overhead.
[0046] During high-speed driving, the vehicle enters a complex urban environment: the vehicle perception system detects heavy rain and activates the infrared imaging sensor, causing a sudden surge in data traffic connected to 10G port 2, consuming a large amount of PCIe uplink bandwidth; the system detects that the shared link bandwidth utilization rate SharedLinkBandwidthUtilization(i,j) of the path where 10G port 2 is located spikes; the algorithm automatically increases the C[i][j] values of all vehicles passing through this path.
[0047] The original path cost from port 2 to a certain computing node was 10. Due to traffic congestion, the cost was dynamically adjusted to 50 based on the β×SharedLinkBandwidthUtilization term. At this time, even if the computing node has a low load, the scheduler will tend to schedule newly arrived, latency-sensitive tasks to other nodes with slightly higher loads but smoother paths, in order to avoid task queuing and blocking on congested links.
[0048] In different operating modes of an autonomous driving system:
[0049] Highway cruise mode (computation-intensive): The system mainly performs long-distance path planning and a small number of obstacle detections; the task type is computationally intensive; the scheduling strategy increases λ2λ2 (load balancing weight); at this time, even if a computing node is physically far away (path cost is slightly higher), as long as the computing resources of that node are idle, the system will still allocate tasks to it to maximize the overall computing power utilization and ensure long-term planning computing throughput.
[0050] Emergency obstacle avoidance mode (I / O intensive and low latency): When the vehicle detects a sudden obstacle ahead, it needs to immediately process multiple high frame rate camera images to make braking decisions; the task type changes to I / O intensive, which is extremely sensitive to latency; the scheduling strategy instantly adjusts λ1λ1 (communication overhead weight) to the maximum.
[0051] The system forcibly selects the NPU node with the smallest value in the path cost matrix (closest physical distance and lowest latency), even if that node already has a certain computational load. This strategy sacrifices local load balancing in exchange for minimizing data transmission latency, ensuring that emergency braking commands can be generated within hundreds of milliseconds, thus meeting the safety requirements of hard real-time systems.
[0052] Therefore, by monitoring each computing node in real time and marking the normal data of each computing node under standard operating conditions, training on each normal data, and constructing a corresponding health prediction model based on the environmental changes during the operation of each computing node, the overall consideration of the environmental changes during the operation of each computing node is taken into account, ensuring the accuracy of the corresponding health prediction model.
[0053] At this point, it is necessary to acquire full-dimensional physical state information of the computing nodes during operation, especially fine-grained parameters that reflect hardware environmental stress and operational stability, to provide a data source for establishing health benchmarks and capturing abnormal fluctuations; using the Intelligent Platform Management Bus (IPMB) or Onboard Management Controller (BMC), all computing nodes are polled at millisecond sampling periods (e.g., 100ms); the collected indicators include not only instantaneous values, but also underlying signal layer data, such as the number of link retransmissions in the PCIe physical layer, the ECC correctable error count of the memory controller, and the ripple value of the voltage regulation module.
[0054] Establish a "health baseline" by collecting data under ideal system operating conditions (such as cold start, low load, and room temperature environment) to define the normal performance and stability of each computing node for subsequent comparison of deviations; during system initialization or maintenance windows, ensure that the cluster is under no-load or standard load conditions; at this time, collect the baseline values of each sensor (such as baseline temperature T0 and baseline power consumption P0) and store them as baseline feature vectors; at the same time, set static thresholds for each indicator.
[0055] The mathematical parameters of the health prediction model are constructed; using benchmark data and expert experience, the weight coefficients of each feature index on the health status are determined through training, so that the model can accurately quantify the "sub-healthy" state; based on historical fault data or simulated injection data, the correlation between various time-series features (such as temperature change rate and power consumption variance) and the health score is calculated; using a weighted fusion algorithm or regression model, the optimal weight coefficient wk and decay factor γ are fitted, so that the health score output by the model can truly reflect the reliability probability of the node.
[0056] This achieves a leap from "static threshold monitoring" to "dynamic trend prediction." Based on changes in the runtime environment, it calculates the node health score H(t) in real time, identifies early signs of performance degradation (sub-health state), and provides decision input for the task scheduler's preventive avoidance strategy. During runtime, it continuously calculates the temporal features (such as temperature slope TrendT and power consumption variance VarP) within the time window W. Substituting these features into the weighted decay formula H(t)=γ×H(t−1)+(1−γ) ×(100−∑wkFk) calculates the real-time score. When a feature trend (such as the second derivative) is detected to indicate deterioration, the model reduces the health score in advance.
[0057] Specifically, in the heterogeneous intelligent computing cluster of the vehicle, the vehicle is driving on a highway in the high temperature of summer; the monitoring process captures the core temperature Tj(t) of the NPU computing card in real time through the IPMB bus; due to the possible limitation of the vehicle's air conditioning system, the node temperature is monitored to rise from 75°C to 90°C in a short period of time; at the same time, the system monitors the retransmission packet count Rj(t) of the PCIe link layer; due to the high-frequency vibration of the vehicle engine being transmitted to the chassis, the VPX connector may become slightly loose, resulting in an increase in the link bit error rate and a non-zero jump in the retransmission count; the monitoring system records these time-series data streams, rather than a single average value, in order to subsequently analyze the slope of the temperature change over time and the cumulative frequency of errors.
[0058] During the vehicle's factory commissioning phase or the initial stage of each cold start: the heterogeneous intelligent computing cluster records that when the ambient temperature is 25℃ and there are no tasks running, the standby power consumption of each NPU module is 15W and the core temperature is 35℃; these values are marked as P0 and T0; the "normal" operating range is defined, for example, the PCIe link retransmission rate should be 0 and the ECC error rate should be 0; the system loads these standard operating condition data into the health assessment module; if the temperature reading of a certain NPU node deviates from the baseline T0 by more than 20℃ while the vehicle is driving at high speed, the system determines that the cooling system of that node may have undergone physical degradation (such as the thermal paste drying out), providing a basis for subsequent health score deduction.
[0059] Engineers simulated vehicle vibration environments and found that the "PCIe link retransmission rate" was far more accurate in predicting task failure rates than the "instantaneous temperature". Therefore, the training algorithm automatically assigns a higher weight (e.g., 0.5) to the retransmission rate feature wR, while the weight of the temperature change rate wT is relatively low (e.g., 0.2).
[0060] The training determined the historical health decay factor γ to be 0.7, which means that the health score has a strong "memory" and will not drop sharply due to a single temperature fluctuation. It must be abnormal for multiple consecutive sampling periods to be judged as a serious fault. Through this training, the model learned to remain robust when vehicle bumps cause instantaneous data jitter, and to react quickly when the temperature continues to rise and causes a thermal throttling trend.
[0061] As the vehicle transitioned from a flat highway to a rugged mountain road, the environment underwent a drastic change. The heterogeneous intelligent computing cluster's model detected that the temperature of NPU node 5, responsible for processing LiDAR data, was rising linearly at a rate of 1.5°C per second within 10 seconds (rendT remained positive), and the PCIe retransmission rate began to appear intermittently. Instead of waiting for the node to reach its overheat shutdown temperature (e.g., 105°C), the model, based on the aforementioned "deterioration trend," dynamically lowered the health score H(t) from 90 to 65 (below the Hoptimal threshold of 85) when the temperature reached 85°C and the retransmission rate increased. This health score was input to the scheduler in real time. The scheduler then redirected subsequent high-priority path planning tasks, originally assigned to node 5, to node 6, which had a perfect health score. This preventative migration of tasks was achieved before node 5 completely failed due to overheating, ensuring the continued availability of the autonomous driving system under adverse road conditions.
[0062] refer to Figure 3 In step S12, the specific steps are as follows:
[0063] S121: The heterogeneous intelligent computing cluster periodically collects data through the IPMB bus and manages the data for each computing node to collect multiple time-series data, including temperature, power consumption, ECC error, and PCIe link retransmission rate; based on the optimization of multiple indicators of multiple time-series data, the corresponding multiple coupling relationships are determined.
[0064] S122: Identify the corresponding electromagnetic interference factors for the control of harsh environments; each computing node performs time window detection on the electromagnetic interference factors and multiple coupling relationships, and outputs its characteristics within the time window;
[0065] S123: Introduce preset health decay content, mark the decay amount along the health decay content and each feature, and calculate the health score of each calculation node in combination with the health prediction model; the health score reflects the real state of the calculation node under specific working conditions and environmental stress.
[0066] In the embodiments of this application, the heterogeneous intelligent computing cluster periodically collects data through the IPMB bus and performs data management and control for each computing node to collect multiple time-series data, including temperature, power consumption, ECC error and PCIe link retransmission rate; based on the multi-index optimization of multiple time-series data, the corresponding multiple coupling relationships are determined, which takes into account the overall consideration of corresponding electromagnetic interference factors and ensures the accuracy of the corresponding multiple coupling relationships.
[0067] At this point, a transparent and real-time perception mechanism for the operating status of computing nodes is established; using a management bus independent of the business data channel, the underlying parameters reflecting the physical status of the hardware are obtained with high reliability and low overhead without interrupting the business flow; the monitoring daemon in the cluster scheduling manager sends polling commands to the Intelligent Platform Management Controller (IPMC) on each computing node at fixed time intervals (such as 100ms) through the Intelligent Platform Management Bus (IPMB); the IPMC reads the local sensor registers, packages the raw data into data frames conforming to the IPMI protocol, and sends them back to the main control module.
[0068] Acquire multi-dimensional state information with time-series characteristics to comprehensively characterize the operational health status of nodes from different physical dimensions (thermal, electrical, signal integrity), providing a data foundation for subsequent feature extraction; establish an independent data cache queue for each computing node; the collected indicators not only include instantaneous values, but also cover key time-series characteristic indicators, including core temperature Tj(t), instantaneous power consumption Pj(t), ECC correctable error count Ej(t), PCIe link layer retransmission packet count Rj(t), and NPU computing core utilization Uj(t).
[0069] By exploring the inherent correlations (multiple coupling relationships) between different physical indicators, we can identify complex fault modes that cannot be reflected by a single indicator. Through multi-indicator correlation analysis, we can improve the accuracy of identifying stress in complex environments. We can also perform correlation analysis and feature fusion on multiple sets of time-series data. We can calculate the covariance or correlation coefficient of different indicators within the same time window to determine whether there are coupling relationships (e.g., whether an increase in temperature leads to an increase in ECC error rate, or whether power consumption fluctuations are synchronized with PCIe retransmission rate). We can use these coupling relationships to optimize the health assessment model so that it can make a comprehensive judgment on complex symptoms.
[0070] Specifically, although the main data channel (PCIe bus) is processing 10Gbps point cloud data from the LiDAR at full capacity, the monitoring process reads the status in parallel through a separate IPMB bus (usually I2C or the physical layer of the system management bus). This design ensures that the monitoring data flow remains unimpeded even if data services become congested.
[0071] The main control module periodically sends the "acquire sensor readings" command to NPU cards 1-12 through dual redundant links of IPMB-A and IPMB-B. The system aligns the polling cycle with the vehicle control cycle (such as 10ms or 100ms) to ensure that the latest hardware status snapshot can be obtained in each control cycle.
[0072] The heterogeneous intelligent computing cluster implements independent data flow control for the 12 NPU computing nodes in the vehicle: it monitors the core temperature sequence of NPU node 5 within a continuous time window, such as {t1:75∘C,t2:76∘C,t3:78∘C}, and captures the trend of temperature change over time; it synchronously collects the physical layer retransmission packet count of the PCIe link; if the vehicle passes through a bumpy road section causing a momentary disconnection of the connector, the retransmission packet count Rj(t) will show a non-zero jump; the system maintains a circular buffer of length WW (e.g., 10 cycles) for each node to store the time-series data within the most recent second. This control mechanism ensures the temporal continuity of the data, enabling the system to distinguish between "momentary interference" and "continuous degradation".
[0073] In addition, when vehicles enter extreme operating conditions of high temperature and high load, a single indicator may not be sufficient for early warning. The heterogeneous intelligent computing cluster optimizes through the coupling of multiple indicators: the system found that although the temperature Tj(t) of NPU node 8 did not exceed the threshold, the variance VarP of its power consumption Pj(t) increased abnormally. At the same time, the PCIe link retransmission rate Rj(t) showed an upward trend, which revealed the coupling relationship of "increased power supply ripple -> decreased signal integrity".
[0074] The health model adjusts the weights based on this coupling relationship; if the system detects that "power fluctuations" and "link retransmissions" are strongly correlated within the time window, it determines that the node is suffering from unstable power noise interference (which may be caused by sudden changes in the load of the vehicle generator), which is more in line with the real risk than simply relying on temperature warnings; the model reduces the health score of the node accordingly to avoid potential calculation errors in advance.
[0075] Furthermore, corresponding electromagnetic interference factors are identified for the control of harsh environments; each computing node performs time window detection on these electromagnetic interference factors and multiple coupling relationships, and outputs their characteristics within the time window, which is compatible with the overall consideration of the control of harsh environments and ensures the accuracy of the corresponding electromagnetic interference factors.
[0076] At this point, in harsh embedded environments (such as vehicles driving in areas with strong electromagnetic interference), the impact of external electromagnetic noise on the signal links of computing nodes can be identified and quantified; invisible environmental stresses can be transformed into monitorable system-level parameters to shield or correct physical layer damage to data transmission caused by electromagnetic interference.
[0077] By utilizing the error statistics mechanism and spectrum monitoring technology of the PCIe link physical layer, the intensity of electromagnetic interference can be inferred by monitoring the abnormal fluctuations in physical layer coding errors (such as 8b / 10b decoding errors), link retraining frequency, and bit error rate (BER). At the same time, the voltage ripple on the power management bus (PMBus) is monitored, because strong electromagnetic interference usually induces high-frequency noise on the power lines.
[0078] Within a specific time sliding window, the multi-coupling relationships between electromagnetic interference factors and the internal state of nodes (such as temperature, power consumption fluctuations, and ECC errors) are comprehensively analyzed; secondary fault modes triggered by environmental interference are identified, such as whether electromagnetic interference causes memory flips or computational logic errors; at the same time, a sliding time window W (e.g., 1 second or 100 milliseconds) is defined; within this window, the correlation analysis of multi-dimensional time-series data collected for each computing node is performed; whether electromagnetic interference indicators (such as a surge in PCIe retransmission rate) have a causal or co-occurring relationship with internal computing indicators (such as an increase in ECC error count and abnormal fluctuations in computing core utilization) in time is detected.
[0079] Representative numerical features are extracted from the detected multi-dimensional coupling data for quantitative calculation of the health prediction model. These features reflect the stability and anti-interference ability of nodes in harsh environments. Key statistical features of each indicator within the time window are calculated. For electromagnetic interference, the "link layer retransmission packet density" is calculated. For coupling relationships, the "error rate increment at the time of interference" or the "cross-correlation coefficient between interference duration and temperature rise" is calculated. These normalized feature values Fk are used as inputs to the penalty term in the health model.
[0080] Specifically, the monitoring module of the heterogeneous intelligent computing cluster detected a sudden increase in the physical layer bit error rate of the PCIe port connected to the external millimeter-wave radar. The system identified this as being caused by a strong external electromagnetic field coupling onto the high-speed differential signal line, resulting in a decrease in signal integrity. At this point, the interference factor was quantified into two key indicators: “link signal-to-noise ratio (SNR) degradation” and “physical layer error burst”, which were then input into the health assessment model as environmental stress.
[0081] In a strong electromagnetic interference environment, the heterogeneous intelligent computing cluster performs time window detection on the NPU node responsible for processing radar data: within the time window [t, t+W], the system observes a peak in the PCIe link retransmission rate R(t); simultaneously, it detects a slight increase in the ECC correctable error count E(t) of the node, revealing the coupled link of "electromagnetic noise -> signal timing violation -> memory read / write data bit flipping"; the system does not regard the ECC error as merely a random defect of the memory chip, but associates it with external electromagnetic interference factors, determining that it is an instantaneous error induced by environmental stress.
[0082] After completing the time window detection, the heterogeneous intelligent computing cluster outputs the following features for scoring: the output feature FEMI is the "percentage of time during which the retransmission rate exceeds the threshold", assumed to be 20%; the output feature ECC_Coupling is the "multiple of the ECC error rate relative to the baseline during the interference period", assumed to be 5 times.
[0083] These features are input into the health prediction formula H(t)=γ×H(t−1)+(1−γ)×(100−∑wkFk); due to the large weight w of retransmission and ECC coupling features caused by electromagnetic interference, the health score of the NPU node will drop rapidly; based on this, the system determines that the node is currently severely affected by electromagnetic interference and is in a "high-risk" state, so as to avoid the node in subsequent task scheduling, or migrate the critical tasks on the node to other computing boards with better shielding, so as to ensure that the perception function of the autonomous driving system is not affected by the electromagnetic environment.
[0084] Therefore, a preset health decay content is introduced, and the decay amount is marked along this health decay content and various features. The health score of each computing node is calculated in combination with the health prediction model. The health score reflects the true state of the computing node under specific working conditions and environmental stress. The introduction of the health score reflects the true state of the computing node under specific working conditions and environmental stress. At the same time, the health of each computing node is controlled and quantified through the health score.
[0085] At this point, a "state persistence" memory mechanism is established to prevent the health score from drastically changing due to instantaneous noise or single sampling anomalies. By introducing a historical decay factor, the system is given the ability to judge changes in node state with "inertia," distinguishing between short-term fluctuations and long-term deterioration trends. A historical health decay factor γ (with a value range of 0~1, such as 0.7) is defined, which determines the extent to which the current health score inherits the score from the previous moment. The design principle of the decay content is: if the node's historical state is good, a single minor anomaly will not cause the score to drop precipitously; conversely, if the node is in an abnormal state for a long time, the score will show an accelerated downward trend.
[0086] The negative impact of specific physical characteristics on health is quantified by transforming multi-dimensional time-series features into a unified numerical "penalty amount." By calculating the normalized feature value Fk, the degree to which indicators such as temperature, power consumption, and ECC error deviate from the standard is determined. At the same time, for each extracted time-series feature (such as the temperature change slope TrendT, power consumption variance VarP, and error rate ErrRateE), it is compared with the preset fault thresholds (such as θT, θP). The penalty value Fk is calculated using a normalization function (such as a linear piecewise function or a sigmoid function), where Fk∈[0,1], where 0 represents no effect and 1 represents reaching the maximum tolerance limit.
[0087] By combining historical status and current real-time characteristics, a quantitative score that accurately reflects the survivability of a node under specific working conditions is output. This score will be directly used for task scheduling decisions to determine whether a node is preferred, a backup, or isolated. The preset decay factor γ, the health status H(t−1) of the previous time step, and all feature penalty values Fk and their weights wk of the current time step are substituted into the weighted decay model formula: H(t)=γ×H(t−1)+(1−γ) ×(100−∑(wk×Fk)). The calculated result H(t) is the health status score of the current time step.
[0088] Specifically, in the heterogeneous intelligent computing cluster of the vehicle, the NPU node may experience a sudden temperature jump due to coolant flow fluctuations caused by a sudden high-speed turn; the system is set to γ=0.7, which means that the current health score is determined by 70% of the historical state and 30% of the current observed features.
[0089] This design prevents instantaneous vibration interference when a vehicle passes over a speed bump (such as brief bit errors in the PCIe link) from causing the node to be misjudged as "faulty"; the attenuation mechanism acts as a low-pass filter, filtering out high-frequency environmental noise and ensuring that only persistent environmental stress (such as continuous high temperature or radiator blockage) will significantly lower the health score.
[0090] The system detected that the temperature change rate TrendT exceeded the threshold θT, and calculated FTemp=0.4 (indicating that the temperature rise was too fast, but not yet critical). Simultaneously, the ECC error rate slightly increased, and FECC=0.1 was calculated. The system marked these Fk values as "attenuation," representing the degree of erosion of node health by the current environmental stress (high temperature, soft errors), and used them as input for weighted summation to deduct from the total score.
[0091] The heterogeneous intelligent computing cluster performs a final score on the NPU nodes affected by high temperatures: Assuming the node's health H(t−1) = 95 (out of 100) at the previous moment; based on the marked decay amount, the formula is used to calculate: H(t) = 0.7 × 95 + 0.3 × (100 − (0.5 × 0.4 + 0.3 × 0.1) × 100) (Note: Here we assume weight wTemp = 0.5, wECC = 0.3, and the calculation within the parentheses deducts the total score); the calculated H(t) may drop to 88 points; since the system's optimal threshold Hoptimal is 85, the node is still in the optimal set, but the score is close to the edge. This health score reflects the node's true state under the specific working condition of "scorching sun exposure + high load computing": still reliable, but the margin is decreasing; when the scheduler allocates the next high-priority task, it will refer to this score and tend to select nodes with higher health, thereby achieving a smooth switch of task flow before hardware performance deteriorates.
[0092] refer to Figure 4 In step S13, the specific steps are as follows:
[0093] S131: Based on the monitoring of the heterogeneous intelligent computing cluster, determine the corresponding new task; based on the parsing of the new task, determine the corresponding descriptor and historical execution mode; based on the multi-factor analysis of the descriptor and historical execution mode of the new task, determine the corresponding output data content;
[0094] S132: In each computing node, a corresponding set framework is determined based on the computing node and the corresponding data content, and the matching coefficient of each computing node is determined by matching the health score of each computing node with the set framework. Based on the matching coefficient of each computing node, it is divided into the preferred set and the alternative set, and some abnormal nodes are excluded.
[0095] S133: Input the preferred set and the candidate set into the virtual environment space, and determine the computing power level of each computing node in the virtual environment space to handle the new task. Based on the comparison of the computing power levels of each computing node, determine the best candidate node. The best candidate node will execute the new task first and perform multi-level execution on the new task.
[0096] In the embodiments of this application, a new task is determined based on the monitoring of the heterogeneous intelligent computing cluster, and a corresponding descriptor and historical execution mode are determined based on the parsing of the new task. The corresponding output data content is determined based on the multi-factor analysis of the descriptor and historical execution mode of the new task, which is compatible with the overall consideration of the multi-factor analysis of the descriptor and historical execution mode of the new task, and ensures the accuracy of the corresponding output data content.
[0097] At this point, by continuously monitoring external I / O ports and message queues, AI inference tasks that require computational resources are identified, providing input sources for subsequent scheduling decisions. The cluster scheduler monitors the status of each data input port (such as 10G Ethernet ports and Gigabit Ethernet ports) through interrupt or polling mechanisms. When an external sensor or data source sends a data frame, the network stack receives the data and triggers the scheduler's event response, recognizing it as a new task to be processed (Taskm).
[0098] Deconstructing the technical attributes of new tasks, extracting the metadata required for scheduling, and retrieving historical execution statistics for this type of task helps the system understand the resource requirements of the task and predict its impact on system load and communication links. The scheduler parses the data header or control signaling of the task and extracts the task descriptor, including the source I / O port pm, estimated computational load Compm (such as FLOPs), estimated data volume Datam (such as MB), and priority Prim. At the same time, it queries the historical database to retrieve pattern data such as the average execution time and peak resource consumption of this type of task in the past.
[0099] Simultaneously, contextual data for task scheduling is generated. This content includes not only the original load but also scheduling auxiliary information derived from multi-factor analysis, such as data distribution characteristics and topology affinity suggestions, to guide the scheduler in selecting the optimal node. The topology information (source port) in the task descriptor and the communication characteristics in the historical pattern are comprehensively analyzed. If the historical pattern shows that the task is "I / O intensive", the "data locality" mark is strengthened in the output content; if it is "computation intensive", the "computing availability" mark is strengthened. Finally, a complete data content containing the original data vector and the scheduling feature vector is formed.
[0100] Specifically, when the vehicle-mounted LiDAR scans a pedestrian ahead, it sends a high-density point cloud data packet through 10G Ethernet port 1. The multi-functional rear panel on the VPX backplane receives the data packet, writes the data to the main control module's memory via DMA (Direct Memory Access), and sends a hardware interrupt to the scheduling manager, indicating that a new high-priority perception task has arrived.
[0101] The analysis revealed that the source port was 10G port 1, the data volume was 5MB (point cloud frame), and it was estimated that 3D target detection inference was required, which involved a large amount of computation. The priority was "critical security task". The system query history showed that such "3D point cloud target detection" tasks usually require access to a large amount of video memory and have specific requirements for data exchange between computing nodes (such as the need to frequently access the shared parameter server).
[0102] The heterogeneous intelligent computing cluster, combined with the characteristics of the LiDAR task, performs multi-factor analysis and outputs the data content for scheduling: The analysis suggests that since the data comes from 10G port 1 and the task is 3D processing, it is a typical combination of "I / O intensive" (large data transfer volume) and "computation intensive" task; the system output includes a payload containing the original point cloud data and an accompanying scheduling tag. It is recommended to prioritize NPU computing nodes that are physically close to 10G port 1 and have high-bandwidth PCIe links to minimize the transmission latency of point cloud data and ensure the real-time response of the autonomous driving system to obstacles ahead.
[0103] Furthermore, in each computing node, a corresponding set framework is determined based on the computing node and the corresponding data content, and the matching coefficient of each computing node is determined by matching the health score of each computing node with the set framework. Based on the matching coefficient of each computing node, it is divided into a preferred set and a candidate set, and some abnormal nodes are excluded.
[0104] At this point, a preliminary matching logic framework between tasks and computing nodes is established, mapping the attributes of new tasks to available computing resources in the cluster, providing a structured evaluation basis for subsequent fine-grained screening and decision-making; the scheduler parses the data content of new tasks (such as I / O port location, data size, and computing resource requirements) and aligns it with the status information of all online computing nodes; an evaluation context containing all candidate nodes is constructed, which defines the basic logical boundaries for determining whether a node is suitable to undertake the task (such as whether it has sufficient remaining computing power and whether it supports a specific instruction set).
[0105] Quantify the comprehensive adaptability of each computing node to execute the task, and introduce the physical health status of the node as the core weight into the matching process; the higher the health score, the stronger the reliability of the node in the current harsh environment (such as vibration, high temperature), and the higher the priority of the task it undertakes; introduce health score thresholds (such as Hoptimal and Hmin); calculate the matching coefficient of each computing node, which not only depends on the computing power matching degree, but more on its real-time health score Hj(t); the health score is directly converted into the gain factor of the matching coefficient. For example, the matching coefficient of a node with a health score of 100 is 1.0 (full score), and the matching coefficient of a node with a decreasing health score decays linearly or non-linearly accordingly.
[0106] Implement hierarchical management of computing nodes, preferentially allocate key tasks to the most reliable nodes, while reserving a certain number of secondary nodes as a buffer, and actively isolate nodes with serious risks to ensure the certainty of task execution and the high availability of the system; classify nodes according to the preset health thresholds (Hoptimal and Hmin): Preferred set (Soptimal): includes nodes with high matching coefficients and health scores ≥ Hoptimal; Alternate set (Sbackup): includes nodes with medium matching coefficients and health scores between Hmin and Hoptimal; Exclusion set: includes abnormal nodes with health scores < Hmin, and these nodes do not participate in this task scheduling.
[0107] Specifically, in the heterogeneous intelligent computing cluster of the vehicle, the system receives a high-definition video stream task from "10G network port 1"; the system identifies that the task has a large amount of data and extremely high real-time requirements, so the collection framework excludes those computing nodes whose current load has exceeded 90%, and those nodes with insufficient physical link bandwidth with 10G port 1, forming a preliminary "node pool for allocation".
[0108] The vehicle is driving at high speed in a high-temperature environment, and the health of the NPU nodes shows differentiation; Node A (NPU node 5): Due to good heat dissipation, the health score HA(t) = 95; due to its high score, its matching coefficient is marked as 1.0 (fully matched); Node B (NPU node 8): Affected by the engine vibration, the PCIe link has occasional retransmissions, and the health score HB(t) = 75; although the computing power is sufficient, due to the health being lower than the preferred threshold (such as 85), its matching coefficient is demoted and marked as 0.8 or lower; Node C (NPU node 12): The temperature is too high to trigger throttling, and the health score HC(t) = 40; its matching coefficient is extremely low, even close to 0.
[0109] The heterogeneous intelligent computing cluster performs the final classification of the above nodes: Optimal Set: Node A (health level 95) is assigned to the optimal set Soptimal; the scheduler will prioritize assigning urgent obstacle detection tasks to this node to ensure absolute stability of computing power supply during vehicle obstacle avoidance; Backup Set: Node B (health level 75) is assigned to the backup set Sbackup; if all nodes in the optimal set are busy, tasks will be scheduled to this node, and the system will prepare corresponding fault tolerance mechanisms; Anomaly Removal: Node C (health level 40) is marked as an "abnormal node" and directly removed from the scheduling queue; the system prohibits the distribution of new tasks to it to prevent the critical path of the autonomous driving system from being interrupted due to possible failures of this node (such as hot crashes), and at the same time triggers the isolation and recovery process of this node.
[0110] Therefore, the preferred set and the candidate set are input into the virtual environment space, and the computing power level of each computing node in the virtual environment space is determined to handle the new task. The best candidate node is determined based on the comparison of the computing power levels of each computing node. The best candidate node is given priority to execute the new task and performs multi-level execution of the new task. This takes into account the overall consideration of the comparison of the computing power levels of each computing node and ensures the accuracy of the best candidate node.
[0111] At this point, a sandbox model for task scheduling is built in a logically isolated virtual environment, mapping the real-time state of the physical cluster (preferred / alternate node set) to a virtual resource pool. This allows the scheduler to perform complex simulations and multi-dimensional resource assessments without interfering with the actual business flow.
[0112] The scheduler instantiates an "environment virtual space" object in memory and loads the state of all nodes (including static attributes such as architecture type and dynamic attributes such as current cache popularity) from the preferred set Soptimal and the alternative set Sbackup, which have been filtered by health, into this space. This space also preloads the computational requirements and prediction models for new tasks.
[0113] The potential performance of each candidate node in executing a specific new task is quantified; not only are general computing power indicators considered, but also the fit between the node and task characteristics (such as algorithm type and data structure) and the constraints of the current operating environment on performance are evaluated; a performance prediction model is run in virtual space, which analyzes the matching degree between task descriptors (such as the number of neural network layers and the number of parameters) and node attributes (such as tensorcore configuration and memory bandwidth); at the same time, an environmental correction factor (such as the frequency reduction coefficient caused by temperature) is introduced to calculate a normalized "computing power level".
[0114] The scheduler selects the target node with the highest overall cost-effectiveness from the candidate pool. By comparing computing power levels horizontally and combining path cost and health risk, it makes a globally optimal scheduling decision to maximize task throughput and minimize latency. The scheduler sorts the computing power levels of all nodes in the virtual space. Combining the constructed path cost matrix and the calculated health score, it performs a multi-objective weighted decision. The node with the highest score in "computing power level" and the lowest scores in "communication overhead" and "health risk" is selected as the best candidate node jopt.
[0115] To ensure that high-priority tasks can immediately obtain the best resources for execution, a phased task execution status tracking mechanism is established to monitor progress in real time and perform fine-grained fault tolerance management during execution. The scheduler sends a task start instruction to the best candidate node, joptjopt, and assigns it the highest bus arbitration priority. At the same time, the execution process of new tasks is divided into multiple logical stages (such as "data loading", "forward inference N layers", and "post-processing"), and a multi-level execution context is established in the task status table.
[0116] Specifically, the vehicle's heterogeneous intelligent computing cluster is handling an emergency obstacle avoidance task; the system loads the NPU nodes with the best health (1-3, the preferred set) and the nodes with average health (4-6, the alternative set) into the virtual environment space; in this virtual space, the system not only maps the hardware state, but also simulates the vehicle's current dynamic state (such as voltage fluctuations caused by rapid acceleration), providing a contextual environment for subsequent evaluation of computing power levels.
[0117] In the virtual environment, the system evaluates the ability of each node to process the "LiDAR point cloud segmentation" task: Node A (NPU node 1): Although it has strong computing power, the cache contains video data, which does not match the point cloud data structure, and its computing power level is rated as "good"; Node B (NPU node 2): The architecture is specially optimized for 3D convolution operations, and the current temperature is suitable and the frequency is not reduced, so its computing power level is rated as "excellent"; Node C (candidate set): Due to abnormal fan speed, it is slightly overheated, and the prediction performance will decrease by 15%, so its computing power level is rated as "medium".
[0118] The system comparison revealed that although node B (NPU node 2) had a slightly longer physical path, it had the highest overall score due to its "excellent" computing power level and perfect health. Node A, on the other hand, had a shorter physical path, but its computing power level was only "good" due to cache mismatch, resulting in the second-highest overall score. The system determined node B to be the best candidate node. This decision ensured optimal task processing efficiency and avoided additional latency caused by architecture mismatch.
[0119] The system issues the following instructions to the selected NPU node 2: Priority execution: Utilizing the PCIeQoS (Quality of Service) mechanism, ensure that node 2 obtains the highest priority memory access bandwidth and immediately begins processing LiDAR point cloud data to guarantee the real-time performance of obstacle avoidance decisions; Multi-level execution: The system divides the obstacle avoidance inference task into K stages; After each stage is completed (e.g., completing the inference of the 50th layer network), node 2 will save the intermediate activation values (feature maps) as checkpoints to shared storage according to the above requirements. This multi-level execution mechanism ensures that even if a node suddenly fails in the later stages of task execution (e.g., the last 10 layers), the system can quickly "continue calculation" on other nodes based on the most recent checkpoint without having to start from scratch, greatly improving the survivability of the autonomous driving system in harsh environments.
[0120] refer to Figure 5 In step S14, the specific steps are as follows:
[0121] S141: Mark the health score of each computing node. The health score of each computing node changes dynamically as the heterogeneous intelligent computing cluster works. Determine the amount of change in the health score of each computing node. Based on the amount of change in the health score of the computing node and the computing load of the computing node, determine the corresponding early warning signal. Trigger node early warning for risky nodes based on the early warning signal.
[0122] S142: When a risk node's warning is triggered, the heterogeneous intelligent computing cluster immediately freezes the distribution of new tasks to the risk node and uses the PCIeDMA link to read the contents of all unfinished tasks on the risk node in parallel to migrate the contents of all unfinished tasks on the risk node. During the migration process, a double buffering mechanism is introduced so that the new node can continue to execute all unfinished tasks on the risk node.
[0123] In the embodiments of this application, the health score of each computing node is marked. The health score of each computing node changes dynamically as the heterogeneous intelligent computing cluster operates. The change in the health score of each computing node is determined. Based on the change in the health score of the computing node and the computing load of the computing node, a corresponding early warning signal is determined. Based on the early warning signal, a node early warning for the risk node is triggered. This approach takes into account both the change in the health score of the computing node and the computing load of the computing node, ensuring the accuracy of the corresponding early warning signal.
[0124] At this point, a real-time, quantified reliability identifier is established for each computing node in the cluster; the health score is the core basis for the scheduler to judge the current status of the node (healthy, sub-healthy, faulty), and is used to dynamically adjust the task allocation strategy; using the health prediction model (as described in step S123), the health score Hj(t) of each computing node is continuously calculated and updated. This score is stored in the shared status table of the scheduler as a key attribute field of the cluster resource view, which can be read by the scheduler at any time.
[0125] By capturing the dynamic evolution trend of node status and calculating the time derivative (rate of change) of the score, potential accelerated degradation or fault precursors can be identified, realizing the transformation from "static assessment" to "dynamic prediction". The monitoring process records the time series Hj(t), Hj(t−1), Hj(t−2)... of the health score at a fixed sampling rate; the change in health score ΔHj=Hj(t)−Hj(t−1) and the trend of change within a specific time window (such as the second derivative or moving average rate of change) are calculated.
[0126] By combining hardware health status with the urgency of current business load, the system intelligently determines the warning level; it avoids false alarms when nodes have slight fluctuations but extremely low load, or overreacts when nodes deteriorate rapidly but are running non-critical tasks, thus achieving accurate risk warnings; it establishes a multi-dimensional warning judgment logic; the input variables include the change in health score ΔHj (or instantaneous score), the current computing load type (such as critical security tasks vs. non-critical entertainment tasks) and load volume, and determines the warning signal level (such as: green - normal, yellow - need attention, red - immediate isolation) through rule engine or model reasoning.
[0127] Before a risk node completely fails, the system's proactive defense mechanism is activated. By triggering an early warning, the scheduler is notified to stop distributing new tasks to the node and prepare to initiate a hot migration process, minimizing the impact of potential failures. When the early warning signal reaches the trigger threshold (such as a red warning), the scheduler immediately updates the node status table and marks the node as "warning" or "isolated". The system sends an interrupt signal to the fault prediction and proactive hot migration module to initiate the subsequent task migration process.
[0128] Specifically, in the vehicle's heterogeneous intelligent computing cluster, the system maintains health indicators for 12 NPU nodes: nodes 1-8 are marked green (score > 90, healthy), nodes 9-10 are marked yellow (score 70-85, sub-healthy), and node 11 is marked red (score < 40, severe risk). These score labels are displayed in real time on the monitoring dashboard of the vehicle's central computing platform, allowing system administrators or autonomous driving decision-makers to intuitively perceive the reliability status of the underlying hardware.
[0129] The vehicle's computing load surged, causing drastic fluctuations in node status: NPU node 9, responsible for image processing, experienced a rapid drop in its health score from 88 to 75 within the past second due to obstructed heat dissipation; the system calculated that the change in node 9 was ΔH9=−13, and the rate of change exceeded the preset threshold θdecay. This significant negative change indicates that the node is experiencing rapid performance degradation.
[0130] The heterogeneous intelligent computing cluster makes a comprehensive judgment based on the status and current task of node 9: Node 9 is currently running the "lane detection" task (which is a critical safety task) and the computing load is as high as 95%; due to the sharp drop in ΔH9 and the carrying of critical high load, the judgment logic is at extremely high risk; the system generates a "red warning" signal, indicating that the node is about to fail and the impact is huge; in contrast, if node 9 is only running the background log analysis task, the system may only generate a "yellow warning".
[0131] In response to the red warning signal from NPU node 9, the heterogeneous intelligent computing cluster executed an emergency response: Upon receiving the red warning signal, the scheduler immediately marked NPU node 9 as "high-risk"; the scheduler instantly stopped sending new lane detection frames to node 9 to prevent the loss of new data due to node failure; at the same time, the system automatically queried the currently executing task queue on node 9, preparing to read its latest checkpoint data so that the task could be seamlessly migrated to the healthy NPU node 1, ensuring that the visual perception function of the autonomous driving system was not interrupted and guaranteeing driving safety.
[0132] Furthermore, when a risk node's warning is triggered, the heterogeneous intelligent computing cluster immediately freezes the distribution of new tasks to the risk node and uses the PCIeDMA link to read the content of all unfinished tasks on that risk node in parallel to migrate the content of all unfinished tasks on that risk node. During the migration process, a double buffering mechanism is introduced, allowing the new node to continue executing all unfinished tasks on the risk node. At the same time, the best candidate node is introduced, enabling applications in various demanding scenarios and triggering the corresponding node warning to achieve the transfer of the task content corresponding to the risk node, thereby improving the task scheduling accuracy of the heterogeneous intelligent computing cluster.
[0133] At this moment, upon confirming that the compute node has entered a high-risk state (such as a sharp deterioration in health), the inflow of subsequent tasks is cut off to prevent more tasks from accumulating before the node experiences an irreversible failure, thereby limiting the impact of the failure to the limited set of currently executing tasks. After receiving the red warning signal, the scheduler immediately changes the status bit of the risk node jfail in the scheduling queue to "frozen" or "stop scheduling". The load balancer removes jfail from all candidate sets (Soptimal and Sbackup) to ensure that new tasks arriving later will not be assigned to this node.
[0134] With the highest data transfer efficiency, critical context data on risky nodes is salvaged; leveraging the high bandwidth and low latency of PCIeDMA (Direct Memory Access), CPU intervention is bypassed to directly capture memory data, enabling rapid transfer of task states and minimizing service interruption time during migration; the main control module initiates DMA transfer requests via the PCIe bus; based on the task status table information on the risky node, the base address and length of all incomplete tasks {Taskf1, Taskf2, ...} in video memory or main memory are locked; the system starts multiple DMA channels in parallel, directly reading the latest checkpoint data of these tasks (such as activation tensors of neural networks, model context) into the main control module's memory buffer or directly transferring it to the memory of the target new node via Peer-to-Peer DMA.
[0135] By employing double buffering technology, new nodes can pre-recover part of their computational context using already arrived data segments while receiving migration data, minimizing task pause time and meeting the extreme service continuity requirements of autonomous driving systems. Two memory regions are allocated on the target new node's `jnew`: a receive buffer and a working buffer. As the data stream from the risk node is continuously written to the receive buffer via PCIeDMA, the runtime library asynchronously monitors the receiving progress. Once a complete computational unit (such as a network layer checkpoint) accumulates in the receive buffer, the system immediately atomically maps or copies its contents to the working buffer and triggers the new node's computational core to resume execution from that checkpoint. Simultaneously, the next batch of data continues to be written to the receive buffer.
[0136] Specifically, NPU node 9 triggers an overheat warning; the scheduler of the heterogeneous intelligent computing cluster immediately stops distributing new camera image frames to node 9; at this time, new frame data in the highway scene will be redirected to other healthy NPU nodes, while node 9 only retains the previous frame image and its corresponding intermediate inference state that has not yet been processed in its memory, ensuring that the fault will not cause the data flow of the perception system to be interrupted.
[0137] For unfinished tasks on NPU node 9: The system detects that two tasks, "lane detection" and "vehicle recognition," are running on node 9; the scheduler immediately initiates PCIeDMA to read the feature map data generated by these two tasks in the 50th layer of the network in parallel; utilizing the high bandwidth of PCIeGen3x8, a large amount of intermediate state data is completely read out within milliseconds. This mechanism avoids the overhead of relaying and copying through the CPU, ensuring that the critical inference context is successfully maintained and processed before node 9 completely crashes.
[0138] The heterogeneous intelligent computing cluster migrates the tasks of NPU node 9 to the healthy NPU node 2: Node 2 allocates BufferA and BufferB; while the activation value of the "lane detection" task is being written to BufferA via PCIe, the computing unit of node 2 is idle and waiting; once BufferA is full, node 2 immediately locks BufferA to perform inference calculations for the 51st layer network; at the same time, the next checkpoint data begins to be written to BufferB.
[0139] Please see Figure 6 , Figure 6 This is a schematic diagram of the structural composition of the task scheduling system for a heterogeneous intelligent computing cluster in an embodiment of the present invention; the task scheduling system for the heterogeneous intelligent computing cluster is applied to the above-mentioned task scheduling method for heterogeneous intelligent computing clusters; the task scheduling system for the heterogeneous intelligent computing cluster includes:
[0140] The health prediction module 21 is used to query the corresponding configuration space of the heterogeneous intelligent computing cluster, determine the connection relationship between each computing node and the data input port, construct a path cost matrix based on the connection relationship, and collect normal data of each computing node under standard working conditions to construct the corresponding health prediction model.
[0141] The health score module 22 is used by the heterogeneous intelligent computing cluster to collect multiple time-series data from each computing node through the IPMB bus, determine its characteristics within the time window based on multiple time-series data, and determine the corresponding health score of each feature and each computing node in the health prediction model.
[0142] The candidate node module 23 is used to collect new tasks, the heterogeneous intelligent computing cluster parses the new tasks and outputs the corresponding data content, and determines the preferred set and the candidate set based on the data content, the corresponding computing nodes and health scores; the best candidate node is determined based on the preferred set and the candidate set, and the new task is executed based on the best candidate node;
[0143] The new node module 24 is used to dynamically monitor the health score of each computing node, trigger corresponding node alerts based on the health score of each computing node, identify risk nodes, migrate the task content corresponding to the risk node, and schedule the task content to the new node.
[0144] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A task scheduling method for a heterogeneous intelligent computing cluster, characterized in that, include: The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and data input port. Based on the connection relationship, it constructs a path cost matrix. At the same time, it collects normal data of each computing node under standard operating conditions and constructs the corresponding health prediction model. The heterogeneous intelligent computing cluster collects multiple time-series data from each computing node through the IPMB bus, determines its characteristics within the time window based on the multiple time-series data, and determines the corresponding health score for each feature and each computing node in the health prediction model. Upon receiving a new task, the heterogeneous intelligent computing cluster analyzes the new task and outputs the corresponding data content. Based on this data content, the corresponding computing nodes, and the health score, it determines the preferred set and the candidate set. Based on the preferred set and the candidate set, it determines the best candidate node and executes the new task based on the best candidate node. The system dynamically monitors the health score of each computing node, triggers corresponding node alerts based on the health score of each computing node, identifies risk nodes, migrates the tasks corresponding to the risk nodes, and schedules the tasks to new nodes.
2. The task scheduling method for heterogeneous intelligent computing clusters according to claim 1, characterized in that, The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and data input port. Based on the connection relationship, a path cost matrix is constructed. Simultaneously, normal data from each computing node under standard operating conditions is collected, and a corresponding health prediction model is constructed, including: The heterogeneous intelligent computing cluster is marked. The heterogeneous intelligent computing cluster actively queries the configuration space and implements multi-dimensional topology awareness to identify each computing node. Based on the tracing of each computing node, the corresponding data input port is determined, and the corresponding connection relationship is determined along each computing node and the corresponding data input port. Network latency data from each computing node is collected. The corresponding path cost matrix is determined by combining the network latency data with the connection relationship. Furthermore, it is dynamically adjusted by combining historical traffic data. The path cost matrix dynamically adjusts the path cost according to the task type and system load to match the corresponding scheduling strategy.
3. The task scheduling method for heterogeneous intelligent computing clusters according to claim 2, characterized in that, The heterogeneous intelligent computing cluster queries the corresponding configuration space and determines the connection relationship between each computing node and data input port. Based on the connection relationship, it constructs a path cost matrix. Simultaneously, it collects normal data from each computing node under standard operating conditions and constructs a corresponding health prediction model. The system also includes: Real-time monitoring of each computing node, marking normal data of each computing node under standard operating conditions, training on each normal data, and constructing corresponding health prediction models based on environmental changes during the operation of each computing node.
4. The task scheduling method for heterogeneous intelligent computing clusters according to claim 1, characterized in that, The heterogeneous intelligent computing cluster collects multiple time-series data points from each computing node via the IPMB bus. Based on these data points, it determines the characteristics of each node within a time window. Each characteristic and each computing node is then used in a health prediction model to determine a corresponding health score, including: The heterogeneous intelligent computing cluster periodically collects data through the IPMB bus and manages the data for each computing node to collect multiple time-series data, including temperature, power consumption, ECC error, and PCIe link retransmission rate; based on the optimization of multiple indicators of multiple time-series data, the corresponding multiple coupling relationships are determined. To control harsh environments, corresponding electromagnetic interference factors are identified; each computing node performs time window detection on these electromagnetic interference factors and multiple coupling relationships, and outputs their characteristics within the time window.
5. The task scheduling method for heterogeneous intelligent computing clusters according to claim 4, characterized in that, The heterogeneous intelligent computing cluster collects multiple time-series data from each computing node via the IPMB bus, determines its features within a time window based on the multiple time-series data, and determines the corresponding health score for each feature and each computing node in the health prediction model. The system also includes: A preset health decay content is introduced, and the decay amount is marked along the health decay content and each feature. The health score of each computing node is calculated in combination with the health prediction model. The health score reflects the real state of the computing node under specific working conditions and environmental stress.
6. The task scheduling method for heterogeneous intelligent computing clusters according to claim 1, characterized in that, Upon receiving a new task, the heterogeneous intelligent computing cluster parses the new task and outputs corresponding data content. Based on this data content, the corresponding computing nodes, and health scores, it determines an optimal set and a candidate set. Based on the optimal set and the candidate set, it determines the best candidate node and executes the new task based on the best candidate node, including: New tasks are identified based on the monitoring of heterogeneous intelligent computing clusters; descriptors and historical execution patterns are determined based on the parsing of new tasks; and the corresponding output data content is determined based on multi-factor analysis of the descriptors and historical execution patterns of the new tasks. In each computing node, a corresponding set framework is determined based on the computing node and the corresponding data content. The matching coefficient of each computing node is determined by matching the health score of each computing node with the set framework. Based on the matching coefficient of each computing node, it is divided into the preferred set and the candidate set, and some abnormal nodes are excluded.
7. The task scheduling method for heterogeneous intelligent computing clusters according to claim 6, characterized in that, Upon receiving a new task, the heterogeneous intelligent computing cluster parses the new task and outputs corresponding data content. Based on this data content, the corresponding computing nodes, and health scores, a preferred set and a candidate set are determined. The optimal candidate node is then selected based on the preferred and candidate sets, and the new task is executed based on the optimal candidate node. The process also includes: The preferred set and the candidate set are input into the virtual environment space, and the computing power level of each computing node in the virtual environment space is determined to handle the new task. The best candidate node is determined based on the comparison of the computing power levels of each computing node. The best candidate node is given priority to execute the new task and performs multi-level execution on the new task.
8. The task scheduling method for heterogeneous intelligent computing clusters according to claim 1, characterized in that, The dynamic monitoring of the health score of each computing node, triggering corresponding node alerts based on the health score of each computing node, identifying risk nodes, migrating the tasks corresponding to the risk nodes, and scheduling the tasks to new nodes includes: The health score of each computing node is marked. The health score of each computing node changes dynamically as the heterogeneous intelligent computing cluster works. The change in the health score of each computing node is determined. Based on the change in the health score of the computing node and the computing load of the computing node, the corresponding early warning signal is determined. The node early warning of the risk node is triggered based on the early warning signal.
9. The task scheduling method for heterogeneous intelligent computing clusters according to claim 8, characterized in that, The process of dynamically monitoring the health scores of each computing node, triggering corresponding node alerts based on the health scores of each computing node, identifying risk nodes, migrating the tasks corresponding to the risk nodes, and scheduling the tasks to new nodes also includes: When a risk node's warning is triggered, the heterogeneous intelligent computing cluster immediately freezes the distribution of new tasks to the risk node and uses the PCIeDMA link to read the contents of all unfinished tasks on the risk node in parallel to migrate the contents of all unfinished tasks on the risk node. During the migration process, a double buffering mechanism is introduced so that the new node can continue to execute all unfinished tasks on the risk node.
10. A task scheduling system for a heterogeneous intelligent computing cluster, characterized in that, The task scheduling system of the heterogeneous intelligent computing cluster is applied to the task scheduling method of the heterogeneous intelligent computing cluster as described in any one of claims 1-9; The task scheduling system of the heterogeneous intelligent computing cluster includes: The health prediction module is used to query the corresponding configuration space of the heterogeneous intelligent computing cluster, determine the connection relationship between each computing node and data input port, construct a path cost matrix based on the connection relationship, and collect normal data of each computing node under standard operating conditions to construct the corresponding health prediction model. The health score module is used by heterogeneous intelligent computing clusters to collect multiple time-series data from each computing node via the IPMB bus, determine the characteristics of each time-series data within a time window based on multiple time-series data, and determine the corresponding health score of each feature and each computing node in the health prediction model. The candidate node module is used to collect new tasks, the heterogeneous intelligent computing cluster parses the new tasks and outputs the corresponding data content, and determines the preferred set and the candidate set based on the data content, the corresponding computing nodes and health scores; the best candidate node is determined based on the preferred set and the candidate set, and the new task is executed based on the best candidate node; The new node module is used to dynamically monitor the health score of each computing node, trigger corresponding node alerts based on the health score of each computing node, identify risk nodes, migrate the task content corresponding to the risk node, and schedule the task content to the new node.