Computing power resource dynamic scheduling system supporting trusted data exchange

By using a multi-layered architecture design and a dynamic scheduling system, the problems of heterogeneous hardware adaptation and trusted data exchange were solved, achieving efficient compatibility and trusted data exchange between heterogeneous hardware, improving computing efficiency and throughput, and reducing device adaptation costs.

CN121900930APending Publication Date: 2026-04-21AOJI (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AOJI (BEIJING) TECHNOLOGY CO LTD
Filing Date
2025-11-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing computing power scheduling systems have poor compatibility with heterogeneous hardware and lack a unified adaptation mechanism, resulting in high equipment matching costs. They are unable to meet the needs of efficient computing power utilization and reliable data exchange, and cannot achieve a balance between low latency and high throughput in complex scenarios.

Method used

It adopts a multi-layer architecture design, including a bottom hardware adaptation layer, a trusted data exchange layer, a core computing power optimization layer, and a dynamic scheduling layer. It realizes instruction set mapping and hardware profiling through a chip migration and adaptation module, combines distributed identity authentication, encrypted data transmission and blockchain evidence storage, adopts a low-latency inference system, an N-dimensional parallel system and heterogeneous memory management, and uses a reinforcement learning scheduling module for dynamic resource allocation.

Benefits of technology

It achieves full compatibility with heterogeneous hardware, reduces device adaptation costs by 60%, ensures the immutability and high reliability of data exchange, reduces inference latency by 40%, improves parallel efficiency by 35%, increases memory utilization to 90%, improves computing efficiency by 50%, and achieves throughput of 1200 tokens/s, realizing the lowest cost and highest efficiency computing power scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900930A_ABST
    Figure CN121900930A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power resource dynamic scheduling system supporting trusted data exchange. The computing power resource dynamic scheduling system comprises an underlying hardware adaptation layer, a trusted data exchange layer, a core computing power optimization layer, a dynamic scheduling layer and a monitoring and feedback layer, wherein the bottom hardware adaptation layer comprises a chip migration adaptation module and is used for realizing instruction set mapping and capability portraying of heterogeneous hardware; the core computing power optimization layer comprises a low-delay reasoning system, an efficient N-dimensional parallel system and a heterogeneous memory management system; and the dynamic scheduling layer outputs a computing power distribution strategy based on the capability portrait and the task characteristics. According to the scheme, through deep fusion of hardware adaptation, credibility guarantee, computing power optimization and intelligent scheduling, a full-stack computing power scheduling system in a credible scene can be constructed, and the scheduling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of computing power scheduling and trusted data processing technology. More specifically, this invention relates to a dynamic scheduling system for computing resources that supports trusted data exchange. Background Technology

[0002] With the deep integration of the digital economy and artificial intelligence technologies, computing power has become a core productive force supporting industrial upgrading, with applications covering diverse fields such as financial risk control, medical image analysis, real-time control of the industrial internet, and large-scale model training and inference. To meet the computing power needs of different scenarios, heterogeneous computing hardware systems are gradually becoming more widespread—CPUs, with their versatility, are suitable for complex logic processing; GPUs, relying on their parallel computing advantages, have become the main force in large-scale model training; TPUs optimize the energy efficiency of AI tasks through dedicated architectures; and FPGAs, with their programmability, are adapted to low-latency and highly customized scenarios. At the same time, the demand for cross-device and cross-entity computing power collaboration and data exchange has significantly increased: for example, the medical field needs to achieve the sharing of image data and computing power resources among multiple institutions to improve diagnostic accuracy; industrial scenarios need to ensure real-time production decisions through computing power scheduling between the device end and the cloud. Such scenarios not only require efficient allocation of computing power resources but also place stringent requirements on the security and reliability of data exchange.

[0003] Currently, the collaborative system for computing power scheduling and data exchange is still in the development stage: on the one hand, hardware manufacturers focus on improving the performance of single chips and lack a unified standard for heterogeneous hardware collaboration; on the other hand, the integration of data security technology and computing power scheduling technology is low, making it difficult to simultaneously meet the dual requirements of "efficient computing power utilization" and "trustworthy data flow", resulting in significant shortcomings of existing systems when dealing with complex scenarios.

[0004] Especially in practical applications of heterogeneous computing power scheduling and trusted data exchange, existing computing power scheduling systems suffer from poor compatibility with underlying hardware. There is a lack of a unified adaptation mechanism for heterogeneous chips such as CPUs, GPUs, TPUs, and FPGAs, resulting in high equipment matching costs. Furthermore, the lack of trusted data exchange capabilities makes data transmission and sharing susceptible to tampering and leakage risks. Consequently, computing power optimization relies on a single dimension, depending heavily on static resource allocation, making it difficult to simultaneously address low latency and high throughput requirements, and failing to achieve a balance between deployment costs and computing efficiency.

[0005] The aforementioned problems make it difficult for existing systems to meet the heterogeneous computing power scheduling requirements in high-reliability scenarios, thus hindering the large-scale application of computing resources and the secure release of data value. Therefore, there is an urgent need for a dynamic scheduling system for computing resources that supports trusted data exchange, in order to integrate dynamic scheduling solutions that combine chip adaptation, trusted exchange, and multi-dimensional computing power optimization. Summary of the Invention

[0006] In order to at least solve one or more of the technical problems mentioned above, the present invention proposes a dynamic scheduling system for computing resources that supports trusted data exchange in several aspects.

[0007] In a first aspect, the present invention provides a dynamic scheduling system for computing resources supporting trusted data exchange, comprising: a bottom-level hardware adaptation layer, a trusted data exchange layer, a core computing power optimization layer, a dynamic scheduling layer, and a monitoring and feedback layer; wherein, the bottom-level hardware adaptation layer includes a chip migration adaptation module for implementing instruction set mapping and capability profiling of heterogeneous hardware. The core computing power optimization layer includes a low-latency inference system, a high-efficiency N-dimensional parallel system, and a heterogeneous memory management system. The dynamic scheduling layer outputs computing power allocation strategies based on hardware profiling and task characteristics. The bottom-level hardware adaptation layer is connected to the core computing power optimization layer, the trusted data exchange layer is connected to both the core computing power optimization layer and the dynamic scheduling layer, and the dynamic scheduling layer communicates bidirectionally with the monitoring and feedback layer.

[0008] In some embodiments, the chip migration adaptation module includes an instruction set mapping unit and a hardware profiling unit; wherein, the instruction set mapping unit constructs a unified instruction set library across CPU, GPU, TPU, and FPGA, and configures the instruction translation latency to be ≤10μs. The hardware profiling unit collects pre-configured chip parameters in real time, updating at a frequency of ≥10 times / second.

[0009] In some embodiments, the trusted data exchange layer includes a distributed identity authentication module, a data encryption transmission module, and a blockchain evidence storage module. The distributed identity authentication module uses the SM9 algorithm for two-way authentication, the data encryption transmission module uses an SM4-SM2 combined encryption scheme, and the blockchain evidence storage module writes the exchange logs to the consortium blockchain.

[0010] In some embodiments, the low-latency inference system employs a dynamic quantization strategy, which can automatically adjust the model quantization precision to 4-bit, 8-bit, or 16-bit, and integrates a custom operator library containing 32 operators.

[0011] In some embodiments, the efficient N-dimensional parallel system constructs a three-dimensional parallel architecture of "task-data-model", which can allocate data parallel tasks to GPUs and model parallel tasks to FPGAs based on hardware profiles.

[0012] In some embodiments, the heterogeneous memory management system includes a data popularity prediction unit and a cache scheduling unit. The data popularity prediction unit employs the LSTM algorithm, and the cache scheduling unit implements high-speed cache binding for high-popularity data.

[0013] In some embodiments, the dynamic scheduling layer includes a task parsing module, a reinforcement learning scheduling module, and a policy execution module. The reinforcement learning scheduling module employs an improved algorithm based on Deep Q-Network (DQN), using hardware profiles and task characteristics as state inputs and computing power allocation policies as action outputs. The policy execution module, based on the allocation policy, realizes the dynamic binding and release of computing power resources.

[0014] In some embodiments, the task parsing module decomposes the artificial intelligence task and extracts characteristic parameters such as task type, data scale, and computational complexity.

[0015] In some embodiments, the monitoring and feedback layer uses a PID controller to collect operating parameters such as throughput, latency, and power consumption in real time, and feeds the parameters back to the dynamic scheduling layer.

[0016] In some embodiments, the dynamic scheduling layer employs an experience replay mechanism and target network technology during the training process of the reinforcement learning scheduling module to reduce the correlation and variance during the training process.

[0017] Through the aforementioned dynamic scheduling system for computing resources supporting trusted data exchange, this invention achieves full compatibility with CPUs, GPUs, TPUs, and FPGAs via dynamic instruction mapping and hardware profiling technology in the chip migration and adaptation module. This reduces device adaptation costs by 60%, addressing the pain point of traditional systems requiring "one hardware, one adaptation." Furthermore, in some embodiments, based on the dual protection of national cryptographic algorithms and blockchain evidence storage, data exchange is fully traceable and tamper-proof, with an identity authentication success rate ≥99.9% and a near-zero risk of data leakage, meeting the requirements of high-trust scenarios. Further, in some embodiments, a low-latency inference system reduces inference latency by 40% according to calculations, an efficient N-dimensional parallel system improves parallel efficiency by 35%, and a heterogeneous memory management system increases memory utilization to over 90%, resulting in a 50% improvement in overall computing efficiency. Even further, in some embodiments, a reinforcement learning scheduler enables precise allocation of computing resources, reducing deployment costs. Simultaneously, in single-card tests of training models, a throughput of over 1200 tokens / s can be achieved, realizing the goal of "lowest cost and highest efficiency." Therefore, the present invention integrates "chip migration adaptation + three-dimensional parallel optimization + reinforcement learning scheduling" for the first time, thereby constructing a full-stack computing power scheduling system in trusted scenarios and improving scheduling efficiency. Attached Figure Description

[0018] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0019] Figure 1 An exemplary structural block diagram of a dynamic scheduling system for computing resources according to an embodiment of the present invention is shown. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0022] Figure 1 An exemplary structural block diagram of a dynamic computing resource scheduling system 100 according to an embodiment of the present invention is shown. Figure 1 As shown in the figure, the system 100 of this embodiment of the invention provides a dynamic scheduling system for computing resources that supports trusted data exchange, including: a bottom hardware adaptation layer 101, a core computing power optimization layer 102, a trusted data exchange layer 103, a dynamic scheduling layer 104, and a monitoring and feedback layer 105. The bottom hardware adaptation layer includes a chip migration adaptation module 1011, used to implement instruction set mapping and capability profiling of heterogeneous hardware. The core computing power optimization layer includes a low-latency inference system 1021, a high-efficiency N-dimensional parallel system 1022, and a heterogeneous memory management system 1023. The dynamic scheduling layer outputs computing power allocation strategies based on hardware profiling and task characteristics. The bottom hardware adaptation layer is connected to the core computing power optimization layer, the trusted data exchange layer is connected to both the core computing power optimization layer and the dynamic scheduling layer, and the dynamic scheduling layer communicates bidirectionally with the monitoring and feedback layer.

[0023] The system embodiments provided by the present invention above achieve dynamic scheduling of heterogeneous computing power in trusted data exchange scenarios through the collaborative design of a five-layer architecture. The following description of the technical details and collaborative logic of each layer enables those skilled in the art to better understand the technical solution of the present invention.

[0024] First, a unified interaction interface for heterogeneous hardware is built at the underlying hardware adaptation layer. This includes:

[0025] The chip migration adaptation module employs a "Hardware Abstraction Layer (HAL) + Dynamic Mapping Engine" architecture to address the differences in instruction sets and capabilities between CPUs, GPUs, TPUs, and FPGAs. This includes an instruction set dynamic mapping unit and a hardware capability profiling unit. Details are as follows:

[0026] 1. Instruction Set Dynamic Mapping Unit: Based on extended LLVM (Low-Level Virtual Machine) intermediate representation (IR) technology, a cross-architecture instruction translation library is built. For native instructions from x86 (CPU), CUDA (GPU), TPU ISA (TPU), Verilog (FPGA), etc., pre-compiled instruction translation rules (containing 1200+ mapping relationships) are used to convert them into a unified system IR, and then the hardware driver layer generates the target instructions. To reduce translation latency, a dedicated ASIC acceleration unit is integrated, achieving a single instruction translation time of ≤10μs (test environment: Intel Xeon 8380 + NVIDIA A100).

[0027] 2. Hardware Capability Profiling Unit: Collects 12 core parameters in real time through a distributed sensor network, including:

[0028] - Computing power metrics: Peak computing power (TFLOPS, precision coverage of FP32 / FP16 / INT8), actual computing power utilization (%).

[0029] - Storage metrics: memory bandwidth (GB / s), L3 cache capacity (MB), heterogeneous storage latency (ns);

[0030] -Physical metrics: Real-time power consumption (W), core temperature (°C), PCIe link speed (GT / s);

[0031] - Reliability metrics: historical failure rate (times / hour), task interruption recovery time (ms).

[0032] An asynchronous DMA sampling mechanism (sampling interval 100ms) is adopted, combined with the Kalman filter algorithm to remove outliers, ensuring that the profile update frequency is ≥10 times / second, providing high-fidelity hardware status data for upper-layer scheduling.

[0033] Secondly, multi-dimensional collaborative improvements to computing efficiency can be implemented at the core computing power optimization layer, which may specifically include:

[0034] 1. Low-latency inference system: Employs a three-level optimization approach of "dynamic quantization - operator fusion - hardware binding". The quantization strategy is dynamically switched based on hardware profiles: 16-bit quantization is used for GPUs (supporting Tensor Cores), and 4-bit quantization is used for FPGAs (resource-constrained). KL divergence calibration (calibration dataset ≥1000 samples) ensures accuracy loss ≤1%. The operator fusion library includes 32 combinations such as convolution + batch normalization and attention + residual connections. It adapts to hardware instruction sets (such as GPU's wmma instruction and TPU's matrix unit instruction) through a TVM automatic code generation tool. Single operator fusion can reduce latency by 40% (test model: ResNet50).

[0035] 2. High-Efficiency N-Dimensional Parallel System: An innovative three-dimensional parallel architecture of "task-data-model". The task parallel layer prioritizes multi-user tasks according to their security level (levels 1-5); the data parallel layer dynamically partitions data based on hardware memory bandwidth (number of partitions = number of cores when GPU bandwidth ≥ 800GB / s, number of partitions = number of cores / 2 when TPU bandwidth ≤ 400GB / s); the model parallel layer, for Transformer-type models, splits the multi-head attention layer across the FPGA array (each FPGA processes 1-2 heads). Through a parallel decision model trained using reinforcement learning (input: hardware profile + task features, output: optimal parallel combination), parallel efficiency is improved by 35% (compared to static data parallelism).

[0036] 3. Heterogeneous Memory Management System: A closed-loop "prediction-scheduling-reclamation" mechanism is designed. An LSTM prediction model (input: access frequency in the past 10 seconds, data size, associated task type) predicts the data popularity in the next 100ms. High-popularity data (access probability > 70%) is preferentially stored in high-speed cache (e.g., GPU L2, FPGA on-chip RAM). Cache scheduling uses an improved LRU algorithm, combined with hardware memory latency (e.g., prioritizing small data blocks when DRAM latency > 100ns), resulting in a stable memory utilization rate of over 90% (test scenario: 10-way concurrent inference task).

[0037] Secondly, a comprehensive "authentication-encryption-proof storage" protection system is built through a trusted data exchange layer. Specifically, this may include:

[0038] 1. Distributed Identity Authentication Module: A two-factor authentication mechanism is designed based on the SM9 national cryptographic algorithm. The device generates a distributed identity identifier (DID) containing a hardware fingerprint (CPU ID + MAC address hash) and the user's private key. The authentication process is as follows:

[0039] Challenge phase: The server generates a 128-bit random number as the challenge value and sends it to the device through an encrypted channel.

[0040] Response phase: The device signs the challenge value with its SM9 private key, attaches the DID and returns the device certificate.

[0041] Two-way verification: After the server verifies the validity of the signature, it returns a challenge response to its own signature. Once the device verifies the signature, a trusted connection is established.

[0042] For weak network scenarios, a retransmission mechanism (up to 3 retransmissions) can be designed to ensure an authentication success rate of ≥99.9% (test sample of 100,000 times).

[0043] 2. Data Encryption Transmission Module: Employs a hybrid architecture of "SM4 symmetric encryption + SM2 key negotiation". Data is divided into 1MB blocks, each generating an independent SM4 session key (generated via a true random number generator). The key is encrypted using the receiver's SM2 public key and appended to the block header. The transport layer is based on the DPDK-accelerated TLS 1.3 protocol, combined with a hardware encryption engine (such as Intel QAT) to achieve an encryption throughput of ≥100Gbps, ensuring zero data leakage during transmission.

[0044] 3. Blockchain Evidence Storage Module: A network of evidence storage nodes is designed based on the Hyperledger Fabric consortium blockchain. Participants can include computing power providers, data requesters, and regulatory nodes. Evidence storage information can include: data exchange hash (SHA-256), timestamp (UTC±1ms), both parties' DIDs, task ID, and hardware resource identifier. By adopting the PBFT consensus mechanism (block time ≤50ms when the number of nodes ≤20), the exchange records are ensured to be tamper-proof and traceable.

[0045] Next, the dynamic scheduling layer makes intelligent decisions to achieve optimal resource allocation. The reinforcement learning scheduler is designed based on the Deep Deterministic Policy Gradient (DDPG) algorithm, and its core process is as follows:

[0046] 1. Task Analysis: Extract task feature vectors, including: task type (training / inference), data size (GB), computational complexity (FLOPs), security level (level 1-5), deadline (ms), and trust requirements (whether proof of identity is required).

[0047] 2. State modeling: A 30-dimensional state space is constructed by integrating hardware profile (12-dimensional), task characteristics (6-dimensional), and system load (current computing power utilization and memory usage).

[0048] 3. Decision Optimization: The action space includes chip type selection (4 types), core allocation (10%-100%), memory allocation (10%-90%), and parallel strategy (three-dimensional parallel parameters). The reward function can be designed as follows:

[0049] Reward = α×(throughput / target throughput) + β×(1-deployment cost / baseline cost) - γ×(latency / target latency)

[0050] In the above formula, α, β, and γ are weights (for example, γ=0.6 when the security level is 5 and β=0.5 in a normal scenario). The scheduling strategy converges through offline training (100,000+ task samples) and online fine-tuning.

[0051] 4. Strategy execution: Scheduling instructions are issued through the gRPC interface, and the hardware driver layer converts the instructions into specific operations (such as GPU core binding and FPGA configuration loading), with a response time of ≤100ms (test environment: 100-node cluster).

[0052] Finally, stability is ensured through real-time closed-loop control via a monitoring and feedback layer. Specifically, this may include:

[0053] A distributed monitoring agent (deployed on each node) is used to collect system operation data (throughput, latency, power consumption, etc., sampling frequency 100Hz), and the time-series data is stored in the InfluxDB database. The PID controller dynamically adjusts according to the deviation: when the single-card training throughput is <1200 tokens / s (such as the LLaMA-7B model), the proportional gain Kp is increased by 0.2 to accelerate computing power expansion; when the power consumption is >80% of the rated value, the integral time Ti is extended to 2s to suppress excessive throttling and ensure stable operation of the system under dynamic load.

[0054] According to the above-described embodiments of the present invention, the deep integration of hardware adaptation standardization, data exchange credibility, and intelligent computing power scheduling significantly differs from the single-dimensional optimization schemes in the prior art, especially in cross-hardware collaboration, high-reliability scenario adaptation, and multi-objective dynamic balancing.

[0055] To help those skilled in the art better understand the logical relationships between the layers, the following embodiment illustrates the logic between the layers:

[0056] To further optimize the technical effectiveness of the solution, the core lies in strengthening the "real-time linkage-closed-loop feedback-dynamic adaptation" logic between layers. This involves eliminating inter-layer barriers through efficient information flow, enabling real-time matching of hardware characteristics, trusted status, computing power requirements, and scheduling strategies. The following explanation will focus on the logical flow relationships and optimization directions.

[0057] like Figure 1 As shown, each layer achieves bidirectional information flow interaction through standardized interfaces, forming a closed-loop logic of "perception-decision-execution-feedback". The specific logical flow relationship (including the closed-loop information flow) of each layer is as follows:

[0058] 1. Relationship between the underlying hardware adaptation layer and other layers.

[0059] - Output: The chip migration and adaptation module generates a "real-time hardware profile" (12 parameters, updated at a frequency of ≥10 times / second), which is then sent synchronously to the core computing power optimization layer (for computing power optimization strategy adaptation) and the dynamic scheduling layer (as the hardware basis for scheduling decisions).

[0060] - Input: Receives "resource allocation instructions" (such as core utilization and memory partitioning) from the dynamic scheduling layer and converts them into hardware-executable operations through the instruction set mapping unit.

[0061] 2. Trusted Data Exchange Layer ↔ Other Layers Relationship

[0062] - Output: Send the "identity authentication result" (pass / fail), "data encryption status" (encryption strength, transmission rate), and "blockchain evidence storage progress" (delay, integrity) to the dynamic scheduling layer (affecting task priority: tasks that fail authentication are suspended, and high-security-level tasks are scheduled first).

[0063] - Input: Receive the "trust level requirement" from the dynamic scheduling layer (e.g., Level 5 security tasks require mandatory evidence storage), and trigger the corresponding encryption / evidence storage strategy (e.g., SM4+SM2 encryption, PBFT consensus acceleration).

[0064] 3. Core computing power optimization layer ↔ Other layers relationship

[0065] - Input: Receive the "hardware profile" (such as GPU memory bandwidth, FPGA logic resources) from the underlying hardware adaptation layer and the "optimization target instructions" (such as latency ≤50ms, throughput ≥1200 tokens / s) from the dynamic scheduling layer, and adjust low-latency inference (quantization accuracy), N-dimensional parallelism (sharding granularity), and heterogeneous memory management (caching strategy).

[0066] - Output: Send actual optimization metrics (such as inference latency and parallel efficiency) to the monitoring and feedback layer (for state assessment).

[0067] 4. Dynamic Scheduling Layer ↔ Other Layers Relationship

[0068] - Input: The underlying hardware profile (hardware capabilities), trusted data exchange layer status (security constraints), task parsing results (task requirements), and monitoring and feedback layer data (system operation deviations) are integrated as the decision input for the reinforcement learning scheduler.

[0069] - Output: Generate "computing power allocation strategy" (hardware selection, resource ratio, optimization parameters) and distribute it to the underlying hardware adaptation layer (resource binding) and the core computing power optimization layer (optimization direction).

[0070] 5. Monitoring and Feedback Layer ↔ Other Layers Relationship

[0071] - Input: Collect operational data from each layer (underlying hardware load, trusted exchange latency, core computing throughput, and dynamic scheduling response time).

[0072] - Output: The PID controller generates "correction signals" (such as triggering computing power expansion when throughput is insufficient, or throttling when power consumption exceeds the limit), which are fed back to the dynamic scheduling layer (adjusting the reward function weight) and the core computing power optimization layer (temporarily improving the parallel granularity).

[0073] Furthermore, in another embodiment, the following optimizations can also be made:

[0074] 1. Real-time linkage between hardware profiling and computing power optimization

[0075] The underlying hardware adaptation layer adds "emergency state warning" (such as sudden rise in GPU temperature or sudden drop in memory bandwidth), and the core computing power optimization layer responds immediately: if the GPU temperature is >85℃, the low-latency inference system automatically switches from 16-bit quantization to 8-bit (reducing computational load), and the heterogeneous memory management system releases, for example, 30% of non-critical cache (reducing power consumption), thereby improving hardware stability.

[0076] 2. Dynamic intervention of trusted states in scheduling

[0077] When the trusted data exchange layer detects a "sudden increase in encrypted transmission latency" (e.g., from 5ms to 50ms), it notifies the dynamic scheduling layer in real time: the scheduler immediately switches the task of the link to backup hardware (with equivalent encryption capabilities) and temporarily increases the "communication efficiency weight" (e.g., β in the reward function increases from 0.3 to 0.5) to ensure that the real-time performance of trusted exchange is not affected.

[0078] 3. Predictive applications of monitoring data

[0079] The monitoring and feedback layer introduces an LSTM prediction model (such as based on historical 30-minute data) to predict "hardware load trends" for example, 10 seconds in advance (e.g., TPU utilization will reach 90% in 10 seconds). The dynamic scheduling layer pre-migrates some tasks to idle GPUs to avoid resource congestion, thereby reducing task response latency.

[0080] 4. Standardization and acceleration of inter-layer interfaces

[0081] The data interaction format of each layer is unified (FlatBuffers is used to replace JSON, which improves serialization speed by 5 times), and dedicated data forwarding chips (such as FPGA-accelerated interface conversion units) are deployed to reduce the inter-layer data transmission latency from 10ms to 1ms, thereby enhancing real-time performance.

[0082] Through the logical optimization of the above embodiments, each layer is upgraded from "independent operation" to "collaborative closed loop", which can further improve the hardware adaptation response speed, reduce the trusted data exchange interruption rate, and improve the stability of single-card training throughput (stabilizing above 1400 tokens / s), ultimately achieving the technical effect of "higher reliability, lower latency, and better cost".

[0083] Furthermore, in one embodiment, the chip migration adaptation module includes an instruction set mapping unit and a hardware profiling unit. The instruction set mapping unit constructs a unified instruction set library across CPU, GPU, TPU, and FPGA, and configures the instruction translation latency to be ≤ a preset time (e.g., a preset time of 10 μs). The hardware profiling unit collects pre-configured chip parameters in real time, updating at a frequency ≥ a preset frequency (e.g., a preset frequency of 10 times / second).

[0084] To better understand the above-described embodiments of the present invention, the embodiments are described in detail below. As can be seen from the above embodiments, the solution focuses on the core module of the underlying hardware adaptation layer—the chip migration adaptation module. Through a collaborative design of "layered instruction mapping architecture + high-precision hardware profiling mechanism," it overcomes the compatibility challenges of heterogeneous hardware (CPU, GPU, TPU, FPGA), providing underlying support for upper-layer trusted data exchange and computing power scheduling. Specific technical details are as follows:

[0085] I. Instruction Set Mapping Unit: Enables "zero-barrier" cross-architecture instruction conversion.

[0086] The instruction set mapping unit adopts a three-tier architecture: "hardware-specific submodule - unified IR conversion layer - accelerated execution layer." Differentiated adaptation strategies are designed for the instruction characteristics of different hardware, and multi-dimensional optimization ensures instruction translation latency is ≤10μs. The specific three-tier architecture design is as follows:

[0087] 1. Hardware-specific submodule design: Dedicated adaptation logic is built to address the instruction set differences of the four types of core hardware.

[0088] - CPU (taking x86-64 architecture as an example): Develop a Complex Instruction Set Computing (CISC) decomposition submodule to break down x86 multi-cycle instructions (such as MOVAPS, VADDPS) into 3-5 atomic operations, and then map them to a unified IR to avoid scheduling delays caused by instruction cycle differences.

[0089] -GPU (taking NVIDIA CUDA architecture as an example): Design a SIMT (Single Instruction Multithreaded) instruction adaptation submodule to convert CUDA block and warp instruction parameters into "parallel granularity identifiers" in a unified IR, ensuring that the parallel characteristics of the GPU are not lost.

[0090] -TPU (using Google TPU v4 as an example): For its dedicated matrix computation instructions (such as MXU instructions), develop a matrix operation mapping submodule to align the TPU's "tensor dimension-computation precision" parameters (such as bf16 precision, 1024×1024 matrix) with the tensor descriptors of the unified IR.

[0091] -FPGA (taking Xilinx UltraScale as an example): Considering its custom instruction characteristics, a configurable instruction parsing submodule is designed to extract the opcode and number of operands of custom instructions by reading the FPGA's configuration file (.bit file) to generate dynamic mapping rules.

[0092] 2. Unified IR Transformation Layer Optimization: A transformation library is built based on the extended LLVM IR (with the addition of a "hardware feature tag field"), pre-stored with over 1500 cross-architecture mapping rules (covering mainstream instructions for four types of hardware). To reduce transformation time, a "pre-compiled caching + dynamic update" mechanism is adopted: frequently used instruction mapping rules (such as GPU convolution calculation instructions and CPU logical operation instructions) are pre-compiled into binary code and stored in a high-speed cache (L2 cache), maintaining a cache hit rate of over 95%; when the hardware model is updated, the mapping rule library is updated via OTA (over-the-air download) without requiring a system restart.

[0093] 3. Accelerated Execution Layer Ensures Latency Performance: An integrated dedicated ASIC acceleration unit (based on 28nm process) is designed with a pipelined circuit for the "match-conversion-output" three-step operation of instruction translation, completing the conversion of one instruction per cycle. Simultaneously, an asynchronous instruction processing mechanism is employed to decouple instruction translation from hardware execution—translated instructions are temporarily stored in a FIFO buffer (depth 1024) to avoid translation blocking due to hardware busy / idle states. In the test environment (Intel Xeon 8380 CPU + NVIDIA A100 GPU + Google TPU v4), the average translation time per instruction is 8.2μs, meeting the ≤10μs requirement.

[0094] II. Hardware Profiling Unit: Constructing a high-fidelity real-time hardware state model.

[0095] The hardware profiling unit collects 12 core chip parameters in real time through a process of "distributed sampling - multi-algorithm processing - lightweight transmission," ensuring an update frequency of ≥10 times / second. This provides accurate hardware status information for upper-layer scheduling. The specific implementation is as follows:

[0096] 1. Definition and Classification of 12 Parameters: Based on the requirements of trusted computing power scheduling, the parameters are divided into four categories. The collection logic and application significance of each category of parameters are as follows:

[0097] - Computing power parameters (3 items): Peak computing power (unit: TFLOPS, covering three precision classes: FP32 / FP16 / INT8), real-time computing power utilization (unit: %), and task queue length (unit: tasks). These are used to determine whether the hardware's computing power matches the task requirements. For example, for large model training tasks with high computing power requirements, GPUs with peak computing power ≥ 200 TFLOPS (FP16) are given priority.

[0098] - Storage parameters (3 items): memory bandwidth (unit: GB / s), L3 cache capacity (unit: MB), heterogeneous storage latency (unit: ns, such as the interaction latency between GPU memory and CPU memory), which support the prefetching decision of the heterogeneous memory management system. For example, for TPUs with memory bandwidth <100 GB / s, the frequency of data interaction needs to be reduced.

[0099] - Physical parameters (3 items): Real-time power consumption (unit: W), core temperature (unit: °C), PCIe link speed (unit: GT / s, such as 16 GT / s for PCIe 4.0), used to avoid hardware overload, for example, dynamically reducing the computing power allocation ratio when the core temperature is >85°C.

[0100] - Reliability parameters (3 items): historical failure rate (unit: times / hour), task interruption recovery time (unit: ms), data transmission error rate (unit: ‰), adapted to trusted data exchange scenarios, for example, prioritizing hardware with a historical failure rate of <0.1 times / hour to process sensitive data.

[0101] 2. Parameter Acquisition and Processing Mechanism: A distributed sampling architecture is adopted, deploying a lightweight sampling agent (consuming <5% of CPU resources) on each hardware device to achieve efficient data acquisition through the following methods:

[0102] - Computing power / storage parameters: Call the underlying API provided by the hardware manufacturer (such as NVIDIA NVML, Intel PCM), and set the sampling interval to 100ms.

[0103] - Physical / reliability parameters: Read directly through hardware sensors (such as temperature sensors and power meters), combined with DMA (direct memory access) technology to avoid CPU intervention, with a sampling interval of 50ms.

[0104] - Data processing: Kalman filtering algorithm is used to remove sampling noise (such as instantaneous power consumption fluctuations), and outliers (such as sudden increases in error rate) are identified through the "3σ criterion" to ensure that the parameter accuracy is ≥99.5%.

[0105] - Lightweight transmission: The 12 parameters are encapsulated into a 64-byte binary data packet (60% compressed compared to JSON format) and transmitted to the image aggregation node via UDP protocol. The aggregation node stores the data in the format of "hardware ID-timestamp-parameter set" and generates a complete image every 100ms, meeting the requirement of an update frequency of ≥10 times / second.

[0106] Through the above design, the chip migration and adaptation module can be distinguished from the existing "static instruction adaptation + low-frequency parameter acquisition" solution, and realize dynamic compatibility and high-fidelity status awareness of heterogeneous hardware. This lays the foundation for accurate scheduling of computing power in trusted data exchange scenarios. For example, in the scenario of cross-institutional processing of medical data, devices with "high reliability (low failure rate) + high computing power (meeting the needs of image inference)" can be selected based on hardware profiles. At the same time, instruction mapping can be used to ensure that heterogeneous hardware from different institutions can work together.

[0107] Furthermore, in one embodiment, the trusted data exchange layer includes a distributed identity authentication module, a data encryption transmission module, and a blockchain evidence storage module. The distributed identity authentication module uses the SM9 algorithm for two-way authentication, the data encryption transmission module uses an SM4-SM2 combined encryption scheme, and the blockchain evidence storage module writes the exchange logs to the consortium blockchain.

[0108] In the trusted data exchange layer of the above embodiment, a full-link mechanism of "distributed identity authentication - hybrid encrypted transmission - blockchain notarization" ensures the trustworthiness of data exchange while adapting to the real-time requirements of heterogeneous computing power scheduling. The specific technical solution is described below:

[0109] I. Distributed Identity Authentication Module: Based on the SM9 algorithm, this module employs a two-way trusted mechanism and a two-factor authentication architecture combining hardware fingerprinting and national cryptographic algorithms to address cross-device and cross-subject identity forgery issues. The process and technology optimizations are as follows:

[0110] 1. Identity Generation: Assign a distributed identity (DID) to each participant (device / user), which contains three parts of information:

[0111] - Unique hardware fingerprint: The SHA-256 algorithm is used to hash physical features such as CPU serial number, MAC address, and FPGA chip ID to generate a 256-bit unalterable hardware identifier.

[0112] - User key pair: Based on the SM9 identifier cryptography algorithm, the system key generation center (KGC) assigns a private key to the user (bound to the DID), and the public key is generated by deducing the private key.

[0113] - Trusted Certificate: Includes DID, hardware fingerprint digest, and validity period (default 1 year), and is issued by a regulatory node in the consortium blockchain.

[0114] 2. Two-way authentication process: A three-step interaction mechanism is designed for the computing power requester and the provider as follows:

[0115] - Challenge Phase: The server generates a 128-bit random number (generated by a hardware true random number generator) as the challenge value, attaches its own DID and certificate, and sends it to the requester.

[0116] - Response phase: After verifying the validity of the server certificate (verifying the signature of the supervisory node), the requester signs the challenge value with its own SM9 private key, generates a new challenge value, and returns it along with its own DID, certificate and signature.

[0117] - Confirmation Phase: The server verifies the requester's signature and certificate. If successful, it returns a signature for the new challenge value, and both parties complete mutual trust.

[0118] For high-concurrency scenarios (such as 1000+ concurrent requests), the module is designed with an authentication result cache pool (TTL=300s). Duplicate authentication requests directly hit the cache, reducing the response time from 50ms to less than 10ms.

[0119] II. Data Encryption Transmission Module: The module adopts a highly efficient and secure SM4-SM2 combined encryption scheme, and innovatively uses a "symmetric encryption + asymmetric key negotiation" architecture to ensure data confidentiality while meeting the low latency requirements of computing tasks.

[0120] 1. Key management mechanism:

[0121] - Session key generation: A 128-bit session key is generated using the SM4 algorithm. The key is automatically rotated every 1GB of data transmitted. The key is generated through a noise source (based on resistor thermal noise), and its randomness conforms to the NIST SP 800-22 standard.

[0122] - Key distribution: The SM2 elliptic curve algorithm is used for key negotiation. The sender encrypts the session key with the receiver's public key, and the receiver decrypts it with the private key, thus avoiding key leakage during transmission.

[0123] 2. Data encryption and transmission optimization:

[0124] - Fragmented Encryption: Data is fragmented into 1MB blocks, each with a 32-bit CRC checksum, and encrypted using SM4-CBC mode (initial vector IV is randomly generated) to ensure that single-block tampering can be detected.

[0125] - Hardware acceleration: The integrated Intel QAT (QuickAssist Technology) encryption engine increases SM4 encryption throughput to 200Gbps (10 times faster than software implementation), while DPDK technology bypasses the kernel protocol stack, reducing transmission latency to 5ms (10Gbps link).

[0126] - Anti-replay design: Each data packet is appended with a monotonically increasing sequence number (32 bits), the receiver maintains a sequence number window (default size 1000), and discards duplicate or expired packets to resist replay attacks.

[0127] III. Blockchain Evidence Storage Module: A full-chain traceability mechanism based on a consortium blockchain. The module builds an evidence storage network based on Hyperledger Fabric to achieve immutable recording of the data exchange process. The specific design is as follows:

[0128] 1. Evidence storage nodes and consensus mechanisms:

[0129] - Node composition: Includes computing power provider nodes, data demander nodes, and regulatory nodes (such as industry regulatory agencies). Node access requires verification through the identity authentication module.

[0130] - Consensus optimization: The improved PBFT algorithm is adopted to shorten the view switching time to 200ms (default number of nodes 10), the block interval is set to 50ms, and each block can store a maximum of 100 evidence records.

[0131] 2. Evidence storage data structure and interaction:

[0132] - Evidence content includes data hash (SHA-256 processing), DID of both parties exchanging data, hardware resource ID (associated with underlying hardware profile), timestamp (accurate to microseconds), and task ID (associated with dynamic scheduling layer).

[0133] - Lightweight storage: Raw data is not written to the blockchain, only hashes and metadata are stored, and the size of a single record is controlled within 512 bytes, reducing the storage pressure on the chain.

[0134] - Source tracing interface: Provides a RESTful API for upper-layer systems to query evidence records, supports searching by task ID and time range, and has a query response time of ≤100ms.

[0135] This module addresses identity trust issues through SM9 two-way authentication (authentication success rate ≥99.9%), achieves zero-leakage transmission through SM4-SM2 combined encryption (encryption throughput ≥200Gbps), and ensures process traceability through consortium blockchain notarization (notarization latency ≤50ms). These three elements work together to form a trusted closed loop of "pre-authentication - in-process encryption - post-notarization." Compared to existing single encryption or static authentication schemes, this solution achieves, for the first time, deep adaptation of a trusted mechanism to heterogeneous computing power scheduling. It not only meets the high security requirements of scenarios such as finance and healthcare but also ensures the real-time performance of computing tasks through hardware acceleration and process optimization, providing secure and reliable data source support for the dynamic scheduling layer.

[0136] Furthermore, in one embodiment, the low-latency inference system employs a dynamic quantization strategy, which can automatically adjust the model quantization precision to 4-bit, 8-bit, or 16-bit, and integrates a custom operator library containing 32 operators.

[0137] The above embodiments target low-latency inference systems for the core computing power optimization layer. Through a collaborative design of "hardware-aware dynamic quantization + instruction-level optimized custom operator library", the inference latency of heterogeneous hardware (GPU, FPGA, TPU) is significantly reduced while ensuring inference accuracy. Specific technical details are as follows:

[0138] I. Dynamic Quantization Strategy: Precision Adaptive Adjustment Based on Hardware Profile. The core of dynamic quantization is to switch the quantization precision (4-bit / 8-bit / 16-bit) in real time based on the "hardware capability profile" output by the underlying hardware adaptation layer, avoiding precision loss or performance waste caused by "one-size-fits-all" quantization. The specific implementation includes two main modules:

[0139] 1. Hardware-Aware Quantization Accuracy Decision Mechanism. The system dynamically matches the optimal accuracy by pre-setting a "hardware characteristic-quantization accuracy" mapping rule base and combining it with real-time hardware status. The key logic is described as follows:

[0140] - GPU Adaptation (e.g., NVIDIA A100 / A10): If the hardware profile indicates support for Tensor Core instructions (FP16 / INT8 acceleration) and the current computing power utilization is <60%, 16-bit quantization will be automatically selected—this precision can fully activate the parallel computing capabilities of Tensor Cores while avoiding the precision loss of 4-bit quantization. If the computing power utilization is ≥80%, it will switch to 8-bit quantization to reduce memory bandwidth usage (reducing memory access by 30%) and alleviate hardware load.

[0141] - FPGA adaptation (such as Xilinx UltraScale+): Due to the limited logic resources of FPGA (usually only supporting quantization below 8 bits), if the "Logic Unit Utilization" in the hardware profile is >70%, select 4-bit quantization; if the resources are sufficient (utilization <50%), switch to 8-bit quantization to balance latency and accuracy.

[0142] - TPU adaptation (such as Google TPU v4): Decision based on "Matrix Computing Unit (MXU) utilization" based on hardware profile. When the utilization is <50%, 16-bit quantization is used (to leverage the high-precision computing power of MXU), and when the utilization is ≥70%, 8-bit quantization is used to improve the inference throughput per unit resource.

[0143] 2. Accuracy Calibration and Error Control Technology. To avoid inference errors exceeding limits due to quantization, the system employs a dual mechanism of "KL divergence calibration + outlier protection":

[0144] - KL divergence calibration: Select 1000-2000 representative samples (covering typical data distributions of inference tasks, such as lesion areas in medical images and long text fragments in NLP), calculate the KL divergence of feature distributions before and after quantization. If the divergence is >0.3 (indicating that the distribution difference is too large), automatically increase the precision by one level (e.g., 4bit → 8bit).

[0145] - Outlier protection: For "key feature channels" in the inference process (such as the confidence output channel of the classification task and the bounding box regression channel of the detection task), 16-bit precision is forcibly retained, and the overall precision loss is controlled to be ≤1% through "mixed precision quantization" (only non-key channels use low precision) (test benchmark: ImageNet dataset classification task, ResNet50 model).

[0146] II. Custom Operator Library Design: Instruction-Level Optimized Operator Fusion Scheme. The system constructs a custom operator library containing 32 core operators. It optimizes the fusion of operator combinations (such as convolution-batch normalization and attention-residual connections) that are computationally intensive and memory-intensive in inference tasks. This is achieved by eliminating intermediate variables, reducing data transfer, and lowering latency. Specific technical methods are as follows:

[0147] 1. Technical Implementation of Core Operator Fusion: Differentiated fusion logic is designed for combinations of three types of high-frequency operators. Key solution examples are provided.

[0148] - Convolutional (Conv) + Batch Normalization (BN) Fusion: In traditional inference, the Conv output needs to be written to memory and then read into the BN module, resulting in two memory accesses. After fusion, the "Conv-BN fused weights" are directly generated through formula derivation (embedding the mean, variance, and scaling factor of BN into the weight matrix of Conv), requiring only one memory access and reducing latency by 35% (test environment: NVIDIA A100, input feature map 224×224×3).

[0149] - Multi-Head Attention + LayerNorm Fusion: For Transformer-type models (such as BERT, LLaMA), the "QKV linear transformation" of attention calculation and the "mean / variance calculation" of subsequent LayerNorm are merged into a single operator. By reusing intermediate results through registers (avoiding intermediate features from being written to global memory), a latency reduction of 42% is achieved on FPGA (Xilinx U280).

[0150] - Fusion of activation function (ReLU / SiLU) + pooling (MaxPool): During the operator compilation stage, the nonlinear calculation of the activation function is embedded into the "window comparison" process of pooling. For example, each element in the MaxPool window is first activated by ReLU and then the maximum value is taken, reducing the overhead of one operator call and shortening the single-step inference time by 18% on TPU v4.

[0151] 2. Hardware instruction-level binding optimization: To ensure that the fusion operator is compatible with the native instructions of heterogeneous hardware, the operator library achieves low-level binding through "TVM automatic code generation + hardware instruction template".

[0152] - GPU Adaptation: For NVIDIA GPU Tensor Core instructions (such as wmma::mma_sync), the computation logic of the Conv-BN fusion operator is split into 16×16×16 matrix blocks, and the Tensor Core instructions are directly called to complete parallel computation, improving the computing power utilization to 90% (30% improvement compared to the general CUDA implementation).

[0153] - FPGA adaptation: The fusion operator is encapsulated into a configurable IP core (such as Conv-BN fusion IP). The "number of logic units and DSP resources" parameters of the FPGA are obtained through hardware profiling, and the parallel granularity of the IP core is automatically adjusted (such as using 8-way parallelism when resources are sufficient and 4-way parallelism when resources are scarce).

[0154] - TPU Adaptation: Based on the TPU MXU instruction set, the "attention score calculation" of the Multi-Head Attention fusion operator is mapped to the matrix multiplication instruction of MXU. A single instruction completes the parallel calculation of 4 attention heads, improving throughput by 25%.

[0155] The above-described low-latency inference system embodiment achieves a triangular balance between accuracy, latency, and hardware adaptation through a two-layer optimization of "dynamic quantization + operator fusion," resulting in the following technical effects:

[0156] 1. Significantly reduced latency: In tests of ResNet50 (image classification) and BERT-base (text inference) models, the inference latency in the GPU environment is reduced by 40% compared to the traditional static 8-bit quantization, and in the FPGA environment it is reduced by 45% compared to the non-fusion solution, fully meeting the low latency requirements of industrial real-time inference (such as industrial quality inspection ≤50ms / frame) and edge computing (such as automotive AI ≤100ms / time).

[0157] 2. Controllable accuracy loss: Through KL divergence calibration and mixed precision quantization, the inference accuracy loss is stably controlled within 1% (ImageNet Top-1 accuracy ≥75.8%, BERT sentiment analysis F1 score ≥92.3%), avoiding task failure caused by low precision quantization.

[0158] 3. Strong hardware adaptability: Unlike the "dedicated quantization solution for a single hardware" in existing technologies, this system achieves cross-architecture adaptation of GPU / FPGA / TPU through hardware profiling linkage. There is no need to repeatedly develop quantization modules for a single hardware, reducing development costs by 50%, while ensuring that the inference performance of different hardware reaches the optimal level.

[0159] Furthermore, in one embodiment, the efficient N-dimensional parallel system constructs a three-dimensional parallel architecture of "task-data-model", which can allocate data parallel tasks to GPUs and model parallel tasks to FPGAs based on hardware profiles.

[0160] The above embodiments, based on the "task-data-model" three-dimensional parallel architecture and hardware adaptation logic of a limited high-efficiency N-dimensional parallel system, break through the computing power potential of heterogeneous hardware (GPU, FPGA, etc.) through the deep coupling of dynamic perception strategy and hardware characteristics. The specific technical details are described as follows:

[0161] I. Hierarchical Design and Collaboration Mechanism of Three-Dimensional Parallel Architecture. The embodiment system provided by this invention breaks through the limitations of traditional single-dimensional parallelism, constructing a three-dimensional parallel system of "task layer - data layer - model layer." Each layer achieves collaborative scheduling through a unified parallel controller. The core design is as follows:

[0162] 1. Task Parallelism Layer: Priority scheduling based on security levels. For multi-user, multi-type tasks (such as medical image inference and financial risk control training), the task parallelism layer implements end-to-end management of "classification-sorting-allocation":

[0163] - Task classification mechanism: Extract the security level (level 1-5, level 5 is the highest, such as patient privacy data processing), computation type (inference / training), and time sensitivity (deadline ≤100ms is highly sensitive) of the task, and generate a three-dimensional task label.

[0164] - Priority sorting algorithm: The improved EDF (Earliest Deadline First) algorithm is adopted, combined with dynamic weight adjustment based on security level. For each security level increase, the priority weight increases by 0.2 (e.g., the weight of a level 5 task is 1.0, and that of a level 3 task is 0.6), ensuring that highly reliable tasks get priority access to computing resources.

[0165] - Hardware pre-allocation logic: Based on the matching degree between task tags and hardware profiles (e.g., level 5 tasks prioritize matching hardware with a failure rate of <0.1 times / hour), tasks are initially allocated to GPU / FPGA / TPU clusters, providing a basic framework for the next level of parallelism.

[0166] 2. Data Parallelism Layer: Dynamic sharding strategy based on memory bandwidth. For massive input data of the same task (such as trillions of tokens for training a large model), the data parallelism layer dynamically adjusts the sharding granularity based on hardware memory bandwidth. Core technologies include:

[0167] - Shard Count Calculation Model: The number of shards N is set as N = α × (hardware memory bandwidth / baseline bandwidth), where α is the task data size coefficient (1-10), and the baseline bandwidth is set to 200GB / s (based on typical GPU values). For example:

[0168] - GPU (memory bandwidth 800GB / s): N=α×(800 / 200)=4α. If α=4 (data size 100GB), then divide into 16 chips to make full use of the high bandwidth advantage.

[0169] - FPGA (100GB / s memory bandwidth): N=α×(100 / 200)=0.5α, to avoid data interaction delays caused by excessive fragmentation due to insufficient bandwidth.

[0170] - Slice synchronization optimization: The "gradient compression + asynchronous update" mechanism is adopted. In the GPU cluster, the All-Reduce communication of slice gradient is implemented through the NCCL library (latency ≤2ms). Due to resource constraints, the FPGA cluster uses quantization compression (INT8) gradient synchronization, which reduces the communication volume by 75%.

[0171] 3. Model Parallel Layer: Based on hardware architecture layer splitting rules, the model parallel layer targets large neural networks (such as the 100+ layer structure of Transformer). It splits model layers according to hardware computing characteristics to achieve inter-layer parallel computing. The key strategies are as follows:

[0172] - GPU model parallelism limitation: Because GPUs are better at data parallelism (SM multi-core is suitable for multi-data processing in the same layer), model parallelism is only enabled when the number of model parameters is greater than the hardware memory (such as A100 80GB). The split granularity is "4 consecutive layers" (such as the first 4 layers of Transformer are allocated to GPU0, and the last 4 layers to GPU1). Inter-layer features are transmitted through PCIe 4.0 link (bandwidth 32GB / s).

[0173] - Parallel Optimization of FPGA Model: Leveraging the customizable computing capabilities of FPGA's programmable logic units (LUTs), the model is broken down into "fine-grained functional blocks"—for example, the Transformer's multi-head attention layer is divided into three functional blocks: "Q-linear transformation," "K-linear transformation," and "attention score calculation." These blocks are mapped to three FPGAs (each responsible for one functional block), and data transfer is achieved through high-speed inter-chip interfaces (such as SRIO, with a rate of 10Gbps), reducing single-layer computation latency by 40%.

[0174] - Inter-layer dependency handling: Directed acyclic graphs (DAGs) are used to describe the dependencies between model layers (e.g., convolutional layer → activation layer is a strong dependency). The parallel controller ensures that dependent layers are executed serially and non-dependent layers (e.g., the two branches of ResNet) are executed in parallel through topological sorting, improving resource utilization to 85%.

[0175] II. Hardware Profile-Driven Parallel Policy Decision System. To achieve the optimal match between "hardware characteristics and parallel dimensions," the system design employs a reinforcement learning-based parallel decision model. The core logic is as follows:

[0176] 1. State Input: Integrate hardware profile (6 key parameters such as memory bandwidth, number of cores, and proportion of logical resources) with task characteristics (data scale, number of model layers, and security level) to construct a 12-dimensional state vector.

[0177] 2. Action Space: Includes 3D parallel parameter combinations (number of parallel tasks, number of data shards, model splitting granularity), with a total of 200+ combinations.

[0178] 3. Reward function: Reward = 0.4 × (parallel efficiency) + 0.3 × (hardware utilization) + 0.3 × (task completion rate), where parallel efficiency = actual throughput / theoretical maximum throughput.

[0179] 4. Training and Deployment: The PPO (Proximity Policy Optimization) model is trained offline using 100,000+ task samples. During online deployment, the strategy is fine-tuned every 100ms based on the real-time hardware profile to ensure that the decision adapts to the dynamic state of the hardware (such as automatically reducing the number of data shards when the GPU suddenly downclocks).

[0180] Through the embodiments provided by this invention, the high-efficiency N-dimensional parallel system achieves three major technological breakthroughs through the collaboration of three-dimensional architecture and hardware-based perception and decision-making:

[0181] 1. Significantly improved parallel efficiency: In LLaMA-7B model training, the GPU cluster adopts a strategy of "data parallelism as the main method + model parallelism as the auxiliary method", which improves throughput by 35% compared with pure data parallelism; when the FPGA processes Transformer inference, the latency of single text inference by model parallelism is reduced from 80ms to 48ms.

[0182] 2. Balanced hardware utilization: By dynamically adjusting the parallel granularity, the utilization of the GPU's SM cores is increased from 60% to 90%, and the utilization of the FPGA's logic units is increased from 55% to 85%, avoiding resource idleness.

[0183] 3. Strong cross-hardware adaptability: Unlike the "single hardware dedicated parallel solution" in existing technologies, this system achieves seamless adaptation of GPU / FPGA / TPU through hardware profile-driven strategy decision-making. No manual configuration of parallel parameters is required, and the deployment efficiency in hybrid hardware clusters is improved by 60%, providing a high-efficiency computing power foundation for the dynamic scheduling layer.

[0184] Furthermore, in one embodiment, the heterogeneous memory management system includes a data popularity prediction unit and a cache scheduling unit. The data popularity prediction unit employs the LSTM algorithm, and the cache scheduling unit implements high-speed cache binding for high-frequency data.

[0185] The above embodiments, by defining the core components and technical implementation of the heterogeneous memory management system, solve the problems of low data interaction efficiency and resource waste between heterogeneous memory systems of CPU / GPU / TPU / FPGA through a collaborative mechanism of "LSTM data heat prediction + hardware-aware cache scheduling". Specific technical details are as follows:

[0186] I. Data Popularity Prediction Unit: Accurate Access Trend Prediction Based on LSTM. The data popularity prediction unit uses a Long Short-Term Memory (LSTM) network model to predict the access probability of data in future windows in real time, providing a basis for cache scheduling decisions. The core technology design is as follows:

[0187] 1. Structure and training of the LSTM prediction model. The model is customized for the "data access timing characteristics" of computing tasks. Key parameters and training mechanisms include:

[0188] - Input feature dimensions: A 12-dimensional input vector is constructed by fusing 6 types of data attributes, including: access frequency in the past 10 seconds (times / s), data block size (MB), last access time (ms since the current time), associated task type (training / inference), task priority (levels 1-5), and memory bandwidth of the hardware node (GB / s, taken from the underlying hardware profile).

[0189] - Network structure: A 3-layer LSTM architecture is adopted (with 64, 32 and 16 hidden units respectively), followed by a fully connected layer to output the "access probability in the next 100ms" (0-100%), and the result is normalized by the Sigmoid activation function.

[0190] - Training strategy: Offline training with 500,000+ task samples (covering scenarios such as image classification and large model training), with cross-entropy as the loss function (labeled by whether an actual access has occurred); Incremental learning is enabled during online deployment (fine-tuning with new samples every hour) to ensure that the model adapts to dynamic access patterns (such as a surge in data access for sudden inference tasks).

[0191] 2. Prediction accuracy optimization mechanism. To avoid cache resource mismatch caused by prediction bias, the unit is designed with a "multi-scale verification + anomaly correction" strategy:

[0192] - Multi-scale verification: Simultaneously output the access probability of 3 time windows (50ms / 100ms / 200ms), and take their weighted average (the weight is dynamically adjusted according to the real-time nature of the task, such as 0.6 for the 100ms window in high real-time tasks), reducing the randomness of single-window prediction.

[0193] - Anomaly Correction: When the predicted access probability deviates from the actual access probability by more than 30% (for 3 consecutive times), feature importance analysis (calculated by SHAP value) is automatically triggered. If the contribution of the "task priority" feature is less than 10%, its weight is temporarily increased (from 0.1 to 0.2) to correct the model's sensitivity to high-priority tasks.

[0194] II. Cache Scheduling Unit: Hardware-Aware High-Hot Data Binding Strategy. Based on predicted access probabilities, the cache scheduling unit dynamically binds high-hot data to high-speed caches in heterogeneous hardware (such as GPU L2 cache, FPGA on-chip RAM, and TPUHBM). Core technologies include:

[0195] 1. Cache Tiering and Hardware Adaptation Rules: Design differentiated caching strategies based on the storage architecture characteristics of different hardware.

[0196] - GPU cache scheduling: NVIDIA GPU's L2 cache (capacity 40-60MB) is suitable for storing medium-sized data blocks (1-8MB). When the data access probability is >70% and the size is ≤8MB, it is bound to the L2 cache. At the same time, the GPU's MIG (Multi-Instance GPU) feature is used to divide independent cache partitions (minimum granularity 256KB) for different tasks to avoid cache pollution between tasks.

[0197] - FPGA cache scheduling: FPGA on-chip RAM (capacity usually <10MB) resources are limited. Small data blocks (≤1MB, such as weight parameters of inference tasks) with an access probability >80% are prioritized and mapped directly to the on-chip RAM address space through the AXI bus, reducing the access latency from 100ns in DRAM to 10ns.

[0198] - TPU cache scheduling: TPU's high-bandwidth memory (HBM, bandwidth ≥ 400GB / s) is suitable for feature tensors of large model training. When the data access probability is > 60% and the size is > 32MB, it is bound to HBM and "tensor compression" (FP16 → BF16) is enabled to reduce storage usage. The compression ratio is 1:1.2 and the accuracy loss is negligible.

[0199] - CPU cache scheduling: For cross-hardware collaborative tasks (such as CPU preprocessing + GPU inference), intermediate data (access probability > 60%) is bound to the CPU L3 cache (capacity ≥ 50MB), and data interaction latency is reduced by 40% by avoiding writing to main memory through Intel DDIO technology (direct data I / O).

[0200] 2. Dynamic Cache Replacement Algorithm: To address the conflict problem due to limited cache capacity, the unit employs an improved LRU (Least Recently Used) algorithm, integrating hardware characteristics and data attributes:

[0201] - Replacement priority calculation: Priority = 0.4 × (access probability) + 0.3 × (data size reciprocal) + 0.3 × (hardware cache latency reciprocal). The lower the priority, the earlier it will be replaced. For example, in the GPU L2 cache, 1MB of low-probability data (30%) has a lower priority than 8MB of high-probability data (70%). Even if the former has been accessed recently, it may still be replaced.

[0202] - Batch replacement optimization: When the cache utilization rate is ≥90%, a batch replacement is triggered (5-10 data blocks are replaced at a time). The number of bus interactions is reduced by pre-calculating the replacement list, and the replacement time is controlled within 5ms (GPU environment).

[0203] Through the above-described embodiment of the heterogeneous memory management system provided by the present invention, the following technical effects are achieved through "precise prediction + hardware-adaptive scheduling":

[0204] 1. Significantly reduced memory access latency: In the ResNet50 inference task, GPU L2 cache binding reduces data access latency from 120ns to 35ns; FPGA on-chip RAM binding reduces weight access latency by 90% and overall inference latency by 50%.

[0205] 2. Significantly improved memory utilization: Through dynamic binding and intelligent replacement, GPU HBM utilization can be increased from 60% to 92%, and CPU L3 cache utilization can be increased from 55% to 88%, avoiding resource waste caused by "mixed storage of hot and cold data".

[0206] 3. Cross-hardware collaborative efficiency optimization: Unlike the existing "static cache partitioning" or "single hardware optimization" solutions, this system achieves heterogeneous memory collaboration between CPU / GPU / TPU / FPGA through LSTM prediction and hardware profiling. In multi-hardware joint training scenarios, data interaction efficiency can be improved by 65%, providing efficient storage support for the core computing power optimization layer.

[0207] Furthermore, in one embodiment, the dynamic scheduling layer includes a task parsing module, a reinforcement learning scheduling module, and a policy execution module. The reinforcement learning scheduling module employs an improved algorithm based on Deep Q-Network (DQN), using hardware profiles and task characteristics as state inputs and computing power allocation policies as action outputs. The policy execution module, based on the allocation policy, realizes the dynamic binding and release of computing power resources.

[0208] The above embodiments provide the core components of the dynamic scheduling layer and the reinforcement learning scheduling logic. Through a closed-loop design of "multi-dimensional task analysis + improved DQN scheduling decision + low-latency policy execution," the scheduling objective of "minimizing deployment cost + maximizing computational efficiency" is achieved, while also adapting to the security requirements of trusted data exchange scenarios. Specific technical details are described below:

[0209] I. Task Parsing Module: Multi-dimensional Feature Extraction and Quantization. The core of the task parsing module is to decompose the input computational task (such as large model training, medical image inference) into structured feature vectors, providing accurate input for reinforcement learning scheduling. Key technologies are implemented as follows:

[0210] 1. Feature Extraction Dimensions and Definitions: To address the needs of trusted computing power scheduling, six core features are extracted, covering task attributes, computational requirements, and security requirements:

[0211] - Task type characteristics: Based on the calculation mode, tasks are divided into training tasks (marked as 1) and inference tasks (marked as 0). Training tasks need to additionally record "number of model parameters (unit: B)" and "number of iterations". Inference tasks need to record "data volume of a single task (unit: MB)".

[0212] - Computational complexity characteristics: By statically analyzing the number of computational operations of the task (such as "number of attention calculations = sequence length² × number of heads" in the Transformer model), we convert them into "theoretical FLOPs (unit: TFLOPs)" to quantify the intensity of the task's computational demand.

[0213] - Data security features: Associate the authentication results of the trusted data exchange layer and extract "data security level (level 1-5, level 5 is the highest, such as image data containing patient privacy)" and "whether evidence needs to be stored (yes=1, no=0)". The security level directly affects the hardware selection priority (e.g., level 5 tasks are only allocated hardware with a historical failure rate of <0.1 times / hour).

[0214] - Time constraint characteristics: Record the task's "deadline (unit: ms)" and "tolerable delay (unit: ms)". High time-sensitive tasks (such as industrial real-time control, with a deadline ≤ 50ms) should be scheduled first.

[0215] - Hardware preference feature: If the task has a clear hardware adaptation requirement (such as an FPGA-accelerated cryptographic inference task), mark it as "preferred hardware type (CPU / GPU / TPU / FPGA)"; otherwise, mark it as "0".

[0216] - Resource usage characteristics: Estimate the "peak memory usage (unit: GB)" and "bandwidth requirement (unit: GB / s)" of the task to avoid hardware resource overflow during scheduling.

[0217] 2. Feature Quantization and Vector Construction: To adapt to the numerical input of the reinforcement learning model, the module standardizes the above features:

[0218] - Numerical features (such as parameter count, FLOPs) are mapped to the [0,1] interval using "min-max normalization" (e.g., parameter count 1B-100B corresponds to 0-1).

[0219] - Categorical features (such as task type and security level) are encoded using "one-hot encoding" (e.g., security level 5 corresponds to the vector [0,0,0,0,1]).

[0220] - Finally, a 28-dimensional feature vector is generated (6 numerical features + 22 one-hot encoded features), which is then encapsulated in ProtoBuf format and transmitted to the reinforcement learning scheduling module. Feature extraction takes ≤10ms (test environment: Intel Xeon 8380 CPU).

[0221] II. Reinforcement Learning Scheduling Module: Improved Decision Optimization of the DQN Algorithm. This module uses an improved algorithm based on Deep Q-Network (DQN) to take "hardware profile + task features" as input and output the optimal computing power allocation strategy. The core technology improvements and implementations are as follows:

[0222] 1. Key improvements to the DQN algorithm: To address the issues of "high empirical correlation" and "unstable target value" in traditional DQN, three optimizations are designed:

[0223] - Priority Experience Replay (PER) Mechanism: Construct an experience pool with a capacity of 100,000 entries. Each experience (state S, action A, reward R, next state S') is assigned priority according to "TD error (temporal difference error)". The larger the TD error (indicating that the experience is more valuable to the model update), the higher the sampling probability. "Priority rearrangement" is triggered every 2,000 entries in the experience pool to avoid low-value experiences occupying resources for a long time.

[0224] - Asynchronous updates of dual-target networks: Set up an "evaluation network" (outputs Q value in real time) and a "target network" (updates at fixed intervals). The target network copies parameters from the evaluation network every 100 steps instead of updating synchronously, reducing the fluctuation of the target Q value. At the same time, a "soft update" strategy is adopted (target network parameters = 0.99 × target network parameters + 0.01 × evaluation network parameters) to further improve stability.

[0225] - Security constraint embedding: Filter out policies that do not meet security requirements in the action space (such as not assigning level 5 security tasks to CPUs without encryption capabilities), avoid invalid exploration, and improve the model training convergence speed by 30% (compared to traditional DQN).

[0226] 2. Reward Function and Training Process: The reward function design is directly related to the dual objectives of "cost-efficiency" while also taking into account security requirements. The formula is as follows:

[0227] Reward = α×(actual throughput / target throughput) + β×(1-actual deployment cost / baseline cost) -γ×(actual latency / tolerable latency).

[0228] - Dynamic weight adjustment: At security level 5, γ (latency penalty weight) = 0.6, α = 0.3, β = 0.1, prioritizing low latency; in normal scenarios (security levels 1-2), β (cost weight) = 0.5, α = 0.3, γ = 0.2, prioritizing cost control.

[0229] - Training process: In the offline phase, the model is trained using 100,000+ simulated task samples (covering different hardware profiles and task feature combinations), and converges after 5,000 iterations (loss function value < 0.05). In the online phase, the model is fine-tuned hourly using real task experience (1,000 samples / hour) to ensure that the strategy adapts to dynamic hardware states (e.g., automatically reducing task allocation for that hardware when GPU load suddenly increases).

[0230] III. Policy Execution Module: Low-latency instruction issuance and hardware linkage. This module transforms the "computing power allocation strategy" (hardware type, number of cores, memory ratio, parallel parameters) output by reinforcement learning into hardware-executable instructions. Through collaboration with the underlying hardware adaptation layer, it achieves rapid scheduling, as described below:

[0231] 1. Instruction generation and standardization: To address the differences in instructions across different hardware, the module constructs a "policy-instruction" mapping library.

[0232] - GPU instructions: Convert "core allocation (e.g., 50%)" to NVIDIA CUDA "thread block configuration (blockDim=1024, gridDim=cores×50% / 1024)" and "memory percentage (e.g., 40%)" to "cudaSetLimit(cudaLimitMallocHeapSize, total memory×40%)".

[0233] - FPGA instruction: Converts "computing power allocation" into "logic unit enable ratio (e.g., 60%)" in the FPGA configuration file (.bit file), and loads the configuration via the JTAG interface.

[0234] - TPU instruction: Converts "memory percentage" into "HBM partition size (hbm_partition_size = total HBM × percentage)" in TPU Runtime.

[0235] 2. Low latency delivery and reliability assurance.

[0236] - Interface selection: The gRPC protocol (based on HTTP / 2) is used to issue commands. Compared with the traditional REST interface, the transmission latency is reduced by 40%, and the time taken to issue a single command is ≤20ms.

[0237] - Asynchronous processing: Construct a task scheduling queue (capacity 1000) and use a thread pool (number of threads = number of hardware nodes × 2) to process instruction issuance in parallel, avoiding single-task blocking.

[0238] - Error retry: If the instruction fails to be issued (e.g., hardware is offline), "policy reselection" is automatically triggered (the reinforcement learning module is called to generate alternative policies). The number of retries is ≤3, and the retry interval is 10ms, ensuring a scheduling success rate of ≥99.8%.

[0239] The dynamic scheduling layer in the above embodiments achieves the following technical effects through multi-dimensional task parsing, improved DQN decision-making, and low-latency execution coordination:

[0240] 1. Significantly improved scheduling accuracy: Task feature extraction accuracy ≥99.5%, the "cost-efficiency" matching degree of reinforcement learning scheduling strategy can be improved by 70% compared with traditional static rules, and the single card throughput can stably reach more than 1200 tokens / s in large model training scenarios;

[0241] 2. Response speed meets real-time requirements: The end-to-end time from task reception to instruction issuance is ≤100ms, which is 67% lower than the existing scheduling system (average 300ms), and can support low-latency scenarios such as industrial real-time inference and edge computing.

[0242] 3. Balancing security and efficiency: Through security feature embedding and hardware priority control, the scheduling compliance rate of Level 5 security tasks reaches 100%, while the deployment cost is reduced by 30% compared to random scheduling, achieving a three-dimensional balance of "trustworthiness, efficiency, and low cost", providing an intelligent decision-making core for the computing power scheduling of the entire system.

[0243] In one embodiment, the task parsing module breaks down the artificial intelligence task and extracts characteristic parameters such as task type, data size, and computational complexity.

[0244] The above embodiments, through a three-level decomposition mechanism of "static analysis - dynamic sampling - feature fusion," accurately extract key feature parameters of artificial intelligence tasks, providing high-fidelity input for the dynamic scheduling layer. Specific technical details are as follows:

[0245] I. Refined Decomposition and Classification of Task Types. The module addresses the heterogeneity of AI tasks (training / inference, image / text / speech, etc.) by designing a multi-dimensional classification framework:

[0246] 1. Basic type classification: By parsing the key fields of the task startup script (such as "train.py" containing the "--epochs" parameter marking it as a training task, and "infer.py" containing the "--batch_size" parameter marking it as an inference task), and combining the API call features (the training task calls "model.fit()", and the inference task calls "model.predict()"), the type recognition accuracy is ≥99.8%.

[0247] 2. Scene Sub-type Segmentation: Further classification based on input data format—image task (input is .jpg / .png, including "Conv2D" operator), text task (input is .token file, including "MultiHeadAttention" operator), and speech task (input is .wav, including "MelSpectrogram" operator), and preset feature extraction templates for each sub-type (e.g., additional extraction of "sequence length" and "vocabulary size" for text task).

[0248] II. Quantitative Extraction of Data Scale and Computational Complexity. Addressing the core computational resource requirements of the task, the module achieves precise quantification through a combination of static analysis and dynamic sampling:

[0249] 1. Data Scale Extraction:

[0250] - Static parsing of data list files (such as a list of sample paths in .txt format), counting the total number of samples and combining it with the single sample size (by reading 10 random samples to calculate the mean), to obtain the total data volume (unit: GB).

[0251] - For streaming data (such as real-time video frames), the frame rate (fps) and single frame size are obtained by parsing the streaming protocol (RTSP / HTTP), and the "data volume per second (MB / s)" is calculated to support the bandwidth resource allocation of the dynamic scheduling layer.

[0252] 2. Quantification of computational complexity:

[0253] - The static analytical model structure of "operator counting method" is adopted: traverse the computation graph of the task (such as GraphDef of TensorFlow and TorchScript of PyTorch), count the number of core operators such as convolution (Conv) and matrix multiplication (MatMul) and the input dimension, and calculate the theoretical computing power requirement (unit: TFLOPs) according to the formula "FLOPs = number of operators × input dimension³".

[0254] - Dynamic sampling verification: 1% of the samples (minimum 100) of the running task are used to collect the actual calculation time and hardware utilization rate, and the static calculation value is corrected (dynamic value is enabled when the deviation is >10%) to ensure that the quantification error of complex metrics is ≤5%.

[0255] III. Feature enhancement for trusted scenario adaptation.

[0256] To adapt to the requirements of trusted data exchange, the module additionally extracts two types of security-related features:

[0257] 1. Data Sensitivity Level: Parse the "data tags" in the task configuration file (such as "PII" for sensitive personal information and "PHI" for medical privacy data) and map them to a sensitivity level of 1-5 (linked to the security level of the trusted data exchange layer).

[0258] 2. Interaction Mode: Determine whether the task involves cross-device data exchange (such as containing "grpc: / / " remote call address), mark the "interaction frequency (times / second)", and provide a basis for the dynamic scheduling layer's allocation of communication resources.

[0259] The task parsing module in the above embodiments achieves the following technical effects through multi-dimensional feature extraction and quantization optimization:

[0260] 1. Comprehensive Feature Extraction: Compared to existing technologies that only extract "data volume + task type", this approach adds eight key features, including computational complexity and sensitivity level, increasing the feature dimensionality by three times and providing richer data for scheduling decisions.

[0261] 2. Improved quantization accuracy: Static analysis combined with dynamic sampling reduces the quantization error of data size and computational complexity to ≤5%, which is a significant improvement over pure static analysis (error of about 20%).

[0262] 3. Enhanced scenario adaptability: By extracting sensitivity levels and interaction patterns, the system achieves deep binding between task characteristics and trusted data exchange scenarios, thereby increasing the security compliance rate of subsequent scheduling strategies to 100% and laying the foundation for the "security-efficiency" balance of the dynamic scheduling layer.

[0263] Furthermore, in one embodiment, the monitoring and feedback layer employs a PID controller to collect operating parameters such as throughput, latency, and power consumption in real time, and feeds these parameters back to the dynamic scheduling layer.

[0264] The embodiments provided by this invention, through a closed-loop mechanism of "distributed real-time acquisition + adaptive PID control + dynamic strategy feedback," can perceive the system's operating status in real time and dynamically correct the scheduling strategy, ensuring the stable and efficient utilization of computing resources in trusted data exchange scenarios. Specific technical details are as follows:

[0265] I. Distributed Operation Parameter Acquisition System: By constructing a monitoring network covering all hardware nodes, high-fidelity acquisition of multi-dimensional operation data is achieved.

[0266] 1. Parameter Acquisition System and Quantification Standards

[0267] To address the dual requirements of computing power scheduling and trusted exchange, six core monitoring parameters are defined, with each parameter having a clearly defined data collection granularity and physical meaning:

[0268] - Throughput parameters: Training tasks are quantized in "tokens / s" (e.g., single-card throughput of LLaMA model), and inference tasks are quantized in "samples / s" (e.g., ResNet50 image inference), with a sampling granularity of 10ms / time, accurately reflecting computational efficiency.

[0269] - Delay parameters: divided into "task response delay" (time from receiving the task to starting execution) and "execution delay" (actual time taken for the task), both in milliseconds. For high-confidence scenarios (such as medical inference), the precision should be up to 100μs.

[0270] - Hardware load parameters: including CPU utilization (%), GPU SM core utilization (%), FPGA logic unit utilization (%), and memory utilization (%), with a sampling frequency of 100Hz to capture hardware resource fluctuations.

[0271] - Power consumption parameters: Real-time power consumption (W) and cumulative energy consumption (kWh) of a single node are collected by hardware sensors at a sampling interval of 50ms for cost control and energy-saving scheduling.

[0272] - Data exchange parameters: Associate with the trusted data exchange layer, collect "encrypted transmission rate (Mbps)", "blockchain evidence storage latency (ms)" and "authentication success rate (%)", and monitor the operation status of the trusted mechanism.

[0273] - Abnormal event parameters: Record events such as hardware failure (e.g., GPU card drop), data verification failure (e.g., hash mismatch), task interruption, etc., with timestamps and node IDs for quick source tracing.

[0274] 2. The data acquisition architecture and preprocessing adopt a distributed architecture of "edge proxy + central aggregation", as detailed below:

[0275] - Edge Proxy: Deploy a lightweight acquisition proxy (written in C++, CPU utilization <1%) on each hardware device to directly read hardware register data through kernel-mode interfaces (such as Linux perf, NVIDIA NVML), avoiding the latency of user-mode calls.

[0276] - Central Aggregation: The agent encapsulates the data in Protocol Buffers format (40% compression compared to JSON) and sends it to the central node via UDP. The central node uses InfluxDB time-series database for storage (write throughput ≥ 100,000 records / second).

[0277] - Preprocessing mechanism: Perform "sliding window filtering" (window size 10 sampling points) and "3σ outlier removal" on the raw data to ensure data smoothness (e.g., throughput fluctuation ≤5%), providing reliable input for PID control.

[0278] II. Adaptive PID Controller Design

[0279] The controller breaks through the limitations of traditional fixed-parameter PID controllers by dynamically adjusting the proportional (Kp), integral (Ki), and derivative (Kd) coefficients to achieve precise correction of system deviations. The technical implementation is as follows:

[0280] 1. Multi-objective control logic, with differentiated PID algorithms designed for different control objectives:

[0281] - Throughput control: When the actual throughput is less than the target value (e.g., 1200 tokens / s for a training task), activate "gain-type PID":

[0282] - Deviation e = target value - actual value, Kp increases linearly with the increase of e (Kp = 0.8 when e = 200 tokens / s, Kp = 1.2 when e = 500 tokens / s), accelerating the expansion of computing power;

[0283] - Setting Ki to 0.05 (to avoid integral saturation) and Kd to 0.1 (to suppress overshoot) allows the throughput to converge quickly to within ±5% of the target value.

[0284] - Power consumption control: When the actual power consumption is greater than 80% of the rated value, the "suppressed PID" is activated:

[0285] - Kp=0.3 (mild adjustment), Ki=0.2 (cumulative inhibition effect), Kd=0.05 (smoothing fluctuations);

[0286] - By calculating the "power consumption deviation integral" (∫e dt), when the integral value is >500W·s, a tiered computing power throttling is triggered (each time the computing power allocation is reduced by 10%).

[0287] - Delay control: For delay constraints in high-confidence scenarios (e.g., ≤50ms), a "derivative-first PID" approach is adopted.

[0288] - Kd=0.5 (rate of change of priority response delay), Kp=0.4, Ki=0.1;

[0289] - When the latency change rate is >10ms / s (latency increases rapidly), additional computing resources will be triggered in advance to avoid exceeding the threshold.

[0290] 2. Parameter adaptive adjustment strategy

[0291] Based on fuzzy control theory, parameter adjustment rules are designed, and Kp, Ki, and Kd are dynamically optimized according to the "deviation magnitude" and "deviation change rate".

[0292] - When the deviation is large and changes rapidly (e.g., throughput drops by 30%): increase Kp (to improve response speed) and decrease Ki (to avoid overshoot);

[0293] - When the deviation is small and stable (e.g., power consumption fluctuation <5%): decrease Kp (reduce oscillation) and increase Ki (eliminate steady-state error);

[0294] - The rules are implemented using an offline-trained fuzzy decision table (10×10 grid), with a response time of ≤1ms when invoked online, ensuring real-time control.

[0295] III. Feedback and Linkage Mechanism with Dynamic Scheduling Layer

[0296] The controller translates adjustment instructions into executable signals, which are then applied to the dynamic scheduling layer through three types of paths:

[0297] 1. Reward function weight adjustment: When power consumption exceeds the limit, a signal is sent to the reinforcement learning scheduler to increase the cost weight β from 0.3 to 0.6, guiding the strategy to prioritize low-power hardware.

[0298] 2. Resource allocation threshold adjustment: When the latency remains high, dynamically increase the "hardware computing power allocation lower limit" (e.g., increase GPU core allocation from 30% to 50%) to force expansion.

[0299] 3. Emergency Dispatch for Abnormal Events: When data exchange authentication failure is detected (3 consecutive times), the "Trusted Priority Mode" is triggered, non-critical tasks are suspended, and computing power is concentrated and allocated to trusted channels for repair (such as restarting the blockchain evidence storage node).

[0300] The monitoring and feedback layer in the above embodiments achieves the following technical effects through multi-dimensional data acquisition, adaptive PID control, and dynamic feedback linkage:

[0301] 1. Significantly improved system stability: Throughput fluctuations can be reduced from ±15% in traditional monitoring to ±5%, and the compliance rate of latency control within the threshold can be increased from 85% to 99.9%, ensuring the continuity of tasks such as large model training.

[0302] 2. Dynamic balance of resource utilization: power consumption is controlled within 80% of the rated value (20% lower than the no-feedback solution), while the hardware load utilization is maintained in the high-efficiency range of 70%-90%, avoiding resource waste.

[0303] 3. Enhanced adaptability to trusted scenarios: Through data exchange parameter monitoring and emergency scheduling, the fault recovery time of the trusted mechanism is shortened from 30 seconds to 5 seconds, and the authentication success rate is stabilized at over 99.9%, providing closed-loop protection for the safe and efficient operation of the entire system.

[0304] Furthermore, in one embodiment, the dynamic scheduling layer employs an experience replay mechanism and target network technology during the training process of the reinforcement learning scheduling module to reduce the correlation and variance during the training process.

[0305] The embodiments provided by the present invention above solve the problems of "high sample correlation" and "large variance of Q-value estimation" in traditional reinforcement learning training by combining the strategy of "priority experience replay + asynchronous update of dual-target network", thereby improving the convergence speed and stability of the scheduling strategy. The specific implementation is as follows:

[0306] I. Experience Replay Mechanism: Breaking Temporal Correlation Between Samples. The experience replay mechanism eliminates temporal dependencies between samples by storing, filtering, and replaying historical training samples. The specific implementation is as follows:

[0307] 1. Structured Design of the Experience Pool. The experience pool adopts a composite structure of "circular buffer + priority index" to store the core samples for reinforcement learning training (state S, action A, reward R, next state S'). Key parameter design:

[0308] - Capacity and storage unit: The experience pool capacity is set at 100,000 samples (covering a combination of 100 hardware profiles × 1,000 task features). Each sample contains: a 30-dimensional state vector (hardware profile + task features), a 5-dimensional action vector (chip type, number of cores, memory ratio, etc.), a 1-dimensional reward value (calculated based on cost-efficiency), and a 30-dimensional next state vector. The size of a single sample is compressed to 512 bytes (using binary encoding).

[0309] - Priority Index Table: Assigns a priority weight (0-1) to each sample, dynamically updated based on TD error (temporal difference error, |Q_target - Q_eval|) - the larger the TD error, the higher the value of the sample for model optimization, and the larger the priority weight (up to 1.0). The index is maintained through a red-black tree structure, supporting priority queries and updates with O(logN) complexity.

[0310] 2. Priority Sampling and Sample Balancing Strategy. To avoid an excessively high proportion of low-value samples due to random sampling, a mechanism of "priority probability sampling + importance sampling weight adjustment" is adopted:

[0311] - Sampling probability calculation: The sampling probability of sample i is P(i) = (priority(i))^α / Σ(priority(j))^α, where α is the priority coefficient (default 0.6, can be dynamically adjusted: α=0.4 in the early stage of training to enhance exploration, α=0.8 in the later stage to focus on high-value samples).

[0312] - Importance weight correction: To mitigate the distribution bias caused by priority sampling, the importance weight of the sample is calculated as w(i) = (1 / (N×P(i)))^β, where β is the compensation coefficient (initially 0.4, increasing linearly to 1.0 with each training round). w(i)×(Q_target - Q_eval)^2 is introduced into the loss function to ensure that the training influence of low-priority samples is appropriately suppressed.

[0313] - Sample replacement rules: When the experience pool is full, a "low-priority replacement" strategy is adopted - samples with a priority weight < 0.1 are removed first, and if they do not exist, the earliest stored sample is removed to ensure that the experience pool always retains high-value samples.

[0314] II. Target Network Technology: Reducing Q-value estimation variance. The target network separates Q-value evaluation from target calculation by constructing a dual-network architecture of "evaluation network - target network." The technical implementation is as follows:

[0315] 1. Structure and Functional Division of Dual Networks

[0316] - Evaluation Network (Online Network): Outputs the Q-value (Q_eval) of the current state in real time, which is used to generate the current action (based on an ε-greedy policy: 90% probability of selecting the action with the highest Q-value, 10% probability of random exploration). The network adopts a 3-layer fully connected structure (256, 128, and 64 neurons in the hidden layer), and the activation function is ReLU.

[0317] - Target Network: The parameters are periodically copied from the evaluation network to calculate the target Q value (Q_target = R + γ×maxQ'(S', a'), where γ is a discount factor of 0.95). The network structure is exactly the same as the evaluation network, but the parameters are updated at a lower frequency (once every 100 steps) to avoid unstable Q value estimation caused by parameter fluctuations.

[0318] 2. Asynchronous parameter updates and stability optimization

[0319] To further reduce the variance of the target Q value, a "soft update + gradient pruning" strategy is adopted:

[0320] - Soft update mechanism: The target network parameters are not directly overwritten, but updated according to the formula θ_target = τ×θ_online+ (1-τ)×θ_target (τ=0.01), which makes the parameter changes smoother. Compared with hard update (direct copy), the Q value fluctuation is reduced by 40%.

[0321] - Gradient clipping: When evaluating the backpropagation of the network, if the gradient norm is >5.0 (calculated by L2 norm), it is clipped to within 5.0 to avoid gradient explosion caused by extreme samples and make the loss function converge more stably.

[0322] - Training batch control: 32 samples are sampled from the experience pool each time to form a training batch (balancing randomness and computational efficiency). The Adam optimizer (learning rate 0.001, decay rate β1=0.9, β2=0.999) is used to update and evaluate the network parameters. The training time per batch is ≤20ms (GPU environment).

[0323] III. Dynamic Adaptation and Convergence Guarantee during Training

[0324] To address the dynamic nature of computing power scheduling scenarios (real-time changes in hardware status and task type), the module is designed with a two-stage training mode of "offline pre-training + online fine-tuning":

[0325] 1. Offline pre-training: Train for 5000 rounds using 100,000+ simulated samples (covering typical hardware profiles of CPU / GPU / TPU / FPGA and high / medium / low priority tasks). When the loss function value is <0.05 and the fluctuation is <1% for 100 consecutive rounds, it is considered as preliminary convergence.

[0326] 2. Online fine-tuning: During system operation, 1,000 samples (filtering outliers) are extracted from actual scheduling experience every hour to supplement the experience pool. Fine-tuning is triggered every 200 samples (training for 10 rounds) to adapt the model to real-time hardware conditions (such as sudden increases in GPU load, FPGA failures, etc.) and ensure the continuous effectiveness of the scheduling strategy.

[0327] Through the above embodiments, the training optimization mechanism of the reinforcement learning scheduling module achieves the following technical effects through the collaboration of experience replay and the target network:

[0328] 1. Significantly reduced sample correlation: Priority experience replay reduces the temporal correlation of training samples from 0.8 (original sequence) to below 0.2, avoiding ineffective temporal noise learning by the model.

[0329] 2. Reduced variance in Q-value estimation: The dual-objective network and soft update mechanism reduce the standard deviation of Q-value estimation from ±0.3 (single network) to ±0.15, improving the convergence speed of the loss function by 50%.

[0330] 3. Enhanced robustness of scheduling strategy: In scenarios of sudden hardware failure (such as GPU failure) or sudden changes in task load (such as a 10-fold increase in inference requests), the stability of strategy adjustment is improved by 60% compared with traditional reinforcement learning, and the fluctuation range of single-card training throughput is controlled within ±5% (target 1200 tokens / s), providing highly reliable decision model support for the dynamic scheduling layer.

[0331] In summary, by integrating multi-layered architecture collaboration and innovative technologies in the embodiments provided by this invention, the following technical effects can be achieved, specifically as follows:

[0332] 1. The chip migration and adaptation module of the underlying hardware adaptation layer achieves full compatibility with CPU, GPU, TPU and FPGA through cross-architecture instruction set dynamic mapping (translation latency ≤10μs) and real-time profiling of 12 parameters (update frequency ≥10 times / second), reducing device adaptation costs by 60% and solving the problem of computing power silos in the traditional "one hardware, one adaptation" approach.

[0333] 2. By constructing a trusted data exchange closed loop, the trusted data exchange layer adopts SM9 two-way authentication (success rate ≥99.9%), SM4-SM2 combined encryption (throughput ≥200Gbps) and consortium blockchain notarization (latency ≤50ms), realizing full traceability and zero leakage of data exchange, meeting the needs of high-trust scenarios such as finance and healthcare.

[0334] 3. The core computing power optimization layer can significantly improve computing power optimization efficiency. By using dynamic quantization (accuracy loss ≤1%) and fusion with 32 operators, inference latency is reduced; efficiency is improved through three-dimensional parallelism of "task-data-model"; and memory utilization is improved through heterogeneous memory management (LSTM prediction + intelligent caching), significantly improving overall computing efficiency.

[0335] 4. By using a reinforcement learning scheduler (DQN + experience replay) in the dynamic scheduling layer with the goal of "throughput-cost-latency", the single-card training throughput is stably above 1200 tokens / s, reducing deployment costs; by using PID control in the monitoring feedback layer, the system fluctuation is kept ≤5%, effectively ensuring stable operation, thus achieving a balance between cost and efficiency.

[0336] In summary, the solution, through the deep integration of hardware adaptation, trust assurance, computing power optimization and intelligent scheduling, is significantly superior to existing single-dimensional optimization solutions, providing a full-stack solution for computing power scheduling in trusted scenarios.

[0337] While numerous embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A dynamic scheduling system for computing resources supporting trusted data exchange, characterized in that, include: The system comprises a low-level hardware adaptation layer, a trusted data exchange layer, a core computing power optimization layer, a dynamic scheduling layer, and a monitoring and feedback layer; among these... The underlying hardware adaptation layer includes a chip migration adaptation module, which is used to realize instruction set mapping and hardware capability profiling of heterogeneous hardware. The core computing power optimization layer includes a low-latency inference system, a high-efficiency N-dimensional parallel system, and a heterogeneous memory management system. The dynamic scheduling layer outputs a computing power allocation strategy based on capability profiles and task characteristics. The underlying hardware adaptation layer is connected to the core computing power optimization layer, the trusted data exchange layer is connected to both the core computing power optimization layer and the dynamic scheduling layer, and the dynamic scheduling layer communicates bidirectionally with the monitoring and feedback layer.

2. The system according to claim 1, characterized in that, The chip migration adaptation module includes an instruction set mapping unit and a hardware capability profiling unit; wherein... The instruction set mapping unit constructs a unified instruction set library across CPU, GPU, TPU, and FPGA, and configures the instruction translation delay to be ≤ preset time; The hardware capability profiling unit collects pre-configured chip parameters in real time, with an update frequency greater than or equal to the preset frequency.

3. The system according to claim 1, characterized in that, The trusted data exchange layer includes a distributed identity authentication module, a data encryption transmission module, and a blockchain evidence storage module; wherein... The distributed identity authentication module uses the SM9 algorithm to implement two-way authentication. The data encryption transmission module adopts the SM4-SM2 combined encryption scheme. The blockchain evidence storage module writes the exchange logs into the consortium blockchain.

4. The system according to claim 1, characterized in that, The low-latency inference system adopts a dynamic quantization strategy, which can automatically adjust the model quantization precision to 4-bit, 8-bit, or 16-bit, and integrates a custom operator library containing 32 operators.

5. The system according to claim 1, characterized in that, The high-efficiency N-dimensional parallel system constructs a three-dimensional parallel architecture of "task-data-model", which can allocate data parallel tasks to GPUs and model parallel tasks to FPGAs based on capability profiles.

6. The system according to claim 1, characterized in that, The heterogeneous memory management system includes a data heat prediction unit and a cache scheduling unit; The data popularity prediction unit uses the LSTM algorithm. The cache scheduling unit enables high-speed cache binding of high-frequency data.

7. The system according to claim 1, characterized in that, The dynamic scheduling layer includes a task parsing module, a reinforcement learning scheduling module, and a policy execution module; wherein... The reinforcement learning scheduling module adopts an improved algorithm based on deep Q-network (DQN), with capability profile and task characteristics as state inputs and computing power allocation strategy as action outputs. The strategy execution module, based on the allocation strategy, enables the dynamic binding and release of computing resources.

8. The system according to claim 7, characterized in that, The task parsing module breaks down artificial intelligence tasks and extracts characteristic parameters such as task type, data scale, and computational complexity.

9. The system according to claim 1, characterized in that, The monitoring and feedback layer uses a PID controller to collect operating parameters such as throughput, latency, and power consumption in real time, and feeds these parameters back to the dynamic scheduling layer.

10. The system according to claim 1, characterized in that, During the training process of the reinforcement learning scheduling module, the dynamic scheduling layer employs an experience replay mechanism and target network technology to reduce the correlation and variance during the training process.