Computing power sensing method and system with integrated computing network
By constructing a multi-level NUMA hierarchical structure diagram and optimizing resource mapping scheme, asynchronous Byzantine consensus processing and two-way knowledge distillation technology are used to solve the problem of lack of end-edge-cloud collaborative perception mechanism in the existing technology, efficient resource management and accurate computing power prediction are achieved, and the system's adaptability and collaboration capabilities are improved.
Patent Information
- Application Number
- CN202510257561.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology lacks a unified perception mechanism of end-edge-cloud collaboration, the synchronous consensus mechanism is inefficient, the computing power prediction accuracy is insufficient, and it is difficult to adapt to dynamically changing business needs.
A multi-level NUMA hierarchical structure diagram is constructed, resource mapping optimization is performed, and resource allocation is optimized through two-part graph structure and genetic algorithm. Asynchronous Byzantine consensus processing is adopted, conditions are generated adversarial networks are used to generate bridge samples, bidirectional knowledge distillation is performed, and model optimization and hierarchical aggregation is used to use meta-learning methods.
It realizes efficient resource allocation and management, improves system performance and response speed, improves computing power prediction accuracy, and enhances the system's adaptability and collaboration capabilities.
Smart Images

Figure CN120335983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed computing technology, and particularly to a computing network integrated computing power perception method and system. Background Art
[0002] The integration of computing and network is an important direction for the future development of computing and network architectures. It deeply integrates computing resources and network resources to achieve unified scheduling and optimal allocation of resources. In the context of computing power networking, how to accurately perceive and manage distributed computing power resources has become a key issue.
[0003] Currently, common computing power perception solutions mainly adopt a centralized management architecture, and the central node collects the status information of each computing node. For example, in a traditional data center, the resource management system monitors the computing power by periodically collecting the usage of resources such as server CPU and memory. At the same time, in a distributed edge computing scenario, some systems adopt a hierarchical computing power perception architecture, deploying resource monitoring agents at the edge nodes to collect local computing power information and report it to the upper management platform.
[0004] A solution in the prior art adopts a distributed resource perception method based on a hierarchical architecture, collecting computing power information by deploying monitoring agents at different levels. This solution uses a lightweight data collection protocol and adopts a hierarchical caching mechanism to reduce data transmission overhead.
[0005] However, the prior art has the following problems: 1) Lack of a unified perception mechanism for end-edge-cloud collaboration, the computing power information at each level is fragmented, and it is difficult to achieve global optimization; 2) In a complex distributed environment, the existing synchronous consensus mechanism is inefficient, affecting the system response speed; 3) The accuracy of computing power prediction is insufficient, and it is difficult to adapt to dynamically changing business requirements. Summary of the Invention
[0006] In view of this, this application provides a computing network integrated computing power perception method and system, which solves the problems of lack of a unified perception mechanism for end-edge-cloud collaboration, low efficiency of the synchronous consensus mechanism, and insufficient accuracy of computing power prediction in the prior art.
[0007] The embodiment of this application provides a computing network integrated computing power perception method, including:
[0008] Based on the NUMA node information, memory capacity information, CPU core number information of multiple computing nodes in the system, and the physical connection relationship information between nodes, construct a multi-level NUMA hierarchical structure diagram and analyze it to obtain the node distance matrix and resource distribution.
[0009] Based on the node - to - node distance matrix and the resource distribution, a mapping mathematical model from virtual resources to physical resources is constructed using a bipartite graph structure, and a resource mapping optimization algorithm is executed to obtain an optimized resource mapping scheme;
[0010] Based on the optimized resource mapping scheme and multiple computing node information, node selection and verification are performed through a verifiable random function to obtain a verified node subset;
[0011] Based on the requests of the verified node subset, asynchronous Byzantine consensus processing is performed using a directed acyclic graph structure to obtain a consensus result that is consistent across all network nodes;
[0012] Based on the consensus result and the data feature distribution information of multiple computing nodes, conditional generative adversarial networks are used to generate bridging samples, and bidirectional knowledge distillation is performed to obtain the model knowledge transfer result between adjacent nodes;
[0013] Based on the model knowledge transfer result, a meta - learning method is used to construct a correction network for quality evaluation and optimization, and collaborative training is performed through a hierarchical aggregation method to obtain the optimization results of the end - edge - cloud three - layer models.
[0014] Optionally, the step of constructing a multi - level NUMA hierarchical structure diagram and analyzing it based on the NUMA node information, memory capacity information, CPU core number information, and node - to - node physical connection relationship information of multiple computing nodes in the system to obtain the node - to - node distance matrix and the resource distribution includes:
[0015] By scanning the system device node directory, the NUMA node information, memory capacity information, and CPU core number information of the multiple computing nodes are obtained; based on the NUMA node information, the memory capacity information, and the CPU core number information, a NUMA control tool is used for analysis to obtain the node - to - node physical connection relationship information; based on the node - to - node physical connection relationship information, system topology parsing is performed to obtain the node - to - node distance matrix and the resource distribution.
[0016] Optionally, the step of constructing a mapping mathematical model from virtual resources to physical resources using a bipartite graph structure and executing a resource mapping optimization algorithm based on the node - to - node distance matrix and the resource distribution to obtain an optimized resource mapping scheme includes:
[0017] Based on the node - to - node distance matrix and the resource distribution, construct a bipartite graph structure including memory access latency constraints, bandwidth limit constraints, and load - balancing constraints to obtain an initial bipartite graph model; based on the initial bipartite graph model, use a genetic algorithm for optimization calculation, where the mapping relationship from virtual resources to physical resources is used as chromosome encoding to obtain a preliminarily optimized mapping scheme; based on the preliminarily optimized mapping scheme, perform optimization through multi - generation iterative optimization and local search strategies to obtain the optimized resource mapping scheme.
[0018] Optionally, the step of, based on the optimized resource mapping scheme and multiple computing node information, performing node selection and verification through a verifiable random function to obtain a verified node subset includes:
[0019] Based on the optimized resource mapping scheme and multiple computing node information, perform periodic calculation using a distributed random beacon to obtain a random seed; based on the random seed and the node private key, perform calculation and verification through a verifiable random function to obtain a node random value and a verification result; based on the node random value and the verification result, perform comprehensive evaluation by combining the node response time, historical credit score, and computing power to obtain the verified node subset.
[0020] Optionally, the step of, based on the requests of the verified node subset, performing asynchronous Byzantine consensus processing using a directed acyclic graph structure to obtain a consensus result consistent among all network nodes includes:
[0021] Based on the requests of the verified node subset, perform analysis using a directed acyclic graph structure to obtain the message - dependency relationship between requests; based on the message - dependency relationship between requests, perform local view merging through an adaptive timeout mechanism to obtain a view - merging result; based on the view - merging result, perform consistency verification and perform packaging processing through a batch - processing technique to obtain a local consensus result; based on the local consensus result, perform propagation using a hierarchical broadcast strategy and perform message deduplication and integrity check through a Bloom filter to obtain the consensus result consistent among all network nodes.
[0022] Optionally, in the step of, based on the consensus result and the data - feature distribution information of multiple computing nodes, using a conditional generative adversarial network to generate bridging samples and performing bidirectional knowledge distillation to obtain the model - knowledge transfer result between adjacent nodes, generating the bridging samples includes:
[0023] Based on the consensus result and the data feature distribution information of multiple computing nodes, configure a sample generation module at the edge, cloud, and terminal layers and perform feature input to obtain an initial feature mapping result; based on the initial feature mapping result, use a multi-objective optimization mechanism to train conditions, generate an adversarial network and perform evaluation to obtain an optimized generation network; based on the optimized generation network, perform sample generation through feature matching technology to obtain bridging samples that capture the data statistical characteristics of each layer.
[0024] Optionally, the step of performing two-way knowledge distillation to obtain the model knowledge transfer result between adjacent nodes includes:
[0025] Based on the bridging samples that accurately capture the data statistical characteristics of each layer, initialize and maintain local models and teacher models at each node to obtain an initialized model group; based on the initialized model group and the bridging samples, construct a cross-layer feature mapping matrix to obtain a feature mapping relationship; based on the feature mapping relationship, use an adaptive temperature parameter to adjust the knowledge transfer softness and introduce an attention mechanism to guide knowledge extraction to obtain a knowledge transfer channel; based on the knowledge transfer channel, perform a two-way knowledge transfer process to obtain the model knowledge transfer result between adjacent nodes.
[0026] Optionally, the step of constructing a correction network using a meta-learning method for quality evaluation and optimization based on the model knowledge transfer result includes:
[0027] Based on the model knowledge transfer result and local data, calculate knowledge consistency to obtain a knowledge consistency score; based on the knowledge consistency score, perform analysis through an uncertainty knowledge screening mechanism to obtain a knowledge deviation identification result; based on the knowledge deviation identification result, use a correction network to perform key knowledge correction to obtain corrected knowledge; based on the corrected knowledge, perform an adaptive knowledge correction process to obtain a quality-optimized knowledge transfer result.
[0028] Optionally, the step of performing collaborative training through a hierarchical aggregation method to obtain the optimization results of the edge, cloud, and terminal layer models includes:
[0029] Based on the quality-optimized knowledge transfer result, design and execute a model parameter mapping mechanism to obtain a parameter mapping relationship between models of different scales; based on the parameter mapping relationship, use a dynamic weighting mechanism to perform weight allocation according to node data quality and computing power to obtain a model update weight; based on the model update weight, execute a progressive learning strategy for edge layer optimization to obtain an optimized edge layer model; based on the optimized edge layer model, perform cloud integration through a hierarchical aggregation method to obtain the optimization results of the edge, cloud, and terminal layer models.
[0030] The embodiment of the present application further provides a computing power perception system for computing-network integration, including:
[0031] A NUMA topology analysis module, which is used to construct a multi-level NUMA hierarchical structure diagram based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system, and perform analysis to obtain a node distance matrix and a resource distribution;
[0032] A resource mapping optimization module, which is used to construct a mapping mathematical model from virtual resources to physical resources using a bipartite graph structure based on the node distance matrix and the resource distribution, and execute a resource mapping optimization algorithm to obtain an optimized resource mapping scheme;
[0033] A consensus node selection module, which is used to perform node selection and verification through a verifiable random function based on the optimized resource mapping scheme and multiple computing node information, and obtain a subset of nodes that pass the verification;
[0034] A consensus processing module, which is used to perform asynchronous Byzantine consensus processing using a directed acyclic graph structure based on the requests of the subset of nodes that pass the verification, and obtain a consensus result that is consistent among all network nodes;
[0035] A knowledge migration module, which is used to generate bridging samples using a conditional generative adversarial network based on the consensus result and the data feature distribution information of multiple computing nodes, and execute bidirectional knowledge distillation to obtain a model knowledge transfer result between adjacent nodes;
[0036] A model optimization module, which is used to construct a correction network using a meta-learning method for quality evaluation and optimization based on the model knowledge transfer result, and perform collaborative training through a hierarchical aggregation method to obtain an optimization result for the three-layer models of the terminal, edge, and cloud.
[0037] The present application has the following technical effects:
[0038] By constructing a multi-level NUMA hierarchical structure diagram and optimizing the resource mapping scheme, efficient resource allocation and management are realized, and the system performance is improved;
[0039] By adopting an asynchronous Byzantine consensus processing mechanism, the system response speed and reliability in a distributed environment are significantly improved;
[0040] Through bridging sample generation and bidirectional knowledge distillation technologies, efficient collaboration among the three layers of the terminal-edge-cloud is realized, and the accuracy of computing power prediction is improved;
[0041] By adopting a meta-learning method and a hierarchical aggregation strategy, the model optimization effect is effectively improved, and the adaptability of the system is enhanced. Description of the Drawings
[0042] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below.
[0043] Figure 1 It is a schematic flowchart of a computing-network integrated computing power perception method provided by an embodiment of the present application;
[0044] Figure 2 It is a schematic flowchart of a resource mapping optimization method based on NUMA topology analysis in an embodiment of the present application;
[0045] Figure 3 It is a schematic flowchart of a node selection and verification method based on asynchronous Byzantine consensus in an embodiment of the present application;
[0046] Figure 4 It is a schematic structural diagram of a computing-network integrated computing power perception system provided by an embodiment of the present application.
[0047] It should be noted that the block diagrams shown in the accompanying drawings are only the logical function blocks of functional entities, and do not necessarily correspond to physically independent entities. It can be implemented by hardware, software, or a combination of both, and can be integrated in one processor or distributed in different processors. Specific Embodiments
[0048] The following refers to the case written in the specification to elaborate in detail on the specific embodiments of the embodiments of the present application.
[0049] As Figure 1 shown, an embodiment of the present application provides a computing-network integrated computing power perception method, including the following steps:
[0050] S1: Based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system, construct a multi-level NUMA hierarchical structure diagram and analyze it to obtain the node distance matrix and resource distribution.
[0051] The Chinese full name of NUMA (Non-Uniform Memory Access) is "Non-Uniform Memory Access Architecture".
[0052] Definition of NUMA node:
[0053] A NUMA node is a hardware architecture concept that represents a hardware unit with independent local memory and processor (CPU) resources. Each NUMA node usually contains:
[0054] One or more processors (CPUs);
[0055] Local memory directly connected to these processors;
[0056] Interconnection channels for connecting to other NUMA nodes;
[0057] May also include other hardware resources such as I / O controllers;
[0058] The main features of the NUMA architecture are:
[0059] Local access: The processor can access the local memory within the same NUMA node with the fastest speed and the lowest latency;
[0060] Remote access: When the processor accesses the memory (remote memory) of other NUMA nodes, it needs to go through the inter-node interconnection channel, which will generate additional access latency;
[0061] Non-uniformity: Different memory access paths will result in different access latencies, which is the origin of "non-uniformity";
[0062] Take a practical example:
[0063] In a dual-processor server, there are usually two NUMA nodes, and each node contains a physical CPU and the memory modules directly connected to it.
[0064] If CPU 0 wants to access the memory on NUMA node 1, it needs to go through interconnection channels such as QPI (Intel) or Infinity Fabric (AMD). The latency of this access will be 20 - 50% or even more higher than accessing local memory.
[0065] Step S1 specifically includes:
[0066] S1.1: Through scanning the system device node directory, obtain the NUMA node information, memory capacity information, and CPU core number information of the multiple computing nodes.
[0067] The system devices referred to in the embodiments of the present application include but are not limited to various servers and / or edge node devices, and the system includes a computing power perception system integrating computing and networking, such as a network node cluster and a core server cluster in a computing power network.
[0068] The system obtains the detailed information of the computing nodes by scanning the system device node directory. Specifically, the system will access the / sys / devices / system / node / directory in the Linux system, which contains the configuration information of all NUMA nodes.
[0069] For each NUMA node, the system reads its node_meminfo file to obtain the memory capacity information, including detailed data such as the total memory size, available memory, and used memory.
[0070] Meanwhile, the CPU core list on this NUMA node is obtained by accessing the node_cpulist file, including the physical core ID and the logical core ID. In addition, the system also parses the node_distance file, which records the memory access distance values between the current node and other nodes. During the actual implementation process, the system uses asynchronous I / O to read this information in parallel to improve the information collection efficiency.
[0071] S1.2: Based on the NUMA node information, the memory capacity information, and the CPU core number information, use the NUMA control tool for analysis to obtain the physical connection relationship information between the nodes.
[0072] The system conducts in-depth analysis using the NUMA control tool based on the collected node information.
[0073] Exemplarily, the system calls the hardware_info interface of the numactl tool to obtain the NUMA topology information at the hardware level, including the distribution of memory controllers, the connection method of QPI buses, etc. By parsing the output of numastat, the system can obtain the memory access statistical information of each NUMA node, including performance metrics such as the number of local accesses and the number of remote accesses.
[0074] It should be noted that the system will establish a dynamically updated data structure to store these analysis results. This data structure adopts the form of an adjacency list, which is convenient for subsequent topology analysis. In addition, the system will also monitor the data transfer bandwidth and latency between nodes in real time through the interfaces provided by the libnuma library.
[0075] S1.3: Based on the physical connection relationship information between the nodes, perform system topology parsing to obtain the distance matrix between the nodes and the resource distribution.
[0076] The system performs system topology parsing on the collected physical connection relationship information.
[0077] During the specific implementation, first construct a weighted directed graph, where the nodes represent NUMA nodes and the weights of the edges represent the memory access latency between the nodes. Through this directed graph, the system calculates the complete distance matrix between the nodes, which includes not only the distances between directly connected nodes but also the aggregated distances between nodes that need to be accessed through multiple hops.
[0078] Meanwhile, the system analyzes the distribution of memory and CPU resources across various nodes to generate a resource distribution vector. In addition, the system also considers the distribution of PCIe devices (such as network cards, GPUs, etc.) because the location of these devices affects the efficiency of cross-node communication. In summary, the system obtains a comprehensive view of the NUMA architecture through topology parsing, which provides an important basis for subsequent resource mapping optimization.
[0079] It should be noted that the above steps form a dynamically updated closed loop: the system periodically executes the information collection process of S1.1, obtains the latest hardware status through the analysis of S1.2, and updates the system's topological awareness in S1.3. This dynamic update mechanism ensures that the system can promptly adapt to changes in hardware configurations, such as events like hot-plugging CPUs or adding / removing memory modules.
[0080] S2: Based on the inter-node distance matrix and the resource distribution, construct a mapping mathematical model from virtual resources to physical resources using a bipartite graph structure, and execute a resource mapping optimization algorithm to obtain an optimized resource mapping scheme.
[0081] As Figure 2 shown, step S2 specifically includes:
[0082] S2.1: Based on the inter-node distance matrix and the resource distribution, construct a bipartite graph structure that includes memory access latency constraints, bandwidth limit constraints, and load balancing constraints to obtain an initial bipartite graph model.
[0083] In step S2.1, the system constructs a bipartite graph structure with specific constraint conditions based on the inter-node distance matrix and resource distribution obtained from the previous steps.
[0084] Specifically, the left node set of this bipartite graph represents virtual resources (including virtual CPU cores and virtual memory blocks), and the right node set represents physical resources (including actual CPU cores and memory modules).
[0085] The system introduces three types of key constraints when constructing this bipartite graph: First, the memory access latency constraint requires that the access latency from the mapped virtual resource to its required physical memory does not exceed a preset threshold, which is usually set to 1.5 times the local access latency; second, the bandwidth limit constraint ensures that the load on the memory controller of each NUMA node does not exceed 80% of its maximum bandwidth to reserve sufficient bandwidth margin; finally, the load balancing constraint requires that the load difference on each physical NUMA node does not exceed 20% to avoid uneven resource allocation. In addition, the system also assigns weights to each edge in the bipartite graph, and this weight is the comprehensive calculation result based on memory access latency, bandwidth occupancy, and load conditions.
[0086] S2.2: Based on the initial bipartite graph model, use the genetic algorithm for optimization calculation. Among them, use the mapping relationship from virtual resources to physical resources as chromosome coding to obtain a preliminarily optimized mapping scheme.
[0087] In step S2.2, the system uses an improved genetic algorithm to optimize the resource mapping for calculation. In this algorithm, each chromosome represents a possible resource mapping scheme, and its coding method uses integer coding. Each gene bit represents which physical resource a virtual resource is mapped to. When initializing the population, the system will generate a part of the individuals according to historical optimization experience, and the remaining individuals are generated randomly to ensure the diversity of the initial population. During the evolution process, the system designs a special crossover operator: when two parent chromosomes are crossed, it will preferentially maintain those mapping relationships that result in lower memory access latency, and recombine other mapping relationships. The mutation operator is designed as intelligent mutation, that is, when mutating, it will refer to the load situation of the current node and tend to map resources to nodes with lighter loads. In addition, the system also implements an adaptive evolution parameter adjustment mechanism, which can dynamically adjust the crossover rate and mutation rate according to the convergence situation of the population.
[0088] S2.3: Based on the preliminarily optimized mapping scheme, perform optimization through multi-generation iterative optimization and local search strategies to obtain the optimized resource mapping scheme.
[0089] In step S2.3, the system further improves the quality of the scheme based on the preliminarily optimized mapping scheme through multi-generation iterative optimization and local search strategies. During the iterative optimization process, the system adopts the elitist retention strategy, and the best 5% of the individuals in each generation will directly enter the next generation. At the same time, the system implements a local search strategy based on the idea of simulated annealing: for high-quality individuals in each generation, the system will explore within its neighborhood and try to obtain a better solution by making small-scale mapping adjustments. This local search process will gradually reduce the search range as the number of iterations increases to achieve fine-grained optimization of the solution space.
[0090] In particular, the system will maintain a historical optimal solution library for storing high-quality mapping schemes found during the optimization process. When the system detects performance degradation, it can quickly roll back to these known high-quality schemes. Considering the dynamic change characteristics of computing resources in practical applications, the system also implements an incremental optimization mechanism: when only a small part of the resources change, the system will use the current mapping scheme as the basis and only perform local optimization on the affected part instead of recalculating the entire mapping scheme, which greatly improves the response speed of the system.
[0091] S3: Based on the optimized resource mapping scheme and multiple computing node information, perform node selection and verification through a verifiable random function to obtain a subset of nodes that pass the verification.
[0092] Step S3 specifically includes:
[0093] S3.1: Based on the optimized resource mapping scheme and multiple computing node information, use distributed random beacons to calculate periodically to obtain a random seed.
[0094] In step S3.1, the system uses the distributed random beacon mechanism to generate random seeds periodically.
[0095] Specifically, the system first establishes a hierarchical beacon node structure in the network, divides all participating nodes into multiple levels according to computing power and network connection quality. Each level of beacon node is responsible for generating local random values within a specific time window. The generation of these local random values adopts an entropy source collection mechanism based on hardware events, including but not limited to the low-order bits of the CPU timestamp counter, the time interval between network packet arrivals, the disk IO completion time, etc.
[0096] To ensure the quality of randomness, the system implements an entropy pool manager, which continuously monitors the replenishment rate and quality of entropy. When the entropy value in the entropy pool is insufficient, the system will automatically trigger additional hardware events to supplement the entropy source. Finally, by distributing and combining the local random values generated by different levels of beacon nodes, a global random seed with high entropy characteristics is obtained.
[0097] In particular, the system adopts a time synchronization protocol to ensure that the clock deviation of each node is within an acceptable range, which is crucial for ensuring the correctness of random beacons.
[0098] S3.2: Based on the random seed and the node private key, perform calculations and verifications through a verifiable random function to obtain a node random value and a verification result.
[0099] In step S3.2, the system performs calculations and verifications through a verifiable random function (VRF) based on the obtained random seed and the node private key.
[0100] In the implementation process, each node first uses its own private key to sign the random seed, and then maps the signature result to a fixed range through a specific hash function to obtain a node random value. This process adopts an improved VRF algorithm, which introduces a segmented verification mechanism and can prove the correctness of the random value without revealing the private key.
[0101] Specifically, the system designs a two-stage verification protocol: in the first stage, the node generates a corresponding proof information while generating a random value; in the second stage, other nodes can use the public verification algorithm and the node's public key to verify the legitimacy of the random value.
[0102] To improve the verification efficiency, the system adopts a batch verification technology that can verify the random values of multiple nodes simultaneously. In addition, the system has also implemented an anti-replay attack mechanism, which ensures that the random values generated in each round are fresh by adding timestamp and round information during the verification process.
[0103] S3.3: Based on the node random values and the verification results, comprehensively evaluate by combining the node response time, historical credit score, and computing power to obtain the subset of nodes that pass the verification.
[0104] In step S3.3, the system comprehensively evaluates based on the node random values and verification results obtained in the previous steps, in combination with multiple key indicators. The system first establishes a dynamic scoring model that includes three main dimensions: node response time, historical credit score, and computing power.
[0105] For the response time, the system collects the response data of nodes through a continuous heartbeat detection mechanism and calculates a stable response time metric using the exponential moving average algorithm.
[0106] The historical credit score is based on the historical behavior performance of the nodes, including the success rate of participating in consensus, the quality of provided services, the frequency of wrong behaviors, etc. These historical data are stored in a distributed credit database, and the weights of historical behaviors are adjusted through a decay function.
[0107] The computing power evaluation adopts a multi-level benchmark test scheme, including CPU performance test, memory access speed test, network bandwidth test, etc. The system combines these indicators through configurable weights to obtain the comprehensive score of the node.
[0108] Finally, the system selects the group of nodes with the highest score as the subset of nodes that pass the verification according to the preset threshold and the required number of nodes.
[0109] To ensure the dynamic adaptability of the system, this evaluation process is executed periodically, allowing new high-quality nodes to join the verification subset while removing nodes with declining performance.
[0110] S4: Based on the requests of the subset of nodes that pass the verification, perform asynchronous Byzantine consensus processing using a directed acyclic graph structure to obtain a consensus result that is consistent among all network nodes.
[0111] As Figure 3 shown, step S4 specifically includes:
[0112] S4.1: Based on the requests of the subset of nodes that pass the verification, analyze using a directed acyclic graph structure to obtain the message dependency relationship between requests.
[0113] In step S4.1, the system analyzes the requests of the verified node subset using a directed acyclic graph (DAG) structure.
[0114] Specifically, the system constructs a transaction node for each request, which contains metadata such as the detailed information of the request, timestamp, initiator signature, etc.
[0115] When constructing the DAG, the system establishes directed edges based on the chronological relationship and data dependency between requests. To accurately capture message dependencies, the system implements a fine-grained dependency analyzer that can identify various dependency types such as read-write conflicts and resource competition between requests.
[0116] In particular, the system adopts a causal relationship tracking mechanism based on vector clocks. By maintaining the local vector clock of each node and attaching relevant clock information during message transmission, the system can accurately identify the partial order relationship of events. In addition, the system also implements a dynamically adjustable maximum out-degree limit mechanism, which can adaptively adjust the number of pre-order transactions referenced by each transaction node according to the network load situation to balance the consensus efficiency and network overhead.
[0117] S4.2: Based on the message dependencies between the requests, perform local view merging through an adaptive timeout mechanism to obtain the view merging result.
[0118] In step S4.2, the system performs local view merging through an adaptive timeout mechanism.
[0119] The system first maintains a local DAG view at each node, which contains all the transactions known to the node and their dependencies.
[0120] To handle network latency and out-of-order messages, the system designs a multi-level timeout mechanism: the first level is the basic waiting time, which is dynamically calculated based on the statistical data of network RTT; the second level is the adaptive extension time, which is adjusted according to the current network congestion degree and message arrival pattern; the third level is the emergency deadline to ensure that the view merging is completed within the maximum tolerable delay. During the view merging process, the system adopts an incremental merging strategy, only processes the newly received transaction nodes, and ensures the consistency of the merged view through an efficient conflict detection algorithm. At the same time, the system implements a priority-based message processing queue, which can give priority to processing those transaction nodes with a large influence range and complex dependency relationships.
[0121] S4.3: Based on the view merging result, perform consistency verification and process it through batch processing technology to obtain the local consensus result.
[0122] In step S4.3, the system performs consistency verification on the merged view and improves the processing efficiency through batch processing technology.
[0123] The consistency verification process includes multiple levels: First is signature verification to ensure that each transaction has a legitimate initiator signature; second is dependency verification to check whether the dependency relationships between transactions satisfy causal consistency; and finally is state verification to ensure the legitimacy of the system state after executing these transactions.
[0124] To improve the verification efficiency, the system implements an adaptive batch processing mechanism that can dynamically adjust the batch size according to the current load. During the batch processing, the system adopts parallel verification technology to assign independent verification tasks to different processing cores for execution. In particular, the system designs an efficient state caching mechanism that can cache frequently accessed state data, significantly reducing the computational overhead of state verification.
[0125] S4.4: Based on the local consensus result, adopt a hierarchical broadcast strategy for dissemination, and perform message deduplication and integrity check through a Bloom filter to obtain the consensus result that is consistent among all network nodes.
[0126] In step S4.4, the system uses a hierarchical broadcast strategy to disseminate the consensus result and implements an efficient message deduplication and integrity check mechanism.
[0127] The hierarchical broadcast strategy organizes network nodes into multiple logical levels, and each level of nodes is responsible for forwarding messages to a specific set of nodes in the next level. This hierarchical structure can significantly reduce the number of redundant messages in network broadcasts. The system implements a Bloom filter pool at each node for quickly detecting duplicate messages. This filter pool adopts a segmented update strategy to periodically clean up expired filters to ensure memory usage efficiency.
[0128] During the message dissemination process, the system uses an incremental verification mechanism based on a Merkle tree, enabling receiving nodes to quickly verify the integrity of the received local data without waiting for the complete dataset. To handle network partition situations, the system implements an automatic recovery mechanism: when a network partition is detected, the node will enter a special catch-up mode and quickly synchronize the missed consensus results by requesting the status summaries of neighboring nodes.
[0129] These four steps constitute a complete asynchronous Byzantine consensus processing flow: S4.1 captures the dependencies between requests through a DAG structure, S4.2 ensures the consistency of the local view, S4.3 guarantees the correctness of the consensus result, and S4.4 realizes efficient and reliable result dissemination. A number of innovative technologies are adopted throughout the process to improve performance and reliability, while maintaining good scalability and fault tolerance. Each component of the system supports dynamic configuration and can be optimized according to the actual deployment environment and performance requirements.
[0130] S5: Based on the consensus result and the data feature distribution information of multiple computing nodes, use a conditional generative adversarial network to generate bridging samples, and perform bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes.
[0131] The steps of generating bridging samples specifically include:
[0132] S5.1: Based on the consensus result and the data feature distribution information of multiple computing nodes, configure sample generation modules at the terminal, edge, and cloud layers and perform feature input to obtain an initial feature mapping result.
[0133] The system configures independent sample generation modules at the terminal, edge, and cloud layers respectively.
[0134] At the terminal layer, the generation module is mainly deployed on terminal devices with certain computing capabilities, such as smartphones, Internet of Things gateways, etc.; at the edge layer, the generation module is deployed on edge servers; at the cloud layer, the generation module is deployed on high-performance servers in the data center. Each level of the generation module contains two core components: a feature extractor and a feature mapper.
[0135] The feature extractor adopts a multi-layer perceptron structure to extract key features from the original data. To adapt to the computing power differences at different levels, the system specifically designs feature extraction networks of different scales: the terminal layer uses a lightweight network to mainly extract low-dimensional features; the edge layer uses a medium-scale network that can extract more complex feature combinations; the cloud layer uses a full-size network that can capture fine-grained feature relationships.
[0136] The feature mapper is responsible for converting the extracted features into standardized feature vectors. The system adopts adaptive batch normalization technology to ensure that the feature distributions from different sources can be aligned to the same feature space. At the same time, a feature importance weighting mechanism is introduced to dynamically adjust the weights of features according to their discriminability and stability.
[0137] S5.2: Based on the initial feature mapping result, use a multi-objective optimization mechanism to train the conditional generative adversarial network and perform evaluation to obtain an optimized generative network.
[0138] The system is based on the Conditional Generative Adversarial Network (CGAN) architecture and implements a training mechanism for multi-objective optimization. The generator adopts an improved U-Net structure, which contains multiple residual blocks and skip connections, and can better preserve the detailed information of features. The discriminator adopts the PatchGAN structure, which can perform fine-grained evaluation on the local features of the generated samples.
[0139] During the training process, the system simultaneously optimizes three objectives: feature fidelity, sample diversity, and cross-level consistency. Feature fidelity is ensured by minimizing the distance between the generated samples and the original samples in the feature space; sample diversity is enhanced by introducing a conditional mutual information regularization term; cross-level consistency is constrained by designing a special loss function to ensure that the generated samples can reflect the common features of data at different levels.
[0140] The system also implements a dynamic batch size adjustment mechanism, which automatically adjusts the number of samples in each batch according to the training stability. To improve the training efficiency, a progressive growth strategy is adopted. First, the model is trained at a low resolution, and then the resolution is gradually increased until the target accuracy is reached.
[0141] S5.3: Based on the optimized generation network, sample generation is performed through feature matching technology to obtain bridging samples that capture the statistical characteristics of data at each layer.
[0142] Based on the optimized generation network, the system generates bridging samples through feature matching technology. First, the system uses Locality-Sensitive Hashing (LSH) to build a feature index and quickly finds similar feature combinations. Then, a smooth feature transition sequence is generated through feature interpolation technology to ensure that the generated samples can capture the continuous changes in the data distribution.
[0143] To ensure the quality of the generated samples, the system implements a multi-level quality control mechanism: first, feature consistency checking is performed to ensure that the feature distribution of the generated samples matches the original data; second, sample diversity evaluation is performed, and the minimum generation distance metric is used to measure the sample diversity; finally, cross-level consistency verification is performed to ensure that the generated bridging samples can effectively connect the knowledge representations at different levels.
[0144] In particular, the system adopts an adaptive sampling strategy during the generation process, dynamically adjusts the sampling ratio according to the data density in different regions, and avoids generating redundant or invalid bridging samples. At the same time, by maintaining a sample pool and updating it regularly, it is ensured that the system can always use the latest and most representative bridging samples for knowledge transfer.
[0145] The above steps provide high-quality bridging samples for the subsequent two-way knowledge distillation process, thus ensuring the effect of knowledge transfer. The whole process forms a closed loop: starting from the feature input, through the optimization of the generation network, and finally obtaining bridging samples that can accurately capture the statistical characteristics of each layer of data.
[0146] The steps for performing two-way knowledge distillation specifically include:
[0147] S5.4: Based on the bridging samples that accurately capture the statistical characteristics of each layer of data, initialize and maintain local models and teacher models at each node to obtain an initialized model group.
[0148] The process of the system initializing and maintaining the model group based on the bridging samples at each node adopts a hierarchical architecture design. At each node, the system simultaneously maintains models with two different roles: local models and teacher models. The local model is responsible for processing the tasks of the node itself, while the teacher model undertakes the responsibility of knowledge transfer. This dual-model architecture provides greater flexibility, enabling the node to not only maintain the performance of local tasks but also effectively participate in knowledge sharing.
[0149] The initialization of the local model adopts a progressive construction method.
[0150] First, the system will select a suitable model architecture according to the computing power and storage resources of the node. For end-layer nodes with limited computing resources, the system adopts lightweight model structures, such as MobileNet or ShuffleNet series architectures; for edge-layer nodes, medium-scale models, such as ResNet or DenseNet variants, are used; while in cloud-layer nodes with sufficient resources, full versions of large-scale models are deployed. The system ensures that the models can operate efficiently on their respective hardware platforms through model pruning and quantization techniques.
[0151] The initialization of the teacher model adopts the method of knowledge distillation pre-training.
[0152] The system first trains a high-capacity teacher model in the cloud using a large-scale dataset, and then transfers the knowledge to the teacher models at the edge layer and end layer step by step through cascaded distillation. In this process, the system uses adaptive knowledge compression techniques to ensure that the transferred knowledge can adapt to the model capacities of different levels. In particular, the system implements a model knowledge base for storing and managing common model parameters and architecture configurations at each level.
[0153] To maintain the stability of the model group, the system implements multiple protection mechanisms:
[0154] First, there is a version control system for tracking the evolution history of the model. Each model update generates a new version identifier and saves a complete snapshot of the parameters. This enables the system to quickly roll back to a stable version when the model performance shows anomalies. The system adopts an incremental storage strategy, only saving the changed parts of the model parameters, effectively reducing the storage overhead.
[0155] Second, there is a health monitoring system that continuously monitors the running state of the model. The monitoring metrics include key performance indicators such as computing latency, prediction accuracy, and resource utilization. The system sets adaptive alarm thresholds and triggers corresponding recovery mechanisms when the metrics are abnormal. In particular, the system implements a performance warning mechanism that can detect potential performance degradation problems in advance.
[0156] In addition, the system also implements a load balancing mechanism that reasonably distributes computing tasks through a dynamic scheduling algorithm. When it detects that a certain node has a high load, the system will automatically migrate some tasks to a node with a lighter load. This process takes into account the network latency and bandwidth limitations between nodes to ensure that task migration does not cause performance bottlenecks.
[0157] To improve the efficiency of model updates, the system adopts an incremental learning strategy:
[0158] First, the system maintains a sample cache pool that stores the recently processed data samples and their features. When new training data is received, the system preferentially uses these cached features to avoid repeated feature extraction calculations. The cache pool is managed using the LRU (Least Recently Used) strategy to ensure memory usage efficiency.
[0159] Second, the system implements a priority mechanism for parameter updates. The model parameters of different layers are assigned different update priorities according to their impact on performance. High-priority parameters are updated more frequently, while low-priority parameters are updated at a lower frequency, which not only ensures model performance but also reduces the computing overhead.
[0160] Finally, the system reduces storage redundancy through a model parameter sharing mechanism. The basic layer parameters are shared among similar model structures, and only independent parameters are maintained at the task-specific layers. This parameter sharing not only reduces the storage overhead but also provides a certain degree of regularization effect, helping to improve the generalization ability of the model.
[0161] S5.5: Based on the initialized model group and the bridging samples, construct a cross-level feature mapping matrix to obtain a feature mapping relationship.
[0162] The system constructs a complex cross - level feature mapping matrix based on the initialized model group and bridging samples. This mapping matrix adopts a sparse representation method, effectively reducing the storage and computational overhead. During the construction process, the system first decomposes the model features at each level, projecting the high - dimensional feature space into multiple low - dimensional sub - spaces. This decomposition uses an improved tensor decomposition algorithm, which can maintain the key correlations between features.
[0163] To establish an accurate feature mapping relationship, the system implements an adaptive feature alignment mechanism. This mechanism first identifies the key features in models at different levels through an attention network, and then uses the optimal transport theory to construct the mapping relationship between features. In particular, the system introduces local sensitivity constraints to ensure that similar input features will be mapped to similar target feature spaces.
[0164] During the optimization process of the mapping matrix, the system adopts an iterative refinement strategy. First, a coarse - grained feature mapping is established, and then, by analyzing the feedback of the mapping effect, the mapping relationship is gradually adjusted and optimized. At the same time, the system maintains a dynamically updated feature dictionary to store common feature patterns and their mapping relationships, which greatly improves the efficiency of feature mapping.
[0165] S5.6: Based on the said feature mapping relationship, adopt an adaptive temperature parameter to adjust the softness of knowledge transfer and introduce an attention mechanism to guide knowledge extraction to obtain a knowledge transfer channel.
[0166] When the system establishes a knowledge transfer channel, it implements an innovative adaptive temperature parameter adjustment mechanism. The temperature parameter controls the smoothness of the soft labels during the knowledge distillation process and has an important impact on the effect of knowledge transfer. The system dynamically adjusts the temperature parameter by analyzing the predicted confidence distribution of the model. When the model prediction is relatively certain, a lower temperature value is used to retain more detailed information; when the prediction is uncertain, a higher temperature value is used to obtain a smoother probability distribution.
[0167] The design of the attention mechanism adopts a multi - head self - attention structure, which can simultaneously focus on knowledge transfer at different semantic levels. Each attention head is responsible for capturing specific types of feature associations, such as low - level texture features, mid - level structural features, and high - level semantic features. The system also implements a feature gating mechanism, which can dynamically adjust the attention weights according to the importance of features.
[0168] To improve the efficiency of knowledge extraction, the system designs a hierarchical knowledge distillation strategy. Knowledge at different levels is organized into a pyramid structure, from the specific features at the bottom to the abstract concepts at the top, and knowledge transfer is carried out layer by layer. Each layer of the transfer process has an independent attention module and temperature parameter to ensure the accuracy of knowledge transfer.
[0169] S5.7: Based on the knowledge transfer channel, perform a two-way knowledge transfer process to obtain the model knowledge transfer result between adjacent nodes.
[0170] When the system performs two-way knowledge transfer, it adopts an innovative peer-to-peer transfer mechanism. A two-way knowledge channel is established between adjacent nodes, and each direction is equipped with an independent transfer controller. The controller is responsible for monitoring the quality of knowledge transfer, including the integrity, accuracy, and real-time nature of the transfer. When a decrease in the quality of knowledge transfer is detected, the system automatically adjusts the transfer parameters or triggers a retransmission mechanism.
[0171] To handle the possible model heterogeneity issues between nodes, the system implements a model adaptation layer. This adaptation layer can automatically identify the differences between different model architectures and generate corresponding conversion functions. These conversion functions not only consider the differences in model structures but also take into account the constraints of computing resources to ensure that the transferred knowledge can be efficiently executed on the target node.
[0172] The system also introduces a knowledge caching mechanism, maintaining a knowledge cache pool of the most recently used at each node. The cache pool is managed using the LRU (Least Recently Used) strategy and can quickly respond to frequently transferred knowledge content. At the same time, the system implements an incremental update mechanism, only transmitting the changed parts of the knowledge, significantly reducing the communication overhead.
[0173] During the knowledge transfer process, the system monitors the transfer effect in real time through a feedback control loop. This loop consists of three key components: performance measurement, quality assessment, and parameter adjustment. The performance measurement module collects various indicators during the transfer process, the quality assessment module analyzes these indicators and generates an assessment report, and the parameter adjustment module dynamically optimizes the transfer parameters based on the assessment results. Through this closed-loop control, the system can maintain a stable and efficient knowledge transfer effect.
[0174] In the above process, not only is the knowledge between the end-edge-cloud three layers ensured to be transferred efficiently and accurately, but also a reliable knowledge foundation is provided for subsequent model optimization. The entire process forms a complete knowledge transfer chain: from the construction of the feature mapping matrix to the establishment of the knowledge transfer channel, ultimately achieving effective knowledge transfer between adjacent nodes.
[0175] S6: Based on the model knowledge transfer result, use the meta-learning method to construct a correction network for quality assessment and optimization, and perform collaborative training through hierarchical aggregation to obtain the optimization results of the end-edge-cloud three-layer models.
[0176] As Figure 4 shown, in step S6, the steps of using the meta-learning method to construct a correction network for quality assessment and optimization specifically include:
[0177] S6.1: Calculate the knowledge consistency based on the model knowledge transfer result and local data to obtain a knowledge consistency score.
[0178] The system adopts a multi-dimensional evaluation method when calculating knowledge consistency.
[0179] First, the system establishes a knowledge feature vector space and represents the model knowledge as high-dimensional feature vectors. This feature vector includes multiple dimensions such as the prediction output distribution of the model, the intermediate layer feature activation pattern, and the key parameter statistical information. By calculating the cosine similarity between feature vectors, the system can quantitatively evaluate the consistency degree between knowledge from different sources.
[0180] To improve the accuracy of consistency calculation, the system implements an adaptive feature weight mechanism. Features in different dimensions are assigned different weights according to their impact on the model performance. These weights are dynamically adjusted through reverse verification: the system evaluates the prediction effects of different feature combinations on the validation set and then updates the feature weights according to the evaluation results. In particular, the system uses a feature importance evaluation method based on information gain to ensure more reasonable weight allocation.
[0181] The system also introduces a temporal consistency evaluation mechanism. By maintaining a sliding time window, the system can monitor the change trend of knowledge consistency. When a significant fluctuation in the consistency score is detected, the system will trigger an in-depth analysis process to identify the specific reasons for the inconsistency. This process uses a change point detection algorithm to accurately locate the occurrence time and influence range of knowledge deviation.
[0182] S6.2: Based on the knowledge consistency score, perform analysis through an uncertain knowledge screening mechanism to obtain a knowledge deviation identification result.
[0183] The system implements a knowledge screening mechanism based on uncertainty. First, the system uses Bayesian deep learning methods to estimate the uncertainty of model predictions, including epistemic uncertainty (the uncertainty of the model for unknown samples) and aleatoric uncertainty (the noise in the data itself). Through the Monte Carlo sampling method, the system can obtain the probability distribution of model predictions, thereby evaluating the reliability of knowledge.
[0184] During the knowledge screening process, the system designs a multi-level screening strategy. The first layer is a coarse screening based on thresholds to quickly filter out knowledge with significantly higher uncertainty; the second layer is a fine screening based on relative ranking to conduct a more detailed analysis of knowledge with moderate uncertainty; the third layer is a comprehensive evaluation based on context, considering the relevance between knowledge for the final screening.
[0185] To improve the accuracy of screening, the system implements a knowledge consistency graph. This graph uses a graph structure to represent the dependencies between knowledge, with nodes representing knowledge units and edges representing the association strength between knowledge. Through graph analysis algorithms, the system can identify abnormal patterns and potential sources of deviation in the knowledge network. In particular, the system adopts a community detection-based method that can cluster relevant knowledge for systematic deviation analysis.
[0186] S6.3: Based on the knowledge deviation identification result, use the correction network to correct key knowledge to obtain the corrected knowledge.
[0187] The system constructs a dynamic correction network based on the meta-learning method. The correction network adopts an architecture similar to the residual network but adds a meta-learning layer to enable it to adaptively generate correction strategies according to different types of knowledge deviations. The training of the network uses a gradient-based meta-learning algorithm. By training on multiple related tasks, the correction network can quickly adapt to new knowledge correction requirements.
[0188] The correction process adopts a phased strategy: first is the coarse-grained correction, which quickly adjusts for obvious knowledge deviations; then is the fine-grained correction, which gradually improves the knowledge quality through iterative optimization. At each stage, the system maintains a correction effect evaluation metric for dynamically adjusting the correction strategy. In particular, the system implements an adaptive learning rate adjustment mechanism that can dynamically adjust the parameter update step size according to the correction effect.
[0189] To handle complex knowledge dependencies, the system implements a hierarchical correction strategy. For different levels of knowledge, the system adopts different correction methods: for low-level feature knowledge, it mainly corrects by fine-tuning network parameters; for middle-level representation knowledge, it uses feature recombination and alignment methods; for high-level semantic knowledge, it corrects through knowledge graph reasoning.
[0190] S6.4: Based on the corrected knowledge, execute an adaptive knowledge correction process to obtain a knowledge transfer result with optimized quality.
[0191] The system implements an adaptive knowledge correction process control mechanism. This mechanism consists of three core components: dynamic evaluation, feedback regulation, and termination judgment. The dynamic evaluation module continuously monitors the correction effect, including multiple dimensions such as the improvement degree of knowledge consistency, computational resource consumption, and correction time. The feedback regulation module dynamically adjusts the correction strategy according to the evaluation result, including modifying the learning rate, updating the batch size, and adjusting the correction network structure. The termination judgment module is responsible for deciding when to end the correction process to avoid performance degradation caused by overcorrection.
[0192] To ensure the stability of the correction process, the system implements a checkpoint mechanism. At key nodes of the correction process, the system saves the current correction state, including information such as network parameters and correction effect metrics. If there is a performance degradation during subsequent corrections, the system can quickly roll back to the previous stable state. This mechanism uses an incremental storage strategy, only saving changes to key parameters, effectively reducing storage overhead.
[0193] The system also introduces a quality assurance mechanism for knowledge correction. By constructing a validation set, the system can timely evaluate the actual effect of the corrected knowledge. The validation set is constructed using stratified sampling to ensure coverage of different types of knowledge application scenarios. In particular, the system implements a progressive validation strategy, gradually increasing the strictness of validation during the correction process, effectively balancing the requirements of correction efficiency and quality assurance.
[0194] The above process ensures the efficiency and reliability of the knowledge correction process. The entire correction process forms a complete closed-loop: starting from knowledge consistency assessment, through uncertainty analysis and correction network processing, ultimately achieving knowledge quality optimization. The system ensures the stability and effectiveness of the correction process through multiple protection mechanisms and adaptive adjustment strategies.
[0195] As Figure 4 shown, in step S6, the steps of collaborative training through hierarchical aggregation specifically include:
[0196] S6.5: Based on the knowledge transfer result after quality optimization, design and execute a model parameter mapping mechanism to obtain the parameter mapping relationship between models of different scales.
[0197] The system designs and implements a model parameter mapping mechanism for handling knowledge conversion between models of different scales. First, the system constructs a parameter mapping table that describes the hierarchical correspondence between different model architectures. This mapping table is organized in a tree structure, which can effectively express the inheritance and dependency relationships between layers. For each model layer, the system records key information such as its functional attributes, parameter dimensions, and computational complexity.
[0198] To handle parameter adaptation between models of different scales, the system implements an adaptive parameter conversion algorithm. For layers with mismatched dimensions, the system uses dynamic kernel functions for parameter interpolation or downsampling. In particular, the system adopts an attention mechanism to guide the parameter conversion process to ensure that important feature information is retained during the conversion. When processing convolutional layers, the system achieves adaptive adjustment of parameters through channel pruning and merging techniques.
[0199] The system has also established a quality assessment mechanism for parameter mapping. By performing inference tasks on the target model, the system can evaluate the effect of parameter mapping. The evaluation metrics include multiple dimensions such as task performance, computational efficiency, and resource consumption. Based on the evaluation results, the system will dynamically adjust the mapping strategy, such as adjusting the parameter compression ratio, modifying the conversion function, etc., to obtain the optimal mapping effect.
[0200] S6.6: Based on the parameter mapping relationship, adopt a dynamic weighting mechanism to allocate weights according to the node data quality and computing power, and obtain the model update weights.
[0201] The system has implemented a dynamic weighting mechanism based on multi-dimensional metrics. First, the system collects the real-time status information of nodes, including data quality metrics (such as data integrity, freshness), computing power metrics (such as CPU utilization, memory usage), and historical performance metrics (such as prediction accuracy, response time), etc. After normalization processing, these metrics form a comprehensive score of the nodes.
[0202] The weight allocation adopts an adaptive hierarchical strategy. The system first determines the base weight based on the hardware configuration of the nodes, and then makes dynamic adjustments according to the data quality and computing power. In particular, the system introduces a credit scoring mechanism to record the historical performance of nodes and uses it as an important basis for weight adjustment. For nodes with stable performance, the system will gradually increase their weights; while for nodes with large performance fluctuations, their weights will be correspondingly reduced.
[0203] To ensure the fairness and efficiency of weight allocation, the system has implemented a feedback adjustment mechanism. By monitoring the effect of model updates, the system can evaluate the rationality of the current weight allocation. If it is detected that the contributions of some nodes are overemphasized or ignored, the system will trigger a weight rebalancing process. This process uses the gradient descent method to gradually adjust the weights of each node until a new equilibrium state is reached.
[0204] S6.7: Based on the model update weights, execute an incremental learning strategy for edge layer optimization to obtain an optimized edge layer model.
[0205] The system adopts an incremental learning strategy during the edge layer optimization process.
[0206] This strategy is divided into multiple stages: first is the warm-up stage, where the system gradually adjusts the model using a small learning rate and simple tasks; then is the acceleration stage, where the system gradually increases the learning rate and task difficulty according to the convergence of the model; finally is the stable stage, where the system precisely optimizes the model performance through fine-tuning.
[0207] To handle resource constraints in edge environments, the system implements a task scheduling optimization mechanism. By analyzing the dependencies and resource requirements of tasks, the system constructs a task execution graph and uses heuristic algorithms for scheduling optimization. In particular, the system adopts the task sharding technique, decomposing large optimization tasks into multiple sub-tasks that can be executed in parallel, thereby improving resource utilization efficiency.
[0208] The system also introduces model compression and quantization techniques to adapt to the computing capabilities of edge devices. Through structured pruning, the system can reduce the number of model parameters and the amount of computation; through quantization-aware training, the system can reduce the storage and computation overhead of the model while ensuring accuracy. In particular, the system implements an adaptive compression strategy that can dynamically adjust the compression ratio according to the specific capabilities of the device.
[0209] S6.8: Based on the optimized edge layer model, perform cloud integration through hierarchical aggregation to obtain the optimization results of the end, edge, and cloud three-layer models.
[0210] The system adopts a hierarchical aggregation strategy during the cloud integration phase. First, the system performs local aggregation within the edge node cluster, summarizing the model updates in the same geographical area or business domain. This process uses a weighted average method, with the weights based on the credibility and data quality of the nodes. The results of local aggregation are temporarily stored in the edge cache, waiting for further global aggregation.
[0211] When performing global aggregation, the system implements a distributed model fusion mechanism. By constructing a fusion tree, the system can complete the aggregation operation of large-scale models with logarithmic complexity. In particular, the system adopts an incremental aggregation strategy, only transmitting and processing the changed model parameters, significantly reducing the communication overhead. To ensure the quality of the aggregation results, the system implements an anomaly detection mechanism that can identify and filter out abnormal model updates.
[0212] Finally, the system migrates the knowledge of the global model back to each layer through knowledge distillation. This process uses an adaptive distillation strategy, adjusting the way and intensity of knowledge transfer according to the model capacity and task characteristics of different layers. In particular, the system implements a progressive distillation mechanism, allowing smaller models to gradually absorb the knowledge of the global model, avoiding potential performance losses caused by one-time migration.
[0213] The implementation of these technical details ensures the efficiency and reliability of the optimization process of the end-edge-cloud three-layer model. The entire process forms a complete optimization closed-loop: starting from parameter mapping, through weight allocation and edge layer optimization, and finally achieving overall optimization in the cloud. The system ensures the stability and effectiveness of the optimization process through multiple protection mechanisms and adaptive adjustment strategies. Especially when dealing with model optimization in a large-scale distributed environment, the system effectively solves the problems of computing efficiency and communication overhead through technologies such as hierarchical processing and incremental updates.
[0214] Taking the big data computing power center of a certain province as an example, explain the specific application scenarios and implementation process of the computing power perception method and system of the integrated computing and network:
[0215] Overview of the infrastructure:
[0216] This computing power center has 100 computing nodes, distributed in three data centers in three prefecture-level cities. Each node is equipped with a high-performance GPU / NPU hardware accelerator. At the same time, 200 edge computing nodes are deployed throughout the province, mainly distributed in places such as government service centers, hospitals, and transportation hubs in various cities. Terminal devices include various Internet of Things devices, mobile terminals, etc., with a quantity exceeding 100,000.
[0217] Application of NUMA virtual resource mapping optimization:
[0218] In the deep learning training cluster of the main data center, 32 8-GPU server nodes are deployed, and each node is a typical multi-level NUMA architecture. The system first masters the internal hardware architecture of each node through topology analysis, including the physical connection relationship between CPU cores, memory channels, and GPU cards. For example, when training a large-scale visual model, the system will allocate data preprocessing tasks to the CPU cores close to the GPU according to the characteristics of the training task, and at the same time cache the training data in the memory area closest to the corresponding GPU, thereby reducing the memory access latency from the original 120ns to 70ns and increasing the memory bandwidth utilization rate by 35%.
[0219] Practice of the asynchronous Byzantine protocol:
[0220] When dealing with cross-city data collaborative computing tasks, such as the provincial medical image intelligent diagnosis project, the system will randomly select 20 nodes from the 200 edge nodes distributed throughout the province to form a consensus subset. When a hospital needs to perform multi-center collaborative medical image analysis, its request will first pass through privacy protection verification (ensuring that patient data is de-identified), and then the selected subset nodes will perform consensus processing. This mechanism shortens the original consensus process that took several minutes to within 10 seconds, while ensuring the security of data collaboration.
[0221] Specific application of the federated learning framework:
[0222] Taking the provincial traffic situation prediction as an example, the edge computing nodes deployed by the system at the transportation hubs in each city collect local traffic data. Through the bridging sample generation technology, a sample set that can represent the local traffic characteristics (but does not include specific vehicle information) is extracted. These bridging samples are used to transfer knowledge between the three regional centers. Each center continuously optimizes its own traffic prediction model based on local data and the received knowledge. Finally, a high-precision traffic prediction model covering the whole province is aggregated on the provincial cloud platform, and the prediction accuracy rate is increased from the original 82% to 91%.
[0223] Such as Figure 4 shown, the embodiment of the present application also provides a computing-network integrated computing power perception system, including:
[0224] The NUMA topology analysis module is used to construct a multi-level NUMA hierarchical structure diagram based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system, and perform analysis to obtain the node distance matrix and resource distribution;
[0225] The resource mapping optimization module is used to construct a mapping mathematical model of virtual resources to physical resources in a bipartite graph structure based on the node distance matrix and the resource distribution, and execute a resource mapping optimization algorithm to obtain an optimized resource mapping scheme;
[0226] The consensus node selection module is used to perform node selection and verification through a verifiable random function based on the optimized resource mapping scheme and multiple computing node information to obtain a subset of nodes that pass the verification;
[0227] The consensus processing module is used to perform asynchronous Byzantine consensus processing in a directed acyclic graph structure based on the requests of the subset of nodes that pass the verification to obtain a consensus result that is consistent among all network nodes;
[0228] The knowledge migration module is used to generate bridging samples by using a conditional generative adversarial network based on the consensus result and the data feature distribution information of multiple computing nodes, and execute bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes;
[0229] The model optimization module is used to construct a correction network for quality evaluation and optimization by using a meta-learning method based on the model knowledge transfer result, and perform collaborative training in a hierarchical aggregation manner to obtain the optimization results of the end, edge, and cloud three-layer models.
Claims
1. A computing power perception method for integrated computing and networking, characterized in that, Including: Based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system, construct a multi-level NUMA hierarchical structure diagram and analyze it to obtain the node distance matrix and resource distribution; Based on the node distance matrix and the resource distribution, construct a mapping mathematical model from virtual resources to physical resources using a bipartite graph structure, and execute a resource mapping optimization algorithm to obtain an optimized resource mapping scheme; Based on the optimized resource mapping scheme and multiple computing node information, perform node selection and verification through a verifiable random function to obtain a verified node subset; Based on the requests of the verified node subset, perform asynchronous Byzantine consensus processing using a directed acyclic graph structure to obtain a consensus result that is consistent across all network nodes; Based on the consensus result and the data feature distribution information of multiple computing nodes, use a conditional generative adversarial network to generate bridging samples and execute bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes; Based on the model knowledge transfer result, use a meta-learning method to construct a correction network for quality evaluation and optimization, and perform collaborative training through a hierarchical aggregation method to obtain the optimization results of the end, edge, and cloud three-layer models.
2. The method according to claim 1, characterized in that, The step of constructing a multi-level NUMA hierarchical structure diagram and analyzing it based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system to obtain the node distance matrix and resource distribution includes: Obtain the NUMA node information, memory capacity information, and CPU core number information of the multiple computing nodes through system device node directory scanning; Based on the NUMA node information, the memory capacity information, and the CPU core number information, use a NUMA control tool for analysis to obtain the physical connection relationship information between the nodes; Based on the physical connection relationship information between the nodes, perform system topology parsing to obtain the node distance matrix and the resource distribution.
3. The method according to claim 1, wherein The step of constructing a mapping mathematical model from virtual resources to physical resources using a bipartite graph structure and executing a resource mapping optimization algorithm based on the node distance matrix and the resource distribution to obtain an optimized resource mapping scheme includes: Based on the node distance matrix and the resource distribution, construct a bipartite graph structure including memory access latency constraints, bandwidth limit constraints, and load balancing constraints to obtain an initial bipartite graph model; Based on the initial bipartite graph model, use a genetic algorithm for optimization calculation, where the mapping relationship from virtual resources to physical resources is used as chromosome coding to obtain a preliminarily optimized mapping scheme; Based on the preliminarily optimized mapping scheme, perform optimization through multi-generation iterative optimization and local search strategies to obtain the optimized resource mapping scheme.
4. The method according to claim 1, characterized in that The step of performing node selection and verification through a verifiable random function based on the optimized resource mapping scheme and multiple computing node information to obtain a verified node subset includes: Based on the optimized resource mapping scheme and multiple computing node information, use distributed random beacons to calculate periodically to obtain a random seed; Based on the random seed and the node private key, perform calculations and verifications through a verifiable random function to obtain a node random value and a verification result; Based on the node random value and the verification result, conduct a comprehensive evaluation by combining the node response time, historical credit score, and computing power to obtain the subset of nodes that pass the verification.
5. The method according to claim 1, characterized in that The step of performing asynchronous Byzantine consensus processing on the requests based on the subset of nodes that pass the verification using a directed acyclic graph structure to obtain a consensus result that is consistent among all network nodes includes: Based on the requests of the subset of nodes that pass the verification, analyze using a directed acyclic graph structure to obtain the message dependency relationship between the requests; Based on the message dependency relationship between the requests, perform local view merging through an adaptive timeout mechanism to obtain a view merging result; Based on the view merging result, conduct consistency verification and perform packaging processing through batch processing technology to obtain a local consensus result; Based on the local consensus result, adopt a hierarchical broadcast strategy for propagation and perform message deduplication and integrity check through a Bloom filter to obtain the consensus result that is consistent among all network nodes.
6. The method according to claim 1, wherein In the step of generating bridging samples using a conditional generative adversarial network based on the consensus result and the data feature distribution information of multiple computing nodes, and performing bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes, generating bridging samples includes: Based on the consensus result and the data feature distribution information of multiple computing nodes, configure sample generation modules at the edge, cloud, and end layers and perform feature input to obtain an initial feature mapping result; Based on the initial feature mapping result, train the conditions of a multi-objective optimization mechanism, generate an adversarial network and conduct evaluation to obtain an optimized generative network; Based on the optimized generative network, generate samples through feature matching technology to obtain bridging samples that capture the data statistical characteristics of each layer.
7. The method according to claim 1, wherein The step of performing bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes includes: Based on the bridging samples that accurately capture the data statistical characteristics of each layer, initialize and maintain local models and teacher models at each node to obtain an initialized model group; Based on the initialized model group and the bridging samples, construct a cross-layer feature mapping matrix to obtain a feature mapping relationship; Based on the feature mapping relationship, adjust the knowledge transfer softness using an adaptive temperature parameter and introduce an attention mechanism to guide knowledge extraction to obtain a knowledge transfer channel; Based on the knowledge transfer channel, perform a bidirectional knowledge transfer process to obtain the model knowledge transfer result between adjacent nodes.
8. The method according to claim 1, wherein The step of constructing a correction network using meta-learning methods for quality evaluation and optimization based on the model knowledge transfer result includes: Based on the model knowledge transfer result and local data, calculate knowledge consistency to obtain a knowledge consistency score; Based on the knowledge consistency score, conduct analysis through an uncertainty knowledge screening mechanism to obtain a knowledge deviation identification result; Based on the identified knowledge deviation results, use a correction network to correct key knowledge and obtain the corrected knowledge; Based on the corrected knowledge, execute an adaptive knowledge correction process to obtain a knowledge transfer result with optimized quality.
9. The method according to claim 8, wherein The step of performing collaborative training through a hierarchical aggregation method to obtain the optimization results of the end, edge, and cloud three-layer models includes: Based on the knowledge transfer result with optimized quality, design and execute a model parameter mapping mechanism to obtain the parameter mapping relationship between models of different scales; Based on the parameter mapping relationship, use a dynamic weighting mechanism to allocate weights according to node data quality and computing power to obtain model update weights; Based on the model update weights, execute an incremental learning strategy to optimize the edge layer and obtain an optimized edge layer model; Based on the optimized edge layer model, perform cloud integration through a hierarchical aggregation method to obtain the optimization results of the end, edge, and cloud three-layer models.
10. A computing power perception system integrating computing and network, characterized in that, It includes: A NUMA topology analysis module, which is used to construct a multi-level NUMA hierarchical structure diagram and analyze it based on the NUMA node information, memory capacity information, CPU core number information, and physical connection relationship information between nodes in the system, so as to obtain the node distance matrix and resource distribution; A resource mapping optimization module, which is used to construct a mapping mathematical model of virtual resources to physical resources in a bipartite graph structure based on the node distance matrix and the resource distribution, and execute a resource mapping optimization algorithm to obtain an optimized resource mapping scheme; A consensus node selection module, which is used to select and verify nodes through a verifiable random function based on the optimized resource mapping scheme and multiple computing node information to obtain a subset of nodes that pass the verification; A consensus processing module, which is used to perform asynchronous Byzantine consensus processing in a directed acyclic graph structure based on the requests of the subset of nodes that pass the verification to obtain a consensus result that is consistent among all network nodes; A knowledge migration module, which is used to generate bridging samples using a conditional generative adversarial network based on the consensus result and the data feature distribution information of multiple computing nodes, and execute bidirectional knowledge distillation to obtain the model knowledge transfer result between adjacent nodes; A model optimization module, which is used to construct a correction network using a meta-learning method for quality evaluation and optimization based on the model knowledge transfer result, and perform collaborative training through a hierarchical aggregation method to obtain the optimization results of the end, edge, and cloud three-layer models.
Citation Information
Cited By
Edge Internet of Things asynchronous agent arrangement method and system
CN122437807A
A method and system for orchestrating asynchronous intelligent agents in edge IoT
CN122437807B