Hyper-converged server resource pooling method and system

Through multi-dimensional resource topology construction, dual-queue scheduling, dynamic sharding and master-slave collaborative optimization, the cross-node delay and resource scheduling lag problems of high-frequency small IO operations in industrial control scenarios are solved, achieving efficient resource utilization and real-time business processing.

CN120547063BActive Publication Date: 2025-09-19BEIJING ZHONGKE JIANYOU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511046541.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-19
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies in industrial control scenarios fail to effectively solve the problems of cross-node delays, resource scheduling lags, and queue switching overheads in high-frequency small IO operations, resulting in real-time business delays and low resource utilization.

Method used

Through multi-dimensional resource topology construction, dual-queue scheduling, dynamic sharding, and master-slave collaborative optimization, spatial abstraction and collaborative scheduling of heterogeneous resources are achieved. Specific measures include constructing a multi-dimensional resource topology map, generating a three-dimensional metadata table, deploying an independent dual-queue architecture, generating and scheduling task priority tags, dynamically adjusting sharding granularity, migrating high-frequency data replicas, and performing master-slave node collaborative optimization.

Benefits of technology

It effectively reduces the cross-node delay of high-frequency small IO operations, improves the efficiency and accuracy of resource scheduling, reduces queue switching overhead, and improves the real-time performance and stability of industrial control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547063B_ABST
    Figure CN120547063B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for pooling resources of a hyper-converged server, belonging to the technical field of resource scheduling. The method comprises the following steps: collecting spatial information, generating a multi-dimensional resource topology map, abstracting edge node and core node resources into virtual units with location tags, building a virtual resource pool and generating a three-dimensional metadata table; building an independent dual-queue architecture and configuring a token bucket, triggering a queue clearing mechanism based on a control cycle, and merging short transactions of edge nodes in micro-batches; binding adjacent virtual units with the physical location of the device as the sharding primary key, monitoring access traffic across virtual units, migrating high-frequency data copies and redirecting write requests, and adjusting the sharding granularity according to data relevance; selecting the virtual unit closest to the device group as the master node within the data shard, merging write operations and selecting a low-latency path for synchronization, and using an asynchronous confirmation mechanism to respond to edge nodes, thereby improving the local resource hit rate and task processing efficiency, and ensuring the real-time and reliability of industrial control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for pooling hyper-converged server resources, and belongs to the technical field of resource scheduling. Background Art

[0002] The development of hyperconverged server resource pooling stems from the growing demand for IT infrastructure flexibility, efficiency, and cost control driven by cloud computing and digital transformation. With the diversification of business scenarios and the explosive growth of data volumes, the low resource utilization and complex management of traditional stovepipe architectures have become increasingly prominent, prompting the industry to explore deep integration of hardware resources through software-defined technologies. This technology leverages core technologies such as virtualization, distributed storage, and software-defined networking to integrate computing, storage, and networking resources onto standardized hardware platforms, creating a unified resource pool that supports dynamic scheduling and elastic expansion.

[0003] However, existing technologies have shortcomings when dealing with the reading and writing of device status data in industrial control scenarios. Specifically, they do not take into account the characteristics of high-frequency small IO operations and short transactions, resulting in real-time and regular services sharing the same queue resources. Real-time data is easily delayed in the queue, resulting in the problem of accumulated scheduling delays; in data sharding, there is no dynamic adjustment based on the physical location and distribution characteristics of industrial equipment, which makes the data interaction path across nodes too long and the network transmission delay amplified. In terms of consistency protocols, there is no optimization for the characteristics of industrial control with more reads and fewer writes and strong data locality. Log synchronization and replica negotiation have unnecessary overhead, which cannot meet the requirements of deterministic response in scenarios sensitive to millisecond-level delays, making it difficult to ensure the efficient and stable operation of industrial control real-time services. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a hyper-converged server resource pooling method and system, which solves the problems of cross-node delay and resource scheduling lag in high-frequency small IO operations in industrial control through multi-dimensional topology construction, dual-queue scheduling, dynamic sharding and master-slave collaboration.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] Hyper-converged server resource pooling method, including:

[0007] Probes are deployed on industrial device gateways to collect spatial information, generate a multi-dimensional resource topology map, abstract the resources of edge and core nodes into virtual units with location tags, build a virtual resource pool, and generate a three-dimensional metadata table.

[0008] An independent dual-queue architecture is built in the virtual unit, and a token bucket is configured. A queue clearing mechanism is triggered based on the control cycle to merge and process short transactions on edge nodes. Short transactions are IO operations whose data volume does not exceed the data volume threshold.

[0009] The system uses the physical location of the device as the sharding primary key to bind adjacent virtual units, monitors cross-virtual unit access traffic, migrates data replicas whose access frequency exceeds the access threshold, redirects write requests, and adjusts the sharding granularity based on data relevance.

[0010] Build a dynamic device group, select the virtual unit closest to the device group as the master node within the data shard, perform collaborative optimization of master and slave nodes, allow slave nodes to respond to local read requests, synchronize incremental status summaries to slave nodes, merge write operations and perform path synchronization, and use an asynchronous confirmation mechanism to respond to edge nodes.

[0011] Specifically, the generation logic of the multi-dimensional resource topology map includes:

[0012] Based on the physical coordinates of the devices, the regions are divided, and the devices with spatial distance less than the preset distance and signal strength greater than the preset strength are grouped into geographic clusters, and a weighted undirected graph is constructed;

[0013] Based on the correlation between edge weights and node functions in the weighted undirected graph, the control subnet topology cluster is identified, and the edge nodes and corresponding core nodes of the same production line are divided into the same cluster;

[0014] Extracting three types of time series features from historical data, namely, device location migration patterns, load fluctuation characteristics, and network quality characteristics, and injecting them into the weighted undirected graph to expand it into a dynamic topology graph with a time dimension;

[0015] A spatiotemporal prediction model is constructed by combining graph neural network with long short-term memory network. The current topological state and historical time series features are input and the topological change prediction results are output.

[0016] Specifically, the steps of generating a three-dimensional metadata table include:

[0017] Build a digital twin for each node, collect underlying hardware parameters, use reinforcement learning models to predict resource usage trends, and generate virtual units; nodes include edge nodes and core nodes;

[0018] Extracting spatiotemporal features from the multidimensional resource topology map to generate virtual unit location tags containing spatiotemporal codes;

[0019] Set up a triggered update mechanism to automatically update location labels when the topology changes;

[0020] Build a knowledge graph containing industrial entities, map key information as hash keys to corresponding nodes in a distributed key-value database based on the DHT indexing mechanism, and establish a multi-dimensional mapping relationship;

[0021] Generate a hybrid index, combine the load balancing characteristics of DHT to optimize index distribution, and pre-store related data with access frequency exceeding the access threshold through DHT's cache strategy based on the access heat map to generate a three-dimensional metadata table.

[0022] Specifically, the triggered update mechanism includes:

[0023] A topology change threshold is set. When a device is predicted to enter a new area, the virtual unit position encoding update is triggered synchronously through the hardware timestamp. This includes recalculating the target area topology features through GNN, remapping the spatiotemporal Hilbert curve encoding, and updating the index of the distributed key-value database.

[0024] Specifically, when building a dual-queue architecture, physically isolated dual queues are created for each virtual unit, including a scheduling queue and a regular queue;

[0025] The scheduling queue is used to execute real-time tasks with timing requirements. The depth of the scheduling queue is dynamically configured to a preset multiple of the total number of industrial devices associated with the corresponding virtual unit. A circular buffer structure is used to implement enqueue or dequeue operations. Task attribute tags including security level, control period, and device correlation coefficient are extracted from the three-dimensional metadata table to generate scheduling priority tags and classify real-time tasks.

[0026] The conventional queue is used to process non-real-time tasks. The maximum depth of the conventional queue is set to a multiple of the preset depth of the scheduling queue and supports elastic expansion. A hierarchical priority linked list structure is adopted, and different proportions of queue resources are occupied according to different task categories. Labels with allowed delay thresholds are added to the conventional queue tasks, where task categories include urgent tasks, important tasks, and ordinary tasks.

[0027] Specifically, the triggering logic of the queue clearing mechanism includes:

[0028] Obtaining and parsing real-time IO request vectors, extracting request features, and simultaneously extracting semantic tags from the three-dimensional metadata table, performing association and standardization processing, and generating a task description set;

[0029] Performing multi-dimensional semantic classification on the tasks in the task description set to divide them into real-time task candidates and regular tasks, and placing the regular tasks into the regular queue;

[0030] Based on the security level and control period corresponding to the real-time task candidates, priority labels are divided;

[0031] Based on the source device location of the real-time task candidate, query the DHT index and sort it in ascending order of physical distance to obtain the top n candidate virtual units, and calculate the comprehensive score of each candidate virtual unit;

[0032] Prioritize virtual units whose comprehensive scores are greater than a preset score threshold and whose device correlation coefficient is not less than a preset correlation threshold as target virtual units;

[0033] If all comprehensive candidate scores are less than the preset score threshold, the fault self-healing mapping is triggered, the multi-dimensional mapping relationship is queried, the set of spare virtual units is obtained, and the nodes with loads lower than the preset load threshold are selected after sorting in descending order of computing power margin to generate the target virtual unit.

[0034] Specifically, the triggering logic of the queue clearing mechanism also includes:

[0035] Constructing a DQN model, generating a base number of tokens based on a control cycle, inputting a state vector into the DQN model to output an adjustment coefficient, calculating a compensation number of tokens, and obtaining a total number of tokens;

[0036] At the start of each control cycle, the scheduling queue is traversed and tasks whose remaining processing time is less than a preset ratio of the deadline are discarded;

[0037] Calculating the token usage rate in the scheduling queue in real time, and performing nonlinear expansion when the token usage rate is less than a first usage threshold for three consecutive periods;

[0038] When the single-cycle token usage rate is greater than a second usage threshold, shrinking the bandwidth of the regular queue, and continuing to shrink if it continues to be greater than a third usage threshold;

[0039] Maintain a transaction dependency matrix for each device group and update it in real time, count the associated devices in the current waiting task, calculate the locality index, and dynamically select the tight merge mode, standard merge mode, or sparse merge mode based on the locality index;

[0040] The ratio of the number of cross-cluster tasks to the total number of tasks is calculated to obtain the cross-cluster ratio. When the cross-cluster ratio is greater than the first cross-cluster ratio threshold, the sub-merge groups are divided to ensure that the cross-cluster ratio of each sub-merge group is less than or equal to the second cross-cluster ratio threshold. The locality index is recalculated for each sub-merge group. If the locality index is less than the preset value, the sub-merge group is further split into single-device group tasks and a globally unique merge package serial number is assigned.

[0041] Specifically, the steps to adjust the sharding granularity include:

[0042] Generate a composite shard primary key based on the physical location of the device;

[0043] Query the three-dimensional metadata table to obtain a set of virtual units, sort them in ascending order by Euclidean distance, select the virtual unit with the closest distance and whose computing power margin exceeds a preset first margin ratio to bind to the device, and if the distance is not satisfied, select the virtual unit with the next closest distance and whose computing power margin exceeds a preset second margin ratio to bind to the device;

[0044] Create shard metadata entries in the distributed key-value store and generate a shard mapping table;

[0045] Collect access traffic from edge nodes to core nodes in real time and obtain statistical indicators. When the statistical indicators meet preset trigger conditions, determine the source virtual unit and the target virtual unit, prioritize the migration of data copies with access frequencies exceeding the access threshold, and update the three-dimensional metadata table.

[0046] Collect device interaction logs and calculate the correlation coefficient. Devices with a correlation coefficient greater than a preset correlation threshold are grouped into correlation groups. The sharding granularity is adjusted based on the size of the correlation group or the interaction frequency.

[0047] When processing a write request, the optimal target virtual unit is predicted. If it is a remote node and there is a local replica, it is redirected. The number of replicas is limited according to the security level, and the write diffusion ratio is monitored to adjust the synchronization strategy.

[0048] Specifically, the specific steps of master-slave node collaborative optimization include:

[0049] Based on the three-dimensional metadata table, the devices whose process correlation exceeds a preset correlation threshold are grouped into a dynamic device group, and a dynamic topology map is updated;

[0050] Monitor the position changes of devices in the device group in real time. Once the device displacement exceeds a preset distance threshold, trigger the device group reaggregation, query the shard mapping table, and filter out virtual units within a preset range from the geometric center of the device group to form a candidate set;

[0051] The election score of candidate virtual units is calculated using a comprehensive scoring model. When the device group's movement distance or the election score change rate exceeds a threshold, a new master node is elected and the status of the new and old master nodes are synchronized.

[0052] When an edge node initiates a read request, it parses the shard location information. If the target data belongs to the current virtual unit, the slave node responds directly. Otherwise, it is marked as a remote read request. Before responding, the slave node checks the status difference with the master node. If the difference exceeds the difference threshold, it synchronizes the latest status first.

[0053] The master node merges write requests according to the control cycle, selects a path based on the network topology, synchronizes to the slave node, and asynchronously confirms;

[0054] When the device group moves, the master node is switched and the synchronization strategy is adaptively adjusted, with differentiated synchronization based on data security levels.

[0055] Hyper-converged server resource pooling system, including: perception module, scheduling control module, data management module and read-write optimization module;

[0056] The perception module is used to collect spatial information and construct a multi-dimensional resource topology map, abstract heterogeneous node resources into virtual units with location tags, establish a multi-dimensional mapping relationship between device data points and virtual units, and generate a three-dimensional metadata table;

[0057] The scheduling control module is used to build a dual-queue architecture, extract task priority tags based on a three-dimensional metadata table, combine a spatiotemporal scoring model with an adaptive token bucket mechanism to implement task scheduling, and dynamically merge and split short transactions;

[0058] The data management module is used to perform initial shard binding using the physical location of the device as the shard primary key, monitor cross-virtual unit access and trigger high-frequency data copy migration, dynamically adjust shard granularity based on device data relevance, redirect write requests and control cross-node write propagation;

[0059] The read-write optimization module is used to aggregate dynamic device groups based on device process relevance, elect a master node to process write requests and allow slave nodes to respond to local read requests, merge write operations and synchronize them to slave nodes through topology-aware paths, and adjust synchronization strategies based on device group movement and load status.

[0060] Beneficial effects of the present invention:

[0061] By constructing a multi-dimensional resource topology, the problem of "location blindness" in traditional resource pools is solved, and spatial abstraction of heterogeneous resources is achieved. This provides a collaborative view for scheduling and sharding, allowing the initial sharding of devices to be as close to the physical location as possible, reducing cross-node access. At the same time, the dual-queue architecture combines time-sensitive token buckets with intelligent merging strategies to ensure that real-time tasks are prioritized, reduce I / O queue switching overhead, and avoid delay accumulation due to priority confusion in real-time tasks. Dynamic data sharding and intelligent migration adjust the sharding granularity based on the physical location of the device and data relevance, migrate high-frequency data copies, redirect write requests, and reduce cross-node write diffusion. Furthermore, master-slave node collaborative optimization improves real-time read and write efficiency and consistency through dynamic master node election, read optimization, and write merging. This method systematically solves problems such as cross-node delay, resource scheduling lag, and queue switching overhead in high-frequency small I / O operations in industrial control, improves local resource hit rate and task processing efficiency, and ensures the real-time and reliability of industrial control. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a schematic diagram of a hyper-converged server resource pooling method;

[0063] Figure 2A flowchart of generating a three-dimensional metadata table in the present invention;

[0064] Figure 3 This is a flow chart of the queue clearing mechanism in the present invention;

[0065] Figure 4 This is a flowchart for adjusting the sharding granularity in the present invention;

[0066] Figure 5 This is a diagram of the hyper-converged server resource pooling system structure. DETAILED DESCRIPTION

[0067] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0068] Example 1:

[0069] refer to Figures 1 to 4 As shown, this embodiment introduces a method for pooling hyper-converged server resources, including the following steps:

[0070] Step S1: Deploy lightweight probes on the industrial device gateway to collect spatial information in real time and generate a multi-dimensional resource topology map containing heterogeneous nodes. The computing resources of edge nodes and the storage resources of core nodes are uniformly abstracted into virtual units with location tags to build a virtual resource pool. Each virtual unit is bound to the topology information of its service area and metadata is associated at the same time to establish a mapping relationship between device data points and virtual units. A three-dimensional metadata table containing device ID, location coordinates, and resource ownership is generated to achieve spatial abstraction of heterogeneous resources. This provides a collaborative view of physical location and load status for subsequent scheduling and sharding, solving the "location blindness" problem of traditional resource pools.

[0071] In this embodiment, the hyper-converged server consists of two types of nodes. One type is the edge node connected to industrial equipment, which adopts the ARM architecture and has real-time control capabilities. The other type is the core node responsible for storage and computing, which adopts the X86 architecture and provides powerful data processing capabilities. At the same time, due to the complex distribution of equipment and strict real-time requirements in industrial control scenarios, such as welding robots and sensors in automobile factories distributed at different workstations, lightweight probes are deployed on the ARM edge node gateway.

[0072] Spatial information includes physical layer parameters, network layer parameters, and resource layer parameters. Physical layer parameters include the physical coordinates of ARM edge nodes, X86 core nodes, and industrial equipment. Network layer parameters include the number of network hops between nodes, link latency, bandwidth utilization, and jitter indicators. Resource layer parameters include the real-time control load of edge nodes, such as interrupt response time, storage IOPS, and CPU utilization of core nodes. This enables comprehensive perception of the operating status of heterogeneous nodes and provides multi-dimensional data support for subsequent topology modeling.

[0073] Specifically, the generation logic of the multi-dimensional resource topology map includes:

[0074] First, a static topology is constructed. Since industrial field equipment layouts are typically divided into independent areas based on production processes, such as the welding and spraying areas of automobile factories, the DBSCAN density clustering algorithm is used to regionalize physical coordinates. Devices with spatial distances less than a preset distance and signal strengths exceeding a preset strength are clustered into a geographic cluster. For example, the welding area cluster contains three robots and five sensors. Based on the clustering results, a weighted undirected graph G = (V, E) is constructed, where the nodes V include edge nodes, core nodes, and key industrial equipment entities. The edges E represent physical connections. The weight of the edge E is calculated by multiplying the inverse of the bandwidth by the link delay to quantify data interaction. The link delay is the average round-trip time over the past period, and the bandwidth is the current link available throughput. This allows the topology to intuitively reflect the time and bandwidth consumption of data transmission.

[0075] Leveraging the Louvain community discovery algorithm, the system identifies industrial control subnet topology clusters based on edge weights and node functional relevance, such as PLC control nodes and data storage nodes on the same production line. Edge nodes and corresponding core nodes on the same production line are grouped into the same cluster, limiting real-time data interaction to within the cluster. This solves the problem of amplified network latency caused by cross-cluster access in traditional solutions.

[0076] Considering the dynamic operating characteristics of industrial equipment, such as AGVs moving along fixed trajectories and hourly production line starts and stops causing periodic fluctuations in I / O peaks, three types of time series features are extracted from the historical data of the topology construction engine. These include: fitting the AGV's motion trajectory using a Kalman filter algorithm as the equipment's location migration pattern; using Fourier transform to analyze the I / O peak period caused by production line starts and stops as a load fluctuation feature; and using a sliding window to analyze the time series correlation of industrial Wi-Fi frequency interference during a preset daily time period (e.g., 8:00-9:00) as a network quality feature.

[0077] The three types of extracted timing features are injected into the static topology graph, expanding it into a dynamic topology graph with a time dimension. This includes adding Transformer-based Time Embedding vectors to nodes, such as encoding the load change trend of the past 10 control cycles, and adding timing transition probabilities to edge weights. For example, using historical data to calculate the probability that a link's future 5ms delay will exceed 10μs. This enables the dynamic topology graph to express the dynamic behavior of industrial equipment in terms of timing, providing a data foundation for predicting topology evolution.

[0078] To implement proactive scheduling of high-frequency, small I / O operations and accurately predict topology change trends to meet the real-time demands of these operations, a spatiotemporal prediction model is constructed using a graph neural network combined with a long short-term memory network. The GNN layer uses a graph attention mechanism to process spatial correlations in the dynamic topology graph and extract positional dependencies between nodes. The LSTM layer uses a gating mechanism to capture the temporal evolution of topological states.

[0079] The system inputs the current topology state and historical time series features, and outputs predictions of topology changes in the future. These include predictions of the time when an AGV enters the target area, warnings when a network link's delay exceeds a threshold due to industrial Wi-Fi interference, and predictions of the arrival time of the IOPS peak of the core node in the welding area. This enables accurate predictions of future topology states, allowing for early adjustments to resource layout and scheduling strategies, effectively resolving cross-node delays and resource scheduling lags caused by dynamic topology changes in industrial control.

[0080] Specifically, the steps of metadata association include:

[0081] To achieve coordinated scheduling of heterogeneous resources between edge nodes and core nodes, it is first necessary to intelligently abstract the hardware resources of ARM edge nodes and X86 core nodes and generate virtual units. However, due to the significant differences in the hardware characteristics of the two types of nodes, such as the different real-time interrupt response capabilities of ARM and the storage throughput quantification standards of X86, traditional unified abstraction methods cannot accurately reflect the real-time needs of industrial control. Therefore, a digital twin is built for each node, and the underlying hardware parameters are collected through the firmware-level interface. A reinforcement learning model is used to predict resource usage trends, and then the physical resources are quantified into virtual units with unified measurements. Nodes include edge nodes and core nodes. Semantic-level alignment of heterogeneous resources is achieved, and the generated virtual units have resource prediction capabilities. For example, when the digital twin predicts that the edge node will process high-frequency I / O, it can automatically pre-allocate resources, solving the problem of disconnection between resource abstraction and actual demand.

[0082] After completing the abstraction of virtual resources, dynamic binding of location tags based on spatiotemporal topology awareness is required to accommodate the dynamic migration of industrial field equipment. Since the dynamic migration of industrial field equipment can cause changes in resource access paths, such as AGV movement, traditional fixed location tags cannot adapt to dynamic topology changes. Therefore, spatiotemporal features, such as the three-dimensional coordinates of nodes, network hop counts, and link quality timing probabilities, are extracted from the multidimensional resource topology graph. An improved Hilbert curve algorithm is used to generate virtual unit location tags containing spatiotemporal coding.

[0083] A triggered update mechanism is also introduced to automatically update location tags when the topology changes. A topology change threshold is set. When the multi-dimensional resource topology map predicts that the AGV will enter a new area after a period of time, the virtual unit location code is synchronously triggered through the hardware timestamp. The update process includes: recalculating the target area topology characteristics through GNN, remapping the spatiotemporal Hilbert curve encoding, and updating the index of the distributed key-value database. This improves the matching accuracy between the virtual unit location tag and the dynamic location of the device, and solves the problem of longer access paths caused by delayed dynamic topology response.

[0084] Since traditional metadata association does not take industrial business semantics into consideration, resulting in a disconnect between resource allocation and business needs, it is necessary to implement semantically enhanced three-dimensional metadata association modeling, including: building a knowledge graph containing industrial entities, using key information as hash keys based on the DHT index mechanism, mapping them to corresponding nodes in the distributed key-value database, realizing distributed storage and efficient retrieval of metadata, establishing multi-dimensional mapping relationships, and storing them in the distributed key-value database. At the same time, generating an R-tree-bitmap hybrid index that supports spatial range queries, optimizing index distribution in combination with the load balancing characteristics of DHT, and pre-storing high-frequency associated data through the DHT cache strategy based on the access heat map, thereby generating a three-dimensional metadata graph containing device ID, location coordinates, resource ownership, and security level. According to the table, the local resource hit rate of high-frequency small IO operations has been improved, and the problem of disconnection between resource scheduling and control logic has been solved; among them, the multi-dimensional mapping relationship includes spatial dimension, functional dimension and security dimension. The spatial dimension includes: using the range query feature of DHT, matching the Euclidean distance between the device physical coordinates and the virtual unit position encoding, establishing a physical mapping between the device data point and the nearest virtual unit, and setting an error threshold to achieve rapid positioning of the physical mapping; the functional dimension includes: using the knowledge graph to infer the functional attributes of the device data point, and matching the corresponding virtual unit through the functional tags embedded in the DHT key-value pair; the security dimension includes: according to the security level of the data point, forcibly associating it with the virtual unit with hardware fault tolerance in the DHT index.

[0085] Step S2: In the virtual unit, an independent dual-queue architecture is constructed for edge nodes and core nodes, including a scheduling queue and a regular queue. A time-sensitive token bucket is configured for the scheduling queue. A queue clearing mechanism is dynamically triggered based on the industrial control cycle to ensure that real-time tasks are processed first. Continuous small-data-volume IO short transactions generated by edge nodes are merged to reduce IO queue switching overhead. Small data volume refers to data volume that does not exceed a preset data volume threshold.

[0086] Specifically, when building a dual-queue architecture in a virtual unit, a physically isolated dual queue is created for each virtual unit, including a scheduling queue and a regular queue;

[0087] The scheduling queue is used to execute real-time tasks with strict timing requirements, including equipment status reading and writing, control instruction issuance, real-time data acquisition and processing, and emergency fault response instruction transmission. Since real-time tasks in industrial control need to complete responses within a preset control cycle, traditional mixed queues are prone to cause key instructions to be blocked by non-real-time tasks. Therefore, the depth of the scheduling queue is dynamically configured to a preset multiple of the total number of industrial devices associated with the corresponding virtual unit, and a circular buffer structure is used to implement O (1) time complexity enqueue or dequeue operations to ensure rapid response to sudden high-frequency requests, such as batch equipment status query when the production line is started; wherein, the total number of associated industrial devices is the number of industrial devices that establish a three-dimensional metadata mapping relationship with the corresponding virtual unit, that is, the total number of devices marked as virtual unit service objects in the three-dimensional metadata table, such as programmable logic controllers, distributed control systems, and sensors; wherein, the operation time of O (1) does not change with the amount of data, and always responds quickly;

[0088] At the same time, semantic labels are extracted from the three-dimensional metadata table, including task attribute labels and scheduling priority labels. Task attribute labels include safety level, control cycle, and equipment correlation coefficient. Safety levels are ranked from low to high as SIL1, SIL2, SIL3, and SIL4, with SIL4 being the safest and SIL1 being the most basic. Control cycles are ranked from short to long as T1, T2, and T3. Scheduling priority labels are generated based on task attributes. For example, tasks with a safety level of SIL3 or above or a control cycle not exceeding T1 are marked with a "red emergency" label, and tasks with a safety level of SIL2 or a control cycle of T2 are marked with a "yellow important" label. This achieves a direct mapping between control logic and scheduling priority, solves the problem of delay accumulation caused by priority confusion in real-time tasks, and improves the processing efficiency of real-time task queues.

[0089] Conventional queues are used to handle non-real-time tasks with a high tolerance for delays, including equipment operation log storage, process parameter backup, non-critical configuration updates, and system maintenance instruction execution. Considering that non-real-time tasks need to avoid resource starvation, their maximum depth is set to a multiple of the preset depth of the scheduling queue and supports elastic expansion. A hierarchical priority linked list structure is used to achieve differentiated processing. The rules of the hierarchical priority linked list structure are as follows:

[0090] When it is an urgent task, such as device configuration update and fault warning log, priority is given to occupying the first preset ratio queue resources, and the processing delay must not exceed the first preset delay threshold;

[0091] When it is an important task, it occupies the second preset ratio of queue resources by default, and the processing delay must not exceed the second preset delay threshold;

[0092] For ordinary tasks, such as system maintenance logs and non-critical configuration updates, the remaining proportion of queue resources is occupied, and the processing delay is not strictly limited.

[0093] At the same time, tags with allowed delay thresholds are added to regular queue tasks to achieve flexible scheduling of regular business while ensuring real-time task resources, avoiding the problem of low resource utilization caused by a single fixed queue.

[0094] Specifically, the triggering logic of the queue clearing mechanism includes:

[0095] Real-time I / O request vectors are acquired and parsed to extract request features, including request type, source device ID, data size, timeout requirement, and transaction sequence number. Semantic tags, such as security level, control period, device relevance coefficient, and locality index, are simultaneously extracted from the three-dimensional metadata table. The parsed request features are then associated with the semantic tags to form a complete task description. This information is then standardized and converted into a unified format recognizable within the system to generate the final task description set.

[0096] First, tasks are semantically classified in multiple dimensions. The timeout requirement field of the IO request in the task description set is parsed. If the timeout requirement value is less than the preset requirement value or the request instruction is a control instruction, it is marked as a real-time task candidate. Otherwise, it is marked as a regular task and directly enters the regular queue.

[0097] The real-time task candidates are further divided, and the security level and control period in the three-dimensional metadata table are queried to divide the priority labels;

[0098] If the security level is not lower than the first preset level or the control period does not exceed the first preset period, it is marked as an emergency task;

[0099] If the security level is the second preset level or the control period is within the first and second preset periods, it is marked as an important task; the rest are marked as ordinary tasks; wherein the first preset level is higher than the second preset level, and the first preset period is shorter than the second preset period;

[0100] Based on the source device location of the real-time task candidate, the DHT index is queried and sorted in ascending order based on physical distance to obtain the top n candidate virtual units. A comprehensive score is calculated for each candidate virtual unit. The comprehensive score includes a spatiotemporal score and a computing power score. The spatiotemporal score takes into account the number of network hops and physical distance, while the computing power score is based on the ratio of the virtual unit's real-time computing power margin to the predicted peak value. Semantic weights and load weights are also set for weighted calculation to obtain a comprehensive score.

[0101] Prioritize virtual units whose comprehensive scores are greater than a preset score threshold and whose device correlation coefficient is not less than a preset correlation threshold. If the comprehensive scores of all candidate virtual units are less than the preset score threshold, trigger fault self-healing mapping, query the multi-dimensional mapping relationship, obtain the set of backup virtual units within the preset range of the source device, and sort the backup virtual units in descending order of computing power margin. Select nodes with loads lower than the preset load threshold, update the DHT index cache, and temporarily map the source device to the backup virtual unit; thus, generate the target virtual unit.

[0102] Build a DQN model and define a four-dimensional state vector, including the ratio of the current queue depth to the maximum depth, the ratio of the number of remaining tokens to the total number of tokens, the average load rate of the past three cycles, and the peak load rate of the next cycle output by the spatiotemporal prediction model. Then, normalize the four-dimensional state vector.

[0103] Define a discrete action space, including token generation rate adjustment coefficients, such as -20%, -10%, 0, 10%, and 20%. Design a reward function to calculate the reward value based on the ratio of actual delay to timeout requirement and the number of queue overflows. If the actual delay does not exceed the preset ratio of the timeout requirement, an additional reward is given.

[0104] To perform adaptive token bucket adjustment, first generate a basic token number based on the control period. That is, the ratio of the control period to the minimum processing time is defined as the basic token number. Then, input the current state vector into the trained DQN model, output the optimal adjustment coefficient, calculate the compensation token number, and obtain the total token number.

[0105] At the beginning of each control cycle, the scheduling queue is traversed. Tasks with remaining processing time less than the preset ratio of the deadline are directly discarded and recorded as timed-out tasks, freeing up resources for new tasks.

[0106] The token usage rate in the scheduling queue is calculated in real time. When the token usage rate for three consecutive cycles is less than the first usage threshold, nonlinear expansion is performed. The expansion amount is calculated using a square function, so that the expansion rate slows down as the load rate decreases, avoiding excessive bandwidth release.

[0107] If the token usage rate in a single cycle exceeds the second usage threshold, the bandwidth in the regular queue is immediately reduced to the first usage ratio of the basic bandwidth value. The bandwidth is monitored continuously for three cycles. If the token usage rate is still greater than the third usage threshold, the bandwidth is further reduced to the second usage ratio.

[0108] Maintain a transaction dependency matrix for each equipment group, such as a welding robot group, and update it in real time through a sliding window;

[0109] Count the associated devices in the current waiting task, calculate the ratio of the sum of the device association coefficients to the number of associated devices, and obtain the locality index of the current waiting task , set the secondary index threshold, respectively, the first index threshold , the second index threshold ,and , dynamically select the merge mode;

[0110] like , using tight merge mode, setting the window size to , forced to merge tasks in the same cluster; when When using the standard merge mode, set the window size to , allowing tasks within the same community to be merged; when When , the sparse merging mode is used and the window size is set to , only merge tasks in the same device group;

[0111] Based on the ratio of the number of cross-cluster tasks to the total number of tasks, the cross-cluster ratio is calculated, and cross-cluster transactions are intelligently split and sequenced. Once the cross-cluster ratio exceeds the first cross-cluster ratio threshold, the sub-merge groups are divided according to the Louvain community. The cross-cluster ratio of each sub-merge group is less than or equal to the second cross-cluster ratio threshold. At the same time, the locality index of each sub-merge group is recalculated. If the locality index is less than the preset value, it is further split into single-device group tasks. A globally unique merge package sequence number is assigned to each sub-merge group to ensure that cross-node processing is executed in sequence.

[0112] Determine the target virtual unit address for processing the request, mark the request as entering the scheduling queue or regular queue type, and add a priority tag, including urgency and spatial locality tags, to generate a scheduling decision result; among them, the target virtual unit is the local virtual unit to which the edge node belongs, obtained through a multi-dimensional mapping relationship.

[0113] Step S3: Using the physical location of the device as the sharding primary key, the initial shard is bound to the adjacent virtual unit. Cross-virtual unit access traffic is monitored in real time. When the edge node's access to the remote core node exceeds the threshold, the high-frequency data replica is proactively migrated to the virtual unit to which the edge node belongs. The write request of the associated device is dynamically redirected to the virtual unit where the target data is located, reducing cross-node write diffusion. Based on the data correlation of the equipment on the same production line, the sharding granularity is adaptively adjusted to ensure that strongly related data is stored in the core nodes of the same virtual unit.

[0114] Specifically, the steps for adaptively adjusting the sharding granularity include:

[0115] Generate a composite sharding primary key based on the physical location of the equipment. The primary key combines the hash value of the workshop, production line, equipment group information, and physical coordinates.

[0116] Query the three-dimensional metadata table to obtain the set of virtual units corresponding to the device, calculate the Euclidean distance between the device and each virtual unit and arrange them in ascending order, and give priority to selecting the virtual unit with the closest distance and a computing power margin exceeding a preset first margin ratio as the initial binding object. If the computing power margin of the closest virtual unit does not exceed the first margin ratio, then select the virtual unit with the second closest distance and a computing power margin exceeding a preset second margin ratio to ensure that the initial sharding of the device is as close to its physical location as possible. If the computing power margin of the second closest distance does not exceed the preset second margin ratio, trigger the selection of the backup unit; among which, the determination of the second closest virtual unit is based only on distance. If the distances are the same, select the virtual unit with the least network hops.

[0117] Create shard metadata entries in the distributed key-value store, record the shard primary key, target virtual unit, and historical access frequency distribution, and generate a shard metadata set and a shard mapping table;

[0118] Collect access traffic from edge nodes to core nodes in real time and obtain statistical indicators, such as the number of cross-virtual unit accesses, cross-node data volume, and access latency quantiles;

[0119] Trigger conditions are set based on statistical indicators. Once the statistical indicators meet the trigger conditions, the migration mechanism is immediately triggered to determine the source virtual unit where the current data is located and the target virtual unit to which the edge node belongs. Data with high access frequency is migrated first, and the migration is carried out in batches in an incremental manner to avoid affecting real-time tasks. After the migration is completed, the three-dimensional metadata table is updated to mark the location of the data copy. The trigger conditions include the number of cross-virtual unit accesses exceeding the preset access threshold, the key latency indicator exceeding the preset reasonable multiple of the local access, or the proportion of cross-node data exceeding the preset ratio.

[0120] Collect historical data interaction logs between devices, analyze historical interaction frequencies, and perform standardization to obtain correlation coefficients between devices. Use the DBSCAN clustering algorithm to cluster devices with correlation coefficients greater than a preset correlation threshold into a correlation group.

[0121] When the number of devices in an associated group exceeds the preset scale or the frequency of interaction between groups is lower than the preset level, the sharding granularity is dynamically adjusted, including merging multiple small shards into a large shard to reduce management overhead, or splitting into multiple sub-shards to avoid excessive load on a single shard. When merging or splitting, data is allocated according to the computing power status of the virtual unit to ensure strong correlation between device data within the shard. For example, when performing a merge operation, data from multiple shards is migrated to the virtual unit with the largest computing power margin, and the shard primary key of the associated device is updated. When performing a split operation, data is allocated to different virtual units according to the strength of the correlation to ensure that the average correlation coefficient of the devices in each sub-shard exceeds the preset average threshold;

[0122] For device write requests, the optimal target virtual unit is predicted by comprehensively considering the data shard location, virtual unit load, and physical distance to the device;

[0123] If the optimal target virtual unit is a remote node and a local data copy exists, the write request is redirected to the local virtual unit, and the number of copies synchronized across nodes is limited based on the data security level. For example, high-security data synchronizes multiple copies, while ordinary data synchronizes a single copy.

[0124] Monitor the cross-node write request diffusion ratio in real time. When the diffusion ratio exceeds the preset reasonable range, delay batch synchronization of non-critical data and use multi-path parallel synchronization for critical data to reduce the network overhead of cross-node write operations.

[0125] Step S4: Within the data shard, the virtual unit closest to the device group is selected as the master node based on the shard location. This node is responsible for processing write requests and triggering the read optimization mechanism, allowing slave nodes to directly respond to read requests from edge nodes within the same virtual unit. This allows for master-slave node collaborative optimization. The master node regularly synchronizes incremental state summaries to slave nodes to reduce real-time read latency. The master node merges multiple write operations within a short cycle into a single log entry, selects a low-latency path through topology awareness, and synchronizes to the slave node. It uses an asynchronous confirmation mechanism to immediately respond to the edge node.

[0126] Specifically, the specific steps of master-slave node collaborative optimization include:

[0127] Based on the process attribute labels and historical interaction frequencies of the equipment in the 3D metadata table, the equipment with process correlation exceeding the preset correlation threshold is clustered into dynamic equipment groups, such as welding robots and their sensors. The dynamic topology map is then updated based on the physical coordinates and communication relationships of the equipment.

[0128] Monitor the position changes of any device in the device group in real time. When the displacement of a mobile device exceeds a preset distance threshold, trigger the device group reaggregation, query the shard mapping table, obtain the set of virtual units corresponding to the data shards associated with the device group, and filter out the virtual units within a preset range from the geometric center of the device group to form a candidate set.

[0129] Candidate virtual units are evaluated using a comprehensive scoring model to calculate their election scores. When the device group's movement distance exceeds a preset distance threshold or the candidate set's election score change rate exceeds a preset change threshold, a new master node is elected and state synchronization is completed between the old and new master nodes to ensure the continuity of write operations. The election score includes a spatial distance score, a network quality score, a load status score, and a historical service score. The spatial distance score is obtained by normalizing the inverse of the Euclidean distance. The network quality score is calculated using link latency and jitter indicators. The load status score is calculated by comparing the current computing power margin to the predicted load. The historical service score is calculated based on the number of service requests and latency performance for the device group during the past control period.

[0130] When an edge node initiates a read request, it parses the shard location information in the three-dimensional metadata table to determine whether the target data belongs to the current virtual unit. If so, the read request is marked as a local read request and directly responded to by the slave node. Otherwise, it is marked as a remote read request.

[0131] Before a slave node responds, it checks the status difference between the replica and the master node. Once the difference exceeds the preset difference threshold, it synchronizes the latest status to the master node before responding. The master node regularly broadcasts incremental status summaries to the slave nodes. The synchronization cycle is dynamically adjusted according to the system load. When the load is high, synchronization is actively triggered, and only the status information that the slave node lacks is sent, reducing the amount of data transmission.

[0132] The master node dynamically sets a time window according to the control cycle, merges write requests from the same device group, with similar operation types and data volumes within a preset range into a single log entry, and then, based on network topology information, selects a path with the fewest network hops, a link real-time delay below a preset delay threshold, and a bandwidth utilization below a preset ratio to synchronize to the slave node.

[0133] After the master node completes writing the local log, it immediately returns an asynchronous confirmation to the edge node. The background synchronizes the log to the slave node in parallel, and sets a timeout retransmission mechanism to ensure consistency.

[0134] When it is detected that the device group has moved beyond the preset distance threshold, the new master node pre-acquires the data status in advance. The new and old master nodes process write requests in parallel until the state switch is completed. At the same time, the synchronization strategy is adaptively adjusted according to the load situation. Under high load, the merge window is increased to reduce the number of synchronizations, and under low load, the window is shortened to increase the synchronization frequency.

[0135] Differentiated processing is implemented for data of different safety levels. High-safety-level data (such as SIL3) is forcibly synchronized to a preset number of slave nodes and undergoes a final consistency check. Ordinary data is returned for confirmation after being synchronized to a minimum number of slave nodes.

[0136] Example 2:

[0137] See also Figure 5 , another embodiment provided by the present invention: a hyper-converged server resource pooling system, comprising: a perception module, a scheduling control module, a data management module and a read-write optimization module;

[0138] The perception module is used to deploy lightweight probes on industrial equipment gateways to collect spatial information in real time and build static topology maps. It then generates dynamic topology maps based on time series features. It then uses graph neural networks combined with long-short-term memory networks to predict topology change trends. It also builds digital twins for heterogeneous nodes, generates virtual unit location labels with spatiotemporal encoding, and constructs an industrial knowledge graph to achieve multi-dimensional mapping between equipment data points and virtual units, forming a three-dimensional metadata table.

[0139] The scheduling control module is used to build an independent dual-queue architecture for scheduling queues and regular queues in virtual units. It extracts semantic tags from the three-dimensional metadata table to classify task priorities, selects target virtual units based on a spatiotemporal scoring model, dynamically adjusts the token generation rate through a deep reinforcement learning model, triggers a queue clearing mechanism at the start of the control cycle, maintains the device group transaction dependency matrix, dynamically merges short transactions based on the task locality index, and splits cross-cluster transactions by community, achieving efficient scheduling and queue switching optimization for real-time and non-real-time tasks.

[0140] The data management module is used to generate a composite shard primary key based on the physical location of the device, query the three-dimensional metadata table to select adjacent virtual units for initial shard binding, create shard metadata and shard mapping tables in distributed storage, monitor cross-virtual unit access traffic in real time, trigger incremental replica migration when the number of accesses, latency, or data volume exceeds a threshold, analyze the frequency of interactions between devices to dynamically adjust the shard granularity, predict the optimal target virtual unit for device write requests, redirect remote requests to local replicas, control the number of replica synchronizations based on the security level, and reduce cross-node write diffusion overhead;

[0141] The read-write optimization module is used to aggregate dynamic equipment groups based on equipment process relevance. The virtual unit closest to the equipment group is selected as the master node through a comprehensive scoring model. Edge node read requests are responded to by slave nodes first and the replica status is synchronized on demand. The master node merges write requests into a single log entry according to the control cycle, selects a low-latency path based on the network topology to synchronize to the slave node, and uses an asynchronous confirmation mechanism to respond to the edge node immediately. When the equipment group moves, the new and old master nodes are coordinated to process requests in parallel. The synchronization strategy is dynamically adjusted according to the load, and multiple copies of high-security level data are forced to be synchronized and verified to ensure the real-time and consistency of read and write operations.

[0142] Working principle and effect:

[0143] Lightweight probes collect spatial information, construct a multi-dimensional topology map that includes physical location, network parameters, and resource status, and generate virtual units and three-dimensional metadata tables with spatiotemporal coding. This enables spatial abstraction of heterogeneous resources and provides a collaborative view of location and load for subsequent scheduling, solving the problem of inefficient cross-node access caused by the "location blindness" of traditional resource pools.

[0144] A dual-queue architecture is also built. The scheduling queue combines a time-sensitive token bucket with a dynamic clearing mechanism to ensure that real-time tasks are processed first. Micro-batch merging reduces I / O switching overhead. Regular queues use elastic scaling strategies to avoid resource starvation, effectively solving the problems of accumulated delays and low resource utilization caused by priority confusion in real-time tasks.

[0145] The physical location of the device is used as the sharding primary key, and the adjacent virtual units are initially bound. Cross-unit access is monitored in real time, and high-frequency data copies are dynamically migrated. At the same time, the sharding granularity is adaptively adjusted based on data relevance to reduce cross-node write diffusion and ensure local storage of strongly associated data. This solves the problems of cross-node delay and unreasonable data sharding caused by dynamic topology changes. The adjacent master node is dynamically elected within the shard, allowing the slave node to directly respond to local read requests. The master node improves read and write efficiency through batch merging of write operations, topology-aware path synchronization, and asynchronous confirmation mechanism. Combined with the smooth switching of the master node when the device group moves and the differentiated processing of security levels, it balances real-time performance and data consistency.

[0146] Overall, through the four-layer architecture of topology awareness, intelligent scheduling, dynamic sharding, and master-slave collaboration, a complete resource pooling system is formed, which deeply integrates the physical location of equipment, network topology and business semantics, and realizes accurate prediction and response to the dynamic behavior of industrial equipment. It effectively solves the real-time and consistency problems caused by dynamic topology changes, resource scheduling lags and cross-node access in industrial control, and provides an efficient resource management solution for intelligent manufacturing.

[0147] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for pooling hyper-converged server resources, characterized in that: include: Probes are deployed on industrial device gateways to collect spatial information, generate a multi-dimensional resource topology map, abstract the resources of edge and core nodes into virtual units with location tags, build a virtual resource pool, and generate a three-dimensional metadata table. An independent dual-queue architecture is built in the virtual unit and configured with a token bucket. A queue clearing mechanism is triggered based on the control cycle to merge and process short transactions on edge nodes. Short transactions are IO operations whose data volume does not exceed the data volume threshold. The dual-queue architecture includes a scheduling queue and a regular queue. The scheduling queue is used to execute real-time tasks, and the regular queue is used to process non-real-time tasks. The system uses the physical location of the device as the sharding primary key to bind adjacent virtual units, monitors cross-virtual unit access traffic, migrates data replicas whose access frequency exceeds the access threshold, redirects write requests, and adjusts the sharding granularity based on data relevance. Build a dynamic device group, select the virtual unit closest to the device group within the data shard as the master node, perform master-slave node collaborative optimization, allow the slave node to respond to local read requests, the master node synchronizes incremental status summaries to the slave nodes, merges write operations and performs path synchronization, and uses an asynchronous confirmation mechanism to respond to edge nodes; The steps of generating a three-dimensional metadata table include: Build a digital twin for each node, collect underlying hardware parameters, use reinforcement learning models to predict resource usage trends, and generate virtual units; nodes include edge nodes and core nodes; Extracting spatiotemporal features from the multidimensional resource topology map to generate virtual unit location tags containing spatiotemporal codes; Set up a triggered update mechanism to automatically update location labels when the topology changes; Build a knowledge graph containing industrial entities, map key information as hash keys to corresponding nodes in a distributed key-value database based on the DHT indexing mechanism, and establish a multi-dimensional mapping relationship; Generate a hybrid index, combine the load balancing characteristics of DHT to optimize index distribution, and pre-store related data with access frequency exceeding the access threshold through DHT's cache strategy based on the access heat map to generate a three-dimensional metadata table.

2. The method for pooling hyper-converged server resources according to claim 1, wherein: The generation logic of the multi-dimensional resource topology map includes: Based on the physical coordinates of the devices, the regions are divided, and the devices with spatial distance less than the preset distance and signal strength greater than the preset strength are grouped into geographic clusters, and a weighted undirected graph is constructed; Based on the correlation between edge weights and node functions in the weighted undirected graph, the control subnet topology cluster is identified, and the edge nodes and corresponding core nodes of the same production line are divided into the same cluster; Extracting three types of time series features from historical data, namely, device location migration patterns, load fluctuation characteristics, and network quality characteristics, and injecting them into the weighted undirected graph to expand it into a dynamic topology graph with a time dimension; A spatiotemporal prediction model is constructed by combining graph neural network with long short-term memory network. The current topological state and historical time series features are input and the topological change prediction results are output.

3. The method for pooling hyper-converged server resources according to claim 2, wherein: The triggered update mechanism includes: A topology change threshold is set. When a device is predicted to enter a new area, the virtual unit position encoding update is triggered synchronously through the hardware timestamp. This includes recalculating the target area topology features through GNN, remapping the spatiotemporal Hilbert curve encoding, and updating the index of the distributed key-value database.

4. The method for pooling hyper-converged server resources according to claim 3, wherein: When building a dual-queue architecture, create physically isolated dual queues for each virtual unit, including a scheduling queue and a regular queue; The scheduling queue is used to execute real-time tasks with timing requirements. The depth of the scheduling queue is dynamically configured to a preset multiple of the total number of industrial devices associated with the corresponding virtual unit. A circular buffer structure is used to implement enqueue or dequeue operations. Task attribute tags including security level, control period, and device correlation coefficient are extracted from the three-dimensional metadata table to generate scheduling priority tags and classify real-time tasks. The conventional queue is used to process non-real-time tasks. The maximum depth of the conventional queue is set to a multiple of the preset depth of the scheduling queue and supports elastic expansion. A hierarchical priority linked list structure is adopted, and different proportions of queue resources are occupied according to different task categories. Labels with allowed delay thresholds are added to the conventional queue tasks, where task categories include urgent tasks, important tasks, and ordinary tasks.

5. The method for pooling hyper-converged server resources according to claim 4, wherein: The triggering logic of the queue clearing mechanism includes: Obtaining and parsing real-time IO request vectors, extracting request features, and simultaneously extracting semantic tags from the three-dimensional metadata table, performing association and standardization processing, and generating a task description set; Performing multi-dimensional semantic classification on the tasks in the task description set to divide them into real-time task candidates and regular tasks, and placing the regular tasks into the regular queue; Based on the security level and control period corresponding to the real-time task candidates, priority labels are divided; Based on the source device location of the real-time task candidate, query the DHT index and sort it in ascending order of physical distance to obtain the top n candidate virtual units, and calculate the comprehensive score of each candidate virtual unit; Prioritize virtual units whose comprehensive scores are greater than a preset score threshold and whose device correlation coefficient is not less than a preset correlation threshold as target virtual units; If all comprehensive scores are less than the preset score threshold, the fault self-healing mapping is triggered, the multi-dimensional mapping relationship is queried, the set of spare virtual units is obtained, and the nodes with loads lower than the preset load threshold are selected after sorting in descending order of computing power margin to generate the target virtual unit.

6. The method for pooling hyper-converged server resources according to claim 5, wherein: The triggering logic of the queue clearing mechanism also includes: Constructing a DQN model, generating a base number of tokens based on a control cycle, inputting a state vector into the DQN model to output an adjustment coefficient, calculating a compensation number of tokens, and obtaining a total number of tokens; At the start of each control cycle, the scheduling queue is traversed and tasks whose remaining processing time is less than a preset ratio of the deadline are discarded; Calculating the token usage rate in the scheduling queue in real time, and performing nonlinear expansion when the token usage rate is less than a first usage threshold for three consecutive periods; When the single-cycle token usage rate is greater than a second usage threshold, shrinking the bandwidth of the regular queue, and continuing to shrink if it continues to be greater than a third usage threshold; Maintain a transaction dependency matrix for each device group and update it in real time, count the associated devices in the current waiting task, calculate the locality index, and dynamically select the tight merge mode, standard merge mode, or sparse merge mode based on the locality index; The ratio of the number of cross-cluster tasks to the total number of tasks is calculated to obtain the cross-cluster ratio. When the cross-cluster ratio is greater than the first cross-cluster ratio threshold, the sub-merge groups are divided to ensure that the cross-cluster ratio of each sub-merge group is less than or equal to the second cross-cluster ratio threshold. The locality index is recalculated for each sub-merge group. If the locality index is less than the preset value, the sub-merge group is further split into single-device group tasks and a globally unique merge package serial number is assigned.

7. The method for pooling hyper-converged server resources according to claim 6, wherein: The specific steps to adjust the sharding granularity include: Generate a composite shard primary key based on the physical location of the device; Query the three-dimensional metadata table to obtain a set of virtual units, sort them in ascending order by Euclidean distance, select the virtual unit with the closest distance and whose computing power margin exceeds a preset first margin ratio to bind to the device, and if the conditions are not met, select the virtual unit with the second closest distance and whose computing power margin exceeds a preset second margin ratio to bind to the device; Create shard metadata entries in the distributed key-value store and generate a shard mapping table; Collect access traffic from edge nodes to core nodes in real time and obtain statistical indicators. When the statistical indicators meet preset trigger conditions, determine the source virtual unit and the target virtual unit, prioritize the migration of data copies with access frequencies exceeding the access threshold, and update the three-dimensional metadata table. Collect device interaction logs and calculate the correlation coefficient. Devices with a correlation coefficient greater than a preset correlation threshold are grouped into correlation groups. The sharding granularity is adjusted based on the size of the correlation group or the interaction frequency. When processing a write request, the optimal target virtual unit is predicted. If it is a remote node and there is a local replica, it is redirected. The number of replicas is limited according to the security level, and the write diffusion ratio is monitored to adjust the synchronization strategy.

8. The method for pooling hyper-converged server resources according to claim 7, wherein: The specific steps of master-slave node collaborative optimization include: Based on the three-dimensional metadata table, the devices whose process correlation exceeds a preset correlation threshold are grouped into a dynamic device group, and a dynamic topology map is updated; Monitor the position changes of devices in the device group in real time. Once the device displacement exceeds a preset distance threshold, trigger the device group reaggregation, query the shard mapping table, and filter out virtual units within a preset range from the geometric center of the device group to form a candidate set; The election score of candidate virtual units is calculated using a comprehensive scoring model. When the device group's movement distance or the election score change rate exceeds a threshold, a new master node is elected and the status of the new and old master nodes are synchronized. When an edge node initiates a read request, it parses the shard location information. If the target data belongs to the current virtual unit, the slave node responds directly. Otherwise, it is marked as a remote read request. Before responding, the slave node checks the status difference with the master node. If the difference exceeds the difference threshold, it synchronizes the latest status first. The master node merges write requests according to the control cycle, selects a path based on the network topology, synchronizes to the slave node, and asynchronously confirms; When the device group moves, the master node is switched and the synchronization strategy is adaptively adjusted, with differentiated synchronization based on data security levels.

9. A hyper-converged server resource pooling system, configured to implement the hyper-converged server resource pooling method according to any one of claims 1 to 8, characterized in that: include: Perception module, scheduling control module, data management module and read-write optimization module; The perception module is used to collect spatial information and construct a multi-dimensional resource topology map, abstract heterogeneous node resources into virtual units with location tags, establish a multi-dimensional mapping relationship between device data points and virtual units, and generate a three-dimensional metadata table; The scheduling control module is used to build a dual-queue architecture, extract task priority tags based on a three-dimensional metadata table, combine a spatiotemporal scoring model with an adaptive token bucket mechanism to implement task scheduling, and dynamically merge and split short transactions; The data management module is used to perform initial shard binding using the physical location of the device as the shard primary key, monitor cross-virtual unit access and trigger high-frequency data copy migration, dynamically adjust shard granularity based on device data relevance, redirect write requests and control cross-node write propagation; The read-write optimization module is used to aggregate dynamic device groups based on device process relevance, elect a master node to process write requests and allow slave nodes to respond to local read requests, merge write operations and synchronize them to slave nodes through topology-aware paths, and adjust synchronization strategies based on device group movement and load status.

Citation Information

Patent Citations

  • Resource scheduling method, system and device in edge cloud computing platform and medium

    CN120123099A

  • Computing power resource scheduling method and system for AI multi-service data center

    CN120196420A