Hyper-converged server multi-resource integration system and scheduling method

Through distributed data collection and graph neural network algorithms, a multidimensional index structure and causal relationship library are constructed, which solves the problem of low efficiency in multi-source data integration in hyper-converged servers, realizes dynamic modeling of resource association relationships and accurate anomaly positioning, and improves the effectiveness of resource scheduling strategies and system reliability.

CN120561343BActive Publication Date: 2025-10-10BEIJING ZHONGKE JIANYOU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511046543.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-10
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Hyper-converged servers have problems such as low efficiency in integrating heterogeneous multi-source data, difficulty in real-time dynamic modeling of resource relationships, insufficient accuracy in locating the root causes of resource anomalies, lack of systematic optimization of resource scheduling strategies, and lack of dynamic closed-loop management mechanisms.

Method used

Multi-source data is collected through distributed collection nodes, and after semantic alignment and format conversion, it is stored in a distributed time series database architecture, and a standard data set with a multi-dimensional index structure is constructed. Combined with graph models and graph neural network algorithms, time series attributes and spatial topological attributes are extracted, directed causal edges are generated, and a causal relationship library is built. Anomalies are monitored in real time and resource scheduling strategies are generated. Scheduling is performed through a multi-objective optimization engine to achieve dynamic closed-loop management.

Benefits of technology

It achieves efficient and standardized integration of heterogeneous multi-source data, captures resource topology evolution and temporal dependencies in real time, improves the accuracy of anomaly root cause location and the effectiveness of resource scheduling strategies, and improves resource integration efficiency and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561343B_ABST
    Figure CN120561343B_ABST
Patent Text Reader

Abstract

The application discloses a super-converged server multi-resource integration system and a scheduling method, belongs to the technical field of computer resource management, and aims to overcome the problems of traditional super-converged server resource scheduling, such as dispersion, low storage efficiency, and difficulty in identifying abnormalities. Multi-source data is acquired by relying on distributed acquisition nodes, stored by using a distributed time sequence database architecture, a standard data set containing multi-dimensional indexes is constructed, and efficient data storage and rapid retrieval are realized. A graph model is constructed by positioning target data, a layered and partitioned distributed graph database is built, and dynamic updating is realized by capturing change events in real time. The causal relationship of resources is mined by using a graph neural network, and effective causal chains are screened. At the same time, the weight of the monitoring node and the connection edge is monitored according to the graph model, and differential monitoring is implemented to accurately identify potential abnormal points, and a resource scheduling strategy is generated in combination with the causal relationship. In execution, real-time feedback adjustment is realized, the database is updated, multi-resource deep integration and intelligent scheduling are achieved, and the system resource utilization rate and stability are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer resource management, and more particularly to a hyper-converged server multi-resource integration system and a scheduling method. BACKGROUND

[0002] As a key architecture for realizing resource intensive management in data centers, hyper-converged servers are facing multi-dimensional technical challenges when dealing with rapid business iteration and cluster scale expansion. With the deep integration of heterogeneous hardware devices and multi-element business systems, the format differences and semantic ambiguity of data interfaces from different manufacturers are increasingly prominent. The traditional centralized data processing architecture lacks a standardized adaptation mechanism, resulting in low efficiency of multi-source data integration and difficulty in forming a unified resource management data base. The dynamic evolution characteristics of resource topology relationships and business logic dependencies make it impossible for static modeling methods to capture the timing changes of physical connections and logical interactions in real time. The results of correlation analysis often lag behind the actual system running state, making it difficult to support fine-grained resource scheduling requirements. In terms of exception management, the existing threshold alarm mechanism cannot effectively distinguish between accidental fluctuations and systematic failures, and lacks the ability to trace the cause of abnormal propagation paths, frequently causing false positives and false negatives or deviations in root cause positioning. In the process of generating resource scheduling strategies, the optimization mode based on experience rules or a single target makes it difficult to achieve a global optimal balance between resource utilization, business continuity and operation cost, resulting in a significant deviation between the actual effect and the expected effect. At the same time, the management system lacks a closed-loop adjustment link from exception identification, strategy execution to effect feedback, and cannot dynamically optimize resource allocation according to real-time running state, which may lead to system performance degradation and resource fragmentation in the long run. Therefore, in order to overcome these limitations, the present application proposes a hyper-converged server multi-resource integration system and a scheduling method. SUMMARY

[0003] In view of the deficiencies in the prior art, the present application aims to provide a hyper-converged server multi-resource integration system and a scheduling method, which solves the problems of low efficiency and difficulty in standardization of heterogeneous multi-source data integration, difficulty in real-time dynamic modeling of resource correlation relationships, insufficient accuracy of resource exception root cause positioning, lack of systematization in resource scheduling strategy optimization, and lack of dynamic closed-loop management mechanism leading to inability of resource allocation to continuously evolve.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0005] The hyper-converged server multi-resource integration system comprises:

[0006] The multi-source data sets are collected by distributed collection nodes, and after semantic alignment and format conversion, they are stored using a distributed time series database architecture to build a standard data set containing a multi-dimensional index structure;

[0007] The multi-dimensional index structure based on a standard data set locates target data of a graph model, constructs the graph model, constructs a distributed graph database by using hierarchical partitioning, and triggers graph database updating by capturing change events;

[0008] The time sequence attribute and the spatial topology attribute of the graph node are extracted by a graph neural network algorithm, a time sequence graph is constructed in combination with a sliding window, a directed causal edge is generated by using a structural causal model, and an effective causal chain is screened and stored in a causal relationship database; when the graph database is updated, a causal relationship incremental updating process is started, and the effective causal chain of the causal relationship database is updated by dividing a local graph model;

[0009] Based on the monitoring weight of the graph node and the connection edge in the graph model, the collection period is adjusted on the basis of the basic monitoring frequency, differential monitoring is implemented through performance index data, potential abnormal points are identified, a target causal set is constructed through the causal relationship database, a multi-objective optimization function is constructed through a multi-objective optimization engine, a target candidate strategy of resource scheduling is generated, and the target node is executed to perform resource scheduling;

[0010] Real-time monitoring of the execution result is performed to implement monitoring feedback adjustment, and the graph database and the causal relationship database are updated.

[0011] Specifically, the specific steps of identifying potential abnormal points include:

[0012] According to the historical performance index data of the graph model, the performance index data of the graph node and the connection edge are statistically analyzed to generate a basic threshold range including the mean and the standard deviation;

[0013] According to the frequency of occurrence of potential abnormal points of the graph node and the connection edge in the graph model, the basic threshold range of the graph node and the connection edge is adjusted to a differential threshold, and a dynamic threshold range of the performance index data of the graph node and the connection edge is constructed;

[0014] The performance index data collected in real time from the graph model and the corresponding dynamic threshold range are compared, and if they are within the corresponding dynamic threshold range, time sequence mutation monitoring is performed, the change amplitude of the performance index data is monitored, and it is judged whether resource abnormality early warning is triggered;

[0015] If the performance index data collected in real time is outside the corresponding dynamic threshold range, the collection point is marked as a potential abnormal point, the potential abnormal point is identified as an outlier, if the potential abnormal point is determined as an outlier, time sequence mutation monitoring is performed, and it is judged whether resource abnormality early warning is triggered;

[0016] If the potential abnormal point is determined as a non-outlier, resource scheduling path analysis is triggered, the graph node where the potential abnormal point is located is marked as an abnormal graph node, and the connection edge where the potential abnormal point is located is marked as an abnormal connection edge.

[0017] Specifically, the specific steps of generating the target candidate strategy of the resource scheduling include:

[0018] When the resource scheduling path analysis is triggered, a double-layer screening mechanism is started in the causal relationship library based on the timestamp of the potential abnormal point, target causal chains are screened, and a target causal chain set is constructed;

[0019] The causal subject and the causal object of the target causal chain are extracted, mapped to the graph node in the current graph model through the resource ID of the graph database, the root causal subject is traced back, and the graph model to be subjected to resource scheduling in the graph database is identified as a target graph model;

[0020] For the graph node of the root causal subject and the graph node associated through the target causal chain in the target graph model, a multi-objective optimization function is constructed through a multi-objective optimization engine to generate a candidate strategy set; the multi-objective optimization function includes minimizing the resource usage rate of the graph node of the root causal subject and the graph node associated through the target causal chain, minimizing the resource scheduling business link loss, and minimizing the execution cost of the resource scheduling operation;

[0021] The candidate strategy is verified through a cluster simulation model, a target candidate strategy is selected, and is decomposed into an execution instruction chain and is issued to a scheduling execution module, and the causal strength of the target causal chain in the causal relationship library is updated according to the execution result.

[0022] Specifically, the specific steps of executing the resource scheduling and monitoring the execution result for monitoring feedback adjustment include:

[0023] The execution instruction chain of the target candidate strategy is parsed into an atomic operation sequence through a resource orchestration engine, an instruction chain is generated according to the operation type and the target node, and is mapped into a call sequence, and the call sequence is issued to the target node through a message queue;

[0024] The call sequence is executed on the target node in sequence, and a scheduling quiet period of the target node is set, a scheduling observation period is configured according to the monitoring weight of the target node, high-frequency abnormal monitoring is performed on the target node in the scheduling observation period of the resource scheduling, and a potential abnormal point is identified;

[0025] The frequency of the occurrence of the potential abnormal point in the scheduling observation period is counted, a scheduling abnormal threshold is configured, it is judged whether a scheduling abnormality warning is triggered, if yes, the resource scheduling path is blocked, the call sequence execution is suspended through the message queue, a rollback operation is performed, and the scheduling quiet period of the target node is removed;

[0026] The pseudo causal chain that causes the triggering of the scheduling abnormality warning is eliminated, the target causal chain of the causal relationship library is re-identified, the target candidate strategy is updated, the call sequence is mapped, and the call sequence is executed on the target node again;

[0027] If the frequency of the potential abnormal point is less than or equal to the scheduling exception threshold, it is confirmed that the exception is removed, the resource scheduling observation is stopped, and the abnormal state label and the scheduling timestamp of the node in the graph database are synchronously updated;

[0028] The positioning call sequence involves the graph nodes and the connection edges in the resource scheduling path, and the causal relationship data volume and the graph database are updated, and the call sequence execution whole-process log is recorded.

[0029] Specifically, the specific steps of the effective causal chain screening include:

[0030] The graph model in the distributed graph database is respectively subjected to feature extraction through a graph neural network algorithm, the time sequence attribute and the spatial topology attribute of the graph node are taken as input, and a time sequence graph structure of the resource state transition is constructed;

[0031] The causal influence direction between the graph nodes of the quantized time sequence graph structure is quantified, and a directed causal edge containing a time lag relationship is generated;

[0032] The extracted directed causal edge is subjected to statistical significance verification, and an effective causal chain is screened through setting a causal strength threshold;

[0033] When the graph database triggers an update, an incremental causal relationship update process is automatically started, the affected local graph model in the graph model is divided, the time sequence attribute and the spatial topology attribute of the graph node in the local graph model are re-extracted, and the directed causal edge is updated through an incremental graph neural network;

[0034] The effective causal chain is converted into a causal record according to a standardized structure and stored in a causal relationship database, and the causal record includes a causal subject, a causal object, a causal type, a causal strength and a storage timestamp;

[0035] Through an aging mechanism, the causal record in the causal relationship database that exceeds a preset time period and is not used is deleted, and a clustering algorithm is used to perform similarity merging on the causal record, and the causal relationship database is optimized.

[0036] Specifically, the specific steps of the graph database include:

[0037] Based on a multi-dimensional index structure of a standard data set, a target data of a graph model is located;

[0038] According to a resource entity in the target data of the graph model, a graph node is constructed, a collection timestamp and a spatial topology label are reused to give the graph node time and space attributes, and real-time performance index data is obtained as a state attribute;

[0039] According to the spatial topology label in the standard data set, a physical connection edge of the graph node is generated, and according to a timestamp sequence of a business interaction record of the graph node, a time sequence graph between the graph nodes is constructed, a logical dependence strength between the graph nodes is calculated, a graph node pair is selected, and a logical dependence edge is generated.

[0040] For the connection edges of the graph nodes, according to the graph model target data of the standard data set, the request times and traffic data of the physical connection edges and the logical dependency edges are extracted, and the edge weights of the physical connection edges and the logical dependency edges are respectively given according to a weighted priority mapping algorithm;

[0041] A hierarchical partitioning strategy is adopted to construct a distributed graph database storage architecture, the graph nodes and the connection edges in the graph model are stored in two levels of shards according to resource types and spatial positions, and a multi-dimensional composite index system is established.

[0042] Specifically, the specific steps of the graph database further include:

[0043] Real-time capture of change events of the standard data set, the change events including resource state change, resource configuration adjustment or business rule change, when the change events of the standard data set are captured, triggering the graph database update process:

[0044] According to the change event type and the resource ID, through the multi-dimensional composite index system of the distributed graph database, the affected graph model target shard area is located;

[0045] For the located target shard area, the graph node attributes of the involved graph model are updated, the connection edges of the graph nodes are dynamically adjusted, and the edge weights of the connection edges are updated;

[0046] After the graph database update process is completed, through the transaction consistency mechanism of the distributed graph database, the integrity of the updated graph model target shard area is checked, and the update result is synchronized to the remaining shard areas of the graph database.

[0047] Specifically, the construction steps of the standard data set include:

[0048] Through the distributed message queue, the multi-source data set is received, the distributed time series database is initialized, and the data docking channel is established;

[0049] Based on the domain ontology, a semantic model is constructed, the standard terminology and relationship attributes of the resource entity are defined, and a standardized semantic system is formed;

[0050] Through the rule engine, the heterogeneous format terms of the multi-source data set are mapped to a unified semantic space, a plug-in adaptive conversion architecture is adopted, and conversion plug-ins are dynamically loaded to convert the heterogeneous format terms into standardized columnar data;

[0051] The distributed node collection timestamp is recorded, and the network topology information is obtained, and the spatial topology label of the multi-source data set is dynamically generated;

[0052] In the distributed time series database, LSM tree index and columnar storage are adopted, and the data of the multi-source data set is classified according to the time dimension, access frequency and business priority of the multi-source data set, so as to adjust the shard storage and the number of replicas;

[0053] The spatial topology label and the resource type identifier are fused to construct a multi-dimensional index structure containing time, space and resource type, and a metadata index is constructed. The sensitive data is desensitized and encrypted, and the data operation behavior is recorded.

[0054] Specifically, the collection period is adjusted based on the basic monitoring frequency, and the specific steps of implementing differentiated monitoring through performance indicator data include:

[0055] Based on the historical activity, business criticality and data interaction frequency of each graph node in the graph database, the weight coefficients of the feature dimensions are configured through multi-dimensional feature fusion, and the node weight in the range of 0 to 1 is generated by weighted summation;

[0056] The weight classification threshold is configured, including the upper weight classification threshold and the lower weight classification threshold, and the differentiated monitoring is implemented according to the monitoring weight, including the node weight and the edge weight;

[0057] When the monitoring weight is greater than the upper weight classification threshold, high-frequency anomaly monitoring is adopted, and the collection period is shortened according to the configured shortening ratio based on the preset basic monitoring frequency;

[0058] When the monitoring weight is less than the lower weight classification threshold, low-frequency anomaly monitoring is adopted, and the collection period is increased according to the configured growth ratio based on the preset basic monitoring frequency;

[0059] When the monitoring weight is less than or equal to the upper weight classification threshold and greater than or equal to the lower weight classification threshold, the collection period is set according to the basic monitoring frequency for anomaly monitoring;

[0060] According to the dynamic collection period, the graph nodes or connection edges are monitored differently, and the performance indicator data of the graph nodes and the connection edges are collected.

[0061] The multi-resource scheduling method of the hyper-converged server includes the following steps:

[0062] Step S1: Collecting multi-source data sets through distributed collection nodes, performing semantic alignment and format conversion on the multi-source data sets, and then storing them in a distributed time series database architecture to construct a standard data set containing a multi-dimensional index structure;

[0063] Step S2: Locating the graph model target data based on the multi-dimensional index structure of the standard data set, constructing a graph model, constructing a distributed graph database through hierarchical partitioning, and triggering graph database updates through capture change events;

[0064] Step S3: The graph neural network algorithm is used to extract the temporal attributes and spatial topological attributes of the graph nodes. A temporal graph is constructed by combining the sliding window. The structural causal model is used to generate directed causal edges, and valid causal chains are selected and stored in the causal relationship library. When the graph database is updated, the causal relationship incremental update process is started to update the valid causal chains in the causal relationship library by dividing the local graph model.

[0065] Step S4: Based on the monitoring weights of the graph nodes and connecting edges in the graph model, the collection period is adjusted on the basis of the basic monitoring frequency. Differentiated monitoring is implemented through performance indicator data to identify potential anomalies. The target causal set is constructed through the causal relationship library. The multi-objective optimization function is constructed through the multi-objective optimization engine. The target candidate strategy for resource scheduling is generated and sent to the target node for resource scheduling.

[0066] Step S5: Monitor the execution results in real time, perform monitoring feedback adjustments, and update the graph database and causal relationship library.

[0067] Beneficial effects of the present invention:

[0068] The present invention realizes efficient standardized integration and elastic storage of heterogeneous multi-source data through a distributed time series database with multi-dimensional indexing, effectively solving the problems of data format heterogeneity and semantic ambiguity; with the help of a hierarchical and partitioned dynamic graph database, it captures resource topology evolution and time series dependencies in real time, realizes dynamic modeling of physical connections and logical interactions, and provides an accurate carrier for resource association analysis; uses graph neural networks and structural causal models to deeply mine time series causal chains, constructs a causal relationship library with both statistical significance and business rationality, improves the accuracy of locating the root cause of anomalies and reduces the misjudgment rate; based on a multi-objective optimization engine combined with cluster simulation, it generates a global optimization strategy that takes into account multi-dimensional needs, breaking through the limitations of traditional single-objective scheduling; accurately distinguishes anomaly types through dynamic threshold adjustment and outlier identification mechanism, and realizes continuous evolution of resource management in combination with a closed-loop feedback regulation system, ultimately greatly improving the resource integration efficiency, scheduling strategy effectiveness and system reliability of hyper-converged servers. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a schematic diagram of the structure of the multi-resource integration system of the hyper-converged server of the present invention;

[0070] Figure 2 A flowchart of the specific steps for constructing the standard data set of the present invention;

[0071] Figure 3 A flowchart of the specific steps of constructing a graph model of the present invention;

[0072] Figure 4 Generate a flow chart of resource scheduling strategy for the present invention;

[0073] Figure 5 Flow chart of the multi-resource scheduling method of the super-converged server of the present application. DETAILED DESCRIPTION

[0074] Embodiment 1

[0075] Please refer to Figure 1 This embodiment introduces a multi-resource integration system of a super-converged server, which includes an edge collection module, a converged storage module, a correlation decision module and a scheduling execution module.

[0076] The edge collection module is used to collect the multi-resource data such as calculation, storage and network of each collection node in the super-converged server in real time through the distributed collection nodes, and perform edge preprocessing to build a multi-source data set for the upper system analysis. A standardized driving interface is built through a hardware abstraction layer to uniformly encapsulate the bottom layer protocols of ARM, X86 servers and storage and network devices, to realize the driving adaptation of heterogeneous access devices, to collect original multi-source data, to use adaptive data analysis algorithms for each distributed node, to match the original multi-source data format through preset templates and dynamic rules based on the data type and collection frequency of the original multi-source data, to complete the format unification of heterogeneous original multi-source data, to clean the invalid data of the original multi-source data through a general rule base and semantic analysis, and to reduce the data volume by combining a hybrid compression algorithm to form a multi-source data set with unified format, compact compression and timeliness.

[0077] In this embodiment, a lightweight collection agent program is deployed as a collection node in each physical node or rack unit in the super-converged server cluster, and each collection node forms a distributed collection network through a high-speed backplane or a gigabit Ethernet to realize parallel data collection. The hardware abstraction layer adopts a layered driving architecture design, which is divided into a device adaptation layer, a protocol conversion layer and a data abstraction layer. The ARM, X86 servers and storage and network devices are adapted through standardized driving plug-ins, the private protocols are uniformly converted and the standard interfaces are provided to realize the unified access of heterogeneous devices. The adaptive data analysis algorithm automatically identifies the data type and collection frequency, selects the corresponding data format template from the pre-defined template library, and converts the data according to the template to unify the heterogeneous data format. The general rule base integrates multiple invalid data identification rules and cleans the data in combination with semantic analysis; the hybrid compression algorithm adopts a hierarchical strategy to use Snappy for fast compression of real-time data and Zstandard for deep compression of non-real-time data, and uses CRC check to guarantee data integrity. Each collection node is built-in with an edge cache queue, which adopts a ring buffer design to cache data when the network is temporarily congested to avoid data loss. At the same time, the collection node supports the breakpoint resume function to automatically supplement the unsent data after the network is restored to ensure the integrity and timeliness of the multi-source data set.

[0078] The fusion storage module is used for semantic alignment and format conversion of the multi-source data set output by the edge collection module, and stores intelligently using a distributed time series database architecture, constructs a multi-dimensional index structure including time, space and resource type dimensions, and forms a standard data set supporting high-speed retrieval and cross-domain correlation analysis. By establishing a semantic model to unify the semantic representation of multi-source data, an adaptive conversion mechanism is used to automatically adapt to data format changes caused by hardware upgrades or business changes; a clock synchronization algorithm is used to eliminate time errors between nodes, and a topology monitoring technology is used to update data space topology labels in real time, to realize data time and space attribute labeling. During storage, the write performance of the time series database is optimized, and a redundancy strategy is used to ensure data reliability; at the same time, based on access modes such as data read-write frequency and real-time requirements, the size of data shards and the number of replicas are dynamically adjusted by algorithms, to reduce storage costs while ensuring data availability. Finally, a standard data set with fusion time and space attributes, efficient storage and standardized structure is output, to provide structured data support for upper intelligent correlation analysis, decision scheduling and other modules.

[0079] In the embodiment, the fusion storage module uses a distributed time series database architecture, and realizes multi-source data access buffering through a message queue. Based on a semantic model to unify data terminology, a plug-in mechanism is used to dynamically adapt to data format changes of hardware and business systems. A hybrid clock synchronization strategy is used to eliminate time errors between nodes, and software-defined network and container orchestration technology are used to update data space topology labels in real time. Index optimization and columnar storage are used to improve write performance, and a hybrid redundancy strategy of multiple replicas and erasure coding is used to ensure data reliability. Based on online learning algorithms to analyze data access patterns, the shard size and replica number are dynamically adjusted to implement differentiated storage strategies for cold and hot data. Finally, the data is stored in a standardized columnar format, a multi-dimensional index structure is constructed, and fast time range queries and cross-resource correlation retrieval are supported.

[0080] See Figure 2 , preferably, the specific steps of constructing the standard data set include:

[0081] The data flow in the multi-source data set output by the edge collection module fluctuates, and the real-time requirements of different resource types differ significantly. The distributed message queue is used to receive the multi-source data set output by the edge collection module, and the flow control mechanism is used to realize dynamic buffering and load balancing of data flow. A message classification mechanism based on resource type priority is used, priority is divided according to real-time requirements, and multi-source data sets are managed by queue according to super-converged server resource types, including computing resources, storage resources, network resources, etc., to ensure that high-timeliness data is processed first; at the same time, a distributed time series database architecture is initialized, a data docking channel with the distributed message queue is established, and a foundation is laid for subsequent data storage.

[0082] Hyper-converged servers involve multiple vendors and multiple types of resources, and the original data terms are heterogeneous. A semantic model is built based on the domain ontology to define the standard terms and relationship attributes of hyper-converged server resource entities. Resource entities include computing resource classes, storage resource classes, network resource classes, security resource classes, etc. Object attributes are used to define the relationship between resource entities, and data attributes are used to define entity characteristics to form a standardized semantic system.

[0083] Hardware vendor upgrades or business system changes can cause dynamic changes in data formats. This requires ensuring data format compatibility with the distributed time series database. Using a rules engine, heterogeneous terms in multi-source datasets are mapped to a unified semantic space. For example, different expressions such as CPU usage and memory utilization can be uniformly mapped to standard concepts such as computing resources, processors, and utilization. Using a plug-in-based adaptive conversion architecture, corresponding conversion plug-ins are dynamically loaded to address data format changes caused by hardware vendor upgrades or business system changes, converting heterogeneous terms in multi-source datasets into standardized columnar data. Ensuring data format compatibility with the distributed time series database architecture provides the prerequisite for subsequent storage and retrieval optimization.

[0084] Hyper-converged server clusters are distributed architectures. Node time errors can cause data time sequence distortion. Spatial topology changes require real-time awareness to support spatiotemporal correlation analysis. High-precision clock servers are deployed to record the acquisition timestamps of distributed nodes. In conjunction with software-defined networking, network topology information is acquired in real time. The container orchestration platform is linked to monitor the spatial coordinates and logical groupings of distributed nodes, dynamically generating spatial topology labels for multi-source datasets. Spatiotemporal encoding is used to integrate acquisition timestamps and spatial coordinates into a composite index key, enabling precise spatiotemporal correlation.

[0085] Hyper-converged server clusters generate a large amount of monitoring data daily. Random writes can cause database performance degradation, and a balance must be struck between data reliability and storage costs. For connected multi-source datasets, LSM tree index optimization and columnar storage layout are used in distributed time-series databases to improve the writing and query performance of massive data. LSM tree indexes convert random writes into sequential writes, improving the writing performance of massive data. Columnar storage is compressed by resource type, encoding and compressing numerical data, and adopting a hybrid redundancy strategy of multiple copies and erasure codes. Specifically, by configuring a dynamic update cycle, copies of multi-source dataset data within the update cycle are retained to ensure high availability, and erasure codes are applied to historical data of multi-source datasets outside the update cycle to reduce storage costs. The dynamic update cycle is determined by the frequency of data access, business real-time requirements, and data lifecycle strategies. A distributed consensus algorithm is used to ensure consistency during the writing process of multi-source dataset data, while supporting data redundancy deployment across availability zones.

[0086] There are significant differences in data access patterns, and a unified storage strategy can lead to resource waste. Based on online learning algorithms, data access patterns are continuously analyzed, and hot and cold classification of multi-source data sets is performed according to the time dimension, access frequency, and business priority of multi-source data sets. For hot data, such as resource monitoring data in the past 7 days, small shards are stored and replicas are configured to improve real-time query efficiency; for cold data, such as log data more than 30 days ago, large shards are merged and the number of replicas is reduced, and shard merging and replica migration are performed through an asynchronous scheduling mechanism to optimize storage resource utilization.

[0087] Upper-layer applications need to support complex queries, and a single index cannot meet the cross-dimension retrieval requirements. Based on time series, spatial topology labels and resource type identifiers are integrated to build a multi-dimensional index structure containing time, space, and resource type. The time dimension uses B+ tree index to implement millisecond-level range query, the space dimension uses R tree index to support topology proximity search, and the resource type dimension builds inverted index to accelerate classification and aggregation query. Through index automatic optimization mechanism, index weight and shard strategy are dynamically adjusted according to actual query mode to improve cross-dimension composite query efficiency.

[0088] Data provenance is the basis for quality control and compliance audit, sensitive data must meet privacy protection requirements, and operation behavior must be traceable. Record the whole link processing process from the original multi-source data set to the standard data set, generate data blood relationship metadata, and build metadata index to assist upper-layer applications to quickly locate data assets. Based on data classification standards, sensitive data is implemented at the field level for desensitization and encryption storage; role-based access control combined with attribute permission management is adopted to dynamically control data access permissions according to user identity, operation scenario, etc. Deploy a full-link audit mechanism to record data access, storage, query, and other operation behaviors to ensure that data usage meets privacy protection and compliance requirements. Form a standard data set with the following characteristics: semantic level unified terminology and relationship modeling through ontology model, eliminating cross-source data ambiguity; standardized columnar storage is adopted to support efficient compression and shard processing; precise timestamp and dynamic spatial topology label are integrated in time and space dimensions to meet the requirements of spatio-temporal correlation analysis; multi-dimensional index and storage strategy optimization are implemented to realize millisecond-level time range query and cross-resource type association retrieval; data quality is guaranteed through full-process verification, and security mechanisms ensure controllable data access, providing standardized and highly available data support for intelligent analysis, decision-making, and scheduling upper-layer applications.

[0089] The association decision module is used for constructing a graph model according to a structured standard data set, mining a causal relationship between resources by using a graph neural network algorithm in combination with time series and spatial topological information, generating a resource scheduling strategy, optimizing and querying a statement of an index of a graph database to improve a complex query response speed, constructing a multi-objective optimization function, generating a resource scheduling strategy in combination with a reinforcement learning scheduling algorithm, and predicting and evaluating the resource scheduling strategy by using a simulation technology to construct a cluster simulation model.

[0090] In the embodiment, the association decision module constructs a graph model containing resource nodes and association edges, mines an implicit causal relationship chain between resources by using a graph neural network algorithm to fuse time series and spatial topological information, and significantly improves a complex query response speed by implementing hybrid index optimization and query statement reconstruction on a graph database. The association decision module constructs a multi-objective optimization function, generates a resource scheduling strategy in combination with a reinforcement learning algorithm, and effectively improves resource utilization efficiency by predicting and evaluating a plurality of load scenarios by using a cluster simulation model. The association decision module detects and eliminates strategy conflicts by using a rule engine, reuses historical strategies by using case reasoning, realizes intelligent matching of the strategy, and finally generates a scheduling strategy that takes into account resource efficiency and business reliability, thereby supporting system dynamic optimization.

[0091] Please refer to Figure 3 , preferably, the specific steps of constructing the graph model include:

[0092] The standard data set covers massive data of a full global resource of a hyper-converged server, and direct retrieval efficiency is extremely low. Resource scheduling and analysis usually focus on data of a specific time period, region, or type, and need to be accurately screened. Based on a multi-dimensional index structure of the standard data set, resource state data in a target time period is screened by time indexing, target region resource topological information is obtained by spatial indexing, and target resource type data is extracted by resource type indexing, so as to locate target data of the graph model;

[0093] The hyper-converged server has various types of resources and different attributes, and physical resources and virtual resources need to be abstracted as graph nodes to uniformly model resource states and characteristics, so as to facilitate subsequent association analysis. Graph nodes are constructed according to resource entities in target data of the graph model, and a unique code of the resource entity is used as an identifier of the graph node. A collection time stamp and a spatial topological label are reused to give the graph node time and space attributes, and real-time performance index data is obtained as a state attribute, so as to complete node construction;

[0094] In the hyper-converged server, resources not only exist physical connection, but also exist business logic dependency. Considering only physical connection cannot fully reflect the essence of resource interaction, and logical dependency relationship needs to be excavated at the same time. According to the spatial topology label in the standard data set, the physical connection edges of the graph nodes are generated, and the edge attributes such as bandwidth and delay are labeled. According to the timestamp sequence of the business interaction record of the graph node, the time sequence graph between the graph nodes is constructed, such as the request initiation time and the response completion time, so as to calculate the logical dependency strength between the graph nodes by Granger causality test, and the logical dependency edges are generated for the graph node pairs with the logical dependency strength greater than the preset logical connection threshold, so as to determine the resource connection relationship. The physical connection edge reflects the physical connectivity between devices, supports fault location, and the logical dependency edge reflects the implicit dependency in business operation, supports performance bottleneck analysis.

[0095] The influence degree of different connection edges on resource scheduling and system operation is different, and the importance needs to be quantified to provide basis for priority decision. For the connection edges of the graph nodes, including physical connection edges and logical dependency edges, according to the graph model target data of the standard data set, the request times and traffic data of the physical connection edges and the logical dependency edges are extracted, and the edge weights of the physical connection edges and the logical dependency edges are respectively given according to the weighted priority mapping algorithm; the weighted priority mapping algorithm constructs a weight calculation factor library, sets different weight calculation factors for the different characteristics of the physical connection edges and the logical dependency edges, and calculates the edge weights by using the weighted summation formula.

[0096] The hyper-converged server is large in scale, and the graph model data volume grows rapidly over time, so single machine storage cannot meet the performance and scalability requirements, and needs to be supported by distributed architecture. A hierarchical partitioning strategy is adopted to construct a distributed graph database storage architecture, the graph nodes and connection edges in the graph model are stored in two levels of shards according to resource type and spatial location, and the complete coverage of the global resource data of the hyper-converged server in the target time period is realized. A multi-dimensional composite index system is established, taking resource ID as the primary key, combining spatial location coding and resource type label to construct a three-level index structure, and through prefix tree algorithm to optimize the index level relationship, to ensure that the node retrieval and relationship query performance can still be achieved in milliseconds under the scale of millions of nodes.

[0097] The resource state and configuration of the hyper-converged server change dynamically, and the graph model needs to be updated in real time to reflect the latest situation and avoid invalid analysis results. The change events of the standard data set are captured in real time, including resource state change, resource configuration adjustment or business rule change, when the change event of the standard data set is captured, the graph database update process is triggered, according to the change event type and resource ID, through the multi-dimensional composite index system of the distributed graph database, the affected target shard area of the graph model is quickly located;

[0098] Resource state changes or resource addition and deletion operations occur frequently, and the graph node attributes need to be synchronized to ensure the accuracy of the model. For the target fragment area located, the graph node attributes of the involved graph model are updated. Specifically, if the change event is a resource state change, such as a sudden change in resource load, the system directly extracts the latest real-time performance indicator data from the updated standard data set to replace the corresponding state attribute of the graph node; if it is a resource addition or deletion event, the corresponding graph node is added or deleted according to the resource entity information in the standard data set, and the associated edge relationship is updated synchronously.

[0099] Network topology reconstruction and business rule changes will cause resource connection relationship changes, and the connection edge needs to be updated synchronously to accurately depict resource interaction. The connection edge of the graph node is dynamically adjusted. If the physical connection changes due to network topology reconstruction, the physical connection edge is regenerated or deleted according to the updated spatial topology label in the standard data set, and the edge attributes such as bandwidth and delay are updated; if the logical dependency relationship changes due to business rule changes, the system reconstructs the time series graph based on the updated business interaction record timestamp sequence, recalculates the logical dependency strength through Granger causality test, and updates the weight or relationship of the logical dependency edge whose strength changes by more than the threshold.

[0100] The data of the connection edge such as traffic and request times changes in real time, and the original weight cannot reflect the latest importance, which needs to be recalculated to adapt to the scheduling demand. The edge weight of the connection edge of the graph node is recalculated, the updated request times and traffic data are extracted from the standard data set, and the weighted priority mapping algorithm is called to calculate the differential weight factor of the physical connection edge and the logical dependency edge, and each edge is assigned a weight.

[0101] Data updates in a distributed environment may be inconsistent due to network failures or node abnormalities. After the graph database update process is completed, the integrity of the updated graph model target fragment area is checked through the transaction consistency mechanism of the distributed graph database, and the update result is synchronized to the remaining fragment areas of the graph database to ensure the consistency and accuracy of the entire graph model data.

[0102] Preferably, the specific steps of mining the causal relationship between resources include:

[0103] The causal relationship between resources in the hyper-converged server is complex and implicit, and traditional methods are difficult to capture nonlinear and time-dependent dependencies. Deep learning is needed to mine potential relationships, and graph neural network algorithm is used to extract features from the graph model in the distributed graph database. The time series attributes and spatial topology attributes of the graph node are used as input, and the sliding window mechanism is used to construct the time series graph structure of resource state transition. The structural causal model in the causal graph neural network quantifies the causal influence direction between graph nodes, generates directed causal edges containing time lag relationships, and realizes the explicit extraction of implicit causal relationship chains.

[0104] Causal edges extracted solely through algorithms may be misjudged and require dual verification using both statistics and business logic to ensure the reliability and practicality of causal relationships. The extracted directed causal edges are statistically significant based on the Granger causality test algorithm. By constructing a vector autoregression model, lagged variables are introduced into the model, and regression analysis is performed on time series data to calculate the causal strength of the directed causal edges. Valid causal chains are then screened by setting a causal strength threshold. For example, for the temporal relationship between a sudden increase in storage IOPS and a database transaction delay, the causal probability value within the lag order is calculated. If it exceeds a preset threshold, it is considered a valid causal relationship. At the same time, valid causal chains are verified in conjunction with the business rule library to eliminate pseudo-causal relationships that contradict actual business logic and ensure the business rationality of the causal chain.

[0105] Hyper-converged server resources change dynamically, and static causal relationships cannot reflect real-time status, requiring real-time updates to prevent analysis failures. When a graph database triggers an update, the causal relationship incremental update process is automatically initiated. The affected local graph model is partitioned, and the temporal and spatial topological properties of the graph nodes in the local graph model are re-extracted. Directed causal edges are then updated using an incremental graph neural network. For example, when a new server is added to the storage cluster, the causal relationship between the server's I / O interactions with the existing storage nodes is recalculated, and the causal chain topology is dynamically corrected to ensure that the causal mining results are consistent with the real-time resource status.

[0106] Massive causal chains require structured storage to support fast queries. Furthermore, causal relationships have time-series characteristics, making single-store storage insufficient for multi-dimensional analysis. Valid causal chains are converted into causal records according to a standardized structure and stored in a causal relationship database. Each causal record contains metadata such as the causal subject, causal object, causal type, causal strength, and storage timestamp.

[0107] Long-term accumulation of causal records consumes a large amount of storage resources, and similar causal chains are stored repeatedly, necessitating optimization of storage efficiency and refinement of general models. An aging mechanism is used to delete causal records from the causal relationship library that have been unused for an extended period of time, freeing up storage resources. A clustering algorithm is then used to merge causal records based on similarity, optimizing the causal relationship library. For example, similar causal chains related to fully loaded CPUs in different racks and response delays in the same storage array can be combined into a generalized model.

[0108] See also Figure 4 Preferably, the specific steps of generating a resource scheduling strategy include:

[0109] The importance and activity level of resources in hyper-converged servers differ significantly. A unified monitoring strategy can lead to insufficient monitoring of critical resources or waste of computing resources for non-critical resources. Based on the historical activity, business criticality and data interaction frequency of each graph node in the graph database, the weight coefficients of the feature dimensions are configured through multi-dimensional feature fusion, and the node weight in the range of 0 to 1 is generated by weighted summation, providing a quantitative basis for subsequent hierarchical monitoring strategies. Historical activity is used to represent the resource usage frequency of the node within a specified time period, which is obtained by statistical analysis of the timestamp sequence of node state change records in the graph database. The number of activities per unit time is calculated by the sliding window algorithm. Business criticality is used to identify the importance of the graph node in the core business link, which is determined by the configured graph node priority label and service impact range. Data interaction frequency is used to quantify the data exchange intensity between the graph node and other graph nodes, which is obtained by statistical analysis of the traffic logs and request-response records in the standard data set. The frequency index is generated by analyzing the interaction frequency and data volume of physical connection edges and logical dependency edges.

[0110] Different weights of resources have different impacts on the system, and different monitoring frequencies need to be matched to balance monitoring accuracy and resource consumption. Configure weight classification threshold values, including weight classification upper threshold and weight classification lower threshold, to implement differentiated monitoring according to monitoring weights, including node weights and edge weights.

[0111] When the monitoring weight is greater than the weight classification upper threshold, it means that the graph node or connection edge belongs to a critical resource in the core business link, or it has recently shown high-frequency resource interaction and high activity. Therefore, high-frequency anomaly monitoring is adopted. On the basis of the pre-set basic monitoring frequency, the collection period is shortened according to the configured shortening ratio, and the performance indicators of the graph node and the connection edge are collected at a high frequency. Performance indicators include CPU load, memory usage, disk IOPS, bandwidth utilization, and transmission delay. The shortening ratio is used to dynamically adjust the monitoring density of critical resources, balancing resource risk and system overhead. By constructing the utility function of monitoring benefit and overhead, taking performance indicator fluctuation amplitude, business impact degree and resource importance as independent variables, and using gradient descent method to solve the optimal shortening ratio.

[0112] When the monitoring weight is less than the weight classification lower threshold, it means that the graph node or connection edge belongs to a non-core business resource, or it has been in a low activity state for a long time and has a low data interaction frequency. Therefore, low-frequency anomaly monitoring is adopted. On the basis of the pre-set basic monitoring frequency, the collection period is increased according to the configured growth ratio, and the performance indicators of the graph node and the connection edge are collected at a low frequency. The growth ratio is used to reduce the monitoring frequency of non-critical resources and reduce resource consumption. By calculating the change rate of performance indicators based on historical data and combining with the resource importance coefficient to construct a linear regression model, the collection period growth amplitude is predicted.

[0113] When the monitoring weight is less than or equal to the upper threshold of the weight classification and greater than or equal to the lower threshold of the weight classification, it means that the graph node or connection edge belongs to a regular business resource, and the activity and data interaction frequency are within the normal fluctuation range. In this case, abnormal monitoring is performed according to the basic monitoring frequency, and the performance indicators of the graph nodes and connection edges are collected.

[0114] The normal fluctuation range of resource performance indicators is affected by factors such as business load and time cycle. Static thresholds cannot adapt to dynamic changes and are prone to false positives or omissions. Based on the performance indicator data collected from the graph model, the historical performance indicator data of the graph model is obtained by configuring a time series data extractor. Based on the historical performance indicator data of the graph model, the performance indicator data of the graph nodes and connection edges are statistically analyzed to generate a basic threshold range containing the mean and standard deviation. Based on the frequency of potential anomalies of the graph nodes and connection edges in the graph model, the basic threshold range of the graph nodes and connection edges is adjusted through an adaptive threshold regulator to differentiate the thresholds. The basic threshold range is scaled to construct a dynamic threshold range for the performance indicator data of the graph nodes and connection edges, forming a threshold system that dynamically adapts to the graph structure.

[0115] Resource anomalies may manifest as performance indicators exceeding thresholds or changing dramatically within a short period of time. A single judgment method cannot fully identify anomalies. The performance indicator data collected in real time by the graph model is compared with the corresponding dynamic threshold range. If it is within the corresponding dynamic threshold range, time series mutation monitoring is performed. The performance indicator data is monitored for change amplitude through a differential filter. If the change amplitude of the performance indicator data exceeds the preset amplitude threshold, a resource anomaly warning is triggered.

[0116] Potential anomalies may be accidental fluctuations or real anomalies, and need to be treated differently to avoid invalid analysis. At the same time, clarifying the anomaly propagation path will help locate the root cause and formulate effective strategies. If the performance indicator data collected in real time is outside the corresponding dynamic threshold range, the collection point will be marked as a potential anomaly point, and the potential anomaly point will be identified as an outlier. If the potential anomaly point is determined to be an outlier, it indicates that the anomaly may be caused by temporary noise or accidental fluctuations and has not yet formed a systemic impact. In this case, time series mutation monitoring will be performed to determine whether to trigger a resource anomaly warning. If the potential anomaly point is determined to be a non-outlier, it indicates that the anomaly may be an early sign of a real fault or has formed a lasting impact. In this case, resource scheduling path analysis will be triggered, and the graph node where the potential anomaly point is located will be marked as an abnormal graph node, and the connection edge where the potential anomaly point is located will be marked as an abnormal connection edge.

[0117] The causal relationship library contains a large number of causal chains, and the part related to the current exception needs to be screened out to narrow the analysis range and improve the efficiency of strategy generation. When the resource scheduling path analysis is triggered, a double-layer screening mechanism is started in the causal relationship library based on the timestamp of the potential abnormal point, the effective causal chain in the retrieval time period is extracted through the time window index, and the target causal chain is screened out based on the target strength threshold set based on the causal strength, and a target causal chain set is constructed;

[0118] The screened target causal chain needs to be mapped to the actual graph model to clearly define the affected resource range and provide specific objects for formulating the scheduling strategy. The causal subject and causal object of the target causal chain are extracted, and through the resource ID of the graph database, they are mapped to the specific graph node in the current graph model. If the causal subject and causal object do not exist in the graph model or are in a failed state, the target causal chain is removed. According to the target causal chain set after removal, the root cause causal subject is traced back, and the graph model to be performed resource scheduling in the graph database is identified as the target graph model;

[0119] Resource scheduling needs to consider resource utilization, business continuity, cost and other factors, and a single strategy cannot meet complex needs. For the graph nodes of the root cause causal subject in the target graph model and the graph nodes associated through the target causal chain, a multi-objective optimization function is constructed through a multi-objective optimization engine, and a multi-dimensional resource load balancing algorithm is used to dynamically plan the resource offloading path of the root cause node and the associated graph node, realize the gradient descent of the resource utilization rate, and minimize the resource utilization rate of the root cause causal subject and the graph nodes associated through the target causal chain; through the business link priority graph, the interruption time and recovery cost of each business link are calculated according to the resource scheduling path, the continuity of high-priority businesses is prioritized, and the loss of resource scheduling business links is minimized to minimize the business interruption loss caused by abnormal propagation; through historical scheduling case similarity matching, combined with the real-time bandwidth of the physical connection edge and the node topology position, the execution overhead of the migration path and the expansion deployment is pre-evaluated, the most cost-effective scheduling scheme is selected, and the execution cost of the resource scheduling operation is minimized; considering the constraints of resource utilization, business link loss, scheduling cost and other constraints, a candidate strategy set is generated, including resource migration, load balancing, link switching and other operations. For example, if the root cause is storage device overload, the generated strategy may include migrating part of the IO-intensive business to other storage nodes and adjusting the storage cache strategy to improve performance.

[0120] Before actual implementation, candidate strategies must be evaluated for feasibility and effectiveness to avoid negative impacts on the system. Candidate strategies are verified using a cluster simulation model, evaluating metrics such as anomaly repair time and resource utilization improvement. The optimal candidate strategy is selected based on overall performance. During the simulation, the graph model dynamically updates data, such as newly added node topology and link bandwidth changes, to ensure that the verification results are consistent with the current system state.

[0121] The target candidate strategy is decomposed into an execution instruction chain and sent to the scheduling execution module. At the same time, the causal strength of the target causal chain in the causal relationship library is updated according to the execution results, the causal strength of the effective causal chain is strengthened, and the invalid association is weakened.

[0122] The scheduling execution module is used to send the received execution instruction chain to the target node through the resource orchestration engine, execute resource scheduling, monitor feedback and adjust the resource scheduling, monitor the execution status and effect in real time, and automatically trigger parameter adjustment or policy rollback when the deviation between the actual indicator and the expected indicator exceeds the threshold. It also continues to track and monitor after the adjustment, and feeds back the execution results to the graph database and causal relationship library to form a scheduling closed loop.

[0123] In this embodiment, the scheduling execution module uses the resource orchestration engine to break down the execution instruction chain into a sequence of executable operations, which are then sent to the target node. Simultaneously, probes are deployed at each node to collect resource status in real time and synchronize it with the graph database to update node attributes. When deviations between the actual execution results and expectations are detected, parameter adjustments or policy rollbacks are automatically triggered. After the adjustments are made, continuous tracking and monitoring are maintained, and node dynamic weights and anomaly thresholds are regularly recalculated. This mechanism effectively improves the success rate of scheduling policy execution, shortens anomaly response time, optimizes resource utilization, and continuously updates the causal relationship library through execution results, ensuring intelligent and efficient resource scheduling for hyperconverged server clusters.

[0124] Preferably, the specific steps of executing resource scheduling and performing monitoring, feedback and adjustment on resource scheduling include:

[0125] Target candidate policies are typically abstract logic that needs to be converted into underlying executable operations. The heterogeneous resources of a hyperconverged server must be processed step by step based on type and dependency relationships. The resource orchestration engine parses the target candidate policy's execution instruction chain into a sequence of executable atomic operations. Instruction chains are generated based on operation type and target node. For example, a storage overload policy can be broken down into subtasks such as source node data migration and target node resource pre-allocation, sorted by dependency relationships. A workflow engine maps the execution instruction chain into a call sequence, which is then dispatched to the target node's execution agent via a message queue to ensure operational order and transaction consistency.

[0126] Frequent scheduling of the same node in a short period of time can easily cause resource shock, and nodes of different importance need to be monitored differently. The calling sequence is executed on the target node in turn, and the scheduling quiet period of the target node is set to avoid frequent scheduling of the same node in a short period of time and trigger system shock. The scheduling observation period is configured according to the monitoring weight of the target node. In the scheduling observation period of resource scheduling, high-frequency abnormal monitoring is performed on the target node, high-frequency collection of performance indicators is performed, potential abnormal points are identified, the frequency of occurrence of potential abnormal points in the scheduling observation period is counted, the scheduling abnormal threshold is configured, and if the frequency of occurrence of potential abnormal points is greater than the scheduling abnormal threshold, the scheduling abnormality warning is triggered, the resource scheduling path is blocked, the calling sequence execution is suspended through the message queue, the rollback operation is performed, and the scheduling quiet period of the target node is released.

[0127] The pseudo-causal chain that triggers the scheduling abnormality warning is removed, the target causal chain of the causal relationship library is re-identified, the target candidate strategy is updated, the calling sequence is mapped, the calling sequence is executed on the target node again, the frequency of occurrence of potential abnormal points in the scheduling observation period is counted, and if the frequency of occurrence of potential abnormal points is less than or equal to the scheduling abnormal threshold, it indicates that the resource scheduling strategy effectively alleviates the abnormal state of the target node and does not trigger new resource fluctuations, the abnormality is confirmed to be resolved, the resource scheduling observation is stopped, and the abnormal state label and the scheduling timestamp of the node in the graph database are updated synchronously; the number of times of updating the target candidate strategy is counted, and if it is greater than a preset update threshold, a scheduling update warning is given, which is used to prompt that the current resource scheduling is in a frequent adjustment state, and there may be problems such as unreasonable system configuration, deviation of causal relationship identification, or continuous effect of external interference factors, and manual intervention is required to troubleshoot and optimize the scheduling strategy or adjust the system parameters to avoid resource scheduling shock and system performance degradation.

[0128] The causal strength of the effective causal chain is updated in batches through the graph database transaction operation to strengthen its effectiveness, the graph nodes and connection edges in the resource scheduling path involved in the calling sequence are located, and the causal relationship data and the graph database are updated, and the monitoring weight is recalculated.

[0129] The complete execution process log of the calling sequence is recorded, including the instruction issuing time, the operation sequence, the index change curve, etc., and a composite index is established according to the timestamp and the graph node ID. An audit trail is generated for sensitive operations, and a distributed log system is used to store the log to ensure that the index supports second-level retrieval and meets the safety audit requirements.

[0130] Embodiment 2:

[0131] See Figure 5 The embodiment introduces a multi-resource scheduling method of a hyper-converged server, including the following steps:

[0132] Step S1: Collect multi-source datasets through distributed collection nodes, perform semantic alignment and format conversion on the multi-source datasets, and then use a distributed time series database architecture to store them and build a standard dataset with a multi-dimensional index structure.

[0133] Step S2: Locate the target data of the graph model based on the multidimensional index structure of the standard data set, build the graph model, use hierarchical partitioning to build a distributed graph database, and trigger the update of the graph database by capturing change events;

[0134] Step S3: The graph neural network algorithm is used to extract the temporal attributes and spatial topological attributes of the graph nodes. A temporal graph is constructed by combining the sliding window. The structural causal model is used to generate directed causal edges, and valid causal chains are selected and stored in the causal relationship library. When the graph database is updated, the causal relationship incremental update process is started to update the valid causal chains in the causal relationship library by dividing the local graph model.

[0135] Step S4: Based on the monitoring weights of the graph nodes and connecting edges in the graph model, the collection period is adjusted on the basis of the basic monitoring frequency. Differentiated monitoring is implemented through performance indicator data to identify potential anomalies. The target causal set is constructed through the causal relationship library. The multi-objective optimization function is constructed through the multi-objective optimization engine. The target candidate strategy for resource scheduling is generated and sent to the target node for resource scheduling.

[0136] Step S5: Monitor the execution results in real time, perform monitoring feedback adjustments, and update the graph database and causal relationship library.

[0137] Specifically, the steps to identify potential anomalies include:

[0138] Based on the historical performance indicator data of the graph model, statistical analysis is performed on the performance indicator data of the graph nodes and connected edges to generate a basic threshold range including the mean and standard deviation;

[0139] Based on the frequency of potential anomalies in the graph nodes and edges in the graph model, differentiated threshold adjustments are made to the basic threshold ranges of the graph nodes and edges to construct dynamic threshold ranges for the performance indicator data of the graph nodes and edges.

[0140] Compare the performance indicator data collected in real time by the graph model with the corresponding dynamic threshold range. If it is within the corresponding dynamic threshold range, perform time series mutation monitoring. Based on the change range of the performance indicator data, determine whether to trigger a resource anomaly warning;

[0141] If the performance indicator data collected in real time is outside the corresponding dynamic threshold range, the collection point is marked as a potential anomaly point, and outlier identification is performed on the potential anomaly point. If the potential anomaly point is determined to be an outlier, time series mutation monitoring is performed to determine whether a resource anomaly warning is triggered;

[0142] If the potential abnormal point is determined to be a non-outlier point, a resource scheduling path analysis is triggered, the graph node where the potential abnormal point is located is marked as an abnormal graph node, and the connection edge where the potential abnormal point is located is marked as an abnormal connection edge.

[0143] Specifically, the specific steps of generating the target candidate strategy of the resource scheduling include:

[0144] When the resource scheduling path analysis is triggered, a double-layer screening mechanism is started in the causal relationship library based on the timestamp at which the potential abnormal point occurs, target causal chains are screened, and a target causal chain set is constructed;

[0145] The causal subject and the causal object of the target causal chain are extracted, and are mapped to the graph node in the current graph model through the resource ID of the graph database, the root causal subject is traced back, and the graph model to be subjected to resource scheduling in the graph database is identified as a target graph model;

[0146] For the graph node of the root causal subject and the graph node associated through the target causal chain in the target graph model, a multi-objective optimization function is constructed through a multi-objective optimization engine to generate a candidate strategy set; the multi-objective optimization function includes minimizing the resource usage rate of the graph node of the root causal subject and the graph node associated through the target causal chain, minimizing the resource scheduling business link loss, and minimizing the execution cost of the resource scheduling operation;

[0147] The candidate strategy is verified through a cluster simulation model, a target candidate strategy is selected, and is decomposed into an execution instruction chain and is issued to a scheduling execution module, and the causal strength of the target causal chain in the causal relationship library is updated according to the execution result.

[0148] Specifically, the specific steps of executing the resource scheduling and monitoring the execution result for monitoring feedback adjustment include:

[0149] The execution instruction chain of the target candidate strategy is parsed into an atomic operation sequence through a resource orchestration engine, an instruction chain is generated according to the operation type and the target node, and is mapped into a call sequence, and the call sequence is issued to the target node through a message queue;

[0150] The call sequence is executed on the target node in sequence, and a scheduling quiet period of the target node is set, a scheduling observation period is configured according to the monitoring weight of the target node, high-frequency abnormal monitoring is performed on the target node in the scheduling observation period of the execution resource scheduling, and a potential abnormal point is identified;

[0151] The frequency of the occurrence of the potential abnormal point in the scheduling observation period is counted, a scheduling abnormal threshold is configured, it is judged whether a scheduling abnormality warning is triggered, if yes, the resource scheduling path is blocked, the call sequence execution is suspended through the message queue, a rollback operation is performed, and the scheduling quiet period of the target node is removed;

[0152] The false causal chain triggering the abnormal scheduling warning is pruned, the target causal chain of the causal relationship library is re-identified, the target candidate strategy is updated, the call sequence is mapped, and the call sequence is executed on the target node again;

[0153] If the frequency of the potential abnormal point is less than or equal to the scheduling exception threshold, it is confirmed that the exception is removed, the resource scheduling observation is stopped, and the abnormal state label and scheduling timestamp of the node in the graph database are synchronously updated;

[0154] The graph nodes and connection edges in the resource scheduling path involved in the call sequence are located, the causal relationship data volume and the graph database are updated, and the call sequence execution whole-process log is completely recorded.

[0155] Working principle and effects:

[0156] The application integrates and elastically stores heterogeneous multi-source data through a multi-dimensional index distributed time sequence database, solves the problems of data format heterogeneity and semantic ambiguity, forms a unified resource data basement, and lays a data foundation for accurate scheduling; with the aid of a hierarchical partitioned dynamic graph database, resource topology evolution and time sequence dependency are captured in real time, physical connection and logical interaction are converted into a calculable graph model, dynamic modeling of resource association relationship is realized, and scheduling strategies can accurately match the actual state of resources; graph neural networks and structural causal models are used to deeply mine time sequence causal chains, a causal relationship library is constructed through statistical significance and business rule dual verification, abnormal root causes are accurately located, and resource waste caused by misjudgment of scheduling strategies is avoided; based on a multi-objective optimization engine, a global optimization strategy is generated by combining cluster simulation, resource utilization, business continuity and scheduling cost, the limitations of traditional single target scheduling are broken through, and global optimization of resource scheduling is realized; abnormal types are accurately distinguished through dynamic threshold adjustment and outlier identification mechanism, resource allocation is continuously optimized in combination with closed-loop regulation, scheduling strategies dynamically evolve with business scenarios, and finally, efficient normalization of heterogeneous data is realized at the resource integration level, the accuracy, global optimization ability and adaptive evolution ability of the strategy are significantly improved at the scheduling level, and the collaborative efficiency and system reliability of multi-resource management of the super-converged server are greatly enhanced.

[0157] The above only describes the preferred embodiments of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solutions falling within the concept of the present application shall fall within the protection scope of the present application. It should be noted that, for ordinary skilled persons in the art, some improvements and refinements without departing from the principles of the present application shall also be considered as the protection scope of the present application.

Claims

1. Hyper-converged server multi-resource integration system, characterized by: include: Collect multi-source data sets through distributed collection nodes, perform semantic alignment and format conversion on the multi-source data sets, and then use a distributed time series database architecture to store them and build a standard data set with a multi-dimensional index structure; Locating the target data of the graph model based on the multidimensional index structure of the standard data set, building a graph model, building a distributed graph database using hierarchical partitioning, and triggering the update of the graph database by capturing change events; The graph neural network algorithm is used to extract the temporal attributes and spatial topological attributes of the graph nodes, and a sliding window is used to construct a temporal graph. The structural causal model is used to generate directed causal edges, and valid causal chains are screened and stored in the causal relationship library. When the graph database is updated, the causal relationship incremental update process is started, and the valid causal chains in the causal relationship library are updated by dividing the local graph model. Based on the monitoring weights of the graph nodes and connecting edges in the graph model, the collection period is adjusted on the basis of the basic monitoring frequency, differentiated monitoring is implemented through performance indicator data, potential anomalies are identified, a target causal set is constructed through the causal relationship library, a multi-objective optimization function is constructed through the multi-objective optimization engine, and a target candidate strategy for resource scheduling is generated and sent to the target node for resource scheduling. Monitor execution results in real time, conduct feedback adjustments, and update the graph database and causal relationship library; The specific steps of identifying potential abnormal points include: Based on the historical performance indicator data of the graph model, statistical analysis is performed on the performance indicator data of the graph nodes and connected edges to generate a basic threshold range including the mean and standard deviation; Based on the frequency of potential anomalies in the graph nodes and edges in the graph model, differentiated threshold adjustments are made to the basic threshold ranges of the graph nodes and edges to construct dynamic threshold ranges for the performance indicator data of the graph nodes and edges. Compare the performance indicator data collected in real time by the graph model with the corresponding dynamic threshold range. If it is within the corresponding dynamic threshold range, perform time series mutation monitoring. Based on the change range of the performance indicator data, determine whether to trigger a resource anomaly warning; If the performance indicator data collected in real time is outside the corresponding dynamic threshold range, the collection point is marked as a potential anomaly point, and outlier identification is performed on the potential anomaly point. If the potential anomaly point is determined to be an outlier, time series mutation monitoring is performed to determine whether a resource anomaly warning is triggered; If a potential outlier is determined to be a non-outlier, resource scheduling path analysis is triggered, and the graph node where the potential outlier is located is marked as an abnormal graph node, and the connection edge where the potential outlier is located is marked as an abnormal connection edge.

2. The hyper-converged server multi-resource integration system according to claim 1, wherein: The specific steps of generating a target candidate strategy for resource scheduling include: When resource scheduling path analysis is triggered, a double-layer screening mechanism is initiated in the causal relationship library based on the timestamp of the potential abnormal point, screening the target causal chain and constructing the target causal chain set; Extract the causal subject and causal object of the target causal chain, map them to the graph nodes in the current graph model using the resource ID of the graph database, trace the root causal subject, and identify the graph model to be scheduled in the graph database as the target graph model; For the graph nodes of the root causal subject and the graph nodes associated through the target causal chain in the target graph model, a multi-objective optimization function is constructed through a multi-objective optimization engine to generate a set of candidate strategies; the multi-objective optimization function includes minimizing the resource utilization rate of the graph nodes of the root causal subject and the graph nodes associated through the target causal chain, minimizing the resource scheduling service link loss, and minimizing the execution cost of the resource scheduling operation; The candidate strategies are verified through the cluster simulation model, the target candidate strategy is selected, and it is decomposed into an execution instruction chain and sent to the scheduling execution module. At the same time, the causal strength of the target causal chain in the causal relationship library is updated according to the execution results.

3. The hyper-converged server multi-resource integration system according to claim 2, wherein: The specific steps of executing resource scheduling and real-time monitoring of execution results and performing monitoring feedback adjustment include: The resource orchestration engine parses the execution instruction chain of the target candidate policy into an atomic operation sequence. It generates an instruction chain based on the operation type and target node and maps it into a call sequence. The call sequence is then sent to the target node through a message queue. Execute the call sequence for the target node in sequence and set the target node's scheduling cool-down period. Configure the scheduling observation period according to the monitoring weight of the target node. During the scheduling observation period of executing resource scheduling, use high-frequency anomaly monitoring on the target node to identify potential anomalies. Count the frequency of potential anomalies within the scheduling observation period, configure the scheduling anomaly threshold, and determine whether a scheduling anomaly warning is triggered. If so, block the resource scheduling path, pause the call sequence execution through the message queue, perform a rollback operation, and release the scheduling cool-down period of the target node. Eliminate the pseudo causal chain that triggers the scheduling exception warning, re-identify the target causal chain in the causal relationship library, update the target candidate strategy, map the call sequence, and execute the call sequence on the target node again; If the frequency of potential anomaly points is less than or equal to the scheduling anomaly threshold, the anomaly is confirmed to be resolved, resource scheduling observation is stopped, and the abnormal status label and scheduling timestamp of the node in the graph database are synchronously updated; Locate the graph nodes and connection edges in the resource scheduling path involved in the call sequence, update the causal relationship data volume and graph database, and record the full process log of the call sequence execution.

4. The hyper-converged server multi-resource integration system according to claim 1, wherein: The specific steps of screening the effective causal chain include: The graph neural network algorithm is used to extract features from the graph models in the distributed graph database. The temporal attributes and spatial topological attributes of the graph nodes are used as input to construct a temporal graph structure of resource state transitions. Quantify the causal influence direction between nodes in the time series graph structure and generate directed causal edges containing time lag relationships; Statistical significance verification is performed on the extracted directed causal edges, and valid causal chains are screened by setting a causal strength threshold; When the graph database triggers an update, the causal relationship incremental update process is automatically started. By dividing the affected local graph model in the graph model, the temporal attributes and spatial topological attributes of the graph nodes in the local graph model are re-extracted, and the directed causal edges are updated through the incremental graph neural network. Convert the effective causal chain into a causal record according to the standardized structure and store it in the causal relationship library. The causal record includes the causal subject, causal object, causal type, causal strength and storage timestamp; The aging mechanism is used to delete causal records in the causal relationship library that have not been used for more than a preset time period, and the clustering algorithm is used to merge the causal records based on similarity to optimize the causal relationship library.

5. The hyper-converged server multi-resource integration system according to claim 1, wherein: The specific steps of building a distributed graph database include: Locate the target data of the graph model based on the multi-dimensional index structure of the standard data set; Build graph nodes based on resource entities in the graph model target data, reuse the acquisition timestamp and spatial topology label to assign time and space attributes to the graph nodes respectively, and obtain real-time performance indicator data as status attributes; Generate physical connection edges of graph nodes based on spatial topology labels in standard datasets, and construct a time series graph between graph nodes based on the timestamp sequence of business interaction records of graph nodes. By calculating the logical dependency strength between graph nodes, select graph node pairs and generate logical dependency edges. For the connection edges of graph nodes, based on the graph model target data of the standard dataset, the request count and traffic data of the physical connection edges and logical dependency edges are extracted, and the edge weights of the physical connection edges and logical dependency edges are assigned respectively according to the weighted priority mapping algorithm; A hierarchical partitioning strategy is adopted to build a distributed graph database storage architecture, in which graph nodes and connection edges in the graph model are stored in two-level shards according to resource type and spatial location, and a multi-dimensional composite index system is established.

6. The hyper-converged server multi-resource integration system according to claim 5, wherein: The specific steps of constructing the distributed graph database also include: Capture change events of standard datasets in real time. These change events include resource status changes, resource configuration adjustments, or business rule changes. When a change event of a standard dataset is captured, the graph database update process is triggered: Based on the change event type and resource ID, the affected target shard area of ​​the graph model is located through the multi-dimensional composite index system of the distributed graph database. For the located target shard area, update the graph node attributes of the involved graph model, dynamically adjust the connection edges of the graph nodes, and update the edge weights of the connection edges; After the graph database update process is completed, the integrity of the updated graph model target shard area is checked through the transaction consistency mechanism of the distributed graph database, and the update results are synchronized to the remaining shard areas of the graph database.

7. The hyper-converged server multi-resource integration system according to claim 1, wherein: The specific steps of constructing the standard dataset include: Receive multi-source data sets through distributed message queues, manage sub-queues, initialize distributed time series databases, and establish data docking channels; Build a semantic model based on domain ontology, define standard terms and relationship attributes of resource entities, and form a standardized semantic system; Mapping heterogeneous format terms of the multi-source data set to a unified semantic space through a rule engine, using a plug-in adaptive conversion architecture to dynamically load conversion plug-ins to convert heterogeneous format terms into standardized columnar data; Record the distributed node acquisition timestamps, obtain network topology information, and dynamically generate spatial topology labels for multi-source datasets; LSM tree indexing and columnar storage are used in distributed time series databases. Data from multi-source datasets are classified into hot and cold categories based on their time dimension, access frequency, and business priority to adjust shard storage and the number of replicas. Integrate spatial topology labels and resource type identifiers to build a multidimensional index structure that includes time, space, and resource types, and build a metadata index; desensitize and encrypt sensitive data for storage, and record data operation behaviors.

8. The hyper-converged server multi-resource integration system according to claim 1, wherein: The specific steps of adjusting the collection period based on the basic monitoring frequency and implementing differentiated monitoring through performance indicator data include: Based on the historical activity, business criticality, and data interaction frequency of each graph node in the graph database, multi-dimensional feature fusion is used to configure the weight coefficients of the feature dimensions, and a weighted summation method is used to generate node weights in the range of 0 to 1; Configure weight classification thresholds, including upper and lower thresholds, and implement differentiated monitoring based on monitoring weights, including node weights and edge weights. When the monitoring weight is greater than the upper threshold of the weight classification, high-frequency abnormal monitoring is adopted. On the basis of the preset basic monitoring frequency, the collection cycle is shortened according to the configured shortening ratio; When the monitoring weight is less than the lower threshold of the weight classification, low-frequency abnormal monitoring is adopted. On the basis of the preset basic monitoring frequency, the collection cycle is increased according to the configured growth ratio; When the monitoring weight is less than or equal to the upper threshold of the weight classification, and greater than or equal to the lower threshold of the weight classification, the collection cycle is set according to the basic monitoring frequency to perform abnormal monitoring; Differentiated monitoring of graph nodes or connection edges is performed according to the dynamic collection cycle, and performance indicator data of graph nodes and connection edges is collected.

9. A method for scheduling multiple resources on a hyper-converged server, which is implemented based on the hyper-converged server multiple resource integration system according to any one of claims 1 to 8, characterized in that: The following steps are involved: Step S1: Collect multi-source datasets through distributed collection nodes, perform semantic alignment and format conversion on the multi-source datasets, and then use a distributed time series database architecture to store them and construct a standard dataset with a multi-dimensional index structure; Step S2: Locate the graph model target data based on the multidimensional index structure of the standard data set, build a graph model, construct a distributed graph database using hierarchical partitioning, and trigger the update of the graph database by capturing change events; Step S3: Extract the temporal attributes and spatial topological attributes of the graph nodes through the graph neural network algorithm, build a temporal graph in combination with the sliding window, generate directed causal edges using the structural causal model, and select valid causal chains to store in the causal relationship library; when the graph database is updated, start the causal relationship incremental update process, and update the valid causal chains in the causal relationship library by dividing the local graph model; Step S4: Based on the monitoring weights of the graph nodes and connecting edges in the graph model, the collection period is adjusted on the basis of the basic monitoring frequency, differentiated monitoring is implemented through performance indicator data, potential anomalies are identified, a target causal set is constructed through the causal relationship library, a multi-objective optimization function is constructed through the multi-objective optimization engine, a target candidate strategy for resource scheduling is generated, and the strategy is sent to the target node for resource scheduling. Step S5: Monitor the execution results in real time, perform monitoring feedback adjustments, and update the graph database and causal relationship library.

Citation Information

Patent Citations

  • Monitoring method and system for data governance process

    CN119202545A

  • HBase client main and standby switching method and system based on fault perception

    CN119537484A