Calculation platform management system based on artificial intelligence

Through the artificial intelligence management system, the hardware performance and path status of the distributed storage system are monitored and analyzed in real time, and the resource allocation weight is dynamically adjusted, which solves the resource allocation oscillation caused by hardware performance differences and dynamic attenuation, and improves system stability and hardware life.

CN120469632AActive Publication Date: 2025-08-12QINGDAO HAIKUOTIANGAO INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510518894.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-12
Estimated Expiration
2045-04-24

Smart Images

  • Figure CN120469632A_ABST
    Figure CN120469632A_ABST
Patent Text Reader

Abstract

The invention discloses a computing platform management system based on artificial intelligence, particularly relates to the technical field of distributed storage resource dynamic allocation, and is used for solving the problem of resource allocation oscillation caused by hardware performance difference and dynamic attenuation in the prior art. Acquiring the hardware performance, the load state and the data path of the storage node through a real-time acquisition module; the entropy judgment module calculates a path topology entropy and triggers a reconstruction instruction; the resonance analysis module analyzes the group association strength of hardware performance degradation and marks a resonance node group; the weight generation module fuses the parameters to generate a dynamic allocation weight; the error compensation module optimizes a weight coefficient in combination with historical error data; the migration execution module executes migration according to the weight and limits the load of the resonance node; through cooperative regulation and control of topological entropy monitoring and hardware resonance analysis, local oscillation diffusion is inhibited, unnecessary migration operation is reduced, system stability is improved, and the service life of hardware is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dynamic allocation of distributed storage resources, and more specifically, to a computing platform management system based on artificial intelligence. Background Art

[0002] In distributed storage systems, storage nodes are typically composed of a variety of heterogeneous hardware, such as high-performance solid-state drives (SSDs), large-capacity mechanical hard disks (HDDs), and memory media. These hardware devices have significant differences in performance indicators such as read and write speed, latency, and lifespan. To improve resource utilization, existing technologies generally adopt dynamic allocation strategies to adjust data distribution based on the real-time load status of the nodes (such as remaining storage space and I / O pressure). However, in an environment where hardware performance varies significantly and changes dynamically, traditional resource allocation methods rely on static rules or single-dimensional load indicators, making it difficult to adapt to complex performance fluctuation scenarios.

[0003] In existing technologies, resource allocation strategies are susceptible to interference from the coupling of multiple metrics due to differences in storage node hardware performance and dynamic attenuation characteristics. For example, when a node's hardware performance (such as SSD write lifespan) suddenly changes, traditional dynamic allocation methods are unable to effectively coordinate performance attenuation with real-time load, leading to continuous accumulation of prediction errors and frequent fluctuations in resource allocation results. This fluctuation manifests as unnecessary and repeated migration of data between nodes, which not only reduces system efficiency but also exacerbates hardware wear and tear, affecting the overall stability of the storage service. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a computing platform management system based on artificial intelligence to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] The AI-based computing platform management system includes the following modules:

[0007] Real-time acquisition module: real-time acquisition of hardware performance indicators, real-time load status and data access paths of each storage node in the distributed storage system;

[0008] Entropy determination module: Calculates the topological entropy of the data access path. When the topological entropy exceeds a preset threshold, it is determined that the migration path is in a chaotic state and triggers a path reconstruction instruction.

[0009] Resonance analysis module: Analyzes the nonlinear resonance correlation strength of hardware performance decay rates between different storage nodes based on hardware performance indicators. When the nonlinear resonance correlation strength exceeds a critical threshold, a resonant node group is determined.

[0010] Weight generation module: Based on the nonlinear resonance correlation strength, resonance node group and path reconstruction instructions, it generates node dynamic allocation weights to suppress oscillation through a dynamic feature weighting model;

[0011] Error compensation module: This module extracts prediction error data caused by hardware performance mutations and path confusion in historical resource allocation, and generates error compensation coefficients based on the dynamic node allocation weights.

[0012] Migration execution module: performs data migration operations across storage nodes based on the error compensation coefficient. The data migration operation limits the high-frequency task allocation of the resonant node group and prioritizes the reconstruction of chaotic paths.

[0013] In a preferred embodiment, real-time acquisition of hardware performance indicators, real-time load status, and data access paths of each storage node in the distributed storage system includes:

[0014] The read and write speed, latency, and storage medium life of storage nodes are collected as hardware performance indicators. The storage medium life is quantified by the remaining erase and write times of the solid-state drive and the cumulative operating time of the mechanical hard drive. The real-time load status of the storage node is collected, and the real-time load status includes the remaining storage space and input and output pressure. The connection relationship between nodes in the data access path and the frequency of data migration are collected.

[0015] In a preferred embodiment, the topological structure entropy value of the data access path is calculated. When the topological structure entropy value exceeds a preset threshold, it is determined that the migration path is in a chaotic state and a path reconstruction instruction is triggered, including:

[0016] Based on the adjacency matrix of the connection relationship between nodes and the time series of data migration frequency, the topological structure entropy value is calculated using the Shannon entropy formula. The Shannon entropy formula states that the topological structure entropy value is equal to the weighted sum of the information entropy of the probability distribution of the connection relationship between nodes and the coefficient of variation of the data migration frequency.

[0017] Compare the topological structure entropy value with a preset threshold value, and dynamically adjust the preset threshold value according to the upper limit of the statistical distribution of the historical topological structure entropy value;

[0018] When the topology entropy value exceeds a preset threshold, a migration path reconstruction instruction is generated. The migration path reconstruction instruction includes a list of path identifiers that need to be reconstructed first and corresponding bandwidth occupancy constraints.

[0019] In a preferred embodiment, the coefficient of variation of the data migration frequency is calculated by dividing the standard deviation of the time series of the data migration frequency by the mean.

[0020] In a preferred embodiment, the nonlinear resonance correlation strength of the hardware performance decay rates between different storage nodes is analyzed based on the hardware performance indicators, and a resonant node group is determined when the nonlinear resonance correlation strength exceeds a critical threshold, including:

[0021] Extract the decline rate of the solid-state drive write lifespan and the abnormal increase in the mechanical hard drive seek time of each storage node as the time series data of the hardware performance degradation rate;

[0022] Construct a covariance matrix of the hardware performance degradation rate between different nodes, and use principal component analysis to extract the maximum eigenvalue of the covariance matrix as the nonlinear resonance correlation strength;

[0023] The nonlinear resonance correlation strength is compared with a critical threshold value, which is dynamically set according to the statistical distribution characteristics of the eigenvalues of the covariance matrix;

[0024] When the nonlinear resonance correlation strength exceeds a critical threshold, the nodes whose eigenvector weights in the covariance matrix exceed a preset ratio are marked as a resonance node group.

[0025] In a preferred embodiment, based on the nonlinear resonance correlation strength, the resonance node group and the path reconstruction instruction, a dynamic feature weighted model is used to generate node dynamic allocation weights for suppressing oscillation, including:

[0026] The nonlinear resonance correlation strength is mapped to the suppression coefficient through a normalization function, and the normalization function ensures that the suppression coefficient is positively correlated with the nonlinear resonance correlation strength;

[0027] Generate a path reconstruction factor according to the bandwidth occupancy constraint in the path reconstruction instruction. The path reconstruction factor is the product of the inverse of the topological structure entropy value of the path to be reconstructed first and the remaining available bandwidth.

[0028] The node dynamic allocation weight is generated based on the linear combination of the suppression coefficient, path reconstruction factor and real-time load status.

[0029] In a preferred embodiment, the weight coefficients of the linear combination are optimized by a gradient descent method based on historical resource allocation error data.

[0030] In a preferred embodiment, the prediction error data caused by hardware performance mutation and path confusion in historical resource allocation is extracted and combined with the node dynamic allocation weight to generate an error compensation coefficient, including:

[0031] The number of data migration errors caused by the sudden decrease in the SSD's write lifespan and the number of invalid migration operations caused by path confusion are extracted from the resource allocation log as historical prediction error data.

[0032] Calculate the sliding time window mean of the number of data migration errors and the number of invalid migration operations as the standardized historical prediction error data;

[0033] The normalized historical prediction error data and the node dynamic allocation weight are input into the error compensation function to generate the error compensation coefficient.

[0034] In a preferred embodiment, the error compensation function is a nonlinear mapping relationship based on a hyperbolic tangent function, ensuring that the error compensation coefficient decreases monotonically with the product value of the historical prediction error data and the node dynamic allocation weight.

[0035] In a preferred embodiment, a data migration operation across storage nodes is performed according to the error compensation coefficient, the data migration operation restricts the high-frequency task allocation of the resonant node group and prioritizes the reconstruction of the chaotic path, including:

[0036] Adjust the task allocation priority of each storage node according to the error compensation coefficient. The task allocation priority is negatively correlated with the error compensation coefficient.

[0037] A high-frequency task allocation ratio limit is imposed on the resonant node group. The high-frequency task allocation ratio limit is dynamically set according to the ratio of the node's dynamic allocation weight to the real-time load status;

[0038] According to the path identification list in the migration path reconstruction instruction, priority migration requires priority reconstruction of the data blocks stored in the path. The data block migration process follows the bandwidth occupancy constraint and updates the topology structure entropy value in real time.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. Through the synergistic mechanism of topological entropy monitoring and nonlinear resonance analysis, a dynamic resource allocation system that links hardware performance degradation with path status has been established, significantly improving the stability and resource utilization of distributed storage systems. By quantifying the group correlation strength between path topological disorder and hardware performance degradation in real time, the system can accurately identify the source of oscillation and dynamically adjust weight distribution: Topological entropy monitoring identifies disordered paths based on node connection relationships and migration frequency, triggering reconstruction instructions to optimize data distribution and block local oscillations caused by path disorder. Nonlinear resonance analysis analyzes the synergistic effects of hardware performance degradation through the covariance matrix, limiting the high-frequency load of resonant node groups and suppressing the risk of the spread of hardware degradation hotspots. The two are jointly regulated through a dynamic feature weighting model, forming a two-way negative feedback loop between path order and hardware status, fundamentally reducing unnecessary data migration, reducing hardware loss, and improving task allocation efficiency.

[0041] 2. Adaptive optimization of resource allocation strategies is achieved through a closed-loop design that combines an error compensation mechanism with dynamic priority adjustment. The error compensation coefficient dynamically modifies migration decisions based on historical prediction errors and real-time weights, ensuring that the compensation logic strictly matches the system disturbance frequency. The high-frequency task allocation ratio is dynamically constrained based on the ratio of node weight to load status to prevent overloaded nodes from experiencing performance degradation due to task accumulation. The coordinated control of path-priority migration and bandwidth occupancy constraints ensures the reconstruction of critical paths while avoiding network congestion, thus balancing topology optimization and business continuity. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the structure of the computing platform management system based on artificial intelligence of the present invention. DETAILED DESCRIPTION

[0043] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] Example: Figure 1 The present invention provides a schematic diagram of the structure of the computing platform management system based on artificial intelligence, which includes the following modules:

[0045] Real-time acquisition module: real-time acquisition of hardware performance indicators, real-time load status and data access paths of each storage node in the distributed storage system;

[0046] Entropy value determination module: calculates the topological structure entropy value of the data access path. When the topological structure entropy value exceeds the preset threshold, it is determined that the migration path is in a chaotic state and triggers the path reconstruction instruction;

[0047] Resonance analysis module: Analyzes the nonlinear resonance correlation strength of hardware performance decay rates between different storage nodes based on hardware performance indicators. When the nonlinear resonance correlation strength exceeds a critical threshold, a resonant node group is determined.

[0048] Weight generation module: Based on the nonlinear resonance correlation strength, resonance node group and path reconstruction instructions, it generates node dynamic allocation weights to suppress oscillation through a dynamic feature weighting model;

[0049] Error compensation module: This module extracts prediction error data caused by hardware performance mutations and path confusion in historical resource allocation, and generates error compensation coefficients based on the dynamic node allocation weights.

[0050] Migration execution module: performs data migration operations across storage nodes based on the error compensation coefficient. The data migration operation limits the high-frequency task allocation of the resonant node group and prioritizes the reconstruction of chaotic paths.

[0051] The AI-based computing platform management system aims to achieve intelligent decision-making on resource allocation through multi-dimensional data perception and adaptive optimization of dynamic models. The system relies on real-time collection of hardware performance, path status and load indicators to build a dynamic feature weighting model and error compensation mechanism to simulate the multi-factor trade-offs and experience learning in human decision-making: the dynamic feature weighting model generates the optimal weight allocation strategy for suppressing oscillations by analyzing the correlation between hardware attenuation resonance intensity and path chaos, simulating the comprehensive analysis and judgment of expert experience; the error compensation module iteratively corrects the weight parameters based on historical error data to achieve strategy optimization similar to reinforcement learning, so that the system approaches the global optimal solution during continuous operation; topological entropy monitoring and nonlinear resonance analysis serve as the core perception layer, providing precise input for upper-level intelligent decision-making, and ultimately controlling resource scheduling through closed-loop control of the migration execution module, forming a complete AI regulation chain from data perception to autonomous decision-making.

[0052] The specific implementation method for obtaining the hardware performance indicators, real-time load status and data access path of each storage node in the distributed storage system in real time is as follows. The read and write speeds of the storage nodes are collected through a standard test instruction set. The standard test instruction set includes sequential read and write and random read and write operations. The measurement results are recorded in megabytes per second. The latency is obtained by calculating the average time interval for the storage node to respond to input and output requests. The time interval unit is milliseconds. The remaining number of erase and write times is obtained by parsing the output of the intelligent monitoring tool of the solid-state hard drive. The remaining number of erase and write times indicates the remaining number of erase and write operations that the solid-state hard drive can withstand. The intelligent monitoring tool is a self-monitoring analysis and reporting technology tool. The cumulative operating time of the mechanical hard drive is obtained by reading the cumulative operating hours recorded by the mechanical hard drive firmware. The cumulative operating hours are accumulated since the hard drive was first enabled.

[0053] Remaining storage space is obtained by calling the distributed storage system's storage management interface, which returns the currently unused storage capacity in gigabytes. Input / output pressure is determined by monitoring the storage node's input / output operation queue depth. Queue depth represents the number of pending input / output requests. Queue depth is collected in real time by the operating system's performance monitoring tool and converted into a pressure index. The mapping rule between the pressure index and queue depth is as follows: when the queue depth is less than or equal to the preset threshold, the pressure index is equal to the queue depth divided by the preset threshold. When the queue depth is greater than the preset threshold, the pressure index is fixed to a maximum value of 1.

[0054] The distributed storage system's routing configuration table is parsed to obtain inter-node connectivity. The routing configuration table records the physical or logical connection topology between storage nodes. During parsing, invalid connection paths using the heartbeat detection protocol are filtered out. Data migration frequency is obtained by counting the number of data block migration operations within a 5-minute sliding window. The migration operation count is extracted and aggregated from the storage system's operation logs using an open-source log processing framework. The sliding window has a step size of 1 minute.

[0055] Hardware performance indicators, real-time load status, and data access path data are encapsulated into structured data packets in a key-value pair format. The key-value pairs include read and write speeds, latency, remaining erase and write times, cumulative operating hours, remaining storage space, input and output pressure index, and data migration frequency. The structured data packets are transmitted to the resource allocation decision module via the Advanced Message Queuing Protocol, which configures a persistent storage and transmission confirmation mechanism. The data is standardized using a normalization formula. The normalization formula is the original value minus the minimum value divided by the difference between the maximum and minimum values. The minimum and maximum values are dynamically set based on the statistical distribution of historical data from the past 30 days. For example, the maximum read and write speed is 1000 megabytes per second, and the minimum is 0 megabytes per second.

[0056] Data validity is ensured by performing an integrity check on standardized data. This integrity check includes checking field integrity and numerical rationality. The numerical rationality determination rules are as follows: the remaining erase / write count of a solid-state drive is greater than or equal to 0 and less than or equal to the manufacturer's nominal maximum erase / write count, for example, the maximum erase / write count for a certain solid-state drive model is 3,000; the cumulative operating hours of a mechanical hard drive is greater than or equal to 0 and less than or equal to the designed lifespan, for example, the designed lifespan of a certain mechanical hard drive model is 100,000 hours. Data anomalies are handled through a retransmission request mechanism. Upon detecting missing or exceeded data, the retransmission request mechanism sends a retransmission instruction containing an anomaly identifier and timestamp to the data acquisition module. Upon receiving the retransmission instruction, the data acquisition module recollects data within the specified time window. Abnormal nodes are eliminated by marking a node as temporarily unavailable after three retransmission failures. Nodes in a temporarily unavailable state do not participate in resource allocation until new valid data is received.

[0057] The specific implementation of calculating the topological entropy of a data access path and triggering a path reconstruction instruction is as follows. The storage node topology is constructed based on an adjacency matrix of inter-node connectivity. The adjacency matrix is a two-dimensional matrix, where each element represents whether there is a direct connection between two nodes. Element values are 1 if a direct connection exists, and 0 otherwise. The adjacency matrix is constructed by parsing the routing configuration table of the distributed storage system. The routing configuration table records the physical or logical connection information between storage nodes. During parsing, invalid connection paths used only for heartbeat detection are filtered out. For example, paths with a heartbeat detection protocol are excluded.

[0058] Path activity is quantified based on a time series of data migration frequency, which is the chronological count of path migration operations with a 5-minute granularity. Migration operation counts are extracted from the storage system's operation logs using an open-source log processing framework. The framework aggregates migration operation counts within a 5-minute window by path identifier, generating key-value pairs with the path as the key and the migration count as the value.

[0059] The Shannon entropy formula is used to calculate the information entropy of the probability distribution of inter-node connections. Information entropy is used to quantify the degree of topological disorder. The probability distribution of inter-node connections is the probability of each node being connected to every other node in the adjacency matrix. The connection probability is equal to the number of direct connections between the node and the total number of nodes minus 1. Information entropy is calculated by taking the logarithm of the connection probability for each node and weighting it together, with the weight being the connection probability itself. The coefficient of variation is used to quantify the degree of fluctuation in the time series of data migration frequency. The coefficient of variation of data migration frequency is equal to the standard deviation of the time series of data migration frequency divided by the mean. The standard deviation and mean are calculated based on the number of migrations in the past 24 hours.

[0060] The topology entropy value is calculated by weighting the information entropy and the coefficient of variation according to preset weights. The preset weights are dynamically adjusted based on path stability requirements. For example, for storage paths with high stability requirements, the information entropy weight is set to 0.7 and the coefficient of variation weight is set to 0.3. For paths with low stability requirements, the information entropy weight is set to 0.5 and the coefficient of variation weight is set to 0.5. Weight adjustment rules are defined in a configuration file, which is set by the system administrator based on business needs.

[0061] The topology entropy value is compared to a preset threshold. The threshold is dynamically adjusted based on the upper bound of the historical topology entropy distribution. The historical topology entropy distribution is the set of entropy values for all paths within the same time period over the past seven days. The upper bound is determined by adding three times the standard deviation to the mean of the historical entropy values. For example, if the mean of the historical entropy values is 0.6 and the standard deviation is 0.15, the upper bound is 0.6 plus three times 0.15, which equals 1.05. This dynamic adjustment process occurs every six hours to ensure that the threshold adapts to changes in system load.

[0062] When the topology entropy value exceeds the preset threshold, a migration path reconstruction instruction is generated. The migration path reconstruction instruction contains a list of path identifiers that need to be reconstructed first and the corresponding bandwidth usage constraints. The path identifier list is generated by filtering paths whose topology entropy values exceed the threshold, excluding paths that have been reconstructed within the last 30 minutes. The bandwidth usage constraints are dynamically set based on the current network bandwidth utilization. For example, if the overall bandwidth utilization exceeds 80%, the reconstruction bandwidth of a single path is limited to 5% of the total bandwidth; if the utilization is less than 80%, the limit is relaxed to 10%. Bandwidth utilization is obtained in real time through a monitoring tool, which periodically reads port traffic data from the network switch.

[0063] The specific implementation method for analyzing the nonlinear resonance correlation strength of the hardware performance decay rate between different storage nodes and determining the resonant node group is as follows. The SSD write life decline rate of each storage node is extracted as a performance decay indicator. The SSD write life decline rate is obtained by parsing the SSD self-monitoring analysis and reporting technology tool log. The log records the change in the SSD remaining erase and write times over time. The decline rate is calculated as the difference between the current remaining erase and write times and the remaining erase and write times in the previous hour divided by the time interval, which is fixed at 1 hour. The abnormal increment of the mechanical hard disk seek time of each storage node is extracted as a performance decay indicator. The abnormal increment of the mechanical hard disk seek time is obtained by monitoring the seek operation time of the mechanical hard disk. The seek operation time is recorded by the hard disk firmware and read through the operating system interface. The abnormal increment is calculated as the difference between the current seek time and the historical average seek time. The historical average seek time is based on the seek time data of the same mechanical hard disk within the past 24 hours.

[0064] A covariance matrix of hardware performance degradation rates between different nodes is constructed. This covariance matrix is used to quantify the correlation between performance degradation trends of different nodes. This covariance matrix is constructed using a covariance matrix construction algorithm. This algorithm uses the SSD write life degradation rate and mechanical hard disk seek time abnormality increment of each node as two-dimensional vectors and calculates the covariance values between these vectors. The covariance value represents the linear correlation between the performance degradation trends of two nodes. A covariance value greater than 0 indicates a positive correlation, while a covariance value less than 0 indicates a negative correlation. The number of rows and columns in the covariance matrix is equal to the total number of nodes, and the matrix elements are the covariance values of corresponding node pairs.

[0065] The maximum eigenvalue of the covariance matrix is extracted as the nonlinear resonance correlation strength through principal component analysis (PCA), a well-known statistical method for dimensionality reduction and feature extraction. PCA performs eigenvalue decomposition on the covariance matrix, yielding eigenvalues and their corresponding eigenvectors sorted by size. The maximum eigenvalue is the variance contribution of the first principal component of the covariance matrix, representing the overall correlation strength of the performance decay trends of different nodes. The component of the eigenvector corresponding to the maximum eigenvalue represents the weight of each node in the overall correlation; a larger absolute value of the weight indicates a higher contribution to the overall correlation strength.

[0066] The nonlinear resonance correlation strength is compared to a critical threshold, which is dynamically set based on the statistical distribution characteristics of the covariance matrix eigenvalues. This statistical distribution characteristic is derived by analyzing a historical covariance matrix eigenvalue dataset, which contains hourly covariance matrix eigenvalues generated over the past seven days. The critical threshold is set to the 95th percentile of the historical eigenvalue dataset. For example, if the maximum eigenvalue in the historical eigenvalue dataset ranges from 0.5 to 2.0 and the 95th percentile is 1.9, the critical threshold is set to 1.9. This dynamic setting process is performed every six hours to ensure that the threshold adapts to changes in system load.

[0067] When the strength of a nonlinear resonance correlation exceeds a critical threshold, nodes whose eigenvector weights in the covariance matrix exceed a preset ratio are marked as a resonant node group. The preset ratio of eigenvector weights is set based on business needs. For example, for highly sensitive business scenarios, a preset ratio of 10% indicates that only the top 10% of nodes by weight are marked; for general business scenarios, the preset ratio is relaxed to 20%. Weight ranking is achieved by sorting the eigenvector components in descending order of absolute value. Nodes marked within the last hour are excluded from the marking process to avoid duplicate processing.

[0068] The specific implementation method for dynamically assigning node weights to suppress oscillations based on the nonlinear resonance correlation strength, resonance node group, and path reconstruction instructions is as follows. The nonlinear resonance correlation strength is mapped to a suppression coefficient using a normalization function. The normalization function uses a linear scaling method to map the original value of the nonlinear resonance correlation strength to a range of 0 to 1, ensuring that the suppression coefficient is positively correlated with the nonlinear resonance correlation strength. For example, if the historical maximum value of the nonlinear resonance correlation strength is 5.0 and the current value is 3.0, the suppression coefficient is calculated as 3.0 divided by 5.0, which is 0.6. The maximum and minimum values of the normalization function are dynamically adjusted based on the statistical value of the nonlinear resonance correlation strength of the same storage node group over the past 24 hours, with the maximum value taking 90% of the historical peak value and the minimum value taking 50% of the historical average value.

[0069] The path reconstruction factor is generated based on the bandwidth usage constraints in the path reconstruction instructions. The path reconstruction factor is the product of the inverse of the topological entropy value of the path requiring priority reconstruction and the remaining available bandwidth. The topological entropy value of the path requiring priority reconstruction is obtained from the calculation results of the previous step. A higher topological entropy value indicates greater path chaos, and the smaller its inverse, the lower the path reconstruction factor. The remaining available bandwidth is obtained in real time by monitoring port traffic data of network switches. The monitoring tool collects bandwidth utilization data every 30 seconds. The remaining available bandwidth is calculated as the total port bandwidth minus the currently used bandwidth. For example, if the topological entropy value of a path is 2.0 and the remaining available bandwidth is 200 Mbps, the path reconstruction factor is 1 / 2.0 multiplied by 200, which equals 100.

[0070] Dynamic node allocation weights are generated based on a linear combination of the suppression coefficient, path reconstruction factor, and real-time load status. These weights are optimized using gradient descent based on historical resource allocation error data. Real-time load status includes the remaining storage space and input / output pressure index of the storage node. Remaining storage space is obtained through the storage management interface, and the input / output pressure index is calculated as the ratio of queue depth to a preset threshold.

[0071] The linear combination expression is: Node dynamic allocation weight = Weight A × Suppression Coefficient + Weight B × Path Reconstruction Factor + Weight C × Remaining Storage Space + Weight D × Input / Output Pressure Index. Weights A, B, C, and D are all initially set to 0.25. They are iteratively optimized using gradient descent combined with historical resource allocation error data. The historical resource allocation error data is a weighted average of the number of migration oscillations and hardware performance loss rates caused by improper weight allocation over the past seven days.

[0072] The gradient descent optimization process is achieved by minimizing a loss function, defined as the mean squared error between the actual oscillation frequency and the expected frequency corresponding to the dynamically assigned node weights. During each optimization iteration, the weights are adjusted based on the partial derivatives of the loss function with respect to the weight coefficients. The adjustment step size is set to 0.01, and the maximum number of iterations is 1000 or until the loss function changes by less than 0.001. For example, if the loss function value decreases by less than 0.001 after a certain iteration, the optimization is terminated and the current weight coefficients are used. The optimized weight coefficients are stored in a configuration file, which is reloaded every 6 hours to adapt to changes in system load.

[0073] The specific implementation method for extracting prediction error data caused by sudden hardware performance changes and path confusion in historical resource allocation and generating error compensation coefficients in combination with the dynamic node allocation weights is as follows. The number of data migration errors caused by the sudden drop in the write life of the solid-state drive is extracted from the resource allocation log. The number of data migration errors is the total number of data block migration failures or retry operations caused by the sudden drop in the remaining erase and write times of the solid-state drive. The log entries are obtained by parsing the event records marked as "migration failure" or "retry" in the storage system operation log. The event records contain timestamps, path identifiers, and error types. The number of invalid migration operations caused by path confusion is extracted from the resource allocation log. The number of invalid migration operations is the number of operations in which data blocks are migrated back again due to excessively high entropy values in the path topology structure. Invalid migration operations are statistically obtained by comparing the timestamps and path identifiers of the migration log and the remigration log. The time window matching tolerance is set to 1 minute.

[0074] Calculate the sliding time window mean of the number of data migration errors and the number of invalid migration operations as the standardized historical prediction error data. The window length of the sliding time window is set to 1 hour, the sliding step is set to 10 minutes, and the mean is calculated as the arithmetic mean of all data migration errors and invalid migration operations in the window. For example, in the time window from 10:00 to 11:00, the number of data migration errors is 5 times and the number of invalid migration operations is 3 times, then the standardized historical prediction error data is (5+3) / 2=4 times. The start time of the sliding time window is aligned with the peak load period of the storage system. The peak load period is determined by analyzing the resource allocation log of the past 7 days. For example, if the log shows that 10:00 to 12:00 every day is the peak period of migration operations, then the window start time is set to 10:00.

[0075] The normalized historical forecast error data and the node's dynamically assigned weights are input into the error compensation function to generate the error compensation coefficient. The error compensation function is a nonlinear mapping relationship based on the hyperbolic tangent function. The hyperbolic tangent function compresses the input value to the range of -1 to 1. By adjusting the function parameters, the error compensation coefficient is ensured to monotonically decrease with the product of the historical forecast error data and the node's dynamically assigned weight.

[0076] The error compensation coefficient is calculated using a hyperbolic tangent function. Specifically, the function is fed with the product of the standardized historical forecast error data and the dynamically assigned node weights. The function's output is compressed to the range of -1 to 1, and the output decreases monotonically as the product increases. The parameters of the hyperbolic tangent function include an attenuation factor and the base of the natural logarithm. The attenuation factor is dynamically adjusted based on the fluctuations in the historical forecast error data. This fluctuation is quantified by calculating the standard deviation of the historical forecast error data. A larger standard deviation indicates greater data fluctuations, and a smaller attenuation factor is used to slow the rate of attenuation of the compensation coefficient. A smaller standard deviation results in a larger attenuation factor to accelerate attenuation. The attenuation factor is adjusted inversely with the standard deviation, and its value range is limited to 0.1 to 1.0 to ensure the stability of the function output.

[0077] The error compensation coefficient is applied as follows: as the product of the standardized historical forecast error data and the node's dynamic allocation weight increases, the error compensation coefficient approaches -1, triggering the system to automatically lower the current node's task allocation priority. As the product decreases, the error compensation coefficient approaches 1, triggering the system to increase the node's priority. For example, if the standardized historical forecast error data is 4 and the node's dynamic allocation weight is 0.6, the product is 2.4. Combined with a decay factor of 0.5, the error compensation coefficient is approximately -0.4, corresponding to a 40% reduction in node priority. The dynamic adjustment of the decay factor is synchronized with the step size of the sliding time window. The decay factor is recalculated and updated every 10 minutes based on the latest historical error data to ensure that the compensation coefficient matches the real-time error fluctuation trend. During the update process, the decay factor value is inversely proportional to the standard deviation of the historical error data. As the standard deviation increases, the decay factor decreases to smooth the compensation coefficient change, and as the standard deviation decreases, the decay factor increases to accelerate the response.

[0078] The specific implementation method for performing data migration operations across storage nodes and limiting the high-frequency task allocation of the resonant node group based on the error compensation coefficient is as follows. The task allocation priority of each storage node is adjusted according to the error compensation coefficient. The implementation method of the negative correlation between the task allocation priority and the error compensation coefficient is as follows: the error compensation coefficient is multiplied by negative one and mapped to a priority range of 0 to 1. The lower limit of the priority range corresponds to the highest task allocation priority, and the upper limit corresponds to the lowest priority. For example, if the error compensation coefficient of a node is -0.4, the task allocation priority is calculated to be 0.4. The larger the priority value, the lower the eligibility of the node to receive high-frequency tasks. The priority adjustment result is synchronized to the task scheduler in real time. The task scheduler allocates high-frequency tasks from low to high according to the priority value to ensure that the proportion of high-frequency task allocation of low-priority nodes is suppressed.

[0079] A high-frequency task allocation ratio limit is imposed on the resonant node group. This limit is dynamically set based on the ratio of the node's dynamic allocation weight to its real-time load status. The node's dynamic allocation weight is derived from the weight value generated in the previous step. The real-time load status includes the node's remaining storage space and input / output pressure index. The remaining storage space is queried in real time through the storage management interface, and the input / output pressure index is calculated as the ratio of the queue depth to a preset threshold. The high-frequency task allocation ratio limit is calculated as follows: Limit ratio = node dynamic allocation weight / (real-time load status + 1), where the "+1" prevents the denominator from being zero. For example, if a node's dynamic allocation weight is 0.6 and its real-time load status is 0.8, the limit ratio = 0.6 / (0.8 + 1) = 0.33, meaning that the node's high-frequency task allocation ratio must not exceed 33%. The limit ratio is enforced by the task scheduler's quota management module, which periodically pulls the latest load data from the node monitoring service and updates the limit ratio. The period is set to 5 minutes.

[0080] Data blocks stored in paths requiring reconstruction are prioritized for migration based on the path identifier list in the migration path reconstruction instruction. The data block migration process adheres to bandwidth usage constraints and updates the topology entropy in real time. The list of paths requiring reconstruction is obtained from the path identifier list generated in the previous step. The bandwidth usage constraint stipulates that the bandwidth occupied by the migration operation of a single path must not exceed a preset percentage of the total available bandwidth. This preset percentage is dynamically adjusted based on the current network bandwidth utilization: if the bandwidth utilization is above 80%, the preset percentage is set to 5%; if it is below 80%, the preset percentage is relaxed to 10%. For example, if the total available bandwidth is 1000 Mbps and the current utilization is 85%, the migration bandwidth cap for path P1 is 1000 × 5% = 50 Mbps. After data block migration is completed, the topology entropy of the path is updated in real time by recalculating the node connectivity and migration frequency of the path. The new entropy value is then synchronized with the path management database to ensure that subsequent resource allocation decisions are based on the latest topology status.

[0081] This embodiment systematically addresses the resource allocation oscillation problem caused by hardware performance mutations and path chaos in traditional technologies through dynamic error compensation and multi-dimensional resource control mechanisms. Task allocation priorities are dynamically mapped based on the error compensation coefficient. A negative correlation design ensures that high-error nodes are automatically downgraded, preventing inefficient task allocation from exacerbating oscillations. Compared to traditional static threshold strategies, the priority value is strictly bound to the real-time error state, blocking the vicious cycle of error accumulation. The ratio constraint of the node's dynamic weight and real-time load considers hardware performance degradation (dynamic weight) and instantaneous load pressure (real-time load) in a coordinated manner, preventing overloaded nodes from further deteriorating due to high-frequency tasks. For example, the task ratio of nodes with low dynamic weight and high load is significantly compressed, reducing the risk of collective hardware degradation. Combining the dual constraints of path entropy and bandwidth occupancy, highly chaotic paths are reconstructed first and migration bandwidth is strictly limited. This not only optimizes the topological order through a closed-loop entropy update, but also prevents migration operations from crowding out service traffic. Compared to traditional random migration or pure load-driven strategies, this improves path stability and service continuity.

[0082] Parameters such as the error compensation coefficient, attenuation factor, and limit ratio are dynamically adjusted based on historical and real-time data, replacing the lack of adaptability caused by fixed parameters. Cross-validation of multiple metrics, including hardware performance, path status, and network load, overcomes the biased nature of single-metric decision-making. Core parameters such as priority, limit ratio, and entropy are updated on a minute-by-minute basis to ensure that response speed matches the frequency of system disturbances. Compared to existing technologies, this combined design of dynamic weight decay, closed-loop path entropy updates, and elastic bandwidth usage constraints reduces the frequency of resource allocation oscillations, shortens hardware lifespan, and reduces ineffective migration operations.

[0083] Through the synergistic mechanism of topological entropy monitoring and nonlinear resonance analysis, a dual suppression barrier is constructed to address the two core issues of path disorder and hardware group degradation. Topological entropy monitoring dynamically calculates the path structure disorder based on the connection relationship between nodes and the migration frequency. When the entropy value exceeds the limit, the path reconstruction instruction is triggered, optimizing the data distribution topology in real time and blocking local oscillations caused by path disorder. Nonlinear resonance analysis quantifies the group correlation strength of hardware performance degradation through the covariance matrix, identifies resonant node groups and limits their high-frequency task load, avoiding single-node performance mutations from becoming group degradation hotspots. The two are jointly controlled through a dynamic weight generation module: the path reconstruction factor suppresses the increase of topological entropy, and the resonance suppression coefficient alleviates hardware resonance, ultimately forming a two-way negative feedback control of path order and hardware status, suppressing local oscillations in the bud and preventing them from spreading globally through topological connections or hardware coupling.

[0084] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.

[0085] It should be noted that the present invention can be deployed on the device itself to realize embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0086] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0087] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0088] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0090] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0091] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A computing platform management system based on artificial intelligence, characterized in that: Includes the following modules: Real-time acquisition module: real-time acquisition of hardware performance indicators, real-time load status and data access paths of each storage node in the distributed storage system; Entropy value determination module: calculates the topological structure entropy value of the data access path. When the topological structure entropy value exceeds the preset threshold, it is determined that the migration path is in a chaotic state and triggers the path reconstruction instruction; Resonance analysis module: Analyzes the nonlinear resonance correlation strength of hardware performance decay rates between different storage nodes based on hardware performance indicators. When the nonlinear resonance correlation strength exceeds a critical threshold, a resonant node group is determined. Weight generation module: Based on the nonlinear resonance correlation strength, resonance node group and path reconstruction instructions, it generates node dynamic allocation weights to suppress oscillation through a dynamic feature weighting model; Error compensation module: This module extracts prediction error data caused by hardware performance mutations and path confusion in historical resource allocation, and generates error compensation coefficients based on the dynamic node allocation weights. Migration execution module: performs data migration operations across storage nodes based on the error compensation coefficient. The data migration operation limits the high-frequency task allocation of the resonant node group and prioritizes the reconstruction of chaotic paths.

2. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: Real-time acquisition of hardware performance indicators, real-time load status, and data access paths for each storage node in the distributed storage system, including: The read and write speed, latency, and storage medium life of storage nodes are collected as hardware performance indicators. The storage medium life is quantified by the remaining erase and write times of the solid-state drive and the cumulative operating time of the mechanical hard drive. The real-time load status of the storage node is collected, and the real-time load status includes the remaining storage space and input and output pressure. The connection relationship between nodes in the data access path and the frequency of data migration are collected.

3. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: Calculate the topological entropy of the data access path. When the topological entropy exceeds a preset threshold, the migration path is determined to be in a chaotic state and a path reconstruction instruction is triggered, including: Based on the adjacency matrix of the connection relationship between nodes and the time series of data migration frequency, the topological structure entropy value is calculated using the Shannon entropy formula. The Shannon entropy formula states that the topological structure entropy value is equal to the weighted sum of the information entropy of the probability distribution of the connection relationship between nodes and the coefficient of variation of the data migration frequency. Compare the topological structure entropy value with a preset threshold value, and dynamically adjust the preset threshold value according to the upper limit of the statistical distribution of the historical topological structure entropy value; When the topology entropy value exceeds a preset threshold, a migration path reconstruction instruction is generated. The migration path reconstruction instruction includes a list of path identifiers that need to be reconstructed first and corresponding bandwidth occupancy constraints.

4. The computing platform management system based on artificial intelligence according to claim 3, characterized in that: The coefficient of variation of data migration frequency is calculated as the standard deviation of the time series of data migration frequency divided by the mean.

5. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: The nonlinear resonance correlation strength of the hardware performance decay rate between different storage nodes is analyzed based on hardware performance indicators. When the nonlinear resonance correlation strength exceeds a critical threshold, a resonant node group is determined, including: Extract the decline rate of the solid-state drive write lifespan and the abnormal increase in the mechanical hard drive seek time of each storage node as the time series data of the hardware performance degradation rate; Construct a covariance matrix of the hardware performance degradation rate between different nodes, and use principal component analysis to extract the maximum eigenvalue of the covariance matrix as the nonlinear resonance correlation strength; The nonlinear resonance correlation strength is compared with a critical threshold value, which is dynamically set according to the statistical distribution characteristics of the eigenvalues of the covariance matrix; When the nonlinear resonance correlation strength exceeds a critical threshold, the nodes whose eigenvector weights in the covariance matrix exceed a preset ratio are marked as a resonance node group.

6. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: Based on the nonlinear resonance correlation strength, the resonance node group, and the path reconstruction instructions, a dynamic feature weighting model is used to generate node dynamic allocation weights for suppressing oscillations, including: The nonlinear resonance correlation strength is mapped to the suppression coefficient through a normalization function, and the normalization function ensures that the suppression coefficient is positively correlated with the nonlinear resonance correlation strength; Generate a path reconstruction factor according to the bandwidth occupancy constraint in the path reconstruction instruction. The path reconstruction factor is the product of the inverse of the topological structure entropy value of the path to be reconstructed first and the remaining available bandwidth. The node dynamic allocation weight is generated based on the linear combination of the suppression coefficient, path reconstruction factor and real-time load status.

7. The computing platform management system based on artificial intelligence according to claim 6, characterized in that: The weight coefficients of the linear combination are optimized by gradient descent method based on historical resource allocation error data.

8. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: Extract the prediction error data caused by hardware performance mutation and path confusion in historical resource allocation, and generate error compensation coefficients based on the dynamic node allocation weights, including: The number of data migration errors caused by the sudden decrease in the SSD's write lifespan and the number of invalid migration operations caused by path confusion are extracted from the resource allocation log as historical prediction error data. Calculate the sliding time window mean of the number of data migration errors and the number of invalid migration operations as the standardized historical prediction error data; The normalized historical prediction error data and the node dynamic allocation weight are input into the error compensation function to generate the error compensation coefficient.

9. The computing platform management system based on artificial intelligence according to claim 8, characterized in that: The error compensation function is a nonlinear mapping relationship based on the hyperbolic tangent function, which ensures that the error compensation coefficient decreases monotonically with the product of the historical prediction error data and the node dynamic allocation weight.

10. The computing platform management system based on artificial intelligence according to claim 1, characterized in that: Data migration operations across storage nodes are performed based on the error compensation coefficient. The data migration operation limits the high-frequency task allocation of the resonant node group and prioritizes the reconstruction of chaotic paths, including: Adjust the task allocation priority of each storage node according to the error compensation coefficient. The task allocation priority is negatively correlated with the error compensation coefficient. A high-frequency task allocation ratio limit is imposed on the resonant node group. The high-frequency task allocation ratio limit is dynamically set according to the ratio of the node's dynamic allocation weight to the real-time load status; According to the path identification list in the migration path reconstruction instruction, priority migration requires priority reconstruction of the data blocks stored in the path. The data block migration process follows the bandwidth occupancy constraint and updates the topology structure entropy value in real time.

Citation Information

Patent Citations

  • Dynamic reconstruction method of computing power network and servo device

    CN115086178A

  • Performance evaluation system and method for multi-stage RAID system

    CN117909198A

  • Artificial intelligence-based text travel uniform light guide illumination fault diagnosis method and system

    CN118981734A

  • Software query information management method and system

    CN119669315A

  • Node load-based dynamic data partitioning system

    WO2021073083A1

Cited By

  • College budget management system fusing new quality productivity theory

    CN121303787A

  • Coordination method and system for data intercommunication of heterogeneous control system of thermal power plant

    CN121728122A

  • Coordinated method and system for data intercommunication of heterogeneous control systems of thermal power plants

    CN121728122B

  • Production line resource scheduling management system based on energy efficiency data

    CN121903328A

  • AI-driven heterogeneous data intelligent hierarchical storage and dynamic optimization system

    CN122195354A