Resource scheduling method and device based on IT asset health assessment, medium and equipment
Through a multimodal fusion IT asset health assessment method, the node health score and business priority are obtained and calculated. Combined with affinity and anti-affinity scheduling rules, the shortcomings of IT asset health assessment and scheduling in existing technologies are solved, and a comprehensive and real-time assessment of IT assets and optimized resource allocation are achieved, thereby improving system stability and reliability.
Patent Information
- Application Number
- CN202510821892.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-17
AI Technical Summary
Existing health monitoring and scheduling methods are unable to conduct comprehensive and real-time comprehensive assessment and scheduling of IT assets, and cannot adapt to the complex and changing cloud environment.
By obtaining resource usage data, network connection data, and abnormal information data of each node in the distributed system, a multimodal fusion health score is calculated, and a comprehensive evaluation value is calculated based on the business priority. Resource scheduling is performed based on the comprehensive evaluation value, and affinity and anti-affinity scheduling rules are used to optimize resource allocation.
It achieves a comprehensive and real-time assessment of the health status of IT assets, and improves the stability, reliability and operation and maintenance automation level of the system.
Smart Images

Figure CN120803695A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource scheduling, in particular to a resource scheduling method and device for IT asset health assessment based on multi-modal fusion, a medium and equipment. BACKGROUND
[0002] With the rapid development of cloud computing and containerization technology, modern IaaS (Infrastructure as a Service) and PaaS (Platform as a Service) platforms widely adopt distributed system architecture to support various business applications. In this context, real-time perception and quantitative evaluation of the health status of underlying IT resources, and intelligent scheduling based on the evaluation results, have become a key link to ensure system high availability, improve resource utilization, and enhance the level of operation and maintenance automation.
[0003] Current mainstream health monitoring and scheduling methods mostly rely on centralized monitoring frameworks such as Prometheus, ELK, etc., to collect performance indicators or log information such as CPU, memory, and disk usage of nodes. At the scheduling level, static rules and threshold-triggered strategies are usually adopted, such as starting expansion or migration operations when the load of a node exceeds the set threshold. Such methods have achieved monitoring and scheduling control of resource usage to some extent, but still have obvious limitations when facing complex and variable cloud environments.
[0004] In summary, existing health monitoring and scheduling are all based on processing one kind of data, which cannot comprehensively and real-time evaluate the health status of IT assets and perform corresponding scheduling processing. SUMMARY
[0005] Therefore, the present application provides a resource scheduling method and device for IT asset health assessment based on multi-modal fusion, a medium and equipment, which mainly aims to solve the problem that existing health monitoring and scheduling cannot comprehensively and real-time evaluate the health status of IT assets and perform corresponding scheduling processing.
[0006] According to one aspect of the present application, a resource scheduling method for IT asset health assessment based on multi-modal fusion is provided, which comprises:
[0007] Obtaining resource usage data, network connection data, and abnormal information data corresponding to each node in a distributed system, performing health score calculation based on the resource usage data, the network connection data, and the abnormal information data to obtain a health score corresponding to each node;
[0008] Obtaining the priority of the business adapted to each node, and performing calculation based on the health score corresponding to each node and the priority of the adapted business to obtain a comprehensive evaluation value corresponding to each node;
[0009] Based on the comprehensive evaluation value corresponding to each node, the resource is scheduled.
[0010] Optionally, the abnormal information data includes a fault domain risk coefficient, an abnormal log frequency and a historical failure rate, the health score calculation based on the resource usage data, the network connection data and the abnormal information data obtains a health score corresponding to each node, and the health score calculation includes:
[0011] The collection time corresponding to the resource usage data, the network connection data and the abnormal information data is divided into a plurality of time windows, and the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window are obtained;
[0012] The resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window are sequentially subjected to abnormality detection, smoothing processing and normalization processing;
[0013] The weights corresponding to the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate are obtained, the normalized resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency and historical failure rate and the weights corresponding thereto are substituted into a health score calculation formula, and a health score corresponding to each node is obtained.
[0014] Optionally, the comprehensive evaluation value corresponding to each node is calculated based on the health score corresponding to each node and the priority of the adapted service, and the comprehensive evaluation value corresponding to each node is obtained.
[0015] The health score total score value is calculated based on the health score corresponding to each node and the weight thereof.
[0016] The service total score value is calculated based on the priority of the service adapted by each node and the weight thereof.
[0017] The sum of the health score total score value and the service total score value of each node is taken as the comprehensive evaluation value corresponding to each node.
[0018] Optionally, the resource is scheduled based on the comprehensive evaluation value corresponding to each node, and the scheduling includes:
[0019] The node with a comprehensive evaluation value greater than a comprehensive value threshold is taken as a candidate node.
[0020] The target resource is allocated to the candidate node based on an affinity scheduling rule, and the related resource of the target resource is allocated to another different candidate node based on an anti-affinity scheduling rule.
[0021] Optionally, after the target resource is allocated to the candidate node based on the affinity scheduling rule, the resource scheduling method based on the multi-modal fusion IT asset health assessment further comprises:
[0022] After the target resource is allocated to the candidate node, the resource usage data, network connection data and abnormal information data corresponding to the candidate node to which the target resource is allocated after a preset period of time are acquired, health score calculation is performed based on the resource usage data, network connection data and abnormal information data corresponding to the candidate node after the preset period of time, and a new health score corresponding to the candidate node to which the target resource is allocated is obtained.
[0023] If the new health score of the candidate node to which the target resource is allocated is less than a preset first health threshold, the target resource is migrated to another candidate node, and if the new health score of the candidate node to which the target resource is allocated is greater than a preset second health threshold, another resource is allocated to the candidate node to which the target resource is allocated.
[0024] Optionally, the following method is used to acquire the fault domain risk coefficient:
[0025] The fault history record is acquired, a plurality of fault occurrence regions are determined according to the fault history record, and the first fault quantity and the fault influence degree corresponding to each fault occurrence region in a first preset time range are determined from the fault history record.
[0026] For each fault occurrence region, the quotient of the first fault quantity and the total running time in the first preset time range is taken as the fault occurrence probability, and the product of the fault occurrence probability and the fault influence degree is taken as the fault risk coefficient of each fault occurrence region.
[0027] The sum of the fault risk coefficients of all fault occurrence regions is taken as the fault domain risk coefficient.
[0028] The following method is used to acquire the historical fault rate:
[0029] The second fault quantity in a second preset time range is determined from the fault history record, and the quotient of the second fault quantity and the total running time in the second preset time range is taken as the historical fault rate.
[0030] Optionally, after the health score calculation based on the resource usage data, the network connection data and the abnormal information data is performed to obtain the health score corresponding to each node, the resource scheduling method based on the multi-modal fusion IT asset health assessment further comprises:
[0031] When the health score of any node configured with a virtual machine is less than a preset third health threshold, virtual machine migration information is output.
[0032] According to another aspect of the present application, a resource scheduling device based on multi-modal fusion IT asset health assessment is provided, comprising:
[0033] a health assessment module configured to acquire resource usage data, network connection data and abnormal information data corresponding to each node in a distributed system, perform health score calculation based on the resource usage data, the network connection data and the abnormal information data, and obtain a health score corresponding to each node;
[0034] a comprehensive assessment module configured to acquire a priority of a service adapted to each node, perform calculation based on the health score corresponding to each node and the priority of the service adapted to each node, and obtain a comprehensive assessment value corresponding to each node;
[0035] a resource scheduling module configured to perform resource scheduling based on the comprehensive assessment value corresponding to each node.
[0036] Optionally, the abnormal information data comprises a fault domain risk coefficient, an abnormal log frequency and a historical failure rate, and the health assessment module is further configured to:
[0037] divide the collection time corresponding to the resource usage data, the network connection data and the abnormal information data into a plurality of time windows, and acquire resource usage data, network connection data, a fault domain risk coefficient, an abnormal log frequency and a historical failure rate corresponding to each time window;
[0038] perform abnormal detection, smoothing processing and normalization processing on the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window in sequence;
[0039] acquire a weight corresponding to each of the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate, and substitute the normalized resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency and historical failure rate and the weight corresponding thereto into a health score calculation formula to obtain the health score corresponding to each node.
[0040] Optionally, the comprehensive assessment module is further configured to:
[0041] calculate a health score total value based on the health score corresponding to each node and the weight thereof;
[0042] calculate a service total value based on the priority of the service adapted to each node and the weight thereof;
[0043] sum the health score total value and the service total value of each node as the comprehensive assessment value corresponding to each node.
[0044] Optionally, the resource scheduling module is further configured to:
[0045] take a node with an integrated evaluation value greater than an integrated value threshold as a candidate node;
[0046] allocate a target resource to the candidate node based on an affinity scheduling rule, and allocate a related resource of the target resource to another different candidate node based on an anti-affinity scheduling rule.
[0047] Optionally, the resource scheduling module is further configured to:
[0048] after allocating the target resource to the candidate node, acquire resource usage data, network connection data and abnormal information data corresponding to the candidate node after a preset period of time, perform health score calculation based on the resource usage data, network connection data and abnormal information data corresponding to the candidate node after the preset period of time, and obtain a new health score corresponding to the candidate node to which the target resource is allocated;
[0049] if the new health score of the candidate node to which the target resource is allocated is less than a preset first health threshold, migrate the target resource to another candidate node, and if the new health score of the candidate node to which the target resource is allocated is greater than a preset second health threshold, allocate another resource to the candidate node to which the target resource is allocated.
[0050] Optionally, the resource scheduling device for IT asset health assessment based on multi-modal fusion further comprises:
[0051] a fault domain risk coefficient acquisition module, configured to acquire a fault history record, determine a plurality of fault occurrence regions from the fault history record, determine a first fault quantity and a fault influence degree corresponding to each fault occurrence region in a first preset time range from the fault history record, take a quotient of the first fault quantity and a total running time in the first preset time range as a fault occurrence probability for each fault occurrence region, take a product of the fault occurrence probability and the fault influence degree as a fault risk coefficient of each fault occurrence region, and take a sum of the fault risk coefficients of all fault occurrence regions as a fault domain risk coefficient;
[0052] a historical fault rate acquisition module, configured to determine a second fault quantity in a second preset time range from the fault history record, and take a quotient of the second fault quantity and a total running time in the second preset time range as a historical fault rate.
[0053] Optionally, the resource scheduling device for IT asset health assessment based on multi-modal fusion further comprises:
[0054] The migration information output module is configured to output virtual machine migration information when the health score of any node configured with the virtual machine is less than a preset third health threshold.
[0055] According to another aspect of the present application, a storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform operations corresponding to the resource scheduling method for IT asset health assessment based on multi-modal fusion.
[0056] According to another aspect of the present application, a computer device is provided, which comprises a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus.
[0057] The memory is configured to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the resource scheduling method for IT asset health assessment based on multi-modal fusion.
[0058] By means of the above technical solution, the technical solution provided by the embodiments of the present application has at least the following advantages:
[0059] The present application provides a resource scheduling method, device, medium and equipment for IT asset health assessment based on multi-modal fusion, the health score of each node is calculated based on the resource usage data, network connection data and abnormal information data of each node, the comprehensive evaluation value corresponding to each node is calculated based on the health score of each node and the priority of the adapted service, and the resources are allocated to the appropriate node based on the comprehensive evaluation value. By calculating the comprehensive evaluation value of multi-modal fusion, the resources are scheduled according to the comprehensive evaluation value, the health status of the IT asset is comprehensively and real-timely evaluated, the corresponding scheduling processing is performed according to the comprehensive evaluation, and the stability, reliability and operation automation level of the system are significantly improved.
[0060] The above description is only a summary of the technical solution of the present application, in order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0061] Various other advantages and benefits will become apparent to those of ordinary skill in the art, upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments, and are not meant to limit the present application. Moreover, the same reference numerals are used throughout the accompanying drawings to represent same or similar components. In the drawings:
[0062] Figure 1A flow chart of a resource scheduling method for IT asset health assessment based on multi-modal fusion is shown.
[0063] Figure 2 A flow chart of another resource scheduling method for IT asset health assessment based on multi-modal fusion is shown.
[0064] Figure 3 A flow chart of still another resource scheduling method for IT asset health assessment based on multi-modal fusion is shown.
[0065] Figure 4 A component block diagram of a resource scheduling device for IT asset health assessment based on multi-modal fusion is shown.
[0066] Figure 5 A structural schematic diagram of a computer device is shown.
[0067] wherein,
[0068] Figure 4 In the figure, 402 is a health assessment module; 404 is a comprehensive assessment module; and 406 is a resource scheduling module.
[0069] Figure 5 In the figure, 502 is a processor; 504 is a communication interface; 506 is a memory; 508 is a communication bus; and 510 is a program. DETAILED DESCRIPTION
[0070] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0071] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined purposes, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "an embodiment" or "embodiments" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0072] In order to solve the problem that the existing health monitoring and scheduling cannot comprehensively and timely assess the health status of IT assets and perform corresponding scheduling processing, the present application provides a resource scheduling method for IT asset health assessment based on multi-modal fusion, as shown in Figure 1 The method comprises the following steps.
[0073] 102: Obtain resource usage data, network connection data, and abnormal information data corresponding to each node in the distributed system, perform health score calculation based on the resource usage data, network connection data, and abnormal information data, and obtain a health score corresponding to each node;
[0074] 104: Obtain the priority of the service adapted to each node, perform calculation based on the health score corresponding to each node and the priority of the adapted service, and obtain a comprehensive evaluation value corresponding to each node;
[0075] 106: Perform resource scheduling based on the comprehensive evaluation value corresponding to each node.
[0076] In this embodiment, the distributed system includes many nodes, such as server nodes, network device nodes, etc. The multi-modal running data of each node is obtained, and resources are allocated to suitable nodes according to the multi-modal running data of each node.
[0077] Resource usage data (including CPU usage, CPU load, memory utilization, disk I / O read / write rate, disk space usage, network bandwidth usage, etc.) is obtained through the Prometheus monitoring framework, network connection data is obtained through network detection tools, such as network connectivity index N (including end-to-end delay (RTT), packet loss rate, bandwidth utilization, port connectivity, etc.), abnormal information data is obtained, and the abnormal information data includes fault domain risk coefficient, abnormal log frequency, and historical failure rate. The fault domain risk coefficient is obtained through the fault domain information acquisition module, the abnormal log frequency is counted through the log collection system, and the historical failure rate is obtained through the historical failure data interface.
[0078] The above multi-modal data is standardized to obtain a unified format, stored in a time series database in time sequence, and labeled with source and call chain labels. Health score calculation is performed based on the resource usage data, network connection data, and abnormal information data, and a health score corresponding to each node is obtained.
[0079] The priority of the service adapted to each node is obtained, calculation is performed based on the health score corresponding to each node and the priority of the adapted service, and a comprehensive evaluation value corresponding to each node is obtained. Resources are allocated to nodes with high health scores and in compliance with affinity rules by using Kubernetes Pod affinity / anti-affinity scheduling, virtual machine drift, or global load balancing rules.
[0080] The application provides a resource scheduling method for IT asset health evaluation based on multi-modal fusion. Compared with the prior art, the health score of each node is calculated based on the resource usage data, network connection data and abnormal information data of each node, the comprehensive evaluation value corresponding to each node is calculated based on the health score of each node and the priority of the adapted service, and the resources are allocated to the appropriate nodes based on the comprehensive evaluation value. Through the calculation of the comprehensive evaluation value of multi-modal fusion, the resources are scheduled according to the comprehensive evaluation value, the health status of the IT asset is comprehensively and real-timely evaluated, the corresponding scheduling processing is performed according to the comprehensive evaluation, and the stability, reliability and operation automation level of the system are significantly improved.
[0081] In one embodiment, the abnormal information data includes a fault domain risk coefficient, an abnormal log frequency and a historical failure rate, as shown in Figure 2 The resource scheduling method for IT asset health evaluation based on multi-modal fusion further includes:
[0082] 202: dividing the collection time corresponding to the resource usage data, the network connection data and the abnormal information data into a plurality of time windows, and obtaining the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window;
[0083] 204: sequentially performing abnormal detection, smoothing processing and normalization processing on the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window;
[0084] 206: obtaining the weight corresponding to the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate respectively, and substituting the normalized resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate and the weights corresponding thereto into a health score calculation formula to obtain the health score corresponding to each node;
[0085] 208: calculating the total score of the health score based on the health score corresponding to each node and the weight thereof, calculating the total score of the service based on the priority of the service adapted to each node and the weight thereof, and taking the sum of the total score of the health score of each node and the total score of the service as the comprehensive evaluation value corresponding to each node;
[0086] 210: scheduling the resources based on the comprehensive evaluation value corresponding to each node.
[0087] Specifically, the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate are written as a matrix vector. The row vector X of the matrix vector X is the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window, and the column vector Y is the weight corresponding to the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate. jAll the data are standardized outputs from Prometheus, network probing tool, fault domain information module, log collection system and historical fault data interface; reference health degree labels The data used for weight training and verification are diversified and standardized (including monitoring, network, fault domain, log and historical fault records) and supplemented by SLA or expert evaluation labels, which greatly improves the accuracy and interpretability of model training.
[0088] The matrix vector X is stored in the time series database in time series, and each standardized index data stored in the time series database is aligned in a fixed time window Δt to form the latest value or sliding window statistical value R(t), N(t), F(t), L(t), H f (t) in the window, which is used to ensure that different indicators are comparable at the same time granularity, that is, the collection time is divided into multiple time windows, and the matrix vector in the same time window is aligned.
[0089] As Figure 3 shown, the 3σ principle is applied to each vector index in each time window for anomaly detection, and the sampling value significantly deviating from the normal value is replaced by the mean or median of the adjacent normal value to obtain the smoothed index for removing monitoring noise and transient jitter and improving the robustness of the score; for each smoothed index , Min-Max normalization is performed according to the historical maximum and minimum value:
[0090]
[0091] In the formula, resource usage data x1=R, network connection data x2=N, fault domain risk coefficient x3=F, abnormal log frequency x4=L, and historical failure rate x5=H f are used to eliminate dimensional differences and map all features to the range [0, 1] for easy weighting.
[0092] The normalized indicators are combined, and the normalized data is formed into a feature vector at time t according to the four windows:
[0093] X(t)=[x1(t),x2(t),x3(t),x4(t),x5(t)] T , which is used to unify the multi-source data into a vector format that can be used for linear model operation.
[0094] The historical sample set is used to solve the weight vector w corresponding to resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency and historical failure rate by least squares method with regularization:
[0095]
[0096] Combine the score error of the last k windows, dynamically update w by simple recursion or sliding least squares, to ensure the generalization ability of the model while supporting real-time fine-tuning of the importance of the indicators.
[0097] At the end of each window, calculate and output the health score H(t), and the mathematical expression of the health score H(t) is:
[0098]
[0099] For quantifying the overall health status of the current node as a key input for priority scheduling.
[0100] Periodically compare the real-time calculated H(t) with the SLA event or manual assessment H ref (t) at the corresponding time, calculate the average error E, and the mathematical expression of the average error E is:
[0101]
[0102] If E exceeds the preset threshold, trigger model retraining or manual calibration to continuously ensure the accuracy and reliability of the scoring model.
[0103] The health score (H) is a comprehensive score (0-100) of node resource utilization, failure rate, network delay, etc.
[0104] The business priority (P) is a weight assigned according to the importance of the business (e.g. core business P=5, non-core business P=1).
[0105] Evaluation value calculation: calculate the comprehensive evaluation value (S) of the node through the weighted formula
[0106] S = α × H + β × P
[0107] Where α is the health weight (e.g. 0.7), and β is the business priority weight (e.g. 0.3).
[0108] In one embodiment, the following method is used to obtain the fault domain risk coefficient:
[0109] Obtain the fault history record, determine a plurality of fault occurrence areas according to the fault history record, and determine the first fault quantity and the fault impact degree corresponding to each fault occurrence area in the first preset time range from the fault history record;
[0110] For each failure occurrence area, the quotient of the first failure number and the total running time within the first preset time range is taken as the failure occurrence probability, and the product of the failure occurrence probability and the failure impact degree is taken as the failure risk coefficient of each failure occurrence area;
[0111] The sum of the failure risk coefficients of all failure occurrence areas is taken as the failure domain risk coefficient.
[0112] In an embodiment, the historical failure rate is obtained by the following method:
[0113] From the failure history record, the second failure number within the second preset time range is determined, and the quotient of the second failure number and the total running time within the second preset time range is taken as the historical failure rate.
[0114] Specifically, the failure domain risk coefficient is obtained by comprehensive calculation based on the collected failure isolation levels of the machine room, rack and server levels. By comprehensively evaluating the failure domain risk coefficient based on the isolation and redundancy information of the machine room, rack and server levels, the potential failure risk brought by physical deployment can be more accurately reflected, thereby improving the comprehensiveness and reliability of the health degree evaluation.
[0115] The failure domain risk coefficient is used to quantify the possibility and impact of a failure occurring in a certain failure domain (such as a service, component, or network area) in the system. It is calculated based on data in the failure history record (past failure events and their impact ranges) and the asset database (which records the dependency relationships between system components).
[0116] Data is collected from the above sources, including the time, duration, impact range, and recovery time of failures. Risk assessment is performed for each failure domain, considering factors such as failure frequency, severity, number of affected users, and business value. Different failure domains are assigned weights based on their importance and impact on the business.
[0117] The calculation formula is:
[0118] Occurrence probability (P): calculated based on historical failure frequency
[0119] P = Number of failures in the past 30 days / Total running time (hours) * 100%.
[0120] Impact degree (C): quantified according to the impact range of the failure
[0121] Level 0 (no impact): the failure does not affect the business.
[0122] Level 1 (partial impact): affects a single service or a small number of users.
[0123] Level 2 (serious impact): causes system degradation or widespread user unavailability.
[0124] The risk coefficient (F) can be expressed as:
[0125]
[0126] The historical failure rate (H) characterizes the frequency of system failures within a specific time period. Failure management tools record the occurrence time, duration, resolution time, etc. of system failures. Monitoring systems provide real-time system status and performance indicators to help identify potential failures.
[0127] Collect failure event data within a specific time range, classify them according to the type, cause, impact range, etc. of the failure, and determine the time window for calculating the historical failure rate, such as the past 30 days, 90 days, or one year.
[0128] The calculation formula is:
[0129] H = Number of failures / Total running time * 100%.
[0130] In one embodiment, based on the comprehensive evaluation value corresponding to each node, the scheduling of resources is carried out, including:
[0131] Nodes with a comprehensive evaluation value greater than the comprehensive value threshold are selected as candidate nodes.
[0132] Based on the affinity scheduling rule, the target resource is allocated to the candidate node, and based on the anti-affinity scheduling rule, the related resources of the target resource are allocated to other different candidate nodes.
[0133] Specifically, the comprehensive value threshold is used as a judgment standard. Only nodes with a comprehensive evaluation value exceeding this threshold are eligible to enter the next resource allocation process and become candidate nodes. For example, if the comprehensive threshold is set to 70 points, a server node with an evaluation value of 85 points will be selected as a candidate node, while a node with an evaluation value of 65 points will be excluded.
[0134] The affinity scheduling rule refers to the strategy of preferring to allocate resources that are related or have cooperative needs to the same node. Target resources (such as application programs, data tasks, etc.) will be preferentially allocated to previously screened candidate nodes, as these nodes have been proven to have better processing capabilities. The advantage of this is that it can reduce the communication cost between resources and improve data processing efficiency.
[0135] The anti-affinity scheduling rule is just the opposite, and is intended to avoid excessive concentration of resources in certain nodes, improving the reliability and fault tolerance of the system. Related resources (such as dependent components, backup data, etc.) of the target resource will be dispersedly allocated to other different candidate nodes. For example, to prevent single-point failure, an application program and its backup data are deployed on different server nodes, so that even if a node fails, the system can still rely on other nodes to continue running.
[0136] In Kubernetes, Pod affinity / anti-affinity scheduling is implemented through affinity and anti-affinity rules. Affinity scheduling (Pod co-location with nodes) is to deploy high-priority services on nodes with high health. Anti-affinity scheduling (Pod scattered deployment) avoids the impact of single node failure on multiple high-priority Pods, and distributes other resources to other candidate nodes.
[0137] In another embodiment, the Kubernetes Pod affinity / anti-affinity scheduling strategy includes: preferentially scheduling services to nodes with high health, in the same machine room or low load nodes; when there is no available node in the cluster, further calling global load balancing for cross-machine room traffic switching, using the Pod affinity / anti-affinity scheduling strategy based on health and service priority, and automatically switching to global load balancing when there is no available node, realizing intelligent traffic scheduling across machine rooms, and significantly enhancing disaster recovery capability and business continuity.
[0138] In one embodiment, after the target resource is allocated to the candidate node based on the affinity scheduling rule, the resource scheduling method based on multi-modal fusion IT asset health assessment further includes:
[0139] After the target resource is allocated to the candidate node, the resource usage data, network connection data and abnormal information data of the candidate node allocated with the target resource after a preset period of time are obtained, and health score calculation is performed based on the resource usage data, network connection data and abnormal information data after the preset period of time, to obtain a new health score corresponding to the candidate node allocated with the target resource.
[0140] If the new health score of the candidate node allocated with the target resource is less than a preset first health threshold, the target resource is migrated to another candidate node, and if the new health score of the candidate node allocated with the target resource is greater than a preset second health threshold, another resource is allocated to the candidate node allocated with the target resource.
[0141] Specifically, after the target resource is allocated to the candidate node, the candidate node runs for a preset period of time (for example, 5 minutes, 10 minutes), and the resource usage data, network connection data and abnormal information data of the candidate node are obtained. Based on the above resource usage data, network connection data and abnormal information data, a new health score of the candidate node is calculated.
[0142] The new health score reflects whether the current running state of the candidate node is good, and the higher the score, the healthier the candidate node, and the lower the score, the higher the load or the risk of failure.
[0143] If the new health score of the candidate node is lower than a first threshold (such as 80), it means that the candidate node is not very stable or is in a high load state. In order to avoid service interruption or performance degradation, the "target resources" originally allocated on this candidate node are migrated to another more healthy candidate node, so as to guarantee the availability and stability of the service.
[0144] If the health score of the candidate node is higher than a second threshold (such as 90), it means that the candidate node is currently very healthy and has idle resources to carry more tasks, so other resources are allocated to the candidate node.
[0145] In an embodiment, after the health score calculation based on the resource usage data, network connection data and abnormal information data, the resource scheduling method based on multi-modal fusion IT asset health assessment further comprises:
[0146] When the health score of any node configured with a virtual machine is less than a preset third health threshold, virtual machine migration information is output.
[0147] Specifically, for the virtual machine case, when the health score of the node is lower than the preset third health threshold, for example, the drift threshold T migrate , the virtual machine migration information is output, so that the cloud platform API automatically migrates the virtual machine to a node with higher health, and synchronously updates the GSLB / DNS routing. The virtual machine can be quickly migrated and the routing can be updated when the health of the node is lower than the third health threshold, so as to minimize the impact of the fault and improve the reliable recovery capability of the system.
[0148] In an embodiment, the call chain information collected by the log system is associated with other indicators to form a unified link health portrait, which is used for automatic root cause positioning. The link health portrait function is repeatedly emphasized, which further highlights the key role in supporting multi-source data correlation analysis and automatic root cause diagnosis, and ensures the accuracy of operation and maintenance decisions.
[0149] Further, as an implementation of the method shown in the above Figure 1 , the embodiment of the present application provides a resource scheduling device based on multi-modal fusion IT asset health assessment, as shown in Figure 4 , the device comprises:
[0150] The health assessment module 402 is configured to obtain resource usage data, network connection data and abnormal information data corresponding to each node in the distributed system, and perform health score calculation based on the resource usage data, network connection data and abnormal information data to obtain the health score corresponding to each node.
[0151] The comprehensive evaluation module 404 is configured to obtain the priority of the service adapted by each node, calculate the health score of each node based on the health score corresponding to each node and the priority of the adapted service, and obtain the comprehensive evaluation value corresponding to each node.
[0152] The resource scheduling module 406 is configured to schedule the resources based on the comprehensive evaluation value corresponding to each node.
[0153] The present application provides a resource scheduling device for IT asset health evaluation based on multi-modal fusion. The health score of each node is calculated based on the resource usage data, network connection data and abnormal information data of each node. The comprehensive evaluation value corresponding to each node is calculated based on the health score of each node and the priority of the adapted service. The resources are allocated to appropriate nodes based on the comprehensive evaluation value. The comprehensive evaluation value of multi-modal fusion is calculated, and the resources are scheduled according to the comprehensive evaluation value. The health status of the IT asset is comprehensively and real-timely evaluated. The corresponding scheduling processing is performed according to the comprehensive evaluation, and the stability, reliability and operation automation level of the system are significantly improved.
[0154] In one embodiment, the abnormal information data includes a fault domain risk coefficient, an abnormal log frequency and a historical failure rate. The health evaluation module is further configured to:
[0155] The collection time corresponding to the resource usage data, the network connection data and the abnormal information data is divided into a plurality of time windows. The resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window are obtained.
[0156] The resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate corresponding to each time window are sequentially subjected to abnormality detection, smoothing processing and normalization processing.
[0157] The weights corresponding to the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate are obtained. The normalized resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency and the historical failure rate and the weights corresponding thereto are substituted into the health score calculation formula to obtain the health score corresponding to each node.
[0158] In one embodiment, the comprehensive evaluation module is further configured to:
[0159] The total score value of the health score is calculated based on the health score corresponding to each node and the weight thereof.
[0160] The total score value of the service is calculated based on the priority of the service adapted by each node and the weight thereof.
[0161] The sum of the health score total value of each node and the service total value is taken as a comprehensive evaluation value corresponding to each node.
[0162] In an embodiment, the resource scheduling module is further configured to:
[0163] The node with a comprehensive evaluation value greater than a comprehensive value threshold is taken as a candidate node.
[0164] Based on the affinity scheduling rule, the target resource is allocated to the candidate node, and based on the anti-affinity scheduling rule, the related resource of the target resource is allocated to another different candidate node.
[0165] In an embodiment, the resource scheduling module is further configured to:
[0166] After the target resource is allocated to the candidate node, the resource usage data, network connection data and abnormal information data corresponding to the candidate node after a preset time period are obtained, the health score is calculated based on the resource usage data, network connection data and abnormal information data corresponding to the candidate node after the preset time period, and a new health score corresponding to the candidate node to which the target resource is allocated is obtained.
[0167] If the new health score of the candidate node to which the target resource is allocated is less than a preset first health threshold, the target resource is migrated to another candidate node, and if the new health score of the candidate node to which the target resource is allocated is greater than a preset second health threshold, another resource is allocated to the candidate node to which the target resource is allocated.
[0168] In an embodiment, the resource scheduling device based on the multi-modal fusion IT asset health assessment further comprises:
[0169] The fault domain risk coefficient acquisition module is configured to acquire a fault history record, determine a plurality of fault occurrence regions from the fault history record, determine a first fault number and a fault influence degree corresponding to each fault occurrence region in a first preset time range from the fault history record, take the quotient of the first fault number and the total running time in the first preset time range as a fault occurrence probability for each fault occurrence region, take the product of the fault occurrence probability and the fault influence degree as a fault risk coefficient of each fault occurrence region, and take the sum of the fault risk coefficients of all fault occurrence regions as a fault domain risk coefficient.
[0170] The historical fault rate acquisition module is configured to determine a second fault number in a second preset time range from the fault history record, and take the quotient of the second fault number and the total running time in the second preset time range as a historical fault rate.
[0171] In an embodiment, the resource scheduling device based on the multi-modal fusion IT asset health assessment further comprises:
[0172] The migration information output module is configured to output virtual machine migration information when the health score of any node configured with the virtual machine is less than a preset third health threshold.
[0173] According to an embodiment of the present application, a storage medium is provided, which stores at least one executable instruction, and the computer executable instruction can execute the resource scheduling method for IT asset health assessment based on multi-modal fusion in any method embodiment.
[0174] Figure 5 A structural schematic diagram of a computer device according to an embodiment of the present application is shown, and the specific embodiments of the present application do not limit the specific implementation of the computer device.
[0175] As shown in Figure 5 The computer device can include a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0176] The processor 502, the communications interface 504, and the memory 506 can communicate with each other through the communications bus 508.
[0177] The communications interface 504 is configured to communicate with network elements of other devices, such as clients or other servers.
[0178] The processor 502 is configured to execute the program 510, and specifically can execute related steps in the resource scheduling method for IT asset health assessment based on multi-modal fusion.
[0179] Specifically, the program 510 can include program code, and the program code includes computer operation instructions.
[0180] The processor 502 can be a central processing unit CPU, or an application specific integrated circuit ASIC, or one or more integrated circuits configured to implement embodiments of the present application. The computer device includes one or more processors, which can be processors of the same type, such as one or more CPUs; or can be processors of different types, such as one or more CPUs and one or more ASICs.
[0181] The memory 506 is configured to store the program 510. The memory 506 can include a high-speed RAM memory, and can also include a non-volatile memory, such as at least one disk memory.
[0182] The program 510 can be specifically used to cause the processor 502 to perform the following operations:
[0183] Obtain resource usage data, network connection data and exception information data corresponding to each node in the distributed system, perform health score calculation based on the resource usage data, the network connection data and the exception information data, and obtain a health score corresponding to each node;
[0184] Obtain the priority of the service adapted by each node, perform calculation based on the health score corresponding to each node and the priority of the adapted service, and obtain a comprehensive evaluation value corresponding to each node;
[0185] Perform resource scheduling based on the comprehensive evaluation value corresponding to each node.
[0186] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. In an embodiment, they can be realized by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in an order different from that shown here, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps thereof can be manufactured into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0187] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application, and the protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the spirit and protection scope of the present application, and such modifications or equivalent replacements should also be considered to fall within the protection scope of the present application.
Claims
1. A resource scheduling method for IT asset health assessment based on multimodal fusion, characterized in that: include: Obtain resource usage data, network connection data, and exception information data corresponding to each node in the distributed system, and calculate a health score based on the resource usage data, the network connection data, and the exception information data to obtain a health score corresponding to each node; Obtaining the priority of the service adapted by each node, and calculating based on the health score corresponding to each node and the priority of the adapted service, to obtain a comprehensive evaluation value corresponding to each node; Resources are scheduled based on the comprehensive evaluation value corresponding to each of the nodes.
2. The resource scheduling method for IT asset health assessment based on multimodal fusion according to claim 1, characterized in that: The abnormal information data includes the fault domain risk coefficient, abnormal log frequency and historical failure rate. The health score calculation based on the resource usage data, the network connection data and the abnormal information data is performed to obtain the health score corresponding to each node, including: Divide the collection time corresponding to the resource usage data, the network connection data, and the abnormal information data into multiple time windows, and obtain the resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency, and historical failure rate corresponding to each time window; Perform anomaly detection, smoothing, and normalization on the resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency, and historical failure rate corresponding to each time window; Obtain the weights corresponding to the resource usage data, the network connection data, the fault domain risk coefficient, the abnormal log frequency, and the historical failure rate, and substitute the normalized resource usage data, network connection data, fault domain risk coefficient, abnormal log frequency, historical failure rate, and their corresponding weights into the health score calculation formula to obtain the health score corresponding to each node.
3. The resource scheduling method for IT asset health assessment based on multimodal fusion according to claim 1, characterized in that: The calculation based on the health score corresponding to each node and the priority of the adapted service to obtain the comprehensive evaluation value corresponding to each node includes: Based on the health score and weight corresponding to each node, a total health score is calculated; Calculate the total score of the service based on the priority assigned to the service adapted by each node and its weight; The sum of the total health score and the total business score of each node is used as the comprehensive evaluation value corresponding to each node.
4. The resource scheduling method for IT asset health assessment based on multimodal fusion according to claim 1, characterized in that: The scheduling of resources based on the comprehensive evaluation value corresponding to each of the nodes includes: Nodes with comprehensive evaluation values greater than the comprehensive value threshold are regarded as candidate nodes; Based on affinity scheduling rules, target resources are allocated to candidate nodes, and based on anti-affinity scheduling rules, related resources of the target resources are allocated to other different candidate nodes.
5. The resource scheduling method for IT asset health assessment based on multimodal fusion according to claim 4, characterized in that: After allocating the target resources to the candidate nodes based on the affinity scheduling rule, the resource scheduling method for IT asset health assessment based on multimodal fusion further includes: After allocating the target resource to the candidate node, obtaining resource usage data, network connection data, and abnormal information data corresponding to the candidate node allocated the target resource after a preset period of time, and calculating a health score based on the resource usage data, network connection data, and abnormal information data corresponding to the preset period of time to obtain a new health score corresponding to the candidate node allocated the target resource; If the new health score of the candidate node for allocating the target resource is less than the preset first health threshold, the target resource is migrated to other candidate nodes. If the new health score of the candidate node for allocating the target resource is greater than the preset second health threshold, other resources are allocated to the candidate node for allocating the target resource.
6. The resource scheduling method for IT asset health assessment based on multimodal fusion according to claim 2, characterized in that: The following method is used to obtain the fault domain risk coefficient: Obtaining fault history records, determining multiple fault occurrence areas based on the fault history records, and determining the number of first faults and the degree of fault impact corresponding to each fault occurrence area within a first preset time range from the fault history records; For each fault occurrence area, the quotient of the first fault number and the total operating time within the first preset time range is used as the fault occurrence probability, and the product of the fault occurrence probability and the fault impact degree is used as the fault risk coefficient of each fault occurrence area; The sum of the failure risk coefficients of all failure-occurring areas is taken as the failure domain risk coefficient; The following method is used to obtain the historical failure rate: A second number of faults within a second preset time range is determined from the fault history records, and a quotient of the second number of faults and the total operating time within the second preset time range is used as a historical failure rate.
7. The resource scheduling method for IT asset health assessment based on multimodal fusion according to any one of claims 1 to 6, characterized in that: After calculating the health score based on the resource usage data, the network connection data, and the abnormal information data to obtain the health score corresponding to each node, the resource scheduling method for IT asset health assessment based on multimodal fusion further includes: When the health score of any node configured with a virtual machine is less than a preset third health threshold, virtual machine migration information is output.
8. A resource scheduling device for IT asset health assessment based on multimodal fusion, characterized in that: include: A health assessment module is used to obtain resource usage data, network connection data, and abnormal information data corresponding to each node in the distributed system, and calculate a health score based on the resource usage data, the network connection data, and the abnormal information data to obtain a health score corresponding to each node; A comprehensive evaluation module is used to obtain the priority of the service adapted by each node, and calculate the comprehensive evaluation value corresponding to each node based on the health score corresponding to each node and the priority of the adapted service; The resource scheduling module is used to schedule resources based on the comprehensive evaluation value corresponding to each of the nodes.
9. A storage medium storing at least one executable instruction, characterized in that: The executable instructions enable the processor to execute operations corresponding to the resource scheduling method for IT asset health assessment based on multimodal fusion according to any one of claims 1 to 7.
10. A computer device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to execute operations corresponding to the resource scheduling method for IT asset health assessment based on multimodal fusion as described in any one of claims 1 to 7.
Citation Information
Cited By
Resource scheduling method and device, equipment, storage medium and computer product
CN121644570A
Laboratory resource scheduling method and device, equipment, medium and program product
CN122066182A