Computer power-off protection system
Through the computer power outage protection system, the power status is monitored in real time and the wolf pack algorithm optimization strategy is used to solve the problems of slow decision-making response and waste of resources under power fluctuations, and fast and accurate response and energy-saving management are achieved, which enhances the reliability of data protection.
Patent Information
- Application Number
- CN202510517202.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In the scenarios of power fluctuation monitoring and data security management, the existing technology fails to quickly determine sudden abnormalities, resulting in slow decision-making response, delayed shutdown or continuous operation of non-essential services, and fails to achieve dynamic priority linkage between task load and shutdown policy, resulting in task backlog and resource waste, and insufficient guarantee of key data integrity.
The computer power outage protection system is adopted, including the power supply status monitoring module, the policy decision module, the data protection module and the system shutdown management module. By monitoring the fluctuations of power supply voltage and frequency in real time, the wolf pack algorithm is used to optimize the energy-saving and shutdown strategies, dynamically allocate task priorities, and identify key data for data backup and storage, and gradually shut down the non-critical service server.
It improves the rapidity and accuracy of response decision-making in abnormal power supply fluctuations, optimizes the orderliness of task processing and server resource utilization, enhances the reliability and risk resistance of data protection, and reduces the overall energy consumption level.
Smart Images

Figure CN120449220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of swarm intelligence technology, and in particular to a computer power-off protection system. Background Art
[0002] Swarm intelligence is a field of artificial intelligence technology that has evolved from the collaborative behavior of biological groups in nature. Common research prototypes include ant colonies, bird flocks, fish schools, and wolf packs.
[0003] In the context of power fluctuation monitoring and data security management, existing technologies lack sufficient data collection and analysis granularity for power fluctuation information, failing to quickly identify unexpected anomalies. This results in slow decision-making and response times during emergencies, leading to delayed shutdowns or the continued operation of non-essential services. Furthermore, existing technologies lack dynamic priority linkage between task loads and shutdown policies, which can easily lead to task backlogs and resource waste. Furthermore, the lack of targeted data protection and storage optimization strategies leaves critical data facing insufficient integrity protection, potentially leading to the risk of data loss or corruption. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a computer power-off protection system.
[0005] In order to achieve the above object, the present invention adopts the following technical solution: a computer power failure protection system includes:
[0006] A power status monitoring module monitors the global power status, analyzes power voltage and frequency fluctuations, obtains power stability data, and based on the power stability data, determines whether to activate an emergency response and generates an emergency status determination result;
[0007] A policy decision module, based on the emergency state judgment result, activates the wolf pack algorithm to optimize the energy saving and shutdown strategy and generates a policy optimization plan; applies the policy optimization plan to the real-time task and load management of the server, formulates a priority allocation plan, and generates a hierarchical priority strategy result;
[0008] A data protection module, based on the hierarchical priority strategy results, identifies key data and services, performs data backup and preservation, and obtains a temporary data preservation state; based on the temporary data preservation state, evaluates data integrity and backup speed, optimizes the data preservation process, and generates an optimized data protection result;
[0009] The system shutdown management module controls the gradual shutdown of servers of non-critical services based on the optimized data protection results, implements energy-saving state conversion, and generates energy-saving state adjustment records.
[0010] Preferably, the steps of acquiring the power supply stability data are:
[0011] Extract the instantaneous voltage and frequency values of all sampling points in every 10-second cycle of the power supply, and record the average voltage, maximum voltage, minimum frequency and frequency variance in each cycle to generate a voltage-frequency statistical parameter set;
[0012] According to the voltage frequency statistical parameter set, the inter-cycle stability fluctuation factor is calculated using the following formula:
[0013]
[0014] Among them, R d is the inter-cycle stability volatility factor, is the maximum voltage in the rth cycle, is the average voltage in the rth cycle, is the frequency variance in the rth period, VMT r is the voltage measurement time interval of the rth cycle, VMP r is the mean value of the voltage measurement point in the rth cycle, is the minimum frequency in the rth cycle, and n is the total number of cycles;
[0015] According to the inter-cycle stability fluctuation factor, determine whether the inter-cycle stability fluctuation factor has a mutation greater than 1.5 times the standard fluctuation threshold within three consecutive cycles, record the number of mutations, and generate power supply stability data.
[0016] Preferably, the steps for obtaining the emergency status judgment result are:
[0017] Determine the severity level of the disturbance event based on the inter-cycle stability fluctuation factor and the number of mutations in the power supply stability data, and generate a disturbance level classification result;
[0018] Based on the disturbance level classification result, the disturbance level classification result is compared with the preset emergency response trigger threshold item by item, and whether the disturbance level classification result exceeds the set emergency trigger threshold for two consecutive periods is determined, and a disturbance trigger determination mark is generated;
[0019] Based on the disturbance trigger determination flag, if the disturbance trigger determination flag meets the continuous triggering condition, the emergency state activation command is executed; if the determination flag does not meet the condition, the normal operation mode is maintained and an emergency state determination result is generated.
[0020] Preferably, the steps for obtaining the strategy optimization solution are:
[0021] According to the emergency state judgment result, it is determined whether the current system energy saving and shutdown strategy needs to be re-optimized, and the server operation log is called to establish the initial server group operation state matrix and task type mapping relationship matrix, and generate the wolf pack algorithm startup mark;
[0022] Based on the wolf pack algorithm startup identifier, the server group operation status and task type mapping relationship is called to set the initial position and movement speed of the wolf pack algorithm. According to the current position of each wolf pack individual, the total power consumption and task operation risk corresponding to each server are calculated to form a complete wolf pack individual fitness matrix and generate an initial wolf pack fitness matrix.
[0023] Based on the initial wolf pack fitness matrix, the leader wolf's leading position is updated, and the scout wolf's detection and follower wolf's collaborative position are iteratively calculated. The initial wolf pack fitness matrix corresponding to the individual wolf pack members and the task allocation plan of each server are continuously updated until the initial wolf pack fitness matrix reaches the global convergence standard. The corresponding position of the optimal solution is recorded to generate a strategy optimization plan.
[0024] Preferably, the steps for obtaining the hierarchical priority strategy result are:
[0025] Based on the strategy optimization scheme, the task priority level index is calculated using the following formula:
[0026]
[0027] Among them, P k is the task priority level index of the kth task, a k is the total CPU usage time after the task is submitted, b k The number of instructions contained in the task, STD k The difference between the task submission time and the current timestamp, TQN k Number the task queue, t k is the earliest submitted task number in the same server task set, ψ k is the frequency of occurrence of this task type in the last month, m k The maximum memory space occupied by the kth task;
[0028] Based on the task priority level index, all tasks to be assigned are sorted from high to low according to the task priority level index, and combined with the processing capacity of the server where the task is located, the tasks are assigned to three execution queues of priority, medium priority and low priority to generate a hierarchical priority strategy result.
[0029] Preferably, the steps for obtaining the temporary data storage status are:
[0030] Based on the hierarchical priority strategy results, the task storage adaptation score is calculated using the following formula:
[0031]
[0032] Among them, Z c is the storage adaptation score of the cth data task, P c The task priority level indicator generated by the hierarchical priority strategy result for the cth task, CBW c is the channel bandwidth of the target node of the task, DPC c is the number of data packets included in the task, β c is the cumulative number of times this task type appears on the current node in the last month, α c is the length of the data block corresponding to the task, ω c The source port number when the task is triggered;
[0033] Based on the task storage adaptation score, a task storage adaptation score threshold is set and all data tasks are screened item by item. Data tasks with task storage adaptation scores higher than the threshold are retained and written into the target node storage area according to the data block structure label classification to obtain the temporary data storage status.
[0034] Preferably, the steps for obtaining the optimized data assurance results are:
[0035] Based on the storage success identifier, write timestamp, read backtest check code, node write path number, check code consistency flag and storage adaptation score of each data block in the temporary data storage state, a data assurance assessment original information set is formed;
[0036] Based on the data assurance assessment original information set, the comprehensive integrity response value is calculated using the following formula:
[0037]
[0038] Among them, G p is the comprehensive integrity response value of the pth data block, χ p is the read backtest check code of the pth data block, ζ p The checksum consistency flag of the pth data block, 0 indicates inconsistency, 1 indicates consistency, WRI p is the time interval from writing to first reading of the pth data block, Z c is the storage adaptation score of the cth task, ∈ p is the number of error bits that occur in the first reading of the p-th data block, δ p is the average number of bad blocks detected during the write phase for the p-th data block;
[0039] Based on the comprehensive integrity response value, data groups are screened from all data blocks, and the data blocks are read back and rewritten to generate optimized data assurance results.
[0040] Preferably, the steps of obtaining the energy-saving state adjustment record are:
[0041] According to the comprehensive integrity response value of each server node data block recorded in the optimized data assurance result, the difference between the comprehensive integrity response value and the predetermined critical standard value is compared one by one, and non-critical service servers whose values are below the predetermined critical standard value are identified and marked, and a shutdown list of non-critical service servers is generated;
[0042] Based on the non-critical service server shutdown list, send safety shutdown instructions to the non-critical service servers one by one, monitor the task termination, resource release and process shutdown execution of the servers after receiving the instructions, and generate a server gradual shutdown monitoring log;
[0043] Based on the server gradual shutdown monitoring log, the change value and time point of the server power consumption reduction during the shutdown process are recorded, the energy-saving conversion effect is evaluated, and the energy-saving state adjustment record is generated.
[0044] Compared with the prior art, the advantages and positive effects of the present invention are:
[0045] In the present invention, by real-time monitoring of the power supply voltage and frequency fluctuation status, power supply stability data is extracted, and emergency status judgment results are formed based on the stability data to improve the speed and accuracy of response decisions when the power supply fluctuates abnormally; at the same time, relying on the wolf pack algorithm to execute the energy-saving and shutdown strategy optimization process, a priority dynamic allocation logic between the optimization scheme and the real-time task load is established, so that the orderliness of task processing and the server resource utilization rate are improved simultaneously, and resource idleness and task accumulation are reduced; further based on the task priority, key data and services are accurately identified, the pertinence and real-time nature of the data backup and storage links are strengthened, and the data integrity and backup speed evaluation process are added to enhance the reliability and risk resistance of data protection; on this basis, the data protection results are optimized to further control the gradual and safe shutdown of non-critical service servers. Through the gradual shutdown mechanism, the overall energy consumption level is reduced, and the efficiency of responding to power anomalies and the economy of energy consumption management are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0048] See also Figure 1 The present invention provides a technical solution: a computer power failure protection system includes:
[0049] The power status monitoring module monitors the global power status, analyzes power voltage and frequency fluctuations, obtains power stability data, and based on the power stability data, decides whether to activate emergency response and generates emergency status judgment results;
[0050] The policy decision module uses the wolf pack algorithm to optimize energy-saving and shutdown strategies based on the emergency status judgment results, and generates a policy optimization plan. The policy optimization plan is applied to the server's real-time task and load management, a priority allocation plan is formulated, and a hierarchical priority strategy result is generated.
[0051] The data protection module identifies key data and services based on the hierarchical priority strategy results, performs data backup and preservation, and obtains the temporary data preservation status. Based on the temporary data preservation status, it evaluates data integrity and backup speed, optimizes the data preservation process, and generates optimized data protection results.
[0052] The system shutdown management module controls the gradual shutdown of servers for non-critical services based on the optimized data protection results, implements energy-saving state transitions, and generates energy-saving state adjustment records.
[0053] The steps to obtain power stability data are:
[0054] Extract the instantaneous voltage and frequency values of all sampling points in every 10-second cycle of the power supply, and record the average voltage, maximum voltage, minimum frequency and frequency variance in each cycle to generate a voltage-frequency statistical parameter set;
[0055] According to the voltage-frequency statistical parameter set, the inter-cycle stability fluctuation factor is calculated using the following formula:
[0056]
[0057] Among them, R d is the inter-cycle stability volatility factor, is the maximum voltage in the rth cycle, is the average voltage in the rth cycle, is the frequency variance in the rth period, VMT r is the voltage measurement time interval of the rth cycle, VMP r is the mean value of the voltage measurement point in the rth cycle, is the minimum frequency in the rth cycle, and n is the total number of cycles;
[0058] According to the inter-cycle stability fluctuation factor, determine whether the inter-cycle stability fluctuation factor has a mutation greater than 1.5 times the standard fluctuation threshold within three consecutive cycles, record the number of mutations, and generate power supply stability data.
[0059] Specifically, based on the pre-established cycle definition data during the continuous monitoring process, every 10 seconds is considered an independent sampling cycle in the system. The value acquisition operation is performed sequentially, and the instantaneous voltage and frequency values at all times within the cycle are read. The collected voltage values are between 0V and 24V, and the corresponding acquisition time sequence number and timestamp are recorded. The frequency values are between 40Hz and 60Hz, and the acquisition time sequence number and timestamp are also recorded. Then all collected samples are arranged in chronological order and an index comparison table is constructed. Each index record contains the voltage and frequency values and their corresponding sampling time. Then, all valid records within the same sampling cycle are centrally compared, and the average voltage value, maximum voltage value, minimum frequency value, and frequency variance value within each cycle are calculated. The statistical process scans the valid voltage and frequency data within each cycle one by one to obtain a complete value distribution. Then, the subsequences are divided according to indicators such as average value, extreme value, and variance, and the subsequences are merged and output. The statistical results of each cycle are appended to the cycle summary list, and finally, the voltage and frequency statistical parameter set is formed in the cycle summary list.
[0060] formula: The benefit of the formula is that it performs hierarchical measurement on multiple fluctuation factors of voltage and frequency, and constructs a cross-cycle stability fluctuation factor through information such as the variance, average value, and maximum value of voltage and frequency, thereby quantifying the stability of the system power supply state in a more comprehensive dimension and reflecting local mutation phenomena.
[0061] parameter Acquisition steps: This parameter represents the maximum voltage collected in the rth cycle, and the value range is usually between 220V and 235V. Since the rated voltage of the equipment under constant load conditions is about 220V, the highest voltage value in each cycle is screened out by statistically analyzing the records of the known system running for 600 hours, and these highest values are written into the maximum value sequence in chronological order to form The acquisition process during this period requires a one-time complete scan of the voltage data and the selection of the peak voltage at all points in the cycle. An interval threshold is set to 180V to 240V. By comparing all sampling points with the interval, abnormal points below 180V or above 240V are eliminated and a complete sampling record is ensured for each cycle. Then, the corresponding peak value is selected within the cycle to form a peak value. For example, in the first cycle, 275 voltage sampling points are collected, and the maximum value of 233.6V is obtained at the 205th sampling point.
[0062] parameter Acquisition steps: This parameter represents the average voltage collected in the rth cycle. For example, in the first cycle, the sum of the voltages of 275 sampling points is 60370V.
[0063] parameter Acquisition steps: This parameter represents the frequency variance in the rth period. First, extract all valid frequency data in this period and calculate its average frequency Then each frequency sample value is compared with The square of the difference is accumulated and finally divided by the total number of frequency samples to obtain the variance. The variance result can be further corrected with reference to the frequency fluctuation range under stable power grid conditions. This correction value is called the variance reference coefficient. Here, 1.0 is used as the reference coefficient and the original variance value is directly retained. When constructing the formula, it is recorded as For example, in the first cycle, the mean of all valid frequency points is 49.98 Hz, the number of sampling points is 275, and the sum of the squares of the differences between the frequency and the mean is 0.9485.
[0064] Parameter VMT r Acquisition steps: This parameter represents the voltage measurement time interval of the rth cycle, which mainly reflects the slight delay or advance caused by some jump points in the actual measurement. If the cycle start time is 100.000 seconds and the cycle end time is 110.042 seconds, then VMT r =110.042-100.000=10.042, and VMT1≈10.042 is obtained in the first cycle.
[0065] Parameter VMP r Acquisition steps: This parameter represents the mean value of the voltage measurement points in the rth cycle. For example, in the first cycle, there are 275 sampling points, of which 269 are classified as normal range, 6 sampling points are regarded as burr values and removed, and the remaining 269 sampling values are accumulated to 58672V.
[0066] parameter Acquisition steps: This parameter represents the minimum frequency value in the rth cycle. For example, 275 valid frequency values are collected in the first cycle, among which the minimum value is 49.91Hz.
[0067] Steps for obtaining parameter n: This parameter represents the total number of cycles. For example, in actual testing, if the system runs continuously for 30 seconds and only the first three cycles are analyzed, then n=3.
[0068] Calculation process:
[0069] When executing the actual example, set n=3 and take the following parameters from the previous one:
[0070] Cycle 1:
[0071] VMT1=10.042,VMP1=218.08,
[0072] Cycle 2:
[0073] VMT2=9.999,VMP2=217.95,
[0074] Cycle 3:
[0075] VMT3=10.011,VMP3=218.70,
[0076] The first step is to calculate the absolute value of the fraction in each period:
[0077]
[0078] Take the first cycle as an example:
[0079]
[0080] The second step is to calculate the radical part in the brackets:
[0081]
[0082] Take the first cycle as an example:
[0083]
[0084] (15.4364) 2 ≈238.344,(168.17) 2 ≈28242.9289;
[0085]
[0086] Multiply the two parts together:
[0087] 14.23×168.823≈2405.39;
[0088] Similarly, calculate the corresponding values of the second and third cycles, then find the sum of the three and divide it by n=3 to get the final R d . Only the summary calculations are listed here:
[0089] S1≈2405.39, S2≈2468.77, S3≈2512.51;
[0090]
[0091] The results show that the inter-cycle stability fluctuation factor is approximately 2462.89, indicating that the overall fluctuation of voltage and frequency in this test scenario is still quite obvious. The larger the value, the more prominent the combined impact of the voltage difference and frequency fluctuation between cycles. When the value is lower than 1000, it means that the fluctuation is relatively small, and when it is higher than 2000, it means that the fluctuation is relatively large. At this time, the corresponding threshold will be combined in the subsequent steps to determine whether a higher level response is triggered.
[0092] According to the cross-cycle stability fluctuation factor list obtained above, read its recorded values in three consecutive cycles one by one and compare them with the pre-set standard fluctuation threshold. The standard fluctuation threshold is determined by analyzing the system's operating data within 700 hours and saved as a fixed reference value of 2000. By traversing each value in the cross-cycle stability fluctuation factor list and calculating 1.5 times the reference value equal to 3000, compare the actual fluctuation factor of each cycle to see if it exceeds 3000. Then record the number of times it exceeds 3000 and accumulate the time serial numbers of all mutations that occur in the cycle. Put these time serial numbers into the mutation mark list, and then count the final number of mutations in the mutation mark list. Finally, write the number of mutations into the processing result data set and output the power supply stability data.
[0093] The steps for obtaining the emergency status judgment result are as follows:
[0094] Determine the severity level of the disturbance event based on the inter-cycle stability fluctuation factor and the number of mutations in the power supply stability data, and generate a disturbance level classification result;
[0095] Based on the disturbance level classification results, the disturbance level classification results are compared with the preset emergency response trigger thresholds one by one to determine whether the disturbance level classification results exceed the set emergency trigger thresholds for two consecutive periods, and a disturbance trigger determination flag is generated;
[0096] Based on the disturbance trigger determination flag, if the disturbance trigger determination flag meets the continuous triggering condition, the emergency state activation command is executed; if the determination flag does not meet the condition, the normal operation mode is maintained and an emergency state determination result is generated.
[0097] Specifically, based on the inter-cycle stability fluctuation factor and the number of mutations in the power supply stability data obtained previously, a severity level determination table is first loaded into the system. The determination table allocates multiple level intervals according to the inter-cycle stability fluctuation factor between 0 and 2500, and performs corresponding mapping for the number of mutations between 0 and 10. The data of the inter-cycle stability fluctuation factor and the number of mutations are strictly compared with the thresholds in the determination table. For example, when the inter-cycle stability fluctuation factor is less than 1000 and the number of mutations is less than 3, it is classified as a low-level disturbance event. When the inter-cycle stability fluctuation factor is greater than or equal to 2000 and the number of mutations is more than 7, it is classified as a low-level disturbance event. It is summarized as a high-level disturbance event, and the intermediate values are divided into the intermediate level range. Then, during the comparison process, the specific disturbance event level is recorded in the mapping result list, and the corresponding level code and level description are marked for each record in turn. In order to prevent omissions in the recording process, the system will compare the original value of the cross-period stability fluctuation factor and the numerical range of the current severity level interval for each disturbance event again to confirm that the judgment result is consistent with the threshold range. The entire step will not call any large-scale automated analysis process, but is completed by a single comparison against the judgment table. After all disturbance events are judged according to this process, the disturbance level classification result is finally formed.
[0098] Based on the disturbance level classification results obtained above, the corresponding level code is compared with the benchmark value in the emergency response trigger threshold table. The benchmark value set in the emergency response trigger threshold table is obtained with reference to the frequency of high-level disturbance events recorded in the past 500 hours. The frequency is approximately at least once every 20 hours when the cross-cycle stability fluctuation factor is above 2000 and the number of mutations exceeds 5. Therefore, the threshold of high-level disturbance events is set to a level code equal to 3. When the level code of the current disturbance event is greater than or equal to 3, it is considered to have reached the emergency response trigger threshold. In order to identify whether the threshold is exceeded in two consecutive cycles, it is necessary to compare the disturbance level classification results of adjacent cycles. First, read the level code of the previous cycle and compare it with the level code of the current cycle at the same time. When both the previous cycle and the current cycle reach or exceed level code 3, this situation is considered to be a continuous trigger. Then, after processing all cycle data, the system generates a corresponding disturbance trigger judgment mark.
[0099] Based on the disturbance trigger judgment mark obtained previously, the system will determine whether its value is in a trigger state or a non-trigger state. First, it checks whether the Boolean value of the judgment mark is true. When the Boolean value of the judgment mark is equal to true, the system will send an emergency state activation command and record the current time as the activation time at the same time. Then, it will quickly read the power supply parameters and load conditions. When the voltage distribution is in the range of 200V to 240V, it is considered to be in the general normal range. When it is lower than 200V or higher than 240V, it is considered to be in the range that needs attention. When the Boolean value of the judgment mark is equal to false, the normal operating mode is maintained. The data of the corresponding period is stored in the system as normal state, and is comprehensively checked with the situation of the previous period. The emergency state judgment result of this period is written into the status summary list. When the judgment of all periods is completed, the emergency state judgment result is output.
[0100] The steps to obtain the strategy optimization plan are:
[0101] Based on the emergency status judgment results, determine whether the current system energy saving and shutdown strategies need to be re-optimized, call the server operation log, establish the initial server group operation status matrix and task type mapping relationship matrix, and generate the wolf pack algorithm startup flag;
[0102] Based on the wolf pack algorithm startup flag, the server group operation status and task type mapping relationship is called to set the initial position and movement speed of the wolf pack algorithm. Based on the current location of each wolf pack individual, the total power consumption and task operation risk of each server are calculated to form a complete wolf pack individual fitness matrix and generate the initial wolf pack fitness matrix.
[0103] Based on the initial wolf pack fitness matrix, the leader wolf's leading position is updated, and the scout wolf's detection and follower wolf's collaborative position are iteratively calculated. The initial wolf pack fitness matrix corresponding to the individual wolf pack members and the task allocation plan of each server are continuously updated until the initial wolf pack fitness matrix reaches the global convergence standard. The corresponding position of the optimal solution is recorded and a strategy optimization plan is generated.
[0104] Specifically, according to the emergency status judgment result obtained above, the CPU usage, memory usage, network throughput and other information recorded in the specified server operation log are read, and each record is sorted by timestamp, and these usage values are compared with the pre-established valid range in the system in turn. For example, the CPU usage is compared with the range of 0% to 100%, the memory usage is compared with the range of 0GB to 64GB, and the network throughput is compared with the range of 0MB / s to 1000MB / s. When it is confirmed that all the values do not exceed the corresponding range, the record is classified according to the server ID and written into the corresponding initial state mapping sequence. If it is found that the CPU usage of a server is higher than 80% or the memory usage is close to 64GB, an additional mark is made in the state mapping sequence, and the mark value is included in an item called warning count. The field is accumulated, and then the cumulative number of these tags and the time distribution are combined to make a horizontal comparison of different servers to determine whether there is a large-scale concentrated period that exceeds the set interval. Then, it is associated with the specific task type data according to the server ID. The task type data is obtained by classifying the process information in the server operation log. Here, the processes running on the same server at different times are mapped to the same or similar types of task IDs to form a task type mapping table, and the task ID corresponding to each server ID is marked. In this way, the initial server group operation status matrix and task type mapping relationship matrix are obtained. Finally, a logical judgment is made based on the overall status of each record and the warning count. If it exceeds the specified threshold, it is considered that the server may have a high load and needs to be optimized, and then the wolf pack algorithm start ID is generated.
[0105] Based on the wolf pack algorithm startup flag obtained above, read each record in the mapping relationship between the server group's operating status and the task type, and arrange these records into several sequences according to the server's hardware configuration. Each record contains parameters such as CPU occupancy, memory usage, number of concurrent tasks, and I / O access. Then, internally set the initial position of the wolf pack algorithm as a multidimensional vector, with each dimension corresponding to a server or a type of resource. At the same time, define the movement speed as another multidimensional vector. The reference speed benchmark value is to move 2% to 5% of the CPU occupancy rate per second and 0.5GB to 2GB of memory usage per second. Combined with the high concurrency of the task type, Jing Ze gives greater speed changes, so that fast search can be carried out under high concurrency conditions. Then, according to the multi-dimensional position of each wolf pack individual, the total power consumption generated by the server is calculated one by one. The total power consumption is obtained by monitoring the power and load distribution recorded on the equipment, and comparing whether the power range between 50W and 800W exceeds the preset threshold. Then, statistics are performed on the task type of each server, and the task operation risk is defined as the sum of the process blocking rate and the number of errors. Its range is controlled in the scoring range of 0 to 100, so as to form a complete wolf pack individual fitness matrix and record the initial fitness value. Finally, these calculation results are integrated to obtain the initial wolf pack fitness matrix.
[0106] Based on the initial wolf pack fitness matrix obtained previously, the individual with the highest fitness is considered the leader. The leader's multidimensional vector position is read and its CPU utilization, memory usage, and number of concurrent tasks are compared with those of the other individuals. By setting the number of scout wolves to 3 to 5, the scout's position is updated in small increments in different directions to explore multiple nearby solution regions. The scout's detection results are compared with the leader's position, and the more advantageous coordinate is selected as the local optimal value. The following wolves are then instructed to migrate toward the local optimal value at their own speed. The fitness values of all individuals are recalculated at each iteration. Global convergence is considered achieved when the average change in fitness over three consecutive iterations is less than 0.1. To account for abnormal server load, it is necessary to check during the iterations for any points with CPU utilization exceeding 95% or memory usage exceeding 60GB. If so, the fitness of the individual is penalized to correct the search process. After continuous iterations, the leader's position is finally identified as the optimal solution, and the operating configuration data of each server associated with this position is recorded to obtain the policy optimization solution.
[0107] The steps to obtain the hierarchical priority strategy results are:
[0108] Based on the strategy optimization scheme, the task priority level index is calculated using the following formula:
[0109]
[0110] Among them, P k is the task priority level index of the kth task, a k is the total CPU usage time after the task is submitted, b k The number of instructions contained in the task, STD k The difference between the task submission time and the current timestamp, TQN k Number the task queue, t k is the earliest submitted task number in the same server task set, ψ k is the frequency of occurrence of this task type in the last month, m k The maximum memory space occupied by the k-th task;
[0111] Based on the task priority level index, all tasks to be assigned are sorted from high to low according to the task priority level index, and combined with the processing capacity of the server where the task is located, the tasks are assigned to three execution queues: priority, medium priority and low priority, to generate a hierarchical priority strategy result.
[0112] Specifically, the formula: The benefit of the formula lies in the coupled calculation of multiple aspects of information, such as the CPU usage time and number of instructions after task submission, the task queue number and the earliest submitted task number, and combined with the frequency of similar tasks in the past month and the maximum memory usage of the task, it can comprehensively reflect the resource pressure and urgency of each task in a single indicator, thereby helping to quickly sort the assigned tasks in a multi-tasking environment.
[0113] Parameter a k Obtaining steps: This parameter indicates the total CPU usage time after the kth task is submitted. For example, within 420 hours of device operation, if the CPU usage of a task is counted and it is found that the task's CPU usage time is 5720 seconds, the parameter can be set to 5720.
[0114] Parameter b k Acquisition steps: This parameter indicates the number of instructions contained in the kth task. This is collected by statically analyzing the uploaded executable or script file before the task starts. This analysis analyzes the target file's binary structure or script content, scans each instruction, and counts them. If, in an analysis report, static analysis indicates that the task contains 125,000 valid instructions, this parameter can be set to 125,000.
[0115] Parameter STD kThe acquisition steps of : This parameter represents the difference between the submission time of the kth task and the current timestamp. The acquisition method is to record its submission timestamp when the task starts to queue, and read the current system timestamp at the moment of scheduling calculation. The difference between the two is the actual waiting time. For example, if a task is submitted at T = 200000 seconds and the current system timestamp is T = 200300 seconds, then its STD k = 300. In this example, this parameter can be set to 300.
[0116] Parameter TQN k Obtaining the parameter: This parameter represents the queue number of the kth task when it is queued. An auto-incrementing counter automatically assigns consecutive numbers to new tasks as they are submitted. This number is then retrieved for subsequent calculations when scheduling instructions are issued. For example, if several tasks were submitted within a certain time period and received queue numbers 12, 13, 14, and so on, then if the kth task's queue number is 14, then this parameter can be set to 14.
[0117] Parameter t k Acquisition steps: This parameter represents the earliest submitted task number in the same server task set. By searching all tasks with the same server ID in the queue list, the queue number of the task submitted first is extracted as t k For example, if the current task number is 14, and the earliest task number of the server is 3, then t k =3.
[0118] Parameter ψ k Acquisition steps: This parameter indicates the frequency of occurrence of the task type in the last month. The acquisition process requires scanning the task history records of the task scheduling system in the past 30 days and counting the number of occurrences of the same type of task. The type information can be read from the task description or execution file identifier. The system writes the counted number of times to the database, and then retrieves the number of times during the scheduling calculation and assigns it to ψ k If there are 19 tasks of the same type in the past 30 days, then k =19.
[0119] Parameter m k The acquisition steps of : This parameter represents the maximum memory space occupied by the kth task. The acquisition process needs to be combined with the memory tracking interface provided by the operating system, and the peak memory usage of the task during execution should be recorded regularly. When the task ends or is scheduled, the peak value should be queried again and the statistics should be kept. For example, after system detection and recording, the maximum memory usage of a task is 1024MB, then m can be set k =1024.
[0120] Calculation process:
[0121] The following example gives a specific numerical value input process, let a k =5720,b k =
[0122] 125000,STD k =300,TQN k =14,t k =3,ψ k =19,m k =1024;
[0123] First calculate the fraction part:
[0124]
[0125] Substitute the above parameters into:
[0126] ln(b k +1)=ln(125000+1)≈ln(125001)≈11.736;
[0127] a k ·ln(b k +1)=5720×11.736≈67263.92;
[0128] STD k +1=300+1=301;
[0129]
[0130] Then square the fraction:
[0131] (223.46) 2 ≈49925.54;
[0132] Then calculate the absolute value part:
[0133]
[0134] Submit the parameters one by one:
[0135] TQN k -t k =14-3=11,ψ k +1=19+1=20;
[0136]
[0137] arctan(0.55)≈0.51,
[0138] 0.51×32≈16.32;
[0139] The absolute value is 16.32, and its square root or remainder does not require any additional operation at this point, because the formula only takes the absolute value of this part and then adds it to the square of the previous fraction:
[0140] Add the two parts together and take the square root of the whole:
[0141]
[0142] From this we can get: P k ≈223.47.
[0143] This result shows that in this example scenario, the task priority level index for the kth task is approximately 223.47. A larger value indicates a greater impact on the current task in terms of resource usage, waiting time, and task size. From a scheduling perspective, if the system sets a high priority threshold of 150, this task will be assigned to the high-priority queue, allowing for faster scheduling response in subsequent execution steps.
[0144] Based on the task priority level index obtained above, the system will queue up all tasks to be assigned from large to small according to the index value, and read the numerical record of each task in the index one by one. If it is found that the task priority level index exceeds the preset threshold of 150, these tasks will be temporarily placed in the priority queue and their CPU usage, memory usage, number of instructions and other detailed data will be recorded. At the same time, for tasks with an index lower than 150 but still above 100, the system will put them into the medium priority queue, and determine the sorting method within the queue based on whether the current waiting time STD exceeds 3600 seconds within the medium priority queue. For those tasks with a priority level index less than 100, the system will directly assign them to the low priority queue, and arrange them in order from small to large according to the task queue number size in the low priority queue. In order to prevent To prevent resource contention on the same server at similar times, tasks with the same server ID need to be checked in parallel. If a server has processed all high-priority or medium-priority tasks in the past 60 minutes, the newly added high-priority tasks will be temporarily placed in a marked position called overload check. Then, the system will check whether the current CPU usage of the server exceeds 80% or the memory usage exceeds 16GB. If these preset standards are not exceeded, the task will be moved back to the high-priority queue. If so, it will be temporarily transferred to the medium-priority queue. The resource utilization of the server will be refreshed every 30 seconds until the CPU usage and memory usage return to the safe range again. The task will be put back to the high-priority queue again. Finally, the system will output the hierarchical priority strategy results based on the above queue information.
[0145] The steps to obtain the temporary data storage status are:
[0146] Based on the hierarchical priority strategy results, the task storage adaptation score is calculated using the following formula:
[0147]
[0148] Among them, Z c is the storage adaptation score of the cth data task, P c The task priority level indicator generated by the hierarchical priority strategy result for the cth task, CBW c is the channel bandwidth of the target node of the task, DPC c is the number of data packets included in the task, β c is the cumulative number of times this task type appears on the current node in the last month, α c is the length of the data block corresponding to the task, ω c The source port number when the task is triggered;
[0149] Based on the task storage adaptation score, set the task storage adaptation score threshold and perform item-by-item screening on all data tasks. Retain data tasks with task storage adaptation scores higher than the threshold and write them into the target node storage area according to the data block structure label classification to obtain the temporary data storage status.
[0150] Specifically, the formula: The benefit of the formula is that it comprehensively measures the task priority level indicators in the hierarchical priority strategy results with information such as the channel bandwidth of the target node, the number of data packets, the historical number of occurrences of the task type, the data block length, and the source port number. It balances the network transmission capacity, data size, and the frequency of task occurrence on the node in the same expression, thereby providing a more targeted quantitative basis for the subsequent process of screening suitable storage locations.
[0151] Parameter P c The acquisition steps of P: This parameter represents the task priority level index generated by the cth task in the hierarchical priority strategy result. If the final calculated priority level index of a task is 300, then P c = 300. This parameter can be directly read from the previously confirmed hierarchical priority strategy results and corresponds to the cth data task.
[0152] Parameter CBW c Acquisition steps: This parameter represents the channel bandwidth of the target node of the cth task. The acquisition method requires reading the real-time bandwidth monitoring value on the node communication device and recording the average value of the allocable bandwidth when the node receives this task. In order to ensure the reliability of this bandwidth measurement, it is necessary to set up a scheduled query on the node system, detect the network traffic every 30 seconds and obtain a bandwidth collection array. Then, the measurement value closest to the current task submission time is selected as the CBW.c For example, when a node submits this task, the actual channel bandwidth measured is 850MB / s, and this value is used as CBW. c .
[0153] Parameter DPC c Acquisition steps: This parameter indicates the number of data packets contained in the cth task. The acquisition method is to parse the uploaded data structure by the system, traverse the data stream and count the number of data packets formed after all the segmentation. Each data packet will be assigned an independent number when it is segmented and written into the temporary index table. The total number of data packets is then summarized in the index table. In order to prevent omissions, any empty data packets or error-marked data packets need to be eliminated. The number of valid data packets retained at the end is used as the DPC. c For example, if a file is split into 2000 valid data packets, then DPC c =2000.
[0154] Parameter β c Acquisition steps: This parameter represents the cumulative number of times this task type has appeared on the current node in the past month. The actual frequency of occurrence is obtained by querying all execution records of the same type of task on this node in the past 30 days. It is required to accurately match the task type label with the node identifier, and then count and accumulate the matching records. If the cumulative execution number of a certain type of task is 37 after querying the database, then β can be assigned c =37.
[0155] Parameter α c The acquisition steps of : This parameter indicates the data block length corresponding to the task, which is used to measure the storage capacity occupied by a single data block. When the task starts to be fragmented, the number of bytes of each data block is accumulated and the peak value is selected to reflect the most significant data block length within the task, thereby coping with the differences caused by dynamic fragmentation. For example, when the system is scanning a task in blocks and detects that the largest block is 512MB, α c =512.
[0156] Parameter ω c Acquisition steps: This parameter represents the source port number when the task is triggered. By reading the source port field in the TCP / UDP header when receiving the data stream, the corresponding record is then made in the database according to the task ID. This value indirectly distinguishes the triggering sources of different services or applications through the port number to reflect the load differences that may be caused at the communication level. If the port obtained when parsing the network header of the task is 4567, then ω c =4567.
[0157] Calculation process:
[0158] The following is a sample calculation example, and all parameters are brought in for calculation. Let P c =300,CBW c =850,DPC c =2000,β c =37,α c =512,ω c =4567.
[0159] The first step is to calculate the combined part of the bandwidth and priority level indicators:
[0160]
[0161] log2(2001)≈10.97;
[0162] 300×850=255000,
[0163] The second step is to calculate
[0164]
[0165] Add the two together:
[0166] 23235.37+37.01=23272.38;
[0167] Step 3: Calculate
[0168] ω c +1=4567+1=4568,
[0169] The fourth step is to complete the overall operation:
[0170] (23272.38) 2 =541249136.0644(≈5.4125×10 8 );
[0171] (30.67) 2 ≈940.05;
[0172]
[0173] This result indicates that, based on the values in this example, the task's storage adaptability score is approximately 23266.06. A higher score indicates a better bandwidth environment for the task on the current node and a higher hierarchical priority level. Furthermore, the number of occurrences on this node over the past month has remained at a certain level, indicating that this node is considered more suitable for writing. If the system subsequently sets a storage adaptability score threshold of 20,000, a score greater than 20,000 will be considered eligible for write priority, placing the task within this priority range.
[0174] Based on the task storage adaptation score obtained previously, it is necessary to first read and summarize the hierarchical priority level indicators, target node channel bandwidth, and data packet number corresponding to each data task in the system, and then form a one-to-one matching relationship for this information according to the data task identifier, and refer to the historical records of the past 30 days to retrieve the number of occurrences of the corresponding task type. Subsequently, the program will traverse all the entries of the matching relationship in turn and compare them one by one with a value called the storage adaptation score threshold according to the storage adaptation score calculation results obtained previously. The threshold is set between 20,000 and 30,000, and is determined by analyzing the average node bandwidth allocation and the data packet size range. If the corresponding node average bandwidth is maintained in the range of 500MB / s to 1GB / s, the threshold is initially set to 20,000. When it is monitored that the node bandwidth can reach more than 2GB / s and the cumulative number of tasks of the same type in the past month is When it exceeds 50, the threshold can be raised to 30,000 accordingly, and then the storage adaptation scores of all tasks are compared with the current threshold. When the storage adaptation score exceeds the current threshold, the task is marked as a qualified entry and its data block size is checked to see if it is between 10MB and 1GB. If the data block size exceeds this range, it is marked as an abnormal entry. If it is within the range, it is considered normal and is marked with a record label. Then, hierarchical classification processing is performed according to the structural labels carried by the task. During the classification processing, the key part and the ordinary part of each data block are split according to the size and label respectively and written into the corresponding target node storage area. After all write actions are completed, the successfully written task mark is written into the temporary state sequence and the storage adaptation score is mapped with the corresponding tag number for subsequent review when the node fails or is busy. Finally, the system outputs the temporary data storage state corresponding to the batch of write actions.
[0175] The steps to obtain the optimized data assurance results are:
[0176] Based on the storage success identification, write timestamp, read backtest check code, node write path number, check code consistency flag and storage adaptation score of each data block in the temporary data storage state, the original information set of data assurance assessment is compiled;
[0177] Based on the original information set of data assurance assessment, the comprehensive integrity response value is calculated using the following formula:
[0178]
[0179] Among them, G p is the comprehensive integrity response value of the pth data block, χ p is the read backtest check code of the pth data block, ζ p The checksum consistency flag of the pth data block, 0 indicates inconsistency, 1 indicates consistency, WRI p is the time interval from writing to first reading of the pth data block, Z c is the storage adaptation score of the cth task, ∈ p is the number of error bits that occur in the first reading of the p-th data block, δ p is the average number of bad blocks detected during the write phase for the p-th data block;
[0180] Based on the comprehensive integrity response value, data groups are filtered from all data blocks, and the data blocks are read back and rewritten to generate optimized data assurance results.
[0181] Specifically, based on the storage success identifier, write timestamp, read backtest check code, node write path number, check code consistency flag and storage adaptation score of each data block in the temporary data storage state obtained above, first search the records corresponding to each data block in the system, and match the write timestamp of each data block with the read backtest check code and other information one by one. During the retrieval, the storage success identifier is first checked according to the data block identifier to confirm that the corresponding data block has indeed completed the write operation. Then the node write path number is read to obtain the directory information actually written, and the information is archived in the node storage index table. Then, the check code consistency flag is obtained one by one and compared with the read backtest check code. If the check code consistency flag is 1, the entry is regarded as a normal record. At the same time, if the consistency flag is detected to be 0, it is necessary to check again whether the value in the read backtest check code is in a fixed format. For example, during the reading process, it is found that the check code does not match the hash value generated during the previous write. If a data block is matched, the entry will be marked as a data block to be checked. After all data blocks are confirmed to be in a normal or to-be-checked state, the score value corresponding to each data block is taken out from the storage adaptation score and these score values are attached to the corresponding entries. Then, the entries are grouped and summarized, and the write timestamp and the read timestamp are subtracted using the server time record to obtain the time difference. If the time difference is negative or abnormal, it is marked as a data abnormality entry. All normal entries are then concentrated into a list and all to-be-checked entries are concentrated into another list. Finally, the normal list and the to-be-checked list are summarized and listed as the original information set for data assurance assessment. The above content is recorded one by one in the list according to the data block identifier. If the data block identifier is repeated or missing, it will be eliminated in the summary stage. Finally, a complete original information set for data assurance assessment is obtained that includes all elements such as storage success identifier, write timestamp, read backtest check code, node write path number, check code consistency flag and storage adaptation score.
[0182] formula: The benefit of the formula is that it couples key indicators such as the data block's checksum, consistency flag, time interval from writing to first reading, and the storage adaptation score obtained in the previous step into the same expression, and constructs an adjustment mechanism sensitive to local errors through the difference between the number of error bits and the number of bad blocks. This can measure the comprehensive integrity level of data blocks during storage and reading in a more detailed manner.
[0183] Parameter χ p Acquisition steps: This parameter represents the read backtest check code of the pth data block. The acquisition steps include performing a complete check operation when reading the data block for the first time, recording the hash value or CRC value as χ pIf you are using SHA-256, you can also convert the hash value itself into a decimal integer as a record. In order to make subsequent calculations more intuitive, this large integer is appropriately mapped. The mapping method is to first treat the original hash value as a long integer, and then perform a logarithmic conversion based on the node's built-in mapping rules. p =log 10 (original hash long integer), the final value is used for subsequent operations. For example, if the hash long integer is about 1.2×10 10 , then after mapping χ p ≈10.079.
[0184] Parameter ζ p Acquisition steps: This parameter indicates the consistency flag of the check code of the [th data block. 0 indicates inconsistency and 1 indicates consistency. The range is fixed in {0,1}. The acquisition process is to compare the check code generated for the data block before writing with the check code generated when reading. If they match completely, p =1, otherwise ζ p =0.
[0185] Parameter WRI p Acquisition steps: This parameter represents the time interval from the writing of the [th data block to the first reading. The acquisition process is to mark the timestamp T when the data block is first written. write , the first read operation is marked with a timestamp T read , then calculate WRI p =T read -T write If the system records show that the timestamp of a data block written is 5100 seconds and the timestamp of the first read is 5200 seconds, then WRI p =5200-5100=100.
[0186] Parameter Z c Steps for obtaining : This parameter is the storage adaptation score of the cth task. The value usually ranges from thousands to tens of thousands. For the specific calculation method, please refer to the previous storage adaptation score formula.
[0187] parameter∈ p Acquisition steps: This parameter represents the number of error bits that occurred in the first read of the p-th data block, which depends on the local error status of the storage medium. The acquisition method is to compare the data at the first read with the original data block content, and accumulate 1 in the error counter every time an error bit is found. Then the final value is read from the counter and written into ∈ p For example, if a data block is found to have 12 bit errors after reading, then ∈ p =12.
[0188] Parameter δ pSteps to obtain: This parameter indicates the average number of bad blocks detected during the write phase of the pth data block, which is related to the aging or initial quality of the storage medium. If flash media is used, the bad block table is used to identify available blocks and unavailable blocks before writing. The number of unavailable blocks is then counted and divided by the total number of blocks to obtain the average number of bad blocks. For example, if 200 unavailable blocks are detected on an SSD with a total number of 50,000 blocks, then If the write range of a data block actually involves 10 bad blocks and a total of 2000 physical blocks, then
[0189] Calculation process:
[0190] The following example gives a specific numerical value input process, let χ p =8.716,ζ p =1,WRI p =100,Z c =23266.06,∈ p =12,δ p =0.005.
[0191] The first step is to calculate the radical part: Will but:
[0192] 75.99+1=76.99,
[0193] The second step is to calculate the logarithmic term of the time interval ln(WRI p +1)=ln(100+1)=ln(101)≈4.615;
[0194] Add the two together: 2.96 + 4.615 = 7.575;
[0195] Step 3: Calculate the numerator 7.575×Z c =7.575×23266.06≈176272.795;
[0196] The fourth step is to calculate the denominator |tanh(∈ p -δ p )|+1=|tanh(12-0.005)|+1=|tanh(11.995)|+1;
[0197] tanh(11.995)≈0.99999999999, and after taking the absolute value, it is still approximately 0.99999999999;
[0198] |tanh(11.995)|+1≈1.9999999999≈2.0;
[0199] Step 5: Get the final result
[0200] The result shows that the comprehensive integrity response value of this data block is high, about 88136.4, which means that under the current parameters, the data block is consistent in read and write verification, the combined impact of the number of error bits and the number of bad blocks is low, and the storage adaptation score is high. If the system sets a comprehensive integrity critical value of 50000 in the subsequent stage, then when G p When the value is >50,000, the integrity status is considered relatively good. When the value is less than 50,000, it is considered to need to be read back or rewritten. Therefore, this data block is in the good range.
[0201] Based on the comprehensive integrity response value obtained above, it is necessary to first summarize the response values of all data blocks one by one and record them in association with their identification numbers. When recording, first scan the number of error bits and bad blocks generated by each data block in the previous write and read operation. When the corresponding comprehensive integrity response value exceeds 30,000, it is marked as an entry that has passed the check and matched with the corresponding data block identification. When it is detected that the response value is lower than 30,000 or between 30,000 and 50,000 but the number of error bits contained is greater than 10, the data block is separately included in a list to be checked. After being included, the data block in the list to be checked is read back. The read back operation needs to compare whether the currently read data is consistent with the original storage content. If the number of error bits continues to increase during the read back stage, the read back operation is performed. A rewrite operation is performed. During the rewrite operation, the node is relocated to the same storage node according to the previous storage node write path number and the available space of the node is checked to see if it is greater than 1GB. At the same time, the number of bad blocks is checked to see if it exceeds the node average failure rate set at 0.01. If so, the block is marked as unavailable on the current node and quickly switched to another node for storage. The written-back data block is then verified and compared. If the verification is consistent, the check code consistency flag of the data block is updated and the rewrite completion is displayed in the flag field. Otherwise, it is added to the fault record list for subsequent manual inspection. When all data blocks are read back or rewritten, the final confirmed write path, error bit update status and new check code of all data blocks are recorded in the system to form a sorted result entry, and finally the optimized data assurance result is output.
[0202] The steps for obtaining energy-saving status adjustment records are as follows:
[0203] Based on the comprehensive integrity response value of each server node data block recorded in the optimized data assurance results, the difference between the comprehensive integrity response value and the predetermined critical standard value is compared one by one, and non-critical service servers with values below the predetermined critical standard value are identified and marked, and a shutdown list of non-critical service servers is generated;
[0204] Based on the shutdown list of non-critical service servers, send safe shutdown instructions to non-critical service servers one by one, monitor the task termination, resource release and process shutdown execution of the server after receiving the instructions, and generate a server gradual shutdown monitoring log;
[0205] Based on the server gradual shutdown monitoring log, the change value and time point of the server power consumption reduction during the shutdown process are recorded, the energy-saving conversion effect is evaluated, and the energy-saving status adjustment record is generated.
[0206] Specifically, according to the comprehensive integrity response value of each server node data block in the optimized data assurance result, the corresponding server node identification and the associated data block list are first read in the system, and the comprehensive integrity response value of each data block is compared with the predetermined critical standard value. The setting of the critical standard value refers to the historical detection data in the past 200 hours. After collecting a large number of comprehensive integrity response values of each server node, statistics are performed and their average value and standard deviation are calculated, and then the average value is shifted upward by a safety margin to obtain the critical standard value. If the critical standard value set by the current system is 40000, the comprehensive integrity response value of each server node data block is retrieved one by one. If it is found that the comprehensive integrity response value is lower than 40000, it is determined that the corresponding server may be a non-critical service server, and its node identification is compared with the critical standard value. The corresponding data blocks are recorded in a list, and further verification is conducted to determine whether the CPU utilization of the server has remained below 30% and the memory usage is below 8GB in the past 72 hours. If these low load judgment conditions are met, the server is marked as a shutdown object in the list. The list is then checked for duplicate node identifiers. If the same server has a comprehensive integrity response value below 40,000 in different data blocks and has exhibited low load in the past 72 hours, the server is merged in the list and included in the shutdown list of non-critical service servers. If a node has a comprehensive integrity response value higher than 40,000 in some data blocks or a load record higher than the above range, the node is not included in the shutdown list. After scanning and recording all nodes, the final shutdown list of non-critical service servers is obtained.
[0207] Based on the shutdown list of non-critical service servers obtained earlier, read the unique identifier of each server in the list and check its network connection status. First, use polling to determine whether the server can respond to ping requests normally. Send a safe shutdown instruction to the server that responds normally. The instruction content includes steps such as terminating all currently running processes and releasing occupied memory and file handles. Then, monitor the server to stop the ongoing tasks after receiving the safe shutdown instruction, observe whether the CPU usage gradually drops to the range of 0% to 5%, and check whether the memory usage drops below 2GB. Compare these moment information with timestamps and Record its decline process. If resources are not released in time or the process cannot be closed, trace back to the judgment basis of the server within the scope of non-critical services to confirm whether the server is accidentally stuck or the system is abnormal. If it is found that the data cannot be written back or there is a fault alarm, the server will be temporarily removed from the shutdown sequence and registered as a fault node separately. For servers that continue to shut down safely and successfully, the start and end time of the shutdown operation are recorded, and the process exit sequence and task queue clearing rate are detected. Finally, the process shutdown information of all servers that have been executed and successfully responded to the safe shutdown instruction are arranged in sequence to form a gradual shutdown monitoring log.
[0208] Based on the server gradual shutdown monitoring log obtained above, the power consumption value sampling record is read for each server that has entered the shutdown process. In actual application, the real-time power consumption of these servers under different loads can be obtained through the power detection interface. Sampling is performed every 5 seconds and the power value of the current server is recorded. Then, a shutdown starting point is marked when the shutdown command starts to execute. The sampled power data is continuously read and the difference between it and the average power consumption before the starting point is calculated. If it is found that the power consumption has dropped by more than 100W, the corresponding moment is marked in the monitoring log and the drop rate is recorded. The preset threshold can be set to 150W or more according to the hardware characteristics of the server such as CPU, memory, and motherboard. 200W. The specific value is obtained through testing and verification on servers of the same model. For example, in the test, in an environment with a CPU main frequency of approximately 2.5GHz, the server power consumption is approximately 500W when fully loaded and approximately 300W when idle. If the decrease reaches 200W, it is considered that the shutdown is almost complete. This moment is included in the key recording point, and then the power consumption sampling sequence of all servers is traversed. The time and power consumption difference from the start of shutdown to the time when the power consumption stabilizes in the idle state are summarized. If the server fails to complete the shutdown within the specified time or the power consumption decreases by less than 150W, it is marked as an exception. Finally, the shutdown duration of all servers and the corresponding power consumption changes are uniformly included in the energy-saving conversion evaluation and generate energy-saving status adjustment records.
Claims
1. Computer power failure protection system, characterized in that: The system comprises: A power status monitoring module monitors the global power status, analyzes power voltage and frequency fluctuations, obtains power stability data, and based on the power stability data, determines whether to activate an emergency response and generates an emergency status determination result; A policy decision module, based on the emergency state judgment result, activates the wolf pack algorithm to optimize the energy saving and shutdown strategy and generates a policy optimization plan; applies the policy optimization plan to the real-time task and load management of the server, formulates a priority allocation plan, and generates a hierarchical priority strategy result; A data protection module, based on the hierarchical priority strategy results, identifies key data and services, performs data backup and preservation, and obtains a temporary data preservation state; based on the temporary data preservation state, evaluates data integrity and backup speed, optimizes the data preservation process, and generates an optimized data protection result; The system shutdown management module controls the gradual shutdown of servers of non-critical services based on the optimized data protection results, implements energy-saving state conversion, and generates energy-saving state adjustment records.
2. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the power supply stability data are as follows: Extract the instantaneous voltage and frequency values of all sampling points in every 10-second cycle of the power supply, and record the average voltage, maximum voltage, minimum frequency and frequency variance in each cycle to generate a voltage-frequency statistical parameter set; According to the voltage frequency statistical parameter set, the inter-cycle stability fluctuation factor is calculated using the following formula: Among them, R d is the inter-cycle stability volatility factor, is the maximum voltage in the rth cycle, is the average voltage in the rth cycle, is the frequency variance in the rth period, VMT r is the voltage measurement time interval of the rth cycle, VMP r is the mean value of the voltage measurement point in the rth cycle, is the minimum frequency in the rth cycle, and n is the total number of cycles; According to the inter-cycle stability fluctuation factor, determine whether the inter-cycle stability fluctuation factor has a mutation greater than 1.5 times the standard fluctuation threshold within three consecutive cycles, record the number of mutations, and generate power supply stability data.
3. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the emergency status judgment result are: Determine the severity level of the disturbance event based on the inter-cycle stability fluctuation factor and the number of mutations in the power supply stability data, and generate a disturbance level classification result; Based on the disturbance level classification result, the disturbance level classification result is compared with the preset emergency response trigger threshold item by item, and whether the disturbance level classification result exceeds the set emergency trigger threshold for two consecutive periods is determined, and a disturbance trigger determination mark is generated; Based on the disturbance trigger determination flag, if the disturbance trigger determination flag meets the continuous triggering condition, the emergency state activation command is executed; if the determination flag does not meet the condition, the normal operation mode is maintained and an emergency state determination result is generated.
4. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the strategy optimization solution are: According to the emergency state judgment result, it is determined whether the current system energy saving and shutdown strategy needs to be re-optimized, and the server operation log is called to establish the initial server group operation state matrix and task type mapping relationship matrix, and generate the wolf pack algorithm startup mark; Based on the wolf pack algorithm startup identifier, the server group operation status and task type mapping relationship is called to set the initial position and movement speed of the wolf pack algorithm. According to the current position of each wolf pack individual, the total power consumption and task operation risk corresponding to each server are calculated to form a complete wolf pack individual fitness matrix and generate an initial wolf pack fitness matrix. Based on the initial wolf pack fitness matrix, the leader wolf's leading position is updated, and the scout wolf's detection and follower wolf's collaborative position are iteratively calculated. The initial wolf pack fitness matrix corresponding to the individual wolf pack members and the task allocation plan of each server are continuously updated until the initial wolf pack fitness matrix reaches the global convergence standard. The corresponding position of the optimal solution is recorded to generate a strategy optimization plan.
5. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the hierarchical priority strategy result are: Based on the strategy optimization scheme, the task priority level index is calculated using the following formula: Among them, P k is the task priority level index of the kth task, a k is the total CPU usage time after the task is submitted, b k The number of instructions contained in the task, STD k The difference between the task submission time and the current timestamp, TQN k Number the task queue, t k is the earliest submitted task number in the same server task set, ψ k is the frequency of occurrence of this task type in the last month, m k The maximum memory space occupied by the kth task; Based on the task priority level index, all tasks to be assigned are sorted from high to low according to the task priority level index, and combined with the processing capacity of the server where the task is located, the tasks are assigned to three execution queues of priority, medium priority and low priority to generate a hierarchical priority strategy result.
6. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the temporary data storage status are: Based on the hierarchical priority strategy results, the task storage adaptation score is calculated using the following formula: Among them, Z c is the storage adaptation score of the cth data task, P c The task priority level indicator generated by the hierarchical priority strategy result for the cth task, CBW c is the channel bandwidth of the target node of the task, DPC c is the number of data packets included in the task, β c is the cumulative number of times this task type appears on the current node in the last month, α c is the data block length corresponding to the task, ω c The source port number when the task is triggered; Based on the task storage adaptation score, a task storage adaptation score threshold is set and all data tasks are screened item by item. Data tasks with task storage adaptation scores higher than the threshold are retained and written into the target node storage area according to the data block structure label classification to obtain the temporary data storage status.
7. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the optimized data assurance results are: Based on the storage success identifier, write timestamp, read backtest check code, node write path number, check code consistency flag and storage adaptation score of each data block in the temporary data storage state, a data assurance assessment original information set is formed; Based on the data assurance assessment original information set, the comprehensive integrity response value is calculated using the following formula: Among them, G p is the comprehensive integrity response value of the pth data block, χ p is the read backtest check code of the pth data block, ζ p The checksum consistency flag of the pth data block, 0 indicates inconsistency, 1 indicates consistency, WRI p is the time interval from writing to first reading of the pth data block, Z c is the storage adaptation score of the cth task, ∈ p is the number of error bits that occur in the first reading of the p-th data block, δ p is the average number of bad blocks detected during the write phase for the p-th data block; Based on the comprehensive integrity response value, data groups are screened from all data blocks, and the data blocks are read back and rewritten to generate optimized data assurance results.
8. The computer power-off protection system according to claim 1, characterized in that: The steps for obtaining the energy-saving state adjustment record are: According to the comprehensive integrity response value of each server node data block recorded in the optimized data assurance result, the difference between the comprehensive integrity response value and the predetermined critical standard value is compared one by one, and non-critical service servers whose values are below the predetermined critical standard value are identified and marked, and a shutdown list of non-critical service servers is generated; Based on the non-critical service server shutdown list, send safety shutdown instructions to the non-critical service servers one by one, monitor the task termination, resource release and process shutdown execution of the servers after receiving the instructions, and generate a server gradual shutdown monitoring log; Based on the server gradual shutdown monitoring log, the change value and time point of the server power consumption reduction during the shutdown process are recorded, the energy-saving conversion effect is evaluated, and the energy-saving state adjustment record is generated.
Citation Information
Patent Citations
Computer system protection method after UPS (Uninterrupted Power Supply) outage
CN107544655A
Relay protection device state evaluation method, system, equipment and medium
CN115864644A
Intelligent automatic power-off power management system
CN118413008A