Test method of server system, storage medium, electronic equipment and program product

By automatically identifying and adapting multi-vendor hardware equipment and dynamically determining the test threshold, the server system testing is achieved efficient, accurate and flexible, and the problems of poor test compatibility and time-consuming manual adaptation in the prior art are solved.

CN120144385AActive Publication Date: 2025-06-13INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510629706.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing server system testing methods have poor compatibility and require a lot of manual adaptation work, which increases the testing cost and time.

Method used

By obtaining the device feature vector of the server system to be tested and fuzzy matches with the feature description vector in the driver library, automatic identification and adaptation of hardware devices for multiple vendors is achieved. If the match is successful, the hardware characteristic parameters are collected, and the test threshold of the hardware status is dynamically determined based on the real-time environmental parameters, and dynamically tested.

Benefits of technology

It significantly reduces the time of manual intervention and the complexity of the drive configuration, reduces the cost of testing, improves the efficiency of testing, and improves the accuracy and flexibility of testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144385A_ABST
    Figure CN120144385A_ABST
Patent Text Reader

Abstract

The invention discloses a test method of a server system, a storage medium, electronic equipment and a program product, and relates to the technical field of computers, and the method comprises the following steps: firstly, obtaining an equipment feature vector of a to-be-tested server system; performing fuzzy matching on the equipment feature vector and a feature description vector capable of performing server system testing in a driver library; if the matching is successful, collecting hardware characteristic parameters of the server system to be tested, and dynamically determining a test threshold value of the hardware state of the server system to be tested according to the environment parameters monitored in real time; and then testing the hardware state of the to-be-tested server system based on the hardware characteristic parameters and the dynamically determined test threshold. Compared with the prior art, automatic identification and adaptation of multi-manufacturer hardware equipment are realized, the time of manual intervention and the complexity of drive configuration are remarkably reduced, the test cost is further reduced, the test efficiency is improved, and meanwhile, the test accuracy and flexibility are also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, a storage medium, an electronic device, and a program product. Background Art

[0002] With the rapid development of information technology, data interaction through a server system has become a common business requirement, and the performance, stability, and reliability of the server system are directly related to the normal operation of the business and the user experience. Therefore, comprehensive testing of the server system is a key link to ensure the stable operation of the server and meet the performance standards.

[0003] However, in the face of server hardware from different manufacturers and with different architectures, the compatibility of related server system testing methods is poor, requiring a large amount of manual adaptation work, which increases the testing cost and time. Summary of the Invention

[0004] The present disclosure provides a testing method, a storage medium, an electronic device, and a program product for a server system. Its main purpose is to solve the problem in the related art that the compatibility of the server system testing method is poor, requiring a large amount of manual adaptation work, which increases the testing cost and time.

[0005] In a first aspect, the present application provides a testing method for a server system, including: Obtaining a device feature vector of the server system to be tested; Performing fuzzy matching between the device feature vector and a feature description vector in a driver library that can perform server system testing; If the matching is successful, collecting hardware feature parameters of the server system to be tested, and dynamically determining a test threshold for the hardware state of the server system to be tested according to the real-time monitored environmental parameters; Testing the hardware state of the server system to be tested based on the hardware feature parameters and the dynamically determined test threshold.

[0006] In a second aspect, the present application provides a testing device for a server system, including: An obtaining module, configured to obtain a device feature vector of the server system to be tested; A matching module, configured to perform fuzzy matching between the device feature vector and a feature description vector in a driver library that can perform server system testing; A determining module, configured to, if the matching is successful, collect hardware feature parameters of the server system to be tested, and dynamically determine a test threshold for the hardware state of the server system to be tested according to the real-time monitored environmental parameters; A test module configured to test the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold.

[0007] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method of the first aspect is implemented.

[0009] In a fifth aspect, the present application provides a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0010] The test method, storage medium, electronic device, and program product of the server system provided by the present disclosure, wherein the method includes: first, obtaining the device feature vector of the server system to be tested; then performing fuzzy matching between the device feature vector and the feature description vectors in the driver library that can perform server system tests; if the matching is successful, collecting the hardware characteristic parameters of the server system to be tested, and dynamically determining the test threshold for the hardware status of the server system to be tested according to the real-time monitored environmental parameters; then testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold. Compared with the current related technologies, the present application realizes automatic identification and adaptation of multi-vendor hardware devices by obtaining the device feature vector of the server system to be tested and performing fuzzy matching with the feature description vectors in the driver library, significantly reducing the time of manual intervention and the complexity of driver configuration, thereby reducing the test cost, improving the test efficiency, and dynamically adjusting the test threshold by collecting the hardware characteristic parameters and combining the real-time environmental parameters after the matching is successful, enabling the test strategy to adapt to the current hardware and environmental status, and improving the accuracy and flexibility of the test.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings

[0012] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0013] Figure 1 The flowchart shows a testing method for a server system provided by an embodiment of the present application; Figure 2 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 3 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 4 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 5 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 6 The flowchart shows a testing method for another server system provided by an embodiment of the present application; Figure 7 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 8 The diagram shows a schematic diagram of an example provided by an embodiment of the present application; Figure 9 The diagram shows a schematic structural diagram of a testing device for a server system provided by an embodiment of the present application. Detailed implementation manners

[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0016] To address the technical problems in the related art where the compatibility of the server system testing method is poor, a large amount of manual adaptation work is required, increasing the testing cost and time. This embodiment provides a testing method for a server system, as Figure 1 shown, the method includes the following steps: Step 101, obtain the device feature vector of the server system to be tested.

[0017] Exemplarily, the device feature vector of the server system to be tested can be obtained through information collection, where the device feature vector includes, but is not limited to, key hardware parameters such as CPU model, memory capacity and frequency, storage type and interface, network card configuration, etc.

[0018] For example, in the test process, first, detailed hardware parameters can be collected through the interfaces provided by the server system to be tested, including, but not limited to, processor model and core count, memory capacity and speed, storage device type and its interface specifications, motherboard model, and configurations of expansion cards (such as network cards and graphics cards), etc.

[0019] Step 102: Perform fuzzy matching between the device feature vector and the feature description vectors in the driver library that can be used for server system testing.

[0020] In some examples, by performing fuzzy matching between the feature vector of the device to be tested (including hardware model, interface type, performance parameters, etc.) and the feature description vectors pre-stored in the driver library that can be used for server system testing, the automatic identification and adaptation of multi-vendor hardware devices are realized, thereby effectively improving the compatibility and flexibility of the test system and reducing the adaptation difficulty and test cost caused by device differences.

[0021] Step 103: If the matching is successful, collect the hardware feature parameters of the server system to be tested, and dynamically determine the test threshold for the hardware state of the server system to be tested according to the real-time monitored environmental parameters.

[0022] Exemplarily, after the matching is successful, the system will collect the hardware feature parameters of the server to be tested, and combine the real-time monitored environmental parameters (including parameters such as temperature T, humidity H, and air pressure P), and dynamically calculate the test threshold (Threshold) required for hardware state evaluation through the non-linear mapping model in the adaptive test module. The model expression is: (Formula 1) Where the coefficients , , and the time constant τ are calibrated online through the gradient descent method to realize the real-time optimization and adjustment of the test threshold, thereby improving the accuracy and adaptability of server hardware state evaluation.

[0023] Step 104: Based on the hardware feature parameters and the dynamically determined test threshold, test the hardware state of the server system to be tested.

[0024] Exemplarily, in the process of intelligent fault prediction and optimization of the server system, this embodiment adopts a variety of advanced learning models to improve the reliability and operational stability of the system. First, the Gradient Boosting Decision Tree (GBDT) model is used to predict potential faults in the server system. Based on the historical test data of the server system, the GBDT model can accurately identify potential problems such as hard disk life attenuation, and the prediction accuracy rate reaches 95%. In addition, the intelligent analysis module also uses the Long Short-Term Memory (LSTM) model to perform time series analysis on data from sensors, and real-time monitors and detects hardware anomalies, such as immediate problems like sudden temperature rise and voltage fluctuations.

[0025] By combining a variety of advanced learning models (such as LSTM time series anomaly detection and GBDT fault prediction), not only can potential faults be early warned, but also specific optimization suggestions can be generated based on historical data. For example, adjust the fan speed threshold according to the analysis results to better control the internal temperature of the server, or optimize the power load distribution to reduce the impact of voltage fluctuations. These measures effectively reduce the incidence of hardware failures. It is estimated that the hardware failure risk can be reduced by more than 30%, thus significantly improving the operational stability and efficiency of the server.

[0026] In some examples, an intelligent evaluation system as shown in Figure 2 can be used to improve the test efficiency of the server system. Specifically, the heterogeneous data acquisition module adopts the IPMI / BMC dual protocol stack, supports real-time acquisition and fusion of multi-source sensor data, and ensures the comprehensiveness and accuracy of the data; the dynamic scheduling module uses a multi-core task allocation algorithm based on reinforcement learning to achieve load-balanced task scheduling, improve resource utilization rate and processing speed; the intelligent analysis module is built with an LSTM-GBDT hybrid model architecture for hardware state prediction and fault diagnosis, combines the advantages of time series analysis and decision tree models, and provides accurate prediction and diagnosis results; the adaptive test module includes a dynamic threshold adjustment mechanism for environmental perception and a fuzzy matching function for multi-vendor driver libraries, can dynamically adjust test standards according to the actual operating environment, and is compatible with driver programs of different vendors, enhancing the flexibility and adaptability of the system; the report generation module with blockchain evidence storage is used to generate tamper-proof test reports to ensure the authenticity and credibility of test results; the security audit module realizes full-process monitoring and auditing of system operations through blockchain solidification of operation logs and trusted execution environment protection, and guarantees the security and compliance of the system. The intelligent evaluation system in this embodiment realizes comprehensive performance monitoring and optimization by integrating multiple functional modules, and achieves safe and efficient server system testing.

[0027] Compared with the related technologies, in this embodiment, the device feature vector of the server system to be tested is first obtained; then the device feature vector is fuzzy-matched with the feature description vectors in the driver library that can perform server system tests; if the match is successful, the hardware feature parameters of the server system to be tested are collected, and the test threshold of the hardware state of the server system to be tested is dynamically determined according to the real-time monitored environmental parameters; then, based on the hardware feature parameters and the dynamically determined test threshold, the hardware state of the server system to be tested is tested. By obtaining the device feature vector of the server system to be tested and performing fuzzy matching with the feature description vectors in the driver library, automatic identification and adaptation of multi-vendor hardware devices are achieved, significantly reducing the time of manual intervention and the complexity of driver configuration, thereby reducing the test cost, improving the test efficiency, and collecting hardware feature parameters and dynamically adjusting the test threshold in combination with real-time environmental parameters after the match is successful, enabling the test strategy to adapt to the current hardware and environmental states, and improving the accuracy and flexibility of the test.

[0028] Further, as a refinement and extension of the above embodiment, to specifically illustrate the test process of the server system, optionally, step 102 may specifically include: calculating the cosine similarity between the device feature vector and the feature description vector; if the cosine similarity exceeds the similarity threshold, it is determined that the match is successful.

[0029] For example, in the process of fuzzy matching of the multi-vendor driver library, when the cosine similarity between the calculated device feature vector and the driver description vector exceeds the similarity threshold (such as 0.85), the candidate driver is automatically loaded and the compatibility test is triggered, thereby realizing the dynamic evaluation and adaptation of the compatibility between the device and the driver. It is ensured that even in the case of no exact match, a suitable driver program can be found through similarity analysis, improving the device compatibility and user experience.

[0030] Optionally, the method of this embodiment may specifically further include: obtaining the first hash value of the device identifier of the server system to be tested and the second hash value of the device identifier that can perform server system tests; if there is a second hash value that matches successfully with the first hash value, it is determined that the match is successful; correspondingly, step 102 may specifically further include: if there is no second hash value that matches successfully with the first hash value, the device feature vector is fuzzy-matched with the feature description vectors in the driver library that can perform server system tests.

[0031] Exemplarily, such as Figure 3As shown, during the matching process of the multi-vendor driver library, the adaptive test module can construct an environment mapping model through dynamic threshold adjustment to reflect the changes in the system operating environment in real time. At the same time, the driver fuzzy matching mechanism uses hash comparison and cosine matching technologies to accurately identify and match the performance of different driver programs in a specific environment, so as to achieve comprehensive monitoring and intelligent diagnosis of the server hardware status and ensure the stability and reliability of the system.

[0032] For example, the hash value of the device identifier can be extracted first and compared with the driver fingerprint library to attempt an exact match. Specifically, obtain the first hash value of the device identifier of the server to be tested and compare it with the second hash value in the device identifier library available for testing. If there is a second hash value that matches the first hash value successfully, it is determined that the match is successful, and then the subsequent test process is executed. If no exact match is found, calculate the cosine similarity between the device feature vector and the driver description vector for fuzzy matching.

[0033] Optionally, the method of this embodiment may further specifically include: obtaining the importance degree of the hardware feature parameters. Correspondingly, step 104 may specifically include: based on the hardware feature parameters and the dynamically determined test threshold, and in combination with the importance degree of the hardware feature parameters, test the hardware status of the server system to be tested.

[0034] Exemplarily, during the hardware test process, first obtain the hardware feature parameters of the server system to be tested and evaluate the importance degree of each parameter. Next, during the test phase, test task scheduling can be carried out according to the hardware feature parameters and the dynamically adjusted test threshold, and in combination with the importance level of each parameter, the test priority and test accuracy requirements are dynamically adjusted, so as to achieve key detection of key hardware modules and improve the pertinence and effectiveness of the overall test.

[0035] Optionally, the hardware feature parameters include: hard disk monitoring parameters, frequency domain characteristics of the power supply ripple coefficient, and bit error rate. Correspondingly, the above-mentioned testing of the hardware status of the server system to be tested based on the hardware feature parameters and the dynamically determined test threshold, and in combination with the importance degree of the hardware feature parameters, may specifically include: determining the weights corresponding to the hard disk monitoring parameters, frequency domain characteristics of the power supply ripple coefficient, and bit error rate respectively according to the importance degree; using the weights to perform time series analysis on the server system to be tested to obtain a time series analysis result; comparing the time series analysis result with the test threshold to obtain the test result of the hardware status of the server system to be tested.

[0036] Exemplarily, first, it can be used Figure 4The heterogeneous data acquisition module shown adopts the IPMI / BMC dual protocol stack to support the real-time acquisition and fusion of multi-source sensor data (such as temperature, voltage, fan speed, etc.), and determines the hard disk monitoring parameters, the frequency domain characteristics of the power supply ripple coefficient, and the weight of the bit error rate according to the importance of the hardware characteristic parameters. Among them, the heterogeneous data acquisition module supports the IPMIv2.0 and BMC Redfish API dual protocol stacks and is compatible with the sensor data acquisition of Intel, AMD, and ARM architectures.

[0037] For example, in the data fusion process, different processing methods can be adopted for different types of sensor data, which specifically include: using a sliding window mean filter for voltage data to reduce noise interference and ensure the stability of voltage measurement values; using an exponentially weighted moving average for temperature data to effectively capture the temperature change trend while smoothing short-term fluctuations; for discrete events (such as fan failures), performing timestamp alignment and correlation analysis to accurately track the time point of the event occurrence and its impact range, etc. These data are input into the intelligent analysis module and the adaptive test module for hardware status prediction, fault diagnosis, and dynamic adjustment of test thresholds. At the same time, some data (such as the CPU core load rate, memory occupancy rate, etc.) are also used by the dynamic scheduling module to construct a core load matrix to support load balancing task scheduling decisions. Sensor data can also be used to construct a multi-dimensional feature mapping model to dynamically adjust the environment-sensitive test threshold. For example, the sensor trigger condition is corrected according to the temperature change, thereby improving the accuracy and adaptability of the test. In this way, the heterogeneous data acquisition module provides basic data support for the entire intelligent evaluation system.

[0038] Exemplarily, the collected sensor data is input into the LSTM-GBDT hybrid model architecture of the intelligent analysis module as shown in Figure 5 and can be used for hardware status prediction and fault diagnosis. These data are analyzed by the LSTM network for time series analysis to detect hardware anomalies in real time, and combined with the GBDT model to predict potential faults, such as hard disk life attenuation, etc.

[0039] In some examples, the input dimension of the LSTM network in the LSTM-GBDT hybrid model architecture is [time step × number of features]. The time step is dynamically configured from 12 to 72 hours according to the hardware type to meet the requirements of different hardware characteristics. The GBDT model adopts a feature importance weighting mechanism, and the key features include the dynamic change rate of the hard disk SMART parameters (dynamic change rate = (current value - reference value) / running hours), the frequency domain characteristics of the power supply ripple coefficient (extracting the energy ratio in the 1kHz~10MHz frequency band through the fast Fourier transform FFT), and the second derivative of the PCIe bit error rate.

[0040] In some examples, weighted hardware feature parameters can be utilized to perform time series analysis through a method suitable for processing time series data (such as an LSTM network). After completing the time series analysis, the obtained results are compared with test thresholds dynamically adjusted based on environmental perception. These thresholds can be set according to historical data and environmental operating conditions and are used to determine whether the hardware status is normal, thereby obtaining the hardware status test results of the server system to be tested, identifying possible risk components or confirming the health status of the system, so as to improve the stability and reliability of the system.

[0041] Optionally, the method of this embodiment may specifically further include: determining idle processor cores according to the workload of the server system to be tested; correspondingly, step 104 may specifically further include: based on the hardware feature parameters and the dynamically determined test thresholds, performing the hardware test task of the server system to be tested on the idle processor cores.

[0042] Exemplarily, according to the workload of the server system to be tested, the currently idle processor cores can be dynamically identified, and based on the hardware feature parameters (such as CPU frequency, cache configuration, instruction set support, etc.) and the test thresholds adjusted in real time, the corresponding hardware test tasks are performed on the selected idle cores. This can not only improve the resource utilization rate and parallel processing ability of the test process, but also achieve adaptive scheduling according to the characteristics of different hardware platforms, thereby improving the test efficiency and accuracy.

[0043] As a refinement of this embodiment, in the related steps of testing the hardware status of the server system to be tested based on the hardware feature parameters and the dynamically determined test thresholds, the following methods can be used but are not limited to dynamically multi-core scheduling the hardware test tasks of the server system to be tested, such as Figure 6 shown Figure 6 is a schematic flow chart of a test method for a server system provided by an embodiment of the present disclosure, including: Step 201, splitting the hardware test task into a set of subtasks that can be processed in parallel, and the granularity of the set of subtasks is dynamically determined by the task type and resource occupancy rate of the hardware test task.

[0044] Exemplarily, as Figure 7 shown, the dynamic scheduling module can work together through a task decomposition unit, a resource monitoring unit, a scheduling strategy unit, and a shared memory communication unit to achieve efficient task management and resource allocation. Among them, the task decomposition unit is responsible for decomposing complex tasks into a set of subtasks {T1, T2,..., Tn} that can be processed in parallel.

[0045] Step 202, identifying idle processor cores by real-time monitoring the resource usage of the server system to be tested.

[0046] Exemplarily, the resource monitoring unit in the dynamic scheduling module monitors the system resource status in real time to ensure the effective utilization of resources.

[0047] Optionally, step 202 may specifically include: determining the load levels of multiple processor cores by collecting the load rates, memory occupancy rates, and input / output data volumes of multiple processor cores in real time; and determining the idle processor cores based on the load levels.

[0048] Exemplarily, by collecting key performance indicators such as the load rates, memory occupancy rates, and input / output data volumes (I / O throughput) of each processor core (CPU core) in real time, a core load matrix L = [l1, l2,..., lm] in a unified format is constructed. The load rate of the CPU core reflects the busy degree of each processor core in executing tasks, and the higher the value, the heavier the task processing burden on the core; the memory occupancy rate represents the percentage of the used memory volume in the total memory capacity in the system, reflecting the tightness of the memory resources; and the I / O throughput measures the speed and efficiency of input / output operations, including disk read / write and network transmission, representing the system's data exchange ability.

[0049] For example, each element li ∈ [0, 100%] in the core load matrix represents the usage ratio of the corresponding CPU core. Combining the memory and I / O status can comprehensively reflect the real-time load situation of the system resources and can be reflected as the load levels of multiple processor cores. By analyzing the core load matrix, the running status of the server can be dynamically grasped, providing data support for optimization decisions such as task scheduling, load balancing, bottleneck prediction, and fault warning, thereby improving the running efficiency and stability of the system.

[0050] Step 203: Dynamically allocate the hardware test tasks of the server system to be tested to the idle processor cores based on the subtask set.

[0051] Exemplarily, the scheduling policy unit in the dynamic scheduling module adopts an improved Q-learning algorithm. By defining the state space, adjusting the decay exploration rate, and handling task preemption, the task scheduling policy is optimized to achieve the purpose of load balancing and improving system performance; the shared memory communication unit ensures the rapid exchange and synchronization of information between units. A double-buffer mechanism can be used to achieve cross-core data exchange, and the buffer size is dynamically allocated according to the task data volume to ensure the coordination and consistency of the entire scheduling process.

[0052] In some examples, the reward function of the improved Q-learning algorithm is defined as: (Formula Two) Among them, α, β, and γ are dynamically adjusted coefficients, which are optimized by analyzing the moving window mean of historical task completion times to improve the real-time performance of scheduling decisions and resource utilization efficiency.

[0053] Optionally, step 203 may specifically include: determining a task scheduling decision based on an improved reinforcement learning algorithm according to the state space and decay exploration rate of the server system to be tested; dynamically allocating the hardware test tasks of the server system to be tested to idle processor cores for execution based on the task scheduling decision and the subtask set.

[0054] In some examples, the state space S in the improved Q-learning algorithm can be defined as the set of CPU core load levels {low, medium, high} to represent the resource usage of each core currently; the action space A is defined as the set of instructions for allocating tasks to target cores, indicating the task scheduling decisions that the agent can take. In terms of policy exploration, the algorithm uses an exploration rate ε based on count decay: (Formula Three) where N is the number of tasks that have been successfully allocated, so that the exploration rate gradually decreases with the accumulation of task scheduling experience, thus achieving a smooth transition from exploration to exploitation.

[0055] Optionally, the method of this embodiment may specifically further include: when a target hardware test task with a priority higher than the priority threshold is detected, triggering a preemption mechanism for the target hardware test task, saving the currently processed hardware test task to temporary storage, and re-executing the task dynamic allocation process of the currently processed hardware test task.

[0056] Exemplarily, when the system detects the arrival of a high-priority task, the dynamic scheduling module will immediately trigger the task preemption mechanism. First, interrupt the currently executing low-priority task and place the high-priority task at the front of the execution queue for priority processing; at the same time, the system will save the execution environment (such as status information, data context, etc.) of the interrupted task to a temporary storage area to ensure that it can be accurately resumed later. After the task switch is completed, the improved Q-learning algorithm will re-evaluate and calculate the optimal scheduling strategy based on the current system state (including CPU core load, task queue length, etc.) to ensure the quick response and efficient execution of high-priority tasks. Then the system updates the value of the corresponding state-action pair in the Q table according to the new strategy to reflect the expected benefits of the latest scheduling decision, thereby continuously optimizing the allocation logic of future tasks. The Q table records the reward values obtained by taking different actions in different states and is an important basis for the reinforcement learning algorithm to make decisions. After the high-priority task is executed, the system automatically loads the previously saved low-priority task environment and resumes execution from the breakpoint to ensure the continuity and data integrity of task processing. Optionally, the method of this embodiment may specifically further include: aggregating the test data corresponding to the server system to be tested into a leaf node set; performing layer-by-layer hash calculation on the test data in the leaf node set to generate a root hash value; and writing the root hash value into the Ethereum blockchain.

[0057] In some examples, such as Figure 8 In the blockchain evidence storage module in [], to achieve secure blockchain storage of test data, the test data of the server system to be tested can be aggregated into a leaf node set, and layer-by-layer hash calculation can be performed on these data according to the Merkle tree structure to generate the final root hash value; this root hash value is written into the Ethereum blockchain through a smart contract to ensure that the data is tamper-proof and traceable. At the same time, the generated test report contains multiple verifiable fields: one is the timestamp chain, which is generated by sequentially concatenating the hash values of the execution times of each test module and is used to verify the time integrity of the test process; the other is the digital signature, which uses the national cryptography SM2 algorithm to sign the report summary to ensure the authenticity and legality of the report source.

[0058] Optionally, the method of this embodiment may specifically further include: packing the test operation log blocks of the server system to be tested in a trusted execution environment based on a preset time period, and verifying the validity of the test operation log blocks through a consensus mechanism, where the session key in the trusted execution environment is dynamically generated by a key derivation function.

[0059] In some examples, such as Figure 8 In the security audit module in [], based on a preset time period, the test operation logs of the server system to be tested are regularly packed into blocks in a trusted execution environment, and the collected log data is organized into a new block according to a preset time period (such as every 10 seconds), and the validity of each block is verified by using an improved Practical Byzantine Fault Tolerance (PBFT) consensus mechanism to ensure the accuracy and integrity of the operation logs. The trusted execution environment adopts hardware-enhanced security technologies (such as Intel SGX technology) to ensure the security of the processing process, where the root key is generated inside the memory area (SGX Enclave) and never leaks, further enhancing data protection. The session key is dynamically generated by a Key Derivation Function (KDF) according to specific requirements to maintain the security and flexibility of communication. This ensures the security and reliability of the blockchain solidification process of the test operation logs, and at the same time meets the high-level key management requirements, thus effectively preventing data tampering and unauthorized access.

[0060] In some embodiments, the heterogeneous data acquisition module integrates the IPMI / BMC dual protocol stack, supports real-time acquisition and fusion of multi-source sensor data, uses Python to implement multi-threaded asynchronous data acquisition, is compatible with both IPMI v2.0 and BMC Redfish API specifications, and writes an adapter middle layer to adapt to the hardware characteristics of different manufacturers and different series.

[0061] In some embodiments, the dynamic scheduling module adopts a multi-core task allocation algorithm based on reinforcement learning to achieve load-balanced task scheduling. A core load matrix L = [l, l,..., l] is constructed, where l ∈ [0, 100%]. The overall system load high threshold is set to 70%. Whether it is overloaded is comprehensively judged according to the load rate of each core. The improved Q-learning algorithm is used to dynamically allocate tasks to idle CPU cores, and the timeout time and running times of tasks are limited to prevent resource deadlocks. The scheduling strategy is designed as follows: The state space S is defined as the set of load levels of CPU cores {low, medium, high}, and the action space A is the instruction to allocate tasks to the target core. Here, the state s ∈ S and the action a ∈ A. The exploration rate ε randomly selects actions according to the formula. The code block represents the specific action content. When a high-priority task is detected, a task preemption mechanism is triggered, and the current task environment is saved to temporary storage and the search process is restarted.

[0062] In some embodiments, the intelligent analysis module has a built-in LSTM-GBDT hybrid model architecture for hardware state prediction and fault diagnosis. The LSTM network is used for sequence prediction, and the prediction is updated on a 24-hour cycle. In the diagnosis process, multi-dimensional data such as temperature, voltage, and fan speed are jointly analyzed. A 13-dimensional covariance function is designed. The sample data set is divided into a training set, a test set, and a validation set according to the ratio of 2:8:2. The outlier is calculated based on the average value and standard deviation of the latest 24-hour sampling.

[0063] In some embodiments, the adaptive testing module includes a dynamic threshold adjustment mechanism for environmental perception and a fuzzy matching function for multi-vendor driver libraries. A multi-dimensional feature mapping model F(X) is constructed, the feature space X ∈ R, and the target space Y ∈ R. The nearest neighbor algorithm is used for clustering fitting, and the neighborhood radius R = 2 is set. In the testing stage, the difference between the real sample point and the model estimated output is calculated as the loss function L(F), and the fitting parameters C, F, and W are continuously iteratively optimized. The threshold is divided into intervals according to the least squares method result, and the statistical probability P is used as the basis for weighting. For the access of new devices, the softmax regression output form is used to train the classifier.

[0064] In some embodiments, a report generation module for blockchain evidence storage is used to generate tamper-proof test reports; a security audit module is used to implement blockchain solidification of operation logs and protection of the trusted execution environment, perform offline backup of the root key, and achieve software-level security protection by comparing snapshots through static and dynamic analysis of the current program control flow. Blockchain technology is used to record sensitive information such as environmental parameters, system software and hardware versions, and root keys to ensure the accuracy and integrity of the information.

[0065] Compared with the current related technologies, this embodiment improves the test efficiency, fault prediction accuracy, and cross-platform compatibility of the server system by integrating dynamic multi-core scheduling, intelligent analysis, and adaptive testing mechanisms. Through a dynamic scheduling strategy based on reinforcement learning, tasks are intelligently allocated to idle CPU cores to achieve load balancing and optimal resource utilization; the intelligent analysis module combines the LSTM network to perform time-series modeling on sensor data, real-time detect anomalies (such as sudden temperature rise, voltage fluctuation), and introduces the GBDT model to predict potential faults such as hard disk life attenuation and power supply stability with high accuracy (up to 95%); the adaptive testing module dynamically adjusts the test threshold according to environmental parameters (such as temperature, humidity, air pressure), and integrates multi-vendor driver libraries, and uses a hash comparison and cosine similarity matching mechanism to achieve rapid adaptation and compatibility verification of driver programs. This embodiment forms a technical closed-loop from multiple levels such as system architecture, scheduling algorithm, machine learning model application to adaptive testing method, has strong innovation and protectability, and is applicable to the intelligent testing and operation and maintenance scenarios of high-reliability server systems.

[0066] An embodiment of the present application also provides a test device for a server system, as Figure 1 a specific implementation of the method shown, as Figure 9 shown, the device includes: an acquisition module 31, a matching module 32, a determination module 33, and a test module 34.

[0067] The acquisition module 31 is configured to acquire the device feature vector of the server system to be tested; The matching module 32 is configured to perform fuzzy matching between the device feature vector and the feature description vector that can perform server system testing in the driver library; The determination module 33 is configured to collect the hardware feature parameters of the server system to be tested if the matching is successful, and dynamically determine the test threshold of the hardware state of the server system to be tested according to the real-time monitored environmental parameters; The test module 34 is configured to test the hardware state of the server system to be tested based on the hardware feature parameters and the dynamically determined test threshold.

[0068] In some examples of this embodiment, the matching module 32 is specifically configured to calculate the cosine similarity between the device feature vector and the feature description vector; if the cosine similarity exceeds the similarity threshold, it is determined that the matching is successful.

[0069] In some examples of this embodiment, obtain the first hash value of the device identifier of the to-be-tested server system and the second hash value of the device identifier of the device that can perform server system testing; if there is a second hash value that matches successfully with the first hash value, it is determined that the matching is successful; correspondingly, the matching module 32 is specifically further configured to, if there is no second hash value that matches successfully with the first hash value, perform fuzzy matching between the device feature vector and the feature description vectors in the driver library that can perform server system testing.

[0070] In some examples of this embodiment, obtain the importance level of the hardware feature parameters; correspondingly, the testing module 34 is specifically configured to test the hardware state of the to-be-tested server system based on the hardware feature parameters and the dynamically determined testing threshold, and in combination with the importance level of the hardware feature parameters.

[0071] In some examples of this embodiment, the hardware feature parameters include: hard disk monitoring parameters, frequency domain features of the power supply ripple coefficient, and bit error rate; correspondingly, the testing module 34 is specifically further configured to determine the weights corresponding to the hard disk monitoring parameters, the frequency domain features of the power supply ripple coefficient, and the bit error rate respectively according to the importance level; perform timing analysis on the to-be-tested server system using the weights to obtain a timing analysis result; compare the timing analysis result with the testing threshold to obtain the test result of the hardware state of the to-be-tested server system.

[0072] In some examples of this embodiment, determine the idle processor cores according to the workload situation of the to-be-tested server system; correspondingly, the testing module 34 is specifically further configured to perform the hardware test task of the to-be-tested server system on the idle processor cores based on the hardware feature parameters and the dynamically determined testing threshold.

[0073] In some examples of this embodiment, the testing module 34 is specifically further configured to split the hardware test task into a set of subtasks that can be processed in parallel, and the granularity of the set of subtasks is dynamically determined by the task type and resource occupancy rate of the hardware test task; identify the idle processor cores by real-time monitoring the resource usage situation of the to-be-tested server system; dynamically allocate the hardware test task of the to-be-tested server system to the idle processor cores for execution based on the set of subtasks.

[0074] In some examples of this embodiment, the test module 34 is specifically further configured to determine the load levels of the multiple processor cores by collecting in real time the load rates, memory occupancy rates, and input / output data volumes of the multiple processor cores; and determine the idle processor cores based on the load levels.

[0075] In some examples of this embodiment, the test module 34 is specifically further configured to determine a task scheduling decision based on an improved reinforcement learning algorithm according to the state space and decay exploration rate of the server system to be tested; and dynamically allocate the hardware test tasks of the server system to be tested to the idle processor cores for execution based on the task scheduling decision and the subtask set.

[0076] In some examples of this embodiment, the test module 34 is specifically further configured to trigger a preemption mechanism for the target hardware test task when detecting a target hardware test task with a priority higher than the priority threshold, save the currently processed hardware test task to temporary storage, and re-execute the task dynamic allocation process for the currently processed hardware test task.

[0077] In some examples of this embodiment, the test module 34 is specifically further configured to summarize the test data corresponding to the server system to be tested into a leaf node set; perform layer-by-layer hashing calculations on the test data in the leaf node set to generate a root hash value; and write the root hash value into the Ethereum blockchain.

[0078] In some examples of this embodiment, the test module 34 is specifically further configured to package the test operation log blocks of the server system to be tested in a trusted execution environment based on a preset time period, and verify the validity of the test operation log blocks through a consensus mechanism, where the session key in the trusted execution environment is dynamically generated through a key derivation function.

[0079] It should be noted that for other corresponding descriptions of each functional unit involved in the test device for a server system provided in this embodiment, reference can be made to the corresponding descriptions in Figure 1 and details are not described herein again.

[0080] Based on the methods as shown in Figure 1 and Figure 6 correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the methods as shown in Figure 1 and Figure 6 are implemented.

[0081] Based on the methods as shown in Figure 1 and Figure 6The method described above, correspondingly, this embodiment also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it implements the method as described above Figure 1 and Figure 6 shown.

[0082] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0083] Based on the method as described above Figure 1 and Figure 6 shown, and Figure 9 the virtual device embodiment shown, in order to achieve the above object, this embodiment of this application also provides an electronic device, such as a personal computer, server. This device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above Figure 1 and Figure 6 shown.

[0084] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0085] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or arrange different components.

[0086] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of an information processing program and other software and / or programs. The network communication module is used to implement communication between the components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current related technologies, this embodiment realizes the intelligent decomposition and load balancing distribution of hardware test tasks through the multi-core parallel computing framework and reinforcement learning algorithm of the dynamic scheduling module, and uses shared memory or Unix domain sockets to achieve low-latency communication, improving the parallel test efficiency of thousands of servers by more than 40% and significantly shortening the test cycle. The intelligent analysis module integrates the LSTM time series modeling and the GBDT fault prediction model. The former detects anomalies such as sudden temperature rise and voltage fluctuations in real time, and the latter has an accuracy rate of up to 95% in predicting potential faults such as hard disk life attenuation, and generates optimization suggestions (such as adjusting the fan threshold and optimizing the power load) based on historical data, effectively reducing the hardware failure rate by more than 30% and improving the system stability. The adaptive test module improves the test condition matching degree through the dynamic threshold adjustment mechanism driven by environmental parameters, and at the same time integrates multi-vendor driver libraries to support plug-and-play testing of mainstream interfaces such as PCIe and NVMe. The compatibility coverage rate is increased by 30%, and the adaptive matching time of the test script is shortened by 50%, greatly reducing the cross-platform adaptation cost, and overall realizing an efficient, intelligent and compatible server hardware test system.

[0088] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0089] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for testing a server system, characterized in that: include: Obtaining a device feature vector of the server system to be tested; Fuzzy matching the device feature vector with a feature description vector in a driver library that can be used for server system testing; If the match is successful, the hardware characteristic parameters of the server system to be tested are collected, and the test threshold of the hardware status of the server system to be tested is dynamically determined according to the environmental parameters monitored in real time; Based on the hardware characteristic parameters and the dynamically determined test threshold, the hardware status of the server system to be tested is tested.

2. The method according to claim 1, characterized in that The fuzzy matching of the device feature vector with a feature description vector in a driver library that can be used for server system testing includes: By calculating the cosine similarity between the device feature vector and the feature description vector; If the cosine similarity exceeds the similarity threshold, it is determined that the match is successful.

3. The method according to claim 1, characterized in that Before fuzzily matching the device feature vector with a feature description vector in a driver library that can be used for server system testing, the method further includes: Obtaining a first hash value of a device identifier of the server system to be tested and a second hash value of a device identifier capable of performing server system testing; If there is a second hash value that successfully matches the first hash value, then it is determined that the match is successful; The fuzzy matching of the device feature vector with a feature description vector in a driver library that can be used for server system testing includes: If there is no second hash value that successfully matches the first hash value, the device feature vector is fuzzily matched with a feature description vector in a driver library that can be used for server system testing.

4. The method according to claim 1, characterized in that: Before testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold, the method further includes: Obtaining the importance of the hardware characteristic parameters; The step of testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold comprises: Based on the hardware characteristic parameters and the dynamically determined test threshold, and in combination with the importance of the hardware characteristic parameters, the hardware status of the server system to be tested is tested.

5. The method according to claim 4, characterized in that The hardware characteristic parameters include: hard disk monitoring parameters, frequency domain characteristics of power supply ripple coefficient and bit error rate; The testing of the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold value and in combination with the importance of the hardware characteristic parameters includes: Determine, according to the importance, weights corresponding to the hard disk monitoring parameter, the frequency domain characteristics of the power supply ripple coefficient, and the bit error rate, respectively; Performing a timing analysis on the server system to be tested using the weight to obtain a timing analysis result; The timing analysis result is compared with the test threshold to obtain a test result of the hardware status of the server system to be tested.

6. The method according to claim 1, characterized in that Before testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold, the method further includes: Determining idle processor cores according to the workload of the server system to be tested; The step of testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold comprises: Based on the hardware characteristic parameters and the dynamically determined test threshold, the hardware test task of the server system to be tested is executed on the idle processor core.

7. The method according to claim 6, characterized in that The step of determining an idle processor core according to the workload of the server system to be tested includes: Splitting the hardware test task into a subtask set that can be processed in parallel, wherein the granularity of the subtask set is dynamically determined by the task type and resource occupancy rate of the hardware test task; Identify the idle processor core by real-time monitoring of resource usage of the server system to be tested; The step of executing the hardware test task of the server system to be tested on the idle processor core based on the hardware characteristic parameter and the dynamically determined test threshold comprises: Based on the subtask set, the hardware test task of the server system to be tested is dynamically allocated to the idle processor core for execution.

8. The method according to claim 7, characterized in that The step of identifying the idle processor core by real-time monitoring of resource usage of the server system to be tested includes: Determine the load levels of the multiple processor cores by collecting the load rates, memory occupancy rates, and input and output data volumes of the multiple processor cores in real time; Based on the load level, the idle processor core is determined.

9. The method according to claim 7, characterized in that: The dynamically allocating the hardware test task of the server system to be tested to the idle processor core for execution based on the subtask set includes: Based on the improved reinforcement learning algorithm, the task scheduling decision is determined according to the state space and the decay exploration rate of the server system to be tested; Based on the task scheduling decision and the subtask set, the hardware test task of the server system to be tested is dynamically allocated to the idle processor core for execution.

10. The method according to claim 7, characterized in that The method further comprises: When a target hardware test task with a priority higher than the priority threshold is detected, the preemption mechanism of the target hardware test task is triggered, and the hardware test task currently being processed is saved to temporary storage, and the task dynamic allocation process of the hardware test task currently being processed is re-executed.

11. The method according to claim 1, characterized in that: After testing the hardware status of the server system to be tested based on the hardware characteristic parameters and the dynamically determined test threshold, the method further includes: Aggregate the test data corresponding to the server system to be tested into a leaf node set; Performing layer-by-layer hash calculation on the test data in the leaf node set to generate a tree root hash value; Write the tree root hash value to the Ethereum blockchain.

12. The method according to claim 1, characterized in that The method further comprises: The test operation log block of the server system under test is packaged in a trusted execution environment based on a preset time period, and the validity of the test operation log block is verified through a consensus mechanism. The session key in the trusted execution environment is dynamically generated through a key derivation function.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

14. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product having a computer program stored thereon, characterized in that: When the computer program product is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Testing method, device and system

    CN108763003A

  • Server performance test method and device and medium

    CN115480973A

  • Testing server, information processing system, and testing method

    US20140298082A1

Cited By

  • SSD test method, system, device and equipment and storage medium

    CN120375895A

  • Server configuration matching method and device, equipment and storage medium

    CN120561618A

  • Server configuration matching method and device, equipment and storage medium

    CN120561618B

  • Server test method and device, storage medium and computer program product

    CN120687266A

  • Method and system for cloud platform to automatically adapt to computing node, terminal and storage medium

    CN121173664A