Server hardware state monitoring method and electronic device
By deploying distributed agent nodes inside the server, the bus load rate is monitored and processed in real time, which solves the problem of response latency and data transmission conflicts caused by too many devices in the server system, and improves data transmission efficiency and system scalability.
Patent Information
- Application Number
- CN202511197463.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-08-26
AI Technical Summary
In server systems, as the number of connected devices increases, the line load rises, leading to response delays, reduced data transmission efficiency, and problems such as data transmission conflicts and excessively long polling times.
Distributed agent nodes are deployed within the server according to physical partitions. These agent nodes collect data from slave devices via the internal integrated circuit bus and monitor the bus load rate in real time through a monitoring agent program. After targeted processing, the data is uploaded to the management controller, reducing the pressure of centralized data transmission and avoiding data transmission conflicts caused by too many devices connected to the bus.
It improves system response efficiency, reduces polling time, avoids data transmission conflicts, enhances overall scalability and real-time response capabilities, and adapts to complex server application scenarios.
Smart Images

Figure CN120687330B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of server hardware monitoring, and particularly relates to a server hardware state monitoring method and an electronic device. BACKGROUND
[0002] In a server system, various external devices are usually connected through a serial communication mode; a two-wire design is generally adopted, a multi-master device and a multi-slave device coexistence architecture mode is supported, and shared access and data transmission of multiple master devices to the same line resource can be realized. However, when a centralized management controller is used to connect various devices through a single line, with the increase of the number of connected devices, the line load will greatly increase, resulting in response delay and reduced data transmission efficiency. When the number of devices is large, data transmission conflicts and long polling time are prone to occur. SUMMARY
[0003] The present application provides a server hardware state monitoring method and an electronic device, at least to solve the problems of low data transmission efficiency, prone to data transmission conflicts and long polling time in the related art.
[0004] The present application provides a server hardware state monitoring method, a distributed proxy node is deployed in a server according to physical partitioning, and the proxy node is provided with a monitoring proxy program; the server hardware state monitoring method comprises the following steps:
[0005] After the proxy node is connected with a corresponding slave device, the proxy node collects data of the corresponding slave device through an internal integrated circuit bus;
[0006] The monitoring proxy program is used to monitor the load rate of the internal integrated circuit bus;
[0007] According to the load rate of the internal integrated circuit bus, the proxy node processes the collected slave device data correspondingly, and uploads the processed data to a management controller.
[0008] The present application further provides an electronic device, comprising a memory for storing a computer program, and a processor for executing the computer program to realize the steps of the above-mentioned server hardware state monitoring method.
[0009] Through the application, since the distributed agent nodes are arranged in the server according to physical partition, the agent nodes are provided with monitoring agent programs, the agent nodes are used to realize accurate collection of slave device data through internal integrated circuit buses, the monitoring agent programs are used to monitor the bus load rate in real time and upload the collected data to the management controller after targeted processing, so that the distributed agent nodes can undertake the data collection and processing tasks, the system response efficiency is greatly improved, the management controller does not need to poll all slave devices, and only needs to access the agent nodes to obtain the required data, the problem of long polling time caused by too many device nodes is effectively solved, the bus load is reasonably distributed, the centralized data transmission pressure is reduced, the data transmission conflict phenomenon caused by too many bus-mounted devices is avoided, and the performance limitation of centralized monitoring is broken. Under the distributed collaborative architecture, the system can fully exert the local data processing capacity of each agent node, significantly enhance the overall expansion performance and real-time response capability, perfectly adapt to the server application scenarios with strict hardware state monitoring requirements, especially suitable for actual needs of the increasingly complex server running environment, and the entire process will not affect the integrity of the distributed system.
[0010] In addition, the application also provides an electronic device corresponding to the server hardware state monitoring method, which has the same or corresponding technical features as the above-mentioned server hardware state monitoring method and the same effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 The flowchart of the server hardware state monitoring method provided by the embodiment of the application is shown in the figure.
[0013] Figure 2 The architecture schematic diagram of the server hardware state monitoring method provided by the embodiment of the application is shown in the figure.
[0014] Figure 3 The signaling interaction diagram among the management controller, the master-slave agent nodes and the slave devices provided by the embodiment of the application is shown in the figure.
[0015] Figure 4 The relationship schematic diagram among the management controller, the master-slave agent nodes and the slave devices provided by the embodiment of the application is shown in the figure.
[0016] Figure 5 The structure schematic diagram of the server hardware state monitoring device provided by the embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0017] In the computer hardware design system, the internal integrated circuit bus is widely used as a kind of serial communication protocol with low cost and low power consumption, mainly used for connecting and managing key hardware components to realize low-speed control and state monitoring between devices. However, due to the fact that all sensors share a single internal integrated circuit bus (such as a bandwidth of only 100-400Kbps), when the number of monitoring devices (such as more than 20) reaches a certain scale, the bus utilization rate will record a high level (such as more than 90%). At the same time, its multi-master device sharing bus characteristics is easy to cause bus collision and operation failure. In the traditional centralized architecture mode, if the master node or the bus itself fails, it will directly cause the entire monitoring system to be in a paralyzed state. The present application provides a server hardware state monitoring method, which can solve the above problems.
[0018] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0019] It should be noted that in the description of the present application, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.
[0020] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0021] In combination with the specific application environment architecture or specific hardware architecture on which the server hardware state monitoring method is executed, the specific application environment architecture or specific hardware architecture is described here.
[0022] The embodiments of the present application provide a server hardware state monitoring method, and the server is internally deployed with distributed proxy nodes according to physical partitioning, and the proxy nodes are provided with monitoring proxy programs. The method is described in detail in combination with the execution flow of the server hardware state monitoring method. Figure 1 The flowchart of the server hardware state monitoring method provided by the embodiments of the present application is shown in Figure 1As shown, the method comprises:
[0023] S101, after the agent node is connected with the corresponding slave device, the agent node is used to collect the data of the corresponding slave device through the internal integrated circuit bus.
[0024] It should be noted that the present application is aimed at the scene that the server environment is becoming more and more complex, and the distributed agent nodes are deployed in the server according to the physical partition, and the collection and reporting of the slave device data are performed through the distributed agent nodes. Each agent node can be connected with the corresponding slave device, and the corresponding slave device is connected through the internal integrated circuit bus. Therefore, the defect of low efficiency of the main device in turn can be avoided, the problem of long device polling time caused by too many device nodes can be solved, and the data conflict caused by too many bus devices can be avoided. The present application upgrades the traditional centralized monitoring architecture to the distributed collaborative mode through the deployment of the distributed agent nodes. The low power consumption and easy deployment characteristics of the internal integrated circuit bus are retained, the system scalability and real-time performance are improved through the local processing capacity of the agent nodes, and the present application is suitable for the server scene with high hardware state monitoring requirements.
[0025] S102, the load rate of the internal integrated circuit bus is monitored by using the monitoring agent program.
[0026] In practical application, the monitoring agent program can be installed in each agent node, and the monitoring agent programs installed in each agent node can communicate based on the related protocol, the monitoring data is summarized, and the load rate of the internal integrated circuit bus is monitored. The load rate of the internal integrated circuit bus directly reflects the busy degree of the communication channel. When the load rate is too high, indiscriminate data uploading can intensify the bus congestion, causing delay or loss of critical information. When the load rate is low, the bandwidth can be reasonably utilized to improve the timeliness of data transmission.
[0027] S103, according to the load rate of the internal integrated circuit bus, the agent node is used to process the collected slave device data correspondingly, and the processed data is uploaded to the management controller.
[0028] In the implementation summary, the present application is based on the load rate of the internal integrated circuit bus, and the collected data of the slave device is processed by the agent node and uploaded to the management controller. The agent node as an intermediate processing link can dynamically adjust the data processing strategy according to the real-time monitored bus load rate. Through the dynamic processing and scheduling of the agent node, the communication resources can be flexibly allocated according to the actual load of the bus, the risk of bus overload is effectively reduced, the data transmission conflict and delay are reduced, the receiving and operation pressure of the management controller is reduced through the processed simplified data, and the overall efficiency of the data processing link is improved.
[0029] In the server hardware state monitoring method provided by the embodiment of the application, the server is internally arranged with distributed proxy nodes according to physical partitions, the proxy nodes are provided with monitoring proxy programs, the proxy nodes are used to realize accurate collection of slave device data through an internal integrated circuit bus, the bus load rate is monitored in real time by using the monitoring proxy programs, and the collected data is uploaded to the management controller after being processed in a targeted manner, so that the distributed proxy nodes bear the data collection and processing tasks, the system response efficiency is greatly improved, the management controller does not need to poll all slave devices, and only needs to access the proxy nodes to obtain the required data, the problem of long polling time caused by too many device nodes is effectively solved, the bus load is reasonably allocated, the centralized data transmission pressure is reduced, the data transmission conflict phenomenon caused by too many bus-mounted devices is avoided, and the performance limitation of centralized monitoring is broken. Under the distributed collaborative architecture, the system can fully exert the local data processing capacity of each proxy node, significantly enhance the overall expansion performance and real-time response capability, perfectly adapt to the server application scenarios with strict hardware state monitoring requirements, and is especially suitable for actual needs of the increasingly complex server running environment, and the entire process does not affect the integrity of the distributed system.
[0030] Further, in specific implementation, in the server hardware state monitoring method provided by the embodiment of the application, the server is internally arranged with distributed proxy nodes according to physical partitions, and specifically can include the following steps: the server is internally divided into multiple physical partitions according to hardware functions, layouts and heat dissipation requirements; the physical partitions include a processor area, a storage area, a power supply area and a fan area; and the distributed proxy nodes are arranged in the divided physical partitions.
[0031] In implementation, the physical partitions are not independent of external units of the server, but are physical sub-regions divided internally in the server according to hardware functions, layouts or heat dissipation requirements. For example, the processor area corresponds to the installation area of processors and peripheral core computing components in the server, the storage area corresponds to the deployment space of storage devices such as hard disks and solid state disks, and the power supply area is an independent area where power supply modules and power supply lines are located. These physical partitions jointly constitute the overall hardware structure of the server and are subdivided parts of the physical entity of the server.
[0032] The application divides the server into physical partitions such as processor area, storage area, power supply area and fan area according to hardware functions, layouts and heat dissipation requirements, and deploys distributed proxy nodes in each partition. The physical partitions have clear function boundaries, and the heat dissipation design of each area can be optimized (for example, independent air ducts are configured for the high-heat processor area and power supply area), the heat interference between different hardware modules is reduced, the local overheating risk is reduced, and the stable operation of the hardware is ensured. At the same time, the partition layout also facilitates the centralized management and maintenance of hardware resources. The deployment of distributed proxy nodes in each partition can realize accurate and real-time monitoring of the hardware state (such as processor load, storage read / write speed, power supply voltage, fan speed, etc.) in the partition, avoiding the limitations and delay problems of single node monitoring. This architecture not only improves the reliability and heat dissipation efficiency of the server hardware operation through physical partitioning, but also enhances the perception ability of the state of each area with the help of distributed proxy nodes.
[0033] Figure 2 The server hardware state monitoring method provided by the embodiment of the application corresponds to the schematic diagram of the architecture. As shown in Figure 2 The plurality of proxy nodes can be connected with the management controller, and each proxy node can be connected with a corresponding sensor. For example, the first sensor can be a central processing unit (CPU) sensor, the second sensor can be a memory sensor, and the third sensor can be a graphics processing unit (GPU) sensor. After the proxy node collects the data of each sensor and performs preliminary processing, the data is transmitted to the management controller. The management controller does not need to poll all slave devices, and the number of monitored slave nodes of the server can be expanded.
[0034] Further, in the server hardware state monitoring method provided by the embodiment of the application, step S102 uses a monitoring proxy program to monitor the load rate of the internal integrated circuit bus, which can specifically include: using the monitoring proxy program to obtain the effective working time of the internal integrated circuit bus within a set statistical period; and obtaining the load rate of the internal integrated circuit bus according to the ratio between the obtained effective working time of the internal integrated circuit bus within the set statistical period and the set statistical period.
[0035] In implementation, the load rate of the internal integrated circuit bus can be obtained by using the following formula:
[0036] L=(T_active / T_total)×100%;
[0037] Wherein, L is a load rate of the internal integrated circuit bus, T_active is an effective working time (i.e. low level time of the serial clock line) of the internal integrated circuit bus in a set statistical period, and T_total is the set statistical period. The application can obtain T_active by directly collecting bus timing data from a state machine register of the internal integrated circuit bus of a micro controller unit (MCU), and then realize real-time monitoring and dynamic regulation and control of the bus load rate.
[0038] Further, in the above-mentioned server hardware state monitoring method provided by the embodiment of the application, in the step S103, the collected slave device data is processed by the agent node according to the load rate of the internal integrated circuit bus, and specifically can include: the collected slave device data is filtered by the agent node through a built-in algorithm according to the load rate of the internal integrated circuit bus; the slave device data after filtering is identified for an unexpected state to obtain data features of the unexpected state; and the data features of the unexpected state are compressed.
[0039] In the implementation, the collected slave device data is filtered by the agent node through a built-in algorithm according to the load rate of the internal integrated circuit bus; then, the filtered data is identified for an unexpected state, data features of the unexpected state are extracted and compressed. After the agent node completes filtering, abnormality detection and compression of the sensor data locally, only valid data or abnormal events are transmitted to the management controller. Meanwhile, the application can support active alarm triggered by the sensor (such as sudden temperature rise), the agent node does not need to wait for polling of the master node, and directly reports emergency data through an interrupt mode, so that the fault response time is much lower than that of the traditional scheme. The agent node has arbitration capability of the hardware-level internal integrated circuit bus, and can realize priority preemption type communication. Through local data processing and selective transmission, the bus data volume is greatly reduced, the high load pressure is relieved, and the bus communication efficiency is improved; the active alarm and interrupt reporting mechanism strengthens the response speed of the system to emergency events, and enhances the timeliness of fault disposal; the hardware-level arbitration and priority communication guarantee the transmission priority of key data, avoid delay of core information caused by bus congestion, and comprehensively improve the reliability and real-time performance of the internal integrated circuit bus in the multi-device cooperative scene.
[0040] Further, in the specific implementation, in the server hardware state monitoring method provided by the embodiment of the application, when the slave device is a temperature sensor, the step S101 uses the agent node to collect the corresponding slave device data through the internal integrated circuit bus, and specifically can include: using the agent node to collect the data of the temperature sensor at a set time interval through the internal integrated circuit bus. For example, the temperature sensor can sample 10 times per second, and the agent node only reports when the temperature change exceeds the threshold value, reducing the bus data volume by more than 80%.
[0041] Correspondingly, in the above step, the filtered slave device data is subjected to unexpected state identification, which specifically can include: comparing the filtered temperature sensor data; if the temperature sensor data is higher than a first preset temperature threshold or lower than a second preset temperature threshold, the temperature sensor data is determined to be in an unexpected state; if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold but sudden change occurs within a set time period, the temperature sensor data is determined to be in an unexpected state; and if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold and no sudden change occurs within a set time period, the temperature sensor data is determined to be in an expected state.
[0042] In the implementation, the filtered temperature sensor data is subjected to multi-dimensional comparison and determination. When the data exceeds the upper limit of the first preset temperature threshold or is lower than the lower limit of the second preset temperature threshold, it is directly determined to be abnormal. Even if the data is within the threshold interval, if sudden change occurs within a set time period, it is still determined to be in an unexpected state, and only when the data is stable within the threshold interval and no sudden change occurs is it determined to be in an expected state. In this way, through threshold boundary and dynamic time dimension monitoring, comprehensive coverage of temperature abnormalities is achieved, which not only avoids risks caused by data exceeding the safe range without being detected, but also timely captures abnormal fluctuations within the threshold (such as sudden temperature rise before equipment failure), greatly improving the accuracy and sensitivity of abnormal identification; and accurate state determination can reduce invalid data upload, so that the system focuses only on the processing and response of unexpected state information, reducing bus transmission pressure and main controller operation load, and providing a reliable basis for equipment failure warning and abnormal tracing.
[0043] Further, in the specific implementation, in the above step, according to the load rate of the internal integrated circuit bus, the agent node uses a built-in algorithm to filter the collected slave device data, including: when the load rate of the internal integrated circuit bus exceeds a first preset load rate threshold, the agent node uses the built-in algorithm to filter the collected slave device data for targeted filtering of non-critical data in a corresponding range; and when the load rate of the internal integrated circuit bus is lower than a second preset load rate threshold, the system enters a sleep mode.
[0044] In implementation, the application can perform differentiated data processing strategies for the load rate of the internal integrated circuit bus: when the load rate exceeds a first preset load rate threshold (such as > 80%), an emergency mode is triggered immediately, the data collected from the slave device is filtered by the built-in algorithm of the agent node, non-critical data is discarded preferentially to reduce the bus transmission pressure, thereby ensuring the continuity and reliability of core data communication; when the load rate is less than a second preset load rate threshold (such as ≤ 30%), the system maintains the current running state, and can enter a sleep mode according to the energy consumption optimization requirement; when the load rate is between the first preset load rate threshold and the second preset load rate threshold (i.e. 30% < load rate ≤ 80%), a normal monitoring mechanism is started, and the efficient use of bus resources is ensured by optimizing the data sampling frequency.
[0045] Further, in specific implementation, in the server hardware state monitoring method provided by the embodiment of the application, the agent nodes include a master agent node and a slave agent node.
[0046] In step S101, the agent node collects corresponding slave device data through the internal integrated circuit bus, and in step S103, the agent node processes the collected slave device data. Specifically, the master agent node collects corresponding slave device data through the internal integrated circuit bus, and the master agent node processes the collected slave device data. At this time, the slave agent node is in a sleep state. When the management controller detects that the heartbeat signal of the master agent node fails, the management controller sends an activation instruction to the slave agent node. After the slave agent node is activated, the slave agent node performs a data collection task and a data processing task.
[0047] Figure 3 A signaling interaction diagram between the management controller, the master-slave agent nodes and the slave device provided by the embodiment of the application is shown in FIG. 2. Figure 3 As shown in FIG. 2, the agent nodes can adopt a dual-node redundancy design. When the master agent node is normally running, the slave agent node is in a sleep state, and all data collection and processing are performed by the master agent node. The slave node in the sleep state can reduce the system energy consumption, and the operation of the master node ensures the stability and efficiency of data processing. The management controller continuously monitors the periodic heartbeat signal of the master agent node. If the heartbeat signal is found to be abnormal (indicating that the master agent node fails), the management controller will immediately send an activation instruction to the slave agent node, so that the slave agent node switches to a working state and takes over the data collection task and the data processing task. Seamless switching of data collection and processing can be realized, data interruption or business stagnation caused by single-point failure can be avoided, and the fault tolerance and operation reliability of the system are greatly improved.
[0048] Further, in the specific implementation, in the server hardware state monitoring method provided by the embodiment of the present application, when the slave agent node takes over the data collection task and the data processing task after being activated, the method can further include: performing a reset operation on the master agent node; if the master agent node recovers, the management controller switches back to the master agent node to perform the data collection task and the data processing task in the next collection cycle, and issues a sleep instruction to the slave agent node.
[0049] In the implementation, when the slave agent node takes over the data collection task and the data processing task, the present application can perform a reset operation on the master agent node. If the master agent node recovers normally after being reset, the management controller will automatically switch back to the master agent node for data collection in the next collection cycle; if the master agent node cannot recover to normal operation, the system will actively report fault information and wait for manual intervention.
[0050] Further, in the specific implementation, the distributed agent nodes can be deployed, which can specifically include: deploying each agent node by using an independent power supply mechanism; wherein the agent node is integrated with a master control micro control unit and a power module; the power module provides power supply for the agent node; and the master control micro control unit controls the agent node to interact with data, coordinates the data collection task and the data processing task of the agent node.
[0051] In the implementation, the distributed agent nodes can use independent power supply and communication link, and when a single agent node fails, only the local sensor area is affected, and other nodes can still work normally. For example, when the power module corresponding to the agent node fails, the monitoring of the processor heat dissipation area will not be affected.
[0052] Figure 4 The relationship between the management controller, the master-slave agent nodes and the slave devices provided by the embodiment of the present application is shown in the schematic diagram. Figure 4As shown, the agent nodes are divided into master agent nodes and slave agent nodes, the master agent nodes and the slave agent nodes have data arrangement functions, and each contains a master control module (optionally an MCU module) and a power module to ensure the operation of each agent node. The master agent nodes and the slave agent nodes can interact with slave devices (including device 1, device 2, device 3, device 4, etc.) through upper and lower interfaces to collect data, and can realize low load and high response speed data acquisition due to the short physical distance from the slave devices and the single data collection target. The master control module supports direct execution of hardware control strategies, for example, when a high temperature of the processor is monitored, the processor can be immediately responded and triggered to reduce the frequency and other measures, and at the same time, the related information is reported to the base management controller for feedback, which improves the response speed. The standardized configuration of such master-slave agent nodes ensures the stability and consistency of data collection and improves the data acquisition efficiency; the local hardware control capability of the master control module greatly shortens the abnormal response time, avoids the delay caused by relying on the upper controller decision, and enhances the rapid disposal capability of device failure.
[0053] The server hardware state monitoring method provided by the application will be described below with specific examples:
[0054] Taking an 8-GPU node as an example, the hardware deployment adopts a partitioned agent node architecture: agent node 1 is deployed in the power supply area for monitoring the power module; agent node 2 is deployed in the CPU area for monitoring the CPU and memory temperature; agent node 3 is deployed in the GPU area for monitoring the core and memory temperature of the 8 GPUs; agent node 4 is deployed in the storage area for monitoring the solid state disk and backplane temperature; and agent node 5 is deployed in the fan area for monitoring the fan node state.
[0055] The above architecture realizes data collection and local management of each partitioned slave device through distributed agent nodes, and related data does not need to be centrally collected to the base management controller (Baseboard Management Controller, BMC) for processing, which significantly improves the response efficiency. At the same time, the BMC does not need to poll all slave devices, but only needs to interact with each agent node to obtain the global state, which greatly reduces the load pressure of itself. In addition, since each agent node monitors a single device, it can be optimized for targeted collection according to the characteristics of the device to improve the data collection efficiency; when the device is abnormal, it can realize fast response and local disposal. For example, when the GPU is overheated, the agent node 3 can first execute the frequency reduction operation, and at the same time, report the overheating event to the BMC, and the BMC receives the information and issues an instruction to the agent node 5 to increase the fan speed, forming a cooperative linkage abnormal disposal mechanism to ensure the stable operation of the node.
[0056] When the master proxy node fails, the BMC discovers that the master proxy node has no heartbeat signal returned through the heartbeat detection mechanism, and after actively inquiring and confirming that it has no response, the slave proxy node will be activated immediately to take over the data collection task to ensure business continuity. At the same time, the BMC performs a restart operation on the master proxy node; if the master proxy node recovers normally after restarting, the BMC will automatically switch to the master proxy node for operation in the next collection period, and at the same time, the slave proxy node enters a standby state, so that the seamless switching mechanism ensures that the data collection process is uninterrupted; if the master proxy node still cannot recover after restarting, the running state of the slave proxy node is maintained until the fault is eliminated. This process realizes the high availability of the proxy node redundant architecture through clear fault detection, switching and recovery logic.
[0057] The proxy node has the ability to monitor its own load data in real time and executes a dynamic regulation and control strategy accordingly: when the load data is below 30%, the proxy node automatically switches to a sleep mode to reduce energy consumption by reducing the sampling frequency; if the load data suddenly rises to above 80%, the proxy node will start a data priority screening mechanism to discard non-critical data such as vendor information, and at the same time, the sampling period of each round is extended to relieve the load pressure, and the abnormal state is reported to the BMC. After receiving the information, the BMC triggers a prompt mechanism, which is manually intervened to determine whether the phenomenon belongs to temporary fluctuation, abnormal working condition or potential server failure, so as to take targeted measures. This dynamic load management mechanism not only realizes energy efficiency optimization, but also guarantees the effective transmission of core data in a high-load scenario through hierarchical response, and at the same time, the accuracy of abnormal handling is improved through manual decision-making.
[0058] Through the description of the above implementation mode, those skilled in the art can clearly understand that the method according to the above embodiment can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better implementation mode.
[0059] The embodiment of the application also provides a server hardware state monitoring device, which is distributed in a physical partition inside the server. Figure 5 The server hardware state monitoring device provided by the embodiment of the application is shown in the structural schematic diagram. The embodiment is based on the perspective of functional modules, as shown in the figure, the device comprises: Figure 5
[0060] The data collection module 11 is used for collecting data of the corresponding slave device through the internal integrated circuit bus by the proxy node after the proxy node is connected with the corresponding slave device;
[0061] The load rate monitoring module 12 is used for monitoring the load rate of the internal integrated circuit bus by the monitoring agent program;
[0062] The data processing module 13 is configured to process the collected slave device data by using the agent node according to the load rate of the internal integrated circuit bus, and upload the processed data to the management controller.
[0063] In the server hardware state monitoring device provided by the embodiment of the application, the server is internally divided into multiple physical partitions according to hardware functions, layouts and heat dissipation requirements, and the distributed agent nodes are arranged in the physical partitions. The agent nodes are provided with monitoring agent programs, the agent nodes are used to accurately collect slave device data through the internal integrated circuit bus, the monitoring agent programs are used to monitor the bus load rate in real time and upload the collected data to the management controller after targeted processing, so that the distributed agent nodes can be used to undertake the data collection and processing tasks, the system response efficiency is greatly improved, the management controller does not need to poll all the slave devices, and only needs to access the agent nodes to obtain the required data, the problem of long polling time caused by too many device nodes is effectively solved, the bus load is reasonably distributed, the centralized data transmission pressure is reduced, the data transmission conflict phenomenon caused by too many bus-mounted devices is avoided, and the performance limitation of centralized monitoring is broken. In the distributed collaborative architecture, the local data processing capacity of each agent node can be fully utilized, the overall expansion performance and real-time response capability are significantly enhanced, the server application scenarios with strict hardware state monitoring requirements are perfectly adapted, and the actual demand for the increasingly complex server running environment is particularly suitable, and the integrity of the distributed system is not affected.
[0064] Since the embodiments of the server hardware state monitoring device part correspond to the embodiments of the server hardware state monitoring method part, the description of the features in the embodiments corresponding to the server hardware state monitoring device can be referred to the related description of the embodiments corresponding to the server hardware state monitoring method, which will not be repeated here. And has the same beneficial effects as the above-mentioned server hardware state monitoring method.
[0065] Further, in specific implementation, in the server hardware state monitoring device provided by the embodiment of the application, the server is internally divided into multiple physical partitions according to hardware functions, layouts and heat dissipation requirements; the physical partitions include a processor area, a storage area, a power supply area and a fan area; and the distributed agent nodes are arranged in the divided physical partitions.
[0066] Further, in specific implementation, in the server hardware state monitoring device provided by the embodiment of the application, the load rate monitoring module 12 can be specifically configured to obtain the effective working time of the internal integrated circuit bus in a set statistical period by using the monitoring agent program; and obtain the load rate of the internal integrated circuit bus according to the ratio between the obtained effective working time of the internal integrated circuit bus in the set statistical period and the set statistical period.
[0067] Further, in specific implementation, in the server hardware state monitoring device provided by the embodiment of the application, the data processing module 13 can be specifically configured to perform corresponding filtering processing on the collected slave device data by using the agent node and through an internal integrated circuit bus according to a load rate of the internal integrated circuit bus; perform unexpected state identification on the slave device data after filtering processing to obtain data features of the unexpected state; and perform compression processing on the data features of the unexpected state.
[0068] Further, in specific implementation, in the server hardware state monitoring device provided by the embodiment of the application, when the slave device is a temperature sensor, the data collection module 11 can be specifically configured to collect data of the temperature sensor at a set time interval by using the agent node and through the internal integrated circuit bus.
[0069] Correspondingly, the data processing module 13 can be specifically configured to compare the temperature sensor data after filtering processing; if the temperature sensor data is higher than a first preset temperature threshold or lower than a second preset temperature threshold, it is determined that the temperature sensor data is in an unexpected state; if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold but a sudden change occurs within a set time period, it is determined that the temperature sensor data is in an unexpected state; and if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold and no sudden change occurs within the set time period, it is determined that the temperature sensor data is in an expected state.
[0070] The data processing module 13 can be specifically configured to, when the load rate of the internal integrated circuit bus exceeds a first preset load rate threshold, perform targeted filtering processing on the collected slave device data by using the agent node and through the built-in algorithm to filter out non-critical data in a corresponding range; and when the load rate of the internal integrated circuit bus is lower than a second preset load rate threshold, enter a sleep mode.
[0071] Further, in specific implementation, in the server hardware state monitoring device provided by the embodiment of the application, each agent node includes a master agent node and a slave agent node. The data collection module 11 can be specifically configured to collect corresponding slave device data by using the master agent node and through the internal integrated circuit bus. The data processing module 13 can be specifically configured to perform corresponding processing on the collected slave device data by using the master agent node; the slave agent node is in a sleep state; when the management controller detects that a heartbeat signal of the master agent node fails, an activation instruction is issued to the slave agent node, and the slave agent node performs a data collection task and a data processing task after being activated. The slave agent node performs a reset operation on the master agent node at the same time of performing the data collection task and the data processing task after being activated; if the master agent node recovers, the management controller switches back to the master agent node to perform the data collection task and the data processing task at a next collection cycle, and issues a sleep instruction to the slave agent node.
[0072] Further, in a specific implementation, in the server hardware state monitoring device provided by the embodiment of the present application, each agent node is deployed using an independent power supply mechanism; a main control micro control unit and a power module are integrated in the agent node; the power module provides power supply for the agent node; and the main control micro control unit controls the agent node to perform data interaction, coordinates data collection tasks and data processing tasks of the agent node.
[0073] The embodiment of the present application also provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is arranged to run the computer program to perform the steps in any of the above server hardware state monitoring method embodiments.
[0074] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, wherein the computer program is arranged to perform the steps in any of the above server hardware state monitoring method embodiments when running.
[0075] In an example embodiment, the above computer readable storage medium can include but is not limited to a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk and various computer program storage media.
[0076] The embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to perform the steps in any of the above server hardware state monitoring method embodiments.
[0077] The embodiment of the present application also provides another computer program product, which comprises a non-volatile computer readable storage medium, the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to perform the steps in any of the above server hardware state monitoring method embodiments.
[0078] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0079] The above describes in detail the server hardware state monitoring method and electronic device provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for ordinary skilled persons in the technical field, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A method of server hardware state monitoring, the method comprising: The server is internally arranged with distributed proxy nodes according to physical partition, each of the proxy nodes is connected with a corresponding slave device through an internal integrated circuit bus; each of the proxy nodes is connected with a management controller; each of the proxy nodes is provided with a monitoring proxy program; the monitoring proxy programs arranged in each of the proxy nodes communicate with each other; Each of the proxy nodes comprises a master proxy node and a slave proxy node; The server hardware state monitoring method comprises: After the proxy node is connected with the corresponding slave device, the proxy node collects the data of the corresponding slave device through the internal integrated circuit bus; when the slave device is a temperature sensor, the proxy node collects the data of the temperature sensor at a set time interval through the internal integrated circuit bus; wherein the master proxy node collects the data of the corresponding slave device through the internal integrated circuit bus; the slave proxy node is in a dormant state; when the management controller detects that the heartbeat signal of the master proxy node fails, the management controller sends an activation instruction to the slave proxy node, and the slave proxy node activates to perform a data collection task; The monitoring proxy program is used to obtain the effective working time of the internal integrated circuit bus within a set statistical period; according to the ratio between the obtained effective working time of the internal integrated circuit bus within the set statistical period and the set statistical period, the load rate of the internal integrated circuit bus is obtained by using the following formula: L=(T_active / T_total)×100%; Wherein, L is the load rate of the internal integrated circuit bus, T_active is the effective working time of the internal integrated circuit bus within the set statistical period, and T_total is the set statistical period; the effective working time of the internal integrated circuit bus within the set statistical period is obtained by collecting bus timing data through the internal integrated circuit bus state machine register of the micro control unit, so as to realize real-time monitoring and dynamic regulation and control of the load rate; When the load rate of the internal integrated circuit bus exceeds a first preset load rate threshold, an emergency mode is triggered, the proxy node is used to perform targeted filtering processing on the collected slave device data through a built-in algorithm, and non-key data in a corresponding range is filtered out; the slave device data after filtering processing is subjected to unexpected state identification to obtain data characteristics of the unexpected state; when the slave device is a temperature sensor, the temperature sensor data after filtering processing is subjected to comparison; if the temperature sensor data is higher than a first preset temperature threshold or lower than a second preset temperature threshold, the temperature sensor data is determined as an unexpected state; if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold but appears sudden change within a set time period, the temperature sensor data is determined as an unexpected state; if the temperature sensor data is between the second preset temperature threshold and the first preset temperature threshold and does not appear sudden change within a set time period, the temperature sensor data is determined as an expected state; The data characteristics of the unexpected state are compressed and processed, and the processed data is uploaded to the management controller; When the load rate of the internal integrated circuit bus is lower than a second preset load rate threshold, then enter a sleep mode according to energy consumption optimization requirements; According to the load rate of the internal integrated circuit bus, the collected slave device data is processed by the agent node, and the processed data is uploaded to the management controller, so that the management controller triggers a prompt mechanism; wherein the collected slave device data is processed by the master agent node; the slave agent node is in a sleep state; when the management controller monitors that the heartbeat signal of the master agent node fails, an activation instruction is issued to the slave agent node, and the slave agent node activates to perform a data processing task; at the same time, the master agent node is reset; if the master agent node recovers, the management controller switches back to the master agent node to perform data collection and data processing tasks in the next collection cycle, and issues a sleep instruction to the slave agent node.
2. The server hardware state monitoring method of claim 1, wherein, The server is internally deployed with distributed agent nodes according to physical partitions, including: The server is internally divided into multiple physical partitions according to hardware functions, layouts and heat dissipation requirements; the physical partitions include processor areas, storage areas, power supply areas and fan areas; Distributed agent nodes are deployed in the divided physical partitions.
3. The method of claim 1, wherein, Deploying distributed agent nodes includes: An independent power supply mechanism is used to deploy each agent node; wherein the agent node integrates a master control micro control unit and a power module; the power module is used to provide power supply for the agent node; the master control micro control unit is used to control the agent node to interact data, coordinate data collection tasks and data processing tasks of the agent node.
4. An electronic device, comprising: Including: A memory for storing a computer program; A processor for executing the computer program to realize the steps of the server hardware state monitoring method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Agent node and sensor network
CN102625486A
Elastic monitoring method for key task computer cluster
CN105024880A
Monitoring alarm method and device based on block chain, electronic equipment and storage medium
CN114416490A