Server cooperative control method, storage medium and electronic equipment
By taking over control permissions by programmable logic devices during the server startup stage and ensuring the functional integrity of the substrate management controller through the multi-communication link synchronization mechanism, the problem of the substrate management controller losing control capabilities in the heterogeneous computing architecture is solved, and the stability and fault tolerance of the server system are improved.
Patent Information
- Application Number
- CN202510401844.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In heterogeneous computing architecture, the substrate management controller loses control capabilities due to communication link failure or program abnormality, resulting in system uncontrollable, execution lag or function interruption, which poses a system-level operation risk, and it is difficult to coordinate the control between heterogeneous units, resulting in difficulty in data synchronization and poor execution coherence.
By actively taking over the control authority of the target device based on the sensor operation parameters during the server startup stage, the programmable logic device ensures that the stable operation of the device before initialization is completed. The preset communication link is used to communicate with the substrate management control device for interactive verification, and the functional integrity and communication reliability are ensured during permission handover through two-way communication verification. Using the parameter synchronization mechanism of multi-communication links parallel, the device control parameters are updated in real time through redundant links to ensure that parameter consistency can be maintained when any link is abnormal.
It realizes seamless control rights and reliable control of equipment status during the server startup stage, ensures the rapid recovery of the system and the robustness of data synchronization when the board management controller is abnormal, and improves the stability and fault tolerance of the server system.
Smart Images

Figure CN119917350A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of firmware technology, and in particular to a server collaborative control method, a storage medium, and an electronic device. Background Art
[0002] In the field of server control and management, traditional solutions usually use a single controller to implement the operation monitoring and parameter adjustment functions of the target device. With the popularization of heterogeneous computing architecture, the number of systems that simultaneously deploy programmable logic chips and baseboard management controllers is increasing, but there are significant defects in the collaborative control process between the two.
[0003] In related technologies, when the baseboard management controller loses its control capability due to a communication link failure or program anomaly, the system cannot be controlled, resulting in delayed execution of the target device or even functional interruption, which may cause system-level operation risks. In addition, the control permissions between heterogeneous units are independently configured, making it difficult for heterogeneous units to be coordinated and controlled. Therefore, data synchronization and execution continuity are difficult to ensure, which has become a key technical obstacle to improving the operating efficiency of high-density servers. Summary of the invention
[0004] The present application provides a server collaborative control method, a storage medium and an electronic device to at least solve the problem in the related art that the system operation risk is relatively high and the collaborative control between heterogeneous units is difficult.
[0005] The present application provides a server collaborative control method, comprising: in response to server startup, loading an operating system, taking over control authority of a target device associated with the server according to operating parameters corresponding to a sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and a baseboard management and control device; in response to the baseboard management and control device having completed startup, interactively verifying with the baseboard management and control device based on a preset communication link, and transferring the control authority to the baseboard management and control device if the verification passes; in response to detecting an abnormality in the baseboard management and control device, taking over the control authority again, wherein the programmable logic device and the baseboard management and control device synchronize the control parameters of the target device via at least one communication link, and the at least one communication link includes the preset communication link.
[0006] The present application also provides a server collaborative control device, including: an initialization module, for responding to server startup, loading an operating system, and taking over the control authority of a target device associated with the server according to the operating parameters corresponding to the sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and a baseboard management and control device; a verification module, for responding to the baseboard management and control device having completed startup, interactively verifying with the baseboard management and control device based on a preset communication link, and transferring the control authority to the baseboard management and control device if the verification passes; a recovery module, for responding to detecting an abnormality in the baseboard management and control device, retaking the control authority, wherein the programmable logic device and the baseboard management and control device synchronize the control parameters of the target device through at least one communication link, and the at least one communication link includes the preset communication link.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server collaborative control methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned server collaborative control methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server collaborative control methods when executed by a processor.
[0010] Through the embodiments of the present application, during the server startup phase, the programmable logic device actively takes over the control authority of the target device based on the operating parameters of the sensor device. By prepending the device management responsibility to the programmable logic device with rapid response capabilities, the stable operation of the device is ensured before the baseboard management control device completes the initialization, thereby avoiding device management vacuum or parameter configuration conflict caused by differences in controller startup sequence, thereby achieving the technical effect of seamless connection of control rights during the server startup phase and reliable management and control of device status.
[0011] In addition, a mechanism for interactive verification between the preset communication link and the baseboard management and control device is adopted, and the functional integrity and communication reliability of the baseboard management and control device are ensured through two-way communication verification during the transfer of authority, so as to eliminate the potential logic errors or link abnormality risks before the transfer of control rights, thereby achieving the technical effect of safe transition of control rights and improved system initialization efficiency. At the same time, through the parameter synchronization mechanism of multiple communication links in parallel between the programmable logic device and the baseboard management and control device, the device control parameters are updated in real time through redundant links to ensure that the parameter consistency can be maintained when any link is abnormal, so as to achieve the purpose of enhancing the robustness of data synchronization between controllers, thereby achieving the technical effect of control logic redundancy fault tolerance and rapid recovery of abnormal scenarios.
[0012] Furthermore, a dynamic authority recovery strategy based on real-time anomaly detection is adopted, and the reverse switching of control rights is triggered through the dual monitoring mechanism of sensor devices and communication link status, so as to achieve the purpose of actively taking over the equipment management responsibilities when the baseboard management control device operates abnormally, thereby realizing the dual improvement of the system fault self-healing capability and equipment control continuity.
[0013] Finally, a complete copy of the device control parameters is retained through a multi-link synchronization mechanism, and the parameters are automatically aligned before and after the control is switched, thereby avoiding device status jumps or configuration loss due to permission changes, thereby achieving long-term protection of device control stability and data consistency throughout the life cycle of the server. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0015] Figure 1 A schematic diagram of an application environment of an optional server collaborative control method provided in an embodiment of the present application;
[0016] Figure 2 A flowchart of an optional server collaborative control method provided in an embodiment of the present application;
[0017] Figure 3 A partial schematic diagram of the hardware structure of an optional server collaborative control method provided in an embodiment of the present application;
[0018] Figure 4 A schematic diagram of the hardware structure design of an optional server collaborative control method provided in an embodiment of the present application;
[0019] Figure 5A schematic diagram of the structure of an optional server collaborative control device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0021] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0023] Optionally, in combination with the specific application environment architecture or the specific hardware architecture on which the execution of the server collaborative control method depends, the specific application environment architecture or the specific hardware architecture is described herein.
[0024] This embodiment may include but is not limited to being deployed in a rack server, and the hardware architecture includes the following modules:
[0025] Programmable logic chip: The first processing unit runs a real-time operating system, responsible for parsing temperature sensor data and generating cooling strategies; the second processing unit is implemented through hardware logic to directly control the modulation signal of the target device (such as a cooling fan); the shared storage area is used to cache historical control parameters and fault logs, and synchronize data with the baseboard management controller through a double buffer mechanism.
[0026] Baseboard management controller: The policy engine dynamically adjusts the long-term operation threshold of the cooling device based on the server load prediction model; the status monitoring module verifies the communication status of the programmable logic chip through the heartbeat signal (interval 100ms).
[0027] Communication link: The main channel uses the I²C protocol to transmit control instructions, and the backup channel uses the SPI protocol to achieve redundant communication; the independent alarm channel is directly connected to the system monitoring terminal through the GPIO (General-Purpose Input / Output) pin.
[0028] The application framework of this application uses programmable logic devices as the core execution body to build a multi-stage dynamic control authority switching system. The specific implementation process is as follows:
[0029] S1, server startup phase control authority initialization:
[0030] After the PLD is powered on on the server, it first completes the loading of its own firmware and the initialization of the operating system, and then collects the operating parameters of the target device in real time through the sensor device, including key indicators such as temperature, voltage, and fan speed. Based on the preset logic rules, the PLD performs threshold comparison and trend analysis on the sensor data. If any parameter is detected to be out of the safe range or there is a risk of mutation, the control authority takeover instruction is immediately triggered. During the takeover process, the PLD switches the driver configuration of the target device (such as the power module and the cooling unit) from the default state to the local management mode, and dynamically adjusts the device operating parameters based on the sensor data to ensure the stability of the hardware environment during the server startup phase. At the same time, the PLD sends an initialization status signal to the baseboard management control device through the preset communication link to start the latter's boot program.
[0031] S2, the authority transfer after the baseboard management control device is started:
[0032] When the baseboard management control device completes startup and enters the standby state, the programmable logic device initiates a two-way verification request through a preset communication link. The verification content includes the firmware version number, operation status code and encryption certificate validity of the baseboard management control device. If the verification is successful, the programmable logic device synchronizes the control parameters of the current target device to the baseboard management control device through at least one communication link, and starts the parameter consistency verification process. After the verification is successful, the programmable logic device sends an authority transfer instruction to the baseboard management control device, closes the local device driver interface, and switches the real-time data stream of the sensor device to the data buffer of the baseboard management control device. During the transfer process, the programmable logic device continuously monitors the response delay of the baseboard management control device and the correctness of the instruction execution to ensure that the authority transfer is uninterrupted.
[0033] S3, control permission recovery and restoration when operation is abnormal:
[0034] During the normal operation of the baseboard management and control device, the programmable logic device monitors the health status of the baseboard management and control device in real time through periodic heartbeat signals and two-way data packet verification of the preset communication link. If a heartbeat timeout is detected, the verification error rate exceeds the threshold, or the key device parameters deviate from the preset range, the programmable logic device determines that an abnormality has occurred in the baseboard management and control device and immediately initiates the permission recovery process. First, the programmable logic device forcibly suspends the device access rights of the baseboard management and control device through a backup communication link (such as a GPIO interrupt channel), and then obtains the latest device control parameters from the local cache or the shared memory of the baseboard management and control device, and reloads the driver configuration. After the permission recovery is completed, the programmable logic device takes over the real-time control of the target device and performs adaptive adjustments based on the sensor data until the baseboard management and control device fault is repaired and passes the secondary verification.
[0035] S4, multi-communication link redundancy and parameter synchronization guarantee:
[0036] At least two heterogeneous communication links are established between the programmable logic device and the baseboard management control device, such as a preset link based on a low-speed serial bus for regular data transmission, and a backup link based on a high-speed parallel interface for emergency control command transmission. During the authority transfer, parameter synchronization and abnormal recovery process, the programmable logic device transmits key data concurrently through multiple links, and uses a difference verification algorithm to ensure data consistency. If a link transmission delay or bit error rate abnormality is detected, the programmable logic device automatically switches to the backup link to maintain the real-time and reliability of control parameter synchronization. In addition, the programmable logic device persists multiple historical versions of the storage device control parameters in the local non-volatile memory to quickly roll back to a safe configuration in extreme failure scenarios.
[0037] Through the above framework, programmable logic devices achieve flexible management of control permissions throughout the entire life cycle of the server, solving the system vulnerability problem caused by dependence on a single controller. At the same time, through multi-link redundancy and dynamic synchronization mechanisms, the continuity of device control and data integrity are significantly improved.
[0038] An embodiment of the present application provides a server collaborative control method, and the method is described in detail in conjunction with the execution process of the server collaborative control method.
[0039] Optionally, as an optional implementation, as Figure 2 As shown, the above server collaborative control method includes:
[0040] S202, in response to the server starting up, loading the operating system, taking over the control authority of the target device associated with the server according to the operating parameters corresponding to the sensor device, wherein the server is deployed with a server controller, and the server controller includes a programmable logic device, a sensor device, and a baseboard management control device;
[0041] S204, in response to the baseboard management control device having completed startup, performing interactive verification with the baseboard management control device based on a preset communication link, and transferring the control authority to the baseboard management control device if the verification passes;
[0042] S206, in response to detecting an abnormality in the baseboard management control device, retaking control authority, wherein the programmable logic device and the baseboard management control device synchronize control parameters of the target device via at least one communication link, and the at least one communication link includes a preset communication link.
[0043] Optionally, in an embodiment of the present application, the baseboard management and control device may include but is not limited to an embedded hardware management unit, which is usually integrated on the mainboard for monitoring and managing the server hardware status, including but not limited to functions such as temperature sensor data acquisition, fan speed control, power supply status monitoring, and fault alarm. The core function of the baseboard management and control device is to achieve real-time monitoring of the underlying hardware of the server through a dedicated chip or module independent of the main processor, and to maintain remote management capabilities even if the main processor is in the shutdown state. For example, during the server startup phase, the baseboard management and control device can independently complete hardware self-test and environmental parameter initialization; during the operation phase, it continuously collects sensor data and dynamically adjusts the cooling strategy through a polling mechanism; in a fault scenario, it triggers an alarm signal or starts a redundant hardware switching process.
[0044] Optionally, in an embodiment of the present application, the above-mentioned programmable logic device may include but is not limited to an integrated circuit chip with hardware reconfiguration capability, such as a field programmable gate array (FPGA) or a complex programmable logic device (CPLD). The core feature of this type of device is that it can dynamically configure the internal logic circuit through a hardware description language to achieve customized design of specific functional modules. For example, in a server management scenario, a programmable logic device can be divided into a dual processing domain of a processing system (PS) and a programmable logic (PL): the PS part runs a real-time operating system to perform data processing and decision-making tasks, and the PL part implements low-latency operations such as hardware signal detection and PWM (Pulse Width Modulation) control signal generation. Its typical application scenarios include power supply timing management, hardware signal status acquisition, rapid response to fault signals, etc., and supports data interaction and collaborative control between PS and PL through an AXI (Advanced eXtensible Interface) bus.
[0045] Optionally, in the embodiment of the present application, the above-mentioned sensor device may include but is not limited to a physical quantity detection device for collecting server operating environment parameters, including but not limited to a temperature sensor, a voltage sensor, a current sensor, a vibration sensor, and a humidity sensor. For example, the temperature sensor can be connected to a multiplexer via an I2C (Inter-Integrated Circuit) bus to achieve time-sharing access to data from multiple temperature measurement points; the voltage sensor can monitor the voltage fluctuations of the CPU (Central Processing Unit), memory, and hard disk power supply circuit; the vibration sensor can detect abnormal vibrations of the fan or hard disk mechanical parts. The deployment locations of these sensor devices can cover key server modules, such as near the processor heat sink, the output end of the power module, the hard disk backplane area, etc., and transmit the data to the baseboard management control device or the programmable logic device in the form of digital signals or analog signals for joint analysis.
[0046] It should be noted that the design of the preset communication link can be expanded according to actual needs, such as using the I2C bus to implement status query of low-speed devices, transmitting hardware interrupt signals through the GPIO interface, or using the PCIe high-speed channel to synchronize large-scale control parameters. In scenarios where hardware resources are limited, a time-sharing multiplexing mechanism can be used to share physical links, such as using a multiplexer to time-share the same set of I2C lines to serve baseboard management control devices and programmable logic devices. In addition, the hierarchical design of the communication protocol can include physical layer signal transmission, data link layer verification mechanism, and application layer instruction encapsulation, such as implementing a CRC (Cyclic Redundancy Check) verification module in the PL logic to ensure signal integrity, or defining instruction interaction specifications in JSON (JavaScript Object Notation) format in the PS system.
[0047] It should be noted that the synchronization mechanism of control parameters can include technical implementations in multiple dimensions, such as using a shared memory area to achieve fast exchange of sensor calibration data, synchronizing PWM duty cycle parameters through hardware register mapping, or establishing a dual-port RAM (Random Access Memory) to achieve secure data transmission under an asynchronous clock domain. In terms of parameter update strategies, an active push mode based on event triggering, a passive update mode based on periodic polling, or an incremental synchronization mode based on difference comparison can be designed. For example, when the baseboard management and control device detects that the ambient temperature changes exceed the threshold, it can actively send the updated cooling strategy parameters to the programmable logic device; and during the system startup phase, the programmable logic device can batch read the initialization configuration parameters of the baseboard management and control device.
[0048] It should be noted that the construction of the fault recovery process can implement differentiated strategies based on the fault level, such as using a retry mechanism for momentary interruptions in the communication link, triggering spare parts switching for permanent damage to the hardware module, or initiating emergency frequency reduction protection for temperature runaway. Specific recovery methods may include hardware watchdog reset of PL logic, software hot restart process of PS system, or power supply switching based on redundant power supply modules. For example, when an abnormality of the baseboard management control device is detected, the programmable logic device can execute a three-level recovery strategy: first try to wake up the baseboard management control device through a hardware reset signal, and take over control of key equipment if it fails three times in a row, and report the fault log to the remote operation and maintenance platform through the out-of-band management interface.
[0049] In an exemplary embodiment, during the construction of the server controller, hardware-level fault isolation and functional coordination can be achieved by dividing the programmable logic device into dual processing domains of the processing system (PS) and the programmable logic (PL). Specifically, after the PS unit loads the real-time operating system, it can perform complex decision-making tasks such as sensor data fusion analysis and heat dissipation algorithm calculation; and the PL unit, with its hardware parallel processing capabilities, can achieve millisecond-level response control of devices such as fan speed and power switch. The advantage of this heterogeneous architecture is that when the baseboard management control device loses response due to software crash or hardware failure, the programmable logic device can seamlessly take over the control of the device through a preset interrupt trigger mechanism to ensure the continuous operation of key server functions.
[0050] During the control transfer process, the design of the interactive verification mechanism directly affects the system stability. After the baseboard management control device completes startup, it will exchange state vectors and check codes with the programmable logic device through multiple types of communication links. For example, the heartbeat signal transmitted by the GPIO pin can be used to verify the smoothness of the physical link, the encrypted hash value stored in the shared memory area can ensure the consistency of the software state, and the PWM phase synchronization operation ensures a smooth transition of the speed when the fan control is switched. This multi-level verification strategy effectively avoids the risk of false switching due to single-point verification failure, and at the same time, through the two-way synchronization mechanism of the control parameters, the two management units always maintain policy consistency.
[0051] When the system detects an abnormal operation of the baseboard management control device, the hardware interrupt response mechanism of the programmable logic device will be activated immediately. After the PL unit identifies the abnormal signal through the status monitoring circuit, it will send an interrupt request to the PS unit to trigger the preset fault recovery process. This process can include standardized operation modules such as sensor data takeover, control strategy switching, and fault log recording. For example, in the fan control scenario, the PS unit will load the backup cooling strategy from the persistent storage, reconfigure the PWM controller parameters of the PL unit through the AXI bus, and start the periodic self-check thread to try to restore the communication connection of the baseboard management control device.
[0052] For example, assuming that a server equipped with a redundant power supply and an intelligent cooling system in a data center is used as an example, the specific implementation process includes but is not limited to the following steps:
[0053] S1, control authority takeover during server startup:
[0054] When the server is powered on, the programmable logic device loads the operating system kernel, initializes the sensor interface and collects the target device operating parameters in real time. For example, if the sensor detects that the temperature of a key component exceeds the preset threshold, or the power supply voltage fluctuates beyond the safe range, the programmable logic device triggers the control authority takeover logic and takes over the control of the heat dissipation module and the power supply unit. At this time, the programmable logic device dynamically adjusts the device operation mode based on real-time data, such as increasing the operating intensity of the heat dissipation component to reduce the temperature, or limiting the power supply output power to maintain voltage stability. At the same time, the programmable logic device sends an initialization instruction to the baseboard management control device through a preset communication link to start its boot process. If the sensor data is within the normal range, the default control strategy is maintained until the baseboard management control device is ready.
[0055] S2, baseboard management control device startup verification and authority transfer:
[0056] After the baseboard management and control device completes startup, the programmable logic device initiates a two-way verification process through a preset communication link, including identity authentication and status verification. For example, the programmable logic device sends an encrypted challenge code, and the baseboard management and control device needs to generate a response code based on a specific algorithm and return it, while verifying whether its firmware version meets the requirements. After the verification is passed, the programmable logic device synchronizes the current device control parameters to the baseboard management and control device through multiple heterogeneous communication links, and starts a data consistency check. If it is detected that the parameter deviation exceeds the preset range, a synchronization retry is triggered until the data is consistent. After the verification is completed, the programmable logic device closes the local control interface, switches the sensor data stream to the designated storage area of the baseboard management and control device, and completes the authority transfer.
[0057] S3, abnormal detection and permission recovery of baseboard management control devices:
[0058] During the operation of the baseboard management and control device, the programmable logic device monitors its health status through periodic heartbeat signals and communication link status. For example, if it detects multiple consecutive heartbeat response timeouts, or key device parameters continue to deviate from the set values, the programmable logic device determines that the baseboard management and control device is abnormal. At this time, the programmable logic device forcibly suspends the device access rights of the baseboard management and control device through the backup communication link, obtains the latest control parameters from the local cache or multi-link synchronization data, and reloads the device driver configuration. If the baseboard management and control device cannot respond due to a serious fault, the programmable logic device directly reads the sensor raw data through the backup communication link, takes over the control of the device and performs adaptive adjustment.
[0059] S4, multi-link parameter synchronization and redundancy recovery:
[0060] Multiple heterogeneous communication links are maintained between the PLD and the BMC, including low-speed basic links and high-speed data links. During the authority transfer phase, the PLD transmits core control parameters through the basic link and batches historical operation data through the high-speed link. When the BMC recovers from an abnormality, the PLD writes back the parameters updated during the takeover to the non-volatile storage area of the BMC through the backup link, and uses a data verification algorithm to ensure transmission integrity. If the verification fails, switch to other available links to retransmit the data until synchronization is successful.
[0061] Through the above embodiments, the programmable logic device dynamically takes over the device control authority based on real-time sensor data during the server startup phase, thereby solving the device management blind spot problem caused by the main controller startup delay in the traditional architecture; through multi-link redundant communication and encryption verification mechanism, the data security and control continuity of the authority transfer process are ensured, and the risk of single point failure is reduced; in abnormal scenarios, combined with multi-modal detection strategy and heterogeneous link redundant synchronization mechanism, the rapid recovery of device control rights and parameter consistency maintenance are realized, thereby significantly improving the stability and fault tolerance of the server system under complex working conditions.
[0062] Through the embodiments of the present application, during the server startup phase, the programmable logic device actively takes over the control authority of the target device based on the operating parameters of the sensor device. By prepending the device management responsibility to the programmable logic device with rapid response capabilities, the stable operation of the device is ensured before the baseboard management control device completes the initialization, thereby avoiding device management vacuum or parameter configuration conflict caused by differences in controller startup sequence, thereby achieving the technical effect of seamless connection of control rights during the server startup phase and reliable management and control of device status.
[0063] In addition, a mechanism for interactive verification between the preset communication link and the baseboard management and control device is adopted, and the functional integrity and communication reliability of the baseboard management and control device are ensured through two-way communication verification during the transfer of authority, so as to eliminate the potential logic errors or link abnormality risks before the transfer of control rights, thereby achieving the technical effect of safe transition of control rights and improved system initialization efficiency. At the same time, through the parameter synchronization mechanism of multiple communication links in parallel between the programmable logic device and the baseboard management and control device, the device control parameters are updated in real time through redundant links to ensure that the parameter consistency can be maintained when any link is abnormal, so as to achieve the purpose of enhancing the robustness of data synchronization between controllers, thereby achieving the technical effect of control logic redundancy fault tolerance and rapid recovery of abnormal scenarios.
[0064] Furthermore, a dynamic authority recovery strategy based on real-time anomaly detection is adopted, and the reverse switching of control rights is triggered through the dual monitoring mechanism of sensor devices and communication link status, so as to achieve the purpose of actively taking over the equipment management responsibilities when the baseboard management control device operates abnormally, thereby realizing the dual improvement of the system fault self-healing capability and equipment control continuity.
[0065] Finally, a complete copy of the device control parameters is retained through a multi-link synchronization mechanism, and the parameters are automatically aligned before and after the control is switched, thereby avoiding device status jumps or configuration loss due to permission changes, thereby achieving long-term protection of device control stability and data consistency throughout the life cycle of the server.
[0066] As an optional solution, in response to the server starting up, the operating system is loaded, and the control authority of the target device associated with the server is taken over according to the sensor data corresponding to the sensor device, including:
[0067] In response to the server starting, outputting a modulation signal through a second processing unit in the programmable logic device to enable the target device to enter an initial operation state;
[0068] Periodically acquiring operating parameters through a first processing unit in a programmable logic device;
[0069] The first processing unit calculates the target parameter based on the operating parameter using a preset algorithm, and writes the target parameter into a control register corresponding to the second processing unit;
[0070] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;
[0071] Before the baseboard management control device completes startup, the target parameters are continuously updated in response to changes in the operating parameters.
[0072] Optionally, in an embodiment of the present application, the above-mentioned modulation signal may include but is not limited to a signal type that realizes energy control by adjusting the pulse width, and its core principle is to adjust the output power equivalently by changing the time ratio of the high level and the low level of the pulse. For example, in the server equipment control scenario, the modulation signal can be used to adjust the fan speed, control the voltage output of the power module, or adjust the cooling efficiency of the heat dissipation system. The specific implementation method includes a hardware timer generating a fixed-frequency fundamental signal, changing the effective output power through a duty cycle adjustment circuit, and dynamically adjusting the pulse parameters based on a feedback mechanism. During the server startup phase, the second processing unit can generate a modulation signal with a specific duty cycle to make the target device (such as a cooling fan) smoothly transition from a static state to a preset speed range, thereby avoiding hardware damage caused by current shock.
[0073] Optionally, in the embodiment of the present application, the above-mentioned initial operating state may include but is not limited to the preset working conditions that the target device needs to achieve during the system startup phase, including but not limited to the voltage stability range, the reference parameters of the mechanical component movement, and the handshake state of the communication link. The construction of this state usually relies on the initialization control sequence output by the programmable logic device, and gradually guides the device into the controllable operating range through phased progressive parameter adjustment.
[0074] Optionally, in an embodiment of the present application, the above-mentioned operating parameters may include but are not limited to physical quantities or logical quantities that reflect the real-time working status of the target device, including but not limited to temperature, voltage, current, speed, delay time and error counter value. For example, in a heat dissipation control scenario, the operating parameters may include the processor package temperature, the surface temperature difference of the heat sink, the current fan speed and the air pressure value of the heat dissipation duct; in a power management scenario, it involves input voltage ripple, output current effective value and power conversion efficiency. These parameters are collected by sensors or built-in monitoring circuits of the device, and after analog-to-digital conversion or digital signal processing, they are periodically read and recorded by the operating system of the first processing unit.
[0075] Optionally, in an embodiment of the present application, the above-mentioned preset algorithm may include but is not limited to a mathematical model or rule set for generating a control strategy, including but not limited to proportional integral differential control, fuzzy logic control and model predictive control. For example, in a temperature control scenario, the preset algorithm can predict the heat dissipation demand based on the historical temperature change rate and dynamically adjust the fan speed curve; in a power load balancing scenario, the optimal power supply phase combination can be calculated based on the current distribution model. The implementation methods of these algorithms include mathematical library function calls, hardware acceleration module implementations, or neural network inference engine deployments, and their output results are used to update the control register parameters of the second processing unit.
[0076] Optionally, in the embodiment of the present application, the control register may include but is not limited to a storage unit for storing device control parameters, and its physical implementation form includes but is not limited to a trigger array, a static random access memory or a dedicated register file. For example, in a fan speed control scenario, the control register may contain a target speed setting value, an acceleration limit parameter and a fault protection threshold; in power module management, the output voltage reference value, the overvoltage protection trigger level and the soft start time configuration are stored.
[0077] Optionally, in an embodiment of the present application, the above-mentioned adaptive control mode may include but is not limited to a working mode that dynamically adjusts the control strategy according to environmental changes, including but not limited to parameter self-tuning, dynamic loading of rule bases, and reconstruction of control structures. For example, in a cooling system abnormality scenario, the adaptive control mode can automatically switch to a backup cooling strategy, increase the upper limit of the fan speed, and enable the auxiliary refrigeration unit; in a power supply fluctuation scenario, the voltage compensation coefficient is dynamically adjusted and the collaborative working logic of the multi-phase power supply module is reconstructed. The activation of this mode depends on the threshold comparison result preset in the control register or the mode switching instruction issued by the first processing unit.
[0078] It should be noted that the generation method of the modulation signal can be expanded according to the hardware resources, such as using a dedicated hardware timer to achieve nanosecond-level precision signal output, building a multi-channel synchronous signal generator through a programmable logic unit, or using a software interrupt service routine to simulate a pulse waveform. The signal frequency selection range can cover the kilohertz to megahertz range to adapt to the response characteristics of different devices, such as low-speed mechanical parts are suitable for low-frequency wide-band regulation, and high-speed digital circuits are suitable for high-frequency fine control.
[0079] It should be noted that the method for constructing the initial operating state can be flexibly designed according to the device type. For example, mechanical devices adopt a progressive parameter loading strategy, first releasing the brake and then gradually increasing the drive signal strength; electronic devices implement a phased power supply strategy, completing the core voltage establishment, clock synchronization and interface initialization in sequence. For heterogeneous device groups, a parallel initialization process can be designed to synchronously control multiple devices through a multiplexing mechanism, or to start them in critical order using a priority queue. Exception handling mechanisms may include timeout retries, alternate parameter loading, and faulty device isolation. For example, when the hard disk home calibration times out, it automatically switches to the alternate head positioning algorithm or marks the device as offline.
[0080] It should be noted that the collection mechanism of operating parameters can include a variety of data fusion methods, such as deploying multiple sensors for redundant measurement of the same physical quantity, eliminating outliers through the median filtering algorithm; implementing joint analysis of related parameters, such as calculating the heat dissipation efficiency by combining ambient temperature and fan speed. The data transmission path can be designed as a direct memory access channel, shared buffer polling, or interrupt-driven mode. For example, high-priority parameters (such as overtemperature alarms) are reported instantly using hardware interrupts, and regular parameters are transferred in batches through DMA. Parameter storage strategies can include circular buffers to record historical data, snapshots to save critical states, or compressed archiving of long-term operation logs.
[0081] Through the embodiments of the present application, the intelligent transition of equipment management is realized through the hierarchical control architecture. The refined adjustment of the modulation signal reduces the impact current of the equipment startup process and prolongs the service life of mechanical components. The multi-dimensional collection and fusion analysis of operating parameters improves the dynamic adjustment accuracy of the control strategy. The coordinated design of the preset algorithm and the control register shortens the delay between strategy calculation and hardware response to the microsecond level. The introduction of the adaptive control mode improves the stability of the system when facing sudden load fluctuations or environmental changes.
[0082] It should be noted that in terms of device type expansion, it can support a smooth transition from traditional mechanical hard disks to all-flash arrays; in terms of algorithm evolution, it allows the deployment of new control models through online updates; in terms of communication interfaces, it is compatible with traditional I2C and GPIO protocols, and can also be expanded to support high-speed SerDes interfaces. For server systems of different sizes, the elastic scaling of control accuracy can be achieved by adjusting the resource configuration of programmable logic devices, and a seamless connection can be established between single-node control and cluster management.
[0083] As an optional solution, periodically obtaining the operating parameters through the first processing unit in the programmable logic device includes:
[0084] Using a multiplexed channel to collect data of a plurality of sensor devices as operating parameters in a time-sharing manner through a first processing unit;
[0085] De-noising the abnormal data in the operating parameters by the first processing unit, and marking invalid data segments;
[0086] The valid data is stored in the shared cache area through the first processing unit, so as to be called by the first processing unit during calculation.
[0087] Optionally, in an embodiment of the present application, the above-mentioned multiplexed channel may include but is not limited to a technical solution for realizing time-sharing transmission of multiple signals through a shared physical line, and its core principle is to realize a single transmission medium carrying multiple independent data streams through time slice rotation or address encoding. For example, in a sensor network scenario, the multiplexed channel can be designed as a cyclic acquisition architecture based on time division multiplexing, and a high-speed switching switch is used to sequentially connect the temperature sensor, voltage sensor and speed sensor to the same analog-to-digital converter. In specific implementation, the programmable logic device can generate a channel selection signal, switch different sensors to access the acquisition circuit at millisecond intervals, and cooperate with the sampling and holding circuit to maintain signal stability. This design can not only reduce the hardware resource usage, but also realize the synchronous timestamp marking of multiple types of sensor data. For example, in a server cabinet monitoring system, cyclic acquisition of 12 temperature monitoring points is realized through 8 multiplexed channels, and each channel is allocated a 2 millisecond exclusive sampling window.
[0088] Optionally, in an embodiment of the present application, the above-mentioned abnormal data may include but is not limited to sensor measurement values that exceed a preset reasonable range or a combination of associated data with logical contradictions, including but not limited to distorted signals caused by transient spike interference, sensor drift errors, and communication link errors. For example, in a temperature acquisition scenario, abnormal data may appear as a single sampling value that suddenly changes from the previous data and exceeds a set Celsius threshold, or a linear increase for three consecutive sampling cycles that contradicts the environmental heat dissipation strategy. In a current monitoring scenario, if the current value of a phase power supply is continuously zero while the loads of other phases are normal, it may be an abnormality caused by sensor failure. The identification of abnormal data can rely on threshold comparison, sliding window variance analysis, or machine learning model inference.
[0089] Optionally, in an embodiment of the present application, the above-mentioned shared cache area may include but is not limited to a common storage area for data exchange between multiple processing units or processes, and its physical implementation may include a dual-port memory, a ring buffer, or a dynamic memory pool with a mutex lock. For example, in a heterogeneous processing architecture, the shared cache area can be designed as an on-chip memory block accessible to both the first processing unit (PS) and the second processing unit (PL), and a ping-pong buffer structure is used to implement the pipeline operation of collection and processing. In a specific application, when the PS unit performs data analysis, the PL unit continuously writes new collection data to the backup buffer, and immediately switches the read and write pointers after the PS processing is completed. The cache management strategy may include data block verification, dynamic allocation of read and write permissions, and overflow protection mechanisms, such as allocating independent storage pages for temperature data and adding cyclic redundancy check codes, and automatically enabling data compression storage when the PL unit detects that the PS has not read the data in time.
[0090] It should be noted that the design of multiplexed channels can be flexibly adjusted according to system requirements, such as adopting a priority channel allocation strategy to reserve fixed time slices for key sensors, while secondary sensors dynamically allocate remaining resources. The abnormal data processing method can be expanded according to the data type and system fault tolerance requirements, such as using sliding window median filtering for transient interference, implementing online calibration compensation for sensor drift, and enabling forward error correction mechanism for communication errors; in the correlation analysis dimension, a device status model can be established to verify the rationality of the data, such as when the fan speed increases, the temperature of the corresponding heat dissipation area should have a downward trend; in terms of disposal strategies, data interpolation reconstruction, abnormal mark ignoring or triggering equipment re-inspection processes can be designed. For example, for occasional current spike data, linear interpolation replacement of adjacent sampling points is used; for persistent abnormal temperature readings, redundant sensor cross-validation is started and a device diagnostic report is generated.
[0091] Through the embodiments of the present application, the intelligent scheduling of multiplexed channels is utilized to improve the efficiency of sensor data collection, while reducing hardware costs. The dynamic management strategy of the shared cache area reduces the risk of data loss, thereby improving the stability of the system under sudden data peaks.
[0092] It should be noted that the data acquisition architecture has good scalability. In terms of the expansion of the number of sensors, multi-level node access can be achieved by increasing the multiplexing channel level or adopting a tree topology structure; in the data processing dimension, it supports smooth upgrades from simple threshold alarms to complex machine learning model analysis; in terms of real-time requirements, it can not only meet the second-level monitoring needs, but also adapt to millisecond-level real-time control scenarios by optimizing buffer strategies. For heterogeneous computing platforms, plug-and-play of acquisition modules and different processing units can be achieved through standardized interfaces.
[0093] As an optional solution, the first processing unit calculates the target parameters based on the operating parameters using a preset algorithm, and writes the target parameters into a control register corresponding to the second processing unit, including: generating an adjustment instruction based on the changing trend of the operating parameters by the first processing unit; converting the adjustment instruction into a control instruction recognizable by the second processing unit by the first processing unit, wherein the target parameters include the control instruction; and writing the control instruction into the control register by the first processing unit.
[0094] Optionally, in an embodiment of the present application, the above-mentioned change trend may include, but is not limited to, characteristic indicators of the rate and direction of change of operating parameters through historical data sequence analysis, which are usually used to predict the state evolution trend of the equipment. For example, in a heat dissipation control scenario, the temperature change trend may be manifested as a linear growth of 0.5 degrees Celsius per minute, or an exponential accelerated temperature rise feature; in a power management scenario, the current fluctuation gradient may include a periodic oscillation amplitude attenuation or a sudden spike surge pattern. Methods for calculating change trends include differential method, sliding window linear regression and wavelet transform analysis, and the output results are used to determine whether the device is in a stable state, transition state or abnormal state. In a specific application, when the surface temperature gradient of the radiator is detected to exceed 0.2 degrees Celsius per second, the system can trigger the fan speed-up strategy in advance to avoid thermal runaway.
[0095] It should be noted that the conversion process of control instructions can adapt to a variety of optimization strategies, such as using a table lookup method to achieve nonlinear instruction mapping, improving instruction resolution through interpolation algorithms, or introducing a dead zone compensation mechanism to eliminate the influence of mechanical clearance of the actuator. Instruction encoding methods may include direct numerical writing, relative value incremental update, or conditional trigger instruction packets, such as encoding the target speed value as a 16-bit binary number with a 2-bit checksum and a 1-bit emergency brake flag. The transmission protocol design can consider the hierarchical processing of real-time requirements, with key instructions directly transmitted through dedicated hardware channels and non-critical instructions transmitted in batches through a shared bus. For example, over-temperature protection instructions are transmitted instantly through GPIO pins, while fan speed fine-tuning instructions are updated periodically in batches through the I2C bus.
[0096] In addition, the implementation form of the hardware acceleration module can be flexibly configured according to system resources, such as deploying a fully parallel computing array when logic resources are sufficient, and adopting a time-division multiplexing architecture in resource-constrained scenarios. The acceleration task allocation strategy may include static task binding, dynamic load balancing, or a hybrid scheduling mode. For example, the fixed-cycle control law calculation is bound to a dedicated hardware unit, while the burst data analysis task is assigned to a reconfigurable logic block. In terms of energy efficiency optimization, dynamic voltage and frequency adjustment, idle module clock gating, or adaptive adjustment of calculation accuracy can be designed. For example, when the system load is less than 30%, the operating voltage of the acceleration module is automatically reduced to save 15% power consumption. The interface standardization design allows modules to be migrated and reused between different processing units, such as defining a unified AXI stream interface specification so that the same convolution acceleration core can serve the PS unit and the PL unit.
[0097] Through the embodiments of the present application, the execution efficiency of complex control algorithms is improved, while the computing load of the main processor is reduced. The standardized conversion mechanism of control instructions enhances the maintainability of the system and shortens the access adaptation time of equipment from different manufacturers.
[0098] As an optional solution, the above method further includes at least one of the following:
[0099] Accelerating the operation of the target parameter by using a parallel computing unit deployed in the first processing unit;
[0100] The data stream corresponding to the operating parameter is received by the first processing unit using a high-speed data interface.
[0101] Optionally, in an embodiment of the present application, the above-mentioned parallel computing unit may include but is not limited to a computing module that implements multi-task synchronous processing through hardware architecture design, and its core principle is to improve the computing throughput by increasing the number of physical computing cores or optimizing the data path structure. For example, in a programmable logic device, the parallel computing unit can be designed as a plurality of groups of multiplication accumulator arrays to simultaneously perform matrix multiplication operations in a closed-loop model; or configured as a pipeline structure to split the iterative calculation of the control algorithm into multiple stages and advance in parallel. In specific implementation, computing resources can be dynamically configured for different control scenarios: in the temperature control scenario, 8 groups of parallel integrators are deployed to simultaneously process the thermodynamic equations of multiple heat dissipation areas; in the power management scenario, a 16-way parallel voltage fluctuation prediction unit is used to analyze the status of multi-phase power supply modules in real time. This design shortens the control model operation that originally required milliseconds to complete to microseconds, significantly improving the efficiency of generating control instructions.
[0102] Optionally, in an embodiment of the present application, the above-mentioned high-speed data interface may include but is not limited to a physical channel or protocol standard that supports high-bandwidth, low-latency data transmission, and its design goal is to minimize the transmission delay of data from the acquisition end to the processing end. For example, it is designed as a wide-bit parallel bus to synchronously transmit multi-sensor fusion data through a 128-bit data path. In specific applications, the interface can integrate a hardware flow control mechanism, such as setting a ping-pong buffer at the data receiving end to ensure uninterrupted processing of continuous data streams; implementing a priority scheduling strategy at the transport layer to give priority to the transmission of key parameters (such as over-temperature alarm signals). In the server power management scenario, such an interface can support real-time transmission of tens of thousands of current sampling values per second, providing high-refresh-rate input data for closed-loop control.
[0103] It should be noted that the implementation form of the high-speed data interface can be adjusted according to the system topology. For example, a star topology is used to connect multiple sensor nodes in a centralized architecture, and a ring bus is deployed in a distributed system to realize data relay transmission. At the protocol level, it can support flexible definition from the underlying electrical signal to the upper layer data encapsulation, such as a custom frame structure containing timestamps, checksums and data payloads. The error handling mechanism may include forward error correction coding, retransmission request strategy and redundant path switching. For example, when three consecutive data packet check failures are detected, it automatically switches to the backup transmission channel and triggers error logging. The bandwidth allocation strategy can be designed as static reservation (fixed allocation of 20% bandwidth for key sensors) or dynamic adjustment (real-time allocation of the remaining bandwidth according to the burstiness of the data flow).
[0104] In addition, the writing mechanism of the control register can adapt to a variety of security protection requirements, such as implementing double-signature verification to ensure the legitimacy of the write instruction, using redundant registers to implement atomic operations, or designing shadow registers to support parameter preloading and seamless switching. Register access permission management can include hierarchical control strategies, such as only allowing the hardware acceleration module to write key control fields, while the debug interface only opens status read permissions; in terms of timing control, the write enable signal can be designed to be strictly aligned with the clock edge to avoid metastable problems. The functional extension of the register group can include state machine control bits, mode switching flags, and exception capture registers, such as setting a 4-bit status code to identify the current control mode, and a 32-bit exception vector register to record the types and occurrence times of the last 10 errors.
[0105] Through the embodiments of the present application, the deployment of parallel computing units is utilized to increase the computing speed and shorten the key control instruction generation cycle. The optimized design of the high-speed data interface reduces the transmission delay and supports the real-time processing of sensor signals. The safe write mechanism of the control register reduces the parameter update error rate and improves the consistency of the overall control response of the system. The synergy of these three enhances the ability to suppress power supply fluctuations and improves the efficiency of the cooling system when the server is running at full load, achieving industry-leading stability indicators.
[0106] It should be noted that the hardware acceleration architecture is highly scalable. In terms of computing power expansion, processing power can be improved by increasing the number of parallel units or upgrading the process technology. For emerging technology needs, such as quantum computing control or photon communication management, adaptation can be achieved by replacing the core algorithm of the computing unit and upgrading the interface optoelectronic conversion module. This modular design concept enables the solution to meet existing server control needs while reserving space for technology upgrades for the evolution of future intelligent computing infrastructure.
[0107] As an optional solution, in response to the baseboard management control device having completed startup, interactive verification is performed with the baseboard management control device based on a preset communication link, and control authority is transferred to the baseboard management control device if the verification passes, including:
[0108] In response to the baseboard management control device having completed startup, receiving a permission transfer request sent by the baseboard management control device through a preset communication link;
[0109] Verify the validity of the authority transfer request, and if the verification result is valid, synchronize the historical control parameters and operating parameters to the baseboard management control device;
[0110] Performing a phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device;
[0111] When the control authority is transferred, the working mode of the programmable logic device is switched to the auxiliary monitoring mode.
[0112] Optionally, in an embodiment of the present application, the above-mentioned authority transfer request may include but is not limited to a standardized signal or data packet for triggering the transfer of control rights, which contains key information such as a request identifier, a timestamp, a summary of the current device status, and an encrypted verification credential. For example, in the server control transfer scenario, the request can be actively initiated by the baseboard management and control device after completing the self-test, and transmitted to the programmable logic device through a preset communication link to trigger the control right verification process. In specific implementation, the request content may include the hardware serial number, firmware version number, and temporary token generated during the startup phase of the baseboard management and control device to ensure the legitimacy and timeliness of the request source. At the data transmission level, a frame encryption strategy can be adopted to split the request into multiple data segments and attach a cyclic redundancy check code to avoid information tampering caused by communication link interference.
[0113] Optionally, in an embodiment of the present application, the above-mentioned historical control parameters may include but are not limited to control strategy records generated by the programmable logic device during the period of holding control authority, including but not limited to device operation mode configuration, dynamic adjustment parameters and exception handling strategies. For example, in a cooling system control scenario, historical control parameters may include fan speed curves corresponding to different temperature ranges, load balancing strategies for power supply modules, and degraded operation plans for hard disk arrays. These parameters are usually stored in a non-volatile memory in the form of a time series, including a timestamp, a parameter version number, and an associated sensor data snapshot, to ensure that the complete control logic evolution process can be traced after the baseboard management control device takes over.
[0114] Optionally, in an embodiment of the present application, the above-mentioned phase matching operation may include but is not limited to a timing alignment mechanism to ensure a smooth transition between the old and new control strategies, the core goal of which is to eliminate parameter jumps or logic conflicts during the transfer of control. For example, in the power supply module control scenario, phase matching needs to ensure that the voltage regulation instructions generated by the baseboard management control device and the historical output values of the programmable logic device remain continuous in amplitude, frequency and phase. Specific operations may include parameter interpolation transition (such as gradually switching to new parameters within the next three control cycles), timing synchronization (aligning the execution rhythm of the two control systems through hardware clocks) and safety threshold verification (ensuring that the new parameters do not exceed the tolerance range of the equipment).
[0115] Optionally, in an embodiment of the present application, the auxiliary monitoring mode may include but is not limited to a lightweight working mode in which the programmable logic device continues to run after the control right is transferred, and its core functions include anomaly detection, data backup, and redundant control preparation. For example, in this mode, the programmable logic device can continue to collect sensor data and compare and analyze the control effect of the baseboard management control device. When it is detected that the temperature adjustment response delay exceeds the threshold, an early warning signal is automatically triggered; at the same time, a shadow copy of the key control parameters is maintained to support rapid control right switching back when the baseboard management control device is abnormal. The resource utilization rate of this mode is usually controlled below 10% to ensure that the main hardware resources can be released for other tasks.
[0116] It should be noted that the triggering conditions of the authority transfer request can be expanded and designed according to system requirements, such as using time-driven triggering (initiated with a fixed delay of 5 seconds after the baseboard management control device is started), event-driven triggering (initiated when the programmable logic device load rate is detected to be lower than 20%), or a mixed triggering strategy (simultaneously satisfying the firmware version matching and network topology consistency conditions). The encapsulation format of the request content can be adapted to different security level requirements, such as using a multi-signature mechanism in a trusted execution environment and using hash verification in a normal environment; the transmission protocol can support unicast, multicast or broadcast modes, and adapt to centralized or distributed control architectures. The exception handling mechanism may include timeout retry (aborting the transfer if there is no response to three consecutive requests), conflict resolution (starting priority arbitration when concurrent requests from multiple nodes are detected) and safe rollback (restoring the original control state after the transfer fails).
[0117] Among them, the synchronization mechanism of historical control parameters can be designed with a variety of optimization strategies, such as incremental synchronization (transmitting only the most recently changed parameters), compressed transmission (runtime encoding of repeated parameters) or difference merging (generating incremental packages by comparing version numbers). The data transmission process can implement end-to-end encryption, fragmentation verification and breakpoint resumption (recording the address interval that has been successfully transmitted). The parameter storage structure can include a time series database, key-value pair storage or a relational table structure. For example, a parameter partition table is established by device type, and each partition contains a timestamp index, parameter value and associated metadata (such as a generation algorithm identifier).
[0118] Through the embodiments of the present application, a standardized authority transfer process is used to improve the security and reliability of the control switching process, and the historical parameters and calibration data are fully synchronized to ensure the continuity of the control strategy while controlling the parameter fluctuation amplitude when the device state switches within an acceptable range. The introduction of the phase matching mechanism reduces the coordination error between the new and old control systems and significantly enhances the stability of equipment operation. The auxiliary monitoring mode builds a dual protection system, which prolongs the system's mean time between failures and shortens the recovery time from major failures.
[0119] It should be noted that the control transfer architecture has good expansion and adaptation capabilities. In terms of device compatibility, it supports smooth access to baseboard management and control devices from traditional x86 architecture to ARM architecture; in terms of security protocols, it can be expanded to support new security mechanisms such as national secret algorithms and quantum encryption; in terms of control scale, it can not only meet the needs of single-node servers, but also support cluster-level control transfer through protocol expansion. For emerging intelligent operation and maintenance scenarios, machine learning models can be integrated to achieve intelligent decision-making on the timing of handover, or digital twin technology can be used to rehearse the risks of the handover process. This flexible design concept lays a technical foundation for the fully automated operation and maintenance of future intelligent data centers.
[0120] As an optional solution, performing a phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device, including:
[0121] Obtain the signal period and duty cycle of the control parameters output to the target device;
[0122] In response to receiving a status confirmation signal, a signal period and duty cycle are sent to the baseboard management control device so that the baseboard management control device generates alternative control parameters and completes the authority switching, wherein the status confirmation signal is used to indicate that the baseboard management control device has completed startup, and the alternative control parameters represent the parameters of the baseboard management control device controlling the target device.
[0123] Optionally, in an embodiment of the present application, the above-mentioned signal period may include but is not limited to the time interval between two adjacent identical phase points in the control parameter waveform, which is usually used to characterize the frequency characteristics of the control signal. For example, in a PWM fan control scenario, the signal period may be set to a 40 microsecond period value corresponding to 25kHz, and its high level duration (duty cycle) determines the fan speed; in a voltage regulation scenario, the signal period may correspond to the chopping frequency of the switching power supply, such as a 5 microsecond period corresponding to 200kHz. The capture of the signal period needs to be achieved through a hardware timer or a digital phase-locked loop, and specific methods include edge-triggered capture, zero-crossing detection, or related function calculation. In the phase matching operation, accurate measurement of the current signal period is a prerequisite for ensuring the frequency consistency of the new and old control parameters.
[0124] Optionally, in an embodiment of the present application, the above-mentioned duty cycle may include but is not limited to a parameter of the ratio of the effective level time of the control signal to the total period, which is the core control variable for adjusting the output power of the device. The capture of the duty cycle usually relies on a high-precision timer, and the effective proportion is calculated by measuring the time difference between the rising edge and the falling edge. During the phase matching process, the consistency of the duty cycle needs to be controlled within the error range to avoid the jump of the output state of the device. For multi-channel control scenarios, the duty cycle parameters may include the phase offset between the master and slave channels, and cross-comparison is required to ensure the overall waveform coordination.
[0125] Optionally, in an embodiment of the present application, the above-mentioned alternative control parameters may include but are not limited to a new parameter set generated by a baseboard management control device that has an equivalent control effect to the current operating parameters. The generation process must meet the requirements of time domain continuity, frequency domain consistency and logical compatibility. For example, in a power supply control scenario, alternative parameters may include equivalent voltage set values, soft start time constants, and overvoltage protection thresholds, but are implemented through different adjustment algorithms; in a mechanical control scenario, it is necessary to ensure that the acceleration curve corresponding to the alternative parameters is smoothly connected with the original parameters. The parameter generation method may include table lookup mapping, model equivalent conversion, or dynamic parameter compensation.
[0126] Optionally, in an embodiment of the present application, the above-mentioned status confirmation signal may include but is not limited to a two-way handshake protocol data packet for verifying the security of the control switching, including key information such as the device's current status summary, parameter check code, and switch-ready flag. For example, in a permission switching scenario, the status confirmation signal may be sent by the baseboard management control device to include a confirmation frame containing the current power supply phase, cooling strategy hash value, and device topology information, and the programmable logic device verifies the state consistency by comparing the shadow parameters in the memory. Signal transmission must meet real-time and reliability requirements, and usually adopts a combination of hardware-level confirmation mechanism (such as dedicated GPIO pin level jump) and software-level confirmation protocol (such as encrypted JSON data packet). The exception handling mechanism may include a three-way handshake retry, automatic calibration of difference parameters, or a switch abort rollback process.
[0127] It should be noted that the generation strategy of alternative control parameters can be adapted to a variety of algorithm frameworks, such as parameter mapping based on equivalent conversion of transfer functions, generation of optimized parameters through reinforcement learning models, or simulation verification of parameter effects using digital twin technology. Parameter equivalence verification can include time domain response comparison (such as step response overshoot difference is less than 2%), frequency domain characteristic analysis (such as cutoff frequency offset is less than 5%), or logic compatibility check (such as protection threshold completely covers the original parameter range). In dynamic environment scenarios, parameter compensation mechanisms can be designed, such as automatically adjusting the gain coefficient of alternative parameters according to temperature changes, or loading degradation compensation parameters according to the degree of equipment aging. The interactive protocol of the status confirmation signal can be extended to support a variety of security mechanisms, such as dynamic token authentication (each session generates a unique verification code), two-way digital certificate verification, or physical unclonable function signature. The transport layer design can include redundant confirmation mechanisms (such as three-way handshake + heartbeat detection), differential data synchronization (transmitting only changed parameters), or hierarchical confirmation strategies (immediate confirmation of key parameters and batch confirmation of non-key parameters). The exception handling process can be designed as a progressive response, for example, the first difference triggers parameter fine-tuning, the second difference starts log analysis, and the third difference executes control rollback. For high-reliability scenarios, a multi-node cross-validation mechanism can be deployed, requiring more than half of the verification nodes to return a success signal before the switch can be executed.
[0128] Through the embodiments of the present application, high-precision signal capture and intelligent parameter generation technology are used to reduce the fluctuation amplitude of the device state during the control right switch, and the average completion time of the phase matching operation is shortened. The double security guarantee of the state confirmation mechanism reduces the switching failure rate to less than one in 100,000, and the system reliability is significantly enhanced. These improvements enable key equipment to achieve "senseless switching" during the control right transfer process, providing a solid technical guarantee for high-availability server systems.
[0129] It should be noted that this phase-matching architecture has great expansion potential. In terms of signal type support, it can be expanded to adapt to various control parameters from analog voltage signals to digital pulse sequences; at the algorithm level, it supports the hybrid deployment of traditional control theory and AI models; in the security authentication dimension, it can seamlessly integrate cutting-edge technologies such as quantum encryption and blockchain verification. For distributed control systems, cross-node parameter synchronization can be achieved through protocol extension, supporting cluster-level control transfer of tens of thousands of devices. This modular design concept enables the solution to meet current server control needs, and also provides a reusable technical framework for precision control in future smart factories, autonomous driving and other fields.
[0130] As an optional solution, if the verification is passed, the historical control parameters and operating parameters are synchronized to the baseboard management control device, including:
[0131] If the verification is passed, the operating parameters are written into the shared storage area after calibration and encryption, wherein the baseboard management control device is configured to read and decrypt the operating parameters from the shared storage area in a memory access manner, and update the control parameters of the target device after verifying the data integrity.
[0132] Optionally, in an embodiment of the present application, the above-mentioned calibration encryption may include but is not limited to a processing flow for implementing encryption protection after error correction on the raw data of the sensor, the core of which is to ensure both data accuracy and transmission security. For example, in the temperature parameter calibration encryption scenario, the original sampling value is first corrected according to the nonlinear error table of the sensor, and then a symmetric encryption algorithm is used to encrypt the calibrated data by adding a timestamp and a sensor identifier. In specific implementation, a sharding encryption mechanism can be designed to split a single piece of data into multiple ciphertext segments and store them in a dispersed manner to prevent the data from being intercepted in its entirety. In the voltage monitoring scenario, calibration encryption may involve zero drift compensation, and the data after gain adjustment is then digitally signed using an elliptic curve encryption algorithm to ensure that the data source is credible and cannot be tampered with.
[0133] Optionally, in the embodiment of the present application, the above memory access may include but is not limited to a data transmission technology between a peripheral device and a memory without the intervention of a central processor, which realizes high-throughput data handling through a dedicated channel. For example, in a parameter synchronization scenario, the baseboard management control device can complete data migration at a rate of gigabytes per second by configuring the source address (shared storage area) and the target address (local cache) of the controller.
[0134] Optionally, in an embodiment of the present application, the above-mentioned verification of data integrity may include but is not limited to technical means for verifying that the data has not been accidentally modified or maliciously tampered with during transmission and storage. For example, in the verification parameter synchronization scenario, a cyclic redundancy check can be used to calculate the check code of the data block and compare it with the check value generated before transmission; or a hash algorithm can be used to generate a data fingerprint, and the legitimacy of the data source can be verified through a digital signature. In scenarios with a higher level of security, a multi-layer verification mechanism can be deployed: the physical layer corrects single-bit errors through Hamming codes, the transport layer verifies multi-bit errors, and the application layer implements digital signatures to verify data authenticity. When a verification failure is detected, the system can automatically trigger disposal strategies such as data retransmission, alarm, or switching to an alternate data source.
[0135] Optionally, in an embodiment of the present application, the update of the above-mentioned control parameters may include but is not limited to the process of deploying verified new parameters to the control system of the target device, and it is necessary to ensure that the update process is smooth and safe. For example, in a fan control scenario, the parameter update may include a gradual transition strategy: within the control cycle, the target speed is linearly transitioned from the current value to the new set value to avoid mechanical stress caused by sudden changes in speed. The update mechanism can be designed as a double buffer structure, in which the new parameters are first written to the shadow register, and then switched to take effect through atomic operations after complete verification. For critical equipment, a rollback mechanism can also be deployed, which automatically restores to the previous stable parameter version when an abnormal device status is detected after the update.
[0136] It should be noted that in terms of algorithm selection, symmetric encryption, asymmetric encryption or lightweight encryption can be used; in terms of accuracy assurance, sensor characteristic curve fitting calibration, multi-sensor data fusion calibration or online self-learning calibration can be deployed; in terms of data encapsulation format, fixed-length data packets, structures or custom binary protocols can be designed. For example, a segmented encryption strategy is used for temperature data, and the integer part and the decimal part are encrypted separately and then combined for transmission; after frequency domain feature extraction of vibration signals, the energy values of key frequency bands are selectively encrypted.
[0137] In addition, in terms of channel allocation, you can set priority preemptive channels (critical data is transmitted first), polling channels (bandwidth resources are evenly distributed), or adaptive channels (dynamically adjusted according to data traffic); in terms of transmission mode, single transmission, block transmission, or hash-aggregate transmission are supported; in terms of exception handling mechanisms, transmission timeout interrupts, data alignment error detection, or buffer overflow protection can be designed. For example, in high-speed data synchronization scenarios, configure the channel to ring buffer mode, and when the baseboard management control device lags behind in reading speed, automatically overwrite the earliest historical data and record overflow events.
[0138] In addition, Hamming code verification is implemented at the physical layer to correct transmission errors, verification is deployed at the link layer to detect data block integrity, and digital signatures are used to verify the source of data at the application layer. The verification failure handling strategy can include hierarchical responses: single-bit errors trigger automatic error correction, multi-bit errors start the retransmission mechanism, and signature verification failure isolates data and generates security alarms. For example, a triple verification mechanism is implemented for key control parameters. When two consecutive verifications fail, it automatically switches to the redundant control module and triggers an operation and maintenance alarm.
[0139] Through the embodiments of the present application, the data error of the parameter synchronization process is reduced by utilizing the coordinated design of calibration encryption and integrity verification, while achieving data transmission anti-tampering capability to meet security standards. The optimization of direct memory access technology improves data handling efficiency and shortens the delay in updating key parameters. The combination of the double buffer update mechanism and the progressive transition strategy controls the fluctuation amplitude during device state switching within a reasonable range, significantly enhancing system stability.
[0140] It should be noted that this parameter synchronization architecture has excellent expansion and adaptation capabilities. At the encryption algorithm level, it can be seamlessly upgraded to a lattice encryption solution that is resistant to quantum computing; in the storage architecture dimension, it supports migration from traditional memory to new non-volatile storage media; in terms of integrity verification, it can be expanded to deploy AI-based data anomaly detection models. For heterogeneous computing environments, the solution can adapt to processors with different instruction set architectures and achieve cross-platform data synchronization through standardized interfaces. This flexibility enables it to meet the management needs of traditional servers, and can also be expanded to precision parameter management scenarios for edge computing nodes and industrial control equipment.
[0141] As an optional solution, in response to detecting that an abnormality occurs in the baseboard management control device, taking back the control authority includes:
[0142] Detecting whether the communication signal associated with the baseboard management control device is lost;
[0143] In case of continuous detection of communication signal loss and the cumulative number of signal losses reaches a preset threshold, the emergency control procedure is initiated;
[0144] In response to the emergency control program being started, reconstructing the control strategy of the target device based on the historical operation data and outputting the emergency control parameters, wherein during the reset of the baseboard management control device, the programmable logic device performs the function control and abnormal alarm of the target device;
[0145] After the baseboard management and control device is restored, interactive verification is performed with the baseboard management and control device based on a preset communication link, and the control authority is transferred to the baseboard management and control device if the verification passes.
[0146] Optionally, in an embodiment of the present application, the above-mentioned communication signal loss may include but is not limited to the data transmission interruption state of the preset communication link between the baseboard management control device and the programmable logic device, and its detection means include heartbeat signal timeout, continuous check code errors and physical layer level abnormalities. For example, in a communication scenario based on the I2C protocol, signal loss may be manifested as the bus being in a low-level state for more than 500 milliseconds, or the master device not receiving a response signal from the slave device for three consecutive times; in an Ethernet communication scenario, signal loss can be determined by request timeout or connection reset. The detection module usually deploys a hardware watchdog timer and a software state machine for joint monitoring, for example, detecting the arrival of a heartbeat packet every 50 milliseconds, while monitoring physical layer signal quality indicators (such as eye opening, bit error rate).
[0147] Optionally, in an embodiment of the present application, the above-mentioned preset threshold may include but is not limited to the cumulative number of abnormal events or duration parameters that trigger fault judgment, and its setting needs to balance the system sensitivity and anti-interference ability. For example, in a communication signal detection scenario, the preset threshold may be set to 5 consecutive heartbeat signal losses or a cumulative 3-second communication interruption; in a temperature abnormality scenario, it may be defined as a sensor that has been overheated for 10 consecutive sampling cycles. The threshold setting method may include static configuration (such as setting according to the equipment manual), dynamic learning (automatic adjustment based on historical operating data) or a mixed mode (basic threshold superimposed on the environmental correction coefficient). In the server heat dissipation control scenario, the threshold may be dynamically reduced as the ambient temperature rises to trigger the protection mechanism in advance.
[0148] Optionally, in an embodiment of the present application, the above-mentioned emergency control program may include but is not limited to a set of fault response logic pre-stored in a programmable logic device, including equipment safety state maintenance, redundant module switching and fault isolation strategy. For example, when a baseboard management control device failure is detected, the emergency program may include: immediately locking the fan speed in a safe range, shutting down the power supply of non-critical peripherals, and enabling the backup sensor data link. The program execution process is usually designed as a state machine mode, including three stages: fault diagnosis, control strategy switching, and safety state maintenance. A timeout fallback mechanism is set for each stage. In the power supply abnormality scenario, the emergency program can perform voltage reduction operation, load unloading and emergency shutdown operations in a hierarchical manner.
[0149] Optionally, in an embodiment of the present application, the reconstruction of the above control strategy may include but is not limited to the process of reversely deducing the current device control parameters based on historical operating data, the core of which is to establish a mapping relationship between the device state and the control effect. For example, in a heat dissipation control scenario, the reconstruction strategy can establish a speed prediction model based on ambient temperature by analyzing the corresponding relationship between temperature and fan speed in the past 24 hours; in power management, a dynamic voltage adjustment table can be generated based on historical load curves. The reconstruction algorithm may include sliding window regression analysis, time series prediction, or case-based reasoning, such as using a neural network to predict the heat dissipation demand in the next 5 minutes and generate preventive control parameters.
[0150] Optionally, in the embodiment of the present application, the above-mentioned function control and abnormal alarm may include but is not limited to the ability set to maintain the basic operation capability of the equipment in an emergency state and report fault information to the outside. For example, function control may include: maintaining the processor core power supply within a safe voltage range, limiting the memory access bandwidth to prevent overheating, enabling the backup network interface, etc.; abnormal alarms are implemented through indicator light mode changes, system log records, out-of-band management interface push alarm codes, etc. The alarm information usually includes the fault type code, occurrence timestamp, impact range and recommended disposal measures.
[0151] It should be noted that in bus-type protocols, the bus idle time and the number of error frames can be detected; in packet switching networks (such as Ethernet), the aging status and response timeout can be monitored; in wireless communication scenarios, a comprehensive judgment must be made based on signal strength and bit error rate. The fault-tolerant mechanism can be designed as multi-path detection (monitoring hardware heartbeat signals and software protocol status simultaneously), time window sliding detection (allowing instantaneous interruptions but continuous abnormalities to trigger alarms), or confidence accumulation detection (the longer the abnormality lasts, the higher the calculated fault probability value). For example, in critical power supply control scenarios, the I2C bus data and the power supply module status register are monitored simultaneously, and an abnormality in any channel triggers a pre-alarm.
[0152] In addition, for transient faults (such as communication interruptions), an automatic recovery mechanism can be designed (trying to reinitialize the link); for permanent faults (such as hardware damage), the equipment is triggered to run at a downgraded level; in uncertain fault scenarios, a tentative recovery process can be enabled (such as increasing the control intensity in stages). The program execution mode can include fully automatic processing, semi-automatic (key operations require operation and maintenance confirmation), or a mixed mode (automatic handling of basic faults, manual intervention for complex problems). For example, when it is detected that the temperature of the baseboard management control device is too high, the heat dissipation intensity is automatically increased first. If the abnormality persists, manual confirmation is requested to determine whether to force a restart.
[0153] In addition, after the first handover fails, the synchronization data verification time can be extended; when the second failure occurs, the communication channel is switched or the redundant interface is enabled; and the third failure triggers device isolation. The handover verification dimension can include firmware version compatibility check, control parameter boundary check, and device status consistency comparison. For example, in the handover preparation stage, the baseboard management control device needs to provide an encrypted token containing the current device topology. After the programmable logic device verifies the validity of the token, the parameter synchronization process is allowed to start.
[0154] Through the embodiments of the present application, the coordinated design of multi-level fault detection and intelligent emergency response is utilized to shorten the system abnormality recovery time, and the control continuity of key equipment meets the high availability standard. The control strategy reconstruction algorithm improves the accuracy of emergency parameter generation, and the performance fluctuation during equipment degradation operation is controlled within a reasonable range. The double-checked authority transfer mechanism improves the success rate of control switching and significantly enhances the overall reliability of the system.
[0155] It should be noted that the fault recovery architecture is highly scalable. In the detection dimension, the AI anomaly prediction model can be integrated to predict potential faults in advance; in the control strategy level, it supports smooth upgrades from rule engines to deep learning models.
[0156] As an optional solution,
[0157] Detecting whether the communication signal associated with the baseboard management control device is lost, including: periodically detecting the heartbeat signal of the baseboard management control device; if no valid heartbeat signal is received for N consecutive times, it is determined that the communication signal is lost, and the fault timestamp and the associated hardware status are recorded, where N is a positive integer greater than 2;
[0158] In response to the emergency control program being started, the control strategy of the target device is reconstructed based on the historical operation data, and the emergency control parameters are output, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future needs based on historical operation data, outputting emergency control parameters, and resetting abnormal communication links associated with the baseboard management control device.
[0159] Optionally, in an embodiment of the present application, the above-mentioned heartbeat signal may include but is not limited to a periodic verification signal for verifying the communication connection status between devices, and its core function is to maintain the active detection mechanism of the control system. For example, in the communication scenario between the baseboard management control device and the programmable logic device, the heartbeat signal can be designed as a specific data packet sent every 100 milliseconds, containing information such as the device identifier, the current operating mode and the check code. The generation method of the heartbeat signal may include hardware timer triggering, software task scheduling or event-driven mode. In the specific implementation, differential Manchester encoding may be used to enhance the anti-interference ability, or additional timestamps may be used to achieve two-way clock synchronization. In the abnormal detection scenario, the loss of the heartbeat signal may be caused by a variety of reasons such as physical link breakage, processor deadlock or buffer overflow, and a comprehensive diagnosis needs to be performed in combination with the associated hardware status.
[0160] Optionally, in an embodiment of the present application, the above-mentioned preset fault recovery algorithm library may include but is not limited to a set of intelligent recovery strategies pre-stored in a non-volatile memory, including core functional modules such as device control parameter generation, communication link repair, and safety boundary maintenance. For example, in a cooling system failure scenario, the algorithm library may include a fan speed prediction algorithm driven by a thermodynamic model, a temperature compensation algorithm based on historical data, and a multi-sensor data fusion algorithm. Its implementation form may include a mathematical function library, a hardware description language module, or a machine learning model, such as deploying a recurrent neural network to predict cooling requirements within the next 5 seconds, and generating emergency speed parameters in combination with fuzzy control rules. The calling mechanism of the algorithm library is usually designed as an event-driven mode, and the corresponding recovery strategy is dynamically loaded according to the fault type code.
[0161] Optionally, in an embodiment of the present application, the above-mentioned abnormal communication link may include but is not limited to a physical / logical channel with persistent transmission errors or interrupted logical connections, and its abnormal manifestations may include bit error rate exceeding the standard, protocol handshake failure, or signal integrity degradation. For example, in a communication scenario based on the I2C bus, the abnormality may be manifested as a clock stretching phenomenon caused by the bus being continuously pulled down. The cause of the link abnormality may involve electromagnetic interference, connector oxidation, or firmware logic errors, and the repair strategy may include hardware reset, protocol renegotiation, or redundant path switching.
[0162] It should be noted that in terms of signal form, fixed frequency square waves, encrypted data packets or modulated carrier signals can be used; in terms of fault tolerance mechanisms, adaptive heartbeat intervals (extending the period when the network is congested), multipath redundancy detection (monitoring the hardware level and protocol status simultaneously) or confidence accumulation models (the longer the abnormality lasts, the higher the probability of failure) can be designed. For example, in a high temperature environment, the heartbeat interval is automatically shortened to 50 milliseconds to improve detection sensitivity, and the hardware watchdog circuit is enabled as a backup monitoring method. The recovery strategy can include a progressive response: the first loss triggers link self-checking, the second loss starts parameter synchronization, and the third loss executes control switching.
[0163] In addition, the construction strategy of the fault recovery algorithm library can adapt to a variety of technical routes. In terms of algorithm type, it can include physical model-based feedforward control, data-driven feedback regulation and hybrid intelligent control; in terms of resource management, static preloading (all algorithms reside in memory), dynamic loading (read from memory on demand) or distributed deployment (part of the algorithm runs on hardware accelerators) can be used. For example, communication repair algorithms with high real-time requirements are solidified in programmable logic units, while complex control parameter generation algorithms run on processing systems. The learning mechanism can be designed as offline training online reasoning, incremental learning or transfer learning, such as continuously optimizing the weights of the neural network model based on the historical fault records of the equipment.
[0164] On the other hand, the implementation method of the communication link reset operation can be differentiated according to the type of abnormality. For abnormalities caused by transient interference, protocol layer reinitialization (renegotiation of baud rate and verification method) can be used; for failures caused by hardware damage, it is possible to switch to a backup communication interface or enable a degraded communication mode (such as switching Gigabit Ethernet to 100M mode). The reset verification mechanism can include link quality assessment (bit error rate test), end-to-end loopback test and load stress test, such as sending a test data packet after the reset is completed and verifying the end-to-end transmission delay and integrity.
[0165] Through the embodiments of the present application, the system's self-healing ability in abnormal conditions is significantly improved by utilizing intelligent heartbeat monitoring and multi-strategy recovery mechanisms. The dynamically loaded fault recovery algorithm achieves accurate matching of different fault scenarios, avoiding the limitations of traditional single strategies. The adaptive repair strategy of the communication link effectively reduces the frequency of manual intervention and ensures the stability of continuous operation of the equipment. The layered exception handling architecture takes into account both real-time response and in-depth analysis requirements to form a complete fault handling closed loop.
[0166] It should be noted that this solution has the ability to expand for future technological evolution. In terms of detection mechanism, it can integrate the collaborative monitoring capabilities of edge computing nodes; at the algorithm level, it supports plug-and-play of new artificial intelligence models; in terms of communication interface, it can adapt to new connection methods such as optical fiber communication and wireless transmission. For large-scale equipment clusters, cross-node fault collaborative processing can be achieved through protocol expansion, building an intelligent recovery network with self-organizing capabilities.
[0167] As an optional solution, the above method also includes:
[0168] During the reset of the baseboard management controller, at least one of the following operations is performed by the programmable logic device:
[0169] Take over the control authority of the target device and adjust the device parameters of the target device based on the real-time data of the sensor device;
[0170] Report fault information through the alarm channel;
[0171] Periodically attempt to reestablish a communication link with the baseboard management controller;
[0172] After the baseboard management controller is restored, the encryption identification verification and control parameter synchronization operations are re-executed to complete the authority transfer.
[0173] Optionally, in an embodiment of the present application, the above-mentioned alarm channel may include but is not limited to a fault information transmission mechanism, which reports the abnormal status to an external management system through a preset interface, including but not limited to sending encoded warning messages through a UART (Universal Asynchronous Receiver Transmitter) serial port, triggering an interrupt signal through a PCIe interface, or transmitting messages through an out-of-band management network.
[0174] It should be noted that the form of reporting fault information on the alarm channel can be expanded according to differences in system architecture. For example, in the transmission protocol dimension, the alarm information can be pushed to the cloud management platform through the protocol, or transmitted to the local monitoring terminal through the protocol, or interacted with dedicated hardware through a custom binary protocol. In the content format dimension, the alarm information may contain a concise fault code, a detailed log summary, or an attached sensor raw data snapshot. In the priority dimension, alarms can be classified into emergency alarms (such as power failures), important alarms (such as temperature exceeding the limit), and prompt alarms (such as communication delays), and associated with different response strategies. This application does not make specific restrictions on this.
[0175] Optionally, in an embodiment of the present application, the above-mentioned encrypted identification verification may include but is not limited to an identity verification process based on an asymmetric encryption algorithm, such as a programmable logic device generates a random challenge code, the baseboard management controller signs it with a private key and returns it, and both parties confirm the legitimacy of the identity by verifying the validity of the signature, including but not limited to supporting multiple encryption standards.
[0176] Optionally, in an embodiment of the present application, the above-mentioned control parameter synchronization may include but is not limited to transmitting device configuration data from a programmable logic device to a baseboard management controller during the authority transfer process, including but not limited to avoiding data loss through a double buffer mechanism, using verification to ensure transmission integrity, and improving synchronization efficiency through multi-link concurrent transmission.
[0177] It should be noted that the mechanism of periodically attempting to rebuild the communication link with the baseboard management controller can adapt to a variety of fault tolerance requirements. For example, in the retry strategy dimension, the programmable logic device can use fixed interval retries, exponential backoff retries, or dynamically adjust the retry frequency based on the load status. In the link selection dimension, the preset main link can be tried first, or the backup link can be automatically switched after the main link fails (such as switching from I2C to SPI (Serial Peripheral Interface)), or multiple links can be detected in parallel to select the optimal path. In the verification dimension, the programmable logic device can perform one-way heartbeat detection, two-way data integrity verification, or encryption handshake protocol verification after the link is rebuilt. This application does not make specific restrictions on this.
[0178] As an optional solution, the above method also includes:
[0179] In response to the server starting up, a modulation signal is outputted by a second processing unit in the programmable logic device to control the heat dissipation device to start up in a preset mode, wherein the target device includes the heat dissipation device;
[0180] Loading an operating system through a first processing unit in a programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes a temperature monitoring device;
[0181] Performing denoising and calibration processing on the operating parameters by a first processing unit to generate a valid temperature parameter set;
[0182] Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register;
[0183] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;
[0184] In response to the baseboard management control device having completed startup, receiving a handover request including an encrypted identifier sent by the baseboard management control device through a preset communication link;
[0185] Verify the validity of the encryption identifier through the second processing unit, and feed back the verification result to the first processing unit;
[0186] If the verification is passed, the historical control parameters, the temperature calibration data and the current speed regulation signal are synchronized to the baseboard management controller through the first processing unit, wherein the baseboard management controller generates a speed regulation signal matching the current heat dissipation state of the target device based on the historical control parameters, the temperature calibration data and the current speed regulation signal;
[0187] When the control authority is transferred, the programmable logic device is switched from the working mode to the auxiliary monitoring mode to continuously verify the operating status of the baseboard management controller;
[0188] When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, it takes over the control authority again, where N is a positive integer greater than 2;
[0189] Reset the baseboard management controller by calling a preset fault recovery algorithm library through the first processing unit, and output a speed regulation signal to the heat dissipation device through the second processing unit;
[0190] During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: reporting fault information through an alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller;
[0191] In response to the baseboard management controller being reset, encryption identification verification and control parameter synchronization operations are performed to complete the authority transfer.
[0192] For example, assuming that a server heat dissipation control system is used as an example, the process includes but is not limited to the following:
[0193] S1: Server power-on initialization phase:
[0194] After the second processing unit (PL) is started, a PWM signal is immediately generated to drive the cooling fan to start in a gradual mode. For different types of cooling devices, the preset mode parameters are dynamically adjusted.
[0195] S2: Operating system loading and data collection:
[0196] After the first processing unit (PS) loads the real-time operating system, it collects temperature data through multiplexed channel interval cycles. The data acquisition module uses sliding window filtering technology to perform median denoising on 8 consecutive sampling values of each channel. The calibration module implements online compensation based on the stored sensor offset table to generate a valid temperature parameter set.
[0197] S3: Dynamic control strategy generation:
[0198] Based on the valid temperature parameter set, the PS unit executes the PID control algorithm: when the CPU temperature exceeds 65°C, the target duty cycle is calculated based on the recent temperature rise rate, and the target value is written to the control register of the PL after boundary constraints.
[0199] S4: Preparation for transfer of control authority:
[0200] After the baseboard management controller (BMC) completes startup, it sends a handover request packet containing an encrypted identifier through the I2C bus. The identifier contains the BMC firmware version number, hardware ID, and timestamp. The PL's verification module completes decryption and hash verification, and after confirming the legitimacy of the request, it feeds back a verification success flag to the PS.
[0201] S5: Parameter synchronization and switching execution:
[0202] The PS transfers historical control parameters, temperature calibration data, and current PWM phase information to the BMC through the DMA channel. The BMC phase alignment module performs the following operations:
[0203] S5-1, lock the PWM clock source of PL and measure the phase difference to be 15ns;
[0204] S5-2, adjust the initial value of the internal timer to achieve waveform synchronization;
[0205] S5-3, complete the duty cycle gradual transition within 3 PWM cycles.
[0206] S6: Exception handling and recovery mechanism:
[0207] When the PL fails to detect the 100ms heartbeat signal of the BMC for 5 consecutive times (N=5), a hardware interrupt is triggered to the PS. The fault recovery algorithm library calls the LSTM prediction model to predict the cooling demand in the next 3 minutes based on the temperature data of the past hour and generate an emergency duty cycle curve. The hardware acceleration module (such as the DSP array of the FPGA) completes the calculation, and the PL outputs a speed control signal.
[0208] During the BMC reset, the PL takes over the thermal control authority and sends an alarm message containing a fault code (such as 0x5A3) every 5 seconds through an independent UART channel. The system attempts to rebuild the I2C link every 10 seconds until the BMC recovers and re-executes the complete handover process.
[0209] This application achieves stable control of the server cooling system throughout its life cycle through the collaborative design of progressive authority transfer and intelligent emergency response. The hardware-level generation of PWM signals ensures the speed of speed regulation response, the fusion processing of multi-source temperature data improves the control accuracy, and the encryption verification mechanism ensures the security of authority transfer. The dynamic algorithm is combined with historical data to effectively respond to sudden load fluctuations and environmental changes. The dual control architecture maintains cooling performance in fault scenarios and avoids the risk of system overheating due to single point failures. The modular design supports flexible expansion and can be adapted to a variety of cooling solutions such as air cooling and liquid cooling.
[0210] The following is a further explanation of this application with reference to specific examples:
[0211] With the large-scale application of servers in fields such as big data and Internet AI technology, the data processed by servers and the scale of server hardware in data centers have also shown explosive growth. The field of server management technology has put forward increasingly higher requirements for server reliability and availability. Traditional server management technology faces three challenges:
[0212] The baseboard management controller (BMC) is the core of server hardware management. It monitors the security of the entire server environment and the hardware status. Therefore, the reliability of the baseboard management controller firmware is the most important part to ensure the function of the entire server. At present, the industry usually uses embedded dedicated chips running the Linux operating system as the hardware management core. The management links relied on in the hardware single board design, such as I2C, ADC, PWM, and sensor management, all rely on the reliability of the BMC chip operation. If the BMC operation is abnormal or the hardware link fails due to software and hardware failures during the operation of the server, it may cause the server to work abnormally, such as the failure of the fan cooling function and the abnormal sensor function monitoring.
[0213] The server motherboard uses programmable logic device (FPGA / CPLD) chips as hardware chip power supply timing management, board hardware signal management and other functions. The baseboard management controller usually relies on CPLD to realize hardware signal detection, fault abnormality identification and other functions. At the same time, the programmable logic device itself supports SOC (System-on-Chip) function implementation, which can realize the integration of multiple functional modules such as storage, processing, logic and interface in CPLD or FPGA without increasing hardware costs.
[0214] Figure 3 A partial schematic diagram of the hardware structure of an optional server collaborative control method provided in an embodiment of the present application, such as Figure 3As shown in the figure, the system on a programmable logic device chip generally consists of two parts: the PS processor system and the PL programmable logic. This architecture enables this series of chips to have the advantages of FPGA hardware programmability while also having the advantages of ASIC in terms of energy consumption, performance and compatibility. At the same time, the processor system of the system on a chip supports running general operating systems such as Linux, which facilitates flexible software programming application development. The PS and PL are connected through a standard AXI interface, and the PL part can be managed as part of the peripheral in the software system.
[0215] Based on the above background content, the present application proposes a server controller management method based on a programmable system on chip. The method can utilize the heterogeneous programmable logic resources on the server motherboard and cooperate with the baseboard management controller chip to manage the server hardware, thereby improving the stability and availability of the server at a low hardware cost.
[0216] The present application proposes a server collaborative management method based on a heterogeneous programmable system on chip, which is mainly used in a typical storage server architecture. The server controller mainboard includes a programmable logic device FPAG that supports the on-chip SOC mode, a baseboard management controller BMC, and several temperature sensors. The hardware link between the BMC and the FPGA includes an I2C data access link, a GPIO hardware signal, and a PWM control signal. When the FPGA is used as an on-chip SOC system, the internal logic resources are divided into two parts, wherein the PL logic module realizes the server's hardware power supply control, hardware pin signal detection and other functions, and the PS part can use a built-in soft core to realize system data processing or a hard core with its own processor core to realize system data processing. In order to ensure the real-time performance of the system, a real-time operating system can be used as the running software system. The peripheral sensors are connected to the multi-channel I2C multiplexer chip of the previous stage through the I2C link. This multiplexer uses an I2C multiplexing chip that supports multi-channel access, and supports BMC / FPGA to access the sensor chip of the subsequent stage in time-sharing. Figure 4 A schematic diagram of the hardware structure design of an optional server collaborative control method provided in an embodiment of the present application, such as Figure 4 As shown, including but not limited to solving the following three scenarios:
[0217] Scenario 1: When the BMC system is not started in the boot scenario, the fan speed control can be completed by the FPGA independently, that is, the real-time operating system of the PS part can read the sensor data after the power-on fast startup (seconds) to perform basic speed regulation (calculate the PWM output of the fan speed according to the speed regulation algorithm), and then access the PWM output module in the PL through the AXI bus to output a reasonable cooling speed regulation fan speed. This process is generally completed within 2 minutes after the server is powered on.
[0218] Scenario 2: During server operation, the baseboard management controller (BMC) is reset and restarted, such as when the firmware is updated online / the software runs abnormally. When the communication between the FPGA and the BMC is interrupted (the state vector verification data fails), the PL module in the FPGA detects the BMC abnormality and resets and repairs the BMC. Before triggering the BMC reset action, the state vector verification module in the PL logic module triggers an interrupt and reports it to the PS operating system through the AXI bus. The application running on the PS triggers the interrupt handler to take over the fan.
[0219] Scenario 3: Solve the problem of low data access polling efficiency in the current baseboard management controller BMC. In traditional server firmware management, BMC accesses external chips or signals mostly through polling test. The above architecture can offload interrupt trigger detection and control hardware to the FPGA chip. The reset detection and repair of the mainboard core chip is completed by the combination of PL and PS, without the need for BMC, thereby improving the efficiency of fault recovery.
[0220] The present application implements a stable and reliable server controller hardware management function under the above-mentioned hardware structure. Due to the use of a new hardware design with heterogeneous redundancy, the server operation reliability and availability are improved without increasing the complexity of the hardware design, thereby improving the stability of the entire server operation.
[0221] This application adopts a three-level control transfer mechanism to achieve system heat dissipation speed management. The first stage is the system initialization stage. In this stage, FPGA takes over the system heat dissipation authority. The specific implementation is as follows:
[0222] based on Figure 4 In the server hardware management design block diagram, after the controller is powered on normally, the FPGA provides a fixed speed control fan by default. After the PS part inside the FPGA is loaded with the real-time operating system and started, the program running in the PS polls the test instance to read the temperature of the external core sensor. The PS application calculates the heat dissipation parameters (PID and PWM) based on the collected temperature data, enables the PL logic to take over the fan speed control logic, and outputs the PWM signal of the target speed.
[0223] Phase 2: BMC takeover phase:
[0224] After waiting for the BMC system to start up (about 90 seconds), the communication with the FPGA state vector verification module is successful. After the FPGA GPIO pin receives the BMC communication mark, it indicates that it has complete heat dissipation speed regulation capabilities. The FPGA PL logic releases the fan speed regulation control right, and the BMC inputs PWM to control and regulate the fan. At the same time, in order to speed up the control transfer process, the FPGA and BMC need to quickly synchronize the heat dissipation sensor data. The synchronization content includes exchanging sensor calibration data through the shared memory area, and then performing PWM phase synchronization to ensure that the BMC can smoothly take over the heat dissipation control. After the handover is completed, the BMC system gradually collects the ambient temperature and all temperature sensor information of the controller, and the FPGA returns to the monitoring mode.
[0225] Phase 3: Troubleshooting process:
[0226] When the FPGA internal PL module state vector verification mechanism detects three consecutive communication failures, it will trigger a Level-2 interrupt to the PS system. The PS starts the hardware controller program (preheating algorithm library) and reconfigures the PWM controller through the AXI-Lite bus. During the BMC reset and restart, the FPGA will take over the fan control function, temperature sensor monitoring and alarm functions;
[0227] In view of the problem of low BMC polling efficiency and slow response in scenario 3, this application proposes to use the advantages of heterogeneous hardware design to solve the problem of slow response of traditional BMC heat dissipation and speed regulation, accelerate fault processing by hardware and take over emergency fault scenarios, and adopt a dual-mode architecture of PL hardware fast response + PS intelligent decision-making. The specific implementation process is as follows:
[0228] S1, initialize the PL module programmable event processing unit (PEPU) and associate event-related interrupt pins;
[0229] S2, PS module initialization procedure, including threshold register set of emergency event signal (temperature / voltage / current, etc.);
[0230] S3, configure the event filtering parameters (sliding window size / confidence threshold) to ensure that emergency events are not triggered by anomalies;
[0231] S4, when the preset emergency event occurs, the PL emergency response action is directly triggered. According to the severity level of the event, the AXI interrupt is triggered to report to the PS module. When the severity level of the event is greater than the software processing, the PL can independently execute the fuse strategy, stop the operation of the abnormal function module and other strategies. If the event verification level requires PS analysis, after reporting the AXI interrupt, the PS will perform fault analysis and then process it, such as the result of fan speed control and other strategies.
[0232] The server collaborative management method based on heterogeneous programmable system-on-chip proposed in this application can bring the following beneficial effects:
[0233] Improved reliability: System switching time is less than 200ms when BMC fails (software restart or thread self-recovery in traditional solutions requires at least 30s);
[0234] Heterogeneous hardware design reuses FPGA resources, reduces the need for dedicated monitoring chips, and reduces BOM costs;
[0235] Noise improvement: Through more refined speed regulation methods and speed control transfer strategies, the server noise can be reduced and the user experience can be better;
[0236] Hardware accelerates fault handling, improves the overall system availability, and prevents hardware faults caused by factors such as low monitoring efficiency from spreading and causing other problems.
[0237] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0238] The embodiment of the present application also provides a server collaborative control device, such as Figure 5 As shown, the device comprises:
[0239] an initialization module, configured to load an operating system in response to a server startup, and take over control authority of a target device associated with the server according to operating parameters corresponding to the sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and a baseboard management control device;
[0240] A verification module, configured to perform interactive verification with the baseboard management control device based on a preset communication link in response to the baseboard management control device having completed startup, and transfer the control authority to the baseboard management control device if the verification passes;
[0241] A recovery module is used to take over the control authority again in response to detecting an abnormality in the baseboard management control device, wherein the control parameters of the target device are synchronized between the programmable logic device and the baseboard management control device via at least one communication link, and the at least one communication link includes the preset communication link.
[0242] As an optional solution, in response to the server starting up, the operating system is loaded, and the control authority of the target device associated with the server is taken over according to the sensor data corresponding to the sensor device, including:
[0243] In response to the server starting, outputting a modulation signal through a second processing unit in the programmable logic device to enable the target device to enter an initial operation state;
[0244] Periodically acquiring operating parameters through a first processing unit in a programmable logic device;
[0245] The first processing unit calculates the target parameter based on the operating parameter using a preset algorithm, and writes the target parameter into a control register corresponding to the second processing unit;
[0246] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;
[0247] Before the baseboard management control device completes startup, the target parameters are continuously updated in response to changes in the operating parameters.
[0248] As an optional solution, periodically obtaining the operating parameters through the first processing unit in the programmable logic device includes:
[0249] Using a multiplexed channel to collect data of a plurality of sensor devices as operating parameters in a time-sharing manner through a first processing unit;
[0250] De-noising the abnormal data in the operating parameters by the first processing unit, and marking invalid data segments;
[0251] The valid data is stored in the shared cache area through the first processing unit, so as to be called by the first processing unit during calculation.
[0252] As an optional solution, the first processing unit calculates the target parameter based on the operating parameter using a preset algorithm, and writes the target parameter into a control register corresponding to the second processing unit, including:
[0253] Generate an adjustment instruction based on a change trend of the operating parameter by the first processing unit;
[0254] converting the adjustment instruction into a control instruction recognizable by the second processing unit by the first processing unit, wherein the target parameter includes the control instruction;
[0255] The control instruction is written into the control register through the first processing unit.
[0256] As an optional solution, the above device further includes at least one of the following:
[0257] Accelerating the operation of the target parameter by using a parallel computing unit deployed in the first processing unit;
[0258] The data stream corresponding to the operating parameter is received by the first processing unit using a high-speed data interface.
[0259] As an optional solution, in response to the baseboard management control device having completed startup, interactive verification is performed with the baseboard management control device based on a preset communication link, and control authority is transferred to the baseboard management control device if the verification passes, including:
[0260] In response to the baseboard management control device having completed startup, receiving a permission transfer request sent by the baseboard management control device through a preset communication link;
[0261] Verify the validity of the authority transfer request, and if the verification result is valid, synchronize the historical control parameters and operating parameters to the baseboard management control device;
[0262] Performing a phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device;
[0263] When the control authority is transferred, the working mode of the programmable logic device is switched to the auxiliary monitoring mode.
[0264] As an optional solution, performing a phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device, including:
[0265] Obtain the signal period and duty cycle of the control parameters output to the target device;
[0266] In response to receiving a status confirmation signal, a signal period and duty cycle are sent to the baseboard management control device so that the baseboard management control device generates alternative control parameters and completes the authority switching, wherein the status confirmation signal is used to indicate that the baseboard management control device has completed startup, and the alternative control parameters represent the parameters of the baseboard management control device controlling the target device.
[0267] As an optional solution, if the verification is passed, the historical control parameters and operating parameters are synchronized to the baseboard management control device, including:
[0268] If the verification is passed, the operating parameters are written into the shared storage area after calibration and encryption, wherein the baseboard management control device is configured to read and decrypt the operating parameters from the shared storage area in a memory access manner, and update the control parameters of the target device after verifying the data integrity.
[0269] As an optional solution, in response to detecting that an abnormality occurs in the baseboard management control device, taking back the control authority includes:
[0270] Detecting whether the communication signal associated with the baseboard management control device is lost;
[0271] In case of continuous detection of communication signal loss and the cumulative number of signal losses reaches a preset threshold, the emergency control procedure is initiated;
[0272] In response to the emergency control program being started, reconstructing the control strategy of the target device based on the historical operation data and outputting the emergency control parameters, wherein during the reset of the baseboard management control device, the programmable logic device performs the function control and abnormal alarm of the target device;
[0273] After the baseboard management and control device is restored, interactive verification is performed with the baseboard management and control device based on a preset communication link, and the control authority is transferred to the baseboard management and control device if the verification passes.
[0274] As an optional solution,
[0275] Detecting whether the communication signal associated with the baseboard management control device is lost, including: periodically detecting the heartbeat signal of the baseboard management control device; if no valid heartbeat signal is received for N consecutive times, it is determined that the communication signal is lost, and the fault timestamp and the associated hardware status are recorded, where N is a positive integer greater than 2;
[0276] In response to the emergency control program being started, the control strategy of the target device is reconstructed based on the historical operation data, and the emergency control parameters are output, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future needs based on historical operation data, outputting emergency control parameters, and resetting abnormal communication links associated with the baseboard management control device.
[0277] As an optional solution, the above device also includes:
[0278] During the reset of the baseboard management controller, at least one of the following operations is performed by the programmable logic device:
[0279] Take over the control authority of the target device and adjust the device parameters of the target device based on the real-time data of the sensor device;
[0280] Report fault information through the alarm channel;
[0281] Periodically attempt to reestablish a communication link with the baseboard management controller;
[0282] After the baseboard management controller is restored, the encryption identification verification and control parameter synchronization operations are re-executed to complete the authority transfer.
[0283] As an optional solution, the above device also includes:
[0284] In response to the server starting up, a modulation signal is outputted by a second processing unit in the programmable logic device to control the heat dissipation device to start up in a preset mode, wherein the target device includes the heat dissipation device;
[0285] Loading an operating system through a first processing unit in a programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes a temperature monitoring device;
[0286] Performing denoising and calibration processing on the operating parameters by a first processing unit to generate a valid temperature parameter set;
[0287] Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register;
[0288] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;
[0289] In response to the baseboard management control device having completed startup, receiving a handover request including an encrypted identifier sent by the baseboard management control device through a preset communication link;
[0290] Verify the validity of the encryption identifier through the second processing unit, and feed back the verification result to the first processing unit;
[0291] If the verification is passed, the historical control parameters, the temperature calibration data and the current speed regulation signal are synchronized to the baseboard management controller through the first processing unit, wherein the baseboard management controller generates a speed regulation signal matching the current heat dissipation state of the target device based on the historical control parameters, the temperature calibration data and the current speed regulation signal;
[0292] When the control authority is transferred, the programmable logic device is switched from the working mode to the auxiliary monitoring mode to continuously verify the operating status of the baseboard management controller;
[0293] When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, it takes over the control authority again, where N is a positive integer greater than 2;
[0294] Reset the baseboard management controller by calling a preset fault recovery algorithm library through the first processing unit, and output a speed regulation signal to the heat dissipation device through the second processing unit;
[0295] During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: reporting fault information through an alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller;
[0296] In response to the baseboard management controller being reset, encryption identification verification and control parameter synchronization operations are performed to complete the authority transfer.
[0297] The description of the features in the embodiment corresponding to the above-mentioned server collaborative control device can be found in the relevant description of the embodiment corresponding to the server collaborative control method, which will not be repeated here one by one.
[0298] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned server collaborative control method embodiments.
[0299] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server collaborative control method embodiments when running.
[0300] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0301] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned server collaborative control method embodiments are implemented.
[0302] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps in any of the above-mentioned server collaborative control method embodiments.
[0303] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0304] The above is a detailed introduction to a server collaborative control method, storage medium and electronic device provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A server collaborative control method, characterized in that: Applied in programmable logic devices, including: In response to the server starting up, the operating system is loaded, and the control authority of the target device associated with the server is taken over according to the operating parameters corresponding to the sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device and the baseboard management control device; In response to the baseboard management control device having completed startup, performing interactive verification with the baseboard management control device based on a preset communication link, and transferring the control authority to the baseboard management control device if the verification passes; In response to detecting an abnormality in the baseboard management control device, the control authority is taken back, wherein the control parameters of the target device are synchronized between the programmable logic device and the baseboard management control device via at least one communication link, and the at least one communication link includes the preset communication link.
2. The server collaborative control method according to claim 1, characterized in that: The method of responding to the server startup, loading the operating system, and taking over the control authority of the target device associated with the server according to the sensor data corresponding to the sensor device includes: In response to the server being started, outputting a modulation signal through a second processing unit in the programmable logic device so as to enable the target device to enter an initial operation state; Periodically acquiring the operating parameter through a first processing unit in the programmable logic device; Calculate the target parameter by the first processing unit based on the operating parameter using a preset algorithm, and write the target parameter into a control register corresponding to the second processing unit; activating the adaptive control mode according to the control register and taking over the control authority by the second processing unit; Before the baseboard management control device completes startup, the target parameter is continuously updated in response to the change of the operating parameter.
3. The server collaborative control method according to claim 2, characterized in that: The periodically acquiring the operating parameter by the first processing unit in the programmable logic device includes: Using the first processing unit to collect data of the plurality of sensor devices in a time-sharing manner using a multiplexed channel as the operating parameter; Performing denoising on abnormal data in the operating parameters by the first processing unit, and marking invalid data segments; The first processing unit stores valid data in a shared cache area for use by the first processing unit during calculation.
4. The server collaborative control method according to claim 2, characterized in that: The calculating the target parameter by the first processing unit based on the operating parameter using a preset algorithm, and writing the target parameter into a control register corresponding to the second processing unit, includes: generating, by the first processing unit, an adjustment instruction based on a change trend of the operating parameter; converting the adjustment instruction into a control instruction recognizable by the second processing unit by the first processing unit, wherein the target parameter includes the control instruction; The control instruction is written into the control register through the first processing unit.
5. The server collaborative control method according to claim 4, characterized in that: The method further comprises at least one of the following: Accelerating the calculation of the target parameter by using a parallel computing unit deployed in the first processing unit; The data stream corresponding to the operating parameter is received by the first processing unit using a high-speed data interface.
6. The server collaborative control method according to claim 1, characterized in that: In response to the completion of the startup of the baseboard management control device, interactive verification is performed with the baseboard management control device based on a preset communication link, and the control authority is transferred to the baseboard management control device if the verification passes, including: In response to the baseboard management control device having completed startup, receiving a permission transfer request sent by the baseboard management control device through the preset communication link; Verifying the validity of the authority transfer request, and if the verification result is valid, synchronizing the historical control parameters and the operating parameters to the baseboard management control device; Performing a phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device; When the control authority is transferred, the working mode of the programmable logic device is switched to the auxiliary monitoring mode.
7. The server collaborative control method according to claim 6, characterized in that: The performing of the phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management control device matches the control parameter generated by the programmable logic device includes: Acquire a signal period and a duty cycle of a control parameter output to the target device; In response to receiving a status confirmation signal, the signal period and the duty cycle are sent to the baseboard management control device so that the baseboard management control device generates alternative control parameters and completes the authority switching, wherein the status confirmation signal is used to indicate that the baseboard management control device has completed startup, and the alternative control parameters represent the parameters of the baseboard management control device controlling the target device.
8. The server collaborative control method according to claim 6, characterized in that: When the verification is passed, synchronizing the historical control parameters and the operating parameters to the baseboard management control device includes: If the verification is successful, the operating parameters are written into the shared storage area after being calibrated and encrypted, wherein the baseboard management control device is configured to read and decrypt the operating parameters from the shared storage area in a memory access manner, and update the control parameters of the target device after verifying the data integrity.
9. The server collaborative control method according to claim 1, characterized in that: In response to detecting that the baseboard management control device is abnormal, retaking the control authority includes: Detecting whether a communication signal associated with the baseboard management control device is lost; In the event that the communication signal is continuously detected to be lost and the cumulative number of times the signal is lost reaches a preset threshold, an emergency control procedure is initiated; In response to the emergency control program being started, reconstructing the control strategy of the target device based on historical operation data and outputting emergency control parameters, wherein during the resetting of the baseboard management control device, the programmable logic device performs function control and abnormal alarm of the target device; After the baseboard management control device is restored, interactive verification is performed with the baseboard management control device based on the preset communication link, and the control authority is transferred to the baseboard management control device if the verification passes.
10. The server collaborative control method according to claim 9, characterized in that: The detecting whether the communication signal associated with the baseboard management control device is lost includes: periodically detecting the heartbeat signal of the baseboard management control device; if no valid heartbeat signal is received for N consecutive times, it is determined that the communication signal is lost, and the fault timestamp and the associated hardware status are recorded, wherein N is a positive integer greater than 2; In response to the emergency control program being started, the control strategy of the target device is reconstructed based on historical operating data, and emergency control parameters are output, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future needs based on historical operating data, outputting the emergency control parameters, and resetting the abnormal communication link associated with the baseboard management control device.
11. The server collaborative control method according to claim 1, characterized in that: The method further comprises: During the resetting of the baseboard management controller, at least one of the following operations is performed by the programmable logic device: Taking over the control authority of the target device, and adjusting the device parameters of the target device based on the real-time data of the sensor device; Report fault information through the alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller; After the baseboard management controller is restored, the encryption identification verification and the control parameter synchronization operation are re-executed to complete the authority transfer.
12. The server collaborative control method according to claim 1, characterized in that: The method further comprises: In response to the server starting up, outputting a modulation signal through the second processing unit in the programmable logic device to control the heat dissipation device to start up in a preset mode, wherein the target device includes the heat dissipation device; Loading an operating system through a first processing unit in the programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes the temperature monitoring device; Performing denoising and calibration processing on the operating parameters by the first processing unit to generate a valid temperature parameter set; Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register; activating the adaptive control mode according to the control register and taking over the control authority by the second processing unit; In response to the baseboard management control device having completed startup, receiving a handover request including an encryption identifier sent by the baseboard management control device through the preset communication link; Verifying the validity of the encryption identifier through the second processing unit, and feeding back the verification result to the first processing unit; If the verification is passed, synchronizing the historical control parameters, the temperature calibration data and the current speed regulation signal to the baseboard management controller through the first processing unit, wherein the baseboard management controller generates a speed regulation signal matching the current heat dissipation state of the target device based on the historical control parameters, the temperature calibration data and the current speed regulation signal; When the control authority is transferred, the programmable logic device is switched from the working mode to the auxiliary monitoring mode, and the operation status of the baseboard management controller is continuously checked; When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, the second processing unit takes over the control authority again, wherein N is a positive integer greater than 2; Reset the baseboard management controller by calling a preset fault recovery algorithm library through the first processing unit, and output a speed regulation signal to the heat dissipation device through the second processing unit; During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: reporting fault information through an alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller; In response to the baseboard management controller being reset, encryption identification verification and control parameter synchronization operations are performed to complete authority transfer.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the server collaborative control method as described in any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the server collaborative control method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server collaborative control method as described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Server management and control method, system and device and computer readable storage medium
CN118152161A
Server control method and device, equipment, medium and computer program product
CN118796010A
Server startup management system and method
WO2024149044A1
Cited By
Test method and device of server power supply board, storage medium and electronic equipment
CN120276925A
Macro server system, application method, electronic equipment and storage medium
CN120316046A
Multi-equipment cooperative control method, system and equipment for feeding of smelting furnace
CN120333148A
Method, system and equipment for collaborative control of multiple devices for furnace loading
CN120333148B
Middle backboard controller processing method and device, storage medium, program product, pin circuit and substrate management controller
CN120429203A