Server Cooperative Control Method, Storage Medium and Electronic Device

By actively taking over control permissions during the server startup stage, and interacting with the substrate management control device to verify and synchronize the control parameters, the problem of collaborative control between heterogeneous units is solved, and stable operation and efficient initialization of the server startup stage is achieved.

CN119917350BActive Publication Date: 2025-07-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510401844.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In heterogeneous computing architecture, when the substrate management controller loses control capabilities due to communication link failure or program abnormality, the system cannot control it, resulting in lag in the execution of the target device or interruption of function, and it is difficult to coordinate the control between heterogeneous units, making it difficult to ensure data synchronization, affecting the operation performance of high-density servers.

Method used

The programmable logic device is used to actively take over the control authority of the target device during the server startup stage, and the preset communication link is interactively verified with the substrate management control device, and the control authority is retaken in the event of an abnormality. The multi-communication link synchronous control parameters are used to achieve seamless connection and redundant fault tolerance in device management responsibilities.

Benefits of technology

It realizes seamless control rights during the server startup stage, ensures stable operation of the equipment, improves system initialization efficiency and data synchronization robustness, and enhances system failure self-healing ability and device control continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917350B_ABST
    Figure CN119917350B_ABST
Patent Text Reader

Abstract

The present application discloses a server collaborative control method, a storage medium, and an electronic device, relating to the technical field of firmware, including: during the server startup phase, a programmable logic device actively takes over the control authority of the target device based on the operating parameters of the sensor device to ensure the stable operation of the device before the baseboard management controller device completes initialization; a mechanism for interacting and verifying with the baseboard management controller device through a preset communication link is adopted, and the functional integrity and communication reliability of the baseboard management controller device during permission handover are ensured through two-way communication verification; a dynamic permission recovery strategy based on real-time anomaly detection is adopted, and the reverse switch of the control right is triggered through the dual monitoring mechanism of the sensor device and the communication link state, achieving the purpose of actively taking over the device management responsibility when the baseboard management controller device runs abnormally, thereby realizing the dual improvement of the system fault self-healing ability and the device control continuity, and solving the problem of difficult collaborative control between heterogeneous units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of firmware, and in particular, to a server cooperative control method, a storage medium, and an electronic device. Background Art

[0002] In the field of server regulation and management, traditional solutions usually use a single controller to implement the functions of running monitoring and parameter adjustment of target devices. With the popularization of heterogeneous computing architectures, the situation where programmable logic chips and baseboard management controllers are simultaneously deployed in a system is gradually increasing, but there are significant defects in the cooperative control process between the two.

[0003] In related technologies, when the baseboard management controller loses its regulation ability due to communication link failures or program anomalies, the system cannot be controlled, resulting in lag or even interruption of the functions of target devices, which may cause system-level operation risks. In addition, the control permissions between heterogeneous units are independently configured, making it difficult to achieve cooperative control between heterogeneous units. Therefore, data synchronization and execution coherence cannot be guaranteed, which has become a key technical obstacle restricting the improvement of the operation efficiency of high-density servers. Summary of the Invention

[0004] The present application provides a server cooperative control method, a storage medium, and an electronic device to at least solve the problems of relatively high system operation risks and difficulty in cooperative control between heterogeneous units in related technologies.

[0005] The present application provides a server cooperative control method, including: in response to the startup of the server, loading an operating system, and taking over the control permissions of the target devices associated with the server according to the operating parameters corresponding to the sensor devices, where the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and the baseboard management control device; in response to the completion of the startup of the baseboard management control device, performing interactive verification with the baseboard management control device based on a preset communication link, and transferring the control permissions to the baseboard management control device when the verification is passed; in response to detecting an anomaly in the baseboard management control device, taking over the control permissions again, where the control parameters of the target devices are synchronized between the programmable logic device and the baseboard management control device through at least one communication link, and the at least one communication link includes the preset communication link.

[0006] The present application also provides a server collaborative control device, including: an initialization module, configured to load an operating system in response to the startup of a server, and take over the control authority of a target device associated with the server according to the operating parameters corresponding to a sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and a baseboard management controller device; a verification module, configured to, in response to the completion of the startup of the baseboard management controller device, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device in the case of successful verification; a recovery module, configured to, in response to detecting an abnormality of the baseboard management controller device, take over the control authority again, wherein the control parameters of the target device are synchronized between the programmable logic device and the baseboard management controller device through at least one communication link, and the at least one communication link includes the preset communication link.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above server collaborative control methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any one of the above server collaborative control methods when executed by a processor.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the above server collaborative control methods when executed by a processor.

[0010] Through the embodiments of the present application, in the server startup stage, the programmable logic device actively takes over the control authority of the target device based on the operating parameters of the sensor device. By pre-positioning the device management responsibility to the programmable logic device with fast response ability, the stable operation of the device before the baseboard management controller device completes initialization is ensured, achieving the purpose of avoiding device management vacuum or parameter configuration conflicts caused by differences in the controller startup sequence, thereby realizing the technical effects of seamless connection of control rights and reliable management of device status in the server startup stage.

[0011] In addition, a mechanism for interacting and verifying with the baseboard management control device through a preset communication link is adopted. Through two-way communication verification, the functional integrity and communication reliability of the baseboard management control device during permission handover are ensured, achieving the purpose of eliminating potential logic errors or link abnormality risks before control handover, thereby realizing the technical effects of secure transition of control authority and improvement of system initialization efficiency. At the same time, through the parameter synchronization mechanism of multiple communication links in parallel between the programmable logic device and the baseboard management control device, the device control parameters are updated in real time through redundant links to ensure parameter consistency even when any link is abnormal, achieving the purpose of enhancing the robustness of data synchronization between controllers, thereby realizing the technical effects of redundant fault tolerance of control logic and rapid recovery in abnormal scenarios.

[0012] Furthermore, a dynamic permission recycling strategy based on real-time anomaly detection is adopted. Through the dual monitoring mechanism of the sensor device and the communication link status, the reverse switching of control authority is triggered, achieving the purpose of actively taking over the device management responsibility when the baseboard management control device runs abnormally, thereby realizing the dual improvement of the system's self-healing ability and device control continuity.

[0013] Finally, through the multi-link synchronization mechanism, a complete copy of the device control parameters is retained, and the parameters are automatically aligned before and after the control authority switch, achieving the purpose of avoiding device state jumps or configuration losses caused by permission changes, thereby realizing the long-term guarantee of device control stability and data consistency throughout the server's life cycle. Description of the Drawings

[0014] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 A schematic diagram of the application environment of an optional server collaborative control method provided by an embodiment of the present application;

[0016] Figure 2 A schematic flowchart of an optional server collaborative control method provided by an embodiment of the present application;

[0017] Figure 3 A partial schematic diagram of the hardware structure of an optional server collaborative control method provided by an embodiment of the present application;

[0018] Figure 4 A schematic diagram of the hardware structure design of an optional server collaborative control method provided by an embodiment of the present application;

[0019] Figure 5Schematic diagram of a server collaborative control device provided in an embodiment of the present application. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0021] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0023] Optionally, in combination with the specific application environment architecture or specific hardware architecture on which the execution of the server collaborative control method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0024] This embodiment may include, but is not limited to, deployment in a rack-mounted server, and the hardware architecture includes the following modules:

[0025] Programmable logic chip: The first processing unit runs a real-time operating system and is responsible for parsing temperature sensor data and generating a heat dissipation strategy; the second processing unit is implemented through hardware logic and directly controls the modulation signal of the target device (such as a cooling fan); the shared storage area is used to cache historical control parameters and fault logs, and synchronizes data with the baseboard management controller through a dual-buffer mechanism.

[0026] Baseboard management controller: The policy engine dynamically adjusts the long-term operation threshold of the heat dissipation device based on the server load prediction model; the status monitoring module verifies the communication status of the programmable logic chip through a heartbeat signal (interval 100 ms).

[0027] Communication link: The main channel uses the I²C protocol to transmit control instructions, and the backup channel uses the SPI protocol to achieve redundant communication; the independent alarm channel is directly connected to the system monitoring terminal through the GPIO (General-Purpose Input / Output) pin.

[0028] The application framework of this application uses programmable logic devices as the core execution body to build a multi-stage dynamic control authority switching system. The specific implementation process is as follows:

[0029] S1, server startup phase control authority initialization:

[0030] After the PLD is powered on on the server, it first completes the loading of its own firmware and the initialization of the operating system, and then collects the operating parameters of the target device in real time through the sensor device, including key indicators such as temperature, voltage, and fan speed. Based on the preset logic rules, the PLD performs threshold comparison and trend analysis on the sensor data. If any parameter is detected to be out of the safe range or there is a risk of mutation, the control authority takeover instruction is immediately triggered. During the takeover process, the PLD switches the driver configuration of the target device (such as the power module and the cooling unit) from the default state to the local management mode, and dynamically adjusts the device operating parameters based on the sensor data to ensure the stability of the hardware environment during the server startup phase. At the same time, the PLD sends an initialization status signal to the baseboard management control device through the preset communication link to start the latter's boot program.

[0031] S2, the authority transfer after the baseboard management control device is started:

[0032] When the baseboard management control device completes startup and enters the standby state, the programmable logic device initiates a two-way verification request through a preset communication link. The verification content includes the firmware version number, operation status code and encryption certificate validity of the baseboard management control device. If the verification is successful, the programmable logic device synchronizes the control parameters of the current target device to the baseboard management control device through at least one communication link, and starts the parameter consistency verification process. After the verification is successful, the programmable logic device sends an authority transfer instruction to the baseboard management control device, closes the local device driver interface, and switches the real-time data stream of the sensor device to the data buffer of the baseboard management control device. During the transfer process, the programmable logic device continuously monitors the response delay of the baseboard management control device and the correctness of the instruction execution to ensure that the authority transfer is uninterrupted.

[0033] S3, control permission recovery and restoration when operation is abnormal:

[0034] During the normal operation of the baseboard management controller device, the programmable logic device monitors the health status of the baseboard management controller device in real time through the two-way data packet verification of the periodic heartbeat signal with the preset communication link. If a heartbeat timeout, a verification error rate exceeding the threshold, or a critical device parameter deviating from the preset range is detected, the programmable logic device determines that an abnormality has occurred in the baseboard management controller device and immediately starts the privilege recovery process. First, the programmable logic device forcibly suspends the device access privilege of the baseboard management controller device through a backup communication link (such as a GPIO interrupt channel), and then obtains the latest device control parameters from the local cache or the shared memory of the baseboard management controller device and reloads the driver configuration. After the privilege recovery is completed, the programmable logic device takes over the real-time control of the target device and performs adaptive adjustment based on the sensor data until the baseboard management controller device fails to be repaired and passes the secondary verification.

[0035] S4, Multi-communication link redundancy and parameter synchronization guarantee:

[0036] At least two heterogeneous communication links are established between the programmable logic device and the baseboard management controller device. For example, a preset link based on a low-speed serial bus is used for regular data transmission, and another backup link based on a high-speed parallel interface is used for the transmission of emergency control instructions. During the privilege handover, parameter synchronization, and exception recovery processes, the programmable logic device transmits critical data through multi-link concurrent transmission and uses a differential verification algorithm to ensure data consistency. If a transmission delay or an abnormal error rate is detected in a certain link, the programmable logic device automatically switches to the backup link to maintain the real-time performance and reliability of the control parameter synchronization. In addition, the programmable logic device persistently stores multiple historical versions of the device control parameters in the local non-volatile memory to quickly roll back to a safe configuration in an extreme failure scenario.

[0037] Through the above framework, the programmable logic device realizes the flexible management of control privileges during the entire life cycle of the server, solves the system vulnerability problem caused by the dependence on a single controller, and at the same time significantly improves the continuity of device control and data integrity through the multi-link redundancy and dynamic synchronization mechanism.

[0038] The embodiments of the present application provide a server collaborative control method, and the method is described in detail in combination with the execution process of the server collaborative control method.

[0039] Optionally, as an alternative implementation, as Figure 2 shown, the above server collaborative control method includes:

[0040] S202. In response to the server startup, load the operating system, and take over the control authority of the target device associated with the server according to the operating parameters corresponding to the sensor device. The server is deployed with a server controller, and the server controller includes a programmable logic device, a sensor device, and a baseboard management controller device.

[0041] S204. In response to the completion of the startup of the baseboard management controller device, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device when the verification is passed.

[0042] S206. In response to detecting an abnormality in the baseboard management controller device, take over the control authority again. The programmable logic device and the baseboard management controller device synchronize the control parameters of the target device through at least one communication link, and the at least one communication link includes a preset communication link.

[0043] Optionally, in the embodiment of the present application, the above baseboard management controller device may include, but is not limited to, an embedded hardware management unit, which is usually integrated on the motherboard for monitoring and managing the hardware status of the server, including, but not limited to, functions such as temperature sensor data acquisition, fan speed regulation, power status monitoring, and fault warning. The core function of the baseboard management controller device is to realize the real-time monitoring of the server's underlying hardware through a dedicated chip or module independent of the main processor, and it can maintain the remote management ability even when the main processor is in the shutdown state. For example, in the server startup phase, the baseboard management controller device can independently complete the hardware self-check and the initialization of environmental parameters; in the running phase, continuously collect sensor data through a polling mechanism and dynamically adjust the heat dissipation strategy; in the fault scenario, trigger an alarm signal or start a redundant hardware switching process.

[0044] Optionally, in the embodiments of the present application, the above programmable logic device may include, but is not limited to, an integrated circuit chip with hardware reconfigurability, such as a field-programmable gate array (FPGA, Field-Programmable Gate Array) or a complex programmable logic device (CPLD, Complex Programmable Logic Device). The core feature of such devices is the ability to dynamically configure internal logic circuits through a hardware description language to achieve customized design of specific functional modules. For example, in a server management scenario, the programmable logic device can be divided into a processing system (PS, Processing System) and a programmable logic (PL, Programmable Logic) dual-processing domain: the PS part runs a real-time operating system to perform data processing and decision-making tasks, and the PL part implements low-latency operations such as hardware signal detection and PWM (Pulse Width Modulation) control signal generation. Its typical application scenarios include power supply timing management, hardware signal state acquisition, and fast response to fault signals, and it supports data interaction and collaborative control between the PS and the PL through the AXI (Advanced eXtensible Interface) bus.

[0045] Optionally, in the embodiments of the present application, the above sensor devices may include, but are not limited to, physical quantity detection devices for collecting server operating environment parameters, including but not limited to temperature sensors, voltage sensors, current sensors, vibration sensors, and humidity sensors. For example, a temperature sensor can be connected to a multiplexer through the I2C (Inter-Integrated Circuit) bus to achieve time-sharing access to data at multiple temperature measurement points; a voltage sensor can monitor voltage fluctuations in the power supply circuits of the CPU (Central Processing Unit), memory, and hard disk; a vibration sensor can detect abnormal vibrations of fan or hard disk mechanical components. The deployment locations of these sensor devices can cover key server modules, such as near the processor heat sink, the output end of the power module, and the hard disk backplane area, and transmit data in the form of digital signals or analog signals to the baseboard management controller or programmable logic device for joint analysis.

[0046] It should be noted that the design of the preset communication link can be extended according to actual requirements. For example, the I2C bus can be used to query the status of low-speed devices, the GPIO interface can be used to transmit hardware interrupt signals, or the PCIe high-speed channel can be used to synchronize large-scale control parameters. In scenarios where hardware resources are limited, a time-division multiplexing mechanism can be adopted to share the physical link. For example, the same set of I2C lines can be used to serve the baseboard management controller device and the programmable logic device through a multiplexer at different times. In addition, the hierarchical design of the communication protocol can include physical layer signal transmission, data link layer verification mechanism, and application layer instruction encapsulation. For example, a CRC (Cyclic Redundancy Check) verification module can be implemented in the PL logic to ensure signal integrity, or a JSON (JavaScript Object Notation) format instruction interaction specification can be defined in the PS system.

[0047] It should be noted that the synchronization mechanism of control parameters can include technical implementations in multiple dimensions. For example, a shared memory area can be used to quickly exchange sensor calibration data, the PWM duty cycle parameter can be synchronized through hardware register mapping, or a dual-port RAM (Random Access Memory) can be established to achieve secure data transfer in an asynchronous clock domain. In terms of parameter update strategies, an active push mode based on event triggering, a passive update mode of periodic polling, or an incremental synchronization mode of differential comparison can be designed. For example, when the baseboard management controller device detects that the environmental temperature change exceeds the threshold, it can actively send the updated heat dissipation strategy parameters to the programmable logic device; during the system startup phase, the programmable logic device can batch-read the initialization configuration parameters of the baseboard management controller device.

[0048] It should be noted that the construction of the fault recovery process can implement different strategies according to the fault level. For example, a retry mechanism can be adopted for momentary communication link interruptions, a spare part switch can be triggered for permanent damage to hardware modules, or an emergency frequency reduction protection can be started for out-of-control temperature. Specific recovery means can include hardware watchdog reset of PL logic, software hot restart process of the PS system, or power supply switching based on redundant power modules. For example, when an abnormality of the baseboard management controller device is detected, the programmable logic device can execute a three-level recovery strategy: first, it tries to wake up the baseboard management controller device through a hardware reset signal. If it fails three times in a row, it takes over the control right of key devices and reports the fault log to the remote operation and maintenance platform through the out-of-band management interface at the same time.

[0049] In an exemplary embodiment, during the construction of the server controller, by dividing the programmable logic device into a processing system (PS) and a programmable logic (PL) dual - processing domain, hardware - level fault isolation and function collaboration can be achieved. Specifically, after the PS unit loads the real - time operating system, it can execute complex decision - making tasks such as sensor data fusion analysis and heat dissipation algorithm calculation; while the PL unit, relying on its hardware parallel processing ability, can achieve millisecond - level response control of devices such as fan speed and power switch. The advantage of this heterogeneous architecture is that when the baseboard management controller device loses response due to software crash or hardware failure, the programmable logic device can seamlessly take over the device control right through a preset interrupt trigger mechanism, ensuring the continuous operation of the server's critical functions.

[0050] During the control right handover process, the design of the interactive verification mechanism directly affects system stability. After the baseboard management controller device completes startup, it will exchange status vectors and check codes with the programmable logic device through multiple types of communication links. For example, the heartbeat signal transmitted through the GPIO pin can be used to verify the smoothness of the physical link, the encrypted hash value stored in the shared memory area can ensure software state consistency, and the PWM phase synchronization operation ensures a smooth transition of the fan speed during the fan control right handover. This multi - level verification strategy effectively avoids the risk of mis - handover caused by the failure of single - point verification. At the same time, through the two - way synchronization mechanism of control parameters, the two management units always maintain policy consistency.

[0051] When the system detects an abnormal operation of the baseboard management controller device, the hardware interrupt response mechanism of the programmable logic device will be immediately activated. After the PL unit identifies the abnormal signal through the status monitoring circuit, it will send an interrupt request to the PS unit, triggering a preset fault recovery process. This process can include standardized operation modules such as sensor data takeover, control strategy switching, and fault log recording. For example, in the fan control scenario, the PS unit will load the backup heat dissipation strategy from persistent storage, re - configure the PWM controller parameters of the PL unit through the AXI bus, and at the same time start a periodic self - check thread to attempt to restore the communication connection of the baseboard management controller device.

[0052] Exemplarily, taking a server equipped with redundant power supplies and an intelligent heat dissipation system in a data center as an example, its specific implementation process includes but is not limited to the following steps:

[0053] S1, Control right takeover during the server startup phase:

[0054] When the server is powered on, the programmable logic device loads the operating system kernel, initializes the sensor interface, and collects the operating parameters of the target device in real time. For example, if the sensor detects that the temperature of a critical component exceeds the preset threshold, or the power supply voltage fluctuation exceeds the safe range, the programmable logic device triggers the control authority takeover logic to take over the control of the heat dissipation module and the power supply unit. At this time, the programmable logic device dynamically adjusts the device operating mode based on the real-time data. For example, it increases the operating intensity of the heat dissipation component to reduce the temperature, or limits the power output of the power supply to maintain voltage stability. At the same time, the programmable logic device sends an initialization instruction to the baseboard management controller through the preset communication link to start its boot process. If the sensor data is within the normal range, the default control strategy is maintained until the baseboard management controller is ready.

[0055] S2. The baseboard management controller starts verification and authority transfer:

[0056] After the baseboard management controller completes startup, the programmable logic device initiates a two-way verification process through the preset communication link, including identity verification and status check. For example, the programmable logic device sends an encrypted challenge code, and the baseboard management controller needs to generate a response code based on a specific algorithm and return it, while verifying whether its firmware version meets the requirements. After the verification passes, the programmable logic device synchronizes the current device control parameters to the baseboard management controller through multiple heterogeneous communication links and starts data consistency verification. If it is detected that the parameter deviation exceeds the preset range, synchronization retry is triggered until the data is consistent. After the verification is completed, the programmable logic device closes the local control interface, switches the sensor data stream to the specified storage area of the baseboard management controller, and completes the authority transfer.

[0057] S3. The baseboard management controller performs anomaly detection and authority recovery:

[0058] During the operation of the baseboard management controller, the programmable logic device monitors its health status through periodic heartbeat signals and communication link status. For example, if it is detected that the heartbeat response times out continuously for multiple times, or the critical device parameters continuously deviate from the set values, the programmable logic device determines that the baseboard management controller is abnormal. At this time, the programmable logic device forcibly suspends the device access authority of the baseboard management controller through the standby communication link, obtains the latest control parameters from the local cache or multi-link synchronized data, and reloads the device driver configuration. If the baseboard management controller cannot respond due to a serious fault, the programmable logic device directly reads the original sensor data through the standby communication link, takes over the device control, and performs adaptive adjustment.

[0059] S4. Multi-link parameter synchronization and redundancy recovery:

[0060] Multiple heterogeneous communication links are maintained between the programmable logic device and the baseboard management control device, including a low-speed basic link and a high-speed data link. During the privilege handover phase, the programmable logic device transmits core control parameters through the basic link and simultaneously bulk-transmits historical operation data through the high-speed link. When the baseboard management control device recovers from an anomaly, the programmable logic device writes back the updated parameters during the takeover period to the non-volatile storage area of the baseboard management control device through a standby link and uses a data verification algorithm to ensure transmission integrity. If the verification fails, the data is retransmitted through other available links until synchronization is successful.

[0061] Through the above embodiments, the programmable logic device dynamically takes over the device control privilege based on real-time sensor data during the server startup phase, solving the problem of device management blind spots caused by the startup delay of the main controller in the traditional architecture; through the multi-link redundant communication and encryption verification mechanism, it ensures the data security and control continuity during the privilege handover process and reduces the single-point failure risk; in abnormal scenarios, combined with the multi-modal detection strategy and the heterogeneous link redundant synchronization mechanism, it realizes the rapid recovery of device control rights and the maintenance of parameter consistency, thus significantly improving the stability and fault tolerance of the server system under complex working conditions.

[0062] Through the embodiments of the present application, the programmable logic device actively takes over the control privilege of the target device based on the operating parameters of the sensor device during the server startup phase. By preposing the device management responsibility to the programmable logic device with fast response ability, it ensures the stable operation of the device before the baseboard management control device completes initialization, achieving the purpose of avoiding device management vacuum or parameter configuration conflicts caused by differences in the controller startup sequence, and thus realizing the technical effect of seamless connection of control rights and reliable control of device status during the server startup phase.

[0063] In addition, by adopting the mechanism of interacting and verifying with the baseboard management control device through a preset communication link, through two-way communication verification, it ensures the functional integrity and communication reliability of the baseboard management control device during privilege handover, achieving the purpose of eliminating potential logic errors or link anomaly risks before the handover of control rights, and thus realizing the technical effects of safe transition of control rights and improvement of system initialization efficiency. At the same time, through the parameter synchronization mechanism with multiple communication links in parallel between the programmable logic device and the baseboard management control device, the device control parameters are updated in real time through redundant links to ensure parameter consistency even when any link is abnormal, achieving the purpose of enhancing the robustness of data synchronization between controllers, and thus realizing the technical effects of redundant fault tolerance of control logic and rapid recovery in abnormal scenarios.

[0064] Furthermore, a dynamic permission recovery strategy based on real-time anomaly detection is adopted. Through a dual monitoring mechanism of sensor devices and communication link status, the reverse switch of control right is triggered, achieving the purpose of actively taking over the device management responsibility when the baseboard management controller device runs abnormally, thereby realizing the dual improvement of the system fault self-healing ability and device control continuity.

[0065] Finally, through the multi-link synchronization mechanism, a complete copy of the device control parameters is retained. Through the automatic alignment of parameters before and after the control right switch, the purpose of avoiding device state jumps or configuration losses caused by permission changes is achieved, thereby realizing the long-term guarantee of device control stability and data consistency throughout the server's full life cycle.

[0066] As an optional solution, in response to the server startup, the operating system is loaded, and the control permission of the target device associated with the server is taken over according to the sensor data corresponding to the sensor device, including:

[0067] In response to the server startup, a modulation signal is output through the second processing unit in the programmable logic device to make the target device enter the initial operating state;

[0068] The operating parameters are periodically obtained through the first processing unit in the programmable logic device;

[0069] Based on the operating parameters, the first processing unit calculates the target parameters using a preset algorithm and writes the target parameters into the control register corresponding to the second processing unit;

[0070] The second processing unit activates the adaptive control mode and takes over the control permission according to the control register;

[0071] Before the baseboard management controller device completes startup, the target parameters are continuously updated to respond to changes in the operating parameters.

[0072] Optionally, in the embodiments of the present application, the above modulation signal may include, but is not limited to, a signal type that realizes energy control by adjusting the pulse width. Its core principle is to equivalently adjust the output power by changing the time ratio of the pulse high level to the low level. For example, in the server device control scenario, the modulation signal can be used to adjust the fan speed, control the voltage output of the power module, or adjust the refrigeration efficiency of the heat dissipation system. The specific implementation methods include generating a fundamental wave signal with a fixed frequency by a hardware timer, changing the effective output power through a duty cycle adjustment circuit, and dynamically adjusting the pulse parameters based on a feedback mechanism. In the server startup phase, the second processing unit can make the target device (such as a cooling fan) smoothly transition from a stationary state to a preset speed range by generating a modulation signal with a specific duty cycle, avoiding hardware damage caused by current impact.

[0073] Optionally, in the embodiments of the present application, the above initial operating state may include, but is not limited to, the preset working conditions required by the target device during the system startup phase, including but not limited to the voltage stable range, the motion reference parameters of mechanical components, and the communication link handshake status. The construction of this state usually depends on the initialization control sequence output by the programmable logic device, and through phased progressive parameter adjustment, the device is gradually guided into the controllable operating range.

[0074] Optionally, in the embodiments of the present application, the above operating parameters may include, but are not limited to, physical quantities or logical quantities reflecting the real-time working state of the target device, including but not limited to temperature, voltage, current, rotational speed, delay time, and the value of the error counter. For example, in the heat dissipation control scenario, the operating parameters may include the processor package temperature, the surface temperature difference of the heat sink, the current rotational speed of the fan, and the air pressure value in the heat dissipation duct; in the power management scenario, it involves the input voltage ripple, the effective value of the output current, and the power conversion efficiency. These parameters are collected by sensor devices or the built-in monitoring circuit of the device, and after analog-to-digital conversion or digital signal processing, they are periodically read and recorded by the operating system of the first processing unit.

[0075] Optionally, in the embodiments of the present application, the above preset algorithms may include, but are not limited to, mathematical models or rule sets for generating control strategies, including but not limited to proportional-integral-derivative control, fuzzy logic control, and model predictive control. For example, in the temperature regulation scenario, the preset algorithm can predict the heat dissipation demand according to the historical temperature change rate and dynamically adjust the fan speed curve; in the power load balancing scenario, the optimal power supply phase combination can be calculated based on the current distribution model. The implementation methods of these algorithms include mathematical library function calls, hardware acceleration module implementations, or neural network inference engine deployments, and their output results are used to update the control register parameters of the second processing unit.

[0076] Optionally, in the embodiments of the present application, the above control registers may include, but are not limited to, storage units for storing device control parameters, and their physical implementation forms include but are not limited to trigger arrays, static random access memories, or dedicated register files. For example, in the fan speed control scenario, the control register may contain the target speed setting value, the acceleration limit parameter, and the fault protection threshold; in the power module management, it stores the output voltage reference value, the overvoltage protection trigger level, and the soft start time configuration.

[0077] Optionally, in the embodiments of the present application, the above adaptive control mode may include, but is not limited to, a working mode that dynamically adjusts the control strategy according to environmental changes, including but not limited to parameter self-tuning, dynamic loading of the rule base, and control structure reconstruction. For example, in the scenario of abnormal heat dissipation system, the adaptive control mode can automatically switch to the backup heat dissipation strategy, increase the upper limit of the fan speed, and enable the auxiliary refrigeration unit; in the scenario of power supply fluctuation, the voltage compensation coefficient is dynamically adjusted, and the cooperative working logic of the multi-phase power supply module is reconstructed. The activation of this mode depends on the preset threshold comparison result in the control register or the mode switching instruction issued by the first processing unit.

[0078] It should be noted that the generation method of the modulation signal can be extended according to the hardware resources. For example, a dedicated hardware timer is used to achieve nanosecond-level precision signal output, a multi-channel synchronous signal generator is constructed through a programmable logic unit, or a pulse waveform is simulated using a software interrupt service routine. The selection range of the signal frequency can cover the interval from kilohertz to megahertz to adapt to the response characteristics of different devices. For example, low-speed mechanical components are suitable for low-frequency wide-range adjustment, and high-speed digital circuits are suitable for high-frequency fine control.

[0079] It should be noted that the construction method of the initial operating state can be flexibly designed according to the device type. For example, mechanical devices adopt a progressive parameter loading strategy, first releasing the brake and then gradually increasing the driving signal intensity; electronic devices implement a phased power supply strategy, successively completing core voltage establishment, clock synchronization, and interface initialization. For heterogeneous device groups, a parallel initialization process can be designed to synchronously control multiple devices through a multiplexing mechanism, or a priority queue can be used to start according to the criticality order. The exception handling mechanism can include timeout retry, backup parameter loading, and faulty device isolation. For example, when the hard disk homing calibration times out, it automatically switches to the backup head positioning algorithm or marks the device as offline.

[0080] It should be noted that the acquisition mechanism of the operating parameters can include various data fusion methods. For example, multiple sensors are deployed for the same physical quantity for redundant measurement, and the median filtering algorithm is used to eliminate outliers; joint analysis is performed on related parameters, such as combining the ambient temperature and the fan speed to calculate the heat dissipation efficiency. The data transmission path can be designed as a direct memory access channel, shared buffer polling, or interrupt-driven mode. For example, high-priority parameters (such as over-temperature alarm) are reported immediately using a hardware interrupt, and regular parameters are batch-transmitted through DMA. The parameter storage strategy can include circular buffer recording of historical data, snapshot saving of key states, or compression and archiving of long-term operation logs.

[0081] Through the embodiments of the present application, the intelligent transition of equipment management is realized through the hierarchical control architecture. The refined adjustment of the modulation signal reduces the impact current of the equipment startup process and prolongs the service life of mechanical components. The multi-dimensional collection and fusion analysis of operating parameters improves the dynamic adjustment accuracy of the control strategy. The coordinated design of the preset algorithm and the control register shortens the delay between strategy calculation and hardware response to the microsecond level. The introduction of the adaptive control mode improves the stability of the system when facing sudden load fluctuations or environmental changes.

[0082] It should be noted that in terms of device type expansion, it can support a smooth transition from traditional mechanical hard disks to all-flash arrays; in terms of algorithm evolution, it allows the deployment of new control models through online updates; in terms of communication interfaces, it is compatible with traditional I2C and GPIO protocols, and can also be expanded to support high-speed SerDes interfaces. For server systems of different sizes, the elastic scaling of control accuracy can be achieved by adjusting the resource configuration of programmable logic devices, and a seamless connection can be established between single-node control and cluster management.

[0083] As an optional solution, periodically obtaining the operating parameters through the first processing unit in the programmable logic device includes:

[0084] Using a multiplexed channel to collect data of a plurality of sensor devices as operating parameters in a time-sharing manner through a first processing unit;

[0085] De-noising the abnormal data in the operating parameters by the first processing unit, and marking invalid data segments;

[0086] The valid data is stored in the shared cache area through the first processing unit, so as to be called by the first processing unit during calculation.

[0087] Optionally, in an embodiment of the present application, the above-mentioned multiplexed channel may include but is not limited to a technical solution for realizing time-sharing transmission of multiple signals through a shared physical line, and its core principle is to realize a single transmission medium carrying multiple independent data streams through time slice rotation or address encoding. For example, in a sensor network scenario, the multiplexed channel can be designed as a cyclic acquisition architecture based on time division multiplexing, and a high-speed switching switch is used to sequentially connect the temperature sensor, voltage sensor and speed sensor to the same analog-to-digital converter. In specific implementation, the programmable logic device can generate a channel selection signal, switch different sensors to access the acquisition circuit at millisecond intervals, and cooperate with the sampling and holding circuit to maintain signal stability. This design can not only reduce the hardware resource usage, but also realize the synchronous timestamp marking of multiple types of sensor data. For example, in a server cabinet monitoring system, cyclic acquisition of 12 temperature monitoring points is realized through 8 multiplexed channels, and each channel is allocated a 2 millisecond exclusive sampling window.

[0088] Optionally, in the embodiments of the present application, the above abnormal data may include, but are not limited to, sensor measurement values exceeding a preset reasonable range or associated data combinations with logical contradictions, including but not limited to distorted signals caused by transient spike interference, sensor drift errors, and communication link bit errors. For example, in a temperature acquisition scenario, the abnormal data may be manifested as a mutation of a single sampling value exceeding the set Celsius threshold compared to the previous data, or a linear increase in three consecutive sampling periods that contradicts the environmental heat dissipation strategy. In a current monitoring scenario, if the current value of a certain phase power supply is continuously zero while the loads of other phases are normal, it may be an abnormality caused by sensor failure. The identification of abnormal data can rely on threshold comparison, sliding window variance analysis, or machine learning model inference.

[0089] Optionally, in the embodiments of the present application, the above shared buffer may include, but are not limited to, a common storage area for data exchange between multiple processing units or processes, and its physical implementation may include a dual-port memory, a circular buffer, or a dynamic memory pool with a mutex. For example, in a heterogeneous processing architecture, the shared buffer can be designed as an on-chip memory block accessible to both the first processing unit (PS) and the second processing unit (PL), and a ping-pong buffer structure is used to implement the pipelining operation of acquisition and processing. In a specific application, when the PS unit performs data analysis, the PL unit continuously writes new acquired data to the standby buffer, and immediately switches the read and write pointers after the PS processing is completed. The buffer management strategy may include data block verification, dynamic allocation of read and write permissions, and overflow protection mechanisms. For example, an independent storage page is allocated for temperature data and a cyclic redundancy check code is added. When the PL unit detects that the PS has not read the data in time, data compression storage is automatically enabled.

[0090] It should be noted that the design of the multiplexing channel can be flexibly adjusted according to system requirements. For example, a priority channel allocation strategy is adopted to reserve a fixed time slice for key sensors, and the remaining resources are dynamically allocated to secondary sensors. The abnormal data processing method can be extended according to the data type and system fault tolerance requirements. For example, sliding window median filtering is used for transient interference, online calibration compensation is implemented for sensor drift, and a forward error correction mechanism is enabled for communication bit errors; in the dimension of correlation analysis, a device state model can be established to verify the data rationality. For example, when the fan speed increases, the temperature in the corresponding heat dissipation area should show a downward trend; in terms of the disposal strategy, data interpolation reconstruction, abnormal marking and ignoring, or triggering the device re-inspection process can be designed. For example, for occasional current spike data, linear interpolation of adjacent sampling points is used for replacement; for continuously abnormal temperature readings, redundant sensor cross-verification is started and a device diagnostic report is generated.

[0091] Through the embodiments of the present application, the intelligent scheduling of the multiplexing channel improves the efficiency of sensor data acquisition, reduces the hardware cost at the same time, and the dynamic management strategy of the shared buffer reduces the risk of data loss, improving the stability of the system under sudden data peaks.

[0092] It should be noted that the data acquisition architecture has good scalability. In terms of expanding the number of sensors, multi-level node access can be achieved by increasing the multiplexing channel level or adopting a tree topology structure; in the dimension of data processing, it supports smooth upgrades from simple threshold alarms to complex machine learning model analysis; in terms of real-time requirements, it can not only meet the second-level monitoring requirements, but also adapt to the millisecond-level real-time control scenario by optimizing the buffer strategy. For heterogeneous computing platforms, the plug-and-play of the acquisition module and different processing units can be realized through standardized interfaces.

[0093] As an alternative solution, the first processing unit calculates the target parameter based on the operating parameter using a preset algorithm and writes the target parameter into the control register corresponding to the second processing unit, including: generating an adjustment instruction by the first processing unit based on the change trend of the operating parameter; converting the adjustment instruction into a control instruction recognizable by the second processing unit by the first processing unit, where the target parameter includes the control instruction; writing the control instruction into the control register by the first processing unit.

[0094] Optionally, in the embodiments of the present application, the above change trend may include, but is not limited to, characteristic indicators for analyzing the change rate and direction of the operating parameter through a historical data sequence, which are usually used to predict the evolution trend of the device state. For example, in the heat dissipation control scenario, the temperature change trend may be a linear increase of 0.5 degrees Celsius per minute or an exponential accelerating temperature rise characteristic; in the power management scenario, the current fluctuation gradient may include a periodic oscillation amplitude decay or a sudden spike increase mode. The methods for calculating the change trend include the difference method, sliding window linear regression, and wavelet transform analysis, and the output results are used to determine whether the device is in a stable state, a transition state, or an abnormal state. In a specific application, when it is detected that the surface temperature gradient of the radiator exceeds 0.2 degrees Celsius per second, the system can trigger the fan speed increase strategy in advance to avoid thermal runaway.

[0095] It should be noted that the conversion process of the control instruction can adapt to various optimization strategies, such as using a look-up table method to implement non-linear instruction mapping, improving the instruction resolution through an interpolation algorithm, or introducing a dead zone compensation mechanism to eliminate the influence of the mechanical clearance of the actuator. The instruction encoding method can include direct numerical writing, relative value incremental update, or conditional trigger instruction packets. For example, the target rotation speed value is encoded as a 16-bit binary number, accompanied by 2-bit check codes and 1-bit emergency braking flag. The design of the transmission protocol can consider hierarchical processing of real-time requirements. Key instructions are directly transmitted through dedicated hardware channels, and non-key instructions are batch-transmitted through a shared bus. For example, the over-temperature protection instruction is immediately transmitted through the GPIO pin, while the fan speed fine-tuning instruction is periodically batch-updated through the I2C bus.

[0096] In addition, the implementation form of the hardware acceleration module can be flexibly configured according to system resources. For example, a fully parallel computing array can be deployed when there are sufficient logic resources, and a time-division multiplexing architecture can be adopted in resource-constrained scenarios. The acceleration task allocation strategy can include static task binding, dynamic load balancing, or a hybrid scheduling mode. For example, the calculation of control laws with a fixed period can be bound to a dedicated hardware unit, while bursty data analysis tasks are allocated to reconfigurable logic blocks. In terms of energy efficiency optimization, dynamic voltage and frequency regulation, idle module clock gating, or adaptive adjustment of computing precision can be designed. For example, when the system load is below 30%, the operating voltage of the acceleration module is automatically reduced to save 15% power consumption. The standardized interface design allows the module to be migrated and reused between different processing units. For example, a unified AXI stream interface specification is defined, enabling the same convolution acceleration core to serve both the PS unit and the PL unit.

[0097] Through the embodiments of the present application, the execution efficiency of complex control algorithms is improved, while the computing load of the main processor is reduced. The standardized conversion mechanism of control instructions enhances the maintainability of the system and shortens the access adaptation time of devices from different manufacturers.

[0098] As an alternative solution, the above method further includes at least one of the following:

[0099] Accelerating the operation of target parameters through the parallel computing unit deployed in the first processing unit;

[0100] Receiving the data stream corresponding to the operation parameters by the first processing unit through the high-speed data interface.

[0101] Optionally, in the embodiments of the present application, the above parallel computing unit may include, but is not limited to, a computing module that realizes multi-task synchronization processing through hardware architecture design. Its core principle is to improve the operation throughput by increasing the number of physical computing cores or optimizing the data path structure. For example, in a programmable logic device, the parallel computing unit can be designed as multiple groups of multiply-accumulate arrays to simultaneously perform matrix multiplication operations in a closed-loop model; or configured as a pipeline structure to split the iterative calculation of the control algorithm into multiple stages and advance in parallel. Specifically, when implementing, the computing resources can be dynamically configured according to different control scenarios: in the temperature regulation scenario, 8 groups of parallel integrators are deployed to simultaneously process the thermodynamic equations of multiple heat dissipation areas; in the power management scenario, a 16-way parallel voltage fluctuation prediction unit is used to analyze the status of the multi-phase power supply module in real time. This design enables the operation of the control model that originally took milliseconds to be shortened to microseconds, significantly improving the generation efficiency of control instructions.

[0102] Optionally, in the embodiments of the present application, the above high-speed data interface may include, but is not limited to, physical channels or protocol standards that support high-bandwidth and low-latency data transmission, and its design goal is to minimize the transmission delay of data from the acquisition end to the processing end. For example, it is designed as a wide-bit parallel bus to synchronously transmit multi-sensor fusion data through a 128-bit data path. In specific applications, the interface can integrate a hardware flow control mechanism. For example, a ping-pong buffer is set at the data receiving end to ensure continuous data stream processing without interruption; a priority scheduling strategy is implemented at the transport layer to give priority to transmitting critical parameters (such as over-temperature alarm signals). In the server power management scenario, such an interface can support the real-time transmission of tens of thousands of current sampling values per second, providing high-refresh-rate input data for closed-loop control.

[0103] It should be noted that the implementation form of the high-speed data interface can be adjusted according to the system topology. For example, in a centralized architecture, a star topology is used to connect multiple sensor nodes, and in a distributed system, a ring bus is deployed to achieve data relay transmission; at the protocol level, it can support flexible definition from the underlying electrical signals to the upper-layer data encapsulation. For example, a custom frame structure includes a timestamp, a check code, and a data payload. The error handling mechanism can include forward error correction coding, retransmission request strategy, and redundant path switching. For example, when it is detected that the checksum of three consecutive data packets fails, the backup transmission channel is automatically switched and an error log is triggered. The bandwidth allocation strategy can be designed as static reservation (fixing 20% of the bandwidth for critical sensors) or dynamic adjustment (allocating the remaining bandwidth in real time according to the burstiness of the data stream).

[0104] In addition, the write mechanism of the control register can adapt to various security protection requirements. For example, double-signature verification is implemented to ensure the legality of the write instruction, redundant registers are used to implement atomic operations, or shadow registers are designed to support parameter preloading and seamless switching. The register access permission management can include a hierarchical control strategy. For example, only the hardware acceleration module is allowed to write to the critical control fields, while the debug interface only opens the status reading permission; in terms of timing control, the write enable signal can be designed to be strictly aligned with the clock edge to avoid metastability problems. The function expansion of the register bank can include state machine control bits, mode switching flags, and exception capture registers. For example, a 4-bit status code is set to identify the current control mode, and a 32-bit exception vector register records the types and occurrence times of the last 10 errors.

[0105] Through the embodiments of the present application, the deployment of the parallel computing unit improves the operation speed and shortens the generation cycle of critical control instructions. The optimized design of the high-speed data interface reduces the transmission delay and supports the real-time processing of sensor signals. The secure write mechanism of the control register reduces the parameter update error rate and improves the overall control response consistency of the system. The synergistic effect of these three enables the server to enhance the power supply fluctuation suppression ability and improve the efficiency of the heat dissipation system when running at full load, reaching the leading stability indicators in the industry.

[0106] It should be noted that this hardware acceleration architecture has high scalability. In terms of the expansion dimension of computing power, the processing capacity can be improved by increasing the number of parallel units or upgrading the manufacturing process. For emerging technology requirements, such as quantum computing control or photonic communication management, adaptation can be achieved by replacing the core algorithms of the computing units and upgrading the optoelectronic conversion module of the interface. This modular design concept enables the solution to meet the existing server control requirements and also reserves room for technology upgrades for the evolution of future intelligent computing power infrastructure.

[0107] As an alternative solution, in response to the fact that the baseboard management controller device has completed startup, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device in the case of successful verification, including:

[0108] In response to the fact that the baseboard management controller device has completed startup, receive the authority transfer request sent by the baseboard management controller device through the preset communication link;

[0109] Verify the validity of the authority transfer request, and in the case where the verification result is valid, synchronize the historical control parameters and operating parameters to the baseboard management controller device;

[0110] Execute the phase matching operation corresponding to the control parameters, so that the control parameters generated by the baseboard management controller device subsequently match the control parameters generated by the programmable logic device;

[0111] In the case where the control authority is successfully transferred, switch the working mode of the programmable logic device to the auxiliary monitoring mode.

[0112] Optionally, in the embodiments of the present application, the above-mentioned authority transfer request may include, but is not limited to, a standardized signal or data packet for triggering the transfer of control rights, which contains key information such as a request identifier, a timestamp, a summary of the current device status, and an encryption verification certificate. For example, in the scenario of server control authority transfer, this request can be actively initiated by the baseboard management controller device after completing self-check, and transmitted to the programmable logic device through the preset communication link to trigger the control right verification process. Specifically, when implemented, the request content may include the hardware serial number, firmware version number of the baseboard management controller device, and a temporary token generated during the startup phase, ensuring the legality and timeliness of the request source. At the data transmission level, a framed encryption strategy can be adopted, splitting the request into multiple data segments and attaching a cyclic redundancy check code to avoid information tampering caused by communication link interference.

[0113] Optionally, in the embodiments of the present application, the above historical control parameters may include, but are not limited to, control strategy records generated by the programmable logic device during the control authority holding period, including but not limited to device operation mode configuration, dynamic adjustment parameters, and exception handling strategies. For example, in the scenario of heat dissipation system control, the historical control parameters may include the fan speed curves corresponding to different temperature ranges, the load balancing strategy of the power supply module, and the degradation operation plan of the hard disk array. These parameters are usually stored in the non-volatile memory in the form of a time series, including timestamps, parameter version numbers, and associated sensor data snapshots, ensuring that the complete evolution process of the control logic can be traced after the baseboard management controller takes over.

[0114] Optionally, in the embodiments of the present application, the above phase matching operation may include, but is not limited to, a timing alignment mechanism to ensure a smooth transition between the old and new control strategies, and its core objective is to eliminate parameter jumps or logical conflicts during the handover of control rights. For example, in the scenario of power supply module control, phase matching is required to ensure that the voltage regulation instructions generated by the baseboard management controller are continuous in amplitude, frequency, and phase with the historical output values of the programmable logic device. The specific operations may include parameter interpolation transition (such as gradually switching to the new parameters within the next 3 control cycles), timing synchronization (aligning the execution beats of the two control systems through the hardware clock), and safety threshold verification (ensuring that the new parameters do not exceed the device tolerance range).

[0115] Optionally, in the embodiments of the present application, the above auxiliary monitoring mode may include, but is not limited to, a lightweight working mode in which the programmable logic device continues to operate after the handover of control rights. Its core functions include exception detection, data backup, and redundancy control preparation. For example, in this mode, the programmable logic device can continue to collect sensor data and compare and analyze it with the control effects of the baseboard management controller. When it detects that the temperature regulation response delay exceeds the threshold, it automatically triggers a warning signal; at the same time, it maintains a shadow copy of the key control parameters to support a quick control right handback in case of an exception of the baseboard management controller. The resource occupancy rate of this mode is usually controlled below 10%, ensuring that the main hardware resources can be released for other tasks.

[0116] It should be noted that the triggering conditions for the permission handover request can be extended and designed according to system requirements. For example, time-driven triggering (initiated after a fixed delay of 5 seconds after the baseboard management controller device starts), event-driven triggering (initiated when the load rate of the programmable logic device is detected to be lower than 20%), or a hybrid triggering strategy (meeting the conditions of firmware version matching and network topology consistency) can be adopted. The encapsulation format of the request content can be adapted to different security level requirements. For example, a multi-signature mechanism is used in the trusted execution environment, and hash verification is used in the ordinary environment; the transmission protocol can support unicast, multicast, or broadcast modes to adapt to centralized or distributed control architectures. The exception handling mechanism can include timeout retry (abort the handover if there is no response to three consecutive requests), conflict resolution (start priority arbitration when multi-node concurrent requests are detected), and secure fallback (restore the original control state after the handover fails).

[0117] Among them, various optimization strategies can be designed for the synchronization mechanism of historical control parameters. For example, incremental synchronization (only transmit the recently changed parameters), compressed transmission (perform run-length encoding on repeated parameters), or differential merging (generate an incremental package by comparing version numbers). End-to-end encryption, shard verification, and resume from breakpoint (record the successfully transmitted address range) can be implemented during the data transmission process. The parameter storage structure can include a time series database, key-value pair storage, or relational table structure. For example, a parameter partition table is established according to device types, and each partition contains a timestamp index, parameter values, and associated metadata (such as generation algorithm identification).

[0118] Through the embodiments of this application, by using a standardized permission handover process, the security and reliability of the control right switching process are improved, the historical parameters and calibration data are completely synchronized, while ensuring the continuity of the control strategy, the parameter fluctuation range during device state switching is controlled within an acceptable range. The introduction of the phase matching mechanism reduces the coordination error between the new and old control systems, and significantly enhances the smooth operation of the device. The auxiliary monitoring mode constructs a double guarantee system, extends the mean time between failures of the system, and shortens the recovery time of major failures.

[0119] It should be noted that this control right handover architecture has good expansion and adaptation capabilities. At the device compatibility level, it supports the smooth access of baseboard management controller devices from traditional x86 architectures to ARM architectures; in terms of security protocols, it can be extended to support new security mechanisms such as national cryptographic algorithms and quantum encryption; in the dimension of control scale, it can meet the needs of single-node servers and also support cluster-level control right handover through protocol extension. For emerging intelligent operation and maintenance scenarios, machine learning models can be integrated to achieve intelligent decision-making for handover timing, or the risks of the handover process can be pre-empted through digital twin technology. This flexible design concept lays a technical foundation for the full automation operation and maintenance of future intelligent data centers.

[0120] As an alternative solution, perform a phase matching operation corresponding to the control parameter to make the control parameter subsequently generated by the baseboard management controller match the control parameter generated by the programmable logic device, including:

[0121] Obtain the signal period and duty cycle of the control parameter output to the target device;

[0122] In response to receiving the status confirmation signal, send the signal period and duty cycle to the baseboard management controller to enable the baseboard management controller to generate alternative control parameters and complete the permission switch, where the status confirmation signal is used to indicate that the baseboard management controller has completed startup, and the alternative control parameter represents the parameter for the baseboard management controller to control the target device.

[0123] Optionally, in the embodiments of the present application, the above signal period may include, but is not limited to, the time interval between two adjacent same-phase points in the control parameter waveform, which is usually used to characterize the frequency characteristics of the control signal. For example, in the PWM fan control scenario, the signal period may be set to the 40 microsecond period value corresponding to 25 kHz, and its high-level duration (duty cycle) determines the fan speed; in the voltage regulation scenario, the signal period may correspond to the chopping frequency of the switching power supply, such as the 5 microsecond period corresponding to 200 kHz. The capture of the signal period needs to be implemented through a hardware timer or a digital phase-locked loop, and the specific methods include edge-triggered capture, zero-crossing detection, or correlation function calculation. In the phase matching operation, accurately measuring the current signal period is a prerequisite for ensuring the frequency consistency of the new and old control parameters.

[0124] Optionally, in the embodiments of the present application, the above duty cycle may include, but is not limited to, the proportional parameter of the effective level time of the control signal in the total period, which is the core control variable for adjusting the output power of the device. The capture of the duty cycle usually depends on a high-precision timer, and the effective duty ratio is calculated by measuring the time difference between the rising edge and the falling edge. During the phase matching process, the consistency of the duty cycle needs to be controlled within the error range to avoid jumps in the output state of the device. For the multi-channel control scenario, the duty cycle parameter may include the phase offset between the master and slave channels, and cross-comparison is required to ensure the overall waveform coordination.

[0125] Optionally, in the embodiments of the present application, the above alternative control parameter may include, but is not limited to, a new parameter set generated by the baseboard management controller and having an equivalent control effect to the current operating parameters. Its generation process needs to meet the requirements of time-domain continuity, frequency-domain consistency, and logical compatibility. For example, in the power supply control scenario, the alternative parameter may include equivalent voltage setting values, soft start time constants, and overvoltage protection thresholds, but is implemented through different adjustment algorithms; in the mechanical control scenario, it is necessary to ensure that the acceleration curve corresponding to the alternative parameter is smoothly connected to the original parameter. The parameter generation method may include look-up table mapping, model equivalent conversion, or dynamic parameter compensation.

[0126] Optionally, in the embodiments of the present application, the above state confirmation signal may include, but is not limited to, a two-way handshake protocol data packet for verifying the security of control right switching, including key information such as a summary of the current device state, a parameter check code, and a switching ready flag. For example, in a permission switching scenario, the state confirmation signal may be an acknowledgment frame sent by the baseboard management controller device containing the current power supply phase, the hash value of the thermal management strategy, and the device topology information. The programmable logic device verifies the state consistency by comparing the shadow parameters in the memory. The signal transmission needs to meet the requirements of real-time and reliability, and usually adopts a combination of a hardware-level confirmation mechanism (such as the level jump of a dedicated GPIO pin) and a software-level confirmation protocol (such as an encrypted JSON data packet). The exception handling mechanism may include three-way handshake retries, automatic calibration of differential parameters, or a switching abort rollback process.

[0127] It should be noted that the generation strategy of alternative control parameters can be adapted to multiple algorithm frameworks, such as parameter mapping based on equivalent transformation of transfer functions, generating optimized parameters through a reinforcement learning model, or using digital twin technology to simulate and verify the parameter effects. The verification of parameter equivalence may include time-domain response comparison (such as the overshoot difference of the step response is less than 2%), frequency-domain characteristic analysis (such as the cut-off frequency offset is less than 5%), or logical compatibility check (such as the protection threshold completely covers the original parameter range). In a dynamic environment scenario, a parameter compensation mechanism can be designed, such as automatically adjusting the gain coefficient of the alternative parameter according to the temperature change, or loading degradation compensation parameters according to the device aging degree. The interaction protocol of the state confirmation signal can be extended to support multiple security mechanisms, such as using dynamic token authentication (generating a unique verification code for each session), two-way digital certificate verification, or physically unclonable function signature. The transport layer design may include a redundant confirmation mechanism (such as three-way handshake + heartbeat detection), differential data synchronization (only transmitting the changed parameters), or a hierarchical confirmation strategy (immediate confirmation of critical parameters, batch confirmation of non-critical parameters). The exception handling process can be designed as a progressive response, such as triggering parameter fine-tuning for the first difference, starting log analysis for the second difference, and performing control right rollback for the third difference. For high-reliability scenarios, a multi-node cross-verification mechanism can be deployed, requiring more than half of the verification nodes to return a successful signal before the switching can be executed.

[0128] Through the embodiments of the present application, by using high-precision signal capture and intelligent parameter generation technologies, the fluctuation amplitude of the device state during control right switching is reduced, and the average completion time of the phase matching operation is shortened. The dual security guarantee of the state confirmation mechanism reduces the switching failure rate to less than one in a hundred thousand, and significantly enhances the system reliability. These improvements enable "seamless switching" of critical devices during the control right transfer process, providing a solid technical guarantee for high-availability server systems.

[0129] It should be noted that this phase - matching architecture has strong expansion potential. In terms of signal type support, it can be extended and adapted to various control parameters from analog voltage signals to digital pulse sequences; at the algorithm level, it supports the hybrid deployment of traditional control theory and AI models; in the dimension of security authentication, it can seamlessly integrate cutting - edge technologies such as quantum encryption and blockchain verification. For distributed control systems, cross - node parameter synchronization can be achieved through protocol extension, supporting the transfer of cluster - level control rights for tens of thousands of devices. This modular design concept enables the solution to meet the current server control requirements and also provides a reusable technical framework for precision control in future fields such as intelligent factories and autonomous driving.

[0130] As an optional solution, in the case of successful verification, synchronize the historical control parameters and operating parameters to the baseboard management controller device, including:

[0131] In the case of successful verification, write the operating parameters to the shared storage area after calibration and encryption. Among them, the baseboard management controller device is set to read and decrypt the operating parameters from the shared storage area in the way of memory access, and after verifying the data integrity, update the control parameters of the target device.

[0132] Optionally, in the embodiments of this application, the above - mentioned calibration and encryption may include, but are not limited to, a processing flow of performing error correction on the original sensor data and then implementing encryption protection. Its core lies in ensuring both data accuracy and transmission security. For example, in the scenario of calibrating and encrypting temperature parameters, first correct the original sampling value according to the non - linear error table of the sensor, and then use a symmetric encryption algorithm to encrypt the calibrated data with an attached timestamp and sensor identifier. When specifically implemented, a sharding encryption mechanism can be designed to split a single piece of data into multiple ciphertext segments and store them dispersedly to prevent the data from being intercepted completely. In the voltage monitoring scenario, calibration and encryption may involve zero - point drift compensation and gain adjustment, and then generate a digital signature for the data through an elliptic curve encryption algorithm to ensure the credibility and non - tampering of the data source.

[0133] Optionally, in the embodiments of this application, the above - mentioned memory access may include, but are not limited to, a data transfer technology between peripherals and memories without the intervention of a central processing unit, which realizes high - throughput data transfer through a dedicated channel. For example, in the parameter synchronization scenario, the baseboard management controller device can complete data migration at a rate of gigabytes per second by configuring the source address (shared storage area) and target address (local cache) of the controller.

[0134] Optionally, in the embodiments of the present application, the above verification of data integrity may include, but is not limited to, technical means for verifying that the data has not been accidentally modified or maliciously tampered with during transmission and storage. For example, in the scenario of verifying parameter synchronization, cyclic redundancy check may be used to calculate the checksum of a data block and compare it with the check value generated before transmission; or a hash algorithm may be used to generate a data fingerprint, and the digital signature is used to verify the legality of the data source. In scenarios with a higher security level, a multi-layer verification mechanism may be deployed: at the physical layer, Hamming codes are used to correct single-bit errors, at the transport layer, verification detection is performed for multi-bit errors, and at the application layer, digital signatures are implemented to verify the authenticity of the data. When a verification failure is detected, the system may automatically trigger disposal strategies such as data retransmission, alarm, or switching to an alternative data source.

[0135] Optionally, in the embodiments of the present application, the above update of control parameters may include, but is not limited to, the process of deploying the verified new parameters to the target device control system, and it is necessary to ensure a smooth and secure update process. For example, in the scenario of fan control, the parameter update may include a progressive transition strategy: within the control cycle, linearly transition the target speed from the current value to the new set value to avoid mechanical stress caused by sudden speed changes. The update mechanism can be designed as a double-buffer structure, first writing the new parameters into the shadow register, and then switching and taking effect through an atomic operation after a complete verification. For critical devices, a rollback mechanism may also be deployed. When an abnormal device state is detected after the update, it will automatically revert to the previous stable parameter version.

[0136] It should be noted that in terms of algorithm selection, symmetric encryption, asymmetric encryption, or lightweight encryption can be adapted; in terms of accuracy guarantee, sensor characteristic curve fitting calibration, multi-sensor data fusion calibration, or online self-learning calibration can be deployed; in terms of data encapsulation format, fixed-length data packets, structures, or custom binary protocols can be designed. For example, a segmented encryption strategy is adopted for temperature data, where the integer part and the decimal part are encrypted separately and then combined for transmission; after the frequency-domain feature extraction of vibration signals, selective encryption is performed on the energy values of key frequency bands.

[0137] In addition, in terms of channel allocation, priority pre-emptive channels (critical data is transmitted first), polling channels (equal bandwidth resources are divided), or adaptive channels (dynamically adjusted according to data traffic) can be set; in terms of transmission mode, single transmission, block transmission, or hash-aggregate transmission are supported; in terms of the exception handling mechanism, transmission timeout interruption, data alignment error detection, or buffer overflow protection can be designed. For example, in the scenario of high-speed data synchronization, the channel is configured as a circular buffer mode. When the read speed of the baseboard management controller lags, the earliest historical data is automatically overwritten and the overflow event is recorded.

[0138] Moreover, Hamming code check is implemented at the physical layer to correct transmission error codes, integrity check is deployed at the link layer to detect the integrity of data blocks, and digital signature is used at the application layer to verify the data source. The check failure handling strategy can include hierarchical response: single-bit error triggers automatic error correction, multi-bit error starts the retransmission mechanism, and signature verification failure isolates the data and generates a security alert. For example, a triple-check mechanism is implemented for critical control parameters. When the check fails twice in a row, it automatically switches to the redundant control module and triggers an operation and maintenance alert.

[0139] Through the embodiments of the present application, by using the collaborative design of calibration encryption and integrity check, the data error in the parameter synchronization process is reduced, and at the same time, the anti-tampering ability of data transmission meets the security standard. The optimization of the direct memory access technology improves the data transfer efficiency and shortens the update delay of key parameters. The combination of the double-buffer update mechanism and the progressive transition strategy controls the fluctuation range within a reasonable range when the device state switches, and significantly enhances the system stability.

[0140] It should be noted that the parameter synchronization architecture has excellent expansion and adaptation capabilities. At the encryption algorithm level, it can be seamlessly upgraded to a lattice-based encryption scheme resistant to quantum computing; in terms of the storage architecture dimension, it supports the migration from traditional memory to new non-volatile storage media; in terms of integrity check, an AI-based data anomaly detection model can be extended and deployed. For heterogeneous computing environments, the solution can adapt to processors with different instruction set architectures and achieve cross-platform data synchronization through standardized interfaces. This flexibility enables it to meet the management needs of traditional servers and can also be extended to the precise parameter management scenarios of edge computing nodes and industrial control devices.

[0141] As an optional solution, in response to detecting an abnormality in the baseboard management controller device, take over the control authority again, including:

[0142] Detect whether the communication signal associated with the baseboard management controller device is lost;

[0143] When it is continuously detected that the communication signal is lost and the cumulative number of signal losses reaches a preset threshold, start the emergency control program;

[0144] In response to the emergency control program being started, reconstruct the control strategy of the target device based on historical operation data and output emergency control parameters. Among them, during the reset of the baseboard management controller device, the programmable logic device performs the function control and exception alarm of the target device;

[0145] After the baseboard management controller device is restored, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device when the verification is passed.

[0146] Optionally, in the embodiments of the present application, the above communication signal loss may include, but is not limited to, the data transmission interruption state of the preset communication link between the baseboard management controller and the programmable logic device, and the detection means includes heartbeat signal timeout, continuous error of check codes, abnormal physical layer level, etc. For example, in the communication scenario based on the I2C protocol, the signal loss may be manifested as the bus continuously being in the low level state for more than 500 milliseconds, or the master device not receiving the response signal from the slave device for 3 consecutive times; in the Ethernet communication scenario, the signal loss can be judged by request timeout or connection reset. The detection module is usually deployed with a hardware watchdog timer and a software state machine for joint monitoring. For example, the arrival situation of heartbeat packets is detected every 50 milliseconds, and at the same time, the physical layer signal quality indicators (such as eye opening degree, bit error rate) are monitored.

[0147] Optionally, in the embodiments of the present application, the above preset threshold may include, but is not limited to, the cumulative number of abnormal events or the duration parameter that triggers the fault determination, and its setting needs to balance the system sensitivity and anti-interference ability. For example, in the communication signal detection scenario, the preset threshold may be set to 5 consecutive heartbeat signal losses or a cumulative communication interruption of 3 seconds; in the temperature anomaly scenario, it may be defined as a certain sensor overheating for 10 consecutive sampling cycles. The threshold setting method may include static configuration (such as setting according to the device manual), dynamic learning (automatically adjusting according to historical operation data), or a hybrid mode (basic threshold plus environmental correction factor). In the server heat dissipation control scenario, the threshold may be dynamically reduced as the environmental temperature rises to trigger the protection mechanism in advance.

[0148] Optionally, in the embodiments of the present application, the above emergency control program may include, but is not limited to, a set of fault response logics pre-stored in the programmable logic device, including device safety state maintenance, redundant module switching, and fault isolation strategies. For example, when it is detected that the baseboard management controller fails, the emergency program may include: immediately locking the fan speed in the safe range, turning off the power supply of non-critical peripherals, and enabling the backup sensor data link. The program execution process is usually designed in the state machine mode, including three stages: fault diagnosis, control strategy switching, and safety state maintenance, and a timeout fallback mechanism is set for each stage. In the power supply anomaly scenario, the emergency program can execute step-down operation, load unloading, and emergency shutdown operations in a hierarchical manner.

[0149] Optionally, in the embodiments of the present application, the reconstruction of the above control strategy may include, but is not limited to, the process of reverse-deriving the current device control parameters based on historical operation data, and its core lies in establishing the mapping relationship between the device state and the control effect. For example, in the heat dissipation control scenario, the reconstruction strategy can establish a rotation speed prediction model based on the ambient temperature by analyzing the correspondence between the temperature and the fan rotation speed in the past 24 hours; in power management, a dynamic voltage regulation table can be generated based on the historical load curve. The reconstruction algorithm may include sliding window regression analysis, time series prediction, or case-based reasoning. For example, a neural network is used to predict the heat dissipation demand in the next 5 minutes to generate preventive control parameters.

[0150] Optionally, in the embodiments of the present application, the above function control and exception alarm may include, but is not limited to, the set of capabilities to maintain the basic operation ability of the device and report fault information externally in an emergency state. For example, the function control may include: maintaining the core power supply of the processor within a safe voltage range, restricting the memory access bandwidth to prevent overheating, enabling the backup network interface, etc.; the exception alarm is implemented by means such as the change of the indicator light mode, system log recording, and pushing the alarm code through the out-of-band management interface. The alarm information usually includes the fault type code, the occurrence timestamp, the affected range, and the recommended handling measures.

[0151] It should be noted that in the bus protocol, the idle duration of the bus and the number of error frames can be detected; in the packet-switching network (such as Ethernet), the aging state and response timeout can be monitored; in the wireless communication scenario, it is necessary to comprehensively judge by combining the signal strength and the bit error rate. The fault tolerance mechanism can be designed as multi-path detection (simultaneously monitoring the hardware heartbeat signal and the software protocol state), time window sliding detection (allowing instantaneous interruptions but triggering alarms for continuous anomalies), or confidence accumulation detection (the longer the duration of the anomaly, the higher the calculated value of the fault probability). For example, in the critical power supply control scenario, the I2C bus data and the status register of the power supply module are simultaneously monitored, and any channel anomaly triggers a pre-alarm.

[0152] In addition, for instantaneous faults (such as communication interruptions), an automatic recovery mechanism can be designed (attempting to re-initialize the link); for permanent faults (such as hardware damage), the device is triggered to operate in a degraded mode; in an uncertain fault scenario, a tentative recovery process can be enabled (such as gradually increasing the control intensity in stages). The program execution mode can include full-automatic processing, semi-automatic (requiring operation and maintenance to confirm key operations), or hybrid mode (automatically handling basic faults and manual intervention for complex problems). For example, when it is detected that the temperature of the baseboard management controller device is too high, the heat dissipation intensity is first automatically increased, and if the anomaly persists, a request is sent for manual confirmation of whether to force a restart.

[0153] Moreover, after the first handover fails, the synchronization data verification time can be extended; when it fails for the second time, switch the communication channel or enable the redundant interface; when it fails for the third time, trigger device isolation. The handover verification dimensions can include firmware version compatibility check, control parameter boundary verification, and device status consistency comparison. For example, in the handover preparation stage, the baseboard management controller device needs to provide an encrypted token containing the current device topology. After the programmable logic device verifies the validity of the token, the parameter synchronization process is allowed to start.

[0154] Through the embodiments of this application, by using the collaborative design of multi-level fault detection and intelligent emergency response, the system abnormal recovery time is shortened, and the control continuity of key devices meets the high availability standard. The control strategy reconstruction algorithm improves the accuracy of emergency parameter generation, and controls the performance fluctuation during device degraded operation within a reasonable range. The dual-verification privilege handover mechanism improves the success rate of control right switching, and significantly enhances the overall reliability of the system.

[0155] It should be noted that this fault recovery architecture has high scalability. In the detection dimension, an AI anomaly prediction model can be integrated to predict potential faults in advance; at the control strategy level, it supports a smooth upgrade from a rule engine to a deep learning model.

[0156] As an optional solution,

[0157] Detect whether the communication signal associated with the baseboard management controller device is lost, including: periodically detecting the heartbeat signal of the baseboard management controller device; if no valid heartbeat signal is received continuously for N times, it is determined that the communication signal is lost, and the fault timestamp and the associated hardware status are recorded, where N is a positive integer greater than 2;

[0158] In response to the start of the emergency control program, reconstruct the control strategy of the target device based on historical operation data, and output emergency control parameters, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future requirements based on historical operation data, outputting emergency control parameters, and resetting the abnormal communication link associated with the baseboard management controller device.

[0159] Optionally, in the embodiments of the present application, the above heartbeat signal may include, but is not limited to, a periodic verification signal for verifying the communication connection status between devices, and its core function is to maintain the activity detection mechanism of the control system. For example, in the communication scenario between a baseboard management controller and a programmable logic device, the heartbeat signal can be designed as a specific data packet sent every 100 milliseconds, containing information such as device identifiers, current operating modes, and check codes. The generation method of the heartbeat signal may include hardware timer triggering, software task scheduling, or event-driven modes. When specifically implemented, differential Manchester coding may be used to enhance anti-interference capabilities, or a timestamp may be added to achieve two-way clock synchronization. In the abnormal detection scenario, the loss of the heartbeat signal may be caused by various reasons such as physical link breaks, processor deadlocks, or buffer overflows, and comprehensive diagnosis needs to be combined with the associated hardware status.

[0160] Optionally, in the embodiments of the present application, the above preset fault recovery algorithm library may include, but is not limited to, a set of intelligent recovery strategies pre-stored in a non-volatile memory, including core functional modules such as device control parameter generation, communication link repair, and security boundary maintenance. For example, in the scenario of a cooling system failure, the algorithm library may include a fan speed prediction algorithm driven by a thermodynamic model, a temperature compensation algorithm based on historical data, and a multi-sensor data fusion algorithm. Its implementation form may include a mathematical function library, a hardware description language module, or a machine learning model. For example, a recurrent neural network is deployed to predict the cooling demand in the next 5 seconds, and emergency speed parameters are generated by combining fuzzy control rules. The call mechanism of the algorithm library is usually designed as an event-driven mode, and the corresponding recovery strategy is dynamically loaded according to the fault type code.

[0161] Optionally, in the embodiments of the present application, the above abnormal communication link may include, but is not limited to, a physical / logical channel with persistent transmission errors or logical connection interruptions, and its abnormal manifestations may include an excessive bit error rate, protocol handshake failure, or signal integrity degradation. For example, in the communication scenario based on the I2C bus, the abnormality may be manifested as a clock stretching phenomenon caused by the bus being continuously pulled low. The reasons for the link abnormality may involve electromagnetic interference, connector oxidation, or firmware logic errors, and the repair strategies may include hardware reset, protocol renegotiation, or redundant path switching.

[0162] It should be noted that in the dimension of signal form, a fixed-frequency square wave, an encrypted data packet, or a modulated carrier signal can be adopted; in terms of the fault tolerance mechanism, an adaptive heartbeat interval (extended cycle during network congestion), multi-path redundancy detection (simultaneously monitoring the hardware level and protocol status), or a confidence accumulation model (the longer the abnormal duration, the higher the failure probability) can be designed. For example, the heartbeat interval is automatically shortened to 50 milliseconds in a high-temperature environment to improve detection sensitivity, and at the same time, a hardware watchdog circuit is enabled as a backup monitoring means. The recovery strategy can include a progressive response: the first loss triggers a link self-check, the second loss initiates parameter synchronization, and the third loss performs a control right switch.

[0163] In addition, the construction strategy of the fault recovery algorithm library can be adapted to multiple technical routes. In the dimension of algorithm type, it can include feedforward control based on a physical model, data-driven feedback regulation, and hybrid intelligent control; at the resource management level, static preloading (all algorithms resident in memory), dynamic loading (reading from the memory on demand), or distributed deployment (part of the algorithms running on a hardware accelerator) can be adopted. For example, a communication repair algorithm with high real-time requirements is solidified in a programmable logic unit, while a complex control parameter generation algorithm runs on a processing system. The learning mechanism can be designed as offline training and online inference, incremental learning, or transfer learning. For example, the weights of a neural network model are continuously optimized according to the historical fault records of the device.

[0164] On the other hand, the implementation method of the communication link reset operation can be processed differently according to the type of abnormality. For an abnormality caused by transient interference, protocol layer re-initialization (re-negotiating the baud rate and parity method) can be adopted; for a fault caused by hardware damage, a backup communication interface can be switched to or a degraded communication mode can be enabled (such as switching gigabit Ethernet to 100-megabit mode). The reset verification mechanism can include link quality assessment (bit error rate test), end-to-end loopback test, and load stress test. For example, a test data packet is sent after the reset is completed and the end-to-end transmission delay and integrity are verified.

[0165] Through the embodiments of this application, by using an intelligent heartbeat monitoring and multi-strategy recovery mechanism, the self-healing ability of the system in an abnormal state is significantly improved. The dynamically loaded fault recovery algorithm can accurately match different fault scenarios, avoiding the limitations of traditional single strategies. The adaptive repair strategy of the communication link effectively reduces the frequency of manual intervention and ensures the stability of continuous operation of the device. The hierarchical abnormal handling architecture takes into account the requirements of real-time response and in-depth analysis, forming a complete closed-loop for fault handling.

[0166] It should be noted that this solution has the ability to expand for future technological evolution. In terms of the detection mechanism, the collaborative monitoring capabilities of edge computing nodes can be integrated; at the algorithm level, the plug-and-play of new artificial intelligence models is supported; in terms of the communication interface dimension, new connection methods such as fiber optic communication and wireless transmission can be adapted. For large-scale device clusters, cross-node fault collaborative processing can be achieved through protocol extension, and an intelligent recovery network with self-organization capabilities can be constructed.

[0167] As an alternative solution, the above method further includes:

[0168] During the reset of the baseboard management controller, perform at least one of the following operations through a programmable logic device:

[0169] Take over the control authority of the target device and adjust the device parameters of the target device based on the real-time data of the sensor device;

[0170] Report the fault information through the alarm channel;

[0171] Periodically attempt to reconstruct the communication link with the baseboard management controller;

[0172] After the baseboard management controller resumes, re-execute the encrypted identity verification and control parameter synchronization operations to complete the authority handover.

[0173] Optionally, in the embodiments of the present application, the above alarm channel may include but is not limited to a fault information transmission mechanism that reports the abnormal state to an external management system through a preset interface, including but not limited to sending an encoded alarm message through a UART (Universal Asynchronous Receiver Transmitter) serial port, triggering an interrupt signal through a PCIe interface, or transmitting a message through an out-of-band management network.

[0174] It should be noted that the form of reporting fault information through the alarm channel can be extended according to the differences in the system architecture. For example, in terms of the transmission protocol dimension, the alarm information can be pushed to the cloud management platform through a protocol, or transmitted to a local monitoring terminal through a protocol, or interact with dedicated hardware through a custom binary protocol. In terms of the content format dimension, the alarm information can include a concise fault code, a detailed log summary, or an attached snapshot of the original sensor data. In terms of the priority dimension, alarms can be classified into emergency alarms (such as power failures), important alarms (such as temperature overlimits), and prompt alarms (such as communication delays), and different response strategies can be associated. The present application does not make specific limitations on this.

[0175] Optionally, in an embodiment of the present application, the above-mentioned encrypted identification verification may include but is not limited to an identity verification process based on an asymmetric encryption algorithm, such as a programmable logic device generates a random challenge code, the baseboard management controller signs it with a private key and returns it, and both parties confirm the legitimacy of the identity by verifying the validity of the signature, including but not limited to supporting multiple encryption standards.

[0176] Optionally, in an embodiment of the present application, the above-mentioned control parameter synchronization may include but is not limited to transmitting device configuration data from a programmable logic device to a baseboard management controller during the authority transfer process, including but not limited to avoiding data loss through a double buffer mechanism, using verification to ensure transmission integrity, and improving synchronization efficiency through multi-link concurrent transmission.

[0177] It should be noted that the mechanism of periodically attempting to rebuild the communication link with the baseboard management controller can adapt to a variety of fault tolerance requirements. For example, in the retry strategy dimension, the programmable logic device can use fixed interval retries, exponential backoff retries, or dynamically adjust the retry frequency based on the load status. In the link selection dimension, the preset main link can be tried first, or the backup link can be automatically switched after the main link fails (such as switching from I2C to SPI (Serial Peripheral Interface)), or multiple links can be detected in parallel to select the optimal path. In the verification dimension, the programmable logic device can perform one-way heartbeat detection, two-way data integrity verification, or encryption handshake protocol verification after the link is rebuilt. This application does not make specific restrictions on this.

[0178] As an optional solution, the above method also includes:

[0179] In response to the server starting up, a modulation signal is outputted by a second processing unit in the programmable logic device to control the heat dissipation device to start up in a preset mode, wherein the target device includes the heat dissipation device;

[0180] Loading an operating system through a first processing unit in a programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes a temperature monitoring device;

[0181] Performing denoising and calibration processing on the operating parameters by a first processing unit to generate a valid temperature parameter set;

[0182] Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register;

[0183] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;

[0184] In response to the baseboard management controller device having completed startup, receive a handover request containing an encrypted identifier sent by the baseboard management controller device through a preset communication link;

[0185] Verify the validity of the encrypted identifier through a second processing unit and feedback the verification result to the first processing unit;

[0186] In the case of successful verification, synchronize historical control parameters, temperature calibration data, and the current speed control signal to the baseboard management controller through the first processing unit. Among them, the baseboard management controller generates a speed control signal that matches the current heat dissipation state of the target device based on the historical control parameters, temperature calibration data, and the current speed control signal;

[0187] In the case of successful handover of control authority, switch the programmable logic device from the working mode to the auxiliary monitoring mode and continuously check the operating state of the baseboard management controller;

[0188] When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, take over the control authority again, where N is a positive integer greater than 2;

[0189] Call a preset fault recovery algorithm library through the first processing unit to reset the baseboard management controller, and output a speed control signal to the heat dissipation device through the second processing unit;

[0190] During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: report fault information through an alarm channel; periodically attempt to reconstruct a communication link with the baseboard management controller;

[0191] In response to the baseboard management controller having been reset, perform encrypted identifier verification and control parameter synchronization operations to complete the handover of authority.

[0192] Exemplarily, assuming a server heat dissipation control system as an example, it includes but is not limited to the following processes:

[0193] S1: Server power-on initialization stage:

[0194] After the second processing unit (PL) starts up, it immediately generates a PWM signal to drive the heat dissipation fan to start in a progressive mode. For different types of heat dissipation devices, the preset mode parameters are dynamically adjusted.

[0195] S2: Operating system loading and data acquisition:

[0196] After the first processing unit (PS) loads the real-time operating system, it periodically collects temperature data through a multiplexing channel in a cyclic manner. Through the data acquisition module, the median denoising process is performed on the continuous 8 sampling values of each channel using the sliding window filtering technology. The calibration module performs online compensation based on the stored sensor offset scale table to generate an effective temperature parameter set.

[0197] S3: Dynamic control strategy generation:

[0198] Based on the effective temperature parameter set, the PS unit executes the PID control algorithm: when the CPU temperature exceeds 65°C, the target duty cycle is calculated in combination with the recent temperature rise rate, and the target value is written into the control register of the PL after boundary constraint.

[0199] S4: Preparation for control authority handover:

[0200] After the baseboard management controller (BMC) completes startup, it sends a handover request packet containing an encrypted identifier through the I2C bus. The identifier content includes the BMC firmware version number, hardware ID, and timestamp. The verification module of the PL completes decryption and hash verification, and feeds back a verification success flag to the PS after confirming the legitimacy of the request.

[0201] S5: Parameter synchronization and switching execution:

[0202] The PS transfers the historical control parameters, temperature calibration data, and current PWM phase information to the BMC through the DMA channel. The phase alignment module of the BMC performs the following operations:

[0203] S5-1, Lock the PWM clock source of the PL and measure the phase difference to be 15ns;

[0204] S5-2, Adjust the initial value of the internal timer to achieve waveform synchronization;

[0205] S5-3, Complete the gradual transition of the duty cycle within 3 PWM cycles.

[0206] S6: Exception handling and recovery mechanism:

[0207] When the PL fails to detect the 100ms periodic heartbeat signal of the BMC for 5 consecutive times (N = 5), a hardware interrupt is triggered to the PS. The fault recovery algorithm library calls the LSTM prediction model to predict the future 3-minute heat dissipation demand based on the temperature data in the past 1 hour, and generates an emergency duty cycle curve. The hardware acceleration module (such as the DSP array of the FPGA) completes the calculation, and the PL outputs a speed regulation signal.

[0208] During the BMC reset, the PL takes over the heat dissipation control authority and sends an alarm message containing a fault code (such as 0x5A3) through an independent UART channel every 5 seconds. The system attempts to reconstruct the I2C link every 10 seconds until the BMC recovers and then re-executes the complete handover process.

[0209] Through the collaborative design of progressive permission transfer and intelligent emergency response, the present application realizes the stable control of the entire life cycle of the server cooling system. The hardware-level generation of PWM signals ensures the speed regulation response speed, the fusion processing of multi-source temperature data improves the control accuracy, and the encryption verification mechanism guarantees the security of permission transfer. The combination of dynamic algorithms and historical data effectively responds to sudden load fluctuations and environmental changes. The dual control architecture maintains the cooling efficiency in case of failures, avoiding the risk of system overheating caused by single-point failures. The modular design supports flexible expansion and can be adapted to various cooling solutions such as air cooling and liquid cooling.

[0210] The following further explains and illustrates the present application with specific examples:

[0211] With the large-scale application of servers in fields such as big data, Internet AI technology, etc., the data processed by servers and the scale of server hardware in data centers have also shown explosive growth. The server management technology field has put forward higher and higher requirements for server reliability and availability, and traditional server management technologies face three challenges:

[0212] The baseboard management controller (BMC) is the core of server hardware management. It monitors the security of the entire server environment and the hardware status. Therefore, the reliability of the BMC firmware is the most important part to ensure the functions of the entire server. Currently, in the industry, an embedded dedicated chip running the Linux operating system is usually used as the core of hardware management. The management links relied on in the hardware single-board design, such as I2C, ADC, PWM, and sensor management, all depend on the reliability of the BMC chip operation. If software and hardware failures occur during server operation, resulting in abnormal BMC operation or hardware link failure, it may cause abnormal server operation, such as the failure of the fan cooling function and the abnormal monitoring of sensor functions.

[0213] The server motherboard uses programmable logic device (FPGA / CPLD) chips for functions such as hardware chip power supply timing management and board hardware signal management. The baseboard management controller usually relies on the CPLD to implement functions such as hardware signal detection and fault abnormality identification. At the same time, the programmable logic device itself supports the implementation of the SOC (System-on-Chip) function, and it can integrate multiple functional modules such as storage, processing, logic, and interfaces in the CPLD or FPGA without increasing the hardware cost.

[0214] Figure 3 A partial schematic diagram of the hardware structure of an optional server collaborative control method provided by an embodiment of the present application, as Figure 3As shown in the figure, a programmable logic device on-chip system generally includes two parts: a PS processor system and a PL programmable logic. This architecture enables this series of chips to have the advantages of FPGA hardware programmability while also possessing the advantages of ASIC in terms of energy consumption, performance, and compatibility. At the same time, the processor system of the on-chip system supports the operation of general operating systems such as Linux, facilitating flexible software programming application development. The PS and PL are connected through a standard AXI interface, and in the software system, the PL part can be regarded as a part of the peripheral for management.

[0215] Based on the above background content, this application proposes a server controller management method based on a programmable on-chip system. This method can utilize the heterogeneous programmable logic resources on the server motherboard and cooperate with the baseboard management controller chip to achieve the management of server hardware, improving the stability and availability of the server at a low hardware cost.

[0216] This application proposes a server collaborative management method based on a heterogeneous programmable on-chip system, which is mainly applied in a typical storage server architecture. The server controller motherboard includes a programmable logic device FPAG that supports the on-chip SOC mode, a baseboard management controller BMC, and several temperature sensors. The hardware link between the BMC and the FPGA consists of an I2C data access link, GPIO hardware signals, and PWM control signals, etc. When the FPGA is used as an on-chip SOC system, its internal logic resources are divided into two parts. Among them, the PL logic module realizes functions such as server hardware power supply control and hardware pin signal detection. The PS part can be implemented using a built-in soft core or a hard core with a self-contained processor core to achieve system data processing. In order to ensure the real-time performance of the system, a real-time operating system can be used as the running software system. The peripheral sensors are respectively connected to the front-stage multi-channel I2C multiplexer chip through the I2C link. This multiplexer uses an I2C multiplexing chip that supports multi-channel access and supports the BMC / FPGA to access the rear-stage sensor chips in a time-sharing manner. Figure 4 It is a schematic diagram of the hardware structure design of an optional server collaborative control method provided by an embodiment of this application, as Figure 4 shown, including but not limited to solving the following three scenarios:

[0217] Scenario 1: When the BMC system has not been started and completed in the power-on startup scenario, the fan speed control can be independently completed by the FPGA. That is, the real-time operating system of the PS part can read the sensor data for basic speed regulation (calculating the PWM output of the fan speed according to the speed regulation algorithm) after a quick power-on startup (in seconds), and then access the PWM output module in the PL through the AXI bus to output a reasonable heat dissipation speed-regulated fan speed. This process is generally completed within 2 minutes after the server is powered on.

[0218] Scenario 2: During the operation of the server, the Baseboard Management Controller (BMC) performs a reset and restart operation. For example, during online firmware update / software operation anomaly, when the communication between the FPGA and the BMC is interrupted (status vector verification data fails), after the PL module in the FPGA detects the BMC anomaly, it will perform a reset and repair operation on the BMC. Before triggering the BMC reset operation, the status vector verification module in the PL logic module triggers an interrupt and reports it to the PS operating system via the AXI bus. The application program running on the PS triggers the interrupt handler to take over the fan.

[0219] Scenario 3: To solve the problem of low data access polling efficiency in the current Baseboard Management Controller (BMC), in traditional server firmware management, the BMC accesses external chips or signals mostly by polling. With the above architecture, the interrupt trigger detection and control hardware can be offloaded to the FPGA chip. The reset detection and repair of the motherboard core chip are completed by the combination of the PL and the PS, without the need for the BMC to adopt it, improving the efficiency of fault recovery.

[0220] This application realizes a stable and reliable server controller hardware management function under the above hardware structure. Due to the adoption of a new heterogeneous redundant hardware design, the reliability and availability of the server operation are improved without increasing the hardware design complexity, thereby enhancing the stability of the entire server operation.

[0221] This application adopts a three-level control right transfer mechanism to realize system heat dissipation speed regulation management. In the first stage, the system initialization stage, in this stage, the FPGA takes over the system heat dissipation permission. The specific implementation is as follows:

[0222] Based on Figure 4 In the server hardware management design block diagram, after the controller is powered on normally, the FPGA defaults to providing a constant speed control fan. After the PS part inside the FPGA loads the real-time operating system and starts, the program running on the PS polls to read the temperature of the external associated core sensors. The PS application program calculates the heat dissipation parameters (PID and PWM) based on the collected temperature data, enables the PL logic to take over the fan speed regulation control logic, and outputs the PWM signal of the target speed.

[0223] The second stage: The BMC takeover stage:

[0224] After waiting for the BMC system to start up (about 90 s), when the communication with the FPGA status vector verification module is successful and the FPGA GPIO pins receive the BMC communication identifier, it indicates that the complete heat dissipation speed regulation ability has been achieved. The FPGA PL logic releases the fan speed regulation control right, and the BMC inputs PWM for fan control and speed regulation. At the same time, in order to accelerate the full transfer process of control, the FPGA and the BMC need to quickly synchronize the heat dissipation sensor data. The synchronization content includes exchanging sensor calibration data through the shared memory area. Secondly, the PWM phase synchronization is executed to ensure that the BMC can smoothly take over the entire heat dissipation control. After the handover is completed, the BMC system gradually collects the ambient temperature and all temperature sensor information of the controller, and the FPGA resumes to the monitoring mode.

[0225] The third stage: the fault handling process:

[0226] When the FPGA internal PL module status vector verification mechanism continuously detects three communication failures, it will trigger a Level-2 interrupt to the PS system. The PS starts the hardware controller program (preheating algorithm library) and reconfigures the PWM controller through the AXI-Lite bus. During the BMC reset and restart, the FPGA will take over functions including fan control, temperature sensor monitoring and alarm functions;

[0227] Aiming at the problem of low polling efficiency and slow response of the BMC proposed in Scenario 3, this application proposes to take advantage of the heterogeneous hardware design to solve the problem of slow response of traditional BMC heat dissipation speed regulation, accelerate hardware fault handling and take over emergency fault scenarios, and adopt a dual-mode architecture of PL hardware fast response + PS intelligent decision-making. The specific implementation process is as follows:

[0228] S1, Initialize the programmable event processing unit (PEPU) of the PL module and associate the event-related interrupt pins;

[0229] S2, The PS module initialization program, including the threshold register group (temperature / voltage / current, etc.) of the emergency event signal;

[0230] S3, Configure the event filtering parameters (sliding window size / confidence threshold) to ensure that emergency events will not be accidentally triggered by anomalies;

[0231] S4, When a preset emergency event occurs, directly trigger the PL emergency response action. According to the severity level of the event, trigger the AXI interrupt to report to the PS module. When the event severity level is greater than the software processing, the PL can independently execute the fusing strategy and stop the operation of abnormal function modules and other strategies. If the event verification level requires PS analysis, after reporting the AXI interrupt, the PS will perform fault analysis and then process it, such as the strategy of taking over the fan speed regulation control right and other strategies.

[0232] The server collaborative management method based on heterogeneous programmable system-on-chip proposed in this application can bring the following beneficial effects:

[0233] Improved reliability: When the BMC fails, the system switching time is < 200ms (in traditional solutions, software restart or thread self-recovery etc. all require at least 30s);

[0234] The heterogeneous hardware design reuses FPGA resources, reduces the demand for dedicated monitoring chips, and reduces the BOM cost;

[0235] Noise improvement: Through more refined speed regulation means and the handover strategy of speed regulation control rights, the server can have lower noise and better user experience;

[0236] Hardware-accelerated fault handling improves the overall availability of the system and prevents the spread of hardware failures caused by factors such as low monitoring efficiency, which may lead to other problems.

[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0238] The embodiment of this application also provides a server collaborative control device, as Figure 5 shown, the device includes:

[0239] An initialization module, configured to, in response to the startup of the server, load the operating system, and take over the control authority of the target device associated with the server according to the operating parameters corresponding to the sensor device. Wherein, the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device, and the baseboard management controller device;

[0240] A verification module, configured to, in response to the completion of the startup of the baseboard management controller device, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device in the case of successful verification;

[0241] A recovery module, configured to, in response to detecting an abnormality of the baseboard management controller device, take over the control authority again. Wherein, the programmable logic device synchronizes the control parameters of the target device with the baseboard management controller device through at least one communication link, and the at least one communication link includes the preset communication link.

[0242] As an optional solution, in response to the startup of the server, loading the operating system and taking over the control authority of the target device associated with the server according to the sensor data corresponding to the sensor device includes:

[0243] In response to the server being started, outputting a modulation signal through a second processing unit in the programmable logic device so as to cause the target device to enter an initial operation state;

[0244] Periodically acquiring operating parameters through a first processing unit in a programmable logic device;

[0245] The first processing unit calculates the target parameter based on the operating parameter using a preset algorithm, and writes the target parameter into a control register corresponding to the second processing unit;

[0246] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;

[0247] Before the baseboard management control device completes startup, the target parameters are continuously updated in response to changes in the operating parameters.

[0248] As an optional solution, periodically obtaining the operating parameters through the first processing unit in the programmable logic device includes:

[0249] Using a multiplexed channel to collect data of a plurality of sensor devices as operating parameters in a time-sharing manner through a first processing unit;

[0250] De-noising the abnormal data in the operating parameters by the first processing unit, and marking invalid data segments;

[0251] The valid data is stored in the shared cache area through the first processing unit, so as to be called by the first processing unit during calculation.

[0252] As an optional solution, the first processing unit calculates the target parameter based on the operating parameter using a preset algorithm, and writes the target parameter into a control register corresponding to the second processing unit, including:

[0253] Generate an adjustment instruction based on a change trend of the operating parameter by the first processing unit;

[0254] converting the adjustment instruction into a control instruction recognizable by the second processing unit by the first processing unit, wherein the target parameter includes the control instruction;

[0255] The control instruction is written into the control register through the first processing unit.

[0256] As an optional solution, the above device further includes at least one of the following:

[0257] Accelerating the operation of the target parameter by using a parallel computing unit deployed in the first processing unit;

[0258] The data stream corresponding to the operating parameter is received by the first processing unit using a high-speed data interface.

[0259] As an alternative solution, in response to the baseboard management controller having completed startup, perform interactive verification with the baseboard management controller based on a preset communication link, and transfer control authority to the baseboard management controller in the case of successful verification, including:

[0260] In response to the baseboard management controller having completed startup, receive a privilege transfer request sent by the baseboard management controller via the preset communication link;

[0261] Verify the validity of the privilege transfer request, and in the case of a valid verification result, synchronize historical control parameters and operating parameters to the baseboard management controller;

[0262] Execute a phase matching operation corresponding to the control parameters, so that the control parameters subsequently generated by the baseboard management controller match the control parameters generated by the programmable logic device;

[0263] In the case of successful transfer of control authority, switch the working mode of the programmable logic device to the auxiliary monitoring mode.

[0264] As an alternative solution, execute a phase matching operation corresponding to the control parameters, so that the control parameters subsequently generated by the baseboard management controller match the control parameters generated by the programmable logic device, including:

[0265] Obtain the signal period and duty cycle of the control parameters output to the target device;

[0266] In response to receiving a status confirmation signal, send the signal period and duty cycle to the baseboard management controller, so that the baseboard management controller generates alternative control parameters to complete the privilege switch, where the status confirmation signal is used to indicate that the baseboard management controller has completed startup, and the alternative control parameters represent the parameters for the baseboard management controller to control the target device.

[0267] As an alternative solution, in the case of successful verification, synchronize historical control parameters and operating parameters to the baseboard management controller, including:

[0268] In the case of successful verification, write the operating parameters to the shared storage area after calibration and encryption, where the baseboard management controller is set to read and decrypt the operating parameters from the shared storage area in a memory access manner, and update the control parameters of the target device after verifying data integrity.

[0269] As an alternative solution, in response to detecting an abnormality in the baseboard management controller, regain control authority, including:

[0270] Detect whether the communication signal associated with the baseboard management controller is lost;

[0271] When a communication signal loss is continuously detected and the cumulative number of signal losses reaches a preset threshold, start the emergency control program;

[0272] In response to the start of the emergency control program, reconstruct the control strategy of the target device based on historical operation data and output emergency control parameters. Among them, during the reset of the baseboard management controller device, the programmable logic device performs function control and exception warning of the target device;

[0273] After the baseboard management controller device is restored, perform interactive verification with the baseboard management controller device based on a preset communication link, and transfer the control authority to the baseboard management controller device in the case of successful verification.

[0274] As an alternative solution,

[0275] Detect whether the communication signal associated with the baseboard management controller device is lost, including: periodically detecting the heartbeat signal of the baseboard management controller device; if a valid heartbeat signal is not received continuously for N times, it is determined that the communication signal is lost, and the fault timestamp and associated hardware status are recorded, where N is a positive integer greater than 2;

[0276] In response to the start of the emergency control program, reconstruct the control strategy of the target device based on historical operation data and output emergency control parameters, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future requirements based on historical operation data, outputting emergency control parameters, and resetting the abnormal communication link associated with the baseboard management controller device.

[0277] As an alternative solution, the above device further includes:

[0278] During the reset of the baseboard management controller, perform at least one of the following operations through the programmable logic device:

[0279] Take over the control authority of the target device and adjust the device parameters of the target device based on the real-time data of the sensor device;

[0280] Report fault information through the alarm channel;

[0281] Periodically attempt to reconstruct the communication link with the baseboard management controller;

[0282] After the baseboard management controller is restored, re-execute the encryption identification verification and control parameter synchronization operations to complete the authority transfer.

[0283] As an alternative solution, the above device further includes:

[0284] In response to the startup of the server, output a modulation signal through the second processing unit in the programmable logic device to control the heat dissipation device to start in a preset mode, where the target device includes the heat dissipation device;

[0285] Loading an operating system through a first processing unit in a programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes a temperature monitoring device;

[0286] Performing denoising and calibration processing on the operating parameters by a first processing unit to generate a valid temperature parameter set;

[0287] Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register;

[0288] activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit;

[0289] In response to the baseboard management control device having completed startup, receiving a handover request including an encrypted identifier sent by the baseboard management control device through a preset communication link;

[0290] Verify the validity of the encryption identifier through the second processing unit, and feed back the verification result to the first processing unit;

[0291] If the verification is passed, the historical control parameters, the temperature calibration data and the current speed regulation signal are synchronized to the baseboard management controller through the first processing unit, wherein the baseboard management controller generates a speed regulation signal matching the current heat dissipation state of the target device based on the historical control parameters, the temperature calibration data and the current speed regulation signal;

[0292] When the control authority is transferred, the programmable logic device is switched from the working mode to the auxiliary monitoring mode to continuously verify the operating status of the baseboard management controller;

[0293] When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, it takes over the control authority again, where N is a positive integer greater than 2;

[0294] Reset the baseboard management controller by calling a preset fault recovery algorithm library through the first processing unit, and output a speed regulation signal to the heat dissipation device through the second processing unit;

[0295] During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: reporting fault information through an alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller;

[0296] In response to the baseboard management controller being reset, encryption identification verification and control parameter synchronization operations are performed to complete the authority transfer.

[0297] For the descriptions of the features in the corresponding embodiments of the above server collaborative control device, reference can be made to the relevant descriptions in the corresponding embodiments of the server collaborative control method, which will not be elaborated here one by one.

[0298] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the server collaborative control method.

[0299] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the server collaborative control method when running.

[0300] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disk, and other various media that can store computer programs.

[0301] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server collaborative control method.

[0302] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server collaborative control method.

[0303] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0304] The above has introduced in detail a server collaborative control method, a storage medium, and an electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A server collaborative control method, characterized in that, Applications in programmable logic devices include: In response to the server starting up, the operating system is loaded, and the control authority of the target device associated with the server is taken over according to the operating parameters corresponding to the sensor device, wherein the server is deployed with a server controller, and the server controller includes the programmable logic device, the sensor device and the baseboard management control device; In response to the baseboard management control device having completed startup, performing interactive verification with the baseboard management control device based on a preset communication link, and transferring the control authority to the baseboard management control device if the verification passes; In response to detecting that an abnormality occurs in the baseboard management control device, retaking the control authority, wherein the programmable logic device and the baseboard management control device synchronize the control parameters of the target device through at least one communication link, and the at least one communication link includes the preset communication link; The method of responding to the server startup, loading the operating system, and taking over the control authority of the target device associated with the server according to the operating parameters corresponding to the sensor device includes: periodically acquiring the operating parameters through the first processing unit in the programmable logic device; calculating the target parameters based on the operating parameters by using a preset algorithm through the first processing unit, and writing the target parameters into the control register corresponding to the second processing unit; activating the adaptive control mode according to the control register and taking over the control authority through the second processing unit; In response to the baseboard management control device having completed startup, interactive verification is performed with the baseboard management control device based on a preset communication link, and the control authority is transferred to the baseboard management control device if the verification passes, including: if the verification passes, synchronizing the historical control parameters and the operating parameters to the baseboard management control device; when the control authority is transferred, switching the working mode of the programmable logic device to the auxiliary monitoring mode.

2. The server collaborative control method according to claim 1, wherein, In response to the server starting up, the operating system is loaded, and the control authority of the target device associated with the server is taken over according to the operating parameters corresponding to the sensor device, including: In response to the server being started, outputting a modulation signal through a second processing unit in the programmable logic device so as to enable the target device to enter an initial operation state; periodically acquiring the operating parameter through the first processing unit in the programmable logic device; Calculate the target parameter by the first processing unit based on the operating parameter using a preset algorithm, and write the target parameter into a control register corresponding to the second processing unit; activating the adaptive control mode and taking over the control authority according to the control register by the second processing unit; Before the baseboard management control device completes startup, the target parameter is continuously updated in response to the change of the operating parameter.

3. The server collaborative control method according to claim 2, wherein The periodically acquiring the operating parameter by the first processing unit in the programmable logic device includes: Using the first processing unit to collect data of the plurality of sensor devices in a time-sharing manner using a multiplexed channel as the operating parameter; The first processing unit denoises the abnormal data in the operating parameters and marks the invalid data segments; The first processing unit stores the valid data in the shared buffer for the first processing unit to call during calculation.

4. The server collaborative control method according to claim 2, wherein The first processing unit calculates the target parameter based on the operating parameters by using a preset algorithm and writes the target parameter into the control register corresponding to the second processing unit, including: The first processing unit generates an adjustment instruction based on the change trend of the operating parameters; The first processing unit converts the adjustment instruction into a control instruction recognizable by the second processing unit, where the target parameter includes the control instruction; The first processing unit writes the control instruction into the control register.

5. The server collaborative control method according to claim 4, wherein The method further includes at least one of the following: The parallel computing unit deployed in the first processing unit accelerates the operation of the target parameter; The first processing unit receives the data stream corresponding to the operating parameters by using a high-speed data interface.

6. The server collaborative control method according to claim 1, wherein In response to the baseboard management controller having completed startup, perform an interactive verification with the baseboard management controller based on a preset communication link, and transfer the control authority to the baseboard management controller when the verification is passed, including: In response to the baseboard management controller having completed startup, receive the authority transfer request sent by the baseboard management controller through the preset communication link; Verify the validity of the authority transfer request, and when the verification result is valid, synchronize the historical control parameter and the operating parameter to the baseboard management controller; Execute the phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management controller matches the control parameter generated by the programmable logic device; When the control authority is completed, switch the working mode of the programmable logic device to the auxiliary monitoring mode.

7. The server collaborative control method according to claim 6, wherein The executing the phase matching operation corresponding to the control parameter so that the control parameter subsequently generated by the baseboard management controller matches the control parameter generated by the programmable logic device includes: Obtain the signal period and duty cycle of the control parameter output to the target device; In response to receiving the status confirmation signal, send the signal period and the duty cycle to the baseboard management controller so that the baseboard management controller generates an alternative control parameter to complete the authority switch, where the status confirmation signal is used to indicate that the baseboard management controller has completed startup, and the alternative control parameter represents the parameter for the baseboard management controller to control the target device.

8. The server collaborative control method according to claim 6, characterized in that, The synchronizing the historical control parameter and the operating parameter to the baseboard management controller when the verification is passed includes: When the verification is passed, write the operating parameter into the shared storage area after calibration and encryption, where the baseboard management controller is configured to read and decrypt the operating parameter from the shared storage area in a memory access manner, and update the control parameter of the target device after verifying the data integrity.

9. The server collaborative control method according to claim 1, characterized in that In response to detecting that the baseboard management control device is abnormal, retaking the control authority includes: Detecting whether a communication signal associated with the baseboard management control device is lost; In the event that the communication signal is continuously detected to be lost and the cumulative number of times the signal is lost reaches a preset threshold, an emergency control procedure is initiated; In response to the emergency control program being started, reconstructing the control strategy of the target device based on historical operation data and outputting emergency control parameters, wherein during the resetting of the baseboard management control device, the programmable logic device performs function control and abnormal alarm of the target device; After the baseboard management control device is restored, interactive verification is performed with the baseboard management control device based on the preset communication link, and the control authority is transferred to the baseboard management control device if the verification passes.

10. The server collaborative control method according to claim 9, characterized in that: The detecting whether the communication signal associated with the baseboard management control device is lost includes: periodically detecting the heartbeat signal of the baseboard management control device; if no valid heartbeat signal is received for N consecutive times, it is determined that the communication signal is lost, and the fault timestamp and the associated hardware status are recorded, wherein N is a positive integer greater than 2; In response to the emergency control program being started, the control strategy of the target device is reconstructed based on historical operating data, and emergency control parameters are output, including: calling a preset fault recovery algorithm library, selecting a control strategy that matches the current operating parameters; predicting future needs based on historical operating data, outputting the emergency control parameters, and resetting the abnormal communication link associated with the baseboard management control device.

11. The server collaborative control method according to claim 1, wherein The method further comprises: During the resetting of the baseboard management controller, at least one of the following operations is performed by the programmable logic device: Taking over the control authority of the target device, and adjusting the device parameters of the target device based on the real-time data of the sensor device; Report fault information through the alarm channel; periodically attempting to reestablish a communication link with the baseboard management controller; After the baseboard management controller is restored, the encryption identification verification and the control parameter synchronization operation are re-executed to complete the authority transfer.

12. The server collaborative control method according to claim 1, wherein The method further comprises: In response to the server starting up, outputting a modulation signal through the second processing unit in the programmable logic device to control the heat dissipation device to start up in a preset mode, wherein the target device includes the heat dissipation device; Loading an operating system through a first processing unit in the programmable logic device, and collecting operating parameters of a plurality of temperature monitoring devices in a time-sharing manner through a multiplexing channel, wherein the sensor device includes the temperature monitoring device; Performing denoising and calibration processing on the operating parameters by the first processing unit to generate a valid temperature parameter set; Calculating a target duty cycle based on the effective temperature parameter set by the first processing unit and writing the target duty cycle into a control register; activating the adaptive control mode according to the control register and taking over the control authority by the second processing unit; In response to the fact that the baseboard management control device has completed startup, receive a handover request containing an encrypted identifier sent by the baseboard management control device through the preset communication link; Verify the validity of the encrypted identifier through the second processing unit, and feedback the verification result to the first processing unit; In the case of successful verification, synchronize historical control parameters, temperature calibration data, and the current speed regulation signal to the baseboard management controller through the first processing unit, where the baseboard management controller generates a speed regulation signal matching the current heat dissipation state of the target device based on the historical control parameters, the temperature calibration data, and the current speed regulation signal; In the case of successful handover of the control authority, switch the programmable logic device from the working mode to the auxiliary monitoring mode, and continuously check the operating state of the baseboard management controller; When the second processing unit fails to detect the heartbeat signal of the baseboard management controller for N consecutive times, take over the control authority again, where N is a positive integer greater than 2; Reset the baseboard management controller by invoking a preset fault recovery algorithm library through the first processing unit, and output a speed regulation signal to the heat dissipation device through the second processing unit; During the reset of the baseboard management controller, the programmable logic device performs at least one of the following operations: report fault information through the alarm channel; periodically attempt to reconstruct a communication link with the baseboard management controller; In response to the fact that the baseboard management controller has been reset, perform encrypted identifier verification and the control parameter synchronization operation to complete the authority handover.

13. An electronic device, characterized in that, Comprising: A memory for storing computer programs; A processor for implementing the steps of the server collaborative control method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the server collaborative control method according to any one of claims 1 to 11 when executed by a processor.

15. A computer program product comprising a computer program, characterized in that, The computer program implements the steps of the server collaborative control method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Server management and control method, system and device and computer readable storage medium

    CN118152161A

  • Server control method and device, equipment, medium and computer program product

    CN118796010A