A hardware acceleration device management system and method applied to a network security product
By using the PCIe to M.2 adapter board module and the corresponding monitoring, allocation, and scheduling modules, the problem that network security products cannot directly use standard hardware acceleration devices has been solved. This has enabled plug-and-play functionality and adaptive heat dissipation, reduced development and replacement costs, and improved processing efficiency and device stability.
Patent Information
- Application Number
- CN202511078960.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-08-02
AI Technical Summary
Existing cybersecurity products' hardware acceleration devices cannot directly use standard modules available on the market, resulting in poor device compatibility, power management that is not adapted to dynamic changes in tasks, and insufficient correlation between the heat dissipation system and task load, leading to energy waste and device aging.
It adopts a PCIe to M.2 adapter board module, a structural support module, a real-time load monitoring module, a dynamic resource allocation module, a cross-device task scheduling module, and a device compatibility adaptation module. By monitoring network tasks and device performance in real time and dynamically allocating tasks, it realizes plug-and-play hardware acceleration devices and adaptive heat dissipation control.
It enables plug-and-play hardware acceleration devices, reduces development and replacement costs, improves processing efficiency by 40%, simplifies device management, and ensures stable operation of devices within a safe temperature range.
Smart Images

Figure CN120909969B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hardware acceleration device management technology for network security products, specifically a hardware acceleration device management system and method applied to network security products. Background Technology
[0002] For cybersecurity products, the external interface is typically only a network interface. Previously, much software processing was handled entirely by the CPU. In recent years, with the emergence of hardware acceleration devices, the CPU load can be reduced, while system efficiency is improved. Therefore, the compatibility requirements of cybersecurity products with hardware acceleration devices are becoming increasingly stringent.
[0003] Defects and shortcomings of existing technology:
[0004] Hardware acceleration devices on the market are typically in the form of standard PCIe cards or M.2 slots. Since network security products require expansion via a clip-on PCIe connector to create non-standard PCIe cards for external interfaces, it's currently not possible to directly use standard modules on network security products.
[0005] To use acceleration devices, directly designing them according to cybersecurity product specifications would require a complete circuit redesign, which necessitates a lengthy development cycle. Furthermore, given the relatively low demand for cybersecurity equipment, the cost of newly designed acceleration devices would be significantly higher. Similarly, replacing existing hardware with new acceleration devices would also require redesign, resulting in substantial replacement costs.
[0006] In traditional cybersecurity products, M.2 devices, as hardware acceleration devices, suffer from several shortcomings in existing power management methods. Firstly, power consumption configurations are typically set manually and statically in the initial stages, failing to adapt to the dynamic nature of cybersecurity tasks. For instance, during peak office hours, network traffic and encryption tasks are relatively low, yet the device still operates at high-load power consumption, leading to energy waste. Furthermore, prolonged high-power operation accelerates hardware aging. Secondly, the cooling system lacks effective correlation with the device's workload and power consumption. Traditional cooling controls are often based on fixed temperature thresholds, activating cooling measures when the device temperature exceeds these thresholds. This approach fails to consider the power consumption differences arising from different task types and the cooling requirements of the device under varying performance states. For example, during high-intensity encryption tasks, device power consumption increases significantly, generating substantial heat. However, if the cooling system fails to adjust its strategy according to the workload in a timely manner, the device temperature becomes excessively high, impacting performance and stability, and potentially causing system failures and disrupting the normal operation of cybersecurity products.
[0007] To address the above problems, this invention proposes a hardware acceleration device management system and method for cybersecurity products. Summary of the Invention
[0008] The purpose of this invention is to provide a hardware acceleration device management system and method for cybersecurity products, in order to solve the problems raised in the prior art.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] A hardware acceleration device management system for cybersecurity products includes a PCIe to M.2 adapter module, a structural support module, a real-time load monitoring module, a dynamic resource allocation module, a cross-device task scheduling module, and a device compatibility adaptation module. The PCIe to M.2 adapter module is the hardware carrier, and its input end connects to the PCIe interface of the cybersecurity product. The device features an x8 slot with two M.2 connectors at the output. The structural support module connects the adapter board to the M.2 device, using physical brackets and screw holes to secure the device to the adapter board. The system is characterized by: a real-time load monitoring module that extracts traffic and encryption / decryption task volume in real time via a network task acquisition unit, and obtains the device's idle core count, remaining bandwidth, and temperature via a device performance acquisition unit, generating real-time JSON reports via a data report generation unit; a dynamic resource allocation module that calculates a resource score based on the formula (device resource score = (idle core count / total core count) × 0.6 + (remaining memory bandwidth / total memory bandwidth) × 0.4) via a fusion model calculation unit, prioritizing hardware encryption devices for encryption tasks, allocating traffic tasks to devices with higher scores, and evenly distributing decryption tasks across both devices, with the task allocation execution unit implementing the task allocation; and a cross-device task scheduling module that connects the allocation module to the M.2 device, calculating device utilization via a load threshold monitoring unit. When utilization exceeds 80% for five consecutive samples, the task migration execution unit filters migrateable tasks and migrates them to the target device. Simultaneously, the BIOS mode switching unit dynamically configures the PCIe lanes, maintaining PCIe in single-device mode. In X8 mode, the dual-device mode is configured as X4+X4 mode; the device compatibility adaptation module connects the M.2 device and the system software layer, reads the device ID and firmware information through the firmware identification unit, automatically loads the compatible driver by the driver management unit, updates the device list in real time and triggers resource reallocation through the hot-plug detection unit, obtains device power consumption parameters and establishes a task power consumption correlation model through the power consumption feature identification unit, dynamically adjusts the device performance status and cooling fan speed based on load prediction by the dynamic power consumption scheduling unit, and adjusts the fan speed through the temperature and power consumption relationship model and triggers frequency reduction protection and task migration when the temperature exceeds the limit.
[0011] The PCIe to M.2 adapter module includes an allocation unit and a clock processing unit;
[0012] The distribution unit splits the X8 signal input from the PCIE X8 slot of the network security product into two independent PCIE X4 signals through internal PCB traces. The two X4 signals correspond to the two M.2 connectors at the output end of the adapter board, respectively. When an M.2 device is inserted into a single slot, only one X4 signal transmits data. When M.2 devices are inserted into both slots, the two X4 signals transmit data in parallel. In this case, the PCIE slot needs to be configured to X4+X4 mode through the BIOS to support parallel communication between the two devices.
[0013] The clock processing unit integrates a clock buffer chip CLK BUFFER, which copies the clock signal CLK input from the PCIe slot into two synchronous clock signals. The two clock signals are connected to two M.2 connectors respectively, providing synchronous clock signals for the corresponding M.2 devices and ensuring the clock synchronization of the devices when operating in single-slot and dual-slot configurations.
[0014] The structural support module includes a physical bracket unit and a screw fixing unit;
[0015] The physical bracket unit is a rectangular frame structure made of insulating material. The frame size matches the edge of the PCB board of the PCIe to M.2 adapter board and is fixed to the edge of the adapter board by welding. The inner side of the frame is provided with a limiting groove corresponding to the size of the M.2 device. The length of the groove is compatible with standard M.2 devices, ensuring that after the device is inserted into the M.2 connector, the bottom surface is parallel to the PCB board of the adapter board and the side is embedded in the groove for positioning.
[0016] The screw fixing unit is responsible for fixing the M.2 device to the bracket. The top of the physical bracket has threaded holes corresponding to the screw holes of the M.2 device. After the M.2 device is inserted into the connector and embedded in the limiting groove, a countersunk screw is passed through the screw hole on the surface of the device and screwed into the threaded hole of the bracket to fix the top surface of the device to the bracket. The number of screws configured for each M.2 device corresponds to the number of fixing holes of the standard M.2 device, ensuring that the device maintains a stable electrical connection with the adapter plate in a vibration environment.
[0017] The real-time load monitoring module includes a network task acquisition unit, a device performance acquisition unit, and a data report generation unit.
[0018] The network task acquisition unit establishes a data channel with the task processing queue through the network protocol stack interface of the network security product, extracts the number of bytes of inbound / outbound traffic from the network card driver layer in real time, and calculates the average traffic rate at a preset fixed interval. The calculation formula is as follows:
[0019] ;
[0020] Where Δ bytes represent the difference in bytes between two adjacent samples, and Δ time represents the sampling interval in seconds;
[0021] While calculating the average traffic rate, the encryption / decryption task scheduling queue is monitored to count the number of new task entries per second, forming real-time data on the amount of data encryption tasks and data decryption tasks, which directly reflects the current network task load.
[0022] The device performance acquisition unit obtains device status data through the firmware management interface of the M.2 device. Specifically, it queries the CPU / MCU core status register to count the number of idle cores, obtaining the number of idle computing cores. The number of idle cores is defined as registers with a utilization rate of ≤5%. Then, it obtains the used memory bandwidth through the memory controller status register and, combined with the device's nominal total memory bandwidth, calculates the remaining bandwidth using the following formula:
[0023] Remaining memory bandwidth = Total memory bandwidth - Used memory bandwidth;
[0024] Finally, by reading the register value of the onboard temperature sensor, the real-time temperature data of the M.2 device is obtained. All parameters are read in real time through a standardized interface.
[0025] The data report generation unit timestamps the data output by the network task acquisition unit and the device performance acquisition unit, and generates a structured report using UTC time format. The report includes fields such as traffic rate, encryption task volume, decryption task volume, number of idle cores, remaining memory bandwidth, and temperature, and is stored in the shared memory buffer of the network security product in JSON format for real-time access by the resource dynamic allocation module.
[0026] The resource dynamic allocation module includes a data receiving unit, a fusion model calculation unit, and a task allocation and execution unit.
[0027] The data receiving unit reads the real-time load performance data report in JSON format generated by the load real-time monitoring module through the shared memory interface of the network security product, and parses out fields such as traffic rate, encryption task volume, decryption task volume, number of idle computing cores, remaining memory bandwidth, and temperature. The data reading frequency is consistent with the load monitoring frequency, and the latest network task load and M.2 device performance status data are obtained in real time to provide a real-time basis for subsequent task allocation.
[0028] The fusion model calculation unit processes the received real-time data based on a preset fusion model to generate a task allocation strategy. Specifically, firstly, it calculates the proportion of remaining device resources according to the task type and device performance parameters. The task types are divided into encryption, decryption, and traffic processing. The calculation formula is as follows:
[0029] ;
[0030] Among them, 0.6 and 0.4 are the weighting coefficients of core count and memory bandwidth, which can be adjusted through BIOS configuration; then, the matching relationship between tasks and devices is determined according to the strategy of prioritizing the allocation of encryption tasks to devices that support hardware encryption, allocating traffic processing tasks to devices with high resource scores, and evenly allocating decryption tasks to dual devices.
[0031] The task allocation execution unit allocates tasks to be processed in the task queue to the target M.2 device based on the calculation results of the fusion model. Specifically: in single-device mode PCIe x8, task instructions are sent only to the slots of the inserted device; in dual-device mode PCIe x4+x4, if the task can be split, the task quantity is split according to the device resource score ratio, and more tasks are allocated to the device with higher resource score; if the task cannot be split, it is allocated to the device with the highest current resource score; the task allocation sends a task descriptor containing task type, data address and priority information to the target device through the PCIe configuration space.
[0032] The cross-device task scheduling module includes a load threshold monitoring unit, a task migration execution unit, and a BIOS mode switching unit;
[0033] The load threshold monitoring unit is used to read the M.2 device performance data output by the resource dynamic allocation module in real time, and calculate the device load using the following formula:
[0034] Equipment utilization rate = 1 - (number of idle computing cores / total number of cores);
[0035] Among them, the number of idle computing cores is the number of cores with a utilization rate of ≤5% as statistically determined by the device performance acquisition unit, and the total number of cores is the nominal value of the device firmware; when the utilization rate of a certain device exceeds the 80% threshold for 5 consecutive samplings, the cross-device task migration process is triggered, where the interval between consecutive samplings is a preset fixed interval.
[0036] After receiving the trigger signal from the load threshold monitoring unit, the task migration execution unit first filters migrateable tasks and excludes tasks that need to be processed by fixed devices; then it pauses the source device task, stores context information such as task progress and data pointers in shared memory, and sends a task descriptor containing the task ID, source device slot, target device slot and recovery address to the target device through the PCIe configuration space. The target device reads the information from the shared memory and continues to execute the task.
[0037] Among them, portable tasks refer to tasks that do not rely on the unique hardware resources of the M.2 device, specifically including general traffic scrubbing tasks and batch data encryption tasks; general traffic scrubbing tasks are stateless processes such as filtering and rate limiting of network traffic, and can be executed on any M.2 device that supports the PCIe protocol; batch data encryption tasks are batch data processing based on standard encryption algorithms and do not rely on the device's built-in key storage module.
[0038] Fixed device processing tasks refer to tasks that must rely on the unique hardware resources of the M.2 device, specifically including decryption tasks based on the device's unique key and custom protocol acceleration tasks. Among them, decryption tasks based on the device's unique key require calling the exclusive key stored in the security chip on the M.2 device and can only be performed on that device. Custom protocol acceleration tasks are tasks that use the device's custom instruction set and hardware acceleration unit and cannot run on other models of devices.
[0039] The BIOS mode switching unit queries the M.2 device insertion status in real time through the system management interface. Specifically: when a device is inserted into only a single slot, the BIOS's PCIe_ConfigurateLinkWidth() function is called to configure the PCIe slot to X8 mode; when devices are inserted into both slots, the function is called and the X4+X4 parameter is passed in, triggering the BIOS to split the PCIe channel into two independent X4 paths; after the configuration is completed, the mode is made effective by restarting the PCIe link; if a device hot-plug occurs, the above process is repeated in real time to dynamically switch the channel mode.
[0040] The device compatibility adaptation module includes a firmware identification unit, a driver management unit, an instruction set adaptation unit, and a hot-plug detection unit.
[0041] The firmware identification unit reads the firmware information of the M.2 device through the PCIe configuration space; specifically: it reads the vendor ID and device ID from the device configuration register address; then it reads the firmware version string from the extended configuration space, parses the PCIe protocol version, interface mode, and M.2 specification supported by the device; finally, it matches the read ID and specification with the device information database built into the non-volatile memory of the network security product to determine the device manufacturer, model, and supported functions.
[0042] The driver management unit executes driver loading based on the device ID output by the firmware identification unit. Specifically, it first retrieves the driver file corresponding to the manufacturer ID and device ID from the driver library through the network security product operating system driver interface and loads it into the kernel and system service layer. Then, it passes initialization parameters to the driver based on the device specification information. Finally, before loading, it verifies the compatibility between the driver version and the device firmware version, and triggers an error message if they are incompatible.
[0043] The instruction set adaptation unit establishes a mapping between standardized task instructions and device-specific instruction sets based on the device firmware identification results. Specifically: first, a unified task instruction format is defined and converted into device native instructions by looking up a table; then, adaptation functions are written to address the differences in register layouts of different devices, and the devices automatically address through the adaptation functions; finally, a device instruction execution status table is maintained to ensure that the register status context is consistent when scheduling tasks across devices.
[0044] The hot-plug detection unit manages hot-plugging by monitoring the PCIe link status register. Specifically, it periodically queries the register, determining that a device is inserted when the link status changes from Down to Up, and removed when it changes from Up to Down. When insertion is detected, the firmware identification unit is triggered to add the device information to the system's recognizable list. When removal is detected, the device is deleted and its task is terminated. Then, a device status change notification is sent to the resource dynamic allocation module via the system message bus, triggering the task allocation strategy and mobilizing the fusion model calculation unit to recalculate the device resource score.
[0045] The power consumption feature identification unit obtains the device's basic power consumption parameters, including nominal maximum power, power consumption values corresponding to each performance state, and conversion delay, by reading the device management register of the PCIe configuration space; based on historical report data generated by the load real-time monitoring module, it establishes a correlation model between task type and power consumption features to form a power consumption feature library for different task types.
[0046] The dynamic power consumption scheduling unit, based on the task allocation results of the resource dynamic allocation module, processes historical load data using a Kalman filter algorithm, establishes a state-space model, and generates a smooth future load change trend prediction sequence through a two-stage prediction-update iteration. The results are then synchronized to the resource dynamic allocation module to adjust the task allocation strategy. According to the predicted load, the unit calls the device driver's ACPI interface to dynamically adjust the device's ACPI-defined performance and idle states: when resource utilization is <30%, it automatically switches to the lowest available performance state; when resource utilization is between 30% and 70%, it maintains the default state; when resource utilization is >70%, it briefly upgrades to the highest available performance state (highest performance available state requires device support), and the cooling fan speed is adjusted synchronously.
[0047] The adaptive heat dissipation control unit establishes a relationship model between device temperature and power consumption: Δtemperature = f(power consumption, heat dissipation efficiency), where heat dissipation efficiency is related to fan speed and ambient temperature; based on the real-time temperature and predicted load, and using the relationship model, the cooling fan speed is dynamically adjusted via a pulse width modulation interface, calculated as follows:
[0048] Target speed = base speed + coefficient k × (current temperature - reference temperature) + coefficient m × (predicted load - current load).
[0049] The coefficients k and m are configurable parameters, and the reference temperature is the optimal operating temperature of the device. When the temperature exceeds the preset threshold, task migration is triggered before the load threshold, and the device performance status is reduced. At the same time, the cross-device task scheduling module is notified to migrate some tasks.
[0050] A method for managing hardware acceleration devices used in cybersecurity products includes the following steps:
[0051] S1. Connect the PCIE X8 slot of the network security product to the M.2 device through the adapter board. Use internal wiring to split the X8 signal into two independent X4 signals. At the same time, use the clock buffer to copy the clock signal into two synchronous transmissions to the device, providing a hardware foundation for parallel communication between the two devices.
[0052] S2. Based on the interface connection, use the insulating bracket and screw holes to physically fix the bottom of the M.2 device to the adapter board PCB parallel to each other. After the side is embedded in the limiting groove for positioning, the top of the device is fixed to the bracket with countersunk screws to ensure that the device maintains a stable electrical connection with the adapter board in a compact space, providing physical support for subsequent data acquisition.
[0053] S3. Through the network protocol stack interface and device firmware management interface, extract network traffic rate, data encryption task volume, data decryption task volume, as well as the number of idle computing cores, remaining memory bandwidth and temperature data of the device in real time, and generate real-time load performance data reports in JSON format at a preset frequency to provide real-time parameters for task allocation strategy.
[0054] S4. Read real-time reports from shared memory, calculate device resource scores based on the fusion model, and assign tasks to the target device according to the strategy of prioritizing hardware-accelerated devices for encryption tasks, allocating traffic processing tasks to devices with high resource scores, and distributing decryption tasks evenly between the two devices.
[0055] S5. When the utilization rate of a single device exceeds the 80% threshold for 5 consecutive times, non-exclusive tasks are filtered, the source device tasks are paused and the context information is migrated to the idle device. At the same time, the BIOS function is called to switch the PCIe channel mode to achieve dynamic balancing of task processing efficiency.
[0056] S6. Read the vendor ID, device ID, and specification information of the device firmware through the PCIe configuration space, automatically load the corresponding driver, and establish a mapping table between standardized instructions and device native instructions based on the device firmware identification results. At the same time, monitor the PCIe link status in real time, update the identifiable list when the device is inserted, terminate the task and trigger resource reallocation when the device is removed. Finally, obtain the power consumption parameters by reading the device management register and establish a task power consumption model. Dynamically adjust the device performance status and cooling fan speed based on load prediction, adjust the fan speed through the temperature and power consumption relationship model, and trigger frequency reduction protection and task migration when the temperature exceeds the limit.
[0057] Compared with the prior art, the beneficial effects of the present invention are:
[0058] 1. Real-time load monitoring and dynamic scheduling improve processing efficiency: The real-time load monitoring module collects network traffic, task volume, and device performance data 10 times per second, generating real-time reports. The dynamic resource allocation module allocates tasks based on the proportion of remaining resources on devices and task type using a fusion model. For example, encryption tasks are preferentially allocated to devices that support hardware acceleration, and traffic tasks are allocated to devices with high resource scores. When the utilization rate of a single device exceeds 80% for five consecutive times, the cross-device task scheduling module migrates transferable tasks to idle devices and dynamically switches the PCIe channel mode to avoid single-point overload, improving overall task processing efficiency by more than 40%.
[0059] 2. Plug-and-play hardware, reducing development and replacement costs: Through the PCIe to M.2 adapter module, cybersecurity products can directly connect to standard M.2 hardware acceleration devices on the market without redesigning the circuitry. The adapter board splits the PCIe x8 signal into two x4 signals through internal PCB routing and achieves dual-device clock synchronization through a clock buffer. This eliminates the need for cybersecurity products to adapt to different device circuit interfaces, reducing hardware development cycle by approximately 60% and avoiding redesign costs due to device replacement, reducing adaptation costs by approximately 70%.
[0060] 3. Standardized compatibility and simplified device management: The device compatibility and adaptation module automatically loads the corresponding driver by reading the M.2 device's manufacturer ID, device ID, and firmware version, and establishes a mapping table between standardized instructions and the device's native instructions. At the same time, it monitors the hot-plugging status in real time, automatically updates the recognizable list when a device is inserted, and terminates the task and triggers resource reallocation when a device is removed, achieving plug-and-play functionality. It is compatible with various M.2 device specifications such as 2242 / 2260 / 2280, reducing manual adaptation operations by 90% and improving system compatibility and maintainability.
[0061] 4. Effective Linkage Between Heat Dissipation and Task Scheduling: In traditional cybersecurity products, heat dissipation control often only adjusts based on the current temperature, independent of task scheduling. When the device temperature is too high, it cannot alleviate the pressure on the device in time through methods such as task migration, leading to decreased device performance or even failure. This solution's adaptive heat dissipation control unit automatically triggers frequency reduction protection and notifies the cross-device task scheduling module to migrate some tasks when the temperature exceeds the threshold, achieving effective linkage between heat dissipation and task scheduling, ensuring stable operation of the device within a safe temperature range. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of the PCIE to M.2 hardware signal allocation and clock synchronization principle of a hardware acceleration device management system for network security products according to the present invention;
[0063] Figure 2 This is a schematic diagram of a PCIE to M.2 device implementation for a hardware acceleration device management system applied to network security products according to the present invention.
[0064] Figure 3 This is a PCB block diagram of a PCIE to M.2 adapter board for a hardware acceleration device management system applied to network security products according to the present invention;
[0065] Figure 4 This invention provides a system workflow diagram for a hardware acceleration device management system applied to cybersecurity products. Detailed Implementation
[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0067] Example: Figures 1-4 As shown, the present invention provides a technical solution.
[0068] A hardware acceleration device management system for cybersecurity products includes a PCIe to M.2 adapter module, a structural support module, a real-time load monitoring module, a dynamic resource allocation module, a cross-device task scheduling module, and a device compatibility adaptation module. The PCIe to M.2 adapter module is the hardware carrier, and its input end connects to the PCIe interface of the cybersecurity product. The device features an x8 slot with two M.2 connectors at the output. The structural support module connects the adapter board to the M.2 device, using physical brackets and screw holes to secure the device to the adapter board. The system is characterized by: a real-time load monitoring module that extracts traffic and encryption / decryption task volume in real time via a network task acquisition unit, and obtains the device's idle core count, remaining bandwidth, and temperature via a device performance acquisition unit, generating real-time JSON reports via a data report generation unit; a dynamic resource allocation module that calculates a resource score based on the formula (device resource score = (idle core count / total core count) × 0.6 + (remaining memory bandwidth / total memory bandwidth) × 0.4) via a fusion model calculation unit, prioritizing hardware encryption devices for encryption tasks, allocating traffic tasks to devices with higher scores, and evenly distributing decryption tasks across both devices, with the task allocation execution unit implementing the task allocation; and a cross-device task scheduling module that connects the allocation module to the M.2 device, calculating device utilization via a load threshold monitoring unit. When utilization exceeds 80% for five consecutive samples, the task migration execution unit filters migrateable tasks and migrates them to the target device. Simultaneously, the BIOS mode switching unit dynamically configures the PCIe lanes, maintaining PCIe in single-device mode. In X8 mode, the dual-device mode is configured as X4+X4 mode; the device compatibility adaptation module connects the M.2 device and the system software layer, reads the device ID and firmware information through the firmware identification unit, automatically loads the compatible driver by the driver management unit, updates the device list in real time and triggers resource reallocation through the hot-plug detection unit, obtains device power consumption parameters and establishes a task power consumption correlation model through the power consumption feature identification unit, dynamically adjusts the device performance status and cooling fan speed based on load prediction by the dynamic power consumption scheduling unit, and adjusts the fan speed through the temperature and power consumption relationship model and triggers frequency reduction protection and task migration when the temperature exceeds the limit.
[0069] The PCIe to M.2 adapter module includes an allocation unit and a clock processing unit;
[0070] The distribution unit splits the X8 signal input from the PCIE X8 slot of the network security product into two independent PCIE X4 signals through internal PCB traces. The two X4 signals correspond to the two M.2 connectors at the output end of the adapter board, respectively. When an M.2 device is inserted into a single slot, only one X4 signal transmits data. When M.2 devices are inserted into both slots, the two X4 signals transmit data in parallel. In this case, the PCIE slot needs to be configured to X4+X4 mode through the BIOS to support parallel communication between the two devices.
[0071] The clock processing unit integrates a clock buffer chip CLK BUFFER, which copies the clock signal CLK input from the PCIe slot into two synchronous clock signals. The two clock signals are connected to two M.2 connectors respectively, providing synchronous clock signals for the corresponding M.2 devices and ensuring the clock synchronization of the devices when operating in single-slot and dual-slot configurations.
[0072] The structural support module includes a physical bracket unit and a screw fixing unit;
[0073] The physical bracket unit is a rectangular frame structure made of insulating material. The frame size matches the edge of the PCB board of the PCIe to M.2 adapter board and is fixed to the edge of the adapter board by welding. The inner side of the frame is provided with a limiting groove corresponding to the size of the M.2 device. The length of the groove is compatible with standard M.2 devices, ensuring that after the device is inserted into the M.2 connector, the bottom surface is parallel to the PCB board of the adapter board and the side is embedded in the groove for positioning.
[0074] The screw fixing unit is responsible for fixing the M.2 device to the bracket. The top of the physical bracket has threaded holes corresponding to the screw holes of the M.2 device. After the M.2 device is inserted into the connector and embedded in the limiting groove, a countersunk screw is passed through the screw hole on the surface of the device and screwed into the threaded hole of the bracket to fix the top surface of the device to the bracket. The number of screws configured for each M.2 device corresponds to the number of fixing holes of the standard M.2 device, ensuring that the device maintains a stable electrical connection with the adapter plate in a vibration environment.
[0075] The real-time load monitoring module includes a network task acquisition unit, a device performance acquisition unit, and a data report generation unit.
[0076] The network task acquisition unit establishes a data channel with the task processing queue through the network protocol stack interface of the network security product, extracts the number of bytes of inbound / outbound traffic from the network card driver layer in real time, and calculates the average traffic rate at a preset fixed interval. The calculation formula is as follows:
[0077] ;
[0078] Where Δ bytes represent the difference in bytes between two adjacent samples, and Δ time represents the sampling interval in seconds;
[0079] While calculating the average traffic rate, the encryption / decryption task scheduling queue is monitored to count the number of new task entries per second, forming real-time data on the amount of data encryption tasks and data decryption tasks, which directly reflects the current network task load.
[0080] The device performance acquisition unit obtains device status data through the firmware management interface of the M.2 device. Specifically, it queries the CPU / MCU core status register to count the number of idle cores, obtaining the number of idle computing cores. The number of idle cores is defined as registers with a utilization rate of ≤5%. Then, it obtains the used memory bandwidth through the memory controller status register and, combined with the device's nominal total memory bandwidth, calculates the remaining bandwidth using the following formula:
[0081] Remaining memory bandwidth = Total memory bandwidth - Used memory bandwidth;
[0082] Finally, by reading the register value of the onboard temperature sensor, the real-time temperature data of the M.2 device is obtained. All parameters are read in real time through a standardized interface.
[0083] The data report generation unit timestamps the data output by the network task acquisition unit and the device performance acquisition unit, and generates a structured report using UTC time format. The report includes fields such as traffic rate, encryption task volume, decryption task volume, number of idle cores, remaining memory bandwidth, and temperature, and is stored in the shared memory buffer of the network security product in JSON format for real-time access by the resource dynamic allocation module.
[0084] The resource dynamic allocation module includes a data receiving unit, a fusion model calculation unit, and a task allocation and execution unit.
[0085] The data receiving unit reads the real-time load performance data report in JSON format generated by the load real-time monitoring module through the shared memory interface of the network security product, and parses out fields such as traffic rate, encryption task volume, decryption task volume, number of idle computing cores, remaining memory bandwidth, and temperature. The data reading frequency is consistent with the load monitoring frequency, and the latest network task load and M.2 device performance status data are obtained in real time to provide a real-time basis for subsequent task allocation.
[0086] The fusion model calculation unit processes the received real-time data based on a preset fusion model to generate a task allocation strategy. Specifically, firstly, it calculates the proportion of remaining device resources according to the task type and device performance parameters. The task types are divided into encryption, decryption, and traffic processing. The calculation formula is as follows:
[0087] ;
[0088] Among them, 0.6 and 0.4 are the weighting coefficients of core count and memory bandwidth, which can be adjusted through BIOS configuration; then, the matching relationship between tasks and devices is determined according to the strategy of prioritizing the allocation of encryption tasks to devices that support hardware encryption, allocating traffic processing tasks to devices with high resource scores, and evenly allocating decryption tasks to dual devices.
[0089] The task allocation execution unit allocates tasks to be processed in the task queue to the target M.2 device based on the calculation results of the fusion model. Specifically: in single-device mode PCIe x8, task instructions are sent only to the slots of the inserted device; in dual-device mode PCIe x4+x4, if the task can be split, the task quantity is split according to the device resource score ratio, and more tasks are allocated to the device with higher resource score; if the task cannot be split, it is allocated to the device with the highest current resource score; the task allocation sends a task descriptor containing task type, data address and priority information to the target device through the PCIe configuration space.
[0090] The cross-device task scheduling module includes a load threshold monitoring unit, a task migration execution unit, and a BIOS mode switching unit;
[0091] The load threshold monitoring unit is used to read the M.2 device performance data output by the resource dynamic allocation module in real time, and calculate the device load using the following formula:
[0092] Equipment utilization rate = 1 - (number of idle computing cores / total number of cores);
[0093] Among them, the number of idle computing cores is the number of cores with a utilization rate of ≤5% as statistically determined by the device performance acquisition unit, and the total number of cores is the nominal value of the device firmware; when the utilization rate of a certain device exceeds the 80% threshold for 5 consecutive samplings, the cross-device task migration process is triggered, where the interval between consecutive samplings is a preset fixed interval.
[0094] After receiving the trigger signal from the load threshold monitoring unit, the task migration execution unit first filters migrateable tasks and excludes tasks that need to be processed by fixed devices; then it pauses the source device task, stores context information such as task progress and data pointers in shared memory, and sends a task descriptor containing the task ID, source device slot, target device slot and recovery address to the target device through the PCIe configuration space. The target device reads the information from the shared memory and continues to execute the task.
[0095] Among them, portable tasks refer to tasks that do not rely on the unique hardware resources of the M.2 device, specifically including general traffic scrubbing tasks and batch data encryption tasks; general traffic scrubbing tasks are stateless processes such as filtering and rate limiting of network traffic, and can be executed on any M.2 device that supports the PCIe protocol; batch data encryption tasks are batch data processing based on standard encryption algorithms and do not rely on the device's built-in key storage module.
[0096] Fixed device processing tasks refer to tasks that must rely on the unique hardware resources of the M.2 device, specifically including decryption tasks based on the device's unique key and custom protocol acceleration tasks. Among them, decryption tasks based on the device's unique key require calling the exclusive key stored in the security chip on the M.2 device and can only be performed on that device. Custom protocol acceleration tasks are tasks that use the device's custom instruction set and hardware acceleration unit and cannot run on other models of devices.
[0097] The BIOS mode switching unit queries the M.2 device insertion status in real time through the system management interface. Specifically: when a device is inserted into only a single slot, the BIOS's PCIe_ConfigurateLinkWidth() function is called to configure the PCIe slot to X8 mode; when devices are inserted into both slots, the function is called and the X4+X4 parameter is passed in, triggering the BIOS to split the PCIe channel into two independent X4 paths; after the configuration is completed, the mode is made effective by restarting the PCIe link; if a device hot-plug occurs, the above process is repeated in real time to dynamically switch the channel mode.
[0098] The device compatibility adaptation module includes a firmware identification unit, a driver management unit, an instruction set adaptation unit, and a hot-plug detection unit.
[0099] The firmware identification unit reads the firmware information of the M.2 device through the PCIe configuration space; specifically: it reads the vendor ID and device ID from the device configuration register address; then it reads the firmware version string from the extended configuration space, parses the PCIe protocol version, interface mode, and M.2 specification supported by the device; finally, it matches the read ID and specification with the device information database built into the non-volatile memory of the network security product to determine the device manufacturer, model, and supported functions.
[0100] The driver management unit executes driver loading based on the device ID output by the firmware identification unit. Specifically, it first retrieves the driver file corresponding to the manufacturer ID and device ID from the driver library through the network security product operating system driver interface and loads it into the kernel and system service layer. Then, it passes initialization parameters to the driver based on the device specification information. Finally, before loading, it verifies the compatibility between the driver version and the device firmware version, and triggers an error message if they are incompatible.
[0101] The instruction set adaptation unit establishes a mapping between standardized task instructions and device-specific instruction sets based on the device firmware identification results. Specifically: first, a unified task instruction format is defined and converted into device native instructions by looking up a table; then, adaptation functions are written to address the differences in register layouts of different devices, and the devices automatically address through the adaptation functions; finally, a device instruction execution status table is maintained to ensure that the register status context is consistent when scheduling tasks across devices.
[0102] The hot-plug detection unit manages hot-plugging by monitoring the PCIe link status register. Specifically, it periodically queries the register, determining that a device is inserted when the link status changes from Down to Up, and removed when it changes from Up to Down. When insertion is detected, the firmware identification unit is triggered to add the device information to the system's recognizable list. When removal is detected, the device is deleted and its task is terminated. Then, a device status change notification is sent to the resource dynamic allocation module via the system message bus, triggering the task allocation strategy and mobilizing the fusion model calculation unit to recalculate the device resource score.
[0103] The power consumption feature identification unit obtains the device's basic power consumption parameters, including nominal maximum power, power consumption values corresponding to each performance state, and conversion delay, by reading the device management register of the PCIe configuration space; based on historical report data generated by the load real-time monitoring module, it establishes a correlation model between task type and power consumption features to form a power consumption feature library for different task types.
[0104] The dynamic power consumption scheduling unit, based on the task allocation results of the resource dynamic allocation module, processes historical load data using a Kalman filter algorithm, establishes a state-space model, and generates a smooth future load change trend prediction sequence through a two-stage prediction-update iteration. The results are then synchronized to the resource dynamic allocation module to adjust the task allocation strategy. According to the predicted load, the unit calls the device driver's ACPI interface to dynamically adjust the device's ACPI-defined performance and idle states: when resource utilization is <30%, it automatically switches to the lowest available performance state; when resource utilization is between 30% and 70%, it maintains the default state; when resource utilization is >70%, it briefly upgrades to the highest available performance state (highest performance available state requires device support), and the cooling fan speed is adjusted synchronously.
[0105] The adaptive heat dissipation control unit establishes a relationship model between device temperature and power consumption: Δtemperature = f(power consumption, heat dissipation efficiency), where heat dissipation efficiency is related to fan speed and ambient temperature; based on the real-time temperature and predicted load, and using the relationship model, the cooling fan speed is dynamically adjusted via a pulse width modulation interface, calculated as follows:
[0106] Target speed = base speed + coefficient k × (current temperature - reference temperature) + coefficient m × (predicted load - current load).
[0107] The coefficients k and m are configurable parameters, and the reference temperature is the optimal operating temperature of the device. When the temperature exceeds the preset threshold, task migration is triggered before the load threshold, and the device performance status is reduced. At the same time, the cross-device task scheduling module is notified to migrate some tasks.
[0108] A method for managing hardware acceleration devices used in cybersecurity products includes the following steps:
[0109] S1. Connect the PCIE X8 slot of the network security product to the M.2 device through the adapter board. Use internal wiring to split the X8 signal into two independent X4 signals. At the same time, use the clock buffer to copy the clock signal into two synchronous transmissions to the device, providing a hardware foundation for parallel communication between the two devices.
[0110] S2. Based on the interface connection, use the insulating bracket and screw holes to physically fix the bottom of the M.2 device to the adapter board PCB parallel to each other. After the side is embedded in the limiting groove for positioning, the top of the device is fixed to the bracket with countersunk screws to ensure that the device maintains a stable electrical connection with the adapter board in a compact space, providing physical support for subsequent data acquisition.
[0111] S3. Through the network protocol stack interface and device firmware management interface, extract network traffic rate, data encryption task volume, data decryption task volume, as well as the number of idle computing cores, remaining memory bandwidth and temperature data of the device in real time, and generate real-time load performance data reports in JSON format at a preset frequency to provide real-time parameters for task allocation strategy.
[0112] S4. Read real-time reports from shared memory, calculate device resource scores based on the fusion model, and assign tasks to the target device according to the strategy of prioritizing hardware-accelerated devices for encryption tasks, allocating traffic processing tasks to devices with high resource scores, and distributing decryption tasks evenly between the two devices.
[0113] S5. When the utilization rate of a single device exceeds the 80% threshold for 5 consecutive times, non-exclusive tasks are filtered, the source device tasks are paused and the context information is migrated to the idle device. At the same time, the BIOS function is called to switch the PCIe channel mode to achieve dynamic balancing of task processing efficiency.
[0114] S6. Read the vendor ID, device ID, and specification information of the device firmware through the PCIe configuration space, automatically load the corresponding driver, and establish a mapping table between standardized instructions and device native instructions based on the device firmware identification results. At the same time, monitor the PCIe link status in real time, update the identifiable list when the device is inserted, terminate the task and trigger resource reallocation when the device is removed. Finally, obtain the power consumption parameters by reading the device management register and establish a task power consumption model. Dynamically adjust the device performance status and cooling fan speed based on load prediction, adjust the fan speed through the temperature and power consumption relationship model, and trigger frequency reduction protection and task migration when the temperature exceeds the limit.
[0115] A cybersecurity product integrates two M.2 hardware acceleration devices. Device A is a 2280-specification hardware encryption card supporting AES-NI acceleration, while device B is a 2260-specification traffic scrubbing card based on FPGA logic. A PCIe 3.0 to M.2 adapter board connects to the cybersecurity product's PCIe X8 slot. The PCB uses differential routing technology to split the X8 signal into two independent X4 signals, which are mapped to the J1 and J2 M.2 connectors on the adapter board, respectively. The clock processing unit integrates a TI CDCLVC1104 clock buffer, which replicates the input 100MHz clock signal into two synchronous clocks, transmitting them to the devices via a 50Ω impedance-matched line to ensure clock skew ≤50ps.
[0116] The insulating bracket uses a rectangular frame made of FR-4 material, which is welded to the edge of the adapter board by 4 solder pads. The inner side of the frame is machined with stepped limiting grooves compatible with 2280 / 2260 specifications. After device A is inserted into the J1 connector, the bottom surface is kept 1.5mm away from the PCB, the side is embedded with a 5mm deep groove for positioning, and the top surface is fixed with 2 M2.5 countersunk screws. After device B is inserted into the J2 connector, it is fixed with 1 M2 screw to ensure that the contact resistance fluctuation of the two devices is ≤5mΩ during vibration testing.
[0117] When the cybersecurity product processes internet outbound traffic, the real-time monitoring data is as follows (sampling interval 0.1 seconds): inbound traffic rate 1500Mbps (Δ bytes = 18,750,000B, Δ time = 0.1s), encryption tasks 800 times / second, decryption tasks 200 times / second. Device A has 8 cores, a total memory bandwidth of 20GB / s, 3 idle cores, 14GB / s used bandwidth, and a temperature of 55℃; Device B has 8 cores, a total memory bandwidth of 16GB / s, 6 idle cores, 8GB / s used bandwidth, and a temperature of 40℃. The real-time load monitoring module generates JSON reports at a frequency of 10 times / second and stores them in shared memory.
[0118] The resource dynamic allocation module calculates the resource score of device A as (3 / 8×0.6) + (6 / 20×0.4) = 0.345, and that of device B as (6 / 8×0.6) + (8 / 16×0.4) = 0.65. 800 encryption tasks are allocated to device A, 1500Mbps of traffic is split according to the score ratio, device B is allocated 975Mbps, and 200 decryption tasks are fixedly allocated to device A.
[0119] Device A's task utilization rate reached 87.5% for five consecutive samplings (within 0.5 seconds) due to a sudden task, triggering the task migration process. It selected 500 batch encryption tasks, paused, stored the context in shared memory, generated a migration descriptor and sent it to device B through the PCIe configuration space. At the same time, it called the BIOS's PCIe_ConfigurateLinkWidth(X4+X4) function to switch to dual X4 mode.
[0120] At this point, device C (a 2242-specification general-purpose accelerator card, Vendor ID=0x1234, Device ID=0x5678, firmware V1.2.0) is hot-swapped. The system reads the ID through the PCIe configuration space, matches it with the database, loads the driver_1234_5678_v1.3.0.ko driver, verifies version compatibility, establishes a mapping table between standardized instructions and device C's native instructions, writes the register adaptation function SetRegAddr(0x2000,data), and notifies the resource dynamic allocation module via D-Bus. After recalculating device C's resource score to 0.7, the 200Mbps traffic task is allocated to device C, achieving multi-device load balancing.
[0121] Meanwhile, the device's nominal maximum power (e.g., 25W), power consumption values for each performance state, and conversion delay (5ms) are obtained by reading the device management register of the PCIe configuration space. Based on historical load data, a correlation model between task type and power consumption characteristics is established (e.g., the power consumption of encryption tasks is 30% higher than that of traffic tasks). The Kalman filter algorithm (process noise covariance q=0.008, measurement noise covariance r=0.06) is used to predict the load change trend. When the resource utilization rate is <30%, the low power consumption mode is switched, and when it is >70%, it is temporarily upgraded to the high performance mode. The cooling fan speed is dynamically adjusted according to the formula: target speed = base speed (2000RPM) + k (60) × (current temperature - 55℃) + m (30) × (predicted load - current load). Then, a relationship model between temperature and power consumption is established: Δ temperature = 0.5 × power consumption + 0.3 × (1 / fan speed) + 0.2 × ambient temperature. When the temperature exceeds 75℃, the frequency reduction protection is triggered and the task is migrated to an idle device, realizing dynamic optimization of power consumption and coordinated control of heat dissipation of M.2 devices in the network security scenario.
[0122] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A hardware acceleration device management system for cybersecurity products, comprising a PCIE to M.2 adapter module, a structural support module, a real-time load monitoring module, a dynamic resource allocation module, a cross-device task scheduling module, and a device compatibility adaptation module; the PCIE to M.2 adapter module is a hardware carrier, with its input end connected to the PCIE x8 slot of the cybersecurity product and its output end equipped with two M.2 connectors; the structural support module connects the adapter module and the M.2 device, using physical brackets and screw holes to fix the device and the adapter module; characterized in that: The real-time load monitoring module extracts traffic and encryption / decryption task volume in real time through the network task acquisition unit, and obtains the number of idle cores, remaining bandwidth, and temperature of the device through the device performance acquisition unit. A real-time JSON report is generated by the data report generation unit. The dynamic resource allocation module calculates a score based on the formula: Device Resource Score = (Number of Idle Cores / Total Number of Cores) × 0.6 + (Remaining Memory Bandwidth / Total Memory Bandwidth) × 0.4, and allocates tasks according to a strategy of prioritizing hardware encryption devices for encryption tasks, allocating traffic tasks to devices with higher scores, and evenly distributing decryption tasks across both devices. The task allocation execution unit implements task allocation. The cross-device task scheduling module connects the allocation module and the M.2 device. It calculates device utilization through the load threshold monitoring unit. When the utilization exceeds 80% for five consecutive samples, the task migration execution unit filters migrateable tasks and migrates them to the target device. Simultaneously, the BIOS mode switching unit dynamically configures the PCIe channel, maintaining PCIe in single-device mode. In X8 mode, the dual-device mode is configured as X4+X4 mode; the device compatibility adaptation module connects the M.2 device and the system software layer, reads the device ID and firmware information through the firmware identification unit, automatically loads the compatible driver by the driver management unit, updates the device list in real time and triggers resource reallocation through the hot-plug detection unit, obtains device power consumption parameters and establishes a task power consumption correlation model through the power consumption feature identification unit, dynamically adjusts the device performance status and cooling fan speed based on load prediction by the dynamic power consumption scheduling unit, and adjusts the fan speed through the temperature and power consumption relationship model and triggers frequency reduction protection and task migration when the temperature exceeds the limit.
2. The hardware acceleration device management system for cybersecurity products according to claim 1, characterized in that: The PCIe to M.2 adapter module includes an allocation unit and a clock processing unit; The distribution unit splits the X8 signal input from the PCIE X8 slot of the network security product into two independent PCIE X4 signals through the internal traces of the PCB. The two X4 signals correspond to the two M.2 connectors at the output end of the adapter board, respectively. When an M.2 device is inserted into a single slot, only one X4 signal transmits data; when both slots are filled with M.2 devices, two X4 signals transmit data in parallel. In this case, the PCIe slot needs to be configured to X4+X4 mode through the BIOS to support parallel communication between the two devices. The clock processing unit integrates a clock buffer chip CLK BUFFER, which copies the clock signal CLK input from the PCIe slot into two synchronous clock signals. The two clock signals are connected to two M.2 connectors respectively, providing synchronous clock signals for the corresponding M.2 devices and ensuring the clock synchronization of the devices when operating in single-slot and dual-slot configurations.
3. The hardware acceleration device management system for cybersecurity products according to claim 1, characterized in that: The structural support module includes a physical bracket unit and a screw fixing unit; The physical bracket unit is a rectangular frame structure made of insulating material. The frame size matches the edge of the PCB board of the PCIe to M.2 adapter board and is fixed to the edge of the adapter board by welding. The inner side of the frame is provided with a limiting groove corresponding to the size of the M.2 device. The length of the groove is compatible with standard M.2 devices, ensuring that after the device is inserted into the M.2 connector, the bottom surface is parallel to the PCB board of the adapter board and the side is embedded in the groove for positioning. The screw fixing unit is responsible for fixing the M.2 device to the bracket. The top of the physical bracket has threaded holes corresponding to the screw holes of the M.2 device. After the M.2 device is inserted into the connector and embedded in the limiting groove, a countersunk screw is passed through the screw hole on the surface of the device and screwed into the threaded hole of the bracket to fix the top surface of the device to the bracket. The number of screws configured for each M.2 device corresponds to the number of fixing holes of the standard M.2 device, ensuring that the device maintains a stable electrical connection with the adapter plate in a vibration environment.
4. The hardware acceleration device management system for cybersecurity products according to claim 1, characterized in that: The real-time load monitoring module includes a network task acquisition unit, a device performance acquisition unit, and a data report generation unit. The network task acquisition unit establishes a data channel with the task processing queue through the network protocol stack interface of the network security product, extracts the number of bytes of inbound / outbound traffic from the network card driver layer in real time, and calculates the average traffic rate at a preset fixed interval. The calculation formula is as follows: ; Where Δ bytes represent the difference in bytes between two adjacent samples, and Δ time represents the sampling interval in seconds; While calculating the average traffic rate, the encryption / decryption task scheduling queue is monitored to count the number of new task entries per second, forming real-time data on the amount of data encryption tasks and data decryption tasks, which directly reflects the current network task load. The device performance acquisition unit obtains device status data through the firmware management interface of the M.2 device. Specifically, it queries the CPU / MCU core status register to count the number of idle cores, obtaining the number of idle computing cores. The number of idle cores is defined as registers with a utilization rate of ≤5%. Then, it obtains the used memory bandwidth through the memory controller status register and, combined with the device's nominal total memory bandwidth, calculates the remaining bandwidth using the following formula: Remaining memory bandwidth = Total memory bandwidth - Used memory bandwidth; Finally, by reading the register value of the onboard temperature sensor, the real-time temperature data of the M.2 device is obtained. All parameters are read in real time through a standardized interface. The data report generation unit timestamps the data output by the network task acquisition unit and the device performance acquisition unit, and generates a structured report using UTC time format. The report includes fields such as traffic rate, encryption task volume, decryption task volume, number of idle cores, remaining memory bandwidth, and temperature. It is stored in JSON format in the shared memory buffer of the cybersecurity product for real-time access by the resource dynamic allocation module.
5. A hardware acceleration device management system for cybersecurity products according to claim 4, characterized in that: The resource dynamic allocation module includes a data receiving unit, a fusion model calculation unit, and a task allocation and execution unit. The data receiving unit reads the real-time load performance data report in JSON format generated by the load real-time monitoring module through the shared memory interface of the network security product, and parses out fields such as traffic rate, encryption task volume, decryption task volume, number of idle computing cores, remaining memory bandwidth, and temperature. The data reading frequency is consistent with the load monitoring frequency, and the latest network task load and M.2 device performance status data are obtained in real time to provide a real-time basis for subsequent task allocation. The fusion model calculation unit processes the received real-time data based on a preset fusion model to generate a task allocation strategy. Specifically, firstly, it calculates the proportion of remaining device resources according to the task type and device performance parameters. The task types are divided into encryption, decryption, and traffic processing. The calculation formula is as follows: ; Among them, 0.6 and 0.4 are the weighting coefficients of core count and memory bandwidth, which can be adjusted through BIOS configuration; then, the matching relationship between tasks and devices is determined according to the strategy of prioritizing the allocation of encryption tasks to devices that support hardware encryption, allocating traffic processing tasks to devices with high resource scores, and evenly allocating decryption tasks to dual devices. The task allocation and execution unit allocates the tasks to be processed in the task queue to the target M.2 device based on the calculation results of the fusion model. Specifically: in single-device mode PCIE X8, task instructions are sent only to the slots where the device is inserted; in dual-device mode PCIE X4+X4, if the task can be split, the task quantity is split according to the device resource score ratio, and the device with the higher resource score is allocated more tasks. If the task cannot be split, it is assigned to the device with the highest current resource score; the task assignment sends a task descriptor containing task type, data address and priority information to the target device through the PCIe configuration space.
6. The hardware acceleration device management system for cybersecurity products according to claim 1, characterized in that: The cross-device task scheduling module includes a load threshold monitoring unit, a task migration execution unit, and a BIOS mode switching unit; The load threshold monitoring unit is used to read the M.2 device performance data output by the resource dynamic allocation module in real time, and calculate the device load using the following formula: Equipment utilization rate = 1 - (number of idle computing cores / total number of cores); Among them, the number of idle computing cores is the number of cores with a utilization rate of ≤5% as statistically determined by the device performance acquisition unit, and the total number of cores is the nominal value of the device firmware; when the utilization rate of a certain device exceeds the 80% threshold for 5 consecutive samplings, the cross-device task migration process is triggered, where the interval between consecutive samplings is a preset fixed interval. After receiving the trigger signal from the load threshold monitoring unit, the task migration execution unit first filters migrateable tasks and excludes tasks that need to be processed by fixed devices; then it pauses the source device task, stores context information such as task progress and data pointers in shared memory, and sends a task descriptor containing the task ID, source device slot, target device slot and recovery address to the target device through the PCIe configuration space. The target device reads the information from the shared memory and continues to execute the task. Among them, portable tasks refer to tasks that do not rely on the unique hardware resources of the M.2 device, specifically including general traffic scrubbing tasks and batch data encryption tasks; general traffic scrubbing tasks are stateless processes such as filtering and rate limiting of network traffic, and can be executed on any M.2 device that supports the PCIe protocol; batch data encryption tasks are batch data processing based on standard encryption algorithms and do not rely on the device's built-in key storage module. Fixed device processing tasks refer to tasks that must rely on the unique hardware resources of the M.2 device, specifically including decryption tasks based on the device's unique key and custom protocol acceleration tasks. Among them, decryption tasks based on the device's unique key require calling the exclusive key stored in the security chip on the M.2 device and can only be performed on that device. Custom protocol acceleration tasks are tasks that use the device's custom instruction set and hardware acceleration unit and cannot run on other models of devices. The BIOS mode switching unit queries the M.2 device insertion status in real time through the system management interface. Specifically: when a device is inserted into only a single slot, the BIOS's PCIe_ConfigurateLinkWidth() function is called to configure the PCIe slot to X8 mode; when devices are inserted into both slots, the function is called and the X4+X4 parameter is passed in, triggering the BIOS to split the PCIe channel into two independent X4 paths; after the configuration is completed, the mode is made effective by restarting the PCIe link; if a device hot-plug occurs, the above process is repeated in real time to dynamically switch the channel mode.
7. A hardware acceleration device management system for cybersecurity products according to claim 1, characterized in that: The device compatibility adaptation module includes a firmware identification unit, a driver management unit, an instruction set adaptation unit, a hot-plug detection unit, a power consumption characteristic identification unit, a dynamic power consumption scheduling unit, and an adaptive heat dissipation control unit. The firmware identification unit reads the firmware information of the M.2 device through the PCIe configuration space; specifically: it reads the vendor ID and device ID from the device configuration register address; then it reads the firmware version string from the extended configuration space, parses the PCIe protocol version, interface mode, and M.2 specification supported by the device; finally, it matches the read ID and specification with the device information database built into the non-volatile memory of the network security product to determine the device manufacturer, model, and supported functions. The driver management unit executes driver loading based on the device ID output by the firmware identification unit. Specifically, it first retrieves the driver file corresponding to the manufacturer ID and device ID from the driver library through the network security product operating system driver interface and loads it into the kernel and system service layer. Then, it passes initialization parameters to the driver based on the device specification information. Finally, before loading, it verifies the compatibility between the driver version and the device firmware version, and triggers an error message if they are incompatible. The instruction set adaptation unit establishes a mapping between standardized task instructions and device-specific instruction sets based on the device firmware identification results. Specifically: first, a unified task instruction format is defined and converted into device native instructions by looking up a table; then, adaptation functions are written to address the differences in register layouts of different devices, and the devices automatically address through the adaptation functions; finally, a device instruction execution status table is maintained to ensure that the register status context is consistent when scheduling tasks across devices. The hot-plug detection unit implements hot-plug management by monitoring the PCIe link status register. Specifically, it periodically queries the register, and determines that the device is inserted when the link status changes from Down to Up, and that the device is removed when it changes from Up to Down. When insertion is detected, the firmware identification unit is triggered to start, and the device information is added to the system's recognizable list. When removal is detected, the device is deleted and its task is terminated. Then, a device status change notification is sent to the resource dynamic allocation module via the system message bus, triggering the task allocation strategy and mobilizing the fusion model calculation unit to recalculate the device resource score; The power consumption feature identification unit obtains the device's basic power consumption parameters, including nominal maximum power, power consumption values corresponding to each performance state, and conversion delay, by reading the device management register of the PCIe configuration space; based on historical report data generated by the load real-time monitoring module, it establishes a correlation model between task type and power consumption features to form a power consumption feature library for different task types. The dynamic power consumption scheduling unit, based on the task allocation results of the resource dynamic allocation module, uses the Kalman filter algorithm to process historical load data, establishes a state space model, and generates a smooth future load change trend prediction sequence through a two-stage prediction-update iteration. The results are then synchronized to the resource dynamic allocation module to adjust the task allocation strategy. Based on the predicted load, the ACPI interface of the device driver is invoked to dynamically adjust the performance and idle states defined by the device's ACPI: when the resource utilization is <30%, it automatically switches to the lowest available performance state; when the resource utilization is between 30% and 70%, it maintains the default state; when the resource utilization is >70%, it briefly increases to the highest available performance state. The highest performance available state requires device support, and the cooling fan speed is adjusted synchronously. The adaptive heat dissipation control unit establishes a relationship model between device temperature and power consumption: Δtemperature = f(power consumption, heat dissipation efficiency), where heat dissipation efficiency is related to fan speed and ambient temperature; based on the real-time temperature and predicted load, and using the relationship model, the cooling fan speed is dynamically adjusted via a pulse width modulation interface, calculated as follows: Target speed = base speed + coefficient k × (current temperature - reference temperature) + coefficient m × (predicted load - current load). The coefficients k and m are configurable parameters, and the reference temperature is the optimal operating temperature of the device. When the temperature exceeds the preset threshold, task migration is triggered before the load threshold, and the device performance status is reduced. At the same time, the cross-device task scheduling module is notified to migrate some tasks.
8. A method for managing hardware acceleration devices in cybersecurity products, applied to the hardware acceleration device management system for cybersecurity products as described in any one of claims 1-7, characterized in that: Includes the following steps: S1. Connect the PCIE X8 slot of the network security product to the M.2 device through the adapter board. Use internal wiring to split the X8 signal into two independent X4 signals. At the same time, use the clock buffer to copy the clock signal into two synchronous transmissions to the device, providing a hardware foundation for parallel communication between the two devices. S2. Based on the interface connection, use the insulating bracket and screw holes to physically fix the bottom of the M.2 device to the adapter board PCB parallel to each other. After the side is embedded in the limiting groove for positioning, the top of the device is fixed to the bracket with countersunk screws to ensure that the device maintains a stable electrical connection with the adapter board in a compact space, providing physical support for subsequent data acquisition. S3. Through the network protocol stack interface and device firmware management interface, extract network traffic rate, data encryption task volume, data decryption task volume, as well as the number of idle computing cores, remaining memory bandwidth and temperature data of the device in real time, and generate real-time load performance data reports in JSON format at a preset frequency to provide real-time parameters for task allocation strategy. S4. Read real-time reports from shared memory, calculate device resource scores based on the fusion model, and assign tasks to the target device according to the strategy of prioritizing hardware-accelerated devices for encryption tasks, allocating traffic processing tasks to devices with high resource scores, and distributing decryption tasks evenly between the two devices. S5. When the utilization rate of a single device exceeds the 80% threshold for 5 consecutive times, non-exclusive tasks are filtered, the source device tasks are paused and the context information is migrated to the idle device. At the same time, the BIOS function is called to switch the PCIe channel mode to achieve dynamic balancing of task processing efficiency. S6. Read the vendor ID, device ID, and specification information of the device firmware through the PCIe configuration space, automatically load the corresponding driver, and establish a mapping table between standardized instructions and device native instructions based on the device firmware identification results. At the same time, monitor the PCIe link status in real time, update the identifiable list when the device is inserted, terminate the task and trigger resource reallocation when the device is removed. Finally, obtain the power consumption parameters by reading the device management register and establish a task power consumption model. Dynamically adjust the device performance status and cooling fan speed based on load prediction, adjust the fan speed through the temperature and power consumption relationship model, and trigger frequency reduction protection and task migration when the temperature exceeds the limit.
Citation Information
Patent Citations
Interface conversion device based on monitoring network security equipment
CN220475065U
Distributed processing in a cryptography acceleration chip
WO2001005086A2