Intelligent hardware cooperative integration method and system based on multi-modal protocol adaptation
Patent Information
- Application Number
- CN202611072709.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-20
AI Technical Summary
[0008]基于本申请提供的实施例,本发明通过在边缘硬件设备上构建沙箱运行环境,并将脚本执行空间与系统关键资源隔离,使得后续加载并执行的脚本程序包无法直接访问或破坏边缘硬件设备的系统关键资源,从而在保障边缘设备安全的前提下,支持动态、灵活地执行来自云端的业务逻辑。本发明当待接入硬件设备与边缘硬件设备建立通信连接时,自动采集其通信数据流特征信息,并基于预训练的协议识别模型判别通信协议类型,无需人工预先配置或干预,即可实现对未知协议类型设备的实时识别。本发明根据识别出的通信协议类型,从协议脚本库中动态加载对应的协议解析脚本,使得边缘硬件设备无需为每种协议预先烧录固件或开发专用驱动程序,即可支持多种通信协议的解析与适配。
Smart Images

Figure CN122578743B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of Internet of Things and edge computing technology, and more specifically, to a method and system for collaborative integration of smart hardware based on multimodal protocol adaptation. Background Technology
[0002] In IoT scenarios, hardware devices from different manufacturers use different communication protocols (Modbus, MQTT, HTTP, proprietary protocols, etc.). These devices are widely deployed in various application areas such as smart homes, industrial automation, smart cities, and agricultural monitoring. Edge hardware devices are typically deployed at the network edge, close to the data source, to connect to various front-end sensing or execution devices and process the collected data locally. As the scale of IoT continues to expand, an edge hardware device may need to interact with multiple hardware devices using different communication protocols, either sequentially or simultaneously. In actual deployment environments, the type of communication protocol often depends on the device manufacturer's choice and the specific industry standards. Meanwhile, cloud-based collaboration centers or management platforms typically need to distribute corresponding control commands, data processing rules, or business strategies to edge hardware devices for execution based on business application layer logic, in order to achieve a collaborative working mechanism between cloud training and edge inference, and between cloud decision-making and edge response.
[0003] In existing technologies, dedicated drivers need to be developed for each device, resulting in high protocol adaptation costs. More importantly, when new features need to be added to already deployed hardware devices, firmware upgrades are usually required, which presents challenges such as upgrade difficulties and complex rollbacks in large-scale deployment scenarios.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a method and system for intelligent hardware collaborative integration based on multimodal protocol adaptation to solve the above-mentioned technical problems.
[0006] This application provides a method for intelligent hardware collaborative integration based on multimodal protocol adaptation, comprising: constructing a sandbox runtime environment on an edge hardware device, isolating the script execution space of the sandbox runtime environment from the critical system resources of the edge hardware device; when a hardware device to be connected establishes a communication connection with the edge hardware device, collecting the communication data stream feature information of the hardware device to be connected, and determining the communication protocol type of the hardware device to be connected based on a pre-trained protocol recognition model; dynamically loading a protocol parsing script corresponding to the communication protocol type from a protocol script library; the cloud collaboration center encapsulating the business control logic into a script package and sending it to the edge hardware device via an encrypted transmission channel; the edge hardware device receiving and verifying the script package, and loading the verified script package into the sandbox runtime environment for execution; when the script package is executed in the sandbox runtime environment, it performs cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type.
[0007] This application provides a smart hardware collaborative integration system based on multimodal protocol adaptation, comprising: an environment construction unit, used to build a sandbox running environment on the edge hardware device, and isolate the script execution space of the sandbox running environment from the system critical resources of the edge hardware device; a discrimination unit, used to collect the communication data stream feature information of the hardware device to be connected when establishing a communication connection with the edge hardware device, and to determine the communication protocol type of the hardware device to be connected based on a pre-trained protocol recognition model; a loading unit, used to dynamically load the protocol parsing script corresponding to the communication protocol type from the protocol script library; a transmission unit, used by the cloud collaboration center to encapsulate the business control logic into a script program package and send it to the edge hardware device through an encrypted transmission channel; an execution unit, used by the edge hardware device to receive and verify the script program package, and load the verified script program package into the sandbox running environment for execution; and an interaction unit, used by the script program package to perform cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type when it is executed in the sandbox running environment.
[0008] Based on the embodiments provided in this application, this invention constructs a sandbox runtime environment on edge hardware devices and isolates the script execution space from critical system resources. This prevents subsequently loaded and executed script packages from directly accessing or damaging the critical system resources of the edge hardware devices, thereby supporting the dynamic and flexible execution of business logic from the cloud while ensuring the security of the edge devices. When the hardware device to be accessed establishes a communication connection with the edge hardware device, this invention automatically collects its communication data stream characteristic information and determines the communication protocol type based on a pre-trained protocol recognition model. This achieves real-time identification of devices with unknown protocol types without manual pre-configuration or intervention. Based on the identified communication protocol type, this invention dynamically loads the corresponding protocol parsing script from the protocol script library, enabling edge hardware devices to support the parsing and adaptation of multiple communication protocols without pre-burning firmware or developing dedicated drivers for each protocol.
[0009] This invention encapsulates business control logic into script packages through a cloud-based collaborative center and distributes them to edge hardware devices via an encrypted transmission channel. The edge hardware devices receive and verify the scripts, then load them into a sandbox environment for execution. This eliminates the need to upgrade the edge device firmware for business logic updates; only a new script package needs to be distributed from the cloud. When the script package executes within the sandbox environment, it calls a pre-loaded protocol parsing script corresponding to the communication protocol type to perform cross-protocol data interaction with the hardware device to be connected. This allows the business logic within the sandbox to leverage dynamically adapted protocol parsing capabilities to achieve unified control and data exchange for devices with different protocol types, without requiring the repeated development of protocol adaptation code for each business function. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0011] Figure 1 This is a flowchart of an optional smart hardware collaborative integration method based on multimodal protocol adaptation according to an embodiment of this application;
[0012] Figure 2 This is a flowchart illustrating an optional communication protocol identification and protocol parsing script dynamic loading method according to an embodiment of this application.
[0013] Figure 3 This is a flowchart illustrating an optional sandbox runtime environment security protection and script update according to an embodiment of this application;
[0014] Figure 4 This is a flowchart illustrating an optional cloud-based script package encapsulation, distribution, and edge device verification execution according to an embodiment of this application.
[0015] Figure 5 This is a structural diagram of an optional intelligent hardware collaborative integration system based on a multimodal protocol adaptive according to an embodiment of this application.
[0016] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] According to one aspect of the embodiments of this application, such as Figure 1 As shown, this application provides a smart hardware collaborative integration method based on multimodal protocol adaptation, including:
[0019] S1, build a sandbox runtime environment on edge hardware devices to isolate the script execution space of the sandbox runtime environment from the critical system resources of the edge hardware devices;
[0020] Critical system resources include memory resources, file system resources, and network resources;
[0021] In one embodiment, a lightweight sandbox runtime environment is built on the edge hardware device. The sandbox runtime environment isolates the script execution space from the critical system resources of the edge hardware device, preventing subsequently loaded script packages from directly accessing or damaging the device's operating system kernel, file system root directory, network protocol stack, hardware drivers, and other core functional modules.
[0022] As a preferred implementation, the sandbox operating environment employs a multi-layered security protection architecture, including:
[0023] Memory isolation layer: Based on WebAssembly runtime or similar memory safety strategies, it restricts the memory access boundaries of scripts through a linear memory model to prevent out-of-bounds access or dangling pointer references.
[0024] File system isolation layer: Utilizes kernel-level security modules (such as Linux's Seccomp, AppArmor, etc.) to restrict the range of directories and files that scripts can access, allowing only reading and writing in specified temporary directories.
[0025] Network isolation layer: Employing methods such as the Extended Berkeley Packet Filter (eBPF), network connection requests initiated by scripts are filtered in kernel space, restricting the target addresses and ports that can be connected.
[0026] In addition, the sandbox runtime environment is also equipped with a static security analysis module, which performs lexical, syntactic and control flow analysis on the script code before the script is loaded into the sandbox to detect whether there are dangerous operations (such as calling system exit functions, infinite loops, resource leaks, etc.) and only allows scripts that pass the security verification to enter the sandbox for execution.
[0027] S2, When the hardware device to be accessed establishes a communication connection with the edge hardware device, the communication data stream feature information of the hardware device to be accessed is collected, and the communication protocol type of the hardware device to be accessed is determined based on the pre-trained protocol recognition model;
[0028] When a hardware device to be connected establishes a communication connection with an edge hardware device (e.g., via physical or wireless means such as serial port, Ethernet, Wi-Fi, Bluetooth, etc.), the edge hardware device automatically initiates the protocol identification process.
[0029] First, collect the communication data stream characteristic information of the hardware device to be connected. The communication data stream characteristic information includes, but is not limited to, the content of the handshake packet, the structure of the data packet header (such as the start byte and the position of the length field), the data packet length distribution statistics, and the time interval pattern of data packet transmission (time-sequential transmission mode), etc.
[0030] Then, the collected communication feature information is input into a pre-trained protocol recognition model, which determines the communication protocol type of the hardware device to be accessed. The protocol recognition model is built based on a deep learning network, preferably a hybrid model combining a convolutional neural network (CNN) and an attention mechanism. The CNN is used to extract local features within a single data packet (such as patterns in specific byte sequences), while the attention mechanism is used to capture temporal dependencies between different data packets (such as patterns in request-response sequences). This model is pre-trained using a large number of communication data samples labeled with protocol types (Modbus, MQTT, HTTP, various proprietary protocols, etc.), enabling it to automatically classify the protocol type of unknown devices. For proprietary protocols not present in the training set, the model can infer the most similar known protocol category or label it as a novel proprietary protocol through feature matching and nearest neighbor analysis.
[0031] S3, dynamically loads the protocol parsing script corresponding to the communication protocol type from the protocol script library;
[0032] After identifying the communication protocol type, the edge hardware device retrieves and dynamically loads the protocol parsing script corresponding to that protocol type from the protocol script library. The protocol script library is stored in the edge hardware device's local non-volatile memory or can be updated synchronously from the cloud via a network.
[0033] If a parsing script for a known protocol does not exist in the protocol script library (e.g., when encountering a new proprietary protocol for the first time), this embodiment further provides an automatic generation mechanism based on a protocol description language engine. The protocol description language is a domain-specific language (DSL) used to define the protocol's data packet format (field offset, length, byte order), field encoding / decoding rules (such as CRC checksum, escape rules), and state machine transition logic. Based on the communication feature information collected in step two, a parsing script conforming to the protocol description language specification is automatically generated using pattern matching and rule inference methods. After verification, this script is stored in the protocol script library for direct use by subsequent similar devices.
[0034] After loading the protocol parsing script, the edge hardware device registers it with the protocol parsing engine. This engine is responsible for parsing all communication data from the hardware device to be connected and converting the raw data of heterogeneous protocol formats into a unified standard data format (such as JSON or Protobuf format) to facilitate unified processing by upper-layer business logic.
[0035] S4, the cloud-based collaboration center encapsulates business control logic into script packages and sends them to edge hardware devices via an encrypted transmission channel;
[0036] Based on business needs (such as the need to add data preprocessing rules, edge AI inference logic, device linkage control strategies, etc.), the cloud-based collaboration center encapsulates the corresponding business control logic into script packages. The script package contains executable script code, dependency declarations (such as the version of the protocol parsing script to be called, runtime libraries), and runtime configuration information (such as maximum memory quota, timeout).
[0037] To ensure transmission security, an encrypted transmission channel (such as a session based on the TLS / DTLS protocol) is established between the cloud collaboration center and the edge hardware device. The script package is digitally signed, and then the signed package is sent to the target edge hardware device through the encrypted channel.
[0038] Upon receiving the script package, the edge hardware device first verifies the legality and integrity of the digital signature (e.g., using a pre-configured cloud public key for signature verification) to confirm that it has not been tampered with and its source is trustworthy. After verification, the edge hardware device loads the script package into a previously built sandbox runtime environment and dynamically starts execution without restarting the device or interrupting existing services. For script updates, this embodiment supports incremental hot updates: a binary differential algorithm is used to calculate the difference data between the old and new scripts, and only the compressed difference data is transmitted to the device. The data is then merged locally on the device and takes effect, thereby saving network bandwidth.
[0039] S5: The edge hardware device receives and verifies the script package, and loads the verified script package into the sandbox runtime environment for execution.
[0040] S6, when the script package is executed in the sandbox environment, it performs cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type.
[0041] When the script package is executed in the sandbox environment, it does not directly manipulate the hardware or protocol stack. Instead, it interacts with the hardware device to be connected by calling the pre-loaded protocol parsing script corresponding to the communication protocol type.
[0042] Specifically, the business script packages in the sandbox call the protocol parsing engine through a standardized application programming interface (API). The protocol parsing engine then selects the corresponding protocol parsing script based on the target device's protocol type to complete the following operations:
[0043] Convert the uniform format commands issued by the business script (such as "read register 0x01") into the byte stream format required by the native protocol of the hardware device to be connected;
[0044] The native protocol response data returned by the hardware device to be connected is converted into a unified format and sent back to the business script.
[0045] In this way, the business script package does not need to worry about the differences in the underlying specific protocols. It only needs to read and write data according to a unified standard interface to achieve seamless communication with hardware devices using different protocols (Modbus, MQTT, HTTP, proprietary protocols, etc.). For example, a business script for temperature monitoring can communicate with a sensor using the Modbus protocol by calling a Modbus protocol parsing script, and at the same time communicate with an actuator using a proprietary protocol by calling a proprietary protocol parsing script, thereby realizing the linkage control between heterogeneous devices. The script itself does not contain any protocol adaptation code.
[0046] Furthermore, as a further optimization, when real-time data conversion is required between devices with different protocols (e.g., forwarding data from a Modbus device to an MQTT device), this embodiment supports direct modification of data packets in kernel mode using the Extended Berkeley Packet Filter (eBPF), avoiding the data copy overhead between user mode and kernel mode, achieving zero-copy protocol conversion, and reducing cross-protocol communication latency.
[0047] During script execution within the sandbox, the edge hardware device monitors the script's resource consumption (CPU time slices, memory usage, network bandwidth utilization) and its own status (remaining battery power, temperature, computing load) in real time. Based on this information, the edge hardware device dynamically adjusts the script's execution level or functional level.
[0048] Full load mode: When device resources are plentiful, the script is allowed to enable all extended functions (such as real-time video inference and high-frequency sampling).
[0049] Low-resource mode: When memory or CPU is scarce, non-critical functions are automatically degraded (such as reducing the sampling frequency and turning off some log output) to prioritize core data interaction;
[0050] Emergency Mode: When the device temperature is too high or the battery is extremely low, only the minimum core monitoring scripts will run, and other business scripts will be suspended.
[0051] Meanwhile, edge hardware devices support parallel execution and rapid rollback of multiple versions of scripts: when business scripts need to be updated, the new version and the old version can run simultaneously in a sandbox (for example, to conduct gray-scale testing on some devices or some data streams), and the old version can be completely replaced after the new version is verified to be stable; if the new version has an anomaly, it can be switched back to the old version immediately, thereby ensuring the high availability of the system.
[0052] In one specific implementation, such as Figure 2 As shown, when a hardware device to be accessed establishes a physical or wireless connection with an edge hardware device, the edge hardware device automatically captures the communication data stream of that device. The collected feature information includes: handshake packet content, packet header structure, packet length distribution, and packet transmission time interval pattern (i.e., time-series transmission mode). These features are used for subsequent protocol type determination. The communication features are input into the protocol identification model to automatically determine the protocol type.
[0053] This model is constructed using a convolutional neural network (CNN) combined with an attention mechanism. The CNN is used to extract local features of data packets (such as specific byte sequence patterns), while the attention mechanism is used to capture the temporal dependencies between data packet sequences. The model outputs a discrimination result to determine the type of communication protocol used by the hardware device to be connected, including standard protocols (such as Modbus, MQTT, and HTTP) or vendor-specific protocols.
[0054] Based on the identified protocol type, the local protocol script library of the edge hardware device is searched. If a corresponding protocol parsing script exists, the process jumps directly to the loading step; otherwise, based on the Protocol Description Language (PDL) engine, a parsing script conforming to the PDL specification is automatically generated using pattern matching and rule inference techniques based on the collected communication feature information. After verification, the script is stored in the protocol script library for later use.
[0055] The protocol parsing script is dynamically loaded into the protocol parsing engine of the edge hardware device. This loading process does not require restarting the device or interrupting existing communication services, and supports hot-swapping to update protocol parsing capabilities. After loading, the engine can use the script to parse subsequent communication data from the device.
[0056] The protocol parsing engine calls the loaded protocol parsing script to decode the received raw communication data, converting heterogeneous protocol formats into a unified standard data format (such as JSON or Protobuf). Based on this unified standard format, real-time cross-protocol data conversion and device interoperability can be achieved between devices using different protocols, such as forwarding data from a Modbus device to an MQTT device without needing to worry about the underlying protocol differences.
[0057] In another specific implementation, such as Figure 3 As shown, a lightweight sandbox runtime environment is built on top of the operating system of the edge hardware device. This environment strictly isolates the script execution space from the device's critical system resources (such as the kernel, file system root directory, network protocol stack, hardware drivers, etc.), ensuring that subsequently loaded scripts cannot directly access or damage the device's core functions. A three-layer security protection mechanism is configured for the sandbox runtime environment.
[0058] Before any script is loaded into the sandbox, it undergoes static code analysis. The analysis includes checks for dangerous operations (such as system exit calls or file deletion), resource leaks (such as unreleased memory or file handles), and infinite loops. If the security check passes, the script is allowed to load and execute; otherwise, loading is refused, and an exception is reported to the management module.
[0059] During script execution, the sandbox environment dynamically allocates and monitors the script's resource usage in real time, including CPU time slices, memory quotas, and network bandwidth. When script resource consumption exceeds preset thresholds, the system can impose restrictions or issue alerts based on policies to prevent a single script from excessively consuming resources and affecting the normal operation of the device's main process.
[0060] For script updates, a binary differential algorithm is used to calculate the differences between the old and new versions. Only the compressed difference data is transmitted from the cloud to the device, where it is merged locally to complete the update, reducing the amount of data transmitted. Simultaneously, the sandbox supports parallel execution of multiple script versions (e.g., coexistence of old and new versions during canary releases), allowing for rapid rollback to a stable version if an anomaly occurs in the new version.
[0061] In another specific implementation, such as Figure 4 As shown, the cloud-based collaboration center encapsulates specific business control logic into micro-script packages based on business needs. These script packages contain the following: executable script code (implementing control algorithms, data preprocessing logic, or edge inference logic), dependency declarations (e.g., the version of the protocol parsing script to be called, runtime libraries), and runtime configuration information (such as maximum memory quota, timeout, etc.).
[0062] The cloud-based collaboration center uses a private key to digitally sign the micro-script package, and then sends the signed package to the target edge hardware device via a TLS encrypted transmission channel. The encrypted channel ensures that the data is not eavesdropped on or tampered with during transmission.
[0063] After receiving the microscript package, the edge hardware device first uses a pre-configured cloud public key to verify the legality and integrity of the digital signature, confirming that the script package's source is trustworthy and has not been tampered with. Once verification is successful, the script package is directly loaded into a pre-built sandbox runtime environment, ready for execution, without requiring a device restart.
[0064] The loaded script runs securely in a sandbox environment, thereby extending the computing power of edge hardware devices.
[0065] During script execution, the edge hardware device monitors the script's running status (such as CPU usage and memory usage) and the device's own resource status (such as remaining battery power, temperature, and computing load) in real time. Based on the device's current status, the system automatically classifies the script's functional level, thereby dynamically adjusting the script's execution behavior.
[0066] The cloud-based collaboration center supports canary deployments of scripts: new script versions are first deployed on a small number of devices for verification, and then gradually rolled out to all devices after stability is confirmed. Simultaneously, the system records complete logs of all script push, loading, execution, and modification operations, and writes key operation records (such as deployment and rollback) to the blockchain network. The immutability of the blockchain ensures the authenticity and integrity of the audit logs, meeting compliance requirements.
[0067] Furthermore, the communication data stream characteristic information of the hardware device to be connected is collected, including:
[0068] On edge hardware devices, the DMA controller's transfer completion interrupt or half-transfer interrupt is used to read the DMA remaining count register value. The remaining count register value is then subtracted from the FIFO depth to obtain the instantaneous water level value of the current receiving FIFO.
[0069] On edge hardware devices, the calculation of the ripple spectrum does not depend on the message payload content of the communication data stream, but is based on the microscopic dynamic process of the FIFO level at the communication interface receiver. This process is directly determined by physical layer attributes such as the frame arrival pattern, baud rate configuration, and master-slave interaction rhythm of the hardware device to be connected. Therefore, it can provide physical layer accompanying characteristics that reflect the essential attributes of the communication protocol even when the data content is unknown or encrypted.
[0070] Edge hardware devices utilize a DMA controller common to embedded platforms to handle communication interface data transfer. Some ARM series microcontrollers have built-in DMA controllers that support memory-to-peripheral, peripheral-to-memory, and memory-to-memory transfer modes. In communication interface receiving scenarios, it is configured for peripheral-to-memory circular mode. In this mode, the DMA controller maintains a remainder counter register (such as a general-purpose NDTR / NDT register), the value of which indicates the number of data units yet to be transferred in the current DMA transfer batch.
[0071] The DMA controller generates two types of interrupts during the transfer process: Half Transfer Interrupt (HTIF) and Transfer Complete Interrupt (TCIF). The Half Transfer Interrupt is triggered when half of the data in the current transfer batch is completed; the Transfer Complete Interrupt is triggered when the entire data in the current transfer batch is completed. Edge hardware devices utilize these two types of interrupts as sampling triggers, reading the remaining count register value within the interrupt service routine.
[0072] The instantaneous water level of the current receive FIFO is calculated by subtracting the remaining count register value from the FIFO depth. The principle is that in DMA loop mode, the receive FIFO depth of the communication interface is a fixed hardware parameter (e.g., the receive FIFO depth of a Universal Asynchronous Receiver / Transmitter (UART) is typically 1 to 16 bytes, depending on the communication interface model used), while the remaining count register value reflects the amount of data remaining in the current DMA batch that has not yet been moved from the FIFO to main memory. Therefore, the difference between the FIFO depth and the remaining count is the number of data bytes currently remaining in the receive FIFO, waiting to be moved by the DMA, i.e., the instantaneous water level value. This calculation process involves only one subtraction operation, requiring no floating-point operations or table lookups, and is suitable for interrupt contexts.
[0073] Read the current count value of the kernel tick counter or debug trace cycle counter of the edge hardware device, divide the current count value by the system clock frequency to obtain the microsecond-level timestamp, and pair the instantaneous water level value with the microsecond-level timestamp and write it into the circular buffer;
[0074] Microsecond-level timestamps are derived from hardware counters already existing in the kernel of edge hardware devices, eliminating the need for additional external timers. Specifically, one of the following two sources can be used:
[0075] First, the kernel tick counter (SysTick). SysTick is a 24-bit decrementing counter that can be configured to generate an interrupt every 1 millisecond by writing the system clock frequency value to the reload register and setting the bit to enable it; its current count value increments and decrements with the granularity of the system clock cycle, and reading the current count value will give the number of clock cycles that have elapsed since the last reload.
[0076] Secondly, debug the tracking cycle counter (DWT_CYCCNT). This counter is a 32-bit incrementing counter located within the Data Watchpoint and Track (DWT) unit, enabled by setting the CYCCNTENA bit in the DWT control register. This counter automatically increments by one in each system clock cycle, and rolls back after overflowing, making it suitable for microsecond-level timestamp acquisition requiring higher precision.
[0077] The current count value is divided by the system clock frequency to obtain a microsecond-level timestamp. The system clock frequency is determined by the reset and clock control register (RCC) configuration of the edge hardware device, such as 72 MHz, 168 MHz, or 480 MHz, and the specific value is written to the clock tree configuration register during system initialization. The division operation is implemented in the firmware using integer division: microsecond-level timestamp = current count value / system clock frequency (unit: MHz). Here, the current count value is the number of CPU clock cycles (e.g., the current value of DWT_CYCCNT). For example, if the system clock frequency is 168 MHz, then every 168 count cycles corresponds to 1 microsecond. This division operation is performed once when an interrupt is triggered, pairing the instantaneous water level value with the microsecond-level timestamp and writing them together into a circular buffer.
[0078] The circular buffer is a fixed-capacity structure with a capacity of 256 entries. Each entry contains a 32-bit timestamp field and an 8-bit watermark field. It uses overwrite writing and rolls back to the beginning position when the writing position reaches the end of the buffer.
[0079] The circular buffer is a fixed-capacity structure, statically allocated in the memory of the edge hardware device, with a capacity of 256 entries. The reason for choosing 256 entries is that subsequent ripple spectrum calculations require backtracking 128 consecutive entries from the current write position as the endpoint. The capacity of 256 entries provides exactly twice the backtracking depth, ensuring that during the overwrite process, there is always an interval of no less than 128 entries between the current write position and the earliest valid data, avoiding data incompleteness when the backtracking window crosses the rollback boundary due to data overwriting.
[0080] The physical memory layout for each entry is as follows: the high 32 bits are the timestamp field, and the low 8 bits are the watermark field, occupying a total of 5 bytes. To accommodate the memory alignment requirements of embedded platforms, the actual storage can be padded with 8-byte alignment, where the low 5 bytes store valid data and the high 3 bytes are reserved or used for verification. A circular buffer maintains a write pointer, whose value ranges from 0 to 255. After each new entry is written, the write pointer is incremented by one, and a modulo 256 rollback is performed through a bitwise AND operation, thereby reusing the oldest historical entry in an overwrite manner.
[0081] A water ripple spectrum calculation is triggered every 128 new entries accumulated. The water ripple spectrum calculation includes:
[0082] Starting from the current write position in the circular buffer, backtrack 128 consecutive entries, perform a first-order difference operation on the water level field of the 128 consecutive backtracked entries to obtain 127 difference values, and trim each difference value to the interval [-8, +7].
[0083] The water ripple spectrum calculation starts from the current write position in the circular buffer and traces back 128 consecutive entries. Let the 128 water level fields obtained from the backtracking be arranged in chronological order as follows: , , ..., ,in As the earliest sample, This is the latest sample. First-order difference operations use forward differencing. For i = 1, 2, ..., 127, the i-th difference value is defined. The difference between two adjacent water level values:
[0084] , where i = 1, 2, ..., 127;
[0085] in, and These are the i-th and (i-1)-th watermark fields in the backtracking sequence (in bytes). The calculated difference values (unit: bytes) result in 127 difference values. The operation to prune to the [-8, +7] interval is defined as: for each ,like If < -8, then take -8; if If the value is greater than +7, then take +7; otherwise, keep the value. Original value.
[0086] The physical basis and data structure constraints of this clipping interval are as follows: The 16-level difference histogram needs to completely cover all possible difference values with 4-bit binary numbers (a total of 16 discrete levels). The range of the 4-bit signed integer is -8 to +7, which corresponds exactly to the 16 discrete values and strictly matches the array dimension of the 16-level difference histogram.
[0087] At the physical level, the sampling interval for ripple spectrum calculation is configured to be less than or equal to the transmission time of a single byte (e.g., through a single-byte transmission interrupt using multiplexed DMA, or by using a timer interrupt to sample at a period not exceeding 100 microseconds). Under this configuration, the change in FIFO water level within a single sampling interval is typically 0 or ±1 byte. However, to cover abnormal fluctuations under non-ideal communication conditions (such as water level misreading caused by electrical noise, multi-byte burst reception, etc.), the clipping interval is extended to ±8 bytes. This preserves most of the effective water level dynamic information while ensuring that the dimensions of the differential histogram are fixed, avoiding dynamic memory allocation at the edge.
[0088] The cropped difference values are mapped to 16-level difference histograms. The 16-level difference histograms are stored in array form. Array indices 0 to 7 correspond to difference values -8 to -1, respectively, and array indices 8 to 15 correspond to difference values 0 to +7, respectively. The array element value is the frequency of occurrence of the corresponding difference value among 127 difference values.
[0089] The 16-level difference histogram is stored as a one-dimensional array of length 16, and there is a fixed mapping relationship between the array index and the clipped difference value.
[0090] The mapping rules are as follows: array index 0 corresponds to a difference value of -8; array index 1 corresponds to a difference value of -7; ... array index 7 corresponds to a difference value of -1; array index 8 corresponds to a difference value of 0; array index 9 corresponds to a difference value of +1; ... array index 15 corresponds to a difference value of +7.
[0091] When calculating the 16-level difference histogram, the 127 cropped difference values are traversed. For each difference value, its corresponding array index is determined according to the mapping rules described above, and the array element value at that index is incremented by 1. The initial values of the array elements are cleared to zero before the water ripple spectrum calculation is triggered. After the calculation is completed, each array element value represents the frequency of occurrence of the corresponding difference level among the 127 difference values.
[0092] This 16-level differential histogram characterizes the instantaneous jitter of the communication flow: the frequency distribution of array indices 0 to 7 (negative levels) reflects events such as a sudden drop in the FIFO level (e.g., data clearing or frame termination caused by DMA bulk transfer); the frequency distribution of array indices 8 to 15 (non-negative levels) reflects events such as a gradual rise or sudden increase in the FIFO level (e.g., data filling caused by frame arrival). Different communication protocols exhibit different distribution centers and degrees of dispersion on this histogram due to differences in frame length, frame interval, and flow control strategies.
[0093] Calculate the pulse energy, which is the sum of squares of 127 difference values, and store it in a 16-bit unsigned integer saturated accumulation mode.
[0094] The pulse energy is stored as a 16-bit unsigned integer, and the accumulation process adopts a saturation accumulation mechanism: when performing the sum of squares accumulation, it is determined whether the current accumulation result exceeds 65535 (i.e., the maximum value of a 16-bit unsigned integer); if it does, the accumulation result is kept at 65535, instead of being rolled back to 0 or the lower 16 bits are truncated.
[0095] The reason for using a saturation mechanism instead of ordinary overflow is that pulse energy is used to characterize the overall intensity of sudden changes in the communication flow. If ordinary overflow and wrap-up are used, the accumulated result in high-energy scenarios will jump to a minimum value due to overflow, causing the energy measurement to lose its monotonicity, which in turn leads to misjudgments in subsequent scheduling mode switching based on pulse energy. The saturation mechanism ensures that the energy value stabilizes at the maximum value after reaching the upper limit, transmitting a conservative signal of continuous strong changes to subsequent modules, which conforms to the principles of edge security design.
[0096] The periodicity index is calculated as follows: the periodicity index is the 8-bit unsigned fixed-point number obtained by dividing the first-order autocorrelation value of the 128 water level fields by the sum of the squares of the 128 water level fields, multiplying by 255, and rounding down.
[0097] Assume 128 water level fields are arranged in chronological order as follows: , ,..., The delayed first-order autocorrelation value R is defined as the sum of the products of adjacent water level fields.
[0098]
[0099] The sum of squares S is defined as the sum of the squares of all water level fields, i.e.
[0100] The calculation of the periodicity exponent P uses a classic idea in signal processing: to look at the similarity between two consecutive steps in a sequence.
[0101] In this specific embodiment, the 128 water level values are considered as a short sequence. First, the sum of the products of two adjacent water levels (R) is calculated, and then the sum of the squares of each water level (S) is calculated. The ratio of these two values, R / S, is naturally between 0 and 1, because it reflects the size of the adjacent product relative to its own square.
[0102] When water level changes are regular and adjacent values are roughly equal, R is close to S, and the ratio is close to 1. When water levels fluctuate randomly or remain at zero, the ratio is very small. Multiplying this ratio by 255 and then rounding down gives an integer P between 0 and 255. Multiplying by 255 is for easy storage using an 8-bit integer and does not change the physical meaning of the ratio. If all water levels are 0 (S=0), P is directly set to 0, which indicates no flow and avoids division by zero.
[0103] The calculation method in this embodiment avoids the variance division required for the standard autocorrelation coefficient, and is achieved solely by the ratio of the sum of the products of adjacent water levels to the sum of their squares. This method is suitable for rapid execution within a microcontroller interrupt service routine. Although P here is not the autocorrelation coefficient in the strict sense, it is sufficient to distinguish between regular polling and random noise communication modes, meeting practical requirements on edge devices.
[0104] Specifically, the formula for calculating the periodicity index P is:
[0105]
[0106] in, This indicates rounding down. Here, the denominator uses the sum of squares S instead of the mean square value to ensure that the ratio R / S does not exceed 1. The result falls within the range of 0 to 255 and can be directly stored as an 8-bit unsigned fixed-point number without additional saturation truncation processing.
[0107] because Since S is a non-negative integer, according to the basic inequality, R ≤ S. Therefore, when S > 0, the value of 255 × R / S falls within the interval [0, 255]; when S = 0, it indicates that there is no valid communication flow, and P is directly set to 0.
[0108] ratio The step self-similarity reflecting the water level sequence is as follows: when the water level changes regularly, the ratio of the product of adjacent water levels to the sum of squares is close to 1, and P approaches 255; when the water level fluctuates randomly or is idle, the ratio approaches 0, and P approaches 0. It should be noted that the naming of the periodic index in this invention is based on its physical effect, and its mathematical essence is the step self-similarity of the water level sequence.
[0109] The 16-level difference histogram, pulse energy, and periodic index are combined into a water ripple spectrum data structure, which is stored in the memory area of the edge hardware device.
[0110] The water ripple spectrum data structure consists of the following three parts, which are sequentially assembled in memory:
[0111] 16-level difference histogram: 16 bytes, stored sequentially with array indices 0 to 15;
[0112] Pulse energy: 2 bytes, stored as a 16-bit unsigned integer in big-endian or little-endian order (consistent with the byte order of the edge hardware device platform).
[0113] Periodicity index: 1 byte, stored as an 8-bit unsigned integer.
[0114] The above three parts occupy a total of 19 bytes. Within the fixed memory region of the edge hardware device, the watermark data structure is allocated as a global static array. Its storage address is specified as a fixed segment (such as a custom .watermark segment) in the linker script, ensuring that the data structure's address remains unchanged after a system reset and is not overwritten by runtime stack or heap memory allocations. The protocol identification constraint, sandbox permission modulator, scheduling mode switcher, and script loading preorderer all directly access this data structure through the fixed memory address, eliminating the need for pointer indirect addressing, thereby reducing access latency and ensuring address consistency across multiple executables.
[0115] It should be noted that the characteristic information of communication data streams includes not only traditional message content characteristics (such as byte sequences, protocol fields, statistical entropy values, etc.), but also accompanying dynamic characteristics generated at the physical layer of the communication interface during message transmission. In this application, the edge hardware device receives communication data from the hardware device to be connected via DMA. The data stream forms a micro-dynamic process of transient accumulation and clearing in the FIFO of the communication interface. This micro-dynamic process reflects the essential attributes of the communication protocol, such as the frame arrival pattern, baud rate stability, and master-slave interaction rhythm of the hardware device to be connected, and is unrelated to the message payload content. Therefore, collecting the instantaneous FIFO water level value at the moment of DMA interrupt triggering and calculating the ripple spectrum according to the time series is a multi-modal acquisition method for the characteristic information of communication data streams; the multi-modality is reflected in the fact that it includes both traditional data layer characteristics (such as the message byte statistical characteristics used in the subsequent protocol identification model) and physical layer accompanying characteristics (i.e., the micro-dynamic characteristics of the FIFO water level).
[0116] Furthermore, the edge hardware device uses either the ripple spectrum data structure or the 3D physical layer hidden state vector derived from the ripple spectrum data structure simultaneously for:
[0117] Prune the candidate protocol space of the pre-trained protocol recognition model;
[0118] Dynamically adjust the write permission bit in the effective permission vector of the sandbox runtime environment;
[0119] Switch the scheduler between fixed-step scheduling mode and flexible-step scheduling mode;
[0120] The protocol parsing script entries in the protocol script library are sorted by matching deviation.
[0121] After the water ripple spectrum calculation is completed, the edge hardware device stores 16-level difference histograms (16 bytes), pulse energy (2 bytes), and periodicity index (1 byte) in its memory. In order to simultaneously support the four subsystems of protocol identification, sandbox permission control, scheduling mode switching, and script pre-sorting, this embodiment stores these three parts of data continuously in a fixed shared memory area.
[0122] To facilitate rapid subsequent judgment, a three-dimensional physical layer hidden state vector can be derived from the original ripple spectrum. This vector has three components: a burst index, a regularity index, and a sparsity index, which are obtained by normalizing the inner product of a 16-level difference histogram with three sets of preset weight templates. The burst index reflects the degree of clustering and arrival jitter of data packets in the communication flow, the regularity index characterizes the stability of adjacent interaction time intervals, and the sparsity index describes the idle proportion of the communication link. The original ripple spectrum and the three-dimensional vector are equivalent: the former retains all details and is suitable for precise matching; the latter provides high-dimensional abstraction and is suitable for threshold judgment. To reduce redundancy, this embodiment synchronously derives the three-dimensional vector and stores it in shared memory after each round of ripple spectrum calculation.
[0123] There are priorities among the four subsystems: security-related operations take precedence over performance optimization. The output of the sandbox permission modulator has the highest priority; once it triggers read-only mode, other subsystems must comply with the permission restrictions. The protocol identification constraint unit is next in priority; its output of the candidate protocol space is a prerequisite for subsequent script loading. The scheduling mode switcher and the script loading preorderer belong to the performance optimization module.
[0124] Furthermore, the communication protocol type of the hardware device to be accessed is determined based on the pre-trained protocol recognition model, including:
[0125] Three sets of weight templates are pre-configured in the read-only storage area of the edge hardware device. Each set of weight templates includes 16 8-bit signed fixed-point numbers. The three sets of weight templates correspond to the burstiness dimension, the regularity dimension, and the sparsity dimension, respectively.
[0126] The edge hardware device's read-only memory contains three pre-defined weight templates, corresponding to the three physical dimensions of burstiness, regularity, and sparsity. Each template contains 16 8-bit signed fixed-point numbers in Q7 format (the highest bit is the sign bit, and the lower 7 bits represent the fractional part, with a value range of -128 / 128 to +127 / 128). These 16 values correspond one-to-one with the levels of a 16-level difference histogram: the first value corresponds to the level with a difference of -8 (i.e., histogram index 0), the second value corresponds to the level with a difference of -7 (index 1), and so on, with the sixteenth value corresponding to the level with a difference of +7 (index 15).
[0127] The values for these three templates are derived from statistical learning during the offline phase. Specifically, in a laboratory environment, various existing industrial communication protocols (including Modbus RTU and Modbus TCP) were used under standard operating conditions to collect water ripple spectrum data and calculate 16-level difference histograms. Then, linear discriminant analysis was performed on the large number of collected histogram samples to extract three projection vectors that best distinguish between high and low burstiness, strong and weak regularity, and high and low sparsity. The 16 components of each projection vector were quantized into 8-bit signed fixed-point numbers according to Q7 format and stored in the burstiness template, regularity template, and sparsity template, respectively. Fixed-point inner product operations were performed between the 16-level difference histograms in the water ripple spectrum data structure and the three weight templates to obtain three inner product results. Each inner product result was right-shifted by 8 bits and normalized to form the three-dimensional physical layer hidden state vector.
[0128] Among them, the three components of the three-dimensional physical layer hidden state vector are all 8-bit unsigned fixed-point numbers, and the three components of the three-dimensional physical layer hidden state vector respectively represent the burstiness index, regularity index and sparsity index of the communication flow of the hardware device to be accessed.
[0129] The protocol identification constraint reads the 16-level difference histogram (i.e., 16 frequency values, each ranging from 0 to 127) obtained from the current ripple spectrum calculation cycle. Then, these 16 frequency values are multiplied and accumulated with three sets of weight templates: for the burst template, the first frequency value of the histogram is multiplied with the first value of the template, the second frequency value is multiplied with the second value of the template, and so on, and the 16 products are added together to obtain the accumulated result; the above process is repeated for the regular template and the sparse template, resulting in three accumulated results.
[0130] Each accumulated result is a signed integer. Then, the three accumulated results are right-shifted by 8 bits (equivalent to dividing by 256), and the results are restricted to the range of 0 to 255 (0 for values less than 0 and 255 for values greater than 255), thus obtaining three 8-bit unsigned integers, which are the burstiness index, the regularity index, and the sparsity index.
[0131] The reason for shifting right by 8 bits is that the maximum value of each of the 16 frequency values is 127, the maximum absolute value of the template value is 16, and the maximum absolute value of the multiplication and accumulation is approximately 127 multiplied by 16 and then multiplied by 16, which is approximately 32512. After shifting right by 8 bits, it is approximately 127, which is much less than 255. Therefore, saturation overflow will not occur, but the saturation truncation logic is retained to deal with abnormal data.
[0132] The candidate protocol space is pruned based on three components of the three-dimensional physical layer hidden state vector, including: if the sparsity index exceeds the first preset threshold, the protocol type with a frame interval less than the preset frame interval threshold is excluded from the candidate protocol space; if the regularity index exceeds the second preset threshold and the burstiness index is lower than the third preset threshold, the matching priority of the protocol type with request-response periodic alternation characteristics in the candidate protocol space is increased.
[0133] It should be noted that the request-response periodic alternation characteristic refers to an interaction mode in which the two communicating parties alternately send data according to a fixed time rhythm. Specifically, one party in the communication (usually called the master station or client) sends request frames at roughly constant time intervals, and the other party (slave station or server), after receiving the request and experiencing a relatively fixed processing delay, replies with the corresponding response frame. The entire process presents a cyclical rhythm of "request-response-wait-request-response-wait," as regular as a pendulum. This characteristic is fundamentally different from continuous burst transmission (one party sends multiple frames consecutively without waiting for a reply) or completely random transmission.
[0134] In the field of industrial communication, typical protocols exhibiting this characteristic include Modbus RTU master-slave polling mode, Modbus TCP periodic data acquisition, and many proprietary protocols that employ a "master query - slave response" mechanism. HTTP / 1.1 short connections also exhibit similar characteristics in scenarios where GET requests are initiated at fixed intervals.
[0135] In this embodiment, before inputting the communication data stream features into the protocol recognition model, the candidate protocol space is dynamically pruned based on three indices in the three-dimensional physical layer hidden state vector. Below are some specific value examples.
[0136] The first preset threshold is used to compare with the sparsity index to determine whether the idle ratio of the current communication flow is sufficiently high. This threshold is calibrated as follows: During the edge hardware device initialization phase, a segment of communication traffic known to have extremely low load (e.g., the device only sends periodic heartbeat packets after power-on) is collected, its sparsity index is calculated, and this value is multiplied by a safety factor (e.g., 1.2 to 1.5 times) as the initial reference value for the first preset threshold. Simultaneously, measured data of the sparse protocol under standard operating conditions can be referenced. As an implementable example, in a steady-state communication, when the measured value of the sparsity index is 170, the first preset threshold can be set to 180. This threshold can be adjusted online by the cloud-based collaboration center based on the fieldbus load.
[0137] A preset frame interval threshold is used to distinguish between protocols with "longer frame intervals" and "extremely short frame intervals." The value of this threshold is related to the type of physical bus to which the edge hardware device is connected and the typical application scenario. As an example, in scenarios primarily using serial communication, the preset frame interval threshold can be set to 5 milliseconds; in scenarios involving CAN bus or real-time Ethernet, the threshold can be set to 1 millisecond. This threshold can be modified by the user through the configuration interface according to the actual communication parameters of the access device.
[0138] The second preset threshold is used to compare with the regularity index to determine whether the current communication flow exhibits a highly stable temporal rhythm. This threshold is calibrated as follows: During the device deployment phase, a device known to use a periodic polling protocol (e.g., a Modbus master station) is connected, its steady-state communication data is collected, and the regularity index is calculated. 0.8 to 0.9 times this value is used as the initial reference value for the second preset threshold. As an example, under the aforementioned Modbus RTU periodic polling condition, when the measured value of the regularity index is 200, the second preset threshold can be set to 180.
[0139] The third preset threshold is used to compare with the burst index to determine whether the current communication flow is in a stable state with low jitter and no congestion. This threshold is calibrated as follows: a data flow known to be a stable, responsive communication stream (e.g., a Modbus slave's response stream to a master's request) is collected, its burst index is calculated, and 1.5 to 2 times this value is used as the initial reference value for the third preset threshold. As an example, under stable conditions during a single interaction, when the measured burst index is 45, the third preset threshold can be set to 70.
[0140] If the calculated sparsity index exceeds the first preset threshold, it indicates that the idle intervals between data packets in the communication flow are generally long. In this case, protocol types with frame intervals smaller than the preset frame interval threshold are excluded from the candidate protocol space. This is because protocols with extremely short frame intervals (such as high-load CAN bus, real-time Ethernet, etc.) will necessarily exhibit tightly packed data packets at the physical layer, and their sparsity index will be much lower than the first preset threshold. There is an essential contradiction between the two, and excluding these types can avoid conflict decisions in subsequent models.
[0141] If the current calculated regularity index exceeds the second preset threshold, and the suddenness index is lower than the third preset threshold, it indicates that the communication flow exhibits a stable, low-frequency, and non-sudden regular rhythm, which is exactly the manifestation of the aforementioned request-response periodic alternation characteristic.
[0142] At this point, the matching priority of protocol types with this characteristic in the candidate protocol space (such as Modbus RTU master-slave polling, Modbus TCP periodic sampling, question-and-answer private protocols, etc.) is increased. The specific implementation of priority increase is as follows: in the probability vector output by the protocol identification model, the original probability values of these protocol types are multiplied by a weighting coefficient (e.g., 1.5), and then all probability values are renormalized, making protocols with periodic alternation characteristics more likely to be selected in the end.
[0143] The pre-trained protocol recognition model performs communication protocol type discrimination within the pruned candidate protocol space and outputs the communication protocol type of the hardware device to be connected.
[0144] After the above pruning, the number of types in the candidate protocol space is typically reduced from dozens to 3 to 8. The pre-trained protocol recognition model only performs discrimination within this reduced space.
[0145] The protocol recognition model used in this embodiment is a hybrid structure of convolutional neural network and attention mechanism. Its input data comes from the raw communication data stream of the hardware device to be connected, specifically including: several consecutively captured complete data packets (usually 8 to 16 packets), the payload byte sequence of each data packet (the first 64 bytes are fixed), and the inter-packet time interval (normalized to the range of 0 to 255). This information is concatenated into a two-dimensional matrix as the model input. The matrix size is: the number of rows is the number of data packets (e.g., 8 rows), and the number of columns is the feature length of each packet (e.g., 64 bytes plus 4 bytes of time interval, for a total of 68 columns).
[0146] The forward inference process of the model is as follows: First, the two-dimensional input matrix passes through three convolutional neural network layers, each using a 3×3 convolutional kernel and the ReLU activation function to progressively extract local features within each data packet (e.g., specific byte sequence patterns, checksum positions, frame header identifiers, etc.). The output of the convolutional layers is flattened into a one-dimensional feature vector. This feature vector is then fed into the attention mechanism module. The attention mechanism captures cross-packet temporal dependencies (e.g., pairwise patterns of request and response packets, periodic patterns in packet intervals) by calculating the correlation weights of features between different data packets. The global feature vector output by the attention module passes through a fully connected layer, mapping it to various types in the candidate protocol space, and outputs the matching probability of each protocol type through the softmax function. The number of output nodes of the fully connected layer equals the number of protocol types in the pruned candidate space, not the total number of known protocols.
[0147] The model was trained offline using a large number of communication data samples labeled with protocol types. During training, the model learned the mapping relationship from message content and temporal features to protocol types. When actually executed on edge devices, because the candidate space has been pruned and reduced, the model only needs to compute a small number of output nodes, significantly improving inference speed and avoiding misjudgments caused by candidate types containing protocols with obviously contradictory physical layer features.
[0148] Furthermore, a sandbox runtime environment is built on the edge hardware device to isolate the script execution space of the sandbox runtime environment from the critical system resources of the edge hardware device, including:
[0149] The sandbox runtime environment maintains a valid permission vector, and each bit field of the valid permission vector corresponds to a system call type.
[0150] The sandbox runtime environment maintains a 32-bit valid permission vector. The lower 8 bits are defined as follows: bit 0 corresponds to file write operations (such as open write, write); bit 1 corresponds to network send operations (such as send); bit 2 corresponds to hardware register write operations (such as iowrite); bit 3 corresponds to system configuration modification operations. Bits 4 to 7 are reserved. The mapping relationship between bit fields and system call numbers is fixed during sandbox initialization using a static lookup table.
[0151] The sandbox runtime environment reads the burst index of the hidden state vector of the three-dimensional physical layer in real time. When the burst index exceeds the preset burst threshold, the sandbox runtime environment clears the bit field representing the write operation type in the effective permission vector to zero, so that the sandbox runtime environment enters read-only mode.
[0152] A preset burst threshold is used to determine whether the communication flow is in an extreme burst state. This threshold can be calibrated according to the scenario: in industrial control scenarios with high real-time requirements, a lower threshold can be set to respond to anomalies earlier, for example, a value of 150; in data acquisition scenarios that tolerate certain fluctuations, a higher threshold can be set to avoid false triggering, for example, a value of 200. The specific value is set by the maintenance personnel through the configuration file according to the fieldbus load characteristics. The sandbox permission modulator reads the burst index in each ripple spectrum calculation cycle. When the index exceeds the preset burst threshold, an atomic operation is performed: interrupts are disabled, the valid permission vector is read, bits 0 to 3 are cleared to zero, written back, and interrupts are enabled. The atomic operation ensures that the permission state is not corrupted in a multi-tasking or interrupted environment. After clearing, the sandbox enters read-only mode. In read-only mode, write operation system calls of script packages and protocol parsing scripts to the hardware devices to be accessed are blocked;
[0153] In read-only mode, the sandbox intercepts write-type system calls in kernel space via seccomp-bpf. Specifically, the BPF filter determines the validity based on the permission vector: if the bit field corresponding to the call number is 0, execution is refused, the error code -EPERM (operation not allowed) is returned, and the blocking event (timestamp, call number, script ID) is recorded in the local log. Read-type calls are unaffected.
[0154] When the suddenness index is below the preset recovery threshold for three consecutive water ripple spectrum calculation cycles, the sandbox operating environment will restore the bit field representing the write operation type in the effective permission vector to its original value before entering read-only mode.
[0155] A preset recovery threshold lower than a preset burst threshold creates hysteresis. For example, if the preset burst threshold is set to 150, the preset recovery threshold can be set to 80; if the preset burst threshold is set to 200, the recovery threshold can be set to 100. The specific value is obtained through calibration: a typical value of the burst index is measured under stable communication conditions, and 1.5 to 2 times this typical value is taken as the preset recovery threshold. The sandbox maintains a recovery counter, initially set to 0. At the end of each ripple spectrum calculation cycle: if the burst index is lower than the preset recovery threshold, the counter is incremented by 1; otherwise, the counter is reset to zero. When the counter reaches 3, an atomic operation is performed to restore bits 0 to 3 to their original values (usually all 1s), and then the read-only mode is exited.
[0156] One water ripple spectrum calculation cycle corresponds to the collection of 128 new water level entries, and its duration depends on the communication rate. Users can accept a certain range of recovery delays based on the actual communication rate. The counting condition of three consecutive cycles ensures that the link has stably recovered from the abnormal state, rather than from momentary fluctuations.
[0157] Furthermore, when the script package is executed within the sandbox environment, it performs cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script, including:
[0158] When the protocol parsing script is loaded, it registers the baseline step length with the scheduler in the sandbox environment. The baseline step length represents the expected bus occupancy time required for the protocol parsing script to complete a single protocol semantic state transition under standard communication conditions. It is stored in a fixed memory area of the sandbox environment in a fixed-point format.
[0159] When the protocol parsing script is loaded, it registers a baseline step length with the scheduler within the sandbox environment. This baseline step length characterizes the expected bus occupancy time required for the script to complete a single protocol semantic state transition under standard communication conditions (i.e., no interference, no congestion). For example, for a Modbus RTU protocol parsing script, the baseline step length for completing a "send request-wait-receive response" state transition can be set to 20 milliseconds. This value can be obtained through statistical testing in a laboratory environment or directly calculated from the frame length and typical baud rate defined in the protocol specification. The baseline step length is stored in a 16-bit unsigned fixed-point number format, in microseconds, in a fixed area of the sandbox environment's memory. This address is allocated during sandbox initialization and remains unchanged.
[0160] The scheduler maintains a step budget table, which records the current step budget of each loaded protocol parsing script.
[0161] The scheduler maintains a step budget table. This table is a linear array with the same number of entries as the number of currently loaded protocol parsing scripts. Each entry contains three fields: a script identifier (e.g., script ID or name), the current state step budget (in microseconds, dynamically updated), and the baseline step time registered for that script (read-only). The step budget table is stored in the memory heap area of the sandbox runtime environment and is exclusively accessed by the scheduler.
[0162] The scheduler reads the regularity index of the hidden state vector of the three-dimensional physical layer in real time:
[0163] When the regularity index exceeds the preset high regularity threshold, the scheduler switches to fixed step scheduling mode. Fixed step scheduling mode uses the base step time as the smallest scheduling unit and maps the status step of each protocol parsing script to the bus time axis according to the base step time.
[0164] When the regularity index is lower than the preset low regularity threshold, the scheduler switches to the elastic step size scheduling mode. The elastic step size scheduling mode no longer determines the state step size budget based on the baseline step size. Instead, it multiplies the pulse energy obtained from the most recent ripple spectrum calculation by the preset conversion coefficient to obtain the estimated value of the actual bus occupancy time, and uses the estimated value of the actual bus occupancy time as the current state step size budget.
[0165] The scheduler reads the regularity index of the hidden state vector of the three-dimensional physical layer in real time and switches the scheduling mode according to the index.
[0166] When the regularity index exceeds a preset high regularity threshold, the scheduler switches to fixed-step scheduling mode. This mode is determined by the communication link exhibiting a highly stable temporal rhythm (e.g., a periodically polling Modbus master), making deterministic time-slice scheduling suitable. In fixed-step scheduling mode, the scheduler uses the baseline step size of each script as the smallest scheduling unit, arranging the state step sizes of each script sequentially along the time axis to form a polling schedule table. For example, if the baseline step size for script A is 20 milliseconds and for script B is 30 milliseconds, the scheduler will allocate 20 milliseconds to A and 30 milliseconds to B sequentially along the time axis, executing them in a loop. The scheduling granularity is at the microsecond level and is driven by system tick interrupts.
[0167] When the regularity index falls below a preset low regularity threshold, the scheduler switches to flexible step-size scheduling mode. This mode is determined by the following criteria: the communication link exhibits random and irregular timing characteristics (e.g., bursty sensor data), making fixed time-slice allocation inefficient. In flexible step-size scheduling mode, the scheduler no longer uses a baseline step length but instead multiplies the pulse energy calculated from the most recent ripple spectrum by a preset conversion coefficient to obtain an estimated actual bus occupancy time. The preset conversion coefficient represents the bus occupancy time corresponding to each unit of pulse energy, and its value can be calibrated experimentally. As an example, for serial communication scenarios, the conversion coefficient can be set to 10, meaning that for every 1 unit increase in pulse energy, the estimated bus occupancy time increases by 10 microseconds; for high-speed bus scenarios, the conversion coefficient can be set to 1. The scheduler uses this estimated value to update the current script's state step-size budget in the step-size budget table, and subsequent scheduling is based on this dynamic budget.
[0168] The preset high and low regularity thresholds are determined through calibration. For example, in periodically polling Modbus communication, the regularity index is typically higher than 180, while in completely random MQTT communication, the regularity index is typically lower than 100. Therefore, the high regularity threshold can be set to 160, and the low regularity threshold to 120. Specific values can be adjusted by the cloud-based collaboration center based on the characteristics of on-site communication. A smooth transition strategy is adopted when switching scheduling modes: the current state step size budget is not reset; instead, the unfinished budget value from the previous cycle is inherited to avoid instruction interruptions or timeouts due to the switch.
[0169] Furthermore, the protocol parsing script corresponding to the communication protocol type is dynamically loaded from the protocol script library, including:
[0170] Each protocol parsing script entry in the protocol script library is associated with a ripple spectrum expectation template. The ripple spectrum expectation template includes the centroid index of the differential histogram of the protocol parsing script under standard communication conditions, the average pulse energy, and the average periodic exponent. The ripple spectrum expectation template is pre-placed in the metadata area of the protocol script library.
[0171] Each protocol parsing script entry in the protocol script library is associated with a ripple spectrum expectation template. This template records the typical physical layer characteristics of the protocol under standard communication conditions and is used to match the current real-time ripple spectrum. The ripple spectrum expectation template is generated as follows: In a laboratory environment, the device to be connected is allowed to run stably under standard operating conditions (e.g., typical baud rate, typical frame interval, no interference), and multiple ripple spectrum samples are collected. The centroid index of the differential histogram, the pulse energy, and the periodicity index are calculated for each sample, and then the average value of each is taken as the template content. The template is pre-stored in the metadata area of the protocol script library, and the storage format is: the centroid index of the differential histogram occupies 1 byte, the average pulse energy occupies 2 bytes, and the average periodicity index occupies 1 byte, for a total of 4 bytes.
[0172] The determination of the centroid index of the difference histogram includes: multiplying each array index of the 16-level difference histogram with its corresponding element value, and summing all the products to obtain the weighted index sum; summing the element values of the 16-level difference histogram to obtain the total frequency sum; and dividing the weighted index sum by the total frequency sum and rounding it down to obtain the centroid index of the difference histogram.
[0173] Before loading the protocol parsing script corresponding to the communication protocol type from the protocol script library, the edge hardware device calculates the matching deviation between the current ripple spectrum and the expected template of the ripple spectrum of each protocol parsing script entry in the protocol script library.
[0174] The matching bias is the weighted sum of the following three absolute differences: the absolute value of the difference between the centroid index of the differential histogram of the current ripple spectrum and the centroid index of the differential histogram of the expected template of the ripple spectrum; the absolute value of the difference between the pulse energy of the current ripple spectrum and the mean pulse energy of the expected template of the ripple spectrum; and the absolute value of the difference between the periodicity index of the current ripple spectrum and the mean periodicity index of the expected template of the ripple spectrum.
[0175] The protocol parsing script entries in the protocol script library are sorted from smallest to largest according to the matching deviation. The protocol parsing script entry with the smallest matching deviation is loaded first, and the protocol parsing script entries with matching deviation greater than the preset exclusion threshold are removed from the candidate list for this loading.
[0176] The calculation method for the centroid index of the differential histogram is as follows: First, multiply each index (0 to 15) of the 16-level differential histogram by the corresponding frequency value, and sum all the products to obtain the weighted index sum; second, sum all the frequency values of the 16 levels to obtain the total frequency sum; finally, divide the weighted index sum by the total frequency sum and round the result to the nearest integer to obtain the centroid index. The physical meaning of this centroid index is the central trend of the differential jitter distribution of the communication flow. For example, a small centroid index (close to 0) indicates that negative differentials (sudden drop in water level) are dominant, while a large centroid index (close to 15) indicates that positive differentials (sudden increase in water level) are dominant.
[0177] Before loading the protocol parsing script from the protocol script library, the edge hardware device calculates the matching deviation between the current ripple spectrum (obtained from the most recent ripple spectrum) and the expected template of the ripple spectrum for each entry in the script library. The matching deviation is a weighted sum of three absolute differences: the first is the absolute value of the difference between the centroid index of the current difference histogram and the centroid index of the expected template; the second is the absolute value of the difference between the current pulse energy and the mean pulse energy of the expected template; and the third is the absolute value of the difference between the current periodicity index and the mean periodicity index of the expected template.
[0178] In this embodiment, the three items are weighted equally, meaning that the differences of each item are directly added together as the matching deviation. The reason for using absolute difference instead of Euclidean distance or relative difference is that: absolute difference is simple to calculate, suitable for real-time execution on embedded platforms, and the dimensions of each physical quantity have been normalized to a similar range (centroid index 0-15, pulse energy 0-65535, periodicity index 0-255), so no additional weighting is required.
[0179] After obtaining the matching deviation of each script entry, the edge hardware device sorts the protocol parsing script entries in ascending order of matching deviation. The sorting algorithm can use linear scan to select the item with the smallest deviation (when only one script needs to be loaded) or insertion sort (when the number of entries is small). At the same time, a preset exclusion threshold is used to remove entries with excessive deviation.
[0180] The exclusion threshold is determined by calculating the matching deviation between the ripple spectrum of a known correct protocol and its own expected template under standard communication conditions, and taking 2 to 3 times this deviation value as the exclusion threshold. As an example, the exclusion threshold can be set to 50. Script entries with a matching deviation greater than 50 are considered mismatched and removed from the current loading candidate list. Ultimately, the device prioritizes loading the protocol parsing script entry with the smallest matching deviation. If multiple entries have the smallest deviation (e.g., multiple entries with the same deviation and less than the exclusion threshold), then any one of them is loaded, or the latest version from the script library is loaded.
[0181] Furthermore, the method also includes:
[0182] Edge hardware devices map the ripple spectrum data structure to a fixed region of shared memory;
[0183] The edge hardware device maps the ripple spectrum data structure to a fixed region of shared memory. This shared memory region is allocated in the device's memory, with its starting address specified by the linker script (e.g., set to 0x20004000), and a length of 32 bytes (the ripple spectrum data structure actually occupies 19 bytes, with 13 bytes reserved for expansion and status word storage). The storage offsets of the ripple spectrum data structure within the region are as follows: offsets 0 to 15 store the 16-level difference histogram, offsets 16 to 17 store the pulse energy, offset 18 stores the periodicity index, and offsets 19 to 31 are reserved.
[0184] The protocol identification constraint, sandbox permission modulator, scheduling mode switcher, and script loading preorderer in the sandbox runtime environment, as concurrent consumption execution entities, all read the water ripple spectrum data structure from the fixed area of shared memory through a read-only access interface;
[0185] The protocol recognition constraint extracts the difference histogram distribution features and periodicity index from the water ripple spectrum data structure, and prunes the candidate protocol space of the pre-trained protocol recognition model;
[0186] The sandbox permission modulator extracts pulse energy from the water ripple spectrum data structure and dynamically adjusts the bit field state representing the write operation type in the effective permission vector of the sandbox operating environment.
[0187] The scheduling mode switcher extracts regularity indices from the water ripple spectrum data structure and switches between fixed step size scheduling mode and flexible step size scheduling mode.
[0188] The script loads a pre-sorter to extract the centroid index, pulse energy, and periodicity index of the difference histogram in the water ripple spectrum data structure, and performs a matching deviation sorting on the protocol parsing script entries in the protocol script library.
[0189] The protocol identification constraint, sandbox permission modulator, scheduling mode switcher, and script loading preorderer each correspond to a status word in a fixed area of shared memory. After updating their output status, each executor writes the current status value to the corresponding status word and triggers a hardware event flag after writing, so that other executors can read each status word and update their decisions synchronously in the next ripple spectrum calculation cycle.
[0190] Hardware event flags are changes in the level of general-purpose input / output pins of edge hardware devices or pending bits in the interrupt controller.
[0191] Within the sandbox runtime environment, the protocol identification constraint, sandbox permission modulator, scheduling mode switcher, and script loading preorderer act as four concurrent consumer executors. All of them read the ripple spectrum data structure from a fixed area of shared memory via a read-only access interface. The read-only access is implemented by mapping the shared memory area to a read-only segment in each executor's virtual address space during sandbox initialization (read-only permissions are set via the MMU, or the compiler marks it as a const area on platforms that do not support the MMU). This read-only design prevents any consumer from accidentally or maliciously tampering with the original ripple spectrum data, ensuring that other executors read consistent information.
[0192] The implementation forms of the four executors are as follows:
[0193] The protocol identification constraint is implemented as a periodically triggered software timer callback function, which is triggered once after each water ripple spectrum calculation is completed, with a medium priority.
[0194] The sandbox permission modulator is implemented as a kernel-mode hook function, which is directly called by the ripple spectrum calculation module at the end of each ripple spectrum calculation cycle. It has the highest priority to ensure that permission adjustment takes precedence over other operations.
[0195] The scheduling mode switcher is implemented as a state machine inside the scheduler. It reads water ripple spectrum data at the beginning of each scheduling cycle, with a normal priority.
[0196] The script loading preorderer is implemented as a thread, which is only awakened when a new protocol parsing script needs to be loaded, and has a low priority.
[0197] The protocol identification constraint extracts the difference histogram distribution features and periodicity index from the ripple spectrum data structure, and prunes the candidate protocol space of the pre-trained protocol identification model. The sandbox permission modulator extracts the pulse energy from the ripple spectrum data structure and dynamically adjusts the state of the write operation bit field in the effective permission vector of the sandbox operating environment. The scheduling mode switcher extracts the regularity index from the ripple spectrum data structure and switches between fixed-step scheduling mode and flexible-step scheduling mode. The script loading pre-sorter extracts the difference histogram centroid index, pulse energy, and periodicity index from the ripple spectrum data structure and performs matching deviation sorting on the protocol parsing script entries in the protocol script library.
[0198] Each of the four executors corresponds to a status word in a fixed region of shared memory. The status words are allocated in reserved offsets 19 to 31 bytes, and each status word is an 8-bit unsigned integer. Specifically: offset 19 is the status word for the protocol identification constraint, offset 20 is the status word for the sandbox permission modulator, offset 21 is the status word for the scheduling mode switcher, and offset 22 is the status word for the script loading preorderer. After updating its output status (e.g., whether the pruned candidate space is empty, whether the permission vector has been modified, the current scheduling mode, and whether sorting is complete), each executor immediately writes its current status value to the corresponding status word and triggers a hardware event flag after writing.
[0199] The hardware event flag is implemented as follows: a level change on a general-purpose input / output (GPIO) pin of the edge hardware device, or a pending bit setting in the interrupt controller. This embodiment uses the interrupt controller's pending bit method: after each executor completes a state update, it writes a specific register to the interrupt controller, setting the corresponding event pending bit to 1. Other executors detect state changes by querying these pending bits at the start of the next ripple spectrum calculation cycle.
[0200] Specifically, after each calculation cycle, the ripple spectrum calculation module sequentially reads four status words and calls the update interface of the corresponding executor based on the value of each status word. This design avoids race conditions caused by multiple executors simultaneously modifying shared data because the writing and reading of status words occur in different time windows: status words are written independently by each executor within the cycle, and at the end of the cycle, the coordination module reads them uniformly and distributes the decisions. The data race avoidance mechanism also includes using double buffering for the ripple spectrum data structure itself. The ripple spectrum calculation module writes the latest data to one buffer, while the consuming executor reads the old data from another buffer, achieving lock-free access by exchanging pointers.
[0201] Furthermore, the cloud-based collaboration center encapsulates business control logic into script packages, including:
[0202] The cloud-based collaboration center receives business control logic described by users in natural language, performs semantic parsing on the business control logic, generates corresponding executable script code, and encapsulates it into a script package.
[0203] Users input natural language descriptions through a web interface provided by the cloud-based collaboration center. For example, a user could input, "Read the temperature sensor value every 5 seconds, and turn off the heater if it exceeds 30 degrees Celsius." The input text is limited to 500 characters in length and must be in the form of a declarative or imperative sentence; no programming syntax is required.
[0204] The semantic parsing is implemented as follows: The cloud-based collaboration center calls a pre-trained small language model (e.g., a BERT-based intent recognition and slot filling model) to parse natural language. This model is trained offline using a large amount of labeled business control command corpus, enabling it to extract triggering conditions (e.g., "every 5 seconds" as a periodic trigger, "if it exceeds 30 degrees" as a conditional judgment), the operation object (e.g., "temperature sensor," "heater"), and the action type (e.g., "read," "close"). The parsed output is a structured intermediate representation, specifically an abstract syntax tree in JSON format, where node types include periodic triggers, conditional judges, read device operations, write device operations, etc.
[0205] The mapping rules from semantic parsing results to executable script code are as follows: The cloud-based collaboration center maintains a rule base that maps each node type to a corresponding function call in the target scripting language. This embodiment supports two target languages: Lua (suitable for lightweight embedded environments) and custom bytecode (suitable for devices with extremely limited resources). Mapping rule examples: periodic triggers are mapped to timer creation functions, conditional statements are mapped to if-then statements, and device read operations are mapped to read API calls from the protocol parsing engine. The generated script code includes necessary header file references, a main loop structure, and error handling logic.
[0206] When packaged into a script program package, the cloud collaboration center adds metadata to the generated script code, including: dependency declarations (such as the version of the protocol parsing script to be called and runtime library identifiers), runtime configuration information (such as maximum memory quota, timeout, and priority), and version information (version number and generation timestamp). Then, the entire program package is digitally signed and distributed to edge hardware devices via an encrypted transmission channel.
[0207] As a feasible example, given the input "Read the temperature sensor value every 5 seconds, and turn off the heater if it exceeds 30 degrees," the generated Lua script code snippet after semantic parsing is as follows: A 5-second periodic timer is created. The timer callback function calls the temperature sensor's read interface (via the protocol parsing engine), compares the return value with 30, and if it is greater, calls the heater's shutdown interface. The entire encapsulation process requires no programming skills from the user, lowering the barrier to deploying edge business logic.
[0208] According to another aspect of the embodiments of this application, a smart hardware collaborative integration system based on multimodal protocol adaptation is provided, comprising:
[0209] The environment building unit is used to build a sandbox runtime environment on the edge hardware device, and isolate the script execution space of the sandbox runtime environment from the critical system resources of the edge hardware device;
[0210] The discrimination unit is used to collect the communication data stream feature information of the hardware device to be accessed when the hardware device to be accessed establishes a communication connection with the edge hardware device, and to determine the communication protocol type of the hardware device to be accessed based on the pre-trained protocol recognition model.
[0211] The loading unit is used to dynamically load the protocol parsing script corresponding to the communication protocol type from the protocol script library;
[0212] The transmission unit is used by the cloud-based collaborative center to encapsulate business control logic into script packages and send them to edge hardware devices via an encrypted transmission channel.
[0213] The execution unit is used by edge hardware devices to receive and verify script packages, and load the verified script packages into the sandbox runtime environment for execution.
[0214] The interaction unit is used when the script package is executed in the sandbox environment to perform cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type.
[0215] In one example, such as Figure 5 As shown in the diagram, this embodiment provides a structural diagram of an intelligent hardware collaborative integration system based on multimodal protocol adaptation, including the following levels and interaction processes:
[0216] Multimodal heterogeneous hardware device access layer: Multiple hardware devices using different communication protocols access the edge hardware device via wired or wireless means, serving as data sources or actuators.
[0217] Plug-in Adaptive Protocol Parsing Layer: Deployed within edge hardware devices, this layer performs the following functional sequence: Feature Acquisition: Detects new device access and automatically captures features of the communication data stream (handshake packets, header structure, packet length distribution, timing patterns). Protocol Recognition: Inputs features into a pre-trained deep learning model to determine the protocol type. Script Loading: Retrieves corresponding protocol parsing scripts from the protocol script library or automatically generates them through the protocol description language engine, and dynamically loads them into the parsing engine. Data Standardization: Converts raw heterogeneous protocol data into a unified standard data format using the loaded scripts. Cross-Protocol Conversion: Enables real-time data interoperability and conversion between devices using different protocols based on the unified standard format. Unified Standard Data Format Output:
[0218] Data processed by the parsing layer is presented in a unified format for use by upper-layer business logic or the cloud-based collaboration center. The cloud-based collaboration center is responsible for the orchestration and management of remote dynamic empowerment, including: Business logic orchestration: Encapsulating control algorithms, data preprocessing, edge inference, etc., into microscript packages according to application requirements. Model aggregation and training: Collecting model parameter gradients uploaded from edge devices through federated learning, aggregating them to generate a global model, and distributing updates. Canary release management: Supporting canary release strategies for script packages to gradually promote new features.
[0219] The edge hardware device incorporates two core modules: a sandbox runtime environment, providing an isolated script execution space, coupled with multi-layered security protection (memory isolation, file system isolation, network isolation) and resource configuration management (dynamic allocation and monitoring of CPU / memory / bandwidth); and remote dynamic empowerment, receiving encrypted script packages from the cloud, verifying their signatures, loading them into the sandbox for execution, enabling flexible expansion and collaborative interaction of device functions without firmware upgrades.
[0220] Flexible device functionality expansion and collaborative interaction (no firmware upgrade required):
[0221] Through this architecture, edge hardware devices can dynamically acquire new protocol parsing capabilities, business processing logic, and edge intelligence functions without upgrading firmware, and seamlessly collaborate and interact with heterogeneous devices using different protocols. This enables rapid iteration and intelligent collaboration of IoT systems.
[0222] It should be noted that the embodiments implemented on the intelligent hardware collaborative integration system side based on multimodal protocol adaptation in this application can be referenced with the embodiments implemented on the intelligent hardware collaborative integration method side based on multimodal protocol adaptation, and will not be described in detail here.
[0223] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for collaborative integration of intelligent hardware based on multimodal protocol adaptation, characterized in that, include: A sandbox runtime environment is built on the edge hardware device, and the script execution space of the sandbox runtime environment is isolated from the critical system resources of the edge hardware device; When the hardware device to be connected establishes a communication connection with the edge hardware device, the communication data stream feature information of the hardware device to be connected is collected, and the communication protocol type of the hardware device to be connected is determined based on the pre-trained protocol recognition model. Dynamically load the protocol parsing script corresponding to the communication protocol type from the protocol script library; The cloud-based collaboration center encapsulates business control logic into script packages and distributes them to the edge hardware devices via an encrypted transmission channel; The edge hardware device receives and verifies the script package, and loads the verified script package into the sandbox runtime environment for execution. When the script package is executed within the sandbox environment, it performs cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type.
2. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 1, characterized in that, The collection of communication data stream characteristic information of the hardware device to be accessed includes: On the edge hardware device, the DMA controller's transmission completion interrupt or half-transmission interrupt is used to read the DMA remaining count register value. The remaining count register value is then subtracted from the FIFO depth to obtain the instantaneous water level value of the current receiving FIFO. Read the current count value of the kernel tick counter or debug trace cycle counter of the edge hardware device, divide the current count value by the system clock frequency to obtain the timestamp, and pair the instantaneous water level value with the timestamp and write it into the circular buffer; The circular buffer is a fixed-capacity structure with a capacity of 256 entries. Each entry contains a 32-bit timestamp field and an 8-bit watermark field.
3. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 2, characterized in that, The step of collecting the communication data stream feature information of the hardware device to be accessed also includes: A water ripple spectrum calculation is triggered every 128 new entries accumulated, and the water ripple spectrum calculation includes: Starting from the current write position of the circular buffer, backtrack 128 consecutive entries, perform a first-order difference operation on the water level field of the 128 consecutive backtracked entries to obtain 127 difference values, and trim each difference value to the interval [-8, +7]. The cropped difference values are mapped to a 16-level difference histogram, which is stored in array form. The array element value is the frequency of occurrence of the corresponding difference value among 127 difference values. Calculate the pulse energy, which is the sum of squares of the 127 difference values; Calculate the periodic index; The 16-level difference histogram, the pulse energy, and the periodic index are combined into a water ripple spectrum data structure, which is stored in the memory area of the edge hardware device.
4. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 3, characterized in that, The edge hardware device uses the water ripple spectrum data structure for: Prune the candidate protocol space of the pre-trained protocol recognition model; Dynamically adjust the write operation permission bit in the effective permission vector of the sandbox runtime environment; Switch the scheduler between fixed-step scheduling mode and flexible-step scheduling mode; The protocol parsing script entries in the protocol script library are sorted by matching deviation.
5. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 4, characterized in that, The pre-trained protocol recognition model determines the communication protocol type of the hardware device to be accessed, including: Three sets of weight templates are pre-set in the read-only storage area of the edge hardware device. Each set of weight templates includes 16 8-bit signed fixed-point numbers. The three sets of weight templates correspond to the burst dimension, the regularity dimension, and the sparsity dimension, respectively. The 16-level difference histogram in the water ripple spectrum data structure is subjected to fixed-point inner product operation with the three sets of weight templates to obtain three inner product results; each inner product result is shifted 8 bits to the right and normalized to form a three-dimensional physical layer hidden state vector. The three components of the three-dimensional physical layer hidden state vector are all 8-bit unsigned fixed-point numbers, and the three components of the three-dimensional physical layer hidden state vector respectively represent the burstiness index, regularity index and sparsity index of the communication flow of the hardware device to be accessed. The candidate protocol space is trimmed based on the three components of the hidden state vector of the three-dimensional physical layer; The pre-trained protocol recognition model performs communication protocol type discrimination within the pruned candidate protocol space and outputs the communication protocol type of the hardware device to be connected.
6. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 5, characterized in that, The step of building a sandbox runtime environment on the edge hardware device, and isolating the script execution space of the sandbox runtime environment from the critical system resources of the edge hardware device, includes: The sandbox runtime environment maintains a valid permission vector, and each bit field of the valid permission vector corresponds to a system call type; The sandbox operating environment reads the burst index of the hidden state vector of the three-dimensional physical layer in real time. When the burst index exceeds the preset burst threshold, the sandbox operating environment clears the bit field representing the write operation type in the effective permission vector to zero, so that the sandbox operating environment enters read-only mode. In the read-only mode, write operation system calls to the hardware device to be connected by the script package and the protocol parsing script are blocked. When the suddenness index is lower than the preset recovery threshold for three consecutive water ripple spectrum calculation cycles, the sandbox operating environment will restore the bit field representing the write operation type in the effective permission vector to its original value before entering read-only mode.
7. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 5, characterized in that, When the script package is executed within the sandbox environment, it performs cross-protocol data interaction with the hardware device to be connected by calling the pre-loaded protocol parsing script corresponding to the communication protocol type, including: When the protocol parsing script is loaded, it registers the baseline step time with the scheduler in the sandbox runtime environment; The scheduler maintains a step budget table, which records the current state step budget of each loaded protocol parsing script; The scheduler reads the regularity index of the hidden state vector of the three-dimensional physical layer in real time: When the regularity index exceeds the preset high regularity threshold, the scheduler switches to the fixed step size scheduling mode, which uses the base step size as the minimum scheduling unit. When the regularity index is lower than the preset low regularity threshold, the scheduler switches to the elastic step size scheduling mode.
8. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 4, characterized in that, The step of dynamically loading the protocol parsing script corresponding to the communication protocol type from the protocol script library includes: Each protocol parsing script entry in the protocol script library is associated with a ripple spectrum expectation template. The ripple spectrum expectation template includes the centroid index of the differential histogram of the protocol parsing script under standard communication conditions, the average pulse energy, and the average periodic exponent. The ripple spectrum expectation template is pre-placed in the metadata area of the protocol script library. Before loading the protocol parsing script corresponding to the communication protocol type from the protocol script library, the edge hardware device calculates the matching deviation between the current ripple spectrum and the expected template of the ripple spectrum of each protocol parsing script entry in the protocol script library. The protocol parsing script entries in the protocol script library are sorted from smallest to largest according to the matching deviation. The protocol parsing script entry with the smallest matching deviation is loaded first, and the protocol parsing script entries with a matching deviation greater than a preset exclusion threshold are removed from the candidate list for loading.
9. The intelligent hardware collaborative integration method based on multimodal protocol adaptation according to claim 1, characterized in that, The cloud-based collaboration center encapsulates business control logic into script packages, including: The cloud-based collaboration center receives business control logic described by the user in natural language, performs semantic parsing on the business control logic, generates corresponding executable script code, and encapsulates it into the script package.
10. A smart hardware collaborative integration system based on multimodal protocol adaptation, characterized in that, include: The environment building unit is used to build a sandbox running environment on the edge hardware device, and to isolate the script execution space of the sandbox running environment from the critical system resources of the edge hardware device. The discrimination unit is used to collect the communication data stream feature information of the hardware device to be accessed when the hardware device to be accessed establishes a communication connection with the edge hardware device, and to determine the communication protocol type of the hardware device to be accessed based on a pre-trained protocol recognition model. The loading unit is used to dynamically load the protocol parsing script corresponding to the communication protocol type from the protocol script library; The transmission unit is used by the cloud-based collaborative center to encapsulate business control logic into script packages and send them to the edge hardware devices via an encrypted transmission channel. An execution unit is used for the edge hardware device to receive and verify the script package, and to load the verified script package into the sandbox runtime environment for execution. The interaction unit is used to perform cross-protocol data interaction with the hardware device to be connected by calling the loaded protocol parsing script corresponding to the communication protocol type when the script package is executed in the sandbox runtime environment.
Citation Information
Patent Citations
Transmission method for improving energy efficiency of NOMA cooperative communication system
CN112261662A
System and method for preserving references in sandboxes
US20120311702A1