Public testing environment access method based on heterogeneous interconnection of industrial control system
By combining DPDK and gRPC technologies, the complex problems of data acquisition and transmission of heterogeneous equipment in industrial control systems are solved, and the efficiency, real-time and stability of data flow are achieved, ensuring that the system operates stably and dynamically allocates resources in large-scale testing scenarios.
Patent Information
- Application Number
- CN202510303047.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-30
AI Technical Summary
In industrial control systems, traditional point-to-point connection methods cannot meet complex needs when data from heterogeneous devices are efficiently collected and transmitted, especially when real-time and system flexibility are required.
By combining the high-speed packet processing capabilities of DPDK with the efficient remote process calling capabilities of gRPC, the system achieves close collaboration in all aspects of data acquisition, protocol analysis, data transmission and task distribution to ensure the efficiency, real-time and stability of data flows.
It realizes the efficiency, real-time and stability of data flow, ensuring that the test platform can obtain and analyze data from different devices in real time, and transmit data to multiple test nodes for processing through the gRPC protocol. The system can operate stably in large-scale test scenarios and dynamically allocate resources according to requirements.
Smart Images

Figure FT_1 
Figure QLYQS_1 
Figure QLYQS_2
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heterogeneous interconnection and crowdsourcing testing of industrial control systems, and specifically provides a solution for efficient data acquisition and transmission for crowdsourcing testing tasks. By combining the high-speed packet processing ability of DPDK with the efficient remote procedure call ability of gRPC, the system realizes close cooperation in all aspects of data acquisition, protocol parsing, data transmission, and task distribution, ensuring the efficiency, real-time performance, and stability of the data stream. Background Art
[0002] With the wide application of industrial control systems and automation devices, more and more devices and systems need to be interconnected and interoperated to build a complex and large network architecture. With the explosion of the types and quantities of devices, as well as the diversification of communication protocols, how to manage and maintain these complex systems has become a huge challenge in the modern industrial field. In this context, devices in the industrial control environment not only need to be able to efficiently process massive amounts of data, but also must have high reliability and security to ensure continuous and stable operation in complex and dynamic working environments. With the in-depth promotion of industrial Internet of Things, intelligent manufacturing, and digital transformation, the networking and intelligence levels of industrial control systems are also continuously improving, which not only promotes system innovation, but also poses higher requirements for system security, real-time performance, and efficiency.
[0003] When facing distributed crowdsourcing testing tasks, how to efficiently acquire and transmit data from different industrial control devices has become a major challenge in system design. Crowdsourcing testing tasks often require real-time acquisition of data packets from multiple heterogeneous devices and transmission of these data to the testing platform through a unified interface for subsequent analysis and verification. However, there are significant differences in the protocol stacks, data formats, and communication methods of these devices, resulting in the inability of traditional point-to-point connection methods to meet the increasingly complex requirements. For example, many old devices use dedicated industrial protocols, while modern devices may use more common communication standards. This protocol incompatibility in the heterogeneous environment makes data acquisition, transmission, and processing a huge technical challenge.
[0004] There are a wide variety of devices in industrial control systems, and they involve multi-level real-time data acquisition and processing. These devices include sensors, actuators, PLCs, SCADA, etc. They usually use different communication protocols and have different data formats. In this case, how to ensure that the system can be compatible with and process multiple protocols from different devices simultaneously during high-concurrency data acquisition is a major problem in system design. Traditional point-to-point data acquisition and protocol conversion devices often cannot cope with the challenges of this heterogeneous environment, especially when real-time performance and system flexibility need to be guaranteed, they often face performance bottlenecks and insufficient adaptability. Summary of the Invention
[0005] The object of the present invention is to provide a method for intervening in a crowdsourcing testing environment based on heterogeneous interconnection of industrial control systems. By combining the high-speed data packet processing ability of DPDK with the efficient remote procedure call ability of gRPC, the efficiency, real-time performance, and stability of the data stream are ensured. Through efficient data collection and rapid parsing, the testing platform can obtain and analyze data from different devices in real time, and transmit the data to multiple testing nodes for processing through the gRPC protocol, and can dynamically allocate resources according to requirements.
[0006] The present invention is achieved through the following steps: Step 1: In the initialization stage, deploy corresponding clients for the industrial control devices that need to conduct crowdsourcing testing, ensuring that the clients can communicate with the industrial control devices and can receive and process multi-protocol data packets; Step 2: In the data collection stage, the multi-protocol data packets generated by the industrial control devices are accessed into the system through the network card hardware. To ensure the efficient capture and transmission of data packets, the system assigns the receiving task to a dedicated receiving logic core. This core quickly captures the data packets by polling the network card receiving queue and stores them in the memory pool; Step 3: In the data shunting and parsing stage, the collected data packets are dynamically allocated to different logic cores for parallel processing according to real-time traffic analysis and protocol characteristics. The system introduces intelligent load balancing and parallel parsing optimization technologies to ensure that each core reasonably allocates tasks according to the complexity and priority of the data packets, and dynamically selects the optimal parsing path through real-time learning and optimization; Step 4: In the format conversion stage, based on the characteristics of the data packets collected in real time and combined with a dynamically optimized rule library based on machine learning, the system automatically selects the most suitable parsing rule. And the parsed data packets are encapsulated into a unified metadata format; Step 5: In the data transmission stage, the data is transmitted to multiple testing nodes for processing through the gRPC protocol. Through a dynamic task allocation algorithm, according to the data type, task priority, and resource status of the nodes, the testing tasks are reasonably allocated to ensure the efficient execution of the tasks; Step 6: In the data distribution and assembly stage, each testing node and client obtain the data packets they need from the gRPC transmission queue according to the task allocation result, and reassemble the data packets into complete test data according to the preset assembly rules. Each node screens out the data packets related to itself according to the type of the task and the identification information of the data packets, and assembles the data packets in sequence through information such as timestamps and sequence numbers. Description of the Drawings
[0007] Figure 1It is a flowchart of the crowdsourcing testing environment intervention method based on the heterogeneous interconnection of industrial control systems, showing the collaborative relationship among various links of data collection, protocol parsing, data transmission, and task distribution. Specific implementation
[0008] The crowdsourcing testing environment access method based on the heterogeneous interconnection of industrial control systems described in the embodiments of the present invention combines the high-speed packet processing ability of DPDK and the efficient remote call mechanism of the gRPC protocol. The steps include: 1) Initialization stage: The system first deploys corresponding client programs for the industrial control devices to be accessed. The client can communicate with industrial control devices of different models and protocol types and can receive and process multi-protocol data packets generated by the devices. Specifically, the client automatically loads the corresponding drivers and configuration files according to the type and communication protocol of the device to ensure seamless docking with the device. This stage also includes the setting of the DPDK environment. First, the received data packets are managed through the memory pool (mempool) of DPDK to ensure the high-speed caching and efficient storage of data. The memory pool is used for quickly accessing data packets, reducing the performance bottleneck of the traditional kernel protocol stack and ensuring the efficient capture of data packets. At the same time, the system also configures the receive queue (RX queue) and transmit queue (TX queue) of the network card to adapt to different scales of data transmission requirements. 2) Data collection stage: In the data collection stage, multi-protocol data packets generated by industrial control devices enter the system through network card hardware. To improve data processing efficiency, the system adopts a parallel processing architecture based on DPDK, and allocates multiple different logical cores to process different tasks respectively. First, the system assigns the task of receiving data packets to a dedicated receive logical core. This core polls data packets from the network card receive queue and quickly stores the received data packets into the memory pool to ensure the efficient capture of data. Subsequently, the processing logical core parses the data packets stored in the memory pool. Through a trained deep neural network model, the data packets are classified according to their characteristics (such as source IP, destination IP, protocol type, etc.). Assume that the characteristics of each data packet are X i =(IP src , IP dst , Port src , Port dst , Length, Protocol), and the deep learning model outputs the probability distribution of the data packets belonging to different categories: P(y i |X i ) = σ(W T X i+b), where σ is the sigmoid activation function, W and b are the model weights and biases obtained through training. After classification, the data packets are assigned to different processing queues according to preset rules. Next, the logical core that pulls data packets is responsible for pulling data packets from the test platform. Each pulling core selects the data packets it needs according to specific pulling rules (using identifiers such as netID and serial number). The pulled data packets will continue to undergo protocol parsing and format conversion to ensure that the data can be transmitted in the predetermined format. Finally, the logical core that pushes data packets is responsible for pushing the processed data packets to the test platform. This core pushes data to the target node through the gRPC protocol, and ensures that the data can be correctly distributed and processed according to the task allocation. 3) Data diversion and parsing stage: After data collection, the system will divert the received raw data packets. Through the set diversion rules, the system assigns different types of protocol data packets to different processing cores for parallel processing. Each core processes the corresponding type of data packet according to the preset rules, ensuring that the system can parse multiple data packets in parallel on multiple cores, improving overall processing efficiency. During the data packet parsing process, the system performs hierarchical parsing based on the protocol type of the data packet (such as IPv4, TCP, UDP, etc.). This process uses a deep learning model to dynamically adjust the parsing priority of the data packet. Assume that the system uses the following function to adjust the parsing priority based on the protocol complexity and real-time requirements of the data packet: P(y i )=α·RealTime(y i )+β·Complexity(y i ). Among them, α and β are weight coefficients, RealTime(y i ) represents the real-time requirement of the data packet, Complexity(y i ) represents the parsing complexity of the protocol. 4) Format conversion stage: The parsed data packet will automatically select the appropriate parsing rules for format conversion according to the characteristics of different protocols. To this end, the system maintains a dynamic rule base that can process data packets of multiple protocols in a heterogeneous environment. Suppose the parsing rule base of the data packet is P t , the incremental learning method continuously updates the parsing rule base: t+1 =P t ∪ΔΡ, where ΔΡ represents the incremental parsing rules extracted from the new data packet, P t+1 The updated rule base. All parsed data packets will be encapsulated in a unified metadata format. This standardized format ensures the compatibility and processability of the data, facilitating subsequent transmission, storage, and analysis. Through this format, the system can ensure that data packets of different protocol types can be uniformly processed and managed while ensuring efficiency and stability. 5) Data Transmission Phase: After the data packet undergoes format conversion, the system transmits it to multiple test nodes via the gRPC protocol for processing. The efficient transmission mechanism of gRPC ensures that data can be transmitted to remote nodes quickly and with low latency. During this process, the system, through a dynamic task allocation algorithm, reasonably allocates test tasks based on the data packet type, task priority, and the resource status of the nodes: Among them, Q i represents the processing queue assigned to the i-th data packet, Load(Q j ) is the current load of the j-th queue, and Latency(Q j ) is the latency of the queue. Each test node retrieves the corresponding data packet from the gRPC transmission queue and assembles and tests the data according to the preset task rules. 6) Data Distribution and Assembly Phase: After data transmission, the test nodes retrieve the required data packets from the gRPC transmission queue according to the task allocation results. These data packets are reassembled into complete test data based on identification information such as timestamps and sequence numbers. The system, through a time synchronization mechanism, ensures that the data packets can be assembled in order, avoiding data loss or errors caused by out-of-order transmission. The assembled data packets will be further processed to generate test reports or perform other analysis operations. After data assembly, the test nodes feedback the processing results to the master node via the gRPC protocol. The master node will make dynamic adjustments and resource optimizations based on the feedback results to ensure the efficient operation of the system. 7) Task Execution and Result Feedback: After data transmission and processing, the test nodes execute relevant tests according to the task requirements and return the test results to the master node. The master node makes dynamic adjustments based on the test feedback results to ensure the efficient operation of the entire system. Dynamic adjustments include task reallocation and resource optimization to ensure load balancing and system stability. For example, if the load of a certain test node is too high, the master node will reallocate some tasks to nodes with lighter loads according to the resource status, thus ensuring load balancing of the entire system. 8) System Monitoring and Maintenance: To ensure the stability and efficiency of the system, the system also implements real-time monitoring functions. Through the monitoring module, the system can track the data packet capture rate, task execution efficiency, and node resource usage in real time. If it is found that the data packet loss rate is too high or the task execution efficiency decreases, the system can automatically make adjustments, such as by adjusting the working frequency of the receive logic cores of DPDK, or reallocating tasks to other nodes, to ensure that the system is always in an efficient operating state. By combining the high-speed packet processing ability of DPDK with the efficient remote procedure call ability of gRPC, the present invention ensures the efficiency, real-time performance, and stability of the data stream. Through efficient data collection and rapid parsing, the test platform can obtain and analyze data from different devices in real time, and transmit the data to multiple test nodes for processing via the gRPC protocol. This flexible distributed task execution architecture ensures that the system can operate stably in large-scale test scenarios, and can dynamically allocate resources according to requirements to ensure the rapid execution of tasks and the efficient processing of data.
Claims
1. A new solution for providing efficient data collection and transmission for crowd-testing tasks in industrial control systems, characterized by: The techniques include: Step 1: In the initialization phase, deploy the corresponding client for the industrial control equipment that needs to be tested, ensure that the client can communicate with the industrial control equipment, and can receive and process multi-protocol data packets. During initialization, the client will load the corresponding driver and configuration according to the type and protocol requirements of the industrial control equipment to ensure seamless connection with the equipment. The initialization phase also includes setting up the DPDK (Data Plane Development Kit) environment, creating a memory pool for storing data packets, and configuring the network card's receive and send queues; Step 2: During the data collection phase, the multi-protocol data packets generated by the industrial control equipment are directly received from the network card hardware, and the high-performance data packet processing capabilities of DPDK are used to ensure efficient capture and transmission of data packets. DPDK bypasses the operating system kernel and directly accesses the network card hardware, greatly improving the capture speed and processing efficiency of data packets. Data packets are obtained from the network card's receive queue (RX queue) through a polling mechanism and stored in the memory pool to ensure efficient capture of data packets; Step 3: In the data distribution and analysis stage, the collected data packets are dynamically allocated to different logical cores for parallel processing based on real-time traffic analysis and protocol characteristics. The system introduces intelligent load balancing and parallel analysis optimization technology to ensure that each core reasonably allocates tasks according to the complexity and priority of the data packet, and dynamically selects the optimal analysis path through real-time learning and optimization; Step 4: In the format conversion phase, based on the real-time collected data packet features and combined with the dynamic rule base optimized by machine learning, the system automatically selects the most suitable parsing rules to improve the parsing accuracy and efficiency of the data packet. At the same time, through protocol-specific compression and optimization technology, the data transmission volume is reduced to achieve efficient data conversion. Finally, the parsed data packets are encapsulated into a unified metadata format to ensure compatibility, scalability, and subsequent efficient processing; Step 5. In the data transmission stage, the data is transmitted to multiple test nodes for processing through the gRPC protocol. The dynamic task allocation algorithm is used to reasonably allocate test tasks according to the data type, task priority, and node resource status to ensure efficient execution of tasks. The high efficiency and low latency characteristics of the gRPC protocol ensure the rapid transmission of data in a distributed environment. Each test node will obtain the data packets it needs from the gRPC transmission queue based on the task allocation results, and reassemble the data packets into complete test data according to the preset assembly rules; Step 6: In the data distribution and assembly phase, each test node or client obtains the required data packets from the gRPC transmission queue according to the task allocation results, and reassembles the data packets into complete test data according to the preset assembly rules. Each node screens out the data packets related to itself according to the type of task and the identification information of the data packet, and assembles the data packets in sequence through information such as timestamps or serial numbers to ensure data integrity and consistency. The assembled data packets will be further processed to generate test reports or perform other analysis operations.
2. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 1 is characterized in that: In step 2, multiple logical cores of DPDK are used to process in parallel the tasks of receiving, processing, pulling and pushing data packets from the network card, specifically including: Step 2.1: When receiving data packets, the receiving logic core (RXlcore) of DPDK directly receives data packets from the NIC hardware queue to ensure efficient data packet capture. The receiving logic core obtains data packets from the NIC hardware queue through a polling mechanism and passes the data packets to the processing logic core (Processinglcore) using a lock-free ring buffer (RingBuffer). The system intelligently selects the most appropriate receiving core for processing based on real-time traffic analysis and the protocol characteristics of the data packet. By predicting and analyzing network traffic, the system can dynamically determine which cores are suitable for receiving specific data packets, thereby avoiding excessive concentration of load on a certain core or queue, improving the parallel efficiency of data packet reception, and ensuring load balancing of the receiving cores; Step 2.2: When processing data packets, the received data packets are initially parsed and classified by the processing logic core. For each received data packet, the system extracts its key features, such as protocol type (IPv4, TCP, UDP, etc.), source IP, destination IP, port number, packet length, etc. This can be expressed as: i =(IP src ,IP dst ,Port src ,Port dst ,Length,Protocol), where X i is the feature vector of the ith packet. The processing logic core dynamically adjusts the task allocation according to the complexity of the packet and the preset rules to ensure that different types of packets can be processed in parallel and improve the processing efficiency of the system. According to the prediction results of the deep learning model, the packets will be assigned to different processing queues. For example, TCP packets can be assigned to a dedicated TCP queue, and UDP packets to a UDP queue. Based on the complexity of the protocol, the system will further adjust the priority or resolution path. Assuming there are N processing queues, the system dynamically adjusts the distribution of data packets based on the load of each queue: Among them, Q i Indicates the processing queue assigned to the i-th data packet, Load(Q j ) is the current load of the jth queue, Latency(Q j ) is the queue delay. Peak network traffic may cause some packets to be delayed or lost. The system uses real-time analysis technology combined with deep learning models to dynamically adjust task priorities. High-priority packets (such as real-time TCP traffic) will be processed first. Assume that the priority of each packet is represented by the following function: P(y i )=α·RealTime(y i )+β·Complexity(y i ). Among them, α and β are weight coefficients, RealTime(y i ) represents the real-time requirement of the data packet, Complexity(y i ) represents the parsing complexity of the protocol. At the same time, for packets with complex protocols, the system intelligently optimizes the parsing path and reduces invalid parsing steps, thereby improving the speed and accuracy of protocol parsing; Step 2.3, when pulling data packets from the transit platform, the pull logic core (Pulllcore) pulls data packets in batches from the processing queue, and performs further protocol parsing and format conversion. The pull logic core automatically selects appropriate parsing rules for data decoding in combination with the dynamic rule base to ensure the efficiency of protocol parsing and format conversion of data packets. In this process, the system can dynamically update the parsing rules and data packet decoding paths in real time according to changes in network traffic to ensure the accuracy and processing efficiency of protocol parsing. The system uses incremental learning methods to continuously optimize the parsing rule base. Whenever a new protocol type appears or the network environment changes, the system can automatically update the parsing rules according to the latest data packet characteristics. Assume that the current parsing rule base is P t , then the incremental learning process is: t+1 =P t ∪ΔΡ. Where ΔΡ represents the incremental parsing rules extracted from the new data packet, P t+1 is the updated rule base. In addition, the system uses parallelization to process packets at multiple protocol levels simultaneously. For example, IPv4 and TCP layers can be decoded simultaneously, significantly improving parsing throughput through parallel processing. For each packet, the system uses multi-core processing technology to execute multiple decoding tasks in parallel. Assuming there are M parallel decoding tasks, the decoding time of each task is: in, represents the time of the mth decoding task, where M is the number of parallel tasks; Step 2.4: When pushing data packets to the transit platform, the push logic core (Pushlcore) pushes the processed data packets to the gRPC transmission queue in preparation for remote transmission. The push logic core pushes the processed data packets to the gRPC transmission queue and transmits the data to multiple test nodes through gRPC. The system monitors the resource status and network latency of each test node in real time, intelligently schedules tasks and data packet push paths, and ensures that high-priority data can be transmitted to the appropriate node in a timely manner. Assume that the resource utilization of node j can be expressed as the following vector: R j =(CPU j ,Mem j ,Bw j ), where CPU j Indicates the CPU usage of node j, Mem j Indicates memory usage, Bw j Indicates the network bandwidth utilization. The overall load of the node can be obtained by weighted summation: Load j =w1·CPU j +w2·Mem j +w3·Bw j , where w1, w2, w3 are weight coefficients of various resources, which are set according to the degree of dependence of the task on different resources. Assuming there are multiple test nodes N = {1, 2, ..., n}, the system will calculate the push priority F of each node j j , which combines the node load, network delay and data packet priority. The priority function can be expressed as: Among them, j is the set of data packets assigned to node j.
3. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 1 is characterized in that in step 5, the dynamic task allocation algorithm performs task allocation based on the following factors: (5.1) Data type: According to the protocol type, content and business requirements of the data packet, the corresponding test tasks are assigned. The system not only relies on fixed protocol type classification rules, but also combines real-time traffic analysis and data packet characteristics to dynamically adjust the task allocation strategy; (5.2) Task priority: High-priority tasks are assigned first based on their urgency and importance. Tasks with high real-time requirements will be assigned to nodes with sufficient resources to ensure that critical tasks can be completed in a timely manner; (5.3) Node resource status: According to the resource status of the test node such as CPU, memory and network bandwidth, reasonably allocate tasks to ensure the load balance of the system.
4. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 1 is characterized in that: The method further comprises: Step 6.1, Task execution and result feedback: After receiving the data, the test node executes the corresponding test task and feeds back the test results to the master node through the gRPC protocol. The master node further analyzes and processes the feedback results. The master node will make dynamic adjustments based on the test results, including task reallocation and resource optimization, to ensure the stable operation of the system.
5. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 4 is characterized in that: In step 6.1, the master control node is dynamically adjusted according to the test results, specifically including: (6.1.1) Task reallocation: Dynamically adjust the task allocation strategy based on the test results and node resource status. For example, when the load of a node is too high, the master node will reallocate some tasks to other idle nodes to ensure the load balance of the system; (6.1.2) Resource optimization: Dynamically adjust the resource allocation of nodes according to the load of the test nodes to ensure the stable operation of the system. For example, when the CPU usage of a node is too high, the master node will adjust the task allocation of the node to reduce its load.
6. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 1 is characterized in that: The method further comprises: Step 6.2: System monitoring and maintenance: By real-time monitoring of the system's operating status, abnormal situations can be discovered and handled in a timely manner to ensure the system's efficiency and stability. System monitoring includes packet capture rate monitoring, task execution efficiency monitoring, and node resource usage monitoring.
7. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 6 is characterized in that: In step 6.2, system monitoring includes: (6.2.1) Packet capture rate monitoring: Real-time monitoring of the packet capture rate to ensure efficient packet capture. If the packet loss rate is found to be too high, the system will automatically adjust the operating frequency of the DPDK receiving logic core to ensure complete capture of the packet; (6.2.2) Task execution efficiency monitoring: Real-time monitoring of task execution efficiency to ensure efficient execution of tasks. If a task is found to take too long to execute, the system will reallocate the task to other nodes to ensure timely completion of the task; (6.2.3) Node resource usage monitoring: Real-time monitoring of the resource usage of the test nodes to ensure reasonable allocation of resources. If the resource usage of a node is found to be too high, the system will dynamically adjust the task allocation of the node to ensure system load balancing.
8. The method for crowd-testing environment intervention based on heterogeneous interconnection of industrial control systems according to claim 1 is characterized in that: The method is applicable to the interconnection and crowd-testing tasks of various heterogeneous devices in industrial control systems, ensuring that the system can operate stably in large-scale test scenarios and can dynamically allocate resources according to demand, ensuring rapid execution of tasks and efficient processing of data.
Citation Information
Cited By
Unified data metadata conversion system
CN121078139A