Large-scale industrial process data aggregation method and device and computing equipment

Through multi-threading and asynchronous I/O architecture combined with Apache Arrow columnar data storage format, efficient aggregation and processing of large-scale industrial process data is achieved, solving the problems of large resource occupation and low processing efficiency, and is suitable for complex industrial big data computing scenarios.

CN120029758APending Publication Date: 2025-05-23SHANXI TAIGANG ENG TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411925175.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems of huge resource occupation and inefficient processing in real-time collection and efficient and low-latency processing of large-scale industrial process data.

Method used

The industrial process data acquisition connection is established through multi-threading methods, data is read through asynchronous I/O architecture, and zero-copy aggregation is performed in the columnar data cache area based on the Apache Arrow open protocol, data is pushed according to subscription information and respond to end users' needs.

Benefits of technology

It effectively solves the problems of large resource occupation and low processing efficiency, reduces data copying and conversion processes, improves the processing efficiency of large-scale industrial process data, and is suitable for industrial big data computing scenarios of different scales and complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029758A_ABST
    Figure CN120029758A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale industrial process data aggregation method and device and computing equipment, and relates to the field of industrial process big data processing.The method comprises the steps that industrial process data collection connection is established in a multi-thread mode, and original data of large-scale industrial process data is obtained; reading original data of the obtained large-scale industrial process data through an asynchronous I / O architecture, and putting the original data into a data cache region after processing; and pushing data to the data application according to the subscription information, and responding to a terminal user demand through the data application program. According to the method, multiple threads and an iouring-based asynchronous I / O architecture are adopted to realize the collection, caching and application processes of industrial process data, a data cache region is realized based on a column memory data format of an Apache Arrow open protocol, and zero-copy sharing of data applications is realized through a standardized data cache region. The problems of huge resource occupation and low processing efficiency in data acquisition, processing, multi-protocol subscription and online analysis in the current large-scale industrial process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial process big data processing, and provides an efficient method, device and computing equipment suitable for large-scale industrial process data aggregation. Background Art

[0002] As the application research of AI technology in the field of industrial production gradually deepens, AI-enabled, data-driven control algorithms have brought more sophisticated and intelligent control methods to industrial automation. Without exception, these applications are based on large-scale, real-time industrial process data. Therefore, the real-time collection and efficient, low-latency processing of large-scale industrial process data are particularly important and urgent.

[0003] Therefore, the real-time collection and efficient low-latency processing of large-scale industrial process data are particularly important and urgent. At present, the acquisition of industrial process data usually adopts the Internet of Things gateway or data communication programs based on multiple industrial Ethernet protocols. The heterogeneous data transmitted by these different devices (the so-called heterogeneous data refers to the data provided by different data acquisition terminals that are not the same in structure) needs to be processed and converted before being forwarded to message queues such as Kafka, and finally processed by big data analysis platforms such as Spark / Flink. In the middle, there are a large number of data copying and conversion operations, involving the connection of multiple complex links, which will lead to huge resource occupation and low processing efficiency. Summary of the invention

[0004] In order to solve some or all of the technical problems existing in the above-mentioned prior art, the present invention provides a large-scale industrial process data aggregation method, device and computing equipment, which can solve the problems of huge resource occupation and low processing efficiency in the current large-scale industrial process data collection, processing and online analysis.

[0005] The technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention provides a large-scale industrial process data aggregation method, comprising:

[0007] Establish industrial process data acquisition connections in a multi-threaded manner to obtain raw data of large-scale industrial process data;

[0008] The raw data of large-scale industrial process data is read through the asynchronous I / O architecture and placed in the data cache after processing;

[0009] Push data to data applications based on subscription information and respond to end-user needs through data applications.

[0010] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, establishing an industrial process data acquisition connection in a multi-threaded manner includes:

[0011] The main thread is defined as a management program thread, which is used to configure the data collection program, data cache area and data application program. Each data collection program and data application program sets an independent running thread. The collection program and application program realize zero-copy aggregation processing of data through a shared columnar data cache area based on the Apache Arrow open protocol.

[0012] The multi-threading method uses the POSIX thread library in the Linux system to create and manage threads.

[0013] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, the data acquisition program is configured as an acquisition connection; the acquisition connection includes an array consisting of several data acquisition points, and the arrays of several data acquisition points form an acquisition channel. The data acquisition program collects data information using the acquisition channel as the minimum execution unit. One acquisition connection corresponds to an independently running thread in the multi-threading, and each acquisition channel under the same acquisition connection is scheduled in the thread based on the asynchronous I / O mechanism.

[0014] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, reading the raw data of the large-scale industrial process data obtained through the asynchronous I / O architecture includes:

[0015] After receiving the collection instruction of large-scale industrial data, the management program dynamically creates a running thread for this collection connection, registers an asynchronous I / O instance in this thread, then creates a loop body in the thread, and calls the asynchronous I / O waiting function in the loop body to detect whether there is an I / O event completed. Each collection channel will register a collection event timer in the collection connection thread. The timing time is set according to the collection frequency of the collection channel. When the timer time is reached, the data collection event is triggered to submit an I / O request to the asynchronous I / O instance. The loop body waiting function in the collection connection thread detects the occurrence of the I / O event, and obtains the original collection data through the asynchronous I / O event result acquisition function.

[0016] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, the data buffer includes at least the following fields: name, data type, value, timestamp, value change flag, value valid flag and metadata, the data cache area is based on the Apache Arrow open protocol format, and the data is stored in columnar format, wherein the data cache area is divided into different Arenas, one acquisition connection corresponds to one Arena, each Arena corresponds to a different acquisition channel, and each Arena is divided into different Tables, and the management program thread initializes the space occupied by Arena and Table according to the number of acquisition points of each acquisition channel and the data application program configuration;

[0017] The acquisition program stores the processed data in a columnar data buffer based on the Apache Arrow open protocol format. The process of reading and storing data in the data buffer includes the following operations:

[0018] Determine whether the acquisition channel reading is valid. If invalid, set an invalid flag for each data point in the buffer;

[0019] Iterate the read data set and convert each data into the data type supported by the Apache Arrow protocol;

[0020] Store the converted data into the data buffer and set the value change flag.

[0021] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, the data application is configured as an application service, and multiple application services can be independently running at the same time, and each application service subscribes to the data content in the data cache area in units of Table;

[0022] The subscription methods of the application service to the data content in the data cache area include:

[0023] PERIODIC mode, which includes a periodic mode, and is used for application services to obtain data for processing by subscribing to messages according to a configured time period;

[0024] NOTIFICATION mode, the NOTIFICATION mode includes a notification mode, which is used for the application service to obtain data through subscription messages when the subscribed data changes;

[0025] SCHEDULE mode, the SCHEDULE mode includes a plan mode, which is used for the application service to obtain data through subscription messages according to the configured acquisition plan;

[0026] DIRECT mode, the DIRECT mode includes a direct memory analysis mode, which is used for application services to directly perform analysis operations on data in the data cache area through the Apache Arrow agent.

[0027] Furthermore, in the above-mentioned large-scale industrial process data aggregation method, the application service is also included to filter the data in the data cache area, and the filtering method includes:

[0028] ALL filtering, wherein the ALL filtering includes subscribing to all data;

[0029] INCLUSIVE filtering, wherein the INCLUSIVE filtering includes subscribing to all selected data;

[0030] EXCLUSIVE filtering, wherein the EXCLUSIVE filtering includes subscribing to all non-selected data;

[0031] Among them, INCLUSIVE filtering and EXCLUSIVE filtering both filter the subscription content through a filter table.

[0032] Further, in the above-mentioned large-scale industrial process data aggregation method, pushing data to the data application according to the subscription information, and responding to the end user's needs through the data application program includes:

[0033] The data collection program processes the raw data and stores it in a columnar data cache based on the Apache Arrow open protocol;

[0034] The management program queries the subscriber information of the current Table;

[0035] Confirm the event notification function registered by the data application that subscribes to this Table;

[0036] Call the event notification function to notify the data application;

[0037] After receiving the subscription message, the data application operates the subscription data according to the subscription method and filtering method.

[0038] In a second aspect, the present invention provides a large-scale industrial process data aggregation device, comprising:

[0039] A management unit, which is used to configure basic information, provide a user interface, create and manage data acquisition connections, create and manage data buffers, and create and manage data application services;

[0040] A data cache unit, which is used to provide a columnar data cache area for the collected data based on the Apache Arrow open protocol, and configure the size and performance parameters of the cache area physical space Arena and Table;

[0041] A data collection unit, which is used to configure data collection connections, collection channels, manage connection properties, and maintain data collection point information;

[0042] The data application unit is used to configure and respond to data application services that meet the requirements of terminal data users, manage application service attributes, and set application service subscription information.

[0043] In a third aspect, the present invention provides a computing device, the computing device comprising a processor and a memory, wherein:

[0044] The memory is used to store program code and transmit the program code to the processor;

[0045] The processor is used to execute the large-scale industrial process data aggregation method as described above according to the instructions in the program code.

[0046] The main advantages of the technical solution of the present invention are as follows:

[0047] The large-scale industrial process data aggregation method of the present invention solves the problems of huge resource usage and low processing efficiency in the current large-scale industrial process data collection, processing, multi-protocol subscription and online analysis by adopting a multi-threaded and asynchronous I / O architecture and a columnar data storage format based on the Apache Arrow open protocol, reduces the data copying and conversion process, and provides the possibility for direct memory analysis of large-scale process data, which can adapt to industrial big data computing scenarios of different scales and complexities. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0049] Figure 1 A flowchart of a method for large-scale industrial process data aggregation provided in an embodiment of the present application.

[0050] Figure 2 A schematic diagram of an exemplary system architecture of a method and processing device for large-scale industrial process data aggregation provided in an embodiment of the present application.

[0051] Figure 3 A structural and functional diagram of a columnar data cache unit based on the Apache Arrow open protocol in a method for large-scale industrial process data aggregation provided in an embodiment of the present application.

[0052] Figure 4 A data storage format based on the Apache Arrow open protocol in a method for large-scale industrial process data aggregation provided in an embodiment of the present application.

[0053] Figure 5 A schematic diagram of an architecture design using multi-threading and asynchronous I / O in a method for large-scale industrial process data aggregation provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0055] The following is combined with Figure 1-5 , describes in detail the technical solution provided by the embodiments of the present invention.

[0056] First, in order to more clearly understand and explain the technical solutions provided by the embodiments of the present invention, the technical terms involved in the embodiments of the present invention are further explained below. Specifically, the technical terms involved in the present invention include:

[0057] Multithreading: In modern operating systems, a process is the smallest unit of system resource allocation, while a thread is the smallest unit of CPU scheduling. Multithreading is a form of parallel execution of computing tasks. By creating multiple threads in a process, these threads can be executed concurrently, thereby improving the execution efficiency of tasks. Threads under the same process share memory space, and the data of one thread can be directly used by other threads. Therefore, multithreading can reduce the cost of system data exchange.

[0058] Asynchronous I / O: In modern computer systems, the CPU performs calculations very quickly. However, when it comes to I / O operations, such as reading and writing files and sending network data, it is necessary to wait for the current I / O operation to complete before proceeding to the next step. Asynchronous I / O provides an efficient processing mechanism. When a time-consuming I / O operation needs to be performed, the CPU only needs to issue an I / O instruction, and can handle other tasks after issuing the instruction. After a period of time, when the I / O operation is completed and the result is returned, the CPU is notified to process subsequent tasks. This processing mechanism can greatly improve the processing efficiency for I / O-intensive tasks.

[0059] Apache Arrow: The Apache Arrow open protocol provides an in-memory columnar data storage format standard suitable for heterogeneous big data systems. On the one hand, it contains a series of technical components that enable data to be quickly processed and moved between heterogeneous big data systems. On the other hand, the Apache Arrow open protocol specifies a standardized columnar data format that is independent of development languages ​​for two-dimensional and multi-dimensional data, making efficient direct in-memory data analysis possible.

[0060] Linux: Linux is a UNIX-like operating system that is free to use and distribute. It is a multi-user, multi-tasking, multi-threaded, and multi-CPU operating system based on POSIX. For example, it supports 32-bit and 64-bit hardware.

[0061] io_uring: io_uring is a high-performance asynchronous I / O framework first introduced in the Linux 5.1 kernel in 2019, which can significantly accelerate the performance of I / O-intensive applications.

[0062] UI: User Interface, the user interface, abbreviated as UI, is the medium for interaction and information exchange between the hardware system or software system and the user, which realizes the conversion between the internal form of information and the form acceptable to humans.

[0063] Electron: Electron is a cross-platform desktop application framework built using JavaScript, HTML, and CSS. Electron is compatible with Mac, Windows, and Linux operating systems and can build applications for all three platforms.

[0064] As Figure 1 shown, the present invention discloses a method for aggregating large-scale industrial process data, which includes the following steps S1 - S3:

[0065] Step S1: Establish an industrial process data acquisition connection in a multi-threaded manner to obtain the raw data of large-scale industrial process data;

[0066] Step S2: Read the raw data of the large-scale industrial process data obtained through an asynchronous I / O architecture, and put it into the data buffer after processing;

[0067] Step S3: Push data to the data application according to the subscription information, and respond to the terminal user's needs through the data application program.

[0068] Specifically, the following will combine the attached Figure 1 - attached Figure 5 , and detail how to aggregate large-scale industrial process data according to the above steps S1 - S3 in the present invention, and how to implement the technical solutions provided by the embodiments.

[0069] It should be noted that in the present invention, the aggregation in a method, device, and computing device for aggregating large-scale industrial process data refers to a single or overall operation of collecting industrial process data, converting the collected industrial data, and analyzing, processing, forwarding, and persisting the converted industrial data.

[0070] Considering that large-scale industrial process data collection and processing is an I / O intensive task, in order to minimize the large amount of inefficient data copying and data conversion processes that may be caused by large-scale data processing, this embodiment adopts multi-threading and a message mechanism based on asynchronous I / O to achieve. Among them, this embodiment uses the POSIX thread library (pthread library) in the Linux system to create and manage threads, and uses the io_uring asynchronous component to implement the asynchronous I / O mechanism.

[0071] The main thread is defined as a management program thread, which is automatically created when the example program of this embodiment is running and is responsible for configuring the data acquisition program, data cache area and data application program; the data acquisition program and the data application program are respectively set up with independent running threads, which are created by the management program thread during the running process according to the configuration information. The data acquisition program and the data application program realize zero-copy aggregation processing of large amounts of data through a shared columnar data cache area based on the Apache Arrow open protocol.

[0072] In one possible design, the method and device are configured with a UI interface in this embodiment, developed using Electron, and can run on Linux or Windows operating systems. The end user can interact with the method manager instance through the UI interface and then operate the process described in the method.

[0073] In order to make the description of the technical solution of the present invention more clear and complete, the technical solution of the present invention is described in detail below in combination with steps S01 to S08.

[0074] Combination Figure 1 , is a flow chart of a large-scale industrial process data aggregation method shown in an embodiment of the present application, and the method is executed by an electronic device. The method includes the following steps:

[0075] Step S01, configuration management program, input configurable parameters.

[0076] This method instance has the attributes of a software product. The configuration parameters of the management program can be provided in the form of a configuration file, that is, the node name, IP address, node description, optional data acquisition program driver component, data application program service component and parameter configuration of these components of the method instance can be pre-set in the configuration file. When the method instance is running, these parameters can be set through the UI interface.

[0077] In a possible design, the UI interface developed by the Electron is configured with a management page, and the above parameters can be configured in the management page. After configuration, the relevant configuration parameters can be changed through the save button or other methods with saving significance.

[0078] Step S02, configure the data collection program and set the information of the points to be collected.

[0079] The configuration of the data acquisition program and the setting of the information of the points to be collected include: a data acquisition program is used to collect data information of a specific device. At the same time, multiple data acquisition programs can be run simultaneously for multiple devices. The method defines a collection program as a collection connection, and a collection connection includes an array composed of several data collection points. The method also defines the array of collection points in the above-mentioned collection connection as a collection channel. When the data acquisition program is running, the collection channel is used as the minimum execution unit to collect data information; the collection channel of the collection connection is set by the pre-configured channel parameters. These parameters include two aspects: first, communication parameters, that is, the configuration parameters and channel name and collection frequency of the communication protocol used by the corresponding collection device; second, collection point attribute parameters, including the name, address, data type, etc. of the collection point. The data acquisition program configuration also includes the following processes: establishing a data acquisition connection, selecting the collection connection communication protocol, setting protocol parameters, such as IP address, port number, etc.; establishing a data acquisition channel, setting the collection frequency; establishing a collection point in the data acquisition channel, and setting the collection point attribute information.

[0080] In a possible design, the UI interface developed by Electron is configured with a data acquisition page. Taking the acquisition object as an S7-1500 series PLC as an example, the acquisition points need to be collected: MW100, MD102, and M104.0. First, through the data acquisition page in the UI interface developed by Electron, establish the acquisition connection: conn_1500, select the communication protocol: s7; set the device IP address: 192.168.0.1; then establish the acquisition channel: channel_1500, set the acquisition frequency: 10Hz, and establish the acquisition points: Tag1 [address: mw100, data type: int, metadata: Kw], Tag2 [address: md102, data type: float, metadata: A], Tag3 [address: m104.0, data type: bool, metadata: none]. After the settings are completed, the configured parameters can be changed through the Save button or other meaningful methods.

[0081] Step S03, perform data cache settings and initialize the data cache area.

[0082] The data cache is set up and the data cache area is initialized, including that the data cache area is based on the Apache Arrow open protocol format and data is stored in columnar format. Figure 3As shown, in order to improve the memory I / O efficiency, the cache area is divided into different Arenas, one acquisition connection corresponds to one Arena, and each Arena corresponds to different acquisition channels, which are divided into different Tables. The management program initializes the space occupied by Arena and Table according to the number of acquisition points of each acquisition channel and the data application configuration. Considering memory alignment, the size of Arena and Table is configured to an integer power of 2 in bytes. The data cache area is set after the acquisition program configuration is completed. It is necessary to set Arena for each acquisition connection, and then set Table for each acquisition channel under this acquisition connection.

[0083] In a possible design, the UI interface developed by Electron is configured with a data cache page. In the page, a new Arena is configured, the name is set to Arena_1500, and the corresponding connection is selected as conn_1500; a new Table is created, the name is set to Arena_1500_Table_1, and the corresponding collection channel is selected as channel_1500;

[0084] Step S04, configuring the data application and setting the application subscription content.

[0085] The configuration data application sets the subscription content for the application, including defining a data application as an application service in the embodiment of the method, and multiple application services can run independently at the same time. Each application service subscribes to the data content in the data cache area in units of Table, and there are 4 subscription methods: PERIODIC, NOTIFICATION, SCHEDULE, and DIRECT. Among them, the PERIODIC method is a periodic mode, and the application service obtains data through subscription messages for processing according to the configured time period; the NOTIFICATION method is a notification mode, that is, the application service obtains data through subscription messages when the subscribed data changes; SCHEDULE is a plan mode, that is, the application service obtains data through subscription messages according to the configured acquisition plan; DIRECT is a direct memory analysis mode, that is, the application service directly analyzes and calculates the data in the data cache area through the Apache Arrow agent. There are three filtering methods for each application service to subscribe to the content in the data cache area: ALL, INCLUSIVE, and EXCLUSIVE. Among them, ALL is to subscribe to all data; INCLUSIVE is to subscribe to all selected data; EXCLUSIVE is to subscribe to all non-selected data. Both INCLUSIVE and EXCLUSIVE filter the subscription content through a filter table. The data application configuration also includes the following process: select the application service type, including but not limited to OPC UA server, MQTT protocol, Websocket protocol, ArrowDirect R / W agent; set application service parameters; select subscription method and set subscription method parameters; select subscription content and set filtering method; set application service operation mode. Among them, the application service operation mode includes manual and automatic modes. No matter which mode is selected, the application service needs to detect whether the collection connection corresponding to the subscribed content is running. When the collection connection is running, the application service is actually put into operation.

[0086] In a possible design, the UI interface developed by Electron is configured with a data application page. A new data application, app_1500_opcua, is created on the page. The application type is OPCUA. The subscription method is PERIODIC. The subscription period is set to 100ms. The subscription content is ALL.

[0087] Step S05, receiving a data collection instruction and starting a corresponding data collection program.

[0088] The receiving of data acquisition instructions and starting of the corresponding data acquisition program include: the data collection process of the embodiment of the method is implemented through an asynchronous I / O mechanism. After receiving the acquisition instruction, the management program dynamically creates a running thread for this acquisition connection, and registers an asynchronous I / O instance in this thread, then creates a loop body in the thread, and calls the waiting function of the asynchronous I / O in the loop body to detect whether there is an I / O event completed. Each acquisition channel will register an acquisition event timer in the acquisition connection thread, and the timing time is set according to the acquisition frequency of the channel. When the timer time arrives, the data acquisition event is triggered to submit an I / O request to the asynchronous I / O instance. The loop body waiting function in the acquisition connection thread detects the occurrence of the I / O event, and the original acquisition data is obtained through the event result acquisition function of the asynchronous I / O. Among them, the management program creates a connection thread through the pthread_create function in the POSIX thread library (pthread library). When the acquisition connection needs to be stopped, the pthread_join function is called to wait for the thread to end.

[0089] Among them, this embodiment uses the io_uring asynchronous I / O component in the Linux operating system kernel to implement the collection, processing, and subscription processes, which are specifically implemented as follows: After the collection connection thread is created, the asynchronous I / O and io_uring structure instances are created through the io_uring_create function, and the io_uring_queue_init function is called to initialize the io_uring structure, and then the IO event environment is configured through the io_uring_setup function and the io_uring_register function. When the registered timer in the thread arrives, the I / O request is submitted to io_uring through the io_uring_submit function. Then, in the thread loop body, the io_uring_wait function is called to detect whether an I / O event occurs. If so, the original data collected by the I / O event is obtained through the io_uring_getevents function.

[0090] In a possible design, the UI interface developed by Electron is configured with a data acquisition page, in which a data acquisition connection can be selected to start or stop. In the page, the running status of the data acquisition connection is marked by color changes or text descriptions.

[0091] Step S06, read the original collected data, and put it into the data buffer after processing.

[0092] The reading of the original collected data and the processing and placing it in the data buffer include, when the asynchronous I / O obtains the original collected data, performing numerical processing and format conversion on it, and storing the processed data in a data buffer based on the Apache Arrow open protocol format, and the data buffer content includes the following fields: name, value, timestamp, value change flag, value valid flag, metadata. The reading and storage process also includes the following operations: judging whether the acquisition channel reading is valid, if not, setting an invalid flag for each data point in the buffer; iterating the read data set and converting each data therein into a data type supported by the Apache Arrow protocol; and storing the converted data in the data buffer at the same time, setting a value change flag.

[0093] like Figure 4 As shown in the figure, the data in the cache is stored in a columnar memory format based on the Apache Arrow open protocol. In the figure, Name represents the name field of the collection point, Value represents the current value field after the collection point is processed, and Changed represents the value change flag field that indicates whether the current value of the collection point has changed compared to the previous moment.

[0094] In a possible design, the UI interface developed by Electron is configured with a data monitoring page, through which the data status in the Arena / Table corresponding to each acquisition channel in each acquisition connection in the current data cache area can be viewed.

[0095] Step S07: Push data to the data application according to the subscription information.

[0096] The method of pushing data to the data application according to the subscription information includes that after the data acquisition program obtains the original data and stores it in the data cache area, the management program notifies the data application through the event notification function registered by the data application, and the data application operates the subscribed data according to the subscription method and the screening method after receiving the subscription message. In this process, the message body in the subscription message carries the address index of the collected data in the data area based on the Apache Arrow open protocol format, and does not involve data copying. This method greatly improves the efficiency of data transmission.

[0097] Take the data application service as an example, select OPCUA and name app_1500_opcua. When selecting the Table in the subscription data cache area, the positioning method is PERIODIC, and the subscription content is ALL, then when the management program notifies the data application through the event notification function registered by the data application, the application service app_1500_opcua indexes the subscription data through the Apache Arrow memory area offset address and updates the subscription data to OPCUA.

[0098] In a possible design, the UI interface developed by Electron is configured with a data application page, through which the subscription data update mark and update timestamp can be viewed.

[0099] Step S08: the data application program responds to the terminal data user's demand.

[0100] The data application program responds to the needs of terminal data users, including receiving terminal user instructions by the management program, forwarding them to the corresponding data application service for processing according to the instruction type, and the data application service responding to the terminal data user needs to determine whether the data is written to the disk through the SQL application service, or read and write data by the OPC UA application service protocol, or directly performing memory data analysis by the Apache Arrow agent.

[0101] In a possible design, the UI interface developed by Electron is configured with a data application page, through which the data application service response terminal data user status can be viewed, including whether the response is successful, the response timestamp, etc.

[0102] Second, Figure 2 is a structural diagram of a large-scale industrial process data aggregation device according to an exemplary embodiment. Figure 2 The device includes: a management unit, a data cache unit, a data acquisition unit and a data application unit, wherein:

[0103] The management unit is used to configure basic information, provide user interfaces, create and manage data collection connections, create and manage data caches, and create and manage data application services; the data cache unit is used to provide a columnar data cache based on the Apache Arrow open protocol for the collected data, and configure the size and performance parameters of the cache physical space Arena and Table; the data collection unit is used to configure data collection connections and collection channels, manage connection properties, and maintain data collection point information. The data collection unit can configure multiple data collection connections at the same time, such as Figure 2 As shown, there are collection connections 1, ..., collection connection N; each collection connection can have multiple collection channels, and in the figure, collection connection 1 and collection connection N respectively establish three collection channels: collection channel 1, collection channel 2, and collection channel 3; the data application unit is used to configure and respond to data application services that meet the requirements of terminal data users, manage application service attributes, and set application service subscription information. The data application unit can configure multiple data application services at the same time, such as Figure 2 As shown, application service 1, application service 2, and application service N are configured.

[0104] like Figure 5As shown, the logical relationship between the various units of the above-mentioned device is as follows: the main thread is the management program thread, which is automatically created when the instance program of this embodiment is running, and is responsible for configuring the data acquisition program, data cache area and data application program; the data acquisition program and the data application program are respectively set up with independent running threads, which are created by the management program thread during the running process according to the configuration information. Figure 5 In the example, two data acquisition threads are created. Taking data thread 2 as an example, after creation, the io_uring structure will be created, initialized and registered in the thread for asynchronous I / O processing. After completion, the timer will be registered according to the acquisition frequency of the acquisition channel. The three timers in the figure correspond to the three acquisition channels respectively. After running, these timers will submit I / O requests to the io_uring structure diagram. In the loop body of the acquisition thread, the original acquisition data is obtained through the io_uring result acquisition function; similarly, Figure 5 In the example, four application threads are created. From the logical relationship in the figure, it can be seen that after the collection thread 2 obtains the data, it will query the data subscription application, and then the notification function of the relevant data application service will push the data address in the data cache to the subscription application, which is application thread 1 and application thread 2 in the figure. The data collection program and the data application program realize zero-copy aggregation processing of large amounts of data through a shared columnar data cache based on the Apache Arrow open protocol.

[0105] Therefore, through the above-mentioned device of the present invention, the collection, conversion and analysis of industrial process data can be realized.

[0106] In a third aspect, the present invention also provides a large-scale industrial process data computing device, in which the above-mentioned large-scale industrial process data aggregation device is provided, and the steps of the above-mentioned method can be implemented or executed by the device, thereby realizing the aggregation processing of industrial data.

[0107] As an achievable method of this embodiment, taking the data aggregation of an automobile production line as an example, when the data acquired by the automobile production line is aggregated, a multi-threaded and asynchronous I / O architecture is used to collect the automobile production line data. The multi-threaded system includes at least one main thread and multiple other threads. The main thread is used to configure the data acquisition program, data cache area and data application program for the automobile production line data aggregation. Each data acquisition program and data application program is set to an independent running thread, and one acquisition program is configured to correspond to one acquisition connection. One acquisition connection includes an array composed of several data acquisition points. An array composed of one data acquisition point is set to an acquisition channel corresponding to an independently running thread in the multi-threaded system, so that when the data acquisition program is running, the acquisition channel is used as the minimum execution unit to collect the automobile production line data, and the main thread collects the original data of the automobile production line in the form of multiple channels; so as to realize the quality control of the automobile and the monitoring of the production line quality.

[0108] As another achievable method of this embodiment, taking the data aggregation of the power system as an example, when the data of the power system is aggregated, the power system data is collected using a multi-threaded and asynchronous I / O architecture, and the multi-threads include at least one main thread and multiple other threads. The main thread is used to configure the data acquisition program, data cache area and data application program of the power system data aggregation, wherein each data acquisition program and data application program is set to an independent running thread, and one acquisition program is configured to correspond to one acquisition connection, and one acquisition connection includes an array composed of several data acquisition points, and an array composed of one data acquisition point is set to an acquisition channel corresponding to an independently running thread in the multi-threads, so that when the data acquisition program is running, the acquisition channel is used as the minimum execution unit to collect the power system data, and the main thread collects the original data of the power system in the form of multiple channels; in this way, through the aggregation processing of the power system data, the analysis of the power system data is realized, and the supervision of the safety of the power system is realized. It can better distribute and regulate electricity and reduce the waste of costs.

[0109] As another achievable method of this embodiment, taking the data aggregation of aircraft engines as an example, when the data of the engine is aggregated, the engine data is collected using a multi-threaded and asynchronous I / O architecture, and the multi-threads include at least one main thread and multiple other threads, and the main thread is used to configure the data acquisition program, data cache area and data application program of the engine data aggregation, wherein each data acquisition program and data application program is set to an independent running thread, and one acquisition program is configured to correspond to one acquisition connection, and one acquisition connection includes an array composed of several data acquisition points, and an array composed of one data acquisition point is set to a collection channel corresponding to an independently running thread in the multi-threads, so that when the data acquisition program is running, the acquisition channel is used as the minimum execution unit to collect engine data, and the main thread collects the original data of the engine in the form of multiple channels; by collecting, converting and analyzing the data of the aircraft engine, the quality of the engine and the operation and service life of the aircraft engine can be detected, so as to ensure the flight safety of the aircraft, and ensure the safety of the aircraft and the pilot. The quality and life of the aircraft engine can be improved, the failure rate can be reduced, the maintenance cost can be reduced, and the safety of the transmitter and its aircraft can be improved.

[0110] The device authentication method provided in the embodiment of the present invention can be executed by one or more programs / softwares, and the program / software can be provided by the network side. The terminal device and the Internet of Things device mentioned in the above embodiments can download the required corresponding program / software to the local non-volatile storage medium, and when it needs to execute the above device authentication method, the program / software is read into the memory by the CPU, and then the CPU executes the program / software to implement the device authentication method provided in the above embodiment. The execution process can refer to the above Figures 1 to 3 Instructions in .

[0111] The above-described apparatus / equipment, and the module / unit embodiments corresponding to or included in the apparatus / equipment are merely illustrative, wherein the units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art may understand and implement the present embodiment without creative effort.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0113] It should be understood that the specific features, operations and details described hereinabove about the method of the present invention may also be similarly applied to the device and system of the present invention, or, vice versa. In addition, each step of the method of the present invention described above may be performed by the corresponding parts or units of the device or system of the present invention.

[0114] It should be understood that each module / unit of the device of the present invention can be implemented in whole or in part by software, hardware, firmware or a combination thereof. Each module / unit can be embedded in the processor of the electronic device or independent of the processor in the form of hardware or firmware, or can be stored in the memory of the electronic device in the form of software for the processor to call to perform the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.

[0115] It will be appreciated by those skilled in the art that the method steps of the present invention may be performed by instructing related hardware such as an electronic device or a processor through a computer program, and the computer program may be stored in a non-temporary computer-readable storage medium, which causes the steps of the present invention to be performed when the computer program is executed. Depending on the circumstances, any reference to memory, storage, or other media herein may include non-volatile or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0116] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In addition, "front", "back", "left", "right", "upper" and "lower" in this article are all referenced to the placement state shown in the accompanying drawings.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large-scale industrial process data aggregation method, characterized in that: include: Establish industrial process data acquisition connections in a multi-threaded manner to obtain raw data of large-scale industrial process data; The raw data of large-scale industrial process data is read through the asynchronous I / O architecture and placed in the data cache after processing; Push data to data applications based on subscription information and respond to end-user needs through data applications.

2. A large-scale industrial process data aggregation method according to claim 1, characterized in that: Establishing industrial process data acquisition connections in a multi-threaded manner includes: The main thread is defined as a management program thread, which is used to configure the data collection program, data cache area and data application program. Each data collection program and data application program sets an independent running thread. The collection program and application program realize zero-copy aggregation processing of data through a shared columnar data cache area based on the Apache Arrow open protocol. The multi-threading method uses the POSIX thread library in the Linux system to create and manage threads.

3. A large-scale industrial process data aggregation method according to claim 2, characterized in that: The data acquisition program is configured as an acquisition connection; the acquisition connection includes an array consisting of several data acquisition points, and the arrays of several data acquisition points form an acquisition channel. The data acquisition program collects data information using the acquisition channel as the minimum execution unit. One acquisition connection corresponds to an independently running thread in the multi-threading. Each acquisition channel under the same acquisition connection is scheduled in the thread based on an asynchronous I / O mechanism.

4. A large-scale industrial process data aggregation method according to claim 3, characterized in that: The raw data of large-scale industrial process data acquired through asynchronous I / O architecture includes: After receiving the collection instruction of large-scale industrial data, the management program dynamically creates a running thread for this collection connection, registers an asynchronous I / O instance in this thread, then creates a loop body in the thread, and calls the asynchronous I / O waiting function in the loop body to detect whether there is an I / O event completed. Each collection channel will register a collection event timer in the collection connection thread. The timing time is set according to the collection frequency of the collection channel. When the timer time is reached, the data collection event is triggered to submit an I / O request to the asynchronous I / O instance. The loop body waiting function in the collection connection thread detects the occurrence of the I / O event, and obtains the original collection data through the asynchronous I / O event result acquisition function.

5. A large-scale industrial process data aggregation method according to claim 4, characterized in that: The data buffer includes at least the following fields: name, data type, value, timestamp, value change flag, value valid flag and metadata. The data cache is based on the Apache Arrow open protocol format, and the data is stored in columns. The data cache is divided into different Arenas, one collection connection corresponds to one Arena, each Arena corresponds to a different collection channel, and each Arena is divided into different Tables. The management program thread initializes the space occupied by Arena and Table according to the number of collection points of each collection channel and the data application program configuration; The acquisition program stores the processed data in a columnar data buffer based on the Apache Arrow open protocol format. The process of reading and storing data in the data buffer includes the following operations: Determine whether the acquisition channel reading is valid. If invalid, set an invalid flag for each data point in the buffer; Iterate the read data set and convert each data into the data type supported by the Apache Arrow protocol; Store the converted data into the data buffer and set the value change flag.

6. A large-scale industrial process data aggregation method according to claim 5, characterized in that: The data application is configured as an application service. Multiple application services can run independently at the same time. Each application service subscribes to the data content in the data cache area in units of Table. The subscription methods of the application service to the data content in the data cache area include: PERIODIC mode, which includes a periodic mode, and is used for application services to obtain data for processing by subscribing to messages according to a configured time period; NOTIFICATION mode, the NOTIFICATION mode includes a notification mode, which is used for the application service to obtain data through subscription messages when the subscribed data changes; SCHEDULE mode, the SCHEDULE mode includes a plan mode, which is used for the application service to obtain data through subscription messages according to the configured acquisition plan; DIRECT mode, the DIRECT mode includes a direct memory analysis mode, which is used for application services to directly perform analysis operations on data in the data cache area through the Apache Arrow agent.

7. A large-scale industrial process data aggregation method according to claim 5, characterized in that: The application service also filters the data in the data cache area, and the filtering method includes: ALL filtering, wherein the ALL filtering includes subscribing to all data; INCLUSIVE filtering, wherein the INCLUSIVE filtering includes subscribing to all selected data; EXCLUSIVE filtering, wherein the EXCLUSIVE filtering includes subscribing to all non-selected data; Among them, INCLUSIVE filtering and EXCLUSIVE filtering both filter the subscription content through a filter table.

8. A large-scale industrial process data aggregation method according to claim 1, characterized in that: Push data to data applications based on subscription information and respond to end-user needs through data applications including: The data collection program processes the raw data and stores it in a columnar data cache based on the Apache Arrow open protocol; The management program queries the subscriber information of the current Table; Confirm the event notification function registered by the data application that subscribes to this Table; Call the event notification function to notify the data application; After receiving the subscription message, the data application operates the subscription data according to the subscription method and filtering method.

9. A large-scale industrial process data aggregation device, characterized in that: include: A management unit, which is used to configure basic information, provide a user interface, create and manage data acquisition connections, create and manage data buffers, and create and manage data application services; A data cache unit, which is used to provide a columnar data cache area for the collected data based on the Apache Arrow open protocol, and configure the size and performance parameters of the cache area physical space Arena and Table; A data collection unit, which is used to configure data collection connections, collection channels, manage connection properties, and maintain data collection point information; The data application unit is used to configure and respond to data application services that meet the requirements of terminal data users, manage application service attributes, and set application service subscription information.

10. A computing device, characterized in that The computing device comprises a processor and a memory, wherein: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute a large-scale industrial process data aggregation method as described in claims 1-8 according to the instructions in the program code.