Coal production data processing method and system based on Internet of Things technology
Through Internet of Things technology, the collection and processing of production data in the coal industry and the selection of appropriate processing methods based on the processing efficiency is solved, the problem of low processing efficiency of coal production data is realized, efficient data collection and accurate calculation are achieved, and support for the intelligent development of the coal industry.
Patent Information
- Application Number
- CN202510323405.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology lacks effective methods to collect and process coal production data, which limits the intelligent production and data application of the coal industry.
Coal production data is collected through edge gateways based on IoT technology, characteristic information associated with production data and response information of the data processing system, the processing efficiency of the data processing system to process production data, and the matching processing method (streaming, batch or mixed processing) is selected for data processing.
It realizes efficient collection and accurate calculation of coal production data, providing a solid foundation for intelligent production, quality monitoring, environmental monitoring, etc. in the coal industry.
Smart Images

Figure CN120218855A_ABST
Abstract
Description
Technical Field
[0001] This application relates to technical fields such as industrial Internet technology, Internet of Things and sensor technology, big data technology, and cloud computing, and particularly relates to a method and system for processing coal production data based on Internet of Things technology. Background Art
[0002] Industrial Internet is a new type of infrastructure, application mode, and industrial ecosystem that deeply integrates new generation information and communication technology with industrial economy. It builds a brand-new manufacturing and service system covering the entire industrial chain and value chain through comprehensive connection of people, machines, things, systems, etc., providing an implementation approach for the digital, networked, and intelligent development of industry and even the entire industry. In the aspect of coal industry production data collection, industrial Internet technology mainly plays an important role.
[0003] The future application scenarios of mining production data are very broad, covering multiple aspects such as intelligent factories and automated production, supply chain optimization, energy management, safety assurance, product quality monitoring and traceability, product innovation and optimization, and sales forecasting and demand management. However, there is currently a lack of effective means for collecting and processing coal production data. Summary of the Invention
[0004] An embodiment of this application provides a method and system for processing coal production data based on Internet of Things technology.
[0005] According to the first aspect of the embodiments of this application, a method for processing coal production data based on Internet of Things technology is provided, which is applied to a data processing system. The method includes:
[0006] Collecting production data of the coal industry from edge gateway access devices based on an edge gateway;
[0007] Obtaining feature information associated with the production data and obtaining the response information of the data processing system;
[0008] Determining the processing efficiency of the data processing system for processing production data according to the feature information and the response information;
[0009] Selecting a processing method that matches the processing efficiency from multiple processing methods to process the production data; among them, the multiple processing methods include a streaming data processing method, a batch data processing method, and a hybrid processing method, and the hybrid processing method refers to a method that mixes streaming data processing and batch data processing.
[0010] According to the second aspect of the embodiments of this application, a data processing system is provided, including:
[0011] An acquisition module, configured to acquire production data of the coal industry from edge gateway access devices based on an edge gateway, and obtain feature information associated with the production data and response information of the data processing system. According to the feature information and the response information, determine the processing efficiency of the data processing system for processing the production data;
[0012] A data processing platform, configured to select a processing method matching the processing efficiency from multiple processing methods to process the production data; wherein, the multiple processing methods include a streaming data processing method, a batch data processing method, and a hybrid processing method, and the hybrid processing method refers to a method of mixing streaming data processing and batch data processing.
[0013] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.
[0017] According to a fourth aspect of the embodiments of the present application, there is provided a storage medium storing instructions that, when run on an electronic device, cause the electronic device to execute the method described in the foregoing first aspect.
[0018] According to a fifth aspect of the embodiments of the present application, there is provided a computer program product that, when instructions in the computer program product are executed by a processor, implements the steps of the method described in the foregoing first aspect.
[0019] According to the technical solution of the present application, the processing efficiency of the data processing system for processing the production data can be determined according to the feature information associated with the collected production data of the coal industry and the response data of the data processing system, which is convenient for selecting a processing method matching the processing efficiency to process the production data, so that the coal production data can be collected and calculated efficiently and accurately, providing a solid foundation for the intelligent production, quality monitoring, environmental monitoring, etc. of the coal industry.
[0020] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:
[0022] Figure 1 Schematic flowchart of the coal production data processing method based on the Internet of Things technology provided by the embodiments of the present application;
[0023] Figure 2 Block diagram of the data processing system provided by the embodiments of the present application;
[0024] Figure 3 Schematic architecture diagram of the data processing system provided by the embodiments of the present application. Detailed implementation manners
[0025] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0026] The following description is made with reference to the accompanying drawings for the exemplary embodiments of the present application, including various details of the embodiments of the present application to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0027] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0029] It should be noted that in the technical solution of this application, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processing comply with the provisions of relevant laws and regulations and do not violate public order and good customs. The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0030] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0031] Next, a coal production data processing method and system based on Internet of Things technology according to embodiments of this application will be described with reference to the accompanying drawings.
[0032] Among them, it should be noted that the execution subject of the coal production data processing method based on Internet of Things technology in the embodiments of this application can be a data processing system, which can be implemented in the form of software and / or hardware, and this system can be configured in an electronic device. Exemplarily, the electronic device can include but not be limited to terminals, server sides, etc.
[0033] Figure 1 It is a schematic flowchart of the coal production data processing method based on Internet of Things technology provided for the embodiments of this application. As Figure 1 shown, the coal production data processing method based on Internet of Things technology can include but not be limited to the following steps.
[0034] In step 101, production data of the coal industry is collected from edge gateway access devices based on an edge gateway.
[0035] In some embodiments, the edge gateway can collect production line measurement point data as the production data from devices or systems such as PLC (Programmable Logic Controller), DCS (Distributed Control System), and SCADA (Supervisory Control And Data Acquisition) at the production site of the coal industry. Exemplarily, the edge gateway can include a collection module, and the production data of the coal industry is collected from edge gateway access devices through this collection module.
[0036] In some embodiments, the edge gateway access device may include, but is not limited to, one or more of the following devices: safety monitoring sub-station, hydrological monitoring sub-station, pressure monitoring sub-station, stress monitor, chromatograph analyzer, seismometer, water pump, etc. Optionally, the edge gateway access device may further include access sensors. Exemplarily, the access sensors may include: oxygen concentration sensor, carbon monoxide concentration sensor, gas concentration sensor, hydrogen sulfide concentration sensor, carbon dioxide concentration sensor, dust concentration sensor, water temperature sensor, water pressure sensor, water level sensor, flow sensor, nitrogen concentration, ethylene concentration, ethane concentration, wind speed, wind pressure, temperature, power feed status, roof pressure, surrounding rock stress, water pump flow rate, water pump power, and various other sensors and data points.
[0037] In some embodiments, the production data collected in the coal industry may include time-series data. Optionally, the production data may further include non-time-series data, and it can be determined whether the collected production data is time-series data, non-time-series data, or both time-series data and non-time-series data according to the actual requirements of the application system. Optionally, when the edge-side basic network cannot guarantee real-time data transmission, an edge time-series library can be deployed on the edge side of the data processing system to cache a large amount of time-series data generated by the production line. Generally, the edge time-series library does not store time-series data for too long a period. Exemplarily, it can store time-series data for a certain period (such as 5 to 7 days), or the configuration can be adjusted according to the actual production situation to achieve longer-period data storage. The data in the edge time-series library can be compressed and transmitted to the data processing platform side regularly according to the network situation.
[0038] In step 102, obtain the feature information associated with the production data and obtain the response information of the data processing system.
[0039] In the embodiments of the present application, during the process of collecting production data of the coal industry from the edge gateway access device based on the edge gateway, the key information elements of the production data can be recorded, and the feature information associated with the production data can be determined based on the key information elements. In some embodiments, the feature information associated with the production data may include, but is not limited to, at least one of the following: processing duration per unit of production data, number of units of production data processed, total amount of production data processed, complexity of production data system processing, maximum delay time of processed production data.
[0040] In the embodiments of the present application, the response information of the data processing system can be read from the data processing system. In some embodiments, the response information of the above data processing system may include, but is not limited to, at least one of the following: benchmark duration for the system to process data, benchmark duration for the system to process per unit of production data, actual system delay.
[0041] In step 103, according to the feature information and the response information, determine the processing efficiency of the data processing system for processing production data.
[0042] In some embodiments, according to the feature information associated with the production data and the response information of the data processing system, the processing efficiency of the data processing system for processing production data can be predicted based on deep learning techniques. Exemplarily, a deep learning model can be trained using training data to obtain a trained deep learning model. The training data can include the feature information associated with historical production data and the response information of the data system. Among them, the input of the deep learning model includes the feature information associated with historical production data and the response information of the data system, and the output is the processing efficiency of the data processing system for processing production data. The deep learning model can be a convolutional neural network model, but is not limited thereto. In model application, the obtained feature information associated with the production data and the response information of the data processing system can be input into the pre-trained deep learning model, so that the processing efficiency of the data processing system for processing production data can be obtained.
[0043] In some embodiments, multiple real-time processing performance information can be determined according to the feature information associated with the production data and the response information of the data processing system; according to the multiple real-time processing performance information, determine the processing efficiency of the data processing system for processing production data. Among them, in some embodiments, the above multiple real-time processing performance information can include but is not limited to at least one of the following: real-time processing rate, real-time reception rate, real-time processed data volume, real-time processing complexity, real-time delay rate.
[0044] In a possible implementation manner, the above real-time processing rate can be determined according to the processing duration of unit production data and the reference duration of the system for processing data. Exemplarily, a logarithmic operation can be performed on the ratio of the processing duration of unit production data to the reference duration of the system for processing data, and the obtained value is determined as the real-time processing rate. For example, the calculation formula of the real-time processing rate can be expressed as follows: r1 = ln(t1 / t3), where r1 is the real-time processing rate, t1 is the processing duration of unit production data, t3 is the reference duration of the system for processing data, and ln() is the logarithmic function with the constant e as the base. When t1 = t3, the value of r1 is 0, which can be used as a threshold point. The longer the processing duration t1 of unit production data, the larger the value of r1, and the higher the real-time processing requirement for this production data.
[0045] In a possible implementation manner, the above real-time reception rate can be determined according to the quantity of production data processed per unit and the benchmark duration for the system to process production data per unit. Exemplarily, the ratio of the quantity of production data processed per unit to the benchmark duration for the system to process production data per unit can be determined as the above real-time reception rate. For example, the calculation formula for this real-time reception rate can be expressed as follows: r2 = n / t4, where r2 is the real-time reception rate, n is the quantity of production data processed per unit, and t4 is the benchmark duration for the system to process production data per unit.
[0046] In a possible implementation manner, the above real-time processed data volume can be determined according to the total volume of production data processed and the benchmark duration for the system to process data. Exemplarily, the ratio of the total volume of production data processed to the benchmark duration for the system to process data can be determined as this real-time processed data volume. For example, the calculation formula for this real-time processed data volume can be expressed as follows: v1 = v / t3, where v1 is the real-time processed data volume, v is the total volume of production data processed, and t3 is the benchmark duration for the system to process data.
[0047] In a possible implementation manner, the above real-time processing complexity can be determined according to the complexity of processing production data by the system and the benchmark duration for the system to process production data per unit. Exemplarily, the ratio of the complexity of processing production data by the system to the benchmark duration for the system to process production data per unit can be determined as this real-time processing complexity. For example, the calculation formula for this real-time processing complexity can be expressed as follows: c1 = c / t4, where c1 is the real-time processing complexity, c is the complexity of processing production data by the system, and t4 is the benchmark duration for the system to process production data per unit.
[0048] In a possible implementation, after obtaining multiple real-time processing efficiency information, based on the multiple real-time processing efficiency information, deep learning technology can be used to predict the processing efficiency of the data processing system for processing production data. Exemplarily, historical data (such as processing efficiency information over a period of time) can be used as training data, and based on this training data, a deep learning model is trained, and the trained deep learning model is used to predict the processing efficiency of the data processing system for processing production data.
[0049] In a possible implementation, after obtaining multiple real-time processing efficiency information, based on the multiple real-time processing efficiency information, deep learning technology can be used to predict the processing efficiency of the data processing system for processing production data. Exemplarily, historical data (such as processing efficiency information over a period of time) can be used as training data, and based on this training data, a deep learning model is trained, and the trained deep learning model is used to predict the processing efficiency of the data processing system for processing production data.
[0050] In another possible implementation, the processing efficiency of the data processing system for processing production data can be calculated using a first formula based on multiple real-time processing performance information; wherein, the first formula can be expressed as follows:
[0051]
[0052] Wherein, E is the processing efficiency of the data processing system for processing production data; r1 is the real-time processing rate, r2 is the real-time reception rate, v1 is the amount of real-time processed data, c1 is the real-time processing complexity, r3 is the real-time delay rate, all of which are real-time processing performance information; p1, p2, p3, and p4 are all processing performance impact factors. Exemplarily, the above p1, p2, p3, and p4 can be preset adjustable parameters, for example, empirical values obtained based on a large number of experiments. Or, exemplarily, the above p1, p2, p3, and p4 can be hyperparameters learned through deep learning techniques.
[0053] In step 104, select a processing method that matches the processing efficiency from multiple processing methods to process the production data. Among them, the multiple processing methods include a streaming data processing method, a batch data processing method, and a hybrid processing method, and the hybrid processing method refers to a method that mixes streaming data processing and batch data processing.
[0054] In the embodiments of the present application, the processing efficiency can be compared with a first threshold and a second threshold, and based on the result of the comparison, a corresponding processing method is selected from multiple processing methods to process the production data.
[0055] In some embodiments, when the processing efficiency is less than the first threshold, a batch data processing method that matches the processing efficiency can be used to process the production data. Exemplarily, a batch computing tool can be used to process and process the production data, and the processed data can be stored in a data warehouse. For example, when the processing efficiency of the data processing system for processing production data is less than the first threshold, it can be considered that the collected production data is non-temporal data (not temporal data) or other reasons cause the production data to need to be processed in batches. At this time, the Spark batch computing tool can be used to process and process the production data, and the processed data is directly stored in the data warehouse for permanent preservation. Exemplarily, the data warehouse data can be applicable to business systems that need to store a large amount of historical data and conduct in-depth analysis.
[0056] In some embodiments, when the processing efficiency is greater than a second threshold, a streaming data processing method matching the processing efficiency may be adopted to process production data. Exemplarily, the production data is stored in a Kafka cluster, and data operations and processing are performed on the production data in the Kafka cluster based on a real-time computing tool, and the processed data is stored in a time series database; wherein, the data in the time series database is periodically summarized and saved in a data warehouse. For example, when the processing efficiency of the data processing system for processing production data is greater than the second threshold, it can be considered that the collected production data is time series data and requires efficient calculation. For example, a streaming data processing method needs to be adopted for processing. For example, a Kafka cluster can be used to receive streaming data (i.e., the collected production data), and real-time computing tools such as Flink are used to perform data operations and processing on the production data in the Kafka cluster. The processed data will be stored in the time series database and periodically summarized and stored in the data warehouse for long-term storage. Exemplarily, the data in the time series database is applicable to business systems with high timeliness requirements but not requiring a large amount of historical data.
[0057] In some embodiments, when the processing efficiency is greater than or equal to a first threshold and less than or equal to a second threshold, a hybrid processing method matching the processing efficiency may be adopted to process production data. In one possible implementation, the data allocation rate of the hybrid processing method can be determined; based on this data allocation rate, the production data is split into a first part of production data and a second part of production data, wherein the first part of production data is to be processed by a streaming data processing method, and the second part of production data is to be processed by a batch data processing method. According to different business requirements, the time series database and the data warehouse can be flexibly combined to meet complex application scenarios.
[0058] In an alternative implementation, the above data allocation rate can be a preset fixed value. In another alternative implementation, the data allocation rate can be determined based on the real-time processed data volume and the real-time processing complexity. Exemplarily, a data processing correspondence table can be maintained, which includes the mapping relationship between v1*c1 and the ratio of the time series data warehouse usage. Exemplarily, the ratio of the time series data warehouse usage can be a value in [0, 1] (i.e., the value range of the ratio of the time series data warehouse usage can be greater than or equal to 0 and less than or equal to 1), and the v1*c1 represents the product value of the real-time processed data volume and the real-time processing complexity. For example, assuming that the ratio of the time series data warehouse usage corresponding to v1*c1 found from the table is 0.25, then the data allocation rate of the hybrid processing method is determined to be 1:4, that is, the allocation rate of the streaming data processing and the batch data processing for hybrid processing of this production data is 1:4. One-fifth of the production data can be used as the above first part of production data for processing by the streaming data processing method, and four-fifths of the production data can be used as the above second part of production data for processing by the batch data processing method.
[0059] In some embodiments, during the process of processing production data by using a processing method matching the processing efficiency, it is possible to determine whether an alarm event is triggered based on a preset warning rule, and when the alarm event is triggered, an alarm is made based on the alarm method associated with the alarm event. Exemplarily, during the process of performing streaming data processing and / or batch data processing on production data, it is possible to further determine whether an alarm event will be triggered based on the warning rule. If the alarm event is triggered, an alarm is made based on the alarm method associated with the alarm event. Among them, the warning rule can be set according to actual needs. For example, the warning rule can include the triggering condition of the alarm event and can also include the alarm method associated with the alarm event, etc. This application does not make specific limitations on this and will not elaborate further.
[0060] In the above embodiments, the processing efficiency of the data processing system for processing the production data can be determined according to the characteristic information associated with the collected production data of the coal industry and the response data of the data processing system, which is convenient for selecting a processing method matching the processing efficiency to process the production data, so that the coal production data can be collected and calculated efficiently and accurately, providing a solid foundation for the intelligent production, quality monitoring, environmental monitoring, etc. of the coal industry.
[0061] Figure 2 It is a block diagram of the data processing system provided by the embodiments of this application. As Figure 2 shown, the data processing system may include: a collection module 201 and a data processing platform 202. Among them, the collection module 201 is used to collect the production data of the coal industry from the edge gateway access device based on the edge gateway, and obtain the characteristic information and the response information associated with the production data, and determine the processing efficiency of the data processing system for processing the production data according to the characteristic information and the response information. The data processing platform 202 is used to select a processing method matching the processing efficiency from multiple processing methods to process the production data; among them, the multiple processing methods include a streaming data processing method, a batch data processing method, and a hybrid processing method, and the hybrid processing method refers to a method of mixing streaming data processing and batch data processing. Exemplarily, the collection module 201 can be deployed in the edge gateway.
[0062] In some embodiments, the acquisition module 201 is configured to: determine a plurality of real-time processing efficiency information according to the feature information and the response information; and determine the processing efficiency of the data processing system for processing production data according to the plurality of real-time processing efficiency information. Exemplarily, the feature information includes at least one of the following: the processing duration of unit production data, the quantity of unit production data processed, the total quantity of production data processed, the complexity of the production data system processed, the maximum latency time of the production data processed; the response information includes at least one of the following: the reference duration for the system to process data, the reference duration for the system to process unit production data, the actual latency of the system; the plurality of real-time processing efficiency information includes at least one of the following: real-time processing rate, real-time reception rate, real-time processed data volume, real-time processing complexity, real-time latency rate.
[0063] In some embodiments, the acquisition module 201 is configured to: determine the real-time processing rate according to the processing duration of unit production data and the reference duration for the system to process data; determine the real-time reception rate according to the quantity of unit production data processed and the reference duration for the system to process unit production data; determine the real-time processed data volume according to the total quantity of production data processed and the reference duration for the system to process data; determine the real-time processing complexity according to the complexity of the production data system processed and the reference duration for the system to process unit production data; determine the real-time latency rate according to the maximum latency time of the production data processed and the actual latency of the system.
[0064] In some embodiments, the acquisition module 201 is configured to: calculate the processing efficiency of the data processing system for processing production data by using a first formula according to the plurality of real-time processing efficiency information; wherein, the first formula is expressed as follows:
[0065]
[0066] wherein, E is the processing efficiency of the data processing system for processing production data; r1 is the real-time processing rate, r2 is the real-time reception rate, v1 is the real-time processed data volume, c1 is the real-time processing complexity, r3 is the real-time latency rate, all of which are real-time processing efficiency information; p1, p2, p3, p4 are all processing efficiency impact factors.
[0067] In some embodiments, the data processing platform 202 is configured to: when the processing efficiency is less than a first threshold, process the production data by using a batch data processing method matching the processing efficiency; or, when the processing efficiency is greater than a second threshold, process the production data by using a streaming data processing method matching the processing efficiency; or, when the processing efficiency is greater than or equal to the first threshold and less than or equal to the second threshold, process the production data by using a hybrid processing method matching the processing efficiency.
[0068] In some embodiments, the data processing platform 202 is configured to: determine the data distribution rate of the hybrid processing mode; and based on the data distribution rate, split the production data into a first part of production data and a second part of production data, wherein the first part of production data is to be processed by the streaming data processing mode, and the second part of production data is to be processed by the batch data processing mode.
[0069] In some embodiments, the data processing platform 202 is configured to: process and process the production data by using a batch computing tool, and store the processed data in a data warehouse. In some embodiments, the data processing platform 202 is configured to: store the production data in a Kafka cluster, and perform data operation and processing on the production data in the Kafka cluster based on a real-time computing tool, and store the processed data in a time series database; wherein the data in the time series database is periodically summarized and saved in the data warehouse.
[0070] In some embodiments, the data processing platform 202 is further configured to: during the process of processing the production data by using a processing mode matching the processing efficiency, determine whether an alarm event is triggered based on a preset warning rule, and when the alarm event is triggered, perform an alarm based on the alarm mode associated with the alarm event.
[0071] It should be noted that the foregoing explanation of the embodiments of the coal production data processing method based on the Internet of Things technology is also applicable to the data processing system of this embodiment, and will not be elaborated here.
[0072] To facilitate those skilled in the art to understand this application more clearly, the following will be combined with Figure 3 for description.
[0073] As Figure 3 shown, the data processing system can be divided into an edge side, a platform side, and an application side. Among them, an edge gateway can be deployed on the edge side. The edge gateway collects production line measurement point data from devices or systems such as PLC, DCS, and SCADA at the production site, and the collected data can be sent to the collection platform side (such as the data processing platform) through a 4G / 5G mobile network or an enterprise internal fiber optic network. For example, business system data can be collected from business systems (such as office automation OA system, engineering production management system PMS, network content control system NCC system, knowledge base system, human resources system, etc.) as the production data.
[0074] When the edge-side basic network cannot guarantee real-time data transmission, an edge time series library can be deployed on the edge side to cache a large amount of time series data generated by the production line. Generally, the edge time series library does not store time series data for too long a period. Usually, it stores time series data within 5 to 7 days, and the configuration can also be adjusted according to the actual production situation to achieve longer-period data storage. The data in the edge time series library will be compressed and transmitted to the platform side regularly according to the network situation.
[0075] Exemplarily, the edge gateway access devices include various devices such as safety monitoring sub-stations, hydrological monitoring sub-stations, pressure monitoring sub-stations, stress monitors, chromatographic analyzers, seismographs, water pumps, etc. The access sensors involved include various sensors and data points such as oxygen concentration sensors, carbon monoxide concentration sensors, gas concentration sensors, hydrogen sulfide concentration sensors, carbon dioxide concentration sensors, dust concentration sensors, water temperature sensors, water pressure sensors, water level sensors, flow sensors, nitrogen concentration, ethylene concentration, ethane concentration, wind speed, wind pressure, temperature, feed status, roof pressure, surrounding rock stress, water pump flow, water pump power, etc.
[0076] As Figure 3 shown, the platform side can calculate and store the collected production data (such as time series data and non-time series data). Based on the production data collected by the edge gateway, it can be divided into two cases according to the actual application scenario. For the case of stream data processing, the data will first be stored in the Kafka cluster, and the Flink real-time computing tool will be used for data operation and processing. The processed data will be stored in the time series database (such as InfluxDB or CirroTimes), and the data in the time series library will be periodically summarized and permanently saved in the data warehouse. For the case of batch data processing, the Spark batch computing tool will be used for data processing and processing, and the processed data will be directly stored in the data warehouse (such as DM database or CirroData) for permanent storage.
[0077] Usually, after the production data is collected on the platform side, there is a need for real-time warning. To achieve this in production, Flink and Spark will read the warning rules in the relational database (such as MySQL) to further determine whether to alarm and the way of alarming, etc.
[0078] For the above coal production data, the following three processing methods are mainly adopted:
[0079] (1) Streaming data processing: Use a Kafka cluster to receive streaming data and use Flink for real-time data calculation and processing. The processed data will be stored in a time series database (such as InfluxDB, CirroTimes) and periodically summarized and stored in a data warehouse (such as DM Database, CirroData) for long-term storage.
[0080] (2) Batch data processing: Use Spark batch computing tools to process and process data, and the processed data is directly stored in the data warehouse for permanent storage.
[0081] (3) Hybrid processing: Use both processing methods in combination.
[0082] Correspondingly, the acquisition results of production data will ultimately be stored in a time series database and a data warehouse for the application system to call according to requirements. Among them, time series database data: suitable for business systems with high timeliness requirements but not requiring a large amount of historical data. Data warehouse data: suitable for business systems that need to store a large amount of historical data and conduct in-depth analysis. Hybrid use: According to different business requirements, the time series database and the data warehouse can be flexibly combined to meet complex application scenarios. Exemplarily, the method embodiments involved in this application can be used to process coal production data, and the implementation method can refer to the optional implementation methods of the above method embodiments, which will not be elaborated here.
[0083] As Figure 3 shown, the application side can use coal production data. The acquisition results of production data will be stored in a time series database and a data warehouse. The application side can select which database to obtain data according to actual usage requirements. Usually, the data in the time series database will be provided to business systems with high timeliness requirements but not very concerned about a large amount of historical data. For business systems with low timeliness requirements but high requirements for a large amount of history, they will choose to obtain data from the data warehouse. According to actual business needs, the time series database and the data warehouse can also be used in combination to meet more complex business needs.
[0084] It should be noted that the embodiments of this application can provide strong support for the acquisition, calculation, and storage of production data (such as production environment time series data). And the future application scenarios of these production data are very broad, not only providing a support basis for data mining and in-depth application, but also playing a crucial role in future intelligent factories, automated production, production process optimization, product quality monitoring and traceability. The realization of these application scenarios will help improve production efficiency, reduce costs, improve product quality and market competitiveness.
[0085] To implement the above embodiments, the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the coal production data processing method provided by the present application based on Internet of Things technology.
[0086] To implement the above embodiments, the present application further provides a storage medium storing instructions that, when run on an electronic device, cause the electronic device to execute the coal production data processing method provided by the present application based on Internet of Things technology.
[0087] To implement the above embodiments, the present application further provides a computer program product that, when the instructions in the computer program product are executed by a processor in an electronic device, implements the steps of the coal production data processing method provided by the present application based on Internet of Things technology.
[0088] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0089] Any process or method description in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a manner not shown or discussed, including substantially concurrently or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0090] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0091] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0092] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0093] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0094] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for processing coal production data based on Internet of Things technology, characterized in that: Applied to a data processing system, the method comprises: Collect production data of the coal industry from edge gateway access devices based on edge gateway; Acquiring characteristic information associated with the production data, and acquiring response information of the data processing system; Determining, based on the characteristic information and the response information, a processing efficiency of the data processing system in processing the production data; A processing method that matches the processing efficiency is selected from a plurality of processing methods to process the production data; wherein the plurality of processing methods include a streaming data processing method, a batch data processing method and a hybrid processing method, and the hybrid processing method refers to a method of mixing streaming data processing and batch data processing.
2. The method according to claim 1, characterized in that The determining, according to the characteristic information and the response information, a processing efficiency of the data processing system in processing the production data includes: Determining a plurality of real-time processing performance information according to the characteristic information and the response information; The processing efficiency of the data processing system in processing the production data is determined according to the plurality of real-time processing performance information.
3. The method according to claim 2, characterized in that The characteristic information includes at least one of the following: the processing time of unit production data, the number of unit production data processed, the total amount of production data processed, the complexity of the production data system processing, and the maximum delay time of the processed production data; The response information includes at least one of the following: a benchmark time for the system to process data, a benchmark time for the system to process unit production data, and an actual system delay; The multiple real-time processing performance information includes at least one of the following: real-time processing rate, real-time receiving rate, real-time processing data volume, real-time processing complexity, and real-time delay rate.
4. The method according to claim 3, characterized in that The determining of a plurality of real-time processing performance information according to the characteristic information and the response information includes: Determining the real-time processing rate according to the processing time of the unit production data and the benchmark time of the system processing data; Determining the real-time receiving rate according to the amount of unit production data processed and the benchmark time length for the system to process the unit production data; Determining the real-time processing data volume according to the total volume of the processed production data and the benchmark time length for the system to process the data; Determining the real-time processing complexity according to the complexity of the production data system processing and the benchmark time length of the system processing unit production data; The real-time delay rate is determined according to the maximum delay time of the processed production data and the actual delay of the system.
5. The method according to any one of claims 2 to 4, characterized in that: Determining the processing efficiency of the data processing system in processing the production data according to the plurality of real-time processing efficiency information includes: According to the plurality of real-time processing efficiency information, a first formula is used to calculate the processing efficiency of the data processing system in processing the production data; wherein the first formula is expressed as follows: Among them, E is the processing efficiency of the data processing system in processing the production data; r1 is the real-time processing rate, r2 is the real-time receiving rate, v1 is the real-time processing data volume, c1 is the real-time processing complexity, and r3 is the real-time delay rate, all of which are the real-time processing performance information; p1, p2, p3, and p4 are all processing performance influencing factors.
6. The method according to claim 1, characterized in that The selecting a processing method that matches the processing efficiency from a plurality of processing methods to process the production data includes: The processing efficiency is less than a first threshold, and the production data is processed using the batch data processing method that matches the processing efficiency; or, The processing efficiency is greater than a second threshold, and the production data is processed using the streaming data processing method that matches the processing efficiency; or, The processing efficiency is greater than or equal to the first threshold and less than or equal to the second threshold, and the production data is processed using the hybrid processing method that matches the processing efficiency.
7. The method according to claim 6, characterized in that The adopting the hybrid processing method matching the processing efficiency to process the production data includes: determining a data allocation rate of the hybrid processing mode; Based on the data allocation rate, the production data is split into a first part of production data and a second part of production data, wherein the first part of production data is to be processed by the stream data processing method and the second part of production data is to be processed by the batch data processing method.
8. The method according to claim 6, characterized in that The adopting the batch data processing method matching the processing efficiency to process the production data comprises: Processing the production data using a batch computing tool, and storing the processed data in a data warehouse; The adopting the streaming data processing method matching the processing efficiency to process the production data includes: The production data is stored in a Kafka cluster, and data operations and processing are performed on the production data in the Kafka cluster based on real-time computing tools, and the processed data is stored in a time series database; wherein the data in the time series database is regularly aggregated into the data warehouse for storage.
9. The method according to claim 1, characterized in that The method further comprises: In the process of processing the production data using a processing method that matches the processing efficiency, it is determined whether an alarm event is triggered based on a preset early warning rule, and when the alarm event is triggered, an alarm is issued based on an alarm method associated with the alarm event.
10. A data processing system, characterized in that: include: A collection module, used to collect production data of the coal industry from an edge gateway access device based on an edge gateway, and obtain characteristic information associated with the production data and response information of the data processing system, and determine the processing efficiency of the data processing system in processing the production data according to the characteristic information and the response information; A data processing platform is used to select a processing method that matches the processing efficiency from a plurality of processing methods to process the production data; wherein the plurality of processing methods include a streaming data processing method, a batch data processing method and a hybrid processing method, and the hybrid processing method refers to a method of mixing streaming data processing and batch data processing.
Citation Information
Patent Citations
Data processing framework and data processing method based on batch processing and stream processing
CN106873945A
Data storage processing method and device for coal mine data center station
CN112650739A
Industrial data processing system and method based on streaming computing engine, and medium
CN116010452A
Stream processing and batch processing switching method and switching device
CN116841753A
Carbon emission data acquisition method based on edge gateway
CN117527855A