Mine auxiliary transportation data processing system
By introducing real-time processing modules and offline processing modules into the mine auxiliary transportation data processing system, the problem of single and low efficiency of mine data processing methods in the existing technology is solved, and efficient processing of mine auxiliary transportation data is achieved.
Patent Information
- Application Number
- CN202510024711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has problems such as single processing methods and low efficiency in mine data processing, especially in real-time and offline processing of mine auxiliary transportation data.
A mine auxiliary transportation data processing system is proposed, including a business data acquisition module, a log data acquisition module, a real-time processing module, an offline processing module and a computing module. The collected data is processed accordingly through the real-time processing module and the offline processing module to achieve stable and reliable data processing and high-efficiency batch processing.
It improves the processing efficiency of business data and log data in the mine, realizes stable and reliable processing of offline data and high-efficiency batch processing of real-time data.
Smart Images

Figure CN119961234A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of mine data processing, and in particular to a mine auxiliary transportation data processing system. Background Art
[0002] With the continuous development and application of technologies such as big data and artificial intelligence, coal mine big data has shown explosive growth. It is necessary to make full use of the large amount of data in mine production, effectively manage, analyze and mine these data, so as to improve mine production efficiency and resource utilization, optimize mine production structure, and achieve sustainable development of mine production. Research on intelligent coal mine big data governance technology is of great significance to improving the data management and utilization level of coal mining enterprises, optimizing coal mine production structure, and improving the intelligence level of coal mines. Domestic and foreign scholars have carried out a lot of research on intelligent coal mine big data governance technology, applying machine learning, artificial intelligence and other technologies to mine big data governance, exploring and practicing from the aspects of big data collection, processing, storage, analysis and application, and improving the data management and utilization level of mining enterprises. However, in actual applications, the processing method of data in mines is relatively single and the processing efficiency is low. Summary of the invention
[0003] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0004] To this end, the first object of this application is to propose a mine auxiliary transportation data processing system.
[0005] The second objective of the present application is to provide an electronic device.
[0006] The third object of the present application is to provide a computer program product.
[0007] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a mine auxiliary transportation data processing system, including:
[0008] Business data collection module, log data collection module, real-time processing module, offline processing module, and calculation module;
[0009] The business data collection module and the log data collection module are connected to the real-time processing module, the business data collection module and the log data collection module are connected to the offline processing module, and the real-time processing module and the offline processing module are both connected to the calculation module;
[0010] The log data collection module is used to collect log data of the mine auxiliary transportation system, wherein the log data is used to record the status of the system;
[0011] The business data collection module is used to collect business data of the mine auxiliary transportation system, wherein the business data is used to record business related information;
[0012] The real-time processing module is used to obtain the real-time processing data in the business data collection module and the log data collection module and perform real-time processing;
[0013] The offline processing module is used to obtain the offline processing data in the business data collection module and the log data collection module and perform offline processing;
[0014] The calculation module is used to analyze and process the real-time processing data and the offline processing data.
[0015] Optionally, the log data collection module includes: a point collection module, a first reverse proxy module, a log server module, a first aggregation transmission module, a first persistence module, and a second aggregation transmission module;
[0016] The buried point collection module is used to collect the log data at the front buried point of the mine auxiliary transportation system;
[0017] The first reverse proxy module is used to send and store the log data of the embedding point collection module in the log server module, and is also used to schedule the log data in the log server module and send it to the first aggregation transmission module;
[0018] The first aggregation transmission module is connected to the tracking point collection module and is used to send the log data to the first persistence module;
[0019] The first persistence module is used to perform persistence processing on the log data and send it to the second aggregation transmission module;
[0020] The second aggregation transmission module is used to send the log data to the real-time processing module or the offline processing module.
[0021] Optionally, the business data collection module includes: a business data collection submodule, a second reverse proxy module, a business server module, a structured query module, a channel module, and a data synchronization module;
[0022] The business data collection submodule is used to intercept business data of the mine system energy auxiliary transportation system;
[0023] The second reverse proxy module sends the business data of the business data collection submodule to the business server module and stores it, and is also used to schedule the business data in the business server module and send it to the structured query module;
[0024] The structured query module is connected to the second reverse proxy module and is used to perform preliminary processing on the business data;
[0025] The channel module is connected to the second reverse proxy module and is used to send the business data in the business server module to the first persistence module;
[0026] The data synchronization module is used to synchronize the business data in the structured query module to the real-time processing module or the offline processing module.
[0027] Optionally, the offline processing module includes: a distributed file module, a data warehouse module, an offline synchronization module, and a database module;
[0028] The distributed file module is used to store the offline processing data;
[0029] The data warehouse module is connected to the distributed file module and is used to perform hierarchical processing on the offline processing data stored in the distributed file module to obtain structured offline processing data;
[0030] The offline synchronization module is connected to the data warehouse module and the database module respectively, and is used to synchronize the structured offline processing data in the data warehouse module to the database module;
[0031] The data warehouse module is also used to send the structured offline processing data to the calculation module.
[0032] Optionally, the real-time processing module includes: a second persistence module, a stream computing module, and a dimension database module;
[0033] The second persistence module is used to receive the real-time processing data, perform persistence processing on the data, and send the data to the stream computing module;
[0034] The dimension database module is connected to the second persistence module, and the dimension database module includes a dimension table and state information;
[0035] The second persistence module is connected to the stream computing module, and the stream computing module is used to receive the data sent by the second persistence module and perform stream computing according to the dimension table and the state information to structure the real-time processing data, and to rewrite the structured real-time processing data into the second persistence module;
[0036] The second persistence module is used to perform data stratification according to the structured real-time processing data;
[0037] The streaming computing module is also used to send the structured real-time processing data to the computing module.
[0038] Optionally, the system further includes a display module, which is connected to the calculation module and is used to visualize the results of analysis and processing by the calculation module.
[0039] To achieve the above-mentioned purpose, a second aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0040] The memory stores computer-executable instructions;
[0041] The processor executes the computer-executable instructions stored in the memory to implement the system as described in any one of the first aspects.
[0042] To achieve the above-mentioned purpose, a third aspect of the present application provides a computer program product, which, when executed by a processor, implements the system described in any one of the first aspects.
[0043] The mine auxiliary transportation data processing method, device, electronic device and storage medium provided in the present application respectively process the collected data through the real-time processing module and the offline processing module, thereby realizing stable and reliable processing of offline data and efficient batch processing of real-time data, and improving the processing efficiency of business data and log data in the mine.
[0044] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0046] Figure 1 A schematic diagram of the structure of a mine auxiliary transportation data processing system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0048] With the continuous development and application of technologies such as big data and artificial intelligence, coal mine big data has shown explosive growth. It is necessary to make full use of the large amount of data in mine production, effectively manage, analyze and mine these data, so as to improve mine production efficiency and resource utilization, optimize mine production structure, and achieve sustainable development of mine production. Research on intelligent coal mine big data governance technology is of great significance to improving the data management and utilization level of coal mining enterprises, optimizing coal mine production structure, and improving the intelligence level of coal mines. Regarding intelligent coal mine big data governance technology, domestic and foreign scholars have carried out a lot of research work, applied machine learning, artificial intelligence and other technologies to mine big data governance, explored and practiced from big data collection, processing, storage, analysis and application, and improved the data management and utilization level of mining enterprises. However, there are still some problems in practical applications, mainly concentrated in the following aspects: ① The "data island" phenomenon is serious. Most of the data exchange is still manual, lacking business collaboration between data processing systems, poor timeliness, data still exists in a scattered and weakly associated manner, the system efficiency is low, it is difficult to analyze the dynamic evolution law in the coal mining process, and it is impossible to realize the integration and application of big data. ② Low data quality and lack of unified data standards. There are many internal application systems in coal mines, and there is no unified standard for various types of data, which leads to low quality of coal mine big data and difficulty in realizing multi-source data fusion application. ③ Lack of data governance system. Although there are some data usage standards at the subsystem level, there is still a lack of governance standards for overall data. ④ Data empowerment is not sufficient, there are too many conceptual studies, and relatively few application practices. Although a large amount of data is generated during the production and operation of coal mines, it is not fully empowered for actual production and operation decisions.
[0049] At present, many intelligent coal mines and intelligent mining faces have been built across the country, generating massive amounts of data. These data can be divided into three types: structured, semi-structured, and unstructured. Structured data refers to data with clearly defined formats and patterns, mainly including production and operation business data and equipment environment IoT data. Semi-structured data refers to text that has no strictly defined format but has tags or labels to help organize information processing, such as contract data in coal mine asset management systems, resume information data of personnel in human resource management systems, and various system background operation log data. Unstructured data refers to any form of data that cannot be parsed or stored in tables using traditional methods, such as images, audio, and video collected by underground sensors. Coal mine big data analysis based on audio and video data is one of the important characteristics that distinguishes intelligent coal mines from traditional coal mines. Intelligent coal mine data has the characteristics of general big data, such as large and dispersed data scale, diverse data types, fast collection and processing speed, and low data value density. In addition, it also has the characteristics of strong data time series and strong data correlation. According to the characteristics of intelligent coal mine big data and the problems existing in current data governance, the basic requirements for intelligent coal mine big data governance are as follows: ① Collect and store data in a unified manner, break down the barriers between subsystems, and provide necessary conditions for data fusion applications. ② Clean and standardize low-quality data to improve data quality and standardization to meet the requirements of application analysis. ③ Unify data asset planning, form standard data assets, and provide standard data sharing services. ④ Carry out data governance practices based on the actual situation of coal mines to form intelligent applications.
[0050] The design of Lambda architecture can provide a flexible and efficient solution for processing data of auxiliary transportation vehicles such as mine trackless rubber-tyred vehicles. The core idea of Lambda architecture is to divide the data processing process into two independent levels: batch processing and real-time processing, thereby combining the advantages of both. In the mine environment, the planning of these auxiliary transportation vehicle data can benefit from Lambda architecture in the following ways:
[0051] 1. Real-time monitoring and response:
[0052] Real-time layer: The real-time layer of the Lambda architecture can process data such as the location and status of vehicles in real time, allowing you to monitor the dynamics of transport vehicles in the mine on an almost real-time basis. This is critical for monitoring and responding to emergency situations, such as accidents, abnormal behavior, etc.
[0053] (1) Real-time event notification: The real-time layer can capture key vehicle events in real time, such as sudden accidents, equipment failures, speeding, and emergency braking. By establishing a real-time event notification system, relevant personnel can receive alerts quickly and take timely actions.
[0054] (2) Geographical Fences and Location Services: Combined with the Geographic Information System (GIS), the real-time layer can accurately monitor the vehicle's location. By setting geographic fences and area restrictions, the system can immediately send a notification once a vehicle enters or leaves a specific area, which is used to monitor the vehicle's movement trajectory and ensure that it is active within the predetermined area.
[0055] (3) Real-time dashboards and visualization: Using the data processed by the real-time layer, real-time dashboards and visualization tools are created. Such dashboards can provide key performance indicators such as vehicle speed, load conditions, fuel consumption, etc., so that operators can understand the status of the fleet in real time.
[0056] (4) Real-time fault diagnosis: Real-time monitoring of vehicle sensor data helps to detect potential faults and anomalies early. Through real-time layer analysis, real-time fault diagnosis reports can be generated to provide timely information to the maintenance team to reduce downtime and improve production efficiency.
[0057] (5) Automated response system: Based on real-time monitoring, an automated response system is established. For some common events, the system can automatically trigger corresponding responses, such as emergency stop, alarm notification, scheduling adjustment, etc., to speed up the resolution of problems.
[0058] (6) Real-time communication platform: Information from the real-time layer can be integrated into the real-time communication platform to ensure that relevant team members can receive instant notifications about vehicle status and events. This instant communication helps to take response measures more quickly.
[0059] 2. Historical data analysis:
[0060] Batch processing layer: The batch processing layer is responsible for processing historical data, including a large amount of vehicle trajectories, transportation efficiency and other information. By analyzing historical data, you can find the vehicle operation mode, peak period, trough period, etc., providing a basis for optimizing transportation plans and equipment maintenance.
[0061] (1) Efficiency and productivity analysis: The batch processing layer can analyze data such as vehicle historical trajectory, speed, load conditions, etc. to evaluate the overall transportation efficiency of the fleet. By comparing data from different time periods, efficient operation patterns can be identified to provide guidance for improving productivity and resource utilization.
[0062] (2) Transportation trend prediction: Using historical data, the batch layer can apply advanced analytical techniques, such as time series analysis or machine learning algorithms, to predict future transportation trends. This helps optimize scheduling, prepare for peak-period challenges in advance, and improve fleet management strategies.
[0063] (3) Equipment maintenance planning: Through batch layer analysis of historical failure data, common failure modes and trends of vehicle equipment can be identified. This provides a basis for formulating effective maintenance plans to reduce downtime and improve vehicle reliability.
[0064] (4) Fuel consumption and energy efficiency: The batch processing layer can analyze the historical fuel consumption and energy utilization of the vehicle. This helps to determine energy-saving strategies, improve driving behavior, reduce carbon footprint, and reduce operating costs.
[0065] (5) Regional analysis: Considering that there may be different areas or work areas in a mine, the batch layer can conduct differentiated analysis of transportation activities in these areas. By understanding the transportation needs and characteristics of different areas, more personalized management and optimization plans can be formulated.
[0066] (6) Transportation risk assessment: Analyzing historical data can help identify potential transportation risks and safety hazards. This analysis helps develop risk management strategies and improve the safety of mine transportation systems.
[0067] 3. Data consistency and integrity:
[0068] Batch processing layer: The batch processing layer can be used to clean, correct and supplement data to ensure data consistency and integrity. This is crucial to ensure data quality in complex environments such as mines, especially when data is collected using devices such as sensors.
[0069] (1) Data cleaning and deduplication: The batch processing layer can detect and process errors, anomalies, or redundancies in the data. This may include removing duplicate records, filling missing values, repairing abnormal data, etc., to ensure that the quality of the original data is improved.
[0070] (2) Format normalization: Since data from different sources may use different formats and units, the batch processing layer can normalize the format of the data so that it remains consistent throughout the system for better unified processing.
[0071] (3) Outlier detection and repair: The batch processing layer can detect and handle outliers by applying statistical methods or machine learning algorithms. This helps eliminate unreasonable data introduced due to equipment failure or other reasons and ensures data accuracy.
[0072] (4) Data association and merging: In a mine environment, there may be multiple data sources involving different aspects of information, such as geological data, transportation data, equipment data, etc. The batch processing layer can perform data association and merging to ensure that data from different data sources can be correctly matched in related fields to maintain data integrity.
[0073] (5) Checksum and Validation: The batch processing layer can implement data checksum and validation rules to ensure that the data complies with the expected business rules and standards. This helps prevent non-compliant data from entering the system and improves data consistency.
[0074] (6) Historical data repair: If problems are found in historical data during real-time processing, the batch processing layer can be responsible for repairing and updating the historical data. This ensures the consistency of historical data and also provides a reliable foundation for real-time data analysis.
[0075] (7) Metadata management: The batch processing layer can maintain metadata about the data, including information such as the source of the data, update time, quality assessment, etc. This helps to establish data traceability and improve the transparency of data management.
[0076] 4. Decision support:
[0077] Real-time layer and batch processing layer: Lambda architecture can provide support for real-time decision-making and long-term strategic planning. The real-time layer can support immediate and urgent decisions, while the batch processing layer supports in-depth analysis of historical data and provides a basis for long-term decision-making.
[0078] (1) Instant Decision-making and Emergency Response (Real-time Layer): The real-time layer processes the location, status, and other key data of vehicles in real time, allowing decision-makers to make immediate decisions quickly, especially in emergency situations such as accidents, equipment failures, or emergencies. This layer also supports real-time notifications and alerts to ensure that decision-makers can respond quickly.
[0079] (2) Real-time operation monitoring: Using the dashboard and visualization tools of the real-time layer, decision makers can monitor the operation status of the fleet in real time. This includes real-time information such as vehicle location, speed, and load status, providing real-time support for scheduling, route planning, and operation optimization.
[0080] (3) Long-term strategic planning and optimization (batch layer): The batch layer is responsible for in-depth analysis of historical data to provide a basis for long-term strategic planning. Decision makers can use the results of historical data analysis to identify transportation trends, equipment performance patterns, and mine operation patterns. This provides key information for long-term planning, resource allocation, and equipment updates.
[0081] (4) Predictive analysis and optimization: The batch processing layer can apply advanced analytical techniques, such as machine learning algorithms, to mine historical data to predict future trends and transportation needs. This enables decision makers to make more accurate predictions, thereby optimizing fleet scheduling and improving production efficiency.
[0082] (5) Resource optimization and cost management: Through historical data analysis at the batch layer, decision makers can better understand resource utilization, transportation costs, and equipment maintenance costs. This helps optimize resource allocation, reduce operating costs, and improve overall benefits.
[0083] (6) Develop strategies and specifications: The batch layer can help develop strategic directions and specifications, such as developing transportation safety strategies, energy efficiency standards, equipment replacement plans, etc. This helps ensure that the mine transportation system remains efficient, safe and sustainable in the long term.
[0084] 5. Flexibility and scalability:
[0085] Lambda architecture: By separating real-time and batch processing, Lambda architecture provides flexibility and scalability. You can increase the processing capacity of the real-time layer or batch layer as needed to adapt to changes in the amount of data and processing requirements in the mine.
[0086] (1) Modular architecture: The modular design of the Lambda architecture allows each layer to be scaled independently. The real-time layer and batch layer are independent components that can be scaled individually based on demand without affecting the stability of the overall system.
[0087] (2) Elastic Scaling: Mine transportation systems may face a dramatic increase in data volume, especially during certain periods, such as peak transportation periods. With the Lambda architecture, you can leverage the elastic scaling features of cloud computing to automatically or manually adjust computing resources as needed to ensure that the system can maintain efficient operation when processing large amounts of real-time and historical data.
[0088] (3) Rapid deployment of new features: The requirements of the mine transportation system may change over time. With the Lambda architecture, you can develop and deploy new real-time processing logic or batch processing jobs relatively independently without making large-scale modifications to the entire system. This helps to quickly adapt to new business needs and technological changes.
[0089] (4) Technology diversity: Lambda architecture does not limit the specific technology stack. You can choose the real-time processing engine and batch processing framework that suits your specific needs. This technology diversity makes the system more flexible and allows you to choose the most suitable technology components according to specific scenarios.
[0090] (5) Horizontal expansion and load balancing: Both the real-time layer and the batch processing layer can increase processing capacity through horizontal expansion, while ensuring system stability through load balancing. This helps to process large-scale data and maintain high performance.
[0091] (6) Fault tolerance: The Lambda architecture takes fault tolerance into consideration, allowing the use of multiple nodes in the real-time layer and batch layer. If a node fails, the system can still maintain some functions without completely crashing, ensuring the availability and stability of the system.
[0092] (7) Integration of new data sources: The mine transportation system may need to integrate data from new sensors, equipment, or other data sources. The flexibility of the Lambda architecture makes it relatively easy to integrate new data sources without affecting the existing system architecture.
[0093] 6. Equipment health monitoring:
[0094] Real-time layer: The real-time layer of the Lambda architecture can be used to monitor the real-time status of transport vehicles, such as temperature, humidity, vibration, etc. This is critical for timely detection of abnormal conditions of equipment and corresponding maintenance and care.
[0095] (1) Real-time sensor data monitoring: Use the real-time layer to monitor the data of device sensors, including but not limited to temperature, humidity, vibration, pressure, etc. These data can reflect the current operating status and environmental conditions of the device in real time.
[0096] (2) Real-time capture of abnormal events: The real-time layer can set thresholds and rules. Once the sensor data exceeds the set range or an abnormality occurs, the system can immediately capture and generate corresponding real-time event notifications so that emergency actions can be taken.
[0097] (3) Real-time feedback on health status: Provide real-time feedback on the health status of equipment to operators or maintenance teams. Through visualization methods, such as real-time dashboards, decision makers can understand the current status of all transport vehicles at a glance.
[0098] (4) Predictive maintenance reminders: Combining real-time data with advanced analytics, the real-time layer can generate predictive maintenance reminders. This means the system can predict when equipment is likely to fail, thereby notifying relevant personnel in advance to perform maintenance and reduce downtime.
[0099] (5) Equipment operation trend analysis: The real-time layer analysis function can analyze the equipment operation trend in real time and find possible problem patterns. This is of great significance for improving equipment design, formulating better maintenance strategies and increasing equipment life.
[0100] (6) Remote monitoring and operation: Based on real-time layer monitoring, remote monitoring and operation of equipment can be realized. In this way, even at the edge of the mine, the equipment can be responded to and controlled in real time, improving the flexibility and response speed of the transportation system.
[0101] (7) Alarm and notification system: An integrated alarm and notification system ensures that when an emergency occurs in the equipment, relevant personnel can be notified immediately so that the problem can be handled in a timely manner.
[0102] Offline data warehouse: To put it simply, offline data warehouse is the traditional data warehouse. Data is calculated and stored in the form of T+1, and the calculated data is provided to various front-end analysis applications. In the era of big data, this model is called "batch processing of big data."
[0103] It’s just that the original single environment tools (Oracle, Informatica, etc.) have basically been replaced by tools within the big data system (Hadoop, Hive, Sqoop, Oozie, etc.).
[0104] Data collection: Flume / logstash+kafka, replacing the FTP of traditional data warehouses;
[0105] Batch data synchronization: Sqoop, Kettle. Like traditional data warehouses, Kettle is used. Some commercial ETL tools also begin to support big data clusters.
[0106] Big data storage: Hadoop HDFS / Hive, TiDB, GP and other MPPs, replacing Oracle, MySQL, MS SQL, DB2 and other traditional data warehouses; Big data computing engines: MapReduce, Spark, Tez, replacing the database execution engines of traditional data warehouses;
[0107] OLAP engine: Kylin / druid (Molap, pre-calculation required), Presto / Impala (Rolap, no pre-calculation required), replacing various BI tools such as BO, Brio, and MSTR.
[0108] Real-time data warehouse: Real-time data warehouse was first widely used in log data analysis business. Later, it was promoted by various real-time battle report screens. Compared with offline computing, real-time computing reduces data landing and replaces the data computing engine. Currently, pure streaming data processing basically only has Spark Streaming, and Flink is a batch-stream integration. After the real-time data calculation results are completed, they can be landed in various databases or directly connected to the screen for display.
[0109] Kappa architecture: The design concept of Kappa architecture is that all data is stream-based. The data source of stream computing is the message queue. All data that needs to be calculated is placed in the message queue, and then the stream computing engine is used to calculate all data. Because all data exists in Kafka, the Flink batch-stream integrated data processing engine calculates the data from Kafka and stores it in table n at the service layer. If the demand changes, the offset of Kafka is adjusted, and Flink restarts a task to recalculate and store it in table N+1. When the data progress of N+1 catches up with that of table n, the task of table n is stopped.
[0110] The amount of data cached by the message middleware and the backtracking data have performance bottlenecks. Usually the algorithm requires data from the past 180 days. If all of this data is stored in the message middleware, it will undoubtedly put a lot of pressure on it. At the same time, backtracking and correcting 180 days of data at one time also consumes a lot of resources for real-time computing.
[0111] When processing real-time data and associating a large number of different real-time streams, it is very dependent on the capabilities of the real-time computing system, and it is very likely that data will be lost due to problems with the order of the data streams.
[0112] When Kappa abandoned the offline data processing module, it also abandoned the more stable and reliable characteristics of offline computing.
[0113] Fault tolerance and consistency:
[0114] The Kappa architecture has only one real-time streaming layer, and any failure in the real-time layer may cause data loss or inconsistency. The Lambda architecture can provide stronger fault tolerance through the combination of batch layer and real-time layer. The batch layer is used to store and process historical data, while the real-time layer is responsible for real-time processing. Since there are two layers, the system can continue to operate even if one layer fails.
[0115] Maintainability and debugging:
[0116] Kappa architecture: Kappa architecture has only one real-time streaming layer, which usually requires online debugging and troubleshooting, which may increase the complexity of debugging. The batch processing layer in Lambda architecture can be used for offline debugging and analysis, which makes troubleshooting and debugging more convenient. At the same time, the data of the batch processing layer stores historical records, which is convenient for backtracking and analyzing historical data.
[0117] Flexibility and evolution:
[0118] Once the real-time streaming layer in the Kappa architecture is deployed, it is difficult to modify historical data because it focuses on processing real-time streams. The coal mine new energy vehicle auxiliary transportation data does not require high data real-time performance. The batch processing layer of the Lambda architecture allows offline processing and modification of data, which is very helpful for the evolution and update of data models. New batch processing tasks can be applied to historical data, while the real-time layer is applicable to real-time data.
[0119] To address this problem, the present application provides a mine auxiliary transportation data processing system. Figure 1 This is a schematic diagram of the structure of a mine auxiliary transportation data processing system provided in an embodiment of the present application. Figure 1 As shown, the system includes: a business data collection module, a log data collection module, a real-time processing module, an offline processing module, and a calculation module;
[0120] The business data collection module and the log data collection module are connected to the real-time processing module, the business data collection module and the log data collection module are connected to the offline processing module, and the real-time processing module and the offline processing module are both connected to the calculation module;
[0121] The log data collection module is used to collect log data of the mine auxiliary transportation system, wherein the log data is used to record the status of the system;
[0122] The business data collection module is used to collect business data of the mine auxiliary transportation system, wherein the business data is used to record business related information;
[0123] The real-time processing module is used to obtain the real-time processing data in the business data collection module and the log data collection module and perform real-time processing;
[0124] The offline processing module is used to obtain the offline processing data in the business data collection module and the log data collection module and perform offline processing;
[0125] The calculation module is used to analyze and process the real-time processing data and the offline processing data.
[0126] Optionally, the log data collection module includes: a point collection module, a first reverse proxy module, a log server module, a first aggregation transmission module, a first persistence module, and a second aggregation transmission module;
[0127] The buried point collection module is used to collect the log data at the front buried point of the mine auxiliary transportation system;
[0128] The first reverse proxy module is used to send and store the log data of the embedding point collection module in the log server module, and is also used to schedule the log data in the log server module and send it to the first aggregation transmission module;
[0129] The first aggregation transmission module is connected to the tracking point collection module and is used to send the log data to the first persistence module;
[0130] The first persistence module is used to perform persistence processing on the log data and send it to the second aggregation transmission module;
[0131] The second aggregation transmission module is used to send the log data to the real-time processing module or the offline processing module.
[0132] Optionally, the business data collection module includes: a business data collection submodule, a second reverse proxy module, a business server module, a structured query module, a channel module, and a data synchronization module;
[0133] The business data collection submodule is used to intercept business data of the mine system energy auxiliary transportation system;
[0134] The second reverse proxy module sends the business data of the business data collection submodule to the business server module and stores it, and is also used to schedule the business data in the business server module and send it to the structured query module;
[0135] The structured query module is connected to the second reverse proxy module and is used to perform preliminary processing on the business data;
[0136] The channel module is connected to the second reverse proxy module and is used to send the business data in the business server module to the first persistence module;
[0137] The data synchronization module is used to synchronize the business data in the structured query module to the real-time processing module or the offline processing module.
[0138] Optionally, the offline processing module includes: a distributed file module, a data warehouse module, an offline synchronization module, and a database module;
[0139] The distributed file module is used to store the offline processing data;
[0140] The data warehouse module is connected to the distributed file module and is used to perform hierarchical processing on the offline processing data stored in the distributed file module to obtain structured offline processing data;
[0141] The offline synchronization module is connected to the data warehouse module and the database module respectively, and is used to synchronize the structured offline processing data in the data warehouse module to the database module;
[0142] The data warehouse module is also used to send the structured offline processing data to the calculation module.
[0143] Optionally, the real-time processing module includes: a second persistence module, a stream computing module, and a dimension database module;
[0144] The second persistence module is used to receive the real-time processing data, perform persistence processing on the data, and send the data to the stream computing module;
[0145] The dimension database module is connected to the second persistence module, and the dimension database module includes a dimension table and state information;
[0146] The second persistence module is connected to the stream computing module, and the stream computing module is used to receive the data sent by the second persistence module and perform stream computing according to the dimension table and the state information to structure the real-time processing data, and to rewrite the structured real-time processing data into the second persistence module;
[0147] The second persistence module is used to perform data stratification according to the structured real-time processing data;
[0148] The streaming computing module is also used to send the structured real-time processing data to the computing module.
[0149] Optionally, the system further includes a display module, which is connected to the calculation module and is used to visualize the results of analysis and processing by the calculation module.
[0150] In this embodiment, first, a method for building a data collection platform for mine auxiliary transportation vehicles such as trackless rubber-wheeled vehicles
[0151] 1. Set the data format of mine vehicle log
[0152] Motor is a nested field, and its meaning is:
[0153] 2. Establish mine vehicle dimension data
[0154] 3. Environment preparation and framework deployment
[0155] 1) Write a cluster distribution script xsync to copy files to the same directory on all nodes in a loop
[0156] 2) SSH password-free login configuration
[0157] 3) JDK, Zookeeper, Hadoop installation, Kafka installation, Flume installation, MySQL installation
[0158] 4) Balance cluster data: balance data between nodes and between disks
[0159] 4. Building a data collection platform
[0160] (1) Batch and stream processing data collection platform
[0161] Batch processing and stream processing share a common collection platform
[0162] (2) Flume mine auxiliary transportation vehicle log data collection
[0163] Select TaildirSource and KafkaChannel, and configure the log verification interceptor. TailDirSource: breakpoint resume, multiple directories. Using Kafka Channel, eliminating the Sink and improving efficiency.
[0164] (2) Kafka Data Collection Flume
[0165] According to the project plan, we need to use Flume to import the topic_log data in Kafka into HDFS.
[0166] To this end, we selected three components, KafkaSource, FileChannel, and HDFSSink, for data processing. Later, we need to use Flume's interceptor to process the data stored in Kafka, and only KafkaSource can be used with the interceptor.
[0167] (3) Dimensional data collection
[0168] Business data is an important source of data for the data warehouse. We need to extract data from the business database on a daily basis, transfer it to the data warehouse, and then analyze and compile statistics on the data.
[0169] Use DataX for data synchronization.
[0170] 5. Offline data warehouse architecture design
[0171] Offline architecture description:
[0172] 1. There are two main ways to collect data on the auxiliary transportation of new energy vehicles in mines. One is to use the front-end Web embedding point of the mine auxiliary transportation system to collect log data such as system operation logs; the other is to use the mine business server to collect business data of the mine energy auxiliary transportation system.
[0173] 2. As a reverse proxy, Nginx can hide the specific architecture and topology of the mine new energy vehicle data collection platform and provide a unified access point to the outside world. External requests first arrive at the Nginx server, and then Nginx forwards the request to the back-end data collection server. In addition, Nginx can distribute data collection requests from new energy vehicles to multiple back-end servers to achieve load balancing. This ensures that the system can effectively handle a large number of concurrent requests and improves the stability and performance of the system.
[0174] 3. For log data collection: Flume-Kafka-Flume data collection architecture is often used to efficiently transfer log and event data from data sources to data storage in distributed systems.
[0175] Specifically:
[0176] Flume cluster:
[0177] The Flume cluster is responsible for collecting data from various devices in mine new energy-assisted transportation, including sensor data of electric transport vehicles, charging pile logs, vehicle GPS data, etc.
[0178] The collected data is transmitted to the Kafka cluster via Flume for subsequent processing and storage.
[0179] Specific process:
[0180] Source: Flume's Source component obtains data from electric transport vehicles, charging piles and other equipment. For example, it obtains information such as battery status, speed, and location from vehicles, and obtains charging logs from charging piles.
[0181] Channel: Source puts the collected data into Flume's Channel. This can be a persistent channel to ensure smooth data transmission even under high load.
[0182] Interceptor: You can use interceptors to process data, such as filtering invalid data, converting data formats, etc.
[0183] Sink (data target): Data is transmitted from Channel to Sink, and Sink sends the data to the corresponding topic in the Kafka cluster.
[0184] Kafka Cluster:
[0185] Kafka, as a message queue, receives data from Flume and persists it for subsequent processing and analysis. Kafka acts as a buffer between Flume and downstream processing systems, ensuring reliable data delivery while decoupling the speed of each component.
[0186] process:
[0187] Producer: Flume Sink acts as a Kafka Producer, sending the data of mine new energy auxiliary transportation to a specific Topic in Kafka.
[0188] Topic: Topics in Kafka are used to organize data. Different types of data are sent to different Topics. For example, one Topic is used to store vehicle sensor data, and another is used to store charging station logs.
[0189] Consumer: Downstream data processing system, which may be some real-time processing application or data storage system, acts as a Kafka Consumer and consumes data from different Topics.
[0190] 4. For business data collection: directly import it into the distributed storage HDFS through Sqoop.
[0191] 5. Perform offline data warehouse stratification in Hive (the stratification part has been explained in detail elsewhere, and building tables and modeling various data is the main work), using SparkSQL.
[0192] 6. Use DataX to synchronize data to the MySQL database
[0193] 7. Use ad hoc query frameworks such as Druid to perform ad hoc data queries.
[0194] 8. The data publishing interface is displayed on the data visualization platform.
[0195] Real-time architecture description:
[0196] 1. The data collection part is consistent with the offline architecture. The difference is that the offline architecture collects data into HDFS, while the real-time database collects data into Kafka.
[0197] 2. Flink consumes data in Kafka and writes back the data to achieve the effect of data stratification.
[0198] As a streaming computing engine, Flink can consume data in real time from message queues such as Kafka. In the data layering scenario, we can consume raw data from Kafka, which may contain information of different levels and types, such as vehicle behavior data, vehicle equipment data, etc.
[0199] The consumed data is processed by Flink and can be cleaned, transformed, aggregated, etc. These processing operations are designed to convert the raw data into a more structured and analyzable format for subsequent analysis and storage.
[0200] The processed data is stored in layers according to business needs (dwd, dws, etc.)
[0201] 3. The dimension table and status information of the mine new energy vehicle auxiliary operation system are maintained in Redis and HBase.
[0202] Specifically:
[0203] Operation process in Redis:
[0204] 1. Operation of dimension table:
[0205] a. Insert dimension table records: In the business, suppose there is a dimension table of vehicle types, including electric vehicles and traditional fuel vehicles. When inserting records, you can use vehicle type, maximum speed, mileage and other information as attributes and store them using Redis's Hash structure. For example, inserting information about electric vehicles and traditional fuel vehicles:
[0206] HMSET vehicle_types:electric_cartype_name"Electric Car"max_speed120mileage_per_charge 300
[0207] HMSET vehicle_types:gasoline_cartype_name"Gasoline Car"max_speed160fuel_efficiency 25
[0208] This way, detailed information for each vehicle type can be easily queried.
[0209] b. Query dimension table records: Through the HGETALL command, you can query all attribute information of electric vehicles and traditional fuel vehicles. This information provides a basis for the interpretation and analysis of business data.
[0210] c. Update dimension table records: If the maximum speed of the electric vehicle changes, you can use the HSET or HMSET command to update this information to ensure that the data in the dimension table remains up to date.
[0211] d. Delete dimension table records: If traditional fuel vehicles are no longer used in the system, you can use the DEL command to delete the corresponding dimension table records to ensure timely update of data.
[0212] Summarize:
[0213] In Redis, dimension table operations are mainly completed through the Hash structure. By inserting, querying, updating, and deleting dimension table records, the system can manage business data related to vehicle types more flexibly. These dimension information provides basic support for the interpretation and analysis of business data.
[0214] Operation process in HBase:
[0215] 1. Operation of dimension table:
[0216] a. Insert dimension table records: In HBase, you can use the Put operation to insert dimension table records into the table. For the vehicle type dimension table, you can use electric vehicles and traditional fuel vehicles as row keys, the column family as the attribute type, and the column name as the specific attribute, for example:
[0217] Put'vehicle_types','electric_car','info:type_name','Electric Car'
[0218] Put'vehicle_types','electric_car','info:max_speed','120'
[0219] Put'vehicle_types','electric_car','info:mileage_per_charge','300'
[0220] In this way, the information of each vehicle type will be stored in the HBase table as a row key.
[0221] b. Query dimension table records: Through the Get operation, you can query detailed information about electric vehicles and traditional fuel vehicles, thereby supporting the interpretation and analysis of business data.
[0222] c. Update dimension table records: Use the Put operation to update existing rows. For example, if the maximum speed of an electric car changes, you can execute:
[0223] Put'vehicle_types','electric_car','info:max_speed','150'
[0224] d. Delete dimension table records: Through the Delete operation, you can delete vehicle types that are no longer in use to ensure that the data in the dimension table remains updated.
[0225] Summarize:
[0226] In HBase, dimension table operations are mainly completed through Put, Get, and Delete operations. Dimension information can be efficiently managed through reasonable row key design and organization of column families and column names. These dimensional information provides basic support for the interpretation and analysis of business data of the mine new energy vehicle auxiliary transportation system.
[0227] 4. Flink consumes the data of the mine new energy vehicle auxiliary transportation system from Kafka to ClickHouse or ES, and builds a DWS layer. Through this design, the Flink job in the real-time data warehouse can process the real-time data from Kafka and write it to ClickHouse and Elasticsearch respectively to build a DWS layer to support real-time query and analysis of the business. ClickHouse provides high-performance, high-throughput real-time data storage and query services, which are suitable for OLAP scenarios, while Elasticsearch can be used for full-text retrieval and complex query support, which is suitable for some scenarios that require real-time text search.
[0228] 5. Publish to the visualization project.
[0229] Lambda overall architecture description:
[0230] 1. Only one data collection platform is built. After the data is collected to Kafka, real-time architecture processing is performed. At the same time, the data in Kafka is synchronized to HDFS for offline processing.
[0231] 2. The data in Kafka is consumed by Flink and real-time data warehouse layered processing is built in Kafka; the data in HDFS is consumed by HiveSpark and offline data warehouse layered architecture is built in Hive.
[0232] In a possible embodiment, the content of the user behavior log mainly includes the user's various behavior information and the environment information of the behavior. The main purpose of collecting this information is to optimize the product and provide data support for various analysis and statistical indicators. The means of collecting this information is usually to embed points.
[0233] Adopt code tracking, which is to call the tracking SDK function, call the interface at the business logic function location where tracking is required, and report tracking data. For example, after we track a button on the page, when the button is clicked, we can call the data sending interface provided by the SDK in the OnClick function corresponding to the button to send data.
[0234] 2. Technical architecture: Flume, Kafka
[0235] Log collection Flume needs to collect the contents of log files, verify the log format (JSON), and then send the verified logs to Kafka. As shown in the following figure
[0236] 1) First, the APP and WEB are used to track data -> generate logs -> send them to the background log server -> perform a load through Nginx -> then to the collection system. Here we use three machines as our log servers; 4. After passing through the Nginx proxy server, each log server will store the log information on its own disk. For log collection, we chose Flume. There are many log collection frameworks. The main reason for choosing Flume is that it has a better data collection effect, and secondly, it also has better support for HDFS and Kafka.
[0237] 2) Flume mainly consists of three components: Source, Channel, and Sink;
[0238] At the Flume layer, we also made an interceptor, which mainly filters the collected logs. Because the logs are stored in the Json format in the background, a simple cleaning of the Json with illegal format is performed in the interceptor.
[0239] We also made a classification interceptor, in which we distinguish the data types. We mainly made a labeling function to label different log data with different labels, and then put the data with different labels into different topics through the subsequent selector Multiplexing, which is convenient for downstream data processing.
[0240] 3) Downstream data transmission uses Kafka as a message queue to transmit messages. The main reason for using Kafka is because of Kafka's high throughput and the ability to classify data into different topics, which is convenient for the next layer to use.
[0241] Kafka is a distributed stream processing platform that uses a publish-subscribe model and is mainly used for real-time data transmission and processing. The following is the workflow of Kafka:
[0242] Producer sends a message:
[0243] The producer is responsible for generating messages and publishing them to topics in the Kafka cluster.
[0244] A topic is a logical container for messages. Producers can publish messages to one or more topics.
[0245] Broker (proxy server) receives the message:
[0246] A Kafka cluster consists of multiple Brokers, each of which is responsible for receiving and storing messages.
[0247] The messages sent by the producer will be distributed to different partitions, and each partition stores a part of the message data.
[0248] ③Partition:
[0249] A topic can be divided into several partitions, each of which stores a specific range of messages.
[0250] Partitioning allows Kafka clusters to scale horizontally, increasing concurrency and capacity.
[0251] Consumer Group:
[0252] Consumers read messages from topics in the form of consumer groups.
[0253] A consumer group can contain one or more consumers, each of which is responsible for processing messages from one or more partitions.
[0254] ⑤Consumer receives the message:
[0255] A consumer subscribes to one or more topics and then pulls messages from its assigned partitions.
[0256] Consumers can track the offset of consumed messages so that they can checkpoint when needed.
[0257] 4) Message storage:
[0258] Kafka uses persistent storage to save messages to ensure the persistence and reliability of messages.
[0259] When messages are stored, they can be deleted after a certain period of time or after reaching a certain size according to the configured retention policy.
[0260] Zookeeper coordination:
[0261] Zookeeper is used to manage and coordinate the nodes of the Kafka cluster.
[0262] It maintains the metadata of the cluster, monitors the health of the Broker, and is responsible for the allocation of partitions and consumers.
[0263] 5) Message distribution and load balancing:
[0264] Kafka achieves message distribution and consumer load balancing through partitioning and partition rebalancing mechanisms.
[0265] When consumers join or leave the consumer group, Kafka redistributes partitions to ensure that each consumer is responsible for processing an even number of partitions.
[0266] 6) Consumer confirmation and offset management:
[0267] The consumer tells Kafka that it has successfully consumed a message by confirming the message.
[0268] The offset of each consumer is maintained by Kafka, and consumers can submit their offsets periodically.
[0269] 7) Horizontal expansion:
[0270] Kafka supports horizontal expansion, which can increase the capacity and throughput of the system by adding Brokers, partitions, and consumers.
[0271] (1)ODS layer design
[0272] The table structure design of the ODS layer relies on the data structure synchronized from the business system.
[0273] The ODS layer needs to save all historical data, so its compression format should be selected with higher compression. Here, gzip is selected.
[0274] The naming convention for ODS layer table names is: ods_table name_single-partition incremental full identifier (inc / full).
[0275] (2) DIM layer design
[0276] The design of the DIM layer is based on dimensional modeling theory, and this layer stores the dimension tables of the dimensional model.
[0277] The data storage format of the DIM layer is orc column storage + snappy compression.
[0278] The naming convention for DIM layer table names is dim_table name_full table or zip table identifier (full / zip)
[0279] (3) DWD layer design
[0280] The design of the DWD layer is based on dimensional modeling theory, and this layer stores the fact table of the dimensional model.
[0281] The data storage format of the DWD layer is orc column storage + snappy compression.
[0282] The naming convention for DWD layer table names is dwd_data domain_table name_single partition incremental full quantity identifier (inc / full)
[0283] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0284] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0285] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0286] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0287] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.
[0288] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.
[0289] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0290] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0291] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0292] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0293] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0294] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0295] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0296] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A mine auxiliary transportation data processing system, characterized in that: include: Business data collection module, log data collection module, real-time processing module, offline processing module, and calculation module; The business data collection module and the log data collection module are connected to the real-time processing module, the business data collection module and the log data collection module are connected to the offline processing module, and the real-time processing module and the offline processing module are both connected to the calculation module; The log data collection module is used to collect log data of the mine auxiliary transportation system, wherein the log data is used to record the status of the system; The business data collection module is used to collect business data of the mine auxiliary transportation system, wherein the business data is used to record business related information; The real-time processing module is used to obtain the real-time processing data in the business data collection module and the log data collection module and perform real-time processing; The offline processing module is used to obtain the offline processing data in the business data collection module and the log data collection module and perform offline processing; The calculation module is used to analyze and process the real-time processing data and the offline processing data.
2. The system according to claim 1, characterized in that The log data collection module includes: a point collection module, a first reverse proxy module, a log server module, a first aggregation transmission module, a first persistence module, and a second aggregation transmission module; The buried point collection module is used to collect the log data at the front buried point of the mine auxiliary transportation system; The first reverse proxy module is used to send and store the log data of the embedding point collection module in the log server module, and is also used to schedule the log data in the log server module and send it to the first aggregation transmission module; The first aggregation transmission module is connected to the tracking point collection module and is used to send the log data to the first persistence module; The first persistence module is used to perform persistence processing on the log data and send it to the second aggregation transmission module; The second aggregation transmission module is used to send the log data to the real-time processing module or the offline processing module.
3. The system according to claim 2, characterized in that The business data collection module includes: a business data collection submodule, a second reverse proxy module, a business server module, a structured query module, a channel module, and a data synchronization module; The business data collection submodule is used to intercept business data of the mine system energy auxiliary transportation system; The second reverse proxy module sends the business data of the business data collection submodule to the business server module and stores it, and is also used to schedule the business data in the business server module and send it to the structured query module; The structured query module is connected to the second reverse proxy module and is used to perform preliminary processing on the business data; The channel module is connected to the second reverse proxy module and is used to send the business data in the business server module to the first persistence module; The data synchronization module is used to synchronize the business data in the structured query module to the real-time processing module or the offline processing module.
4. The system according to claim 1, characterized in that The offline processing module includes: a distributed file module, a data warehouse module, an offline synchronization module, and a database module; The distributed file module is used to store the offline processing data; The data warehouse module is connected to the distributed file module and is used to perform hierarchical processing on the offline processing data stored in the distributed file module to obtain structured offline processing data; The offline synchronization module is connected to the data warehouse module and the database module respectively, and is used to synchronize the structured offline processing data in the data warehouse module to the database module; The data warehouse module is also used to send the structured offline processing data to the calculation module.
5. The system according to claim 1, characterized in that The real-time processing module includes: a second persistence module, a stream computing module, and a dimension database module; The second persistence module is used to receive the real-time processing data, perform persistence processing on the data, and send the data to the stream computing module; The dimension database module is connected to the second persistence module, and the dimension database module includes a dimension table and state information; The second persistence module is connected to the stream computing module, and the stream computing module is used to receive the data sent by the second persistence module and perform stream computing according to the dimension table and the state information to structure the real-time processing data, and to rewrite the structured real-time processing data into the second persistence module; The second persistence module is used to perform data stratification according to the structured real-time processing data; The streaming computing module is also used to send the structured real-time processing data to the computing module.
6. The system according to claim 1, characterized in that The system further comprises a display module, which is connected to the calculation module and is used for visualizing the results of the analysis and processing by the calculation module.
7. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the system according to any one of claims 1 to 6.
8. A computer program product, characterized in that The invention comprises a computer program, which implements the system according to any one of claims 1 to 6 when being executed by a processor.