Equipment data storage Dores method based on industrial internet
Through the Doris method of equipment data storage based on the industrial Internet, the problems of high concurrent writes and real-time queries are solved, low-latency, high-quality data storage and automated management are realized, and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510873588.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing technology is difficult to take into account high concurrent writes, second-level analysis, controllable storage costs and flexible partition management, resulting in the problem of writing performance bottlenecks, complex system architecture and high operation and maintenance costs in equipment data storage.
The Doris method based on the industrial Internet is adopted to realize high concurrent write and real-time query through real-time acquisition, preprocessing, and uploading to Kafka and Doris consumption and storage, combined with dynamic partitioning and automated partition management.
Significantly reduce data processing delays, improve real-time and data quality, consistency, reduce operation and maintenance costs, and improve system scalability and reliability.
Smart Images

Figure CN120386780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a Doris method for device data storage based on the industrial Internet. Background Art
[0002] With the development of the industrial Internet and the Internet of Things, the amount of device data collected has increased sharply, and the requirements for real-time performance, reliability, and horizontal scalability are increasing day by day.
[0003] Existing solutions mainly include: directly writing data into a relational database, big data storage based on the Hadoop ecosystem, the integration of a message middleware + database, and a dedicated time series database. However, these solutions have obvious deficiencies: relational databases are prone to performance bottlenecks under high-concurrency writes, and need to be sharded or expanded, resulting in complex costs and operations; big data solutions require the deployment of multiple clusters such as Kafka, HDFS, YARN, Hive / Spark, etc., with complex architectures and high operation and maintenance costs, and HDFS multi-copy storage wastes resources and has poor real-time performance; the message middleware + database solution requires additional consumer programs, with complex consumer-side logic, limited scalability, and inconsistent historical data archiving, which is error-prone; although the dedicated time series database is optimized for time series data, the cluster performance is prone to bottlenecks in the scenario of massive concurrent writes, and the retention policy in the TTL method lacks flexibility and cannot meet the differentiated needs of different devices and business scenarios. In addition, the time series database may also have problems such as data skew and query timeouts, increasing the risk of system instability.
[0004] Generally speaking, the existing technology is difficult to simultaneously balance high-concurrency writes, second-level analysis, controllable storage costs, and flexible partition management. Therefore, there is an urgent need for a new method for storing device data collection data to solve the above defects. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems existing in the storage of device data collection data in the prior art, such as write performance bottlenecks, complex system architectures, high storage costs, and inflexible partition management, and to propose a Doris method for device data storage based on the industrial Internet, which realizes high-concurrency writes, real-time queries, and automated data archiving management.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A Doris method for device data storage based on the industrial Internet includes the following specific steps: S1: Obtain industrial device node data: Real-time collect the data of on-site PLC or gateway devices to prepare for subsequent cleaning and storage; S2: Preprocess the obtained data: Remove duplicates, fill in missing values, denoise, and standardize the original collected data to ensure the data quality and consistency of subsequent data storage; S3: Upload the preprocessed data to Kafka: Implement data buffering and decoupling through a distributed message queue. The downstream can consume in parallel and has good scalability; S4: Doris consumes Kafka data and stores it: Let Doris pull JSON data from Kafka in real time and write it into a relational table. At the same time, use dynamic partitioning to achieve automatic partitioning and expiration cleaning of data; S5: Monitoring and performance optimization: Monitor the end-to-end link from data collection to storage, promptly discover bottlenecks and anomalies, and perform targeted optimizations; S6: Expansion: Secondary utilization and downstream analysis. On the premise of ensuring data quality and real-time performance, support subsequent data warehouse docking, real-time computing, and machine learning applications.
[0007] As a further technical solution of the present invention, the S1 specifically includes: S11: Selection of collection method: Direct collection by PLC or collection through gateway forwarding; S111: Direct collection by PLC: The on-site PLC establishes communication protocols such as Modbus (TCP / RTU), PROFINET, EtherNet / IP, etc. with the upper computer through Ethernet or serial port. The system sends a read request to the PLC according to the preset sampling frequency (such as once per second or once per ten seconds) to obtain the real-time values within the specified register or data block; It is necessary to confirm the register range (starting address and ending address) read each time of collection, which is configured by the equipment manufacturer or project implementers according to actual requirements; S112: Collection through gateway forwarding: If some on-site devices do not support traditional PLC protocols, it is necessary to collect and parse the data of the device's own protocol through an industrial edge gateway (such as an OPC UA gateway, Modbus gateway, Siemens S7 gateway, etc.), and then convert the parsed original information (usually key-value pairs or custom binary formats) into a unified data format (such as JSON), and forward it to the backend system; S12: Transmission protocol and upload method: Upload through MQTT or upload through HTTP interface; S13: Timestamp and unique identifier processing: In the collection end and the gateway, a consistent time reference should be used. Generally, a millisecond-level Unix timestamp is adopted, that is, the number of milliseconds from January 1, 1970 to the current moment. This can ensure that the data uploaded by different devices has a unified standard in time sorting; For the unique identifier of the device, if the device itself does not have a globally unique ID, it can be composed of a fixed prefix (such as the gateway number) and the device's local number spliced on the gateway. For example, "GW01_PLC05" can be used as the globally unique number of the device. This can ensure that there will be no device confusion during subsequent processing and storage.
[0008] As a further technical solution of the present invention, S12 specifically includes: S121: MQTT upload: Using the lightweight MQTT protocol, the gateway or the acquisition end publishes the processed data in the form of a message to a specified topic. The topic can be hierarchically organized according to "enterprise number / factory number / workshop number / equipment number". For example: / CompanyA / Factory1 / Line2 / PLC_001; The QoS (Quality of Service) usually has two optional levels: "at least once" or "exactly once". "At least once" ensures that the message will not be lost, but may be repeated; "exactly once" consumes more resources; The retain message flag is generally set to "not retain" to avoid existing subscribers receiving outdated historical messages; The information carried during reporting includes the equipment number, timestamp, and sensor readings, etc.; If the connection is disconnected or the server returns an error, the number of retries and the retry interval can be set, and an alarm or local backup can be written after multiple attempts fail to prevent data loss; S122: HTTP interface upload: Send the collected data to the backend service through an HTTP POST request; Example of the interface address: http: / / <backend domain name>:<port> / api / v1 / device / Data; Specify Content-Type: application / json in the HTTP request header, and the request body is in the standard JSON format; If the request fails due to network problems or backend exceptions, the client will try again according to the pre-set retry policy (such as retrying at most three times, with a half-second interval each time); After all attempts fail, record the error information in the local log and trigger an alarm to notify the operation and maintenance personnel.
[0009] As a further technical solution of the present invention, S2 specifically includes: S21: Duplicate data filtering and invalid data discarding; S211: Duplicate data filtering: If the same device uploads the same timestamp and the same data value to the backend multiple times within a short period of time, it is regarded as duplicate data. After receiving the data, the system first checks whether the device number and the timestamp are exactly the same as the most recent one. If they are the same, the subsequent data can be directly discarded to avoid redundant storage; S212: Discard invalid data: The messages uploaded by the device often contain a "status code" field, which is used to indicate whether the sampling status is normal. If the status code is not "normal", it means the data is invalid and should be directly discarded without entering the subsequent processing flow. In addition, for key physical quantities such as temperature and pressure, if the values significantly exceed the reasonable range defined in the product manual or business requirements (e.g., the temperature is greater than its rated upper limit or lower than the rated lower limit), they are also regarded as abnormal data, discarded and logged in the log. S22: Fill in and interpolate missing data: To ensure the continuity of time-series data, it is sometimes necessary to fill in the missing time points during the acquisition process. A common approach is to check the time span between two adjacent valid samplings of the same device. If the time interval does not exceed the pre-set maximum interval (e.g., within 5 seconds), the missing points are estimated using the linear relationship between the two previous and next valid values. If the interval is too large, interpolation is no longer performed, and instead, the previous valid value or a null value is directly filled to avoid error accumulation. For simple cases, "nearest value filling" is adopted, directly using the nearest valid sampling value to the missing time point. If more accuracy is required, advanced interpolation algorithms (such as polynomial interpolation, spline interpolation, etc.) can also be used at the application layer, but the implementation cost and computational overhead are relatively higher and are less used in general industrial IoT scenarios. S23: Data cleaning and standardization: It includes unit unification, noise filtering, and field renaming and supplementation. S231: Unit unification: All physical quantities need to be converted to standard units. For example, if the temperature unit uploaded by the on-site device is Fahrenheit, the system should convert it to Celsius before proceeding to the next step to ensure consistency in subsequent analysis and storage. The same concept applies to various physical quantities such as pressure, humidity, current, and voltage. S232: Noise filtering: Due to noise interference in on-site acquisition, it is sometimes necessary to smooth the data. Common methods include moving average or median filtering. For example, the values of the most recent N samplings are averaged to eliminate short-term pulse interference, or the median of N samplings is taken to further remove the influence of extreme values. S233: Field renaming and supplementation: If the downstream system (such as Kafka producer) requires fixed field names and formats, the original field names must be uniformly renamed. For example, the "temp" field is changed to "temperature" after cleaning, and "press" is changed to "pressure". If the message lacks some predefined fields, default values can be supplemented (e.g., the numerical value is set to 0 or the string is set to "unknown"), or null values can be set, depending on the business requirements. S24: Package data into JSON: After the above cleaning, the data is organized into a unified JSON structure, which usually contains the following parts: device_id: The unique identifier of the device; timestamp: Unix timestamp in milliseconds when the server receives it; data: key-value pairs of core sensor readings (such as temperature, pressure, humidity, etc.); status: device status code, used to indicate whether the sampling is normal; metadata: optional information, such as auxiliary fields like gateway number, signal strength, collection end IP, etc.; Example JSON (for illustration only): { "device_id": "GW01_PLC05", "timestamp": 171XXXXXXX000, "data": { "temperature": 42.8, "pressure": 1.05, "humidity": 58.2}, "status": 0, "metadata": { "gateway_id": "GW01", "signal_strength": -65}}; After adopting this fixed format, when sending data to Kafka or calling an HTTP interface later, the data in JSON structure can be directly used as the message body, keeping the meanings of each field consistent.
[0010] As a further technical solution of the present invention, the specific steps of S3 include: S31: Kafka cluster and Topic planning: Set up at least 3 Kafka Brokers at the backend to ensure high availability of the system; Design Topic names according to business logic, such as "companyA_ factory1_line2_equipment_data", and gather data from the same production line or the same type of device under the same Topic for centralized processing; According to the peak message volume and the system concurrency ability, configure an appropriate number of partitions for this Topic. The more partitions there are, the stronger the parallel consumption ability of the downstream, but it will also increase the management overhead of the Broker. Usually, first evaluate the system peak message rate (such as 100,000 messages per second), and then combine the maximum throughput of each partition to finally determine an appropriate number of partitions such as 8, 16, 32, etc.; Set the replication factor to 3, that is, each message will have a copy on 3 Brokers to ensure that data can still be read normally even if any two nodes fail; S32: Kafka Producer configuration; S33: Message Sending and Exception Handling: The Producer calls the asynchronous sending interface and listens for the sending result in the callback function. If the sending is successful, record the success log or ignore it. If the sending fails, record the reason for the failure (such as network timeout, Broker unavailable, etc.) in the log and retry according to the retry policy. If it still fails after multiple retries, put the message into the local temporary buffer queue and retry the delivery after the background resumes or manual intervention. Throughput Optimization: Appropriately set the batch sending parameters. For example, accumulate a certain number of messages (such as 500 messages) or reach a certain number of bytes (such as 32KB) each time and then send them all at once to avoid the network overhead caused by frequent small - batch sending. At the same time, "linger.ms" can be set to dozens of milliseconds to let the Producer try to pack the messages arriving in a short period of time before sending.
[0011] As a further technical solution of the present invention, the S32 specifically includes: S321: Broker List: Write the addresses of all Brokers into the Producer's configuration, such as broker1:9092, broker2:9092, broker3:9092; S322: Message Confirmation Mechanism (acks): It is recommended to set it to "all", indicating that the message is considered successfully sent only after it is written to all In - Sync replicas to ensure data is not lost; S323: Number of Retries: Automatically retry when the sending fails, for example, set it to 3 times. If it still fails after multiple retries, trigger an alarm or write to the local log to prevent silent data loss; S324: Serializer: Use string serialization when sending the device number as the Key, and also use UTF - 8 encoding serialization when using the JSON string as the Value; S325: Partition Allocation Policy: Adopt the default hash partition policy, that is, take the modulus of the total number of partitions after hashing the device number. In this way, all data of the same device can fall into the same partition, ensuring the timeliness and sequential consumption of data of the same device.
[0012] As a further technical solution of the present invention, the S4 specifically includes: S41: Doris Table Structure and Dynamic Partition Design: Create a wide table in Doris with fields including: device_id, event_time, temperature, pressure, humidity, status, gateway_id, etc.; The table uses the method of partitioning by date range (Range partitioning) and enables the dynamic partitioning function; The system will automatically create a total of 5 partitions from "3 days before the current date" to "1 day after the current date" according to the current date. The partition naming method is pYYYYMMDD. For example, p20250606 represents the partition for June 6, 2025. This can ensure that there are corresponding partitions for the data of the current day and the recent past days, making queries more efficient; Set the partition retention period, for example, 30 days. That is, when a partition has been created for a certain number of days (such as 30 days), the system will automatically delete the partition to achieve automatic cleaning of expired data, making the management of storage space more flexible; At the same time, perform bucketing (Hash bucketing) design on the table. For example, hash bucket according to device_id and set 16 buckets. In this way, when querying and writing, data from different devices or different partitions can be scattered to multiple nodes in parallel to improve performance; S42: Create a Routine Load Task: Execute a specific SQL statement on Doris to create a Routine Load task, specifying from which Kafka Topic the task pulls data, what concurrency number to use, how many rows or how many bytes to pull at most per batch, and how to parse JSON; The mapping from JSON fields to table fields is provided by a JSONPaths file, and the corresponding relationship is roughly as follows: $.device_id corresponds to device_id in the table; $.timestamp corresponds to event_time, and the timestamp needs to be converted to the DATETIME type of Doris when writing; $.data.temperature, $.data.pressure, $.data.humidity correspond to the temperature, pressure, and humidity fields in the table respectively; $.status corresponds to status in the table; $.metadata.gateway_id corresponds to gateway_id in the table; In the Routine Load task configuration, the number of concurrent consumers (such as 4 concurrent threads), the maximum number of rows (such as 5000 rows) or the maximum number of bytes (such as 10MB) per batch, and the strict mode switch can be specified. After enabling the strict mode, if a field in a certain JSON is missing or the type does not match, it will be considered "dirty data" and discarded, and a log will be recorded to ensure data consistency; S43: Data Loading Process and Dirty Data Handling: Doris Routine Load pulls messages from Kafka partitions in parallel according to the specified concurrency. Whenever messages meeting the threshold (number of rows or bytes) are pulled, they are stored in a temporary file. The system decompresses and parses the temporary file, extracts fields from the JSON according to JSONPaths and writes them into the in-memory table. The data in the in-memory table is automatically allocated to the corresponding partitions according to the partitioning strategy, and then the backend nodes batch-persist the data to the column store. For records with parsing failures, mismatched field types, or missing required fields, Routine Load will put the record into the "dirty data" queue and decide whether to ignore or alarm according to the "maximum dirty data ratio" threshold. If the dirty data ratio of a certain batch exceeds the threshold (such as 5%), the system will be alarmed and manual inspection is required to check whether the JSON format or JSONPaths mapping is incorrect. S44: Dynamic Partitioning and Expired Cleaning: Since the dynamic partitioning function is enabled, Doris will automatically judge the current date every day and create 5 partitions from "3 days before the current date" to "1 day after the current date". For example, if today is June 6, 2025, then partitions p20250603, p20250604, p20250605, p20250606, and p20250607 will be created. When a partition has been created for a certain number of days (such as 30 days) (for example, after p20250506 expires), the system will automatically delete the partition without manual intervention, thus saving storage space. If the production environment needs to retain historical data for a longer time, just change the "partition expiration duration" parameter from 30 days to 60 days, 90 days or longer.
[0013] As a further technical solution of the present invention, the S5 specifically includes: S51: End-to-End Monitoring Metrics: Including collection-end monitoring, Kafka monitoring, and Doris monitoring; S511: Collection-End Monitoring: Whether the collection frequency is stable: Regularly count the number of samples per second or per minute. If the number of samples suddenly decreases or interrupts, it is necessary to check the PLC connectivity or gateway status; Network packet loss rate and retry times: When the gateway delivers data to the backend, if the network condition is poor, retries or packet losses will occur. Monitor the ratio of the number of retries to the total number of transmissions. When it exceeds a certain threshold (such as 5%), an alarm is issued; S512: Kafka Monitoring: Consumer Lag: Measures the gap between the consumption speed of downstream consumers (Routine Load) and the sending speed of producers (Producer). Specifically, it checks the difference between the "current write offset" and the "committed offset" for each partition. When the lag value increases and continuously exceeds a certain threshold (such as 10,000 messages), it indicates that the downstream consumption cannot keep up, and it is necessary to expand the capacity or adjust the concurrency count; Throughput: Counts the rate of messages (messages / second) and byte rate at which the Producer writes messages to Kafka to confirm whether it reaches the maximum bearing capacity of the cluster; Partition distribution: Checks whether the message distribution of each partition is balanced. If the load of a certain partition is much higher than that of other partitions, consider increasing the number of partitions or using a custom partitioning strategy; S513: Doris Monitoring: Routine Load Latency: Measures the time difference between the latest data written in Kafka and the time it is written to Doris. You can periodically query the maximum event time (event_time) of the inserted data in the table and compare it with the current system time. If the latency continuously exceeds 5 minutes, there may be a performance bottleneck in downstream writes, and it is necessary to troubleshoot the problem; Write Throughput (Write QPS / TPS): Counts the number of queries per second or the number of transactions per second when each BE node is writing. When a certain node is continuously in a high-load state (such as CPU usage exceeding 80%), it is necessary to consider expanding the capacity or optimizing the parameters; Partition Size: Monitors the data volume of each partition. When the data volume of a certain partition far exceeds the expectation (such as exceeding 100GB), consider adjusting the number of buckets or further refining the historical partitions; S52: Performance Optimization Strategies: Include Kafka optimization and Doris optimization.
[0014] As a further technical solution of the present invention, the S52 specifically includes: S521: Kafka Optimization: Batch Sending and Compression: Set an appropriate batch waiting time (such as dozens of milliseconds) at the Producer side, aggregate multiple messages that arrive in a short time and send them together, and enable a compression algorithm, such as Snappy or LZ4, which can not only reduce network bandwidth occupancy but also improve the Broker write speed; Reasonably Configure the Number of Concurrent Consumers: If the concurrency count of Routine Load is too small, the consumption speed will not be able to keep up with the production speed; while if the concurrency count is too large, the write tasks of Doris will compete for resources on the backend nodes; it is necessary to find the optimal balance of the concurrency count based on the CPU, memory, and disk I / O metrics of the BE nodes; S522: Doris Optimization: Columnar Storage Encoding and Compression: For frequently used fields in the table (such as numerical columns like temperature and pressure), columnar compression encoding methods (such as RLE, dictionary compression) are adopted to significantly reduce storage occupancy and improve scanning efficiency; Small File Merging: If frequent writes result in a large number of small files, which will affect query speed, the partition file merging operation can be executed regularly to automatically merge small files into large files; Statistical Information Maintenance: Analyze the table regularly to update the histogram statistical information of columns, enabling the query optimizer to select a better execution plan.
[0015] As a further technical solution of the present invention, the S6 specifically includes: S61: Data Warehouse Docking: In Doris, wide tables are constructed into fact tables, and then dimension tables are designed according to business needs. If more complex online analytical processing (OLAP) is required, data can be synchronized to downstream data warehouses (such as Hive / Hadoop, ClickHouse, Presto, etc.) through ETL or scheduled tasks for offline reporting and multi-dimensional analysis; in the same Doris cluster, different granularity data can also be provided to BI tools in a timely manner by combining views or hierarchical table structures; S62: Real-time Computing and Alarming: Between Kafka and Doris, or in parallel outside Doris, use streaming processing engines such as Apache Flink and Spark Streaming to directly consume the preprocessed data for complex event processing (such as anomaly detection, trend judgment). When it is found that indicators such as temperature and pressure exceed the preset thresholds, SMS or email alarms are immediately triggered; based on the CEP (Complex Event Processing) mode, multiple consecutive states or pattern recognitions can be set, such as "temperature exceeds the warning value three times in a row" to trigger an alarm, avoiding false alarms caused by short-term fluctuations; S63: Machine Learning and Deep Learning Applications: For historical time-series data stored in Doris, perform feature engineering according to time windows (such as daily, weekly, monthly), aggregate data such as temperature, pressure, and vibration into statistical features (maximum value, minimum value, average value, variance, etc.), and construct a feature table required for model training; use Python frameworks (such as TensorFlow, PyTorch) to train time-series prediction models (such as LSTM or Transformer-based models) to predict the device health status or failure probability in the short term. When training, the data of the past N days and N moments can be used as input, and the predicted values for a period of time in the future can be output; after the model training is completed and the effect is verified, it is deployed as an online inference service, regularly obtain the latest time-series data from Doris, generate prediction results and compare them with the current thresholds to give early warnings of potential failures.
[0016] The beneficial effects of the present invention are as follows: 1. Significantly reduce data processing latency and improve real-time performance: By uploading data to Kafka in real-time through the gateway and then consuming it in parallel by Doris, the end-to-end processing latency is reduced from the minute level to within 30 seconds and does not exceed 1 minute during peak hours; in high-concurrency scenarios, the data delivery success rate can reach 99.9%, effectively avoiding the "data empty window period" and achieving near-real-time monitoring and early warning; it solves the problem that traditional industrial systems rely mostly on batch imports, resulting in a 5- to 10-minute delay from on-site collection to queryability and prone to data packet loss or backlog when there is no buffer.
[0017] 2. Improve data quality and consistency, and reduce missing and abnormal data: Duplicate removal, interpolation, and filtering are completed before uploading. The duplicate upload rate is reduced to less than 0.1%, the success rate of missing point interpolation exceeds 95%, and only null values are filled for loss periods exceeding 5 seconds; in the strict mode of the backend, the dirty data write rate is lower than 0.5%, and the overall data integrity and consistency are significantly better than traditional solutions; it solves the problem in the prior art that the direct write-to-database mode lacks duplicate removal, interpolation, and exception interception mechanisms, with a common duplicate record rate as high as 2%, serious loss of data and noise interference, and subsequent need for manual cleaning and backfilling.
[0018] 3. Automated partitioning and storage cleaning reduce operation and maintenance costs and improve scalability: Utilizing the dynamic partitioning function of Doris, automatically maintain partitions from 3 days before to 1 day after the current time and set to be automatically deleted after 30 days without manual intervention; in actual operation, the data volume of a single partition is stably controlled at 50-80 GB, and the query and write performance remains stable; when the number of devices increases, only expand Kafka Broker and Doris BE nodes, and the operation and maintenance partition management workload is reduced by approximately 90%; it solves the problem that traditional database partitioning requires manual addition, archiving, and deletion, with a large workload as data accumulates. If partition management is omitted, the size of a single partition can exceed 100 GB, resulting in performance bottlenecks. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of a Doris method for storing device data based on industrial Internet proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] To make the technical means, creative features, achieved purposes, and functions of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0021] Please refer to the attached Figure 1 , a Doris method for storing device data based on industrial Internet, includes the following specific steps: S1: Obtain industrial equipment node data: Collect data from on-site PLC or gateway devices in real time to prepare for subsequent cleaning and storage. S11: Selection of collection method: Direct connection collection by PLC or collection through gateway forwarding. S111: Direct connection collection by PLC: The on-site PLC establishes communication protocols such as Modbus (TCP / RTU), PROFINET, EtherNet / IP, etc. with the host computer through Ethernet or serial port. The system sends a read request to the PLC at a pre-set sampling frequency (e.g., once per second or once every ten seconds) to obtain the real-time values within the specified register or data block. It is necessary to confirm the register range (starting address and ending address) read each time, which is configured by the equipment manufacturer or project implementers according to actual requirements. S112: Collection through gateway forwarding: If some on-site devices do not support traditional PLC protocols, it is necessary to collect and parse the data of the device's own protocol through an industrial edge gateway (such as OPC UA gateway, Modbus gateway, Siemens S7 gateway, etc.), and then convert the parsed original information (usually key-value pairs or custom binary format) into a unified data format (such as JSON) and forward it to the backend system. S12: Transmission protocol and uploading method: Upload through MQTT or HTTP interface. S121: Upload through MQTT: Use the lightweight MQTT protocol. The gateway or collection end publishes the processed data in the form of messages to a specified topic. The topic can be hierarchically organized according to "enterprise number / factory number / workshop number / equipment number", for example: / CompanyA / Factory1 / Line2 / PLC_001; The QoS (Quality of Service) usually has two optional levels: "at least once" or "exactly once". "At least once" ensures that the message will not be lost, but may be repeated; "exactly once" consumes more resources; The retain message flag is generally set to "not retain" to prevent existing subscribers from receiving outdated historical messages; The information carried during reporting includes equipment number, timestamp, and sensor readings, etc.; If the connection is disconnected or the server returns an error, the number of retry attempts and retry intervals can be set, and an alarm or local backup can be written after multiple attempts fail to prevent data loss. S122: Upload through HTTP interface: Send the collected data to the backend service through an HTTP POST request; Example of interface address: http: / / <backend domain name>:<port> / api / v1 / device / Data; Specify Content-Type: application / in the HTTP request header json, the request body is in the standard JSON format; if the request fails due to network problems or backend exceptions, the client will retry according to the pre-set retry policy (for example, retry up to three times, with a half-second interval between each retry); after all failures, the error information will be recorded in the local log and an alarm will be triggered to notify the operation and maintenance personnel; S13: Timestamp and unique identifier processing: In the acquisition end and the gateway, a consistent time reference should be used. Generally, the millisecond-level Unix timestamp is adopted, that is, the number of milliseconds from January 1, 1970 to the current moment. This can ensure that the data uploaded by different devices has a unified standard in time sorting; for the unique identifier of the device, if the device itself does not have a globally unique ID, it can be composed of a fixed prefix (such as the gateway number) and the local device number on the gateway. For example, "GW01_PLC05" can be used as the globally unique number of the device. This can ensure that there will be no device confusion in subsequent processing and storage; S2: Preprocess the acquired data: Dedup, fill in the blanks, denoise, and standardize the original acquired data to ensure the data quality and consistency of the subsequent data storage; S21: Filter duplicate data and discard invalid data; S211: Filter duplicate data: If the same device uploads the same timestamp and the same data value to the backend multiple times within a short period, it is regarded as duplicate data. After the system receives the data, it first checks whether the device number and timestamp are exactly the same as the most recent one. If they are the same, the subsequent data can be directly discarded to avoid redundant storage; S212: Discard invalid data: The message uploaded by the device often contains a "status code" field, which is used to indicate whether the sampling status is normal. If the status code is not "normal", it means the data is invalid and should be directly discarded without entering the subsequent processing process; in addition, for key physical quantities such as temperature and pressure, if the value significantly exceeds the reasonable range defined in the product manual or business (for example, the temperature is greater than its rated upper limit or lower than the rated lower limit), it is also regarded as abnormal data, discarded and recorded in the log; S22: Data Missing Completion and Interpolation: To ensure the continuity of time-series data, it is sometimes necessary to fill in the missing time points during the acquisition process. A common approach is to check the time span between two adjacent valid samplings of the same device. If the time interval does not exceed a pre-set maximum interval (e.g., within 5 seconds), the missing points are estimated using the linear relationship between the two valid values before and after, and filled in; if the interval is too large, interpolation is no longer performed, but instead the previous valid value or a null value is directly filled to avoid error accumulation; for simple cases, "nearest value filling" is adopted, directly using the most recent valid sampling value closest to the missing time point to replace it; if more accuracy is required, advanced interpolation algorithms (such as polynomial interpolation, spline interpolation, etc.) can also be used at the application layer, but the implementation cost and computational overhead are relatively higher and are less used in general industrial IoT scenarios; S23: Data Cleaning and Standardization: This includes unit unification, noise filtering, and field renaming and supplementation; S231: Unit Unification: All physical quantities need to be converted to standard units; for example, if the temperature unit uploaded by on-site equipment is Fahrenheit, the system should convert it to Celsius before proceeding to the next step to ensure consistency in subsequent analysis and storage; the same concept also applies to various physical quantities such as pressure, humidity, current, voltage, etc.; S232: Noise Filtering: Due to noise interference in on-site acquisition, it is sometimes necessary to smooth the data. Common methods include moving average or median filtering; for example, taking the average of the most recent N samplings to eliminate short-term pulse interference; or taking the median of N samplings to further remove the influence of extreme values; S233: Field Renaming and Supplementation: If the downstream system (such as a Kafka producer) requires fixed field names and formats, the original field names must be uniformly renamed; for example, the "temp" field is changed to "temperature" after cleaning; "press" is changed to "pressure"; if the message lacks some predefined fields, default values (such as setting numerical values to 0 or strings to "unknown") can be supplemented, or null values can also be set, depending on business requirements; S24: Encapsulating Data into JSON: After the above cleaning is completed, the data is organized into a unified JSON structure, which usually includes the following parts: device_id: the unique identifier of the device; timestamp: the millisecond-level Unix timestamp when the server receives it; data: the key-value pairs of the core sensor readings (such as temperature, pressure, humidity, etc.); status: the device status code used to indicate whether the sampling is normal; metadata: optional information, such as gateway number, signal strength, acquisition-end IP, and other auxiliary fields; Example JSON (for illustration only): { "device_id": "GW01_PLC05", "timestamp": 171XXXXXXX000, "data": { "temperature": 42.8, "pressure": 1.05, "humidity": 58.2}, "status": 0, "metadata": { "gateway_id": "GW01", "signal_strength": -65}}; After adopting this fixed format, when sending data to Kafka or calling the HTTP interface later, the data in JSON structure can be directly used as the message body, keeping the meanings of all fields consistent; S3: Upload the preprocessed data to Kafka: Implement data buffering and decoupling through a distributed message queue, enabling downstream parallel consumption and good scalability; S31: Kafka cluster and Topic planning: Set up at least 3 Kafka Brokers at the backend to ensure high availability of the system; Design Topic names according to business logic, such as "companyA_ factory1_line2_equipment_data", aggregating data from the same production line or the same type of equipment under the same Topic for centralized processing; Configure an appropriate number of partitions for this Topic based on the peak message volume and system concurrency. The more partitions, the stronger the downstream parallel consumption ability, but it will also increase the management overhead of the Broker. Usually, first evaluate the system peak message rate (such as 100,000 messages per second), and then combine the maximum throughput of each partition to finally determine an appropriate number of partitions such as 8, 16, 32, etc.; Set the replication factor to 3, that is, each message will have a copy on 3 Brokers to ensure that data can still be read normally even if any two nodes fail; S32: Kafka Producer configuration; S321: Broker list: Write the addresses of all Brokers into the Producer configuration, such as broker1:9092,broker2:9092,broker3:9092; S322: Message confirmation mechanism (acks): It is recommended to set it to "all", indicating that the message is considered successfully sent only after it is written to all In-Sync replicas, ensuring data is not lost; S323: Number of retries: Automatically retry when sending fails, for example, set it to 3 times; If it still fails after multiple retries, an alarm will be triggered or written to the local log to prevent silent data loss; S324: Serializer: When sending the device number as the Key, use string serialization, and when using the JSON string as the Value, also use UTF-8 encoding for serialization; S325: Partition Allocation Strategy: Adopt the default hash partition strategy, that is, use the device number for hashing and then take the modulus of the total number of partitions. This can ensure that all data of the same device falls into the same partition, guaranteeing the timeliness and sequential consumption of data of the same device; S33: Message Sending and Exception Handling: The Producer calls the asynchronous sending interface and listens for the sending result in the callback function: If the sending is successful, record the success log or ignore it; If the sending fails, record the reason for the failure (such as network timeout, Broker unavailable, etc.) in the log and retry according to the retry policy; If it still fails after retrying multiple times, put the message into the local temporary buffer queue and try to deliver it again after the background resumes or manual intervention; Throughput Optimization: Appropriately set the batch sending parameters. For example, when accumulating a certain number of messages (such as 500 messages) or reaching a certain number of bytes (such as 32KB) each time, send them all at once to avoid the network overhead caused by frequent small - batch sending; At the same time, "linger.ms" can be set to dozens of milliseconds to let the Producer pack the messages that arrive within a short period of time and then send them; S4: Doris Consumes Kafka Data and Stores It: Let Doris pull JSON data from Kafka in real - time and write it into a relational table, and at the same time, achieve automatic partitioning and expiration cleaning of data with the help of dynamic partitioning; S41: Doris Table Structure and Dynamic Partition Design: Create a wide table in Doris, with fields including: device_id, event_time, temperature, pressure, humidity, status, gateway_id, etc.; This table uses the method of partitioning by date range (Range partitioning) and enables the dynamic partitioning function; The system will automatically create partitions from "3 days before the current date" to "1 day after the current date", a total of 5 days. The partition naming method is pYYYYMMDD. For example, p20250606 represents the partition for June 6, 2025, which can ensure that there are corresponding partitions for the data of the current day and the recent historical days, and the query is more efficient; Set the partition retention period, for example, 30 days, that is, when a certain partition is created for a certain number of days (such as 30 days), the system will automatically delete the partition to achieve automatic cleaning of expired data, making the management of storage space more flexible; At the same time, perform a bucketing (Hash bucketing) design on the table. For example, hash - bucket according to the device_id and set 16 buckets. In this way, when querying and writing, data of different devices or different partitions can be scattered to multiple nodes in parallel to improve performance; S42: Create a Routine Load task: Execute a dedicated SQL statement on Doris to create a Routine Load task, specifying from which Kafka Topic the task pulls data, what concurrency to use, how many rows or bytes to pull per batch at most, and how to parse JSON; the mapping from JSON fields to table fields is provided by a JSONPaths file, and the corresponding relationships are roughly as follows: $.device_id corresponds to device_id in the table; $.timestamp corresponds to event_time, and the timestamp needs to be converted to the DATETIME type of Doris when writing; $.data.temperature, $.data.pressure, and $.data.humidity correspond to the temperature, pressure, and humidity fields in the table respectively; $.status corresponds to the status in the table; $.metadata.gateway_id corresponds to gateway_id in the table; in the Routine Load task configuration, the number of concurrent consumers (such as 4 concurrent threads), the maximum number of rows (such as 5000 rows) or the maximum number of bytes (such as 10MB) per batch, and the strict mode switch can be specified. After enabling the strict mode, if a field in a certain JSON is missing or the type does not match, it will be considered "dirty data" and discarded, and a log will be recorded to ensure data consistency; S43: Data loading process and dirty data handling: Doris Routine Load pulls messages from Kafka partitions in parallel according to the specified concurrency. Whenever a message that meets the threshold (number of rows or bytes) is pulled, it will be stored in a temporary file; the system decompresses and parses the temporary file, extracts fields according to the JSONPaths mapping from the JSON and writes them into the in-memory table; the data in the in-memory table is automatically allocated to the corresponding partitions according to the partitioning strategy, and then the backend nodes batch-persist the data to the column store; for records with parsing failures, field type mismatches, or missing required fields, Routine Load will put the record into the "dirty data" queue and decide whether to ignore or alarm according to the "maximum dirty data ratio" threshold. If the dirty data ratio in a certain batch exceeds the threshold (such as 5%), the system will be alarmed, and manual inspection is required to check whether the JSON format or the JSONPaths mapping is incorrect; S44: Dynamic Partitioning and Expired Cleaning: Since the dynamic partitioning function is enabled, Doris will automatically determine the current date every day and create 5 partitions from "3 days before the current date" to "1 day after the current date". For example, if today is June 6, 2025, it will create p20250603, p20250604, p20250605, p20250606, and p20250607; when a partition has been created for a certain number of days (such as 30 days) (for example, after p20250506 expires), the system will automatically delete the partition without manual intervention, thus saving storage space; if the production environment needs to retain historical data for a longer time, simply change the "partition expiration duration" parameter from 30 days to 60 days, 90 days, or longer; S5: Monitoring and Performance Optimization: Monitor the end-to-end link from data collection to storage, promptly detect bottlenecks and anomalies, and perform targeted optimizations; S51: End-to-End Monitoring Metrics: Include collection-end monitoring, Kafka monitoring, and Doris monitoring; S511: Collection-End Monitoring: Whether the collection frequency is stable: Regularly count the number of samples per second or per minute. If the number of samples suddenly decreases or interrupts, check the PLC connectivity or gateway status; Network packet loss rate and retry count: When the gateway delivers data to the backend, if the network condition is poor, retries or packet loss may occur. Monitor the proportion of the retry count to the total number of sent packets. When it exceeds a certain threshold (such as 5%), an alarm is issued; S512: Kafka Monitoring: Consumer Lag: Measures the gap between the consumption speed of downstream consumers (Routine Load) and the sending speed of producers (Producer). The specific method is to check the difference between the "current write offset" and the "committed offset" of each partition. When the lag value increases and continuously exceeds a certain threshold (such as 10,000 records), it means that the downstream consumption cannot keep up, and it is necessary to expand the capacity or adjust the concurrency; Throughput: Statistically analyze the rate of messages (messages / second) and byte rate written by the Producer to Kafka, and confirm whether it reaches the maximum capacity of the cluster; Partition distribution: Check whether the message distribution of each partition is balanced. If the load of a certain partition is much higher than that of other partitions, consider increasing the number of partitions or changing to a custom partition strategy; S513: Doris Monitoring: Routine Load Latency: Measures the time difference between the latest data written to Kafka and the time it reaches Doris. You can periodically query the maximum event_time of the inserted data in the table and compare it with the current system time. If the latency persists for more than 5 minutes, there may be a downstream write performance bottleneck, and the problem needs to be investigated; Write Throughput (Write QPS / TPS): Counts the number of queries per second or transactions per second when each BE node is writing. When a node is continuously in a high-load state (e.g., CPU usage exceeds 80%), capacity expansion or parameter optimization needs to be considered; Partition Size: Monitors the data volume of each partition. When the data volume of a partition far exceeds the expectation (e.g., exceeds 100GB), the number of buckets should be adjusted or the historical partitions should be further refined; S52: Performance Optimization Strategies: Include Kafka optimization and Doris optimization; S521: Kafka Optimization: Batch Sending and Compression: Set an appropriate batch waiting time (e.g., dozens of milliseconds) at the Producer side, aggregate multiple messages arriving within a short time and send them together, and enable a compression algorithm such as Snappy or LZ4, which can not only reduce network bandwidth usage but also improve the Broker write speed; Reasonably Configure the Number of Concurrent Consumers: If the number of concurrent Routine Load is too small, the consumption speed will not keep up with the production speed; if the number of concurrent consumers is too large, the write tasks in Doris will compete for resources on the backend nodes; it is necessary to find the best balance of the number of concurrent consumers based on the CPU, memory, and disk I / O metrics of the BE nodes; S522: Doris Optimization: Columnar Storage Encoding and Compression: For frequently used fields in the table (e.g., numeric columns such as temperature and pressure), adopt columnar compression encoding methods (such as RLE, dictionary compression), which can significantly reduce storage occupancy and improve scanning efficiency; Small File Merging: If frequent writes result in a large number of small files, it will affect the query speed. You can periodically perform partition file merging operations to automatically merge small files into large files; Statistical Information Maintenance: Periodically perform analysis operations on the table to update the histogram statistical information of columns, enabling the query optimizer to select a better execution plan; S6: Expansion: Secondary Utilization and Downstream Analysis. On the premise of ensuring data quality and real-time performance, support subsequent data warehouse docking, real-time computing, and machine learning applications; S61: Data Warehouse Docking: In Doris, construct a wide table into a fact table, and then design dimension tables according to business needs. If more complex online analytical processing (OLAP) is required, data can be synchronized to downstream data warehouses (such as Hive / Hadoop, ClickHouse, Presto, etc.) through ETL or scheduled tasks for offline reporting and multi-dimensional analysis; in the same Doris cluster, data at different granularities can also be provided to BI tools in a timely manner by combining views or hierarchical table structures; S62: Real-time Computing and Alarming: Between Kafka and Doris, or in parallel outside Doris, use streaming processing engines such as Apache Flink and Spark Streaming to directly consume the preprocessed data for complex event processing (such as anomaly detection and trend judgment). When it is found that indicators such as temperature and pressure exceed the preset thresholds, SMS or email alarms are immediately triggered; based on the CEP (Complex Event Processing) mode, multiple consecutive states or pattern recognitions can be set, such as "temperature exceeding the warning value three times in a row" to trigger an alarm, avoiding false alarms caused by short-term fluctuations; S63: Machine Learning and Deep Learning Applications: For historical time-series data stored in Doris, perform feature engineering according to time windows (such as daily, weekly, monthly), aggregate data such as temperature, pressure, and vibration into statistical features (maximum value, minimum value, average value, variance, etc.), and construct a feature table required for model training; use Python frameworks (such as TensorFlow, PyTorch) to train time-series prediction models (such as LSTM or Transformer-based models) to predict the device health status or failure probability in the short term. When training, the data of the past N days and N moments can be used as input, and the predicted values for a period of time in the future can be output; after the model training is completed and the effect is verified, it is deployed as an online inference service, regularly obtain the latest time-series data from Doris, generate prediction results and compare them with the current thresholds to give early warnings of potential failures.
[0022] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: significantly reducing data processing latency and improving real-time performance: By uploading data in real time to Kafka through the gateway and then consuming it in parallel by Doris, the end-to-end processing latency is reduced from the minute level to within 30 seconds and does not exceed 1 minute during peak hours; in high-concurrency scenarios, the data delivery success rate can reach 99.9%, effectively avoiding the "data blank period" and achieving near-real-time monitoring and warning; solving the problem that traditional industrial systems mostly rely on batch imports, resulting in a 5-10 minute delay from on-site collection to queryability and prone to data packet loss or backlog when there is no buffer.
[0023] Improve data quality and consistency, and reduce missing values and anomalies: Deduplication, interpolation, and filtering are completed before uploading. The duplicate upload rate is reduced to less than 0.1%, the success rate of interpolating missing points exceeds 95%, and only null placeholders are made for loss periods exceeding 5 seconds. In the strict mode of the backend, the dirty data write rate is lower than 0.5%, and the overall data integrity and consistency are significantly better than traditional solutions. It solves the problems in the prior art that the direct write library mode lacks deduplication, interpolation, and anomaly interception mechanisms, the common duplicate record rate is as high as 2%, the missing data and noise interference are serious, and subsequent manual cleaning and backfilling are required.
[0024] Automated partitioning and storage cleaning reduce operation and maintenance costs and improve scalability: Using the dynamic partitioning function of Doris, automatically maintain partitions from 3 days before to 1 day after the current time, and set to be automatically deleted after 30 days without manual intervention. In actual operation, the data volume of a single partition is stably controlled within 50 - 80 GB, and the query and write performance remain stable. When the number of devices increases, only Kafka Broker and Doris BE nodes need to be expanded, and the operation and maintenance workload of partition management is reduced by about 90%. It solves the problems of traditional database partitioning that require manual addition, archiving, and deletion. As the data accumulates, the workload is large. If partition management is missed, the size of a single partition can exceed 100 GB, resulting in performance bottlenecks.
[0025] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention is limited to these examples. Under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity.
[0026] The present invention is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the specification. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A Doris method for storing device data based on the industrial Internet, characterized in that It includes the following specific steps: S1: Obtain industrial equipment node data: Real-time collect the data of on-site PLC or gateway devices to prepare for subsequent cleaning and storage; S2: Preprocess the obtained data: Dedup, fill in the blanks, denoise, and standardize the originally collected data to ensure the data quality and consistency of subsequent data storage; S3: Upload the preprocessed data to Kafka: Implement data buffering and decoupling through a distributed message queue, which allows downstream parallel consumption and has good scalability; S4: Doris consumes Kafka data and stores it: Let Doris pull JSON data from Kafka in real time and write it into a relational table, and at the same time achieve automatic partitioning and expiration cleaning of data with the help of dynamic partitioning; S5: Monitoring and performance optimization: Monitor the end-to-end link from data collection to storage, promptly discover bottlenecks and anomalies, and perform targeted optimizations; S6: Expansion: Secondary utilization and downstream analysis. On the premise of ensuring data quality and real-time performance, support subsequent data warehouse docking, real-time computing, and machine learning applications.
2. The Doris method for device data storage based on industrial Internet according to claim 1, characterized in that, The specific content of S1 includes: S11: Selection of collection methods: Direct connection collection of PLC or collection through gateway forwarding; S12: Transmission protocol and upload method: Upload via MQTT or HTTP interface; S13: Timestamp and unique identifier processing: At the collection end and the gateway, a consistent time reference should be used, and the Unix timestamp in milliseconds should be adopted; For the unique identifier of the device, if the device itself does not have a globally unique ID, it can be composed of a fixed prefix and the local device number spliced on the gateway.
3. The Doris method for device data storage based on industrial Internet according to claim 2, wherein, The specific content of S12 includes: S121: Upload via MQTT: Use the lightweight MQTT protocol. The gateway or the collection end publishes the processed data in the form of messages to the specified topic, and the topic can be hierarchically organized according to "enterprise number / factory number / workshop number / device number"; QoS can be selected from two levels: "at least once" or "exactly once"; The retain message flag is set to "not retained"; If the connection is disconnected or the server returns an error, the number of retries and the retry interval can be set, and an alarm or local backup can be written after multiple attempts fail; S122: Upload via HTTP interface: Send the collected data to the backend service through an HTTP POST request; Specify the Content-Type in the HTTP request header; If the request fails due to network problems or backend exceptions, the client will try again according to the pre-set retry policy; After all attempts fail, the error information will be recorded in the local log and an alarm will be triggered to notify the operation and maintenance personnel.
4. The Doris method for storing device data based on industrial Internet according to claim 1, characterized in that The specific content of S2 includes: S21: Filter duplicate data and discard invalid data; S22: Data Missing Filling and Interpolation: Check the time span between two adjacent valid samplings of the same device. If the time interval does not exceed the preset maximum interval, estimate using the linear relationship between the two valid values before and after to fill the missing points; if the interval is too large, stop interpolation and directly fill with the previous valid value or null value to avoid error accumulation; for simple cases, use "nearest value filling", directly using the valid sampling value closest to the missing time to replace; if more accuracy is required, advanced interpolation algorithms can also be used at the application layer. S23: Data Cleaning and Standardization: Including unit unification, noise filtering, and field renaming and supplementation. S24: Encapsulating Data into JSON: After the above cleaning is completed, organize the data into a unified JSON structure, including the following parts: device_id, Timestamp, status, metadata; after adopting this fixed format, when sending to Kafka or calling the HTTP interface later, the data in the JSON structure can be directly used as the message body to keep the meanings of each field consistent.
5. A Doris method for device data storage based on industrial Internet according to claim 1, characterized in that, The specific steps of S3 are as follows: S31: Kafka Cluster and Topic Planning: Set up at least 3 Kafka Brokers in the backend, design the Topic name according to the business logic, and gather data from the same production line or the same type of device under the same Topic; configure the appropriate number of partitions for this Topic according to the peak message volume and system concurrency; set the replication factor to 3, that is, each message will have a copy on 3 Brokers. S32: Kafka Producer Configuration S33: Message Sending and Exception Handling: The Producer calls the asynchronous sending interface and listens for the sending result in the callback function: if the sending is successful, record the success log or ignore it; if the sending fails, record the reason for the failure in the log and retry according to the retry policy; if it still fails after multiple retries, put this message into the local temporary buffer queue and try to deliver it again after the background resumes or manual intervention.
6. The Doris method for device data storage based on industrial Internet according to claim 5, characterized in that, The specific steps of S32 are as follows: S321: Broker List: Write the addresses of all Brokers into the Producer's configuration. S322: Message Confirmation Mechanism: It is recommended to set it to "all", indicating that the message is considered successfully sent only after being written to all In-Sync replicas to ensure data is not lost. S323: Number of Retries: Automatically retry when the sending fails. If it still fails after multiple retries, trigger an alarm or write to the local log. S324: Serializer: Use string serialization when sending the device number as the Key, and also use UTF-8 encoding serialization when using the JSON string as the Value. S325: Partition Allocation Strategy: Adopt the default hash partition strategy, that is, take the modulus of the total number of partitions after hashing the device number.
7. A Doris method for device data storage based on industrial Internet according to claim 1, characterized in that The specific steps of S4 are as follows: S41: Doris Table Structure and Dynamic Partition Design: Create a wide table in Doris. The table uses the method of partitioning by date range and enables the dynamic partition function. The system will automatically create partitions for a total of 5 days from "3 days before the current date" to "1 day after the current date" according to the current date. Set the partition retention period, that is, when a partition is created for a certain number of days, the system will automatically delete the partition. At the same time, perform bucket design on the table. S42: Create a Routine Load Task: Execute specific SQL statements on Doris to create a Routine Load task, specifying from which Kafka Topic the task pulls data, what concurrency number to use, how many records or how many bytes to pull at most per batch, and how to parse JSON. The mapping from JSON fields to table fields is provided by a JSONPaths file. In the RoutineLoad task configuration, the number of concurrent consumers, the maximum number of rows or maximum bytes per batch, and the strict mode switch can be specified. S43: Data Loading Process and Dirty Data Handling: Doris Routine Load pulls messages from Kafka partitions in parallel according to the specified concurrency number. Whenever messages that meet the threshold are pulled, they will be stored in a temporary file. The system decompresses and parses the temporary file, extracts fields according to the JSONPaths mapping from the JSON, and writes them into an in-memory table. The data in the in-memory table is automatically allocated to the corresponding partitions according to the partition strategy, and then the backend nodes batch-persist the data into the column store. For records with parsing failures, field type mismatches, or missing required fields, Routine Load will put the record into the "dirty data" queue and decide whether to ignore or alarm according to the "maximum dirty data ratio" threshold. If the dirty data ratio of a certain batch exceeds the threshold, the system will be triggered to alarm, and manual inspection is required to check whether the JSON format or JSONPaths mapping is incorrect. S44: Dynamic Partition and Expired Cleaning: Since the dynamic partition function is enabled, Doris will automatically judge the current date every day and create 5 partitions from "3 days before the current date" to "1 day after the current date". When a partition is created for a certain number of days, the system will automatically delete the partition.
8. The Doris method for storing device data based on industrial Internet according to claim 1, wherein The specific content of S5 includes: S51: End-to-End Monitoring Metrics: Include collection-end monitoring, Kafka monitoring, and Doris monitoring. S52: Performance Optimization Strategies: Include Kafka optimization and Doris optimization.
9. The Doris method for storing device data based on industrial Internet according to claim 8, characterized in that, The specific content of S52 includes: S521: Kafka Optimization: Batch Sending and Compression: Set an appropriate batch waiting time at the Producer end, aggregate multiple messages that arrive within a short time and send them together, and enable the compression algorithm. Reasonably Configure the Number of Concurrent Consumers: It is necessary to find the best balance point of the concurrency number according to the metrics of BE node CPU, memory, and disk I / O. S522: Doris Optimization: Column Store Encoding and Compression: For the commonly used fields in the table, adopt the column store compression encoding method, which can greatly reduce the storage occupancy and improve the scanning efficiency. Small file merging: The partition file merging operation can be executed at a scheduled time to automatically merge small files into large files; Statistical information maintenance: Regularly perform analysis operations on tables to update the histogram statistical information of columns, enabling the query optimizer to select a better execution plan.
10. The Doris method for device data storage based on industrial Internet according to claim 1, characterized in that The specific steps of S6 are as follows: S61: Data warehouse docking: In Doris, construct a wide table into a fact table, and then design dimension tables according to business needs. If more complex online analytical processing is required, data can be synchronized to the downstream data warehouse through ETL or scheduled tasks for offline reporting and multi-dimensional analysis; in the same Doris cluster, different granularity data can also be provided to BI tools in a timely manner by combining views or hierarchical table structures; S62: Real-time calculation and alerting: Between Kafka and Doris, or in parallel outside Doris, use streaming processing engines such as Apache Flink and Spark Streaming to directly consume the preprocessed data for complex event processing. When it is found that temperature and pressure indicators exceed the preset thresholds, immediately trigger SMS or email alerts; based on the CEP mode, multiple consecutive states or pattern recognitions can be set; S63: Machine learning and deep learning applications: For historical time series data stored in Doris, perform feature engineering according to time windows, aggregate the data into statistical features, and construct a feature table required for model training; use the Python framework to train a time series prediction model to predict the device health status or failure probability in the short term. When training, the data of the past N days and N moments can be used as input, and the predicted values for a period of time in the future can be output; after the model training is completed and the effect is verified, deploy it as an online inference service, regularly obtain the latest time series data from Doris, generate prediction results, and compare them with the current thresholds to early warn of potential failures.
Citation Information
Patent Citations
Internet of Things data platform based on big data and AI
CN111787066A
Real-time data bin design method, device and equipment based on Dores and storage medium
CN114942916A
ClickHouse-based fused CDN (Content Delivery Network) service data processing method, device and equipment
CN119271888A
Buried point data transmission and storage method and system based on Kafka
CN119561860A
High-reliability lossless data acquisition and processing system and method
CN119961085A
Cited By
Multi-source heterogeneous JSON data stream processing method based on Flink SQL
CN121456002A