Internet of Things data processing system and method based on data lake
Through the IoT data processing system based on the data lake, the problems of high cost of IoT data storage and insufficient sharing capabilities are solved, data sharing and integration in multiple application scenarios are realized, and the accuracy and intelligence level of data analysis are improved.
Patent Information
- Application Number
- CN202510520332.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the IoT data storage cost is high and the sharing capability is insufficient, so it is difficult for traditional databases to effectively integrate and process unstructured data, which affects the accuracy and intelligence of data analysis.
The data lake-based IoT data processing system is adopted, including collection equipment, data lake equipment and digital warehouse equipment. Through the data lake, the IoT data in multiple formats is stored, and aggregation and analysis is carried out in the digital warehouse equipment to realize data sharing and integration in multiple application scenarios.
It reduces storage costs, improves data sharing capabilities and analysis accuracy, fully taps the potential value of IoT data, and improves the intelligence level of IoT applications.
Smart Images

Figure CN120336435A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of Internet of Things technology, and particularly to an Internet of Things data processing system and method based on a data lake. Background Art
[0002] The Internet of Things has entered a new stage of development of everything being interconnected. Whether it is for consumer, enterprise or industrial scenarios, ubiquitous Internet of Things devices will generate a large amount of data. Internet of Things data is characterized by diversity, real-time nature and massive volume. For example, Internet of Things data can include various types such as sensor data, device status data, location data, image data, audio data, etc. This poses new challenges to the storage, processing and analysis of Internet of Things data.
[0003] In the prior art, a common way to process Internet of Things data is to store Internet of Things data in a traditional database and then perform analysis and processing. However, this way has problems such as high storage cost and insufficient sharing ability of Internet of Things data. Summary of the Invention
[0004] Embodiments of this specification provide an Internet of Things data processing system and method based on a data lake, which are used to reduce storage costs, and realize the cross-application of Internet of Things data in multiple application scenarios, and improve the sharing ability of Internet of Things data.
[0005] Embodiments of this specification provide an Internet of Things data processing system based on a data lake. The Internet of Things data processing system based on a data lake includes a variety of collection devices, a data lake device and a data warehouse device;
[0006] The variety of collection devices are used to collect a variety of Internet of Things data and send the Internet of Things data to the data lake device;
[0007] The data lake device is used to store the received Internet of Things data into the data lake;
[0008] The data warehouse device is used to, for each of multiple application scenarios, obtain a variety of Internet of Things data required for this application scenario from the data lake, perform aggregation analysis on the obtained variety of Internet of Things data according to the business logic of this application scenario, and store the result data after aggregation analysis into the data warehouse to provide the result data to downstream application devices.
[0009] Embodiments of this specification also provide an Internet of Things data processing method based on a data lake, including:
[0010] Using a variety of collection devices to collect a variety of Internet of Things data;
[0011] Using the data lake device to store the collected variety of Internet of Things data into the data lake;
[0012] For each of multiple application scenarios, a data warehouse device obtains various Internet of Things (IoT) data required for the application scenario from the data lake, performs aggregated analysis on the obtained various IoT data, and stores the result data after the aggregated analysis in the data warehouse to provide the result data to downstream application devices.
[0013] An embodiment of this specification also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned IoT data processing method based on the data lake is implemented.
[0014] An embodiment of this specification also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned IoT data processing method based on the data lake is implemented.
[0015] An embodiment of this specification also provides a computer program product including a computer program, and when the computer program is executed by a processor, the above-mentioned IoT data processing method based on the data lake is implemented.
[0016] In the technical solution of the embodiment of this specification, the data lake can store IoT data collected by various collection devices; the data warehouse is built based on the data lake, can perform aggregated analysis on the IoT data in the data lake according to application scenarios, and can store result data in multiple application scenarios. Thus, the separation of computing and storage is achieved, and the storage cost can be reduced. In addition, for each application scenario, the data warehouse device can perform aggregated analysis on various IoT data required for the application scenario to obtain the result data of the application scenario. Thus, the effective integration of various IoT data is achieved, which is beneficial to improving the accuracy of data analysis in the application scenario. In addition, the IoT data collected by each collection device can be applied to multiple different application scenarios, so that the IoT data collected by the collection device can be shared in multiple application scenarios. In the technical solution of the embodiment of this specification, through the inflow of IoT data into the lake and into the warehouse, the storage cost is reduced, and the co-construction and sharing of IoT data are achieved. Through the co-construction and sharing of IoT data, the consistency and integrity of data can be effectively improved, the potential value of IoT data can be fully explored, strong support can be provided for decision-making, thereby improving the intelligent level of IoT applications and effectively meeting user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. The accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic structural diagram of the Internet of Things data processing system based on a data lake in the embodiments of the present specification;
[0019] Figure 2 It is a schematic diagram of the Internet of Things data processing process based on a data lake in the embodiments of the present specification;
[0020] Figure 3 It is a schematic diagram of the Internet of Things data processing process based on a data lake in the embodiments of the present specification;
[0021] Figure 4 It is a schematic flowchart of the Internet of Things data processing method based on a data lake in the embodiments of the present specification. Detailed implementation manners
[0022] The following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. The specific embodiments described herein are only used to explain the present disclosure, rather than limiting the present disclosure. Based on the described embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present disclosure. In addition, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0023] In the above prior art, on the one hand, traditional databases require expensive hardware devices and software licenses, and usually adopt a centralized storage architecture, making it difficult to cope with the rapid growth of Internet of Things (IoT) data. As the amount of data increases, the storage and processing costs of traditional databases will continue to rise. On the other hand, traditional databases are mainly optimized for structured data and have weak processing capabilities for unstructured data. IoT data contains a large amount of unstructured data (such as images, audio, etc.). It is difficult for traditional databases to efficiently store this data, resulting in a further increase in storage costs. On the other hand, IoT data comes from a wide range of sources and has diverse formats. It is difficult for traditional databases to effectively integrate IoT data from different sources and formats. This will lead to difficulties in ensuring data consistency and integrity, affecting the accuracy of data analysis. In addition, it is also difficult to fully explore the potential value of IoT data and provide strong support for decision-making. This will result in a low level of intelligence in IoT applications and fail to meet user needs.
[0024] Please refer to Figure 1 、 Figure 2 and Figure 3 。This embodiment of the specification provides an IoT data processing system based on a data lake. The IoT data processing system may include a variety of collection devices, data lake devices, and data warehouse devices.
[0025] In some embodiments, each collection device can collect IoT data and send the collected IoT data to the data lake device. The collection device may include information sensors, radio frequency identification devices, global positioning devices, infrared sensors, laser scanners, smart meters, video collection devices, audio collection devices, etc. The IoT data may be data obtained by the collection device, such as sound, light, heat, electricity, mechanics, chemistry, biology, location, image, audio, and other data.
[0026] The multiple collection devices can belong to multiple Internet of Things (IoT) scenarios. The IoT scenarios can be scenarios for providing IoT data. The IoT scenarios can include office premises, homes, factories, and farms. Through the multiple collection devices, it is possible to monitor multiple IoT scenarios. Thus, it is possible to analyze and mine IoT data across regions, time, and devices, providing capacity support for intelligent decision-making and control. Among them, each IoT scenario can have one or more collection devices. Each collection device can collect data in a streaming manner. Different collection devices can collect the same or different types of IoT data. In some scenario examples, IoT data can include employee perception data, device sensing data, customer perception data, application behavior perception data, internal environment perception data, external data, etc. Employee perception data includes employees' work data, behavior data, service data, etc. Device sensing data includes the basic information of devices, utilization information, fault information, alarm information, etc. Customer perception data includes the trajectory data of customers entering the venue, device operation behavior data, stay duration, location data, etc. Application behavior perception data includes operation data, abnormal data, running data, resource data, etc. Internal environment perception data includes physical environment information (such as temperature data, humidity data, pressure data, light data, etc.). By collecting the physical environment information in the venue, a series of intelligent services such as automatically turning on and off lights and automatically adjusting the air conditioner temperature can be realized. External data includes data received through external data components (such as remote sensing satellite data).
[0027] IoT data has characteristics such as huge data volume, temporality, and real-time nature. For example, IoT data has strong temporality, a strong trend line, low value of individual data points (allowing loss), and high requirements for real-time calculation. For another example, IoT data is heterogeneous data, which can include structured data such as numerical temperature and humidity, and can also include unstructured data such as video and audio, and can also include semi-structured data. For another example, the data source of IoT data is unique. The IoT data collected by different collection devices is independent. For another example, IoT data can be without operations such as addition, deletion, modification, and query.
[0028] Optionally, multiple collection devices can be classified according to the type of IoT data collected by the collection devices, to obtain one or more collection device sets. Each collection device set contains one or more collection devices. Each collection device set corresponds to a type of IoT data, and the collection devices in the collection device set can collect IoT data of this data type. For example, the collection devices in one collection device set can collect temperature data, the collection devices in another collection device set can collect humidity data, and the collection devices in another collection device set can collect video data. Since the collection devices in one collection device set can collect the same type of IoT data, the IoT device management platform can configure the same collection model for the collection devices in the same collection device set; different collection models can be configured for the collection devices in different collection device sets. So that the collection devices in one collection device set can uniformly collect IoT data according to one collection model. The collection model is used to constrain the collection frequency, the format of IoT data, etc. when the collection device collects IoT data.
[0029] In some embodiments, the IoT data processing system based on the data lake may further include an IoT device management platform. The IoT device management platform can be implemented in the form of hardware, software, or a combination of hardware and software. The multiple collection devices can be connected to the IoT device management platform, and each type of collection device sends the collected IoT data to the IoT device management platform. The IoT device management platform pushes the received IoT data to the data lake device. For example, the collection device can be integrated with an SDK (Software Development Kit). The collection device accesses the IoT device management platform through the SDK. The forms of the collection devices are diverse. The IoT device management platform realizes the access of the collection devices by providing a device management platform SDK that adapts to multiple operating system platforms and supports multiple development languages. The IoT device management platform implements an adapter for each supported protocol. The adapter inherits from a base class, and this base class defines a unified data reception and processing interface. The adapter is dynamically loaded into the system as a plug-in. After receiving the data, the adapter performs syntax parsing, and then converts the data into the internal standard format of the IoT device management platform according to the predefined conversion rules.
[0030] In some embodiments, a data lake is deployed on the data lake device. The data lake is a large-scale raw data storage and processing architecture that can store data of various types (such as structured, semi-structured, unstructured) and formats, has strong scalability and flexibility, and supports fast data storage and query. The data lake device adopts distributed storage technology to store the data in the data lake on multiple nodes to improve storage capacity and reliability. The data lake device can receive Internet of Things data collected by various collection devices; it can store the received Internet of Things data into the data lake.
[0031] The data lake can store data in multiple formats without complex structuring processes, reducing storage costs. For example, traditional databases require structured storage of data, which means that for unstructured data, additional processing and conversion are needed, increasing storage costs. In contrast, the data lake can directly store not only structured data but also unstructured data without additional processing of unstructured data, thus reducing storage costs. In addition, the distributed storage technology can flexibly expand the storage capacity according to the growth of data volume, avoiding the expansion costs of traditional databases.
[0032] In some embodiments, the data lake can be implemented using a message queue (such as Kafka). The message queue can store data in multiple formats without complex structuring processes, reducing storage costs. After receiving Internet of Things data, the data lake device can store the received Internet of Things data into the message queue.
[0033] In some other embodiments, the data lake can also be implemented through a combination of multiple storage media. After receiving Internet of Things (IoT) data, the data lake device can identify the data type of the IoT data; it can determine the corresponding storage policy according to the data type of the IoT data; and it can store the IoT data in the corresponding storage medium according to the storage policy. Among them, the data types of IoT data can include structured data (such as temperature data, humidity data), semi-structured data (such as device status data in JSON or XML format), unstructured data (such as video data, audio data), etc. The data lake device can be pre-configured with the corresponding relationship between data types and storage policies. Each data type corresponds to a storage policy. The storage policies can include storing in a relational database (such as MySQL, PostgreSQL) or a distributed database (such as MPP), storing in a NoSQL database, storing in a distributed file system (such as HDFS) or an object storage system (such as COS). For example, structured data corresponds to a relational database or a distributed database. The data lake device can store structured data in a relational database or a distributed database. Another example is that semi-structured data corresponds to a NoSQL database. The data lake device can store semi-structured data in a NoSQL database. Another example is that unstructured data corresponds to a distributed file system or an object storage system. The data lake device can store unstructured data in a distributed file system or an object storage system. The storage medium can include a database. Thus, the data lake device can achieve hierarchical data storage and reduce storage costs.
[0034] In some embodiments, the data lake device can also identify the popularity of IoT data in the data lake and the corresponding storage medium; and it can select IoT data for archived storage according to the popularity of the IoT data and the corresponding storage medium.
[0035] Heat refers to the access frequency or usage frequency of data within a certain time period (such as one month, half a year, etc.). The data lake device can count the heat of each IoT data in the data lake at preset time intervals; it can judge whether the IoT data meets the archiving storage conditions according to the heat of each IoT data and the storage medium where the IoT data is located; if so, the IoT data can be archived and stored. The archiving storage conditions can include: the IoT data is located in a relational database and the heat is greater than or equal to a first value; the IoT data is located in a NoSQL database and the heat is greater than or equal to a second value; the IoT data is located in a distributed file system or an object storage system and the heat is greater than or equal to a third value. Among them, the first value is less than or equal to the second value, and the second value is less than or equal to the third value. The magnitude of the heat is positively correlated with the access frequency or usage frequency of the IoT data. Through archiving storage, the IoT data can be stored in a low-cost storage medium (such as cold storage), which can reduce the storage cost. In addition, in the archiving storage conditions, different databases are configured with different thresholds. Databases with relatively high costs (such as relational databases) are configured with smaller thresholds; databases with relatively low costs (such as distributed file systems or object storage systems) are configured with larger thresholds. This can further reduce the storage cost.
[0036] In some embodiments, a data warehouse is deployed on the data warehouse device. The data warehouse can be a cloud-based data warehouse. The data warehouse device can include an electronic device with computing and network interaction functions, or can also include software running in the electronic device to provide support for data processing and network interaction. The data warehouse device can include an independent device. Or, the data warehouse device can also include multiple physical subsystems. Each physical subsystem can include an electronic device with computing and network interaction functions, or can also include software running in the electronic device to provide support for data processing and network interaction. The multiple physical subsystems correspond to multiple application scenarios. Each physical subsystem corresponds to one application scenario. The application scenario can refer to the usage scenario of IoT data. As an example, the application scenario includes but is not limited to: intelligent healthcare, intelligent agriculture, intelligent security, intelligent banking, non-financial asset applications, unmanned banking, and intelligent parks. The data warehouse device can implement functions such as IoT data services, physical data development, physical data statistics, physical data analysis, and physical device model management.
[0037] In some embodiments, the data warehouse device can obtain various Internet of Things (IoT) data required for each of multiple application scenarios from a data lake, aggregate and analyze the obtained multiple IoT data according to the business logic of the application scenario, and store the result data after aggregation and analysis in a data warehouse. The data lake can store various types and formats of IoT data. For each application scenario, the data warehouse device can effectively integrate the multiple IoT data stored in the data lake to achieve aggregation and analysis of the multiple IoT data in each application scenario. Thus, based on the data lake and the data warehouse, co-construction and sharing of IoT data in multiple application scenarios can be realized. For example, the IoT data collected by a certain collection device can be applied to multiple different application scenarios to generate result data in multiple different application scenarios. Thus, the potential value of IoT data can be fully explored, providing strong support for decision-making, thereby improving the intelligent level of IoT applications and effectively meeting user needs. The multiple application scenarios correspond to the multiple application devices.
[0038] The data lake device can send the IoT data in the data lake to a message queue (such as a Kafka message queue). The data warehouse device can obtain IoT application data from the message queue. For example, the data warehouse device can use a Kafka message queue as an input source and provide an API interface. The data lake device actively calls the API interface to transmit the IoT data to the Kafka message queue. Alternatively, the data warehouse device can also send a request to the data lake device at regular time intervals. The data lake device can send the IoT application data received within the set time interval to the data warehouse device according to the received request. Of course, the data warehouse device can also obtain IoT data from the data lake in other ways, and this embodiment does not make specific limitations on this.
[0039] The Internet of Things data required for each application scenario can be pre-configured. The data warehouse device can obtain various Internet of Things data required for each application scenario from the data lake according to the pre-configuration. By performing aggregation analysis on the various Internet of Things data in each application scenario, the integration of multiple Internet of Things data can be achieved, which is conducive to providing accurate, scientific, and efficient data support for the application scenario. The data warehouse device can directly perform aggregation analysis on the Internet of Things data. Or, the data warehouse device can also preprocess the Internet of Things data; it can perform aggregation analysis on the preprocessed Internet of Things data. The preprocessing can include cleaning, transformation, etc. The cleaning can be implemented based on a rule engine. For example, the data warehouse device can be pre-configured with rules. The rules are used for data range checking, format verification, etc. The data warehouse device can delete data points that do not meet the rules. Or, the cleaning can also be implemented based on machine learning. For example, the data warehouse device can use clustering analysis to identify abnormal data points, so as to delete the abnormal data points in the Internet of Things data. The data transformation can include ETL transformations such as extract, transform, and load. The data warehouse device can perform ETL transformations on the cleaned Internet of Things data.
[0040] The aggregation algorithm for each business scenario can be pre-configured according to the business logic of each application scenario. The data warehouse device can use the aggregation algorithm to perform aggregation analysis on the various Internet of Things data in each application scenario. For example, the data warehouse device can use window functions such as Spark Streaming (such as sliding window average calculation) for data aggregation. Another example is that the data warehouse device can also use MapReduce jobs in Hive for batch processing aggregation. Another example is that the data warehouse device can also input the various Internet of Things data in each application scenario into a model to obtain the result data output by the model. The model can include artificial intelligence models such as neural network models. Each of the result data can include one or more metric data.
[0041] For example, for the unmanned bank application scenario, the various Internet of Things data required for this application scenario include, but are not limited to, temperature data collected by temperature sensors, humidity data collected by humidity sensors, light data collected by light sensors, air quality data collected by air quality sensors, etc. The metric data for this application scenario can include air quality metrics. The air quality metrics can be obtained, for example, by weighted fusion of temperature data, humidity data, light data, air quality data, etc.
[0042] Optionally, the data warehouse device can also select a corresponding database according to the type of the result data, so as to store the result data into the selected database. The types of the result data can include device status data (such as temperature, humidity, light sense, air quality, etc.) and archive data (such as device basic information, device specification information, device type information, etc.). For example, when the data type of the result data is device status data, the data warehouse device can store the result data into a time series database; when the data type of the result data is archive data, the data warehouse device can store the result data into a relational database. Thus, for the result data that often changes in chronological order, it can be stored into a time series database (such as elasticsearch database). And for the result data that does not often change, it can be stored into a relational database (such as mpp database).
[0043] In some embodiments, for each of multiple application scenarios, the data warehouse device can obtain multiple pieces of Internet of Things data in the application scenario from the data lake; and can generate derivative data according to the obtained multiple pieces of Internet of Things data. The data warehouse device can generate derivative data by combining the obtained multiple pieces of Internet of Things data in a certain way. For example, some specific algorithms (such as addition, subtraction, multiplication, division, Cartesian product, one-hot encoding, etc.) can be used to generate derivative data. The result data in the application scenario can include derivative data. Of course, it can also include other data, such as metric data.
[0044] In some embodiments, the system may further include one or more downstream application devices. By accessing the data warehouse device, the one or more application devices can consume the result data. Each application device may correspond to one or more application scenarios. The data warehouse device may obtain various Internet of Things data under each application scenario from the data lake for each of the multiple application scenarios; may perform aggregation analysis on the obtained various Internet of Things data; and may store the result data after aggregation analysis in the data warehouse. In this way, each application device can obtain the corresponding result data from the data warehouse according to its own application scenario. For example, the data warehouse device may obtain various Internet of Things data under each application scenario from the data lake for each of the multiple application scenarios; may process and analyze the data through functions such as stream computing, offline computing, and data synchronization and store it in databases such as mpp, kafka, and elasticsearch. For non-financial asset applications, the data warehouse device may assemble the corresponding result data into files. The application devices corresponding to non-financial asset applications can obtain the files by subscribing. For smart campuses, the data warehouse device may store the corresponding result data in kafka. The application devices corresponding to smart campuses can obtain the data by subscribing. For intelligent vaults, the data warehouse device may encapsulate the corresponding result data into an interface for the application devices corresponding to the intelligent vaults to call. For unmanned banks, the data warehouse device may encapsulate the corresponding result data into an interface for the application devices corresponding to the unmanned banks to call.
[0045] Optionally, the data warehouse device may include multiple physical subsystems. Each physical subsystem may be configured with a data warehouse. Each physical subsystem can, according to its own application scenario, obtain various Internet of Things data under this application scenario from the data lake, perform aggregation analysis on the obtained various Internet of Things data, and store the result data after aggregation analysis in the data warehouse of this physical subsystem. Each application device may correspond to one application scenario. Each physical subsystem corresponds to an application scenario. Thus, there is a corresponding relationship between the physical subsystems and the application scenarios. Each application device can obtain the result data from the data warehouse of the corresponding physical subsystem. The data warehouse of the physical subsystem includes a time series database and a relational database.
[0046] In some embodiments, for each of multiple application scenarios, the data warehouse device can determine whether the application scenario meets the real-time condition; if the application scenario meets the real-time condition, it can actively obtain multiple pieces of Internet of Things data in the application scenario from the data lake at regular intervals; it can perform aggregation analysis on the obtained multiple pieces of Internet of Things data to obtain result data; and it can store the result data and the corresponding application scenario in the data warehouse. Thus, for application scenarios with relatively high real-time requirements, the data warehouse device can pre-generate result data. In this way, when the application device in a high-real-time application scenario needs the result data, the result data can be provided to the application device in a timely manner. The data warehouse device can query whether each application scenario is configured with a real-time flag to determine whether the application scenario meets the real-time condition. If the application scenario is configured with a real-time flag, it can be determined that the application scenario meets the real-time condition; otherwise, it can be determined that the application scenario does not meet the real-time condition.
[0047] Optionally, the data warehouse device can also send the result data obtained by aggregation analysis to the corresponding application device. The corresponding application device can be the application device corresponding to the application scenario. For example, multiple application scenarios can correspond to multiple message queues. Each application scenario can correspond to one message queue. The data warehouse device can send the result data obtained by aggregation analysis to the corresponding message queue. The application device can obtain the result data from the message queue.
[0048] In some embodiments, the data warehouse device can create an index for the result data in the data warehouse to improve the access speed of the result data. The index can include a multi-level index, a hash index, a full-text index, etc. The multi-level index classifies the data at multiple levels, thereby accelerating the data query speed. The hash index quickly queries the data through a hash algorithm. The full-text index performs full-text retrieval on text data to meet more complex query requirements.
[0049] In some embodiments, the system may further include one or more downstream application devices. By accessing the data warehouse device, the one or more application devices can consume the result data. Each application device may correspond to one or more application scenarios. The data warehouse device can provide the result data to the one or more application devices. Each application device can obtain the corresponding result data from the data warehouse according to its own application scenario. For example, the application device can obtain the OpenAPI interface of the data warehouse device by accessing the data warehouse device. The application device can call the data warehouse device through the OpenAPI interface to obtain the result data. Thus, the application devices in multiple application scenarios can obtain the result data aggregated from the Internet of Things data in the data lake by accessing the data warehouse device. It enables the Internet of Things data collected by each collection device to be applied to multiple different application scenarios, thereby realizing the co-construction and sharing of the Internet of Things data collected by the collection device in multiple application scenarios.
[0050] In some embodiments, each application device may send a result data acquisition request to the data warehouse device. The result data acquisition request includes a scenario identifier for identifying the application scenario. The data warehouse device can receive the data acquisition request; it can detect whether the result data corresponding to the scenario identifier is stored in the data warehouse; if so, it indicates that the data warehouse device has pre-generated the result data, and it can obtain the result data corresponding to the scenario identifier in the data warehouse and feedback the obtained result data to the application device; if not, it indicates that the data warehouse device has not pre-generated the result data, and it can obtain multiple Internet of Things data corresponding to the scenario identifier from the data lake; it can perform aggregation analysis on the obtained multiple Internet of Things data; it can store the result data after aggregation analysis in the data warehouse; it can send the result data after aggregation analysis to the application device. The application device can receive the result data. Thus, for application scenarios with relatively high real-time requirements (such as real-time monitoring, real-time alerting, etc.), the data warehouse device can pre-generate the result data. When the application device needs the result data, it can promptly provide the result data to the application device. For application scenarios with relatively low real-time requirements (such as historical data analysis, report generation, etc.), the data warehouse device can generate the result data after receiving the data acquisition request from the application device and provide the generated result data to the application device. This can flexibly select the result data generation strategy according to different real-time requirements, thereby effectively meeting the real-time requirements of different application scenarios and improving the efficiency and accuracy of data processing.
[0051] In some embodiments, each of the multiple application devices can obtain result data; and can display result data. In addition, each of the multiple application devices can also detect whether the obtained result data is abnormal. For example, the application device can detect whether the result data meets the preset conditions; if so, it is determined that the result data is normal; if not, it is determined that the result data is abnormal. For another example, the application device can also input the result data into an artificial intelligence model. The output of the artificial intelligence model is used to indicate whether the result data is abnormal.
[0052] In the following, an application device that detects that the result data is abnormal may be used as a first application device. The first application device may send a first prompt message to a data warehouse device. The data warehouse device may receive the first prompt message. The first prompt message is used to prompt that the result data is abnormal. Optionally, the first prompt message may include abnormal result data.
[0053] The data warehouse device can receive the first prompt information. Since the result data in each application scenario can be generated by multiple IoT data. Therefore, the data warehouse device can obtain multiple IoT data used to generate abnormal result data from the data lake; can analyze the multiple IoT data used to generate abnormal result data to select suspected abnormal IoT data; can obtain suspected abnormal result data generated by suspected abnormal IoT data; can determine a second application device that receives the suspected abnormal result data from the multiple application devices; and can send a second prompt information to the second application device.
[0054] The data warehouse device can receive various IoT data fed back by the data lake device by sending a request to the data lake device. For example, the data warehouse device can determine the generation time of the abnormal result data; can determine the time interval based on the generation time; and can send a request to the data lake device. The data lake device can feed back various IoT data within the time interval to the data warehouse device. Of course, the data warehouse device can also use other methods to obtain various IoT data used to generate abnormal result data.
[0055] The data warehouse device can determine whether the multiple IoT data used to generate abnormal result data meet the abnormal conditions; IoT data that meet the abnormal conditions can be selected as suspected abnormal IoT data. Abnormal conditions may include, for example: IoT data is greater than a first threshold or less than a second threshold. Greater than the first threshold indicates that the IoT data is too large, and less than the second threshold indicates that the IoT data is too small. Alternatively, the data warehouse device can input the multiple IoT data used to generate abnormal result data into the artificial intelligence model. The artificial intelligence model outputs suspected abnormal IoT data.
[0056] The Internet of Things data collected by each collection device can be applied to a variety of different application scenarios, thereby realizing the sharing of the Internet of Things data collected by the collection device in multiple application scenarios. Thus, multiple result data can be generated from the suspected abnormal Internet of Things data. The multiple result data may include abnormal result data (the result data obtained by the first application device). The result data other than the abnormal result data in the multiple result data can be used as suspected abnormal result data.
[0057] Each application scenario can correspond to an application device, and this application device can receive the result data in this application scenario. For this purpose, the data warehouse device can determine a second application device that receives the suspected abnormal result data among multiple application devices; and can send a second prompt message to the second application device. The second prompt message is used to prompt that the result data received by the second application device may be abnormal. The second prompt message may include the suspected abnormal result data. After receiving the second prompt message, the second application device can detect the received suspected abnormal result data to confirm whether the received suspected abnormal result data is abnormal; and can send the detection result of the suspected abnormal result data to the data warehouse device. The method for the second application device to detect whether the suspected abnormal result data is abnormal is similar to the method for the first application device to detect whether the result data is abnormal, which will not be elaborated here. Alternatively, after receiving the second prompt message, the second application device can also prompt the business personnel for confirmation. The business personnel can review and confirm whether the suspected abnormal result data is abnormal, and can input the detection result of the suspected abnormal result data in the second application device. The second application device can receive the detection result of the suspected abnormal result data.
[0058] Optionally, after receiving the second prompt message, the second application device may have detected the suspected abnormal result data or may not have detected the suspected abnormal result data. In the case of not detecting the suspected abnormal result data, the second application device can use a method similar to the method for the first application device to detect whether the result data is abnormal to detect whether the suspected abnormal result data is abnormal. In the case of having detected the suspected abnormal result data, the second application device can prompt the business personnel for confirmation so that the business personnel can review and confirm whether the suspected abnormal result data is abnormal.
[0059] The detection result of the suspected abnormal result data can be selected from abnormal and normal. When the detection result of the suspected abnormal result data indicates normal, it means that the suspected abnormal IoT data is normal, and the abnormal result data received by the first application device occurs accidentally. The data warehouse device can ignore the first prompt message. When the detection result of the suspected abnormal result data indicates abnormal, it means that the suspected abnormal IoT data is indeed abnormal. The data warehouse device can send the suspected abnormal IoT data to the data lake device. The data lake device can receive the suspected abnormal IoT data; it can determine the target collection device that collects the suspected abnormal IoT data among multiple collection devices; and it can send the third prompt message to the target collection device. The target collection device can receive the third prompt message; and it can issue an alarm according to the third prompt message.
[0060] It can be understood that the number of the suspected abnormal result data is one or more, and the number of the second application devices is one or more. The data warehouse device can determine that the suspected abnormal IoT data is indeed abnormal when the detection result of any second application device is abnormal. Or, the data warehouse device can also determine that the suspected abnormal IoT data is indeed abnormal when the detection results of all second application devices are abnormal. This can further improve the accuracy of determining abnormal IoT data.
[0061] Based on the characteristic that the IoT data collected by each collection device can be applied to multiple different application scenarios, by mutually verifying the result data in multiple application scenarios, the abnormal IoT data can be accurately determined. So that the target collection device for collecting the abnormal IoT data can issue an alarm. It is convenient for the maintenance of business personnel.
[0062] Please refer to Figure 4 ... Based on the foregoing IoT data processing system, an embodiment of this specification further provides an IoT data processing method based on a data lake. The IoT data processing method based on a data lake can be applied to the foregoing IoT data processing system based on a data lake, and specifically may include the following steps.
[0063] Step 41, use multiple collection devices to collect multiple types of IoT data.
[0064] Step 42, use the data lake device to store the multiple types of IoT data collected into the data lake.
[0065] Step 43, for each application scenario among multiple application scenarios, use the data warehouse device to obtain the multiple types of IoT data required for this application scenario from the data lake, perform aggregation analysis on the multiple types of IoT data obtained according to the business logic of this application scenario, and store the result data after the aggregation analysis into the data warehouse to provide the result data to the downstream application devices.
[0066] In the Internet of Things (IoT) data processing method according to the embodiments of this specification, a data lake can store IoT data collected by various collection devices; a data warehouse is built based on the data lake, can perform aggregation analysis on the IoT data in the data lake according to application scenarios, and can store result data under various application scenarios. Thus, the separation of computing and storage is achieved, and the storage cost can be reduced. In addition, for each application scenario, the data warehouse device can perform aggregation analysis on various IoT data required for that application scenario to obtain the result data for that application scenario. Thus, the effective integration of various IoT data is achieved, which is conducive to improving the accuracy of data analysis in application scenarios. In addition, the IoT data collected by each collection device can be applied to multiple different application scenarios, so that the IoT data collected by the collection device can be shared under multiple application scenarios. The technical solution of the embodiments of this specification reduces the storage cost through the lake and warehouse entry of IoT data, and realizes the co-construction and sharing of IoT data. Through the co-construction and sharing of IoT data, the consistency and integrity of data can be effectively improved, the potential value of IoT data can be fully explored, strong support can be provided for decision-making, thereby improving the intelligent level of IoT applications and effectively meeting user needs.
[0067] The embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned IoT data processing method based on a data lake is implemented.
[0068] The embodiments of this specification also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned IoT data processing method based on a data lake is implemented.
[0069] The embodiments of this specification also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned IoT data processing method based on a data lake is implemented.
[0070] Those skilled in the art can understand that this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. The computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0072] Each functional unit in the embodiments of this specification can be integrated into a processing unit, or each functional unit can exist physically alone, or two or more functional units can be integrated into a processing unit.
[0073] Those skilled in the art can understand that the descriptions of the embodiments in this specification each have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Additionally, it can be understood that after reading this specification document, those skilled in the art can, without creative labor, think of arbitrarily combining some or all of the embodiments listed in this specification, and these combinations are also within the scope of disclosure and protection of this specification.
[0074] Although this specification is depicted through embodiments, those of ordinary skill in the art know that the above embodiments are only used to help understand the core idea of this specification. Those skilled in the art can understand that this specification has many variations and changes. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification.
Claims
1. An Internet of Things data processing system based on a data lake, characterized in that The Internet of Things (IoT) data processing system based on a data lake includes multiple types of collection devices, a data lake device, and a data warehouse device; The multiple types of collection devices are used to collect multiple types of IoT data and send the collected IoT data to the data lake device; The data lake device is used to store the received IoT data into the data lake; The data warehouse device is used for each of multiple application scenarios to obtain the multiple types of IoT data required for that application scenario from the data lake, perform aggregation analysis on the obtained multiple types of IoT data according to the business logic of that application scenario, and store the result data after the aggregation analysis into a data warehouse to provide the result data to downstream application devices.
2. The Internet of Things data processing system based on a data lake according to claim 1, characterized in that The IoT data processing system based on a data lake further includes an IoT device management platform. The multiple types of collection devices are connected to the IoT device management platform, and each collection device sends the collected IoT data to the IoT device management platform; The IoT device management platform pushes the received IoT data to the data lake device.
3. The IoT data processing system based on a data lake according to claim 1, wherein the data lake device is used to identify the data type of the IoT data; determine a corresponding storage strategy according to the data type of the IoT data; and store the IoT data into a corresponding storage medium according to the storage strategy.
4. The IoT data processing system based on a data lake according to claim 1, wherein the data lake device is used to identify the popularity of the IoT data in the data lake and the corresponding storage medium; and select the IoT data for archival storage according to the popularity of the IoT data and the corresponding storage medium.
5. The IoT data processing system based on a data lake according to claim 1, wherein when the data type of the result data is device status data, the data warehouse device stores the result data into a time series database; when the data type of the result data is archive data, the data warehouse device stores the result data into a relational database.
6. The IoT data processing system based on a data lake according to claim 1, wherein the data warehouse device includes multiple physical subsystems corresponding to multiple application scenarios; each physical subsystem is used to obtain the multiple types of IoT data required for that application scenario from the data lake according to its own application scenario, perform aggregation analysis on the obtained multiple types of IoT data, and store the result data after the aggregation analysis into the data warehouse of the physical subsystem.
7. The IoT data processing system based on a data lake according to claim 1, wherein the data warehouse device is used for each of multiple application scenarios to obtain the multiple types of IoT data required for that application scenario from the data lake and generate derivative data according to the obtained multiple types of IoT data; the result data includes the derivative data.
8. The Internet of Things data processing system based on a data lake according to claim 1, characterized in that The data warehouse device is used to determine whether each of multiple application scenarios meets the real-time condition; if an application scenario meets the real-time condition, it obtains multiple types of Internet of Things data in this application scenario from the data lake at regular intervals, performs aggregation analysis on the obtained multiple types of Internet of Things data, and stores the result data after aggregation analysis in the data warehouse.
9. The Internet of Things data processing system based on a data lake according to claim 1, wherein, The Internet of Things data processing system based on the data lake further includes at least one downstream application device, and each application device corresponds to an application scenario; The application device is used to send a result data acquisition request to the data warehouse device; The data warehouse device is used to receive the result data acquisition request; detect whether the result data corresponding to the application scenario of the application device is stored in the data warehouse; if so, obtain the result data corresponding to the application scenario of the application device in the data warehouse and send the obtained result data to the application device; if not, obtain multiple types of Internet of Things data required for the application scenario of the application device from the data lake, perform aggregation analysis on the obtained multiple types of Internet of Things data, store the result data after aggregation analysis in the data warehouse, and send the result data after aggregation analysis to the application device; The application device is further used to receive the result data.
10. An Internet of Things data processing method based on a data lake, characterized in that, Applied to the Internet of Things data processing system based on the data lake according to any one of claims 1-9 above, it includes: Using multiple collection devices to collect multiple types of Internet of Things data; Using the data lake device to store the collected multiple types of Internet of Things data in the data lake; Using the data warehouse device to obtain multiple types of Internet of Things data required for each of multiple application scenarios from the data lake, perform aggregation analysis on the obtained multiple types of Internet of Things data according to the business logic of this application scenario, store the result data after aggregation analysis in the data warehouse, and provide the result data to the downstream application device.