Data processing system and method based on multi-lake one-in-one platform architecture
The data processing system, built on a multi-lake, single-platform architecture, solves the problem that traditional data centers cannot meet the unified management and cross-domain sharing needs of group enterprises. It achieves unified data classification and standardized management, improves data governance efficiency and sharing capabilities, and promotes cross-domain data circulation and the activation of business value.
Patent Information
- Application Number
- CN202411036155.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-03
AI Technical Summary
Traditional split data centers cannot meet the unified management and cross-domain sharing needs of enterprise data processing. They lack a unified data resource catalog and data service list, resulting in long data application development cycles and the inability to reuse valuable data.
It adopts a multi-lake, single-platform architecture, including a data platform and multi-level data lakes, to provide unified data classification, unified data resource catalog and data standard management. It achieves unified data management and cross-domain sharing through data governance module, data service module and platform control module.
It enables data sharing and circulation across the entire domain, improves data governance efficiency and quality, supports cross-domain and cross-enterprise data sharing, empowers business operation decisions and innovative development, and activates the commercial value of data.
Smart Images

Figure CN121456049A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and specifically to a data processing system and method based on a multi-lake, single-platform architecture. Background Technology
[0002] Large group enterprises face the requirement of data transformation. However, due to factors such as the different geographical locations of the group's application systems and the security of production data in various subsidiaries, it is necessary to establish different data lakes in a decentralized manner. At the same time, a unified data platform needs to be built to achieve unified management of data resource catalogs, data standards, and data services, and support the circulation and sharing of data. Traditional decentralized data centers cannot meet these needs. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data processing system and method based on a multi-lake, single-platform architecture to solve the problem that traditional discrete data centers in the prior art cannot meet enterprise data processing tasks.
[0004] According to a first aspect, embodiments of the present invention provide a data processing system based on a multi-lake, single-platform architecture, the system comprising:
[0005] The data platform is deployed at the enterprise group headquarters to provide unified data classification, unified data resource catalog and unified data standard management functions;
[0006] Several levels of data lakes are deployed in various departments of the enterprise group. Each level of data lake is connected to the data lakes of the adjacent levels, and each level of data lake is connected to the data platform for data interaction.
[0007] In conjunction with the first aspect, in the first embodiment of the first aspect, the data platform specifically includes:
[0008] The data governance module provides technical components for data catalog management, data model management, data standard management, metadata management, master data management, data quality management, and data indicator management.
[0009] The data service module is used to store data;
[0010] The data security module is used to execute preset security policies and procedures to protect data;
[0011] The middle platform management module is used to manage and control the data stored in the data service module and the data lake at each level, according to the various levels of organizations in the enterprise group.
[0012] In conjunction with the first aspect, in the second embodiment of the first aspect, the data lakes at several levels include a headquarters data lake and hierarchical data lakes. The headquarters data lake is set as one and is the highest level data lake. The hierarchical data lakes are set as multiple levels, wherein each level of the data lake has at least one, and each level of the data lake has corresponding level permissions.
[0013] In conjunction with the second implementation of the first aspect, in the third implementation of the first aspect, the data lake includes a headquarters data lake, a module data lake, and an enterprise data lake. The module data lake is connected to the enterprise data lake, the headquarters data lake, and the data platform, and the enterprise data lake is connected to the data platform.
[0014] In conjunction with the second embodiment of the first aspect, in the fourth embodiment of the first aspect, each of the data lakes includes:
[0015] The first processing module is used to store structured data;
[0016] The second processing module is used to store unstructured data;
[0017] The third processing module is used to store time-series data.
[0018] In conjunction with the fourth embodiment of the first aspect, in the fifth embodiment of the first aspect, the first processing module uses FLinkCDC and Kafka components to collect structured data.
[0019] In conjunction with the fourth embodiment of the first aspect, in the sixth embodiment of the first aspect, the second processing module uses Flume and FTPS components to collect semi-structured data.
[0020] In conjunction with the fourth embodiment of the first aspect, in the seventh embodiment of the first aspect, the third processing module uses IoT and FLink components to collect industrial field data.
[0021] According to a second aspect, the present invention provides a data processing method based on a multi-lake, single-platform architecture, applied to a data lake, the method comprising:
[0022] Collect raw data and store it in the source layer of the data lake;
[0023] The changes in the source layer are determined, and the changed data is cleaned according to the data resource catalog and the corresponding data processing standards pre-stored in the data platform connected to the data lake. The cleaned data is then stored in the governance layer of the data lake. The data platform is deployed at the enterprise group headquarters and is used to acquire data stored in the data lake and process the data.
[0024] According to a third aspect, the present invention provides a data processing method based on a multi-lake, single-platform architecture, applied to a data platform, the method comprising:
[0025] The production task logs and messages generated by the data lake connected to the data platform are identified, and data changes in the data lake are determined based on the production task logs and messages; the data lake is deployed in various levels of the enterprise group and is used to store the data of the corresponding organizations;
[0026] The task message in the data is updated based on data changes; the task information includes at least one of the following: task number, task type change, source data model, target data model, and task execution status;
[0027] Consume task messages and generate a full-chain task tree based on the task messages, then start the full-chain task tree to monitor the data lake information.
[0028] The data processing system and method based on a multi-lake, single-platform architecture provided by this invention achieves unified data classification, unified data resource catalog, and unified data standard management capabilities. During the data entry process at all levels of data lakes, it completes data governance and data assetization, enabling big data applications for groups, business units, and enterprises. It also achieves unified data service management and multi-lake control capabilities, providing a unified sharing process and mechanism at the group level, completing the sharing and circulation of data across the entire domain, and realizing cross-domain and cross-enterprise data sharing. Attached Figure Description
[0029] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings:
[0030] Figure 1 This is a schematic diagram of the structure of a data processing system based on a multi-lake, single-platform architecture according to an embodiment of this application;
[0031] Figure 2 This is a schematic diagram of the specific structure of the data platform in the data processing system based on the multi-lake-one-platform architecture according to an embodiment of this application;
[0032] Figure 3 This is a schematic diagram illustrating the interaction between the data lake and the data platform in a data processing system based on a multi-lake, single-platform architecture according to an embodiment of this application.
[0033] Figure 4 This is a flowchart illustrating the application of the data processing method based on a multi-lake, single-platform architecture according to an embodiment of this application to a data lake.
[0034] Figure 5 This is a flowchart illustrating the application of the data processing method based on a multi-lake, single-platform architecture according to an embodiment of this application to a data platform.
[0035] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Large group enterprises face the requirement of data transformation, with the goal of building a reusable data asset center and data capability center across the entire domain, realizing full lifecycle management of all types of data, and making data a resource like water and electricity, available on demand, enabling digital business operations.
[0038] Due to factors such as the different geographical locations of the group's application systems and the security of production data of various subsidiaries, it is necessary to establish different data lakes in a decentralized manner. At the same time, a unified data platform needs to be built to achieve unified management of data resource catalogs, data standards, and data services, and support the circulation and sharing of data.
[0039] Traditional split-type data centers cannot meet the above requirements, and the challenges and problems they face are:
[0040] 1. The lack of a unified sharing process and mechanism at the group level means that there are still certain barriers to the sharing and circulation of data across the entire domain, making it difficult to share data across domains and enterprises;
[0041] 2. The level of data governance varies among companies, there is a lack of a unified data asset catalog and data service list for the group, the data application development cycle is long, and valuable data cannot be reused.
[0042] To address the aforementioned issues, this embodiment provides a data processing system based on a multi-lake, single-platform architecture. Figure 1 This is a schematic diagram of the structure of a data processing system based on a multi-lake, single-platform architecture according to an embodiment of the present invention, as shown below. Figure 1 As shown, the system includes:
[0043] Data Platform 10, specifically, is centrally deployed at the enterprise group headquarters. Located there, it provides unified data classification, a unified data resource catalog, and unified data standard management capabilities. This enables unified data classification, a unified data resource catalog, and unified data standard management capabilities, supporting data governance and assetization during the data ingestion process at all levels of data lakes. It also strengthens data modeling for business scenarios, deeply mines data value, and enhances data insight capabilities. Simultaneously, it achieves unified data service management and multi-lake control capabilities, providing a unified sharing process and mechanism at the group level to complete data sharing and circulation across all domains, enabling cross-domain and cross-enterprise data sharing.
[0044] Several levels of data lakes 20 are deployed in various levels of organizations within the enterprise group. Each level of data lake 20 is connected to the corresponding level of organization, and each level of data lake 20 is connected to the adjacent level of data lake 20. Furthermore, each level of data lake 20 is connected to the data platform 10 for data interaction and to receive control from the data platform 10.
[0045] The group-level data resource system based on a multi-lake, single-platform architecture specifically refers to:
[0046] "Multi-lake": A three-tiered data lake for the group, business units, and enterprises, designed to provide data acquisition capabilities for multi-source heterogeneous, real-time / batch data, as well as data processing capabilities such as multi-modal data storage and data processing;
[0047] "One-stop data platform": The data platform centrally deployed at headquarters aims to provide unified data governance capabilities, platform control capabilities, data security capabilities, and data service output capabilities.
[0048] The system builds six core capabilities: data acquisition (multi-source heterogeneous, real-time / batch data), data processing (multi-modal data storage and processing), data governance (metadata management, master data management), data service (API, data subscription), data security protection, and middle platform management. These capabilities empower the data application of the group, business units, and enterprises.
[0049] The data lake in this system provides multiple data source access capabilities, data synchronization capabilities, task scheduling capabilities, and data transmission optimization capabilities, enabling the rapid and secure transfer of data from the source system to the lake.
[0050] Data lakes provide massive data storage and management capabilities from multiple heterogeneous sources, including: data file storage capabilities based on the distributed file storage system HDFS, supporting multiple data file formats such as CSV, RCF, Parquet, Avro, and Sequence File; and massive semi-structured and unstructured data storage and fast access capabilities based on object storage such as OSS / MinIO, including video, audio files, electronic certificates, electronic documents, and logs.
[0051] Data governance is a core element of the construction, management, and use of a data platform. It provides management and display capabilities for data asset catalogs, metadata, data quality, data lineage, and data lifecycle, presenting the enterprise's data assets in a more intuitive way.
[0052] This system also provides rapid service generation capabilities, as well as service control, authentication, and measurement functions, transforming data into a service capability that allows data to easily participate in business operations and bring value to those operations. Data security is the prerequisite and guarantee for the normal operation of the data platform, including classification and grading, and security protection and monitoring throughout the data lifecycle. Platform management is the foundation for the healthy and continuous operation of the data platform, including supervision of process standard execution, monitoring and optimization of platform resource usage, assessment of data value, and statistical analysis of data service applications.
[0053] Please see Figure 2 The data platform 10 in this system specifically includes:
[0054] Data governance module 11 provides technical components such as data catalog management, data model management, data standard management, metadata management, master data management, data quality management, and data indicator management, supporting the entire group using the system to complete data governance through a 7-step method.
[0055] Data service module 12 is used to provide data-as-a-service capabilities to the outside world in the form of RESTful API by constructing SQL statements, thereby configuring datasets, sorting rules and query conditions in a visual way to form custom data services.
[0056] Data security module 13 is used to plan, formulate, and implement relevant security policies and procedures, build layers and provide graded protection, and ensure that data and information assets have appropriate authentication, authorization, access and audit measures during use. It aims to ensure that the right people use and update data in the right way through effective data security policies and procedures, and to restrict all inappropriate access and updates of data.
[0057] The middle platform management module 14 is used to differentiate by organization and permissions to achieve unified management and control of middle platform capabilities, headquarters lake, and enterprise lake. Through standardized data service processes, it promotes the efficient operation of the four-party mechanism and achieves continuous improvement in the service quality of the middle platform.
[0058] In this embodiment of the invention, the data lake 20 includes a headquarters data lake and multiple levels of hierarchical data lakes. The headquarters data lake is set as one and is the highest level data lake. The hierarchical data lakes are set as several levels. Each level of data lake is set with at least one. Data at different levels has corresponding level permissions. Each level of data lake is connected to the adjacent level of data lake, and each hierarchical data lake is connected to the headquarters data lake.
[0059] Of course, the number of levels in Data Lake 20 can be adjusted according to the overall architecture of the enterprise, such as 4 levels, 5 levels, etc.
[0060] It should be noted that there is only one highest-level Data Lake 20, while the number of non-highest-level Data Lake 20s is set according to the enterprise's situation.
[0061] Please see Figure 3 Specifically, in one possible embodiment of the present invention, for example, the data lakes 20 at various levels may include: the highest-level data lake, the next-lower-level data lake, and the lowest-level data lake. Specifically, the data lake 20 includes a headquarters data lake, module data lakes, and enterprise data lakes. The module data lakes are connected to the enterprise data lake, the headquarters data lake, and the data platform 10, and the enterprise data lake is connected to the data platform 10.
[0062] All data lakes 20 in this system include:
[0063] The first processing module is used to store structured data. The structured data is collected from various levels of structured data and put into the lake using technologies such as FLinkCDC and Kafka. The distributed file storage system HDFS provides data file storage capabilities and supports multiple data file formats such as CSV, RCF, Parquet, Avro, and Sequence File.
[0064] The second processing module is used to store unstructured data. The unstructured data is collected into the lake using technologies such as Flume and FTPS to obtain semi-structured data, files, and images. Based on object storage such as OSS / MinIO, it provides massive semi-structured and unstructured data storage and fast access capabilities for video, audio files, electronic certificates, electronic documents, logs, etc.
[0065] The third processing module is used to store time-series data, which is collected from real-time industrial sites using technologies such as IoT and FLink. Based on a distributed time-series database, it provides high-throughput real-time writing of massive amounts of real-time data that change rapidly over time, precise time-series querying, and ultra-high data compression capabilities.
[0066] The data processing system based on a multi-lake, single-platform architecture provided by this invention achieves unified data classification, unified data resource catalog, and unified data standard management capabilities. During the data entry process at all levels of data lakes, it completes data governance and data assetization, enabling big data applications for groups, business units, and enterprises. It also achieves unified data service management and multi-lake control capabilities, providing a unified sharing process and mechanism at the group level, completing the sharing and circulation of data across the entire domain, and realizing cross-domain and cross-enterprise data sharing.
[0067] Please see Figure 4 This invention also provides a data processing method based on a multi-lake, single-platform architecture, which is applied to data lakes deployed at various levels of an enterprise. The method includes:
[0068] A10. Collect raw data and store it in the source layer of the data lake.
[0069] Each level of the data lake uses components such as Flink CDC to collect data from various source systems and store it in the source-attached layer of the data lake.
[0070] A20. Determine the data changes in the source layer, clean the changed data according to the data resource catalog pre-stored in the data platform connected to the data lake and the data processing standards corresponding to the data resource catalog, and save the cleaned data in the governance layer of the data lake.
[0071] Components such as Flink CDC capture changes in source layer data, and clean the data and store it in the governance layer based on the unified data resource catalog, master data, and data standards in the data platform.
[0072] Similarly, the various computing services of the data lake, based on the data resource catalog, aggregate, calculate, and store data from the aggregation layer and application layer. The aforementioned computing processes, production task logs, and messages are all collected and monitored uniformly by the data platform.
[0073] The data processing method based on a multi-lake, single-platform architecture of the present invention follows a unified technical route, builds a multi-lake, single-platform system for group enterprises, creates a data hub that connects internally and externally, improves data governance efficiency and data quality internally, realizes real-time data sharing, empowers business operation decisions and innovative development, and supports the digital transformation of the group company; externally, it promotes the orderly flow and innovative use of data, activates the commercial value of data, empowers partners to enhance value, and cultivates a "shared and win-win" digital ecosystem.
[0074] Please see Figure 5 This invention also provides a data processing method based on a multi-lake, single-platform architecture, which is applied to a data platform deployed at the headquarters of an enterprise group. The method includes:
[0075] B10. Identify the production task logs and messages generated by the data lake connected to the data platform, and determine the data changes in the data lake based on the production task logs and messages.
[0076] B20. Update task messages in the data based on data changes; the task information includes at least one of the following: task number, task type change, source data model, target data model, and task execution status.
[0077] B30. Consume task messages, generate a full-chain task tree based on the task messages, and start the full-chain task tree to monitor data lake information.
[0078] The above describes the data service access process. When deploying data collection, governance computing, and aggregation computing operators (including FLink and IoT) at all levels of the data lake, they need to register with the access task monitoring service and provide information such as task type number, source data model, and target data model.
[0079] The task monitoring and registration service updates the lineage analysis information in the metadata management of the data platform based on the source data model and the target data model, and generates a full-chain task tree triggered when the corresponding data in the source system changes, based on the full-chain analysis capability of the metadata.
[0080] After FLink CDC captures changes in source system data, it generates a task message and pushes it into Kafka. The message content includes the task number, task type change, source data model, target data model, and task execution status.
[0081] The task monitoring service consumes task messages and initiates full-chain task tree process monitoring, thus enabling data lakes at all levels to provide access to data resources to the outside world through a service approach.
[0082] Data services are uniformly registered with the service gateway of the data platform for unified management and control. Service deployment locations are chosen based on the location of the data source they access, and are deployed near each data lake. The service gateway uses the Kong component to build a two-level deployment, and service access addresses are based on domain names.
[0083] Data consumers across the lake can search for data directories and access directories on the data platform, apply for access permissions, and obtain access rights after approval by the data owner.
[0084] After completing authentication and authorization on the data platform, data consumers can directly access data services from the nearest location via routing.
[0085] The data processing method based on a multi-lake, single-platform architecture of the present invention follows a unified technical route, builds a multi-lake, single-platform system for group enterprises, creates a data hub that connects internally and externally, improves data governance efficiency and data quality internally, realizes real-time data sharing, empowers business operation decisions and innovative development, and supports the digital transformation of the group company; externally, it promotes the orderly flow and innovative use of data, activates the commercial value of data, empowers partners to enhance value, and cultivates a "shared and win-win" digital ecosystem.
[0086] Example 2
[0087] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical commands in the memory 530 to execute a data processing method based on a multi-lake, single-platform architecture, applied to a data lake. This method includes:
[0088] Collect raw data and store it in the source layer of the data lake;
[0089] The changes in the source layer are determined, and the changed data is cleaned according to the data resource catalog and the corresponding data processing standards pre-stored in the data platform connected to the data lake. The cleaned data is then stored in the governance layer of the data lake. The data platform is deployed at the enterprise group headquarters and is used to acquire data stored in the data lake and process the data.
[0090] Alternatively, when applied to a data platform, the method includes:
[0091] The production task logs and messages generated by the data lake connected to the data platform are identified, and data changes in the data lake are determined based on the production task logs and messages; the data lake is deployed in various levels of the enterprise group and is used to store the data of the corresponding organizations;
[0092] The task message in the data is updated based on data changes; the task information includes at least one of the following: task number, task type change, source data model, target data model, and task execution status;
[0093] Consume task messages, generate a full-chain task tree based on the task messages, and start the full-chain task tree to monitor data lake information.
[0094] Furthermore, the logical commands in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent media, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software medium. This computer software medium is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0095] Example 3
[0096] On the other hand, the present invention also provides a computer program medium, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data processing method based on the multi-lake-one-platform architecture provided by the above methods, applied to a data lake. The method includes:
[0097] Collect raw data and store it in the source layer of the data lake;
[0098] The changes in the source layer are determined, and the changed data is cleaned according to the data resource catalog and the corresponding data processing standards pre-stored in the data platform connected to the data lake. The cleaned data is then stored in the governance layer of the data lake. The data platform is deployed at the enterprise group headquarters and is used to acquire data stored in the data lake and process the data.
[0099] Alternatively, when applied to a data platform, the method includes:
[0100] The production task logs and messages generated by the data lake connected to the data platform are identified, and data changes in the data lake are determined based on the production task logs and messages; the data lake is deployed in various levels of the enterprise group and is used to store the data of the corresponding organizations;
[0101] The task message in the data is updated based on data changes; the task information includes at least one of the following: task number, task type change, source data model, target data model, and task execution status;
[0102] Consume task messages, generate a full-chain task tree based on the task messages, and start the full-chain task tree to monitor data lake information.
[0103] Example 4
[0104] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the data processing method based on a multi-lake, single-platform architecture provided by the above methods, applied to a data lake, the method comprising:
[0105] Collect raw data and store it in the source layer of the data lake;
[0106] The changes in the source layer are determined, and the changed data is cleaned according to the data resource catalog and the corresponding data processing standards pre-stored in the data platform connected to the data lake. The cleaned data is then stored in the governance layer of the data lake. The data platform is deployed at the enterprise group headquarters and is used to acquire data stored in the data lake and process the data.
[0107] Alternatively, when applied to a data platform, the method includes:
[0108] The production task logs and messages generated by the data lake connected to the data platform are identified, and data changes in the data lake are determined based on the production task logs and messages; the data lake is deployed in various levels of the enterprise group and is used to store the data of the corresponding organizations;
[0109] The task message in the data is updated based on data changes; the task information includes at least one of the following: task number, task type change, source data model, target data model, and task execution status;
[0110] Consume task messages, generate a full-chain task tree based on the task messages, and start the full-chain task tree to monitor data lake information.
[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of software media. This computer software media can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing system based on a multi-lake, single-platform architecture, characterized in that, The system includes: The data platform is deployed at the enterprise group headquarters to provide unified data classification, unified data resource catalog and unified data standard management functions; Several levels of data lakes are deployed in various departments of the enterprise group. Each level of data lake is connected to the data lakes of the adjacent levels, and each level of data lake is connected to the data platform for data interaction.
2. The data processing system based on a multi-lake, single-platform architecture as described in claim 1, characterized in that, The data platform specifically includes: The data governance module provides technical components for data catalog management, data model management, data standard management, metadata management, master data management, data quality management, and data indicator management. The data service module is used to store data; The data security module is used to execute preset security policies and procedures to protect data; The middle platform management module is used to manage and control the data stored in the data service module and the data lake at each level, according to the various levels of organizations in the enterprise group.
3. The data processing system based on a multi-lake, single-platform architecture according to claim 1, characterized in that, The data lakes at several levels include a headquarters data lake and hierarchical data lakes. The headquarters data lake is set as a single, highest-level data lake, and the hierarchical data lakes are set as multiple levels. Each level of the data lake has at least one data lake, and each level of the data lake has corresponding level permissions.
4. The data processing system based on a multi-lake, single-platform architecture according to claim 3, characterized in that, The data lake includes a headquarters data lake, a module data lake, and an enterprise data lake. The module data lake is connected to the enterprise data lake, the headquarters data lake, and the data platform. The enterprise data lake is connected to the data platform.
5. The data processing system based on a multi-lake, single-platform architecture according to claim 3, characterized in that, Each of the data lakes includes: The first processing module is used to store structured data; The second processing module is used to store unstructured data; The third processing module is used to store time-series data.
6. The data processing system based on a multi-lake, single-platform architecture according to claim 5, characterized in that: The first processing module uses FLinkCDC and Kafka components to collect structured data.
7. The data processing system based on a multi-lake, single-platform architecture according to claim 5, characterized in that, The second processing module uses Flume and FTPS components to collect semi-structured data.
8. The data processing system based on a multi-lake, single-platform architecture according to claim 5, characterized in that, The third processing module uses IoT and FLink components to collect industrial field data.
9. A data processing method based on a multi-lake, single-platform architecture, characterized in that, Applied to data lakes, the method includes: Collect raw data and store it in the source layer of the data lake; The changes in the source layer are determined, and the changed data is cleaned according to the data resource catalog and the corresponding data processing standards pre-stored in the data platform connected to the data lake. The cleaned data is then stored in the governance layer of the data lake. The data platform is deployed at the enterprise group headquarters and is used to acquire data stored in the data lake and process the data.
10. A data processing method based on a multi-lake, single-platform architecture, characterized in that, Applied to a data platform, the method includes: The production task logs and messages generated by the data lake connected to the data platform are identified, and data changes in the data lake are determined based on the production task logs and messages; the data lake is deployed in various levels of the enterprise group and is used to store the data of the corresponding organizations; The task message in the data is updated based on data changes; the task information includes at least: One of the following: task number, task type change, source data model, target data model, or task execution status; Consume task messages and generate a full-chain task tree based on the task messages, then start the full-chain task tree to monitor the data lake information.