Distributed data lake intelligent treatment and optimization system and method thereof
By integrating data governance, resource optimization, intelligent analysis and deep learning technologies in the data lake management system, we will build a distributed data lake intelligent governance and optimization system, which solves the problems of data quality management, resource utilization inefficiency and lack of intelligent capabilities in the existing data lake management methods, and realizes efficient management and value maximization of data lakes.
Patent Information
- Application Number
- CN202510071257.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing data lake management methods lack a unified and intelligent governance framework, resulting in problems such as data quality management, inefficient resource utilization, data value assessment and life cycle management, and lack of intelligence and adaptability, making it difficult to cope with complex and changeable data environments.
A distributed data lake intelligent governance and optimization system is proposed, and the efficient management and value maximization of data lakes are achieved by integrating technologies such as data governance, resource optimization, intelligent analysis and deep learning. The system includes a data governance module, a data optimization module, an intelligent analysis engine and a deep reinforcement learning decision-making system, and works in concert to provide a comprehensive, intelligent, and adaptive data lake management solution.
It has achieved continuous improvement of data quality, optimization of resource utilization efficiency, and maximization of data value, enhanced the overall efficiency of the data lake, reduced the need for manual intervention, improved the scalability and long-term effectiveness of the system, and solved the challenges of data privacy and security.
Smart Images

Figure CN119988362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a distributed data lake intelligent governance and optimization system and method. Background Art
[0002] With the advent of the big data era, enterprises and organizations are facing unprecedented data management challenges. Traditional data warehouses and data management systems have been unable to cope with massive, diverse, and high-speed data flows. In this context, data lakes have emerged to provide organizations with flexible and scalable data storage and management solutions. However, as the scale and complexity of data lakes continue to expand, a series of new problems have also emerged.
[0003] At present, although data lake technology is developing rapidly, it still has many shortcomings. First, data quality management has always been a thorny issue. Most existing data lake solutions lack effective data quality control mechanisms, resulting in the formation of "data swamps" and making valuable data drowned by low-quality data. Secondly, inefficient resource utilization is also a common problem. Existing resource allocation strategies are often static and cannot be dynamically adjusted according to real-time data usage, resulting in resource waste or performance bottlenecks. In addition, data value assessment and lifecycle management are also facing challenges. Many organizations find it difficult to accurately assess the actual value of data, resulting in important data being ignored or deleted prematurely, while useless data occupies storage resources for a long time.
[0004] Existing data lake management methods usually deal with these issues in a separate way and lack a unified and intelligent governance framework. For example, some methods focus on data quality control but ignore resource optimization; some methods focus on resource scheduling but fail to fully consider the value and life cycle of data. This fragmented approach leads to low overall system efficiency and fails to fully realize the potential of the data lake.
[0005] Another significant problem is that most existing data lake solutions lack intelligence and adaptability. They usually rely on preset rules and strategies and cannot effectively cope with complex and changing data environments. When faced with new data types, changing business needs or sudden system loads, these systems often fail to cope and require a lot of manual intervention and adjustment.
[0006] In addition, data privacy and security are also major challenges faced by existing data lake solutions. With the increasing demand for data sharing and collaborative analysis, how to achieve effective analysis of cross-source data while protecting data privacy has become an urgent problem to be solved. Existing methods are either too conservative, limiting the value of data mining, or have security risks and cannot meet the increasingly stringent data protection regulations. Summary of the invention
[0007] In response to these problems, the present invention proposes a distributed data lake intelligent governance and optimization system and method. The system aims to provide a comprehensive, intelligent, and adaptive data lake management solution, which integrates advanced technologies such as data governance, resource optimization, intelligent analysis, and deep learning to achieve efficient management and maximize the value of the data lake.
[0008] The present invention proposes a distributed data lake intelligent governance and optimization system, including:
[0009] Data governance module for:
[0010] Define and implement data governance rules;
[0011] Control access to data;
[0012] Provide application-oriented interfaces and services;
[0013] Maintain and manage the logical and physical structure of the data lake;
[0014] Specify and enforce data standards for the data lake;
[0015] A data optimization module, in communication with the data governance module, is used to:
[0016] Assess the health of your data lake;
[0017] Provide metadata and data tracking capabilities;
[0018] Execute data management strategies and provide data lifecycle management functions;
[0019] Define optimization goals and constraints to provide decision support for data management strategies and resource allocation;
[0020] Obtain data value scores through data value assessment models;
[0021] An intelligent analysis engine, in communication with the data governance module and the data optimization module, is used to:
[0022] Select the corresponding intelligent analysis method according to the data type to analyze the data in the data lake;
[0023] Generate analysis reports and compare with data expectations;
[0024] Dynamically adjust the weight of the analysis method based on the comparison results;
[0025] A deep reinforcement learning decision system is connected to the data optimization module and the intelligent analysis engine for:
[0026] Analyze data usage patterns by combining deep learning networks with reinforcement learning;
[0027] Formulate optimization and adjustment strategies;
[0028] Optimize and adjust storage, network and computing resources, and data processing tasks through the data lake management platform.
[0029] Preferably, the data governance module includes:
[0030] Data Quality Unit for:
[0031] Classify input data into high-quality and low-quality data;
[0032] Document the quality of input data;
[0033] Automatically discover and fix data quality issues based on data governance rules;
[0034] A data security unit is communicatively connected to the data quality unit and is used to:
[0035] Provide user rights management function;
[0036] Verify account permissions and control data access permissions;
[0037] A data service unit, in communication with the data security unit, is used to:
[0038] Provide external interface;
[0039] Provide specific data services based on application needs;
[0040] A data architecture unit, in communication with the data service unit, is used to:
[0041] Maintain and manage the logical and physical structure of the data lake;
[0042] A data standard unit is communicatively connected with the data architecture unit and is used for:
[0043] Specify and enforce data standards for your data lake.
[0044] Preferably, the data optimization module comprises:
[0045] Data Lake Health Unit, used to:
[0046] Assess the health of your data lake;
[0047] Calculate the storage capacity ratio index value, data quality ratio index value, resource utilization ratio index value, storage cost ratio index value, data access volume ratio index value and data value assessment ratio index value;
[0048] Calculate the health of the data lake based on the weight of each indicator;
[0049] A data lifecycle management unit is communicatively connected to the data lake health unit and is used to:
[0050] Provide metadata and data tracking capabilities;
[0051] Implement data management strategies;
[0052] Provide data lifecycle management functions;
[0053] The multi-agent decision-making unit is communicatively connected with the data lifecycle management unit and is used to:
[0054] Define optimization objectives and constraints;
[0055] Provide decision support for data management strategies and resource allocation through analysis and learning;
[0056] A data value assessment unit, which is in communication with the multi-agent decision-making unit, is used to:
[0057] Obtain data value scores through the data value assessment model.
[0058] Preferably, the intelligent analysis engine comprises:
[0059] Data classification unit for:
[0060] Categorize data according to data type;
[0061] An analysis method selection unit, in communication with the data classification unit, is used to:
[0062] Select the corresponding intelligent analysis method according to the data type;
[0063] A data analysis unit is communicatively connected to the analysis method selection unit, and is used to:
[0064] Analyze the data using the selected analytical methods;
[0065] A result comparison unit is connected in communication with the data analysis unit and is used for:
[0066] Compare analysis results with data expectations;
[0067] A weight adjustment unit is connected to the result comparison unit for:
[0068] Dynamically adjust the weight of the analysis method based on the comparison results.
[0069] Preferably, the deep reinforcement learning decision system comprises:
[0070] Deep learning network units for:
[0071] Learning data usage patterns through multi-layer neural networks;
[0072] A reinforcement learning unit, communicatively connected to the deep learning network unit, is used to:
[0073] Develop optimization and adjustment strategies based on the learned data usage patterns;
[0074] A decision execution unit, which is in communication with the reinforcement learning unit, is used to:
[0075] Deliver optimization and adjustment strategies to the data lake management platform;
[0076] A resource optimization unit, in communication with the decision execution unit, is used to:
[0077] Optimize storage, network and computing resources according to optimization adjustment strategies;
[0078] A task scheduling unit, in communication with the resource optimization unit, is used to:
[0079] Optimize and schedule data processing tasks based on optimization adjustment strategies.
[0080] Preferably, the data lake health unit is further used for:
[0081] Send an alert signal when the health of the data lake does not meet the standard;
[0082] The multi-agent decision-making unit is also used for:
[0083] receiving the alarm signal;
[0084] Based on the alarm signal, a resource reallocation decision is triggered.
[0085] Preferably, the data services provided by the data service unit include:
[0086] At least one of data access, data analysis, data retrieval, data processing and data mining; wherein the data mining adopts an algorithm combining supervised and unsupervised models.
[0087] Preferably, the multi-agent decision-making unit comprises:
[0088] Resource scheduling module, used to:
[0089] Determine the real-time load situation by analyzing load data;
[0090] Allocate appropriate cluster resources to workloads;
[0091] Realize workload migration and scheduling;
[0092] A reinforcement learning decision module is connected in communication with the resource scheduling module and is used to:
[0093] Through agent training, we can obtain the distribution representation of data value, workload, and resource conditions;
[0094] Allocate resources to different workloads;
[0095] Supports workload migration.
[0096] Preferably, the data value assessment unit is further used for:
[0097] Establish a data value assessment model, including:
[0098] Use historical usage records of data lake data as model input;
[0099] Use data value-related features as model input;
[0100] By constructing a loss function and training the model, a quantitative assessment of the value of the data is provided.
[0101] The intelligent governance method based on the distributed data lake intelligent governance and optimization system includes the following steps:
[0102] S1. Acquire and collect data resources, form data survey reports, and provide decision makers with an overview of data resources;
[0103] S2. According to the data investigation report, set data quality indicators for the governance of data resources, analyze the data usage pattern through the deep learning network, and form a data governance rule set in combination with the data quality indicators;
[0104] S3. Apply the data governance rule set to data governance, and regularly evaluate the governance results to optimize the data governance rule set;
[0105] S4. Collect description information of data resources and generate resource catalog;
[0106] S5. Collaborative analysis of cross-source data through a federated learning framework to protect data privacy;
[0107] S6. Develop data lifecycle management strategies for data resources and implement metadata tracking based on the data management strategies;
[0108] S7. Set the optimization goal of multi-agent decision-making and complete the decision-making through reinforcement learning to dynamically allocate storage resources;
[0109] S8. Obtain data value scores through data value assessment models and guide data archiving according to data archiving strategies.
[0110] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0111] The system architecture of the present invention fully considers all aspects of data lake management, and the modules work closely together to form an organic whole. The data governance module and the data optimization module jointly ensure the quality and effective use of data, while the intelligent analysis engine and the deep reinforcement learning decision system inject intelligence and adaptive capabilities into the entire system. This all-round design not only solves the fragmentation problem in the existing technology, but also achieves synergy between various functions.
[0112] From a macro perspective, the biggest advantage of the system of the present invention lies in its adaptability and intelligence. By introducing deep learning and reinforcement learning technologies, the system can continuously learn and optimize to adapt to the ever-changing data environment and business needs. This self-evolutionary ability greatly reduces the need for manual intervention and improves the scalability and long-term effectiveness of the system.
[0113] At the micro level, the present invention has achieved technological breakthroughs in multiple key points. For example, the data lake health assessment model adopts a multi-dimensional, dynamically adjustable assessment method, which can fully reflect the operating status of the data lake. The adaptive weight adjustment algorithm of the intelligent analysis engine ensures the continuous optimization of the analysis method and improves the accuracy and efficiency of data analysis. In addition, the deep reinforcement learning decision system based on the Actor-Critic architecture provides powerful decision support for resource scheduling and task management, and realizes the refined management of resource utilization.
[0114] The invention also cleverly resolves the contradiction between data privacy and collaborative analysis. By introducing the federated learning framework, the system can achieve collaborative analysis of cross-source data without directly sharing the original data, which not only protects data privacy but also maximizes the analytical value of the data. This innovation opens up new possibilities for the application of data lakes in sensitive fields.
[0115] In general, the distributed data lake intelligent governance and optimization system and method of the present invention effectively solves many problems in existing data lake management through its comprehensive, intelligent and adaptive design. It not only improves data quality and resource utilization efficiency, but also enhances data value assessment and life cycle management capabilities, while ensuring data security and privacy. The implementation of this system will significantly improve the organization's data management level, provide strong support for data-driven decision-making and innovation, and promote the entire industry to transform to a smarter and more efficient data management model. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 It is a logic block diagram of the whole system of the present invention.
[0117] Figure 2 This is a logical block diagram of the data governance module of the present invention.
[0118] Figure 3 It is a logical block diagram of the data optimization module of the present invention.
[0119] Figure 4 It is a logical block diagram of the multi-agent decision-making unit of the present invention.
[0120] Figure 5 It is a logic block diagram of the intelligent analysis engine of the present invention.
[0121] Figure 6 This is a logical block diagram of the deep reinforcement learning decision-making system of the present invention. DETAILED DESCRIPTION
[0122] See also Figure 1-6 The present invention provides a distributed data lake intelligent governance and optimization system and method thereof, aiming to solve the problems of uneven data quality, low resource utilization efficiency, and difficulty in evaluating data value in existing data lake management. The present invention will be described in detail below in conjunction with specific implementation methods.
[0123] like Figure 1 As shown, the distributed data lake intelligent governance and optimization system of the present invention includes a data governance module 1, a data optimization module 2, an intelligent analysis engine 3 and a deep reinforcement learning decision system 4. These modules work together to achieve intelligent governance and optimization of the distributed data lake.
[0124] The data governance module 1 is one of the core components of the system, which is mainly responsible for defining and implementing data governance rules, controlling data access rights, providing application-oriented interfaces and services, maintaining and managing the logical and physical structure of the data lake, and specifying and enforcing the data standards of the data lake. In a preferred embodiment of the present invention, the data governance module 1 adopts a rule-based governance strategy, combined with machine learning technology, and can adaptively adjust the governance rules to adapt to the ever-changing data environment.
[0125] The data optimization module 2 is in communication with the data governance module 1, and its main functions include evaluating the health of the data lake, providing metadata and data tracking capabilities, executing data management strategies, providing data lifecycle management functions, defining optimization goals and constraints, providing decision support for data management strategies and resource allocation, and obtaining data value scores through data value assessment models. Preferably, the data optimization module 2 adopts a multi-objective optimization algorithm, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II), to balance multiple optimization goals such as data quality, storage efficiency, and access performance.
[0126] The health status of the data lake is evaluated using the weight adjustment algorithm of the data lake health evaluation model: the initial weight setting is determined by the analytic hierarchy process (AHP). First, a hierarchical model is established to divide the evaluation indicators into six dimensions: storage, quality, efficiency, cost, access, and value. Then, a judgment matrix is constructed through expert scoring, and the eigenvector is calculated to obtain the initial weight.
[0127] The dynamic adjustment of weights adopts the exponential smoothing method. The specific steps are as follows: a) Set the initial weight to w i (0); b) For each time period t, calculate the actual health H(t) and the target health H target The gap; c) Update weights: Where α is the smoothing coefficient, usually 0.7-0.9, Represents the sign of the partial derivative of health with respect to the weight.
[0128] The intelligent analysis engine 3 is connected to the data governance module 1 and the data optimization module 2 in communication. Its core function is to select the corresponding intelligent analysis method according to the data type to analyze the data in the data lake, generate an analysis report and compare it with the data expectations, and dynamically adjust the weight of the analysis method according to the comparison results. In one embodiment of the present invention, the intelligent analysis engine 3 integrates a variety of machine learning and deep learning algorithms, such as decision trees, random forests, support vector machines, convolutional neural networks, etc., and can automatically select the most suitable analysis method according to data characteristics.
[0129] The deep reinforcement learning decision system 4 is connected to the data optimization module 2 and the intelligent analysis engine 3 in communication. By combining the deep learning network with reinforcement learning, the data usage pattern is analyzed, an optimization adjustment strategy is formulated, and the storage, network and computing resources and data processing tasks are optimized and adjusted through the data lake management platform. Preferably, the deep reinforcement learning decision system 4 adopts the DQN (Deep Q-Network) algorithm, combined with the priority experience playback mechanism, which can effectively process high-dimensional state space and improve learning efficiency.
[0130] The data governance module 1 includes a data quality unit 11, a data security unit 12, a data service unit 13, a data architecture unit 14, and a data standard unit 15. These units work together to ensure that the data in the data lake is of high quality, secure, reliable, and easy to access and use.
[0131] The data quality unit 11 is responsible for classifying the input data into high-quality and low-quality data, recording the quality of the input data, and automatically discovering and repairing data quality issues according to data governance rules. In one embodiment of the present invention, the data quality unit 11 uses a hybrid method based on rules and statistics to evaluate data quality. For example, for structured data, indicators such as data integrity, consistency, accuracy, and timeliness can be checked; for unstructured data, natural language processing technology can be used to evaluate text quality, or computer vision technology can be used to evaluate image quality.
[0132] The data security unit 12 is connected to the data quality unit 11 in communication, and mainly provides user authority management functions, verifies account authority and controls data access rights. Preferably, the data security unit 12 adopts a role-based access control (RBAC) model, combined with multi-factor authentication and encryption technology, to ensure the security and privacy of data.
[0133] The data service unit 13 is in communication with the data security unit 12, and is responsible for providing an external interface and providing specific data services according to the needs of the application. In one embodiment of the present invention, the data service unit 13 adopts a microservice architecture, provides a RESTful API, and supports multiple services such as data query, analysis, and visualization.
[0134] The data architecture unit 14 is in communication with the data service unit 13 and is mainly responsible for maintaining and managing the logical and physical structure of the data lake. Preferably, the data architecture unit 14 adopts a layered architecture, including an original data layer, a cleaned data layer, an aggregated data layer, and an application data layer to meet the data requirements of different scenarios.
[0135] The data standard unit 15 is in communication with the data architecture unit 14 and is responsible for specifying and executing the data standards of the data lake. In one embodiment of the present invention, the data standard unit 15 formulates metadata standards, data quality standards, and data exchange standards based on industry standards (such as ISO / IEC 11179) to ensure data consistency and interoperability.
[0136] The data optimization module 2 includes a data lake health unit 21, a data lifecycle management unit 22, a multi-agent decision unit 23, and a data value assessment unit 24. These units work together to achieve continuous optimization and value maximization of the data lake.
[0137] The data lake health unit 21 is responsible for evaluating the health of the data lake, calculating multiple key indicators, and calculating the health of the data lake according to the weight of each indicator. Specifically, the calculation formula for the health of the data lake is as follows:
[0138]
[0139] Among them, H is the health of the data lake, w i is the weight of the i-th indicator, I i is the value of the ith indicator, and n is the total number of indicators. In a preferred embodiment of the present invention, the following six key indicators are considered: storage capacity proportion indicator value I1, data quality proportion indicator value I2, resource utilization proportion indicator value I3, storage cost proportion indicator value I4, data access volume proportion indicator value I5 and data value assessment proportion indicator value I6.
[0140] The calculation method of each indicator is as follows:
[0141] 1. Storage capacity ratio index value I1:
[0142]
[0143] 2. Data quality ratio index value I2:
[0144]
[0145] 3. Resource utilization ratio index value I3:
[0146]
[0147] 4. Storage cost ratio index value I4:
[0148]
[0149] 5. Data access volume ratio index value I5:
[0150]
[0151] 5. Data value assessment ratio index value I6:
[0152]
[0153] Weight w i The choice of can be adjusted according to specific business needs and data lake characteristics. For example, in a scenario focusing on data quality, w2 can be given a higher weight; while in a scenario focusing on cost control, the weight of w4 can be increased.
[0154] The data lifecycle management unit 22 is in communication with the data lake health unit 21, and is responsible for providing metadata and data tracking capabilities, executing data management policies, and providing data lifecycle management functions. In one embodiment of the present invention, the data lifecycle management unit 22 adopts a state machine-based management method to divide the data lifecycle into four stages: creation, use, archiving, and deletion, and applies corresponding management policies in each stage.
[0155] The multi-agent decision unit 23 is in communication with the data lifecycle management unit 22, and is mainly responsible for defining optimization goals and constraints, and providing decision support for data management strategies and resource allocation through analysis and learning. Preferably, the multi-agent decision unit 23 adopts a multi-agent system based on a market mechanism, where each agent represents a resource or task, and achieves optimal resource allocation through bidding and negotiation.
[0156] The data value evaluation unit 24 is in communication with the multi-agent decision-making unit 23, and obtains the data value score through the data value evaluation model. In one embodiment of the present invention, the data value evaluation adopts a multi-dimensional evaluation method, taking into account factors such as the timeliness, completeness, accuracy, uniqueness and relevance of the data. The calculation formula for the data value score is as follows:
[0157]
[0158] Among them, V is the data value score, v j is the weight of the jth factor, F j is the score of the jth factor, and m is the total number of evaluation factors.
[0159] Through the collaborative work of the above modules and units, the distributed data lake intelligent governance and optimization system of the present invention can achieve continuous improvement of data quality, optimization of resource utilization efficiency, and maximization of data value, thereby providing enterprises and organizations with high-quality and high-value data assets.
[0160] The intelligent analysis engine 3 of the present invention includes a data classification unit 31, an analysis method selection unit 32, a data analysis unit 33, a result comparison unit 34 and a weight adjustment unit 35. These units work together to achieve intelligent analysis and continuous optimization of data in the data lake.
[0161] The data classification unit 31 is responsible for classifying the data according to the data type. Preferably, the present invention adopts a multi-level data classification method, firstly classifying the data into structured data, semi-structured data and unstructured data according to the data structure, and then further subdividing the data according to the data content. For example, for unstructured data, it can be further divided into text data, image data, audio data and video data, etc. This detailed classification method is helpful for selecting the most suitable analysis method later.
[0162] The analysis method selection unit 32 is connected to the data classification unit 31 for selecting a corresponding intelligent analysis method according to the data type. The system of the present invention integrates a variety of advanced analysis methods, including but not limited to statistical analysis, machine learning, deep learning, and natural language processing. For example, for structured data, regression analysis, cluster analysis, or decision tree methods can be selected; for text data, topic modeling, sentiment analysis, or named entity recognition methods can be selected; for image data, convolutional neural networks or target detection algorithms can be selected.
[0163] The data analysis unit 33 is in communication with the analysis method selection unit 32 and is responsible for analyzing the data using the selected analysis method. In one embodiment of the present invention, the data analysis unit 33 uses a distributed computing framework, such as ApacheSpark or Flink, to improve the efficiency of large-scale data analysis. In addition, the system of the present invention also implements an automated parameter tuning function, which automatically selects the optimal model parameters through algorithms such as Bayesian optimization to further improve the analysis effect.
[0164] The result comparison unit 34 is connected to the data analysis unit 33 for comparing the analysis result with the data expectation. The present invention introduces a novel multi-dimensional comparison method, which not only compares the accuracy of the analysis results, but also considers factors such as time efficiency, resource consumption and interpretability. The comparison result is represented by a comprehensive score, and the calculation formula is as follows:
[0165] S=α·A+β·E+γ·R+δ·I,
[0166] Among them, S is the comprehensive score, A is the accuracy score, E is the efficiency score, R is the resource consumption score, and I is the interpretability score. α, β, γ and 8 are the corresponding weight coefficients, and α+β+γ+δ=1 is satisfied.
[0167] The weight adjustment unit 35 is connected to the result comparison unit 34 in communication, and dynamically adjusts the weight of the analysis method according to the comparison result. The present invention adopts an adaptive weight adjustment algorithm, based on the idea of reinforcement learning, which regards each analysis method as an action and the comparison result as a reward signal. Through continuous trial and learning, the selection probability of various analysis methods is optimized.
[0168] The adaptive weight adjustment algorithm adopts Q-learning algorithm to realize adaptive weight adjustment. It is defined as follows:
[0169] State space S: the weighted combined action space of each current analysis method;
[0170] A: Increase, decrease or maintain the weight reward of a certain analysis method;
[0171] R: the improvement of the comprehensive score S;
[0172] Q value update formula:
[0173]
[0174] Among them, α is the learning rate and γ is the discount factor. Specific steps: Initialize the Q table; for each round: observe the current state s; use the ε-greedy strategy to select action a; perform the action, observe the reward r and the new state s ′ ; Update Q value s←s ′ ; c) repeat until convergence;
[0175] The deep reinforcement learning decision system 4 of the present invention includes a deep learning network unit 41, a reinforcement learning unit 42, a decision execution unit 43, a resource optimization unit 44 and a task scheduling unit 45. These units together constitute an advanced decision system that can adaptively optimize resource allocation and task scheduling of the data lake.
[0176] The deep learning network unit 41 uses a multi-layer neural network to learn the data usage pattern. Preferably, the present invention uses an improved long short-term memory network (LSTM) structure, which can effectively capture the long-term dependencies and periodic patterns of data usage. The input of the network includes historical data access records, resource usage, and task execution logs, and the output is a high-dimensional representation of the data usage pattern. In this way, the system can automatically discover the potential laws of data usage and provide an important basis for subsequent optimization decisions.
[0177] The reinforcement learning unit 42 is connected to the deep learning network unit 41 for communication, and an optimization adjustment strategy is formulated according to the learned data usage pattern. The present invention adopts a policy gradient-based reinforcement learning algorithm, and defines the state space of the data lake as a high-dimensional vector, which includes multiple indicators such as storage utilization, network bandwidth, and computing resource utilization. The action space includes operations such as resource allocation, task scheduling, and data migration. The reward function is designed as the weighted sum of the data lake health indicators, which encourages the system to improve resource utilization efficiency while ensuring performance.
[0178] The decision execution unit 43 is in communication with the reinforcement learning unit 42 and is responsible for delivering the optimization adjustment strategy to the data lake management platform. The decision execution unit 43 of the present invention adopts a progressive strategy implementation method, which gradually adjusts the system configuration in small steps to avoid system instability caused by large adjustments.
[0179] The resource optimization unit 44 is connected to the decision execution unit 43 in communication, and optimizes and adjusts the storage, network and computing resources according to the optimization adjustment strategy. The resource optimization unit 44 of the present invention adopts a multi-objective optimization algorithm to minimize resource waste and energy consumption while meeting performance requirements. For example, for storage resources, the system will automatically migrate data between different storage layers according to the frequency and importance of data access; for network resources, the system will dynamically adjust the network topology and bandwidth allocation according to the data traffic pattern.
[0180] The task scheduling unit 45 is connected to the resource optimization unit 44 in communication, and optimizes and schedules the data processing tasks according to the optimization adjustment strategy. The task scheduling unit 45 of the present invention adopts an intelligent scheduling algorithm based on priority and dependency, which can automatically identify the dependency between tasks and dynamically adjust the task execution order and resource allocation according to the urgency of the task and resource requirements.
[0181] According to claim 6, the data lake health unit 21 of the present invention is also used to send an alarm signal when the health of the data lake does not meet the standard. Preferably, the present invention sets a multi-level health threshold, for example, the health can be divided into five levels: excellent (90-100 points), good (80-89 points), general (70-79 points), poor (60-69 points) and severe (below 60 points). When the health drops to the "general" level, the system will send a warning signal; when it drops to the "poor" level, the system will send a serious warning signal; when it drops to the "severe" level, the system will trigger an emergency response mechanism.
[0182] After receiving the alarm signal, the multi-agent decision unit 23 will immediately trigger the resource reallocation decision. The present invention adopts a multi-agent negotiation algorithm based on an auction mechanism. Each agent represents different resources or tasks, and a new resource allocation plan is quickly reached through bidding. This method can improve the efficiency and flexibility of decision-making while ensuring the global optimum.
[0183] The data service unit 13 of the present invention provides data services including at least data access, data analysis, data retrieval, data processing and data mining. Preferably, the present invention adopts a microservice architecture to modularize and containerize these services to improve the scalability and maintainability of the system.
[0184] The data mining service uses an algorithm that combines supervised and unsupervised models. For example, for customer classification problems, the system can first use unsupervised clustering algorithms (such as K-means or DBSCAN) to discover potential customer groups, and then use supervised classification algorithms (such as random forests or support vector machines) to accurately classify new customers. This combination can make full use of existing labeled data, while discovering potential patterns in the data, improving the accuracy and interpretability of mining results.
[0185] Through the collaborative work of the above modules and units, the distributed data lake intelligent governance and optimization system of the present invention can achieve efficient management, intelligent analysis and continuous optimization of data, and provide powerful data support and decision-making assistance for enterprises and organizations. According to claim 8, the multi-agent decision unit 23 of the present invention includes a resource scheduling module 231 and a reinforcement learning decision module 232. These two modules work together to achieve intelligent scheduling and optimization decisions of data lake resources.
[0186] The resource scheduling module 231 is responsible for determining the real-time load situation by analyzing the load data, allocating appropriate cluster resources to the workload, and realizing the migration and scheduling of the workload. Preferably, the present invention adopts a load prediction algorithm based on time series analysis, combining historical load data and current system status to accurately predict the load changes in the future. The prediction model adopts a hybrid structure of the ARIMA (autoregressive integrated moving average) model and the LSTM (long short-term memory) neural network, which can capture the periodic changes and long-term trends of the load at the same time.
[0187] The core algorithm of the resource scheduling module 231 can be expressed as follows:
[0188] L t =f(H t ,S t ,E t ),
[0189] Among them, L t represents the predicted load at time t, H t Represents historical load data, S t Indicates the current system status, E t Represents external factors (such as holidays, marketing activities, etc.). Function f is defined by the ARIMA-LSTM hybrid model.
[0190] The ARIMA-LSTM hybrid model structure is as follows: a) ARIMA model is used to capture linear trends and seasonality; b) LSTM is used to model nonlinearities and long-term dependencies;
[0191] The specific steps are as follows: 1. Use the ARIMA model to fit the original time series and get the predicted value Yarima ; 2. Calculate the residual: e t =Y t -Y arima,t ; 3. Use LSTM to model the residual sequence LSTM structure: input layer → LSTM layer (64 units) → LSTM layer (32 units) → fully connected layer → output layer; final prediction value = Y arima + Residuals predicted by LSTM; Training method: Maximum likelihood estimation is used for the ARIMA part; Stochastic gradient descent is used for the LSTM part, and the loss function is mean square error.
[0192] Based on the predicted load, the resource scheduling module 231 uses an improved Best Fit algorithm to allocate resources. The algorithm not only considers the utilization of resources, but also introduces a load balancing factor to avoid the situation where some nodes are overloaded while other nodes are idle.
[0193] The reinforcement learning decision module 232 is in communication with the resource scheduling module 231. Through the training of the intelligent agent, the distribution representation of the data value, workload and resource situation is obtained, resources are allocated to different workloads, and the migration of workloads is supported. The present invention adopts a deep reinforcement learning algorithm based on the Actor-Critic architecture, which can learn complex decision strategies in a high-dimensional state space.
[0194] The Actor network is responsible for generating actions (i.e., resource allocation strategies) based on the current state. Its structure is a multi-layer perceptron, with the system state vector as input and the resource allocation probability distribution as output. The Critic network is responsible for evaluating the value of the current state-action pair. Its structure is also a multi-layer perceptron, with the state-action pair as input and the estimated value as output. The two networks are continuously optimized through alternating training and eventually converge to a near-optimal decision strategy.
[0195] The policy gradient algorithm in the deep reinforcement learning decision system adopts the Actor-Critic architecture, as follows:
[0196] Actor network structure: input layer → fully connected layer (256 neurons, ReLU activation) → fully connected layer (128 neurons, ReLU activation) → output layer (action dimension, Softmax activation);
[0197] Critic network structure: input layer → fully connected layer (256 neurons, ReLU activation) → fully connected layer (128 neurons, ReLU activation) → output layer (1 neuron, linear activation);
[0198] Policy gradient update formula:
[0199]
[0200] Where A(s,a) is the advantage function, estimated by the Critic network.
[0201] The value network is updated using mean square error loss:
[0202] L = (r + γV(s′) - V(s)) 2 .
[0203] The data value assessment unit 24 of the present invention is responsible for establishing a data value assessment model, taking the historical usage records of the data lake data and data value-related features as model inputs, and providing a quantitative assessment of the value of the data by constructing a loss function and training the model.
[0204] Preferably, the present invention adopts a multi-factor based data value assessment model. The model takes into account multiple factors that affect the value of data, including but not limited to: data freshness, usage frequency, business relevance, data quality, data scarcity, etc. Each factor has its corresponding scoring function, and the final data value score is the weighted sum of these factor scores.
[0205] The data value assessment model can be expressed as:
[0206]
[0207] Among them, V represents the data value score, w i represents the weight of the i-th factor, f i represents the scoring function of the i-th factor, x i represents the features associated with the i-th factor.
[0208] To train this model, the present invention designs a semi-supervised learning method based on historical usage records. First, a small amount of manually annotated high-value data is used as a seed, and then other potential high-value data is automatically discovered by analyzing the usage patterns of these data. This method greatly reduces the workload of manual annotation while ensuring the accuracy and interpretability of the model.
[0209] The semi-supervised learning method adopts a graph-based semi-supervised learning algorithm. The specific steps are as follows:
[0210] a) Build a data relationship diagram:
[0211] Nodes represent data items, and edges represent the similarity between data items. Similarity calculation:
[0212] b) Label propagation algorithm:
[0213] Initialization: The label of the labeled data is y i, the label of unlabeled data is 0 iteration propagation:
[0214] where w ij is the edge weight
[0215] c) Model training:
[0216] Use the propagated labels as weak supervision signals to train the data value assessment model
[0217] Loss function: L = L ce +λL consistency ;
[0218] Where L ce is the cross entropy loss, L consistency is the consistency regularization term;
[0219] d) Iterative optimization: Use the trained model to predict the unlabeled data, select the prediction results with high confidence and add them to the labeled data set, and repeat steps bd until the model converges or reaches the preset number of iterations.
[0220] The loss function of the model is designed as a combination of cross entropy loss and regularization term:
[0221]
[0222] Among them, y j represents the true label of the jth sample, Represents the data value score predicted by the model, and λ is the regularization coefficient. By minimizing this loss function, the model can strike a balance between fitting the training data and avoiding overfitting.
[0223] The present invention also provides a distributed data lake intelligent governance and optimization method based on the above system. The method includes the following steps:
[0224] S1. Acquire and collect data resources, form data survey reports, and provide decision makers with an overview of data resources;
[0225] S2. According to the data investigation report, set data quality indicators for the governance of data resources, analyze the data usage pattern through deep learning network, and form a data governance rule set in combination with data quality indicators;
[0226] S3. Apply the data governance rule set to data governance, regularly evaluate the governance results, and optimize the data governance rule set;
[0227] S4. Collect description information of data resources and generate resource catalog;
[0228] S5. Collaborative analysis of cross-source data through a federated learning framework to protect data privacy;
[0229] S6. Develop data lifecycle management strategies for data resources and implement metadata tracking based on the data management strategies;
[0230] S7. Set the optimization goal of multi-agent decision-making and complete the decision-making through reinforcement learning to dynamically allocate storage resources;
[0231] S8. Obtain data value scores through data value assessment models and guide data archiving according to data archiving strategies.
[0232] This approach achieves comprehensive governance and optimization of the data lake through systematic and intelligent means, which can significantly improve the quality, availability and value of data, and provide strong data support and decision-making assistance for enterprises and organizations.
[0233] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. Distributed data lake intelligent governance and optimization system, characterized by: include: Data governance module for: Define and implement data governance rules; Control access to data; Provide application-oriented interfaces and services; Maintain and manage the logical and physical structure of the data lake; Specify and enforce data standards for the data lake; A data optimization module, in communication with the data governance module, is used to: Assess the health of your data lake; Provide metadata and data tracking capabilities; Execute data management strategies and provide data lifecycle management functions; Define optimization goals and constraints to provide decision support for data management strategies and resource allocation; Obtain data value scores through data value assessment models; An intelligent analysis engine, in communication with the data governance module and the data optimization module, is used to: Select the corresponding intelligent analysis method according to the data type to analyze the data in the data lake; Generate analysis reports and compare with data expectations; Dynamically adjust the weight of the analysis method based on the comparison results; A deep reinforcement learning decision system is connected to the data optimization module and the intelligent analysis engine for: Analyze data usage patterns by combining deep learning networks with reinforcement learning; Formulate optimization and adjustment strategies; Optimize and adjust storage, network and computing resources, and data processing tasks through the data lake management platform.
2. The system according to claim 1, characterized in that The data governance module includes: Data Quality Unit for: Classify input data into high-quality and low-quality data; Document the quality of input data; Automatically discover and fix data quality issues based on data governance rules; A data security unit is communicatively connected to the data quality unit and is used to: Provide user rights management function; Verify account permissions and control data access permissions; A data service unit, in communication with the data security unit, is used to: Provide external interface; Provide specific data services based on application needs; A data architecture unit, in communication with the data service unit, is used to: Maintain and manage the logical and physical structure of the data lake; A data standard unit is communicatively connected with the data architecture unit and is used for: Specify and enforce data standards for your data lake.
3. The system according to claim 1, characterized in that The data optimization module includes: Data Lake Health Unit, used to: Assess the health of your data lake; Calculate the storage capacity ratio index value, data quality ratio index value, resource utilization ratio index value, storage cost ratio index value, data access volume ratio index value and data value assessment ratio index value; Calculate the health of the data lake based on the weight of each indicator; A data lifecycle management unit is communicatively connected to the data lake health unit and is used to: Provide metadata and data tracking capabilities; Implement data management strategies; Provide data lifecycle management functions; The multi-agent decision-making unit is communicatively connected with the data lifecycle management unit and is used to: Define optimization objectives and constraints; Provide decision support for data management strategies and resource allocation through analysis and learning; A data value assessment unit, which is in communication with the multi-agent decision-making unit, is used to: Obtain data value scores through the data value assessment model.
4. The system according to claim 1, characterized in that The intelligent analysis engine comprises: Data classification unit for: Categorize data according to data type; An analysis method selection unit, in communication with the data classification unit, is used to: Select the corresponding intelligent analysis method according to the data type; A data analysis unit is communicatively connected to the analysis method selection unit, and is used to: Analyze the data using the selected analytical methods; A result comparison unit is connected in communication with the data analysis unit and is used for: Compare analysis results with data expectations; A weight adjustment unit is connected to the result comparison unit for: Dynamically adjust the weight of the analysis method based on the comparison results.
5. The system according to claim 1, characterized in that The deep reinforcement learning decision-making system comprises: Deep learning network units for: Learning data usage patterns through multi-layer neural networks; A reinforcement learning unit, communicatively connected to the deep learning network unit, is used to: Develop optimization and adjustment strategies based on the learned data usage patterns; A decision execution unit, which is in communication with the reinforcement learning unit, is used to: Deliver optimization and adjustment strategies to the data lake management platform; A resource optimization unit, in communication with the decision execution unit, is used to: Optimize storage, network and computing resources according to optimization adjustment strategies; A task scheduling unit, in communication with the resource optimization unit, is used to: Optimize and schedule data processing tasks based on optimization adjustment strategies.
6. The system according to claim 3, characterized in that The data lake health unit is also used to: Send an alert signal when the health of the data lake does not meet the standard; The multi-agent decision-making unit is also used for: receiving the alarm signal; Based on the alarm signal, a resource reallocation decision is triggered.
7. The system according to claim 2, characterized in that The data services provided by the data service unit include: at least one of data access, data analysis, data retrieval, data processing, and data mining; Wherein, the data mining adopts an algorithm combining supervised and unsupervised models.
8. The system according to claim 3, characterized in that The multi-agent decision-making unit comprises: Resource scheduling module, used to: Determine the real-time load situation by analyzing load data; Allocate appropriate cluster resources to workloads; Implement workload migration and scheduling; A reinforcement learning decision module is connected in communication with the resource scheduling module and is used to: Through agent training, we can obtain the distribution representation of data value, workload, and resource conditions; Allocate resources to different workloads; Supports workload migration.
9. The system according to claim 3, characterized in that The data value assessment unit is also used for: Establish a data value assessment model, including: Use historical usage records of data lake data as model input; Use data value-related features as model input; By constructing a loss function and training the model, a quantitative assessment of the value of the data is provided.
10. The intelligent governance method of the distributed data lake intelligent governance and optimization system according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Acquire and collect data resources, form data survey reports, and provide decision makers with an overview of data resources; S2. According to the data investigation report, set data quality indicators for the governance of data resources, analyze the data usage pattern through the deep learning network, and form a data governance rule set in combination with the data quality indicators; S3. Apply the data governance rule set to data governance, and regularly evaluate the governance results to optimize the data governance rule set; S4. Collect description information of data resources and generate resource catalog; S5. Collaborative analysis of cross-source data through a federated learning framework to protect data privacy; S6. Develop data lifecycle management strategies for data resources and implement metadata tracking based on the data management strategies; S7. Set the optimization goal of multi-agent decision-making and complete the decision-making through reinforcement learning to dynamically allocate storage resources; S8. Obtain data value scores through data value assessment models and guide data archiving according to data archiving strategies.
Citation Information
Cited By
Automatic data governance strategy generation method and system based on reinforcement learning
CN120492828A
Automatic data governance strategy generation method and system based on reinforcement learning
CN120492828B
Data quality evaluation method and system based on rule engine and machine learning
CN121030265A