Adaptive data management method and device, computer equipment, readable storage medium and program product
Through adaptive data governance methods, multi-source data sets are collected and analyzed, data quality problems are identified and repaired, and data governance strategies are dynamically adjusted, which solves the problem of insufficient flexibility in the data governance platform in the existing technology, and the ability to effectively respond to business needs and data quality problems is achieved.
Patent Information
- Application Number
- CN202510089212.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
Smart Images

Figure CN120011353A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data management technology, and in particular to an adaptive data governance method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of information technology and the advent of the big data era, all kinds of enterprises generate massive amounts of data every day. These data cover all aspects of business operations, including customer information, financial data, market dynamics, and production data. If these data cannot be managed effectively, they will not only fail to provide beneficial value to the enterprise, but may also bring many risks, especially in terms of data quality and data security.
[0003] In traditional technologies, data integration and data governance are mainly performed through data governance platforms such as Collibra and Talend. However, existing data governance platforms can usually only perform data governance for a specific scenario or a certain type of data, and are slow to respond to complex and changing business needs in daily business. This has the problem of insufficient flexibility and is difficult to deal with, for example, emerging business needs or sudden data quality issues. Summary of the invention
[0004] Based on this, it is necessary to provide an adaptive data governance method, apparatus, computer equipment, computer-readable storage medium and computer program product to address the above-mentioned technical problems.
[0005] In a first aspect, the present application provides an adaptive data governance method, comprising:
[0006] Collecting a multi-source data set in a current business scenario, analyzing current data features of the multi-source data set, and identifying data quality issues in the multi-source data set based on the current data features;
[0007] Determine a corresponding repair solution according to the type of the data quality problem, and repair the multi-source data set based on the repair solution to obtain a repaired target data set;
[0008] Analyze the data change information of the target data set, and if the data change information meets the policy adjustment condition, adjust the data governance policy under the current business scenario to obtain an adjusted target data governance policy;
[0009] Repair feedback is performed on the target data set through a preset feedback mechanism, and adjustment feedback is performed on the target data governance strategy.
[0010] In one embodiment, determining a corresponding repair solution according to the type of the data quality problem includes:
[0011] When the data formats in the multi-source data set are inconsistent, the data in the multi-source data set are uniformly standardized; when there is missing data in the multi-source data set, the location of the missing value is determined and the missing value is filled; when there is data anomaly in the multi-source data set, the location of the outlier value is determined and the outlier value is corrected.
[0012] In one embodiment, the method further comprises:
[0013] Acquire historical data records, perform feature selection and feature extraction on the historical data records to obtain historical data features; predict future data usage trends through a prediction model based on the historical data features and the data change information; and generate data usage recommendations based on the future data usage trends.
[0014] In one embodiment, after generating the data usage suggestion according to the future data usage trend, the method further includes:
[0015] According to the data usage suggestion, the data storage strategy is adjusted to obtain the adjusted target data storage strategy; according to the access frequency of each data in the multi-source data set, the high-frequency access data is screened out, and based on the target data storage strategy, the high-frequency access data is saved to the high-performance storage.
[0016] In one of the embodiments, feedback data of the prediction model is obtained, and multiple groups of candidate model parameter combinations are generated through grid search based on the feedback data; a target model parameter combination is determined from the multiple groups of candidate model parameter combinations based on the prediction accuracy after model training of each group of the candidate model parameter combinations; the prediction model is updated based on the target model parameter combination to obtain an updated prediction model; and the updated prediction model is used to replace the prediction model.
[0017] In one embodiment, the method further comprises:
[0018] Acquire historical user behavior information, and calculate the predicted probability of potential abnormal user behavior through a behavior analysis model based on the historical user behavior information; identify target abnormal user behavior from the potential abnormal user behavior based on the predicted probability and the warning threshold, and generate warning information for the target abnormal user behavior; and generate security protection information for the target abnormal user behavior in response to a user's request to process the warning information.
[0019] In a second aspect, the present application also provides an adaptive data governance device, including:
[0020] A data acquisition module is used to collect a multi-source data set in a current business scenario, analyze the current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features;
[0021] A data repair module, used to determine a corresponding repair solution according to the type of the data quality problem, and repair the multi-source data set based on the repair solution to obtain a repaired target data set;
[0022] A policy adjustment module is used to analyze the data change information of the target data set, and if the data change information meets the policy adjustment conditions, adjust the data governance policy under the current business scenario to obtain an adjusted target data governance policy;
[0023] The information feedback module is used to provide repair feedback to the target data set through a preset feedback mechanism, and to provide adjustment feedback to the target data governance strategy.
[0024] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0025] Collect a multi-source data set in the current business scenario, analyze the current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features; determine the corresponding repair plan based on the type of the data quality problem, and repair the multi-source data set based on the repair plan to obtain a repaired target data set; analyze the data change information of the target data set, and adjust the data governance policy in the current business scenario if the data change information meets the policy adjustment conditions to obtain an adjusted target data governance policy; provide repair feedback for the target data set and adjustment feedback for the target data governance policy through a preset feedback mechanism.
[0026] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0027] Collect a multi-source data set in the current business scenario, analyze the current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features; determine the corresponding repair plan based on the type of the data quality problem, and repair the multi-source data set based on the repair plan to obtain a repaired target data set; analyze the data change information of the target data set, and adjust the data governance policy in the current business scenario if the data change information meets the policy adjustment conditions to obtain an adjusted target data governance policy; provide repair feedback for the target data set and adjustment feedback for the target data governance policy through a preset feedback mechanism.
[0028] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0029] Collect a multi-source data set in the current business scenario, analyze the current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features; determine the corresponding repair plan based on the type of the data quality problem, and repair the multi-source data set based on the repair plan to obtain a repaired target data set; analyze the data change information of the target data set, and adjust the data governance policy in the current business scenario if the data change information meets the policy adjustment conditions to obtain an adjusted target data governance policy; provide repair feedback for the target data set and adjustment feedback for the target data governance policy through a preset feedback mechanism.
[0030] The above-mentioned adaptive data governance method, apparatus, computer equipment, computer-readable storage medium and computer program product start from selecting the current business scenario according to business needs, and go through multiple links such as data collection, data quality problem repair, dynamic adjustment of data governance strategy, repair feedback and adjustment feedback under the current business scenario. It can quickly respond to the complex and changeable business needs in the daily business of the enterprise, effectively avoid the disconnection between data governance and business needs, and realize intelligent adaptive data governance functions, thereby improving the flexibility and automation of data governance, so that it can flexibly respond to emerging business needs or sudden data quality problems, and adjust data governance strategies in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 A diagram of an application environment of an adaptive data governance method in an embodiment;
[0033] Figure 2 A schematic diagram of a process of an adaptive data governance method in one embodiment;
[0034] Figure 3 A schematic diagram of a process flow of a trend prediction step in an embodiment;
[0035] Figure 4 A schematic diagram of a flow chart of a model updating step in one embodiment;
[0036] Figure 5 A schematic diagram of a process of abnormal warning steps in an embodiment;
[0037] Figure 6 is a structural block diagram of an adaptive data management device in one embodiment;
[0038] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] Existing data governance platforms are usually optimized for a specific scenario or a certain type of data, and are slow to respond to the complex and ever-changing business needs of daily business. Although some platforms have integrated intelligent functions, most still rely on rule-driven and manual intervention. Existing intelligent technologies are mostly based on predefined rule sets and have weak adaptability to complex scenarios. Especially in links that require flexible judgment, such as data anomaly detection and data cleaning, intelligent systems often lack sufficient processing capabilities.
[0041] Enterprise data is usually distributed in different systems and platforms, such as CRM systems, ERP systems and data warehouses. When dealing with cross-platform data integration, the existing governance system may face data island problems and lack efficient data fusion and interoperability. This makes data sharing and coordination between different departments complicated, resulting in reduced efficiency of data governance.
[0042] Although existing technologies can provide certain governance in the data generation and storage stages, they lack adequate monitoring and management mechanisms in the later stages of the data life cycle, such as data use, sharing, and destruction. Data governance is often limited to the data storage and cleaning stages, ignoring the long-term monitoring and optimization of data.
[0043] In this regard, the adaptive data governance of this application effectively solves the pain points of existing platforms through four key advantages. First, flexibility is improved by dynamically adjusting data strategies to respond to changes in business needs. Second, intelligence and automation are enhanced using artificial intelligence and big data technologies to automatically detect and repair data quality issues. Third, cross-platform integration capabilities enable data sharing and collaboration between different systems, breaking down data silos and improving efficiency. Finally, data governance runs through the entire life cycle, and the entire process from data storage, use to destruction can be intelligently monitored and optimized to ensure continuous adaptation to business needs.
[0044] The adaptive data governance method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown in FIG. 1 , the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be integrated on the server or placed on the cloud or other network servers. Figure 1 In the application environment shown, the terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, and tablet computers. The server can be implemented as an independent server or a server cluster consisting of multiple servers.
[0045] In one embodiment, Figure 2 As shown, an adaptive data governance method is provided, which is applied to Figure 1 The terminal in is used as an example to illustrate, including the following steps:
[0046] Step S201 , collect a multi-source data set in a current business scenario, analyze current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features.
[0047] Among them, the current business scenario can be adjusted by the user according to actual needs, such as urban traffic management.
[0048] Among them, multi-source datasets obtain data from various data sources (such as sensors, external APIs, log files, and databases) through APIs (Application Programming Interfaces), database connections, and IoT devices. These data can be structured data (such as database tables) or unstructured data (such as text and images).
[0049] Specifically, the terminal can use data stream processing frameworks (such as Apache Kafka and Apache Flink) to monitor the status and changes of various data sources in real time. The system can perform real-time monitoring by setting thresholds (such as data volume and data quality). When an abnormal data source access or a decline in data quality is found, the system will automatically alarm. The characteristics of the data source are analyzed through machine learning algorithms to automatically identify potential data quality issues (such as missing values, outliers, and duplicate data). Based on the historical data training model, the system can predict the change trend of certain data sources and warn of possible quality problems in advance.
[0050] Step S202, determining a corresponding repair solution according to the type of data quality problem, and repairing the multi-source data set based on the repair solution to obtain a repaired target data set.
[0051] Data quality issues include, but are not limited to, missing values, format errors, and inconsistent data.
[0052] Specifically, the terminal uses artificial intelligence and big data analysis technology to automatically identify data quality problems in newly connected data sources, determine the corresponding repair plan according to the type of data quality problem, and repair the multi-source data set based on the repair plan to obtain the repaired target data set, ensuring that data quality and security are improved simultaneously.
[0053] Step S203: Analyze the data change information of the target data set, and when the data change information meets the policy adjustment conditions, adjust the data governance policy in the current business scenario to obtain the adjusted target data governance policy.
[0054] Among them, data change information can be new business requirements and regulatory changes, etc.
[0055] Specifically, the terminal can automatically adjust the data governance strategy by analyzing the data change information of the target data set. For example, according to new requirements, it may be necessary to adjust the data storage format and field requirements; as data quality issues change, the monitoring rules will be adjusted according to the new data issues (for example, adding new data anomaly detection indicators). Based on historical data and business change trends, the terminal recommends the best data governance strategy to users through AI algorithms (such as decision trees and reinforcement learning). These strategies include data collection rules, data cleaning methods, data storage strategies, etc. The terminal will also regularly evaluate the effectiveness of the implemented governance strategies and provide a feedback mechanism to automatically backtrack and adjust the strategies. This can adjust the existing strategies by comparing indicators such as data quality, compliance, and performance before and after governance. In addition, the terminal can also regularly evaluate the effectiveness of the implemented strategies, compare the data quality, compliance, and performance indicators before and after governance, and backtrack and adjust the strategies when necessary.
[0056] Step S204: providing repair feedback to the target data set and adjusting feedback to the target data governance strategy through a preset feedback mechanism.
[0057] Specifically, the repaired target data set will be repaired and fed back through the preset feedback mechanism, and the adjusted target data governance strategy will also be adjusted and fed back through the preset feedback mechanism to ensure that data quality is continuously optimized throughout the entire process.
[0058] In the above-mentioned adaptive data governance method, it starts with selecting the current business scenario according to business needs, and goes through multiple links such as data collection, data quality problem repair, dynamic adjustment of data governance strategy, and repair feedback and adjustment feedback under the current business scenario. It can quickly respond to the complex and changeable business needs in the daily business of the enterprise, effectively avoid the disconnection between data governance and business needs, and realize intelligent adaptive data governance functions, thereby improving the flexibility and automation of data governance, so that it can flexibly respond to emerging business needs or sudden data quality problems, and adjust data governance strategies in real time.
[0059] In one embodiment, in the above step S202, a corresponding repair solution is determined according to the type of data quality problem, which specifically includes the following steps:
[0060] When the data formats in multi-source data sets are inconsistent, the data in the multi-source data sets are uniformly standardized; when there is missing data in the multi-source data sets, the location of the missing values is determined and the missing values are filled; when there are data anomalies in the multi-source data sets, the location of the outliers is determined and the outliers are corrected.
[0061] Specifically, the terminal automatically identifies quality issues in the data, such as missing values, format errors, and inconsistent data. Common repair methods include:
[0062] Missing value handling: fill missing values (such as mean filling and most frequent value filling, etc.), or make inferences based on data patterns.
[0063] Outlier detection: Detect outliers through statistical methods (such as Z-score and IQR) or machine learning methods (such as isolation forest and DBSCAN), and correct or remove them.
[0064] Data standardization and formatting: Standardize data from different data sources to ensure consistency in data format. For example, standardize dates and unify data units.
[0065] Automated repair: Data repair is performed through an automated process, where the system automatically processes the data according to preset repair rules (such as business rules, industry standards, etc.). If the repair rules fail to solve the problem, the system triggers a manual review process.
[0066] In one embodiment, if Figure 3 As shown, the method of the present application also includes the following steps:
[0067] Step S301, obtain historical data records, perform feature selection and feature extraction on the historical data records, and obtain historical data features.
[0068] Step S302: predicting future data usage trends through a prediction model based on historical data characteristics and data change information.
[0069] Step S303: Generate data usage suggestions based on future data usage trends.
[0070] Specifically, the terminal uses machine learning and time series analysis (such as ARIMA and LSTM) to predict future data usage trends. For example, it analyzes historical data access frequency and historical data storage requirements to predict the data lifecycle. Based on the prediction results, it generates data lifecycle optimization suggestions. For example, which data should be stored first, which should be archived or deleted, etc. The decision support system provides users with suggestions to ensure effective use of data.
[0071] For example, the terminal extracts historical data usage records from the system, including but not limited to data access frequency, query mode, and storage requirements.
[0072] Feature selection: Identify features that are helpful for prediction, such as timestamp, user ID, data type, operation type (read / write), and response time.
[0073] Feature engineering: Create new features to enhance model performance, such as calculating the average number of visits in each time period and the maximum and minimum response time differences.
[0074] Model selection: The ARIMA model is suitable for time series data with obvious periodicity and trends. It can capture the trend and seasonality of data changing over time. LSTM (Long Short-Term Memory Network) model: It is particularly effective for nonlinear and complex time series data, can handle long-term dependency problems, and is very suitable for predicting future behavior patterns.
[0075] The terminal divides the data into training sets and test sets to ensure the generalization ability of the model on unknown data. The optimal parameter combination is found through grid search or random search to improve the accuracy of the model. The trained model is used to predict the data access volume and storage requirements in the future. Considering the changes in business scenarios, the input features can be adjusted dynamically to make the prediction closer to the actual needs.
[0076] Anomaly detection: In addition to conventional predictions, anomaly detection mechanisms can also be introduced to identify abnormal situations that may affect normal operations and prepare countermeasures in advance.
[0077] In one embodiment, after generating the data usage suggestion according to the future data usage trend, the method of the present application further includes the following steps:
[0078] According to the data usage suggestion, the data storage strategy is adjusted to obtain the adjusted target data storage strategy; according to the access frequency of each data in the multi-source data set, the high-frequency access data is screened out, and based on the target data storage strategy, the high-frequency access data is saved to the high-performance storage.
[0079] Specifically, for the prediction results, different priorities are set for different types of data. Frequently accessed data should be stored in high-performance storage; occasionally used data can be migrated to storage media with lower costs but slower access speeds; and historical data that is rarely used should be archived or deleted. Develop corresponding lifecycle management strategies for data of different priorities. This includes determining the data retention period, migration timing, and final destruction conditions.
[0080] In addition, when some data changes from high-frequency access to low-frequency access, their historical access frequency characteristics change significantly. These new features will be incorporated into the time series analysis model to help more accurately predict which data may become hot again in the future. As data is migrated or destroyed, the data set used to train the prediction model will also be updated accordingly. This means that the model can adjust its parameters based on the latest data usage to improve prediction accuracy.
[0081] The terminal can also reasonably allocate hardware resources based on the prediction results, such as purchasing additional storage space in advance or optimizing the existing storage architecture. Use automation tools (such as Kubernetes and Docker, etc.) to ensure the consistency of data status and governance policies, and automatically perform data migration tasks through scheduling tools (such as Apache Airflow) to ensure the effective implementation of policies. Establish a closed-loop feedback mechanism to regularly evaluate the prediction effect and policy execution, and adjust model parameters or retrain the model when necessary to adapt to the changing business environment.
[0082] In one embodiment, if Figure 4 As shown, the method of the present application also includes the following steps:
[0083] Step S401, obtaining feedback data of the prediction model, and generating multiple groups of candidate model parameter combinations through grid search according to the feedback data.
[0084] Step S402, determining a target model parameter combination from multiple groups of candidate model parameter combinations according to the prediction accuracy of each group of candidate model parameter combinations after model training.
[0085] Step S403, updating the prediction model according to the target model parameter combination to obtain an updated prediction model; the updated prediction model is used to replace the prediction model.
[0086] The prediction model may be an ARIMA model or an LSTM (Long Short-Term Memory) model.
[0087] The feedback data of the prediction model may be historical output data obtained through calculation and processing by the prediction model.
[0088] Specifically, the terminal generates multiple groups of candidate model parameter combinations according to the feedback data through grid search or random search, and after training the prediction model based on each group of candidate model parameter combinations, the prediction accuracy of the trained prediction model is tested, and then the target model parameter combination is screened out according to the prediction accuracy. Finally, the prediction model is updated according to the target model parameter combination to obtain an updated prediction model, thereby effectively improving the accuracy of the prediction model.
[0089] In one embodiment, if Figure 5 As shown, the method of the present application also includes the following steps:
[0090] Step S501 : acquiring historical user behavior information, and calculating the predicted probability of potential abnormal user behavior through a behavior analysis model based on the historical user behavior information.
[0091] Step S502: identifying target abnormal user behaviors from the potential abnormal user behaviors according to the predicted probability and the warning threshold, and generating warning information for the target abnormal user behaviors.
[0092] Step S503: In response to the user's request to process the warning information, generate security protection information for the target abnormal user behavior.
[0093] Specifically, the terminal can use behavioral analysis models (such as anomaly detection algorithms and cluster analysis models, etc.) to analyze user operation patterns. Based on historical user behavior information, it can identify potential abnormal user behaviors (such as unauthorized access and large amounts of data export, etc.). When potential abnormal user behavior is detected, the terminal automatically triggers an early warning message to inform relevant personnel to take measures. The early warning can be issued through emails, text messages, or platform notifications.
[0094] For example, the terminal collects and stores the daily operation patterns of all users, groups similar behaviors through cluster analysis models, and builds a "normal" behavior model for each user. These models can capture the typical activity characteristics of users in different scenarios. Then, machine learning or deep learning techniques (such as Isolation Forest and Autoencoder) are used to compare the current user behavior with the predefined normal pattern. If behavior deviates from the normal range, it is marked as potential abnormal behavior. In addition to statistical and model-based methods, clear rules are set to capture known types of abnormal behavior, such as unauthorized access attempts and data export volumes that exceed the norm. Once abnormal behavior is confirmed, the terminal will immediately send an alert to the relevant responsible person through pre-configured methods (such as email, SMS, and platform notifications). This ensures that information can be obtained in a timely manner even during non-working hours. Based on preset conditions, the terminal may automatically take certain protective actions, such as temporarily locking suspicious accounts and restricting access rights of specific IP addresses, to prevent possible risk spread. For complex or uncertain situations, the terminal will trigger a manual review process, and a professional team will further investigate and decide whether more in-depth processing is required. All detected security incidents and their processing results are recorded so that the latest security status can be taken into account during data extraction, conversion, and loading, so that corresponding policies can be adjusted. As more security incidents occur and processing experience is accumulated, the terminal continuously optimizes its anomaly detection model and rule set to improve its ability to identify similar problems in the future, thereby ensuring the security and stability of the entire data governance system.
[0095] In addition, in this embodiment, ETL (Extract-Transform-Load) tools such as Apache Nifi or Talend can also be integrated on the terminal to intelligently select the best data transmission solution. The data extraction, transformation and loading processes are automatically performed through the scheduling system. Metadata is automatically managed during the ETL process to ensure the integrity and consistency of the data. The metadata management platform is integrated to achieve consistency management of data quality and standards. With the help of scheduling systems such as Apache Airflow, ETL tasks can be executed at the most appropriate time to avoid conflicts with other high-load operations. In addition, considering security monitoring, any operations involving sensitive data will be subject to additional protection. Relevant metadata records are automatically generated and updated throughout the ETL process to maintain consistency in data standards. This is not only crucial for subsequent data storage strategies, but also supports the automated repair of data quality issues.
[0096] As an application example, the scenario of smart city traffic management is taken as an example to specifically explain how the adaptive data governance method in this application is applied in actual scenarios. In this embodiment, it is necessary to process data from multiple sources, including traffic cameras, sensors, GPS devices on public transportation, citizens' mobile application reports, and social media, etc., in order to optimize urban traffic flow, reduce congestion, and improve traffic safety.
[0097] Step 1: Real-time monitoring and identification of new data sources:
[0098] Collect data from road sensors, traffic lights, weather stations, etc. in real time, and monitor new information from social media and citizen feedback. Automatically identify and access new data sources (such as newly installed sensors or traffic condition reports provided by third parties) through embedded intelligent algorithms, and evaluate their quality.
[0099] Step 2: Repair data quality issues:
[0100] When the vehicle speed data on certain sections of road is detected to be abnormally high or low, or a camera video stream is interrupted, the system automatically identifies these problems and attempts to repair them. AI technology is used to identify and correct incorrect or missing data, such as inferring a reasonable vehicle speed value through data from adjacent sensors; the repaired data is fed back to subsequent steps to ensure continuous improvement in data quality.
[0101] Step 3: Dynamically adjust data governance strategies:
[0102] Dynamically adjust traffic management rules based on factors such as weather changes, holidays, and special events, such as temporarily changing traffic light durations or suggesting alternative routes. Regularly monitor changes in business scenarios, such as extending the green light time at specific intersections during peak hours or setting up dedicated lanes during large-scale events.
[0103] Step 4: Trend prediction and optimization suggestions:
[0104] Predict traffic conditions in the coming days or even weeks based on historical traffic patterns and current trends, and plan response measures in advance. Use machine learning models to analyze historical data trends, predict future traffic hotspots, and provide citizens with travel suggestions, such as avoiding peak hours or choosing better routes.
[0105] Step 5: Security assurance:
[0106] Protect the network security of the transportation system to prevent hacker attacks from causing traffic paralysis or misleading the public. Implement behavioral analysis components to monitor user operation patterns in real time, trigger early warning mechanisms when abnormalities are found, and ensure the security of the system.
[0107] Step 6: Third-party application support:
[0108] Map service providers and online car-hailing platforms are allowed to access the system to share real-time traffic information and improve overall service quality. Cross-platform data sharing and collaboration are achieved through standardized API interfaces, ensuring data consistency and efficient collaboration between different departments.
[0109] Step 7: Intelligent decision support:
[0110] Help city managers make more informed decisions, such as whether to invest in new infrastructure projects. The integrated intelligent analysis platform processes key data in real time, generates visual reports for decision makers to refer to, and supports more scientific resource allocation.
[0111] Step 8: Intelligent ETL process:
[0112] Automatically extract, transform and load data from different sources to ensure that all relevant information can be integrated into the central database in a timely manner. Use ETL tools to intelligently select the best data transmission solution to ensure data integrity and consistency.
[0113] Step 9: Data storage and lifecycle management:
[0114] Arrange the storage location of traffic-related data to ensure fast response to frequently accessed data and proper archiving of expired data. Dynamically adjust storage strategies based on the importance and access frequency of data to optimize resource utilization efficiency.
[0115] Step 10: Application of blockchain technology:
[0116] Ensure the security and traceability of traffic data, especially in cases involving accident liability determination or insurance claims. Combined with blockchain technology to record data access and modification history, provide tamper-proof log records, and enhance trust.
[0117] Step 11: Audit and compliance monitoring:
[0118] Ensure that traffic management complies with local and national regulations, such as privacy protection laws, traffic management regulations, etc. Use AI audit technology to monitor data operation behaviors in real time, automatically generate compliance reports, and promptly detect and correct any violations.
[0119] Step 12: User-friendly interface:
[0120] Provide an intuitive operation panel for traffic management departments to easily view various indicators and parameters and make necessary adjustments. Design an easy-to-use control panel that integrates all data governance functions, allowing users to easily manage and optimize traffic system performance.
[0121] The beneficial effects brought by the above embodiments are as follows:
[0122] 1) Through the application of the adaptive data governance method in this application in the smart city traffic management scenario, it not only improves the traffic management efficiency and reduces the incidence of traffic accidents, but also promotes the sustainable development of the city, reflecting the value of the adaptive data governance solution in practical applications.
[0123] 2) It effectively solves the problems existing in traditional data governance platforms, such as insufficient flexibility, limited intelligence, difficulty in cross-platform collaboration, and lack of full life cycle management, and provides modern enterprises with more advanced, reliable and flexible data governance solutions.
[0124] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0125] Based on the same inventive concept, the embodiment of the present application also provides an adaptive data governance device for implementing the adaptive data governance method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more adaptive data governance device embodiments provided below can refer to the limitations of the adaptive data governance method above, and will not be repeated here.
[0126] In an exemplary embodiment, Figure 6 As shown, an adaptive data governance device is provided, comprising:
[0127] The data acquisition module 601 is used to collect multi-source data sets in the current business scenario, analyze the current data features of the multi-source data sets, and identify data quality problems in the multi-source data sets based on the current data features;
[0128] The data repair module 602 is used to determine a corresponding repair solution according to the type of data quality problem, and repair the multi-source data set based on the repair solution to obtain a repaired target data set;
[0129] The policy adjustment module 603 is used to analyze the data change information of the target data set, and when the data change information meets the policy adjustment conditions, adjust the data governance policy under the current business scenario to obtain the adjusted target data governance policy;
[0130] The information feedback module 604 is used to provide repair feedback on the target data set and adjust the target data governance strategy through a preset feedback mechanism.
[0131] In one embodiment, the data repair module 602 is also used to perform unified standardization processing on the data in the multi-source data sets when the data formats in the multi-source data sets are inconsistent; to determine the location of the missing values and fill the missing values when there are missing data in the multi-source data sets; and to determine the location of the outliers and correct them when there are data anomalies in the multi-source data sets.
[0132] In one embodiment, the adaptive data governance device also includes a trend prediction module, which is used to obtain historical data records, perform feature selection and feature extraction on the historical data records, and obtain historical data features; predict future data usage trends through a prediction model based on historical data features and data change information; and generate data usage recommendations based on future data usage trends.
[0133] In one embodiment, the adaptive data governance device also includes a storage adjustment module, which is used to adjust the data storage strategy according to the data usage recommendation to obtain an adjusted target data storage strategy; filter out high-frequency access data according to the access frequency of each data in the multi-source data set, and save the high-frequency access data to high-performance storage based on the target data storage strategy.
[0134] In one embodiment, the adaptive data governance device also includes a model update module, which is used to obtain feedback data of the prediction model, and generate multiple groups of candidate model parameter combinations through grid search based on the feedback data; determine the target model parameter combination from the multiple groups of candidate model parameter combinations based on the prediction accuracy after model training of each group of candidate model parameter combinations; update the prediction model based on the target model parameter combination to obtain an updated prediction model; the updated prediction model is used to replace the prediction model.
[0135] In one embodiment, the adaptive data governance device also includes an abnormal warning module, which is used to obtain historical user behavior information, calculate the predicted probability of potential abnormal user behavior through a behavior analysis model based on the historical user behavior information; identify target abnormal user behavior among potential abnormal user behaviors based on the predicted probability and the warning threshold, and generate warning information for the target abnormal user behavior; and generate security protection information for the target abnormal user behavior in response to the user's request to process the warning information.
[0136] Each module in the above-mentioned adaptive data governance device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. When the computer program is executed by the processor, an adaptive data governance method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0138] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0139] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0141] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0143] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0144] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0145] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. An adaptive data governance method, characterized in that: The method comprises: Collecting a multi-source data set in a current business scenario, analyzing current data features of the multi-source data set, and identifying data quality issues in the multi-source data set based on the current data features; Determine a corresponding repair solution according to the type of the data quality problem, and repair the multi-source data set based on the repair solution to obtain a repaired target data set; Analyze the data change information of the target data set, and if the data change information meets the policy adjustment condition, adjust the data governance policy under the current business scenario to obtain an adjusted target data governance policy; Repair feedback is performed on the target data set through a preset feedback mechanism, and adjustment feedback is performed on the target data governance strategy.
2. The method according to claim 1, characterized in that Determining a corresponding repair solution according to the type of the data quality problem includes: In the case where the data formats in the multi-source data sets are inconsistent, performing unified standardization processing on the data in the multi-source data sets; In the case where there is missing data in the multi-source data set, determining the location of the missing value and filling the missing value; When data anomalies exist in the multi-source data set, the location of the anomaly is determined and the anomaly is corrected.
3. The method according to claim 1, characterized in that The method further comprises: Acquire historical data records, perform feature selection and feature extraction on the historical data records, and obtain historical data features; Predicting future data usage trends through a prediction model based on the historical data characteristics and the data change information; Generate data usage recommendations based on the future data usage trends.
4. The method according to claim 3, characterized in that After generating the data usage suggestion according to the future data usage trend, the method further includes: According to the data usage suggestion, the data storage strategy is adjusted to obtain an adjusted target data storage strategy; According to the access frequency of each data in the multi-source data set, the frequently accessed data is screened out, and based on the target data storage strategy, the frequently accessed data is saved in the high-performance storage.
5. The method according to claim 3, characterized in that: The method further comprises: Acquiring feedback data of the prediction model, and generating multiple sets of candidate model parameter combinations through grid search according to the feedback data; Determining a target model parameter combination from among the plurality of groups of candidate model parameter combinations according to the prediction accuracy of each group of candidate model parameter combinations after model training; The prediction model is updated according to the target model parameter combination to obtain an updated prediction model; the updated prediction model is used to replace the prediction model.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining historical user behavior information, and calculating the predicted probability of potential abnormal user behavior through a behavior analysis model based on the historical user behavior information; According to the prediction probability and the warning threshold, target abnormal user behavior is identified from the potential abnormal user behavior, and warning information for the target abnormal user behavior is generated; In response to the user's request to process the warning information, security protection information for the target abnormal user behavior is generated.
7. An adaptive data management device, characterized in that: The device comprises: A data acquisition module is used to collect a multi-source data set in a current business scenario, analyze the current data features of the multi-source data set, and identify data quality problems in the multi-source data set based on the current data features; A data repair module, used to determine a corresponding repair solution according to the type of the data quality problem, and repair the multi-source data set based on the repair solution to obtain a repaired target data set; A policy adjustment module is used to analyze the data change information of the target data set, and if the data change information meets the policy adjustment conditions, adjust the data governance policy under the current business scenario to obtain an adjusted target data governance policy; The information feedback module is used to provide repair feedback to the target data set and adjust feedback to the target data governance strategy through a preset feedback mechanism.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Industrial equipment data recovery method, system, equipment, medium and product
CN120653641A
Self-evolution emergence type data governance method and system
CN120973783A