Unified data comprehensive management system

Through a unified comprehensive data management system, using databases and data lakes to manage data from different data sources, the information island problem caused by data dispersion is solved, efficient data governance and management is achieved, and business efficiency is improved.

CN119961340APending Publication Date: 2025-05-09HUANENG ZHAOCAI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411777852.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate and manage data scattered in different systems and equipment, resulting in information silos and inefficient business.

Method used

Provide a unified data comprehensive management system, which stores data from different data sources through the database and adopts dynamic management of data lakes; encrypts data, dynamic classification access rights management, tracks the entire life cycle of data assets, monitors system logs in real time, and data visualization.

Benefits of technology

Effectively improve the efficiency and accuracy of unified data governance and management, and thus improve business efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961340A_ABST
    Figure CN119961340A_ABST
Patent Text Reader

Abstract

The invention provides a unified data comprehensive management system, and relates to the technical field of data management, and the system comprises a storage management module which is used for storing data from different data sources and carrying out the dynamic management through a data lake; the security management module is used for encrypting data and dynamically classifying data access authority; the quality management module is used for tracking the full life cycle of the data assets and monitoring system logs in real time; and the visualization module is used for realizing data visualization by utilizing a set visualization tool. A database is used for storing data of different data sources, and data lake dynamic management is adopted; according to data encryption, dynamic classification of data access authority, tracking of full life cycle of data assets, real-time monitoring of system logs and data visualization, the efficiency and accuracy of unified data governance and management can be effectively improved, and then the service efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular to a unified data integrated management system. Background Art

[0002] In recent years, with the promotion of the digital wave, data has become one of the most valuable assets of enterprises. However, data is scattered in different systems and devices, forming information islands, which are difficult to integrate and analyze effectively, resulting in low business efficiency. Therefore, how to break the information islands, realize centralized data management and sharing, and improve business efficiency has become one of the current research centers.

[0003] Therefore, the present invention provides a unified data comprehensive management system. Summary of the invention

[0004] The present invention provides a unified data comprehensive management system, which is used to store data from different data sources by utilizing a database and adopting data lake dynamic management; data encryption, dynamic classification of data access rights, tracking of the entire life cycle of data assets, real-time monitoring of system logs and data visualization, can effectively improve the efficiency and accuracy of unified data governance and management.

[0005] The present invention provides a unified data integrated management system, comprising: Storage management module: used to store data from different data sources and use data lake dynamic management; Security management module: used to encrypt data and dynamically classify data access rights; Quality management module: used to track the entire life cycle of data assets and monitor system logs in real time; Visualization module: used to realize data visualization by using set visualization tools.

[0006] Preferably, the storage management module includes: Warehouse-storage unit: used to establish a target database on a target server according to business requirements, and encrypt and store first data from different data sources into the target database according to a set import method; The target database creates an index for the structured data to provide a fast query function; Data lake-management unit: used to deploy the target data lake using the access block, the target data lake accesses and processes the first data to obtain key data; uses the annotation block to annotate the key data with a sensitivity level; and uses the storage block to store the key data; Access block: used to deploy a target data lake on a target server according to business requirements. The target data lake accesses first data from various data sources and performs data preprocessing on the first data to obtain key data. Annotation block: used to divide the key data according to data characteristics, analyze the sensitivity of the key data, obtain the sensitivity level and make corresponding annotations; Storage block: used to match the corresponding storage policy according to the sensitivity level of key data and store it in the corresponding storage area of ​​the target data lake.

[0007] Preferably, data sources include various business systems, government data, Internet data, and industry data, wherein business systems include human resources, finance, supply chain, production, marketing, and customer service.

[0008] Preferably, the annotation block includes: After dividing the key data according to data characteristics, first characteristic data is obtained; Acquire a data access record of the first characteristic data, and acquire a first access mode of the current first characteristic data and a first access count within a preset time period from the data access record, and output them as first analysis data; According to the first access mode, determining whether the current first characteristic data has specific frequent accesses, and the corresponding first average access frequency when the specific frequent accesses exist, and outputting them as first analysis data; Determine whether there is specific rare access in the current first characteristic data, and the corresponding second average access frequency when the specific rare access exists, and output it as the first analysis data; Using the set important evaluation index to evaluate the importance of the current first characteristic data, and obtain an important evaluation value; The acquired first analysis data is combined with the important evaluation value to calculate the data sensitivity coefficient; Taking the data sensitivity coefficient as a matching condition, a sensitivity level is matched from a preset sensitivity level mapping table to label the current first characteristic data.

[0009] Preferably, the calculation formula of the data sensitivity coefficient is as follows: ; In the formula, Expressed as the sensitivity coefficient of the current first characteristic data; represents a preset time period; L1 represents the first access count of the current first characteristic data within the preset time period T; ln represents a natural logarithm; e represents a constant, and its value is 2.7; It is expressed as the weight of the impact of access frequency on the sensitivity of the assessed data; The total number of times a specific frequent access is represented as the current first characteristic data; It is represented as the first average access frequency corresponding to the a-th specific frequent access of the first characteristic data, where a=1, 2, 3, , n1; an average value of corresponding first average access frequencies of all specific frequent accesses of the first characteristic data; The total number of specific rare accesses to the current first characteristic data is represented; It is represented as the corresponding second average access frequency of the specific rare access of the first characteristic data for the bth time, where b=1, 2, 3, , n2; an average value of corresponding first average access frequencies of all specific rare accesses of the first characteristic data; It is expressed as the weight of the impact of data importance on the assessment of data sensitivity; It is represented by the importance evaluation value obtained by using the i-th set important evaluation index to evaluate the importance of the current first characteristic data, where i=1, 2, 3; It is represented as the evaluation weight of the importance of the i-th important evaluation indicator to the evaluation data.

[0010] Preferably, the security management module includes: Transmission encryption unit: used to encrypt key data to be transmitted using a set security protocol; Life management unit: used to dynamically manage the life of key data based on the sensitivity level and data characteristics of the current key data; Access control unit: used to define a preset access permission list for the key data, and based on a pre-established permission analysis model, regularly output new access permissions for the key data, and update the preset access permission list for the current key data based on the new access permissions.

[0011] Preferably, the life management unit comprises: Retention block: used to determine the initial retention strategy for key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is rarely sensitive or slightly sensitive, the initial retention policy is output as the first retention policy; If the sensitivity level of the current key data is moderately sensitive or severely sensitive, the corresponding sensitivity impact coefficient of the current sensitivity level is obtained, and the initial retention period extracted from the initial retention policy is adjusted to obtain the first retention period; After replacing the initial retention period in the initial retention policy with the first retention period, the first retention policy is outputted as the first retention policy; Performing data retention management on current key data using the first retention policy; Encryption block: when the sensitivity level of key data is mildly sensitive, moderately sensitive or severely sensitive, the first retention period is extracted from the first retention policy currently applied to the key data, and combined with the real-time storage duration of the key data for analysis to determine the first life stage coefficient of the current key data; The first life stage coefficient is calculated by combining the sensitivity impact coefficient corresponding to the sensitivity level of the current key data to obtain the encryption analysis coefficient; Using the encryption analysis coefficient, obtaining the encryption analysis level of the current key data; Determine the expected encryption scope and expected encryption algorithm for the current key data based on the obtained encryption analysis level; The expected encryption range of the current key data is encrypted using the expected encryption algorithm; Destruction block: used to determine the initial destruction strategy of key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is extremely low sensitivity, the initial destruction strategy is output as the first destruction strategy; If the sensitivity level of the current key data is slightly sensitive, moderately sensitive, or severely sensitive, the first destruction time is obtained using the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; The initial destruction time in the initial destruction strategy is replaced with the first destruction time and then output as the first destruction strategy; The first destruction strategy is used to manage data destruction of current key data.

[0012] Preferably, the quality management module includes: Asset lifecycle management unit: used to select metadata management tools that match business needs; Use the collection tool in the metadata management tool to collect all metadata information in the target database; Combine the acquired metadata information with the pre-established asset lifecycle management model to record and track the entire lifecycle of data assets and generate asset lifecycle management reports; Log monitoring unit: used to collect system log entries in real time using a set log collection tool, pre-process and aggregate the system log entries, obtain aggregated logs and store them accordingly; According to business needs, select a log monitoring tool that is suitable for the current system, and configure log anomaly analysis rules and corresponding alarm conditions in the log monitoring tool; Use the currently set log monitoring tool to monitor system logs and abnormal alarms in real time.

[0013] Preferably, the visualization module includes: Visualization unit: used to establish a connection with the target database using a set visualization tool, and then create a data visualization interface for the current system according to the preset interface requirements; Use visualization components in the data visualization interface to provide multiple report analysis functions for different business types of data; Provides user interaction functions to enable drilling, filtering, exporting and sharing of multiple styles of reports.

[0014] Compared with the prior art, the present invention has the following beneficial effects: By utilizing databases to store data from different data sources and adopting data lake dynamic management; encrypting data, dynamically classifying data access rights, tracking the entire life cycle of data assets, real-time monitoring of system logs, and data visualization, the efficiency and accuracy of unified data governance and management can be effectively improved, thereby improving business efficiency.

[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0016] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a structural diagram of a unified data comprehensive management system in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0019] The embodiment of the present invention provides a unified data integrated management system, such as Figure 1 As shown, including: Storage management module: used to store data from different data sources and use data lake dynamic management; Security management module: used to encrypt data and dynamically classify data access rights; Quality management module: used to track the entire life cycle of data assets and monitor system logs in real time; Visualization module: used to realize data visualization by using set visualization tools.

[0020] In this embodiment, data sources include various business systems, government data, Internet data, and industry data, among which business systems include human resources, finance, supply chain, production, marketing, and customer service; data lake refers to a storage system used to store and manage large amounts of data; set visualization tools are predetermined tools for creating visualization interfaces, such as Tableau, Power BI, and ECharts tools.

[0021] The beneficial effects of the above technical solution are: by using the database to store data from different data sources and adopting data lake dynamic management; encrypting data, dynamically classifying data access rights, tracking the entire life cycle of data assets, real-time monitoring of system logs and data visualization, it can effectively improve the efficiency and accuracy of unified data governance and management, thereby improving business efficiency.

[0022] The embodiment of the present invention provides a unified data integrated management system, wherein the storage management module comprises: Warehouse-storage unit: used to establish a target database on a target server according to business requirements, and encrypt and store first data from different data sources into the target database according to a set import method; The target database creates an index for the structured data to provide a fast query function; Data lake-management unit: used to deploy the target data lake using the access block, the target data lake accesses and processes the first data to obtain key data; uses the annotation block to annotate the key data with a sensitivity level; and uses the storage block to store the key data; Access block: used to deploy a target data lake on a target server according to business requirements. The target data lake accesses first data from various data sources and performs data preprocessing on the first data to obtain key data. Annotation block: used to divide the key data according to data characteristics, analyze the sensitivity of the key data, obtain the sensitivity level and make corresponding annotations; Storage block: used to match the corresponding storage policy according to the sensitivity level of key data and store it in the corresponding storage area of ​​the target data lake.

[0023] In this embodiment, business requirements include business type, data type, data format, data scale, support for business decisions and query requirements; the target server refers to a physical or virtual server selected for deploying a specific database or data lake; the target database refers to a database established on the target server for storing first data from different data sources, which can be designed and configured according to business requirements and can support functions such as encrypted data storage and fast query.

[0024] In this embodiment, data sources include various business systems, government data, Internet data, and industry data, among which business systems include human resources, finance, supply chain, production, marketing, and customer service; the set import method is a predetermined data import method, such as batch import and streaming import; the target data lake refers to a storage system deployed on the target server for storing and managing large amounts of data.

[0025] In this embodiment, key data refers to data obtained by preprocessing the first data; data characteristics refer to time, region and business type, where business types include human resources, finance, supply chain, production, marketing and customer service; sensitivity levels include four levels: extremely sensitive, slightly sensitive, moderately sensitive and severely sensitive; the storage strategy is composed of the storage area, storage method (such as separation of hot and cold data, data compression, data sharding), whether to encrypt the storage and the corresponding encryption algorithm, such as AES and other encryption algorithms.

[0026] The beneficial effects of the above technical solution are: by establishing a database to store data from different data sources and adopting data lake dynamic management, data security can be effectively improved, and strong data support can be provided for the rapid development of the business.

[0027] The embodiment of the present invention provides a unified data integrated management system, wherein the annotation block includes: After dividing the key data according to data characteristics, first characteristic data is obtained; Acquire a data access record of the first characteristic data, and acquire a first access mode of the current first characteristic data and a first access count within a preset time period from the data access record, and output them as first analysis data; According to the first access mode, determining whether the current first characteristic data has specific frequent accesses, and the corresponding first average access frequency when the specific frequent accesses exist, and outputting them as first analysis data; Determine whether there is specific rare access in the current first characteristic data, and the corresponding second average access frequency when the specific rare access exists, and output it as the first analysis data; Using the set important evaluation index to evaluate the importance of the current first characteristic data, and obtain an important evaluation value; The acquired first analysis data is combined with the important evaluation value to calculate the data sensitivity coefficient; Taking the data sensitivity coefficient as a matching condition, a sensitivity level is matched from a preset sensitivity level mapping table to label the current first characteristic data.

[0028] In this embodiment, data characteristics refer to time, region and business type; first characteristic data refers to data after key data is divided according to data characteristics; data access records are composed of data access mode, access time and access times, wherein the access mode includes the identity of the visitor, the type of accessed device, the geographical location of the access, and the way and time distribution of data access. For example, data may be frequently accessed at a specific time (such as working hours on weekdays), but less frequently accessed at other times; specific frequent access refers to frequent access within a specific time; the first average access frequency refers to the access frequency of frequent access to the first characteristic data within a specific time; specific rare access refers to less access within a specific time; the second average access frequency refers to the access frequency of extremely rare access to the first characteristic data within a specific time; the first access mode refers to the access mode extracted from the data access record of the current first characteristic data; the preset time period is predetermined; the first access number refers to the total number of accesses to the first characteristic data within the preset time period extracted from the data access record of the current first characteristic data.

[0029] In this embodiment, the setting of important evaluation indicators is predetermined, including the frequency of data usage, the number of data-driven decisions, and the types of business involved in the data; the data sensitivity coefficient is used to express the sensitivity and importance of the current first characteristic data; the preset sensitivity level mapping table is pre-established, consisting of the data sensitivity coefficient range and the corresponding sensitivity level.

[0030] The beneficial effect of the above technical solution is: by dividing the key data according to the data characteristics, analyzing the sensitivity of the key data to obtain the sensitivity level and making corresponding annotations, it can provide a data basis for the subsequent data storage strategy matching.

[0031] The embodiment of the present invention provides a unified data integrated management system, and the calculation formula of the data sensitivity coefficient is as follows: ; In the formula, Expressed as the sensitivity coefficient of the current first characteristic data; represents a preset time period; L1 represents the first access count of the current first characteristic data within the preset time period T; ln represents a natural logarithm; e represents a constant, and its value is 2.7; It is expressed as the weight of the impact of access frequency on the sensitivity of the assessed data; The total number of times a specific frequent access is represented as the current first characteristic data; It is represented as the first average access frequency corresponding to the a-th specific frequent access of the first characteristic data, where a=1, 2, 3, , n1; an average value of corresponding first average access frequencies of all specific frequent accesses of the first characteristic data; The total number of specific rare accesses to the current first characteristic data is represented; It is represented as the corresponding second average access frequency of the specific rare access of the first characteristic data for the bth time, where b=1, 2, 3, , n2; an average value of corresponding first average access frequencies of all specific rare accesses of the first characteristic data; It is expressed as the weight of the impact of data importance on the assessment of data sensitivity; It is represented by the importance evaluation value obtained by using the i-th set important evaluation index to evaluate the importance of the current first characteristic data, where i=1, 2, 3; It is represented as the evaluation weight of the importance of the i-th important evaluation indicator to the evaluation data.

[0032] In this embodiment, the weights assigned to access frequency and data importance are obtained by solving a matrix constructed after pairwise comparison and relative importance scoring using the hierarchical analysis method; the weights assigned to set important evaluation indicators are obtained by solving a matrix constructed after pairwise comparison and relative importance scoring using the hierarchical analysis method.

[0033] The beneficial effect of the above technical solution is that by calculating the data sensitivity coefficient, a data basis can be provided for matching the sensitivity level of the first characteristic data, which helps to accurately match the storage strategy of subsequent data, thereby improving the reliability of data storage.

[0034] The embodiment of the present invention provides a unified data integrated management system, wherein the security management module includes: Transmission encryption unit: used to encrypt key data to be transmitted using a set security protocol; Life management unit: used to dynamically manage the life of key data based on the sensitivity level and data characteristics of the current key data; Access control unit: used to define a preset access permission list for the key data, and based on a pre-established permission analysis model, regularly output new access permissions for the key data, and update the preset access permission list for the current key data based on the new access permissions.

[0035] In this embodiment, the set security protocol is a predetermined protocol for encrypting transmitted data, such as the SSL / TLS security protocol; dynamic life management includes data retention, encryption and destruction management; the preset access permission list refers to a predetermined permission list that clearly defines which users or user groups can access key data; the new access permission refers to the access permission output after the permission is updated according to a preset time period using the permission analysis model; the permission analysis model is established based on data security and access control requirements, and the establishment steps include defining roles and permissions, analyzing data sensitivity, formulating access rules and implementing optimization, among which, defining roles and permissions refers to defining different user roles according to business needs and personnel structure, and assigning corresponding data access permissions to each role; analyzing data sensitivity refers to analyzing the sensitivity coefficient and sensitivity level of data; formulating access rules refers to formulating detailed access rules based on the sensitivity of roles and data, including who can access which data and under what conditions can they access it; implementing optimization refers to continuously optimizing the permission analysis model according to changes in business needs and the analysis results of access rules.

[0036] The beneficial effects of the above technical solution are: by encrypting data and dynamically classifying data access rights, it is possible to enhance the security of data transmission, protect data privacy, adapt to changes in data usage scenarios and changes in user needs, and ensure the flexibility and security of data access.

[0037] The embodiment of the present invention provides a unified data integrated management system, wherein the life management unit includes: Retention block: used to determine the initial retention strategy for key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is rarely sensitive or slightly sensitive, the initial retention policy is output as the first retention policy; If the sensitivity level of the current key data is moderately sensitive or severely sensitive, the corresponding sensitivity impact coefficient of the current sensitivity level is obtained, and the initial retention period extracted from the initial retention policy is adjusted to obtain the first retention period; After replacing the initial retention period in the initial retention policy with the first retention period, the first retention policy is outputted as the first retention policy; Performing data retention management on current key data using the first retention policy; Encryption block: when the sensitivity level of key data is mildly sensitive, moderately sensitive or severely sensitive, the first retention period is extracted from the first retention policy currently applied to the key data, and combined with the real-time storage duration of the key data for analysis to determine the first life stage coefficient of the current key data; The first life stage coefficient is calculated by combining the sensitivity impact coefficient corresponding to the sensitivity level of the current key data to obtain the encryption analysis coefficient; Using the encryption analysis coefficient, obtaining the encryption analysis level of the current key data; Determine the expected encryption scope and expected encryption algorithm for the current key data based on the obtained encryption analysis level; The expected encryption range of the current key data is encrypted using the expected encryption algorithm; Destruction block: used to determine the initial destruction strategy of key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is extremely low sensitivity, the initial destruction strategy is output as the first destruction strategy; If the sensitivity level of the current key data is slightly sensitive, moderately sensitive, or severely sensitive, the first destruction time is obtained using the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; The initial destruction time in the initial destruction strategy is replaced with the first destruction time and then output as the first destruction strategy; The first destruction strategy is used to manage data destruction of current key data.

[0038] In this embodiment, the initial retention policy is a retention policy selected from the preset retention policy mapping table using the data characteristics of key data as the screening condition, and is composed of the data retention principle, period, storage method and access rights; the preset retention policy mapping table is composed of data characteristics and corresponding retention policies; the sensitive impact coefficient is predetermined and has a value range of , and the higher the sensitivity level, the greater the sensitivity impact coefficient, that is, 4, among which, The sensitivity impact coefficient expressed as a very small sensitivity level; The sensitivity impact coefficient expressed as a mild sensitivity level; The sensitivity impact coefficient is expressed as a moderate sensitivity level; 4 represents the sensitivity impact coefficient of the severe sensitivity level.

[0039] In this embodiment, the first retention period is obtained by adjusting the initial retention period extracted from the initial retention policy using the sensitive impact coefficient. The retention period refers to the length of time the data is retained. The formula for the first retention period is: ;in, It is indicated as the first retention period; Expressed as the initial retention period; It is expressed as the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; e is expressed as a constant with a value of 2.7.

[0040] In this embodiment, the first retention policy refers to the initial retention policy (when the sensitivity level of the current key data is extremely sensitive or mildly sensitive), or the retention policy generated by replacing the initial retention period in the initial retention policy with the first retention period (when the sensitivity level of the current key data is moderately sensitive or severely sensitive).

[0041] In this embodiment, the first life stage coefficient is obtained by dividing the real-time storage duration of the key data by the first retention period, where the real-time storage duration refers to the storage time of the current key data; the encryption analysis coefficient is used to evaluate the encryption degree of the current key data, and the formula is: , where J1 represents the encryption analysis coefficient of the current key data; It is represented as the first life stage coefficient of the current key data; w1 is represented as the influence weight of the first life stage coefficient on the degree of encryption required for analysis; It is expressed as the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; It is expressed as the weight of the sensitive impact coefficient on the degree of encryption required for analysis; the weights assigned to the sensitive impact coefficient and the first life stage coefficient are obtained by solving the matrix constructed after pairwise comparison and relative importance scoring using the hierarchical analysis method.

[0042] In this embodiment, the encryption analysis level is an encryption level obtained by matching the currently acquired encryption analysis coefficient from a set encryption level table, wherein the encryption level includes general encryption, special encryption and strict encryption; the set encryption level table is composed of the encryption analysis coefficient range and the corresponding encryption level.

[0043] In this embodiment, the expected encryption range refers to the data encryption range of the current key data determined according to the encryption analysis level; the expected encryption algorithm refers to the data encryption algorithm of the current key data determined according to the encryption analysis level, such as symmetric encryption (AES) and asymmetric encryption (RSA).

[0044] In this embodiment, the initial destruction strategy consists of a destruction method, time, responsible person, and data backup plan; the initial destruction time refers to the destruction time point of the key data extracted from the initial destruction strategy; the first destruction time is a destruction time point determined by combining the dynamic destruction time period with the storage start time point of the key data after determining the dynamic destruction time period based on the sensitivity impact coefficient corresponding to the sensitivity level of the current key data, wherein the formula of the dynamic destruction time period is expressed as ;in, Represented as a dynamic destruction time period; It is expressed as the initial destruction time period, which is determined based on the storage start time point of the key data and the initial destruction time; It is expressed as the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; e is expressed as a constant with a value of 2.7.

[0045] In this embodiment, for example, if the storage start time of key data 1 is 13:00 on the 22nd and the dynamic destruction time period is 15 hours, then the first destruction time of key data 1 is 4:00 on the 23rd.

[0046] In this embodiment, the first destruction strategy refers to the initial destruction strategy (when the sensitivity level of the current key data is extremely sensitive), or the destruction strategy generated by replacing the initial destruction time in the initial destruction strategy with the first destruction time (when the sensitivity level of the current key data is slightly sensitive, moderately sensitive or severely sensitive).

[0047] The beneficial effect of the above technical solution is: by dynamically managing the life of key data based on the sensitivity level and data characteristics of the current key data, the security of the data can be effectively enhanced and the efficiency and accuracy of data governance and management can be improved.

[0048] The embodiment of the present invention provides a unified data integrated management system, wherein the quality management module includes: Asset lifecycle management unit: used to select metadata management tools that match business needs; Use the collection tool in the metadata management tool to collect all metadata information in the target database; Combine the acquired metadata information with the pre-established asset lifecycle management model to record and track the entire lifecycle of data assets and generate asset lifecycle management reports; Log monitoring unit: used to collect system log entries in real time using a set log collection tool, pre-process and aggregate the system log entries, obtain aggregated logs and store them accordingly; According to business needs, select a log monitoring tool that is suitable for the current system, and configure log anomaly analysis rules and corresponding alarm conditions in the log monitoring tool; Use the currently set log monitoring tool to monitor system logs and abnormal alarms in real time.

[0049] In this embodiment, the metadata management tool refers to a software tool for collecting, storing, managing and analyzing metadata; the acquisition tool is used to collect all metadata information in the data warehouse, including database structure information, data dictionary, associations between data tables, etc.; the steps for establishing the asset lifecycle management model include data collection, data storage, data use, data archiving and data destruction, and each step records the corresponding metadata and information; among them, data collection refers to collecting data from various data sources and preprocessing and cleaning it, data storage refers to storing data in a suitable storage medium, such as a data lake, and data use refers to querying, analyzing, mining and other operations on data according to business needs; data archiving refers to archiving data that is no longer frequently used to save storage space, and data destruction refers to destroying data that is no longer needed according to the destruction policy.

[0050] In this embodiment, the full life cycle includes collecting, storing, managing and maintaining metadata; the asset life cycle management report is generated based on the asset life cycle management model, and is a report used to display the full life cycle status and changes of data assets. It consists of a data asset list, data life cycle status (the current life cycle stage of the data asset), data usage (such as data usage frequency, access mode), data quality issues, etc.; system log entries refer to records generated when the system is running, which are used to record the system's operating status, error messages, user operations, etc.; aggregate logs refer to logs obtained by merging and organizing multiple system log entries; setting log monitoring tools is predetermined, such as Elasticsearch, Logstash, and Kibana; log anomaly analysis rules refer to predetermined rules for detecting and analyzing abnormal behavior in system logs. When content that meets these rules appears in the system log, an abnormal alarm will be triggered to remind staff to handle it.

[0051] The beneficial effect of the above technical solution is that by tracking the entire life cycle of data assets and monitoring, aggregating, analyzing and alarming system logs, it can help ensure data compliance and improve data security.

[0052] The embodiment of the present invention provides a unified data integrated management system, wherein the visualization module comprises: Visualization unit: used to establish a connection with the target database using a set visualization tool, and then create a data visualization interface for the current system according to the preset interface requirements; Use visualization components in the data visualization interface to provide multiple report analysis functions for different business types of data; Provides user interaction functions to enable drilling, filtering, exporting and sharing of multiple styles of reports.

[0053] In this embodiment, the visualization tool is set to be a predetermined tool for creating a visualization interface, such as Tableau, Power BI, and ECharts tools; the preset interface requirements refer to the formulated interface design requirements, such as interface size, color, and font; the data visualization interface refers to an interface designed based on the preset interface requirements for displaying data and analysis results, including various visualization components and tools, such as charts, maps, dashboards, and user interaction functions, such as drilling, filtering, and exporting; the visualization component is set to be used to create a visualization component suitable for displaying different types of data; the multi-style report analysis function refers to the ability to provide multiple types of reports and analysis functions, such as trend analysis, comparative analysis, correlation analysis, cluster analysis, etc.; the user interaction function refers to the function provided for interacting with users, such as user login, data query, report generation, and result export.

[0054] The beneficial effects of the above technical solution are: by visualizing data and providing user interaction functions, the efficiency and quality of data analysis and decision-making can be significantly improved, while promoting user participation, which helps to promote the optimization and continuous improvement of business processes.

[0055] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A unified data integrated management system, characterized in that: include: Storage management module: used to store data from different data sources and use data lake dynamic management; Security management module: used to encrypt data and dynamically classify data access rights; Quality management module: used to track the entire life cycle of data assets and monitor system logs in real time; Visualization module: used to realize data visualization by using set visualization tools.

2. A unified data integrated management system according to claim 1, characterized in that: The storage management module comprises: Warehouse-storage unit: used to establish a target database on a target server according to business requirements, and encrypt and store first data from different data sources into the target database according to a set import method; The target database creates an index for the structured data to provide a fast query function; Data lake-management unit: used to deploy the target data lake using the access block, the target data lake accesses and processes the first data to obtain key data; uses the annotation block to annotate the key data with a sensitivity level; and uses the storage block to store the key data; Access block: used to deploy a target data lake on a target server according to business requirements. The target data lake accesses first data from various data sources and performs data preprocessing on the first data to obtain key data. Annotation block: used to divide the key data according to data characteristics, analyze the sensitivity of the key data, obtain the sensitivity level and make corresponding annotations; Storage block: used to match the corresponding storage policy according to the sensitivity level of key data and store it in the corresponding storage area of ​​the target data lake.

3. A unified data integrated management system according to claim 2, characterized in that: Data sources include various business systems, government data, Internet data, and industry data. Business systems include human resources, finance, supply chain, production, marketing, and customer service.

4. A unified data integrated management system according to claim 2, characterized in that: The marking block includes: After dividing the key data according to data characteristics, first characteristic data is obtained; Acquire a data access record of the first characteristic data, and acquire a first access mode of the current first characteristic data and a first access count within a preset time period from the data access record, and output them as first analysis data; According to the first access mode, determining whether the current first characteristic data has specific frequent accesses, and the corresponding first average access frequency when the specific frequent accesses exist, and outputting them as first analysis data; Determine whether there is specific rare access in the current first characteristic data, and the corresponding second average access frequency when the specific rare access exists, and output it as the first analysis data; Using the set important evaluation index to evaluate the importance of the current first characteristic data, and obtain an important evaluation value; The acquired first analysis data is combined with the important evaluation value to calculate the data sensitivity coefficient; Taking the data sensitivity coefficient as a matching condition, a sensitivity level is matched from a preset sensitivity level mapping table to label the current first characteristic data.

5. A unified data integrated management system according to claim 4, characterized in that: The calculation formula of the data sensitivity coefficient is as follows: ; In the formula, Expressed as the sensitivity coefficient of the current first characteristic data; represents a preset time period; L1 represents the first access count of the current first characteristic data within the preset time period T; ln represents a natural logarithm; e represents a constant, and its value is 2.7; It is expressed as the weight of the impact of access frequency on the sensitivity of the assessed data; The total number of times a specific frequent access is represented as the current first characteristic data; It is represented as the first average access frequency corresponding to the a-th specific frequent access of the first characteristic data, where a=1, 2, 3, , n1; an average value of corresponding first average access frequencies of all specific frequent accesses of the first characteristic data; The total number of specific rare accesses to the current first characteristic data is represented; It is represented as the corresponding second average access frequency of the specific rare access of the first characteristic data for the bth time, where b=1, 2, 3, , n2; an average value of corresponding first average access frequencies of all specific rare accesses of the first characteristic data; It is expressed as the weight of the impact of data importance on the assessment of data sensitivity; It is represented by the importance evaluation value obtained by using the i-th set important evaluation index to evaluate the importance of the current first characteristic data, where i=1, 2, 3; It is represented as the evaluation weight of the importance of the i-th important evaluation indicator to the evaluation data.

6. A unified data integrated management system according to claim 1, characterized in that: The security management module comprises: Transmission encryption unit: used to encrypt key data to be transmitted using a set security protocol; Life management unit: used to dynamically manage the life of key data based on the sensitivity level and data characteristics of the current key data; Access control unit: used to define a preset access permission list for the key data, and based on a pre-established permission analysis model, regularly output new access permissions for the key data, and update the preset access permission list for the current key data based on the new access permissions.

7. A unified data integrated management system according to claim 6, characterized in that: The safety management module and the life management unit include: Retention block: used to determine the initial retention strategy for key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is rarely sensitive or slightly sensitive, the initial retention policy is output as the first retention policy; If the sensitivity level of the current key data is moderately sensitive or severely sensitive, the corresponding sensitivity impact coefficient of the current sensitivity level is obtained, and the initial retention period extracted from the initial retention policy is adjusted to obtain the first retention period; After replacing the initial retention period in the initial retention policy with the first retention period, the first retention policy is outputted as the first retention policy; Performing data retention management on current key data using the first retention policy; Encryption block: when the sensitivity level of key data is mildly sensitive, moderately sensitive or severely sensitive, the first retention period is extracted from the first retention policy currently applied to the key data, and combined with the real-time storage duration of the key data for analysis to determine the first life stage coefficient of the current key data; The first life stage coefficient is calculated by combining the sensitivity impact coefficient corresponding to the sensitivity level of the current key data to obtain the encryption analysis coefficient; Using the encryption analysis coefficient, obtaining the encryption analysis level of the current key data; Determine the expected encryption scope and expected encryption algorithm for the current key data based on the obtained encryption analysis level; The expected encryption range of the current key data is encrypted using the expected encryption algorithm; Destruction block: used to determine the initial destruction strategy of key data based on the data characteristics of the current key data; If the sensitivity level of the current key data is extremely low sensitivity, the initial destruction strategy is output as the first destruction strategy; If the sensitivity level of the current key data is slightly sensitive, moderately sensitive, or severely sensitive, the first destruction time is obtained using the corresponding sensitivity impact coefficient of the sensitivity level of the current key data; The initial destruction time in the initial destruction strategy is replaced with the first destruction time and then output as the first destruction strategy; The first destruction strategy is used to manage data destruction of current key data.

8. A unified data integrated management system according to claim 1, characterized in that: The quality management module comprises: Asset lifecycle management unit: used to select metadata management tools that match business needs; Use the collection tool in the metadata management tool to collect all metadata information in the target database; Combine the acquired metadata information with the pre-established asset lifecycle management model to record and track the entire lifecycle of data assets and generate asset lifecycle management reports; Log monitoring unit: used to collect system log entries in real time using a set log collection tool, pre-process and aggregate the system log entries, obtain aggregated logs and store them accordingly; According to business needs, select a log monitoring tool that is suitable for the current system, and configure log anomaly analysis rules and corresponding alarm conditions in the log monitoring tool; Use the currently set log monitoring tool to monitor system logs and abnormal alarms in real time.

9. A unified data integrated management system according to claim 1, characterized in that: The visualization module comprises: Visualization unit: used to establish a connection with the target database using a set visualization tool, and then create a data visualization interface for the current system according to the preset interface requirements; In the data visualization interface, visualization components are set to provide various report analysis functions for different business types of data; Provides user interaction functions to enable drilling, filtering, exporting and sharing of multiple styles of reports.

Citation Information

Cited By

  • Database management system with data stream security and real-time data protection

    CN120763216A