Data asset management system and method

By collecting data asset information from enterprise databases, file systems, and API data sources in real time, performing label structure sub-column division and time series synchronization processing, combining usage frequency and sensitivity assessment, dynamically controlling access rights, and formulating full life cycle management rules, the problem of low efficiency in traditional data asset management is solved, and efficient and secure data asset management is achieved.

CN119599255BActive Publication Date: 2025-09-12GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411553016.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-09-12
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

Traditional data asset management methods cannot efficiently handle large-scale and dynamically changing data assets, resulting in inefficient management.

Method used

By collecting data asset information in real time from enterprise databases, file systems, and API data sources, performing tag structure sub-column division and time series synchronization processing, combining usage frequency and sensitivity assessment, dynamically controlling access rights, and formulating full life cycle management rules, automated management and reporting feedback are achieved.

Benefits of technology

Ensure the real-time and accuracy of data asset management, improve data retrieval efficiency, identify high-value data assets, dynamically adjust access strategies, provide full life cycle operational guidance, reduce human errors, and improve management efficiency and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119599255B_ABST
    Figure CN119599255B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data asset management technology, and in particular to a data asset management system and method. The method comprises the following steps: by collecting data asset information sets from different data sources in real time and dividing them into label structure sub-columns to obtain data asset information label description structured sub-columns, while creating time series synchronization processing and integrating and storing them in a data asset management database; performing usage frequency and sensitivity evaluation analysis on each data asset sub-information in the data asset management database, and performing data asset access utility evaluation analysis to obtain the data access utility of each data asset sub-information; performing dynamic access permission control and full life cycle management design on each data asset sub-information in the data asset management database, and performing automated management and report feedback recording to generate a data asset information management report. The present invention can improve the transparency, security and efficiency of data asset management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data asset management, and in particular to a data asset management system and method. Background Art

[0002] Data assets have become core assets in modern enterprises and organizations, and the complexity and importance of their management are constantly increasing. Machine learning algorithms can be used to automatically classify and label data assets. Through in-depth analysis of data content and attributes, data types and sensitivity levels can be automatically identified, improving the accuracy and efficiency of data classification. Dynamic storage management technologies can also be introduced to automatically adjust storage resources based on the frequency and importance of data assets. Real-time monitoring of data access and storage requirements can be used to optimize storage space allocation and utilization. Furthermore, real-time data monitoring systems can be deployed to continuously track the use and changes of data assets. Intelligent analysis and early warning mechanisms can promptly identify potential anomalies and risks, automatically generate alerts, and provide corrective action suggestions. However, traditional data asset management methods often rely on static rules and manual operations, which are unable to effectively handle large-scale and dynamically changing data assets, resulting in inefficient data asset management. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a data asset management system and method to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a data asset management method includes the following steps:

[0005] Step S1: Data asset information sets are collected in real time from enterprise databases, file systems, and API data sources, and the data asset information sets are divided into tag structure sub-columns to obtain data asset information tag description structured sub-columns; each data asset information tag description structured sub-column is created and synchronized, and integrated and stored in the data asset management database within the data processing unit;

[0006] Step S2: Performing a usage frequency and sensitivity assessment analysis on each piece of data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each piece of data asset sub-information; performing a data asset access utility assessment analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity of each piece of data asset sub-information to obtain the data access utility of each piece of data asset sub-information;

[0007] Step S3: Dynamically control the access rights of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information to obtain dynamic access control rights of the data asset management database;

[0008] Step S4: Perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database to generate full life cycle management rules for the data asset information; perform automated management and report feedback records for the corresponding data asset information in the data asset management database based on the full life cycle management rules for the data asset information to generate a data asset information management report.

[0009] Furthermore, step S1 includes the following steps:

[0010] Step S11: obtaining a database source data asset information subset by collecting a data asset-related information subset from an enterprise database data source in real time;

[0011] Step S12: obtaining a file system source data asset information subset by collecting a data asset-related information subset from the file system data source in real time;

[0012] Step S13: obtaining an API source data asset information subset by collecting a data asset-related information subset from the API data source in real time;

[0013] Step S14: performing information set processing on the database source data asset information subset, the file system source data asset information subset, and the API source data asset information subset to obtain a data asset information set; dividing the data asset information set into label structure sub-columns to obtain data asset information label description structured sub-columns;

[0014] Step S15: Create a time series synchronization process for each data asset information tag description structured sub-column and integrate and store it in the data asset management database.

[0015] Furthermore, the division of the data asset information set into tag structure sub-columns in step S14 includes the following steps:

[0016] Performing information content semantic analysis on each piece of data asset information in the data asset information set to obtain content semantic data corresponding to the data asset content of each piece of data asset information;

[0017] Perform data format attribute mining and analysis on each data asset information in the data asset information set to obtain the data format attribute of each data asset information corresponding to the data asset content;

[0018] Based on the data format attributes of the data asset content corresponding to each data asset information, the content semantic data of the data asset content corresponding to each data asset information is classified and labeled to obtain various data content attribute labels corresponding to the data asset content, wherein the various data content attribute labels include creation time, data type, storage location, data usage and modification time;

[0019] Divide each data asset information in the data asset information set into an attribute label description according to various data content attribute labels corresponding to the data asset content, and obtain a data asset information content attribute label description sub-column;

[0020] Perform structural conversion on each data asset content in the data asset information content attribute label description sub-column to obtain a data asset information label description structured sub-column.

[0021] Furthermore, step S15 includes the following steps:

[0022] Step S151: Perform creation time label sub-column screening processing on each data asset information label description structured sub-column to obtain a data asset information creation time label structured sub-column and data asset information remaining label structured sub-columns;

[0023] Step S152: performing creation time sequence synchronization processing on the remaining tag structured sub-columns of the data asset information based on the data asset information creation time tag structured sub-column, so as to obtain various data asset information tag structured sub-columns under the same creation time sequence dimension;

[0024] Step S153: Designing a data integration storage format for each data asset information tag structured sub-column under the same creation time sequence dimension through the data asset management database within the data processing unit to obtain a data asset information tag sub-column integration storage format specification;

[0025] Step S154: Integrate and store each data asset information tag structured sub-column under the same creation time sequence dimension into the data asset management database in the data processing unit according to the data asset information tag sub-column integration storage format specification.

[0026] Furthermore, step S2 includes the following steps:

[0027] Step S21: Perform access log recording processing on each data asset sub-information in the data asset management database to obtain data access log record information of each data asset sub-information;

[0028] Step S22: Analyze the access count and time distribution of the data access log record information of each data asset sub-information to obtain the data access count and data access time distribution of each data asset sub-information;

[0029] Step S23: performing a usage frequency statistical analysis on the number of data accesses to each data asset sub-information based on the data access time distribution of each data asset sub-information to obtain the data usage frequency of each data asset sub-information;

[0030] Step S24: Using the data asset information sensitivity calculation formula, perform information sensitivity evaluation calculation on each data asset sub-information in the data asset management database to obtain the data sensitivity level of each data asset sub-information;

[0031] Step S25: Based on the data usage frequency and data sensitivity of each data asset sub-information, a data asset access utility evaluation analysis is performed on each data asset sub-information corresponding to the data asset management database to obtain the data access utility of each data asset sub-information.

[0032] Furthermore, the data asset information sensitivity calculation formula in step S24 is specifically:

[0033]

[0034] Where S i is the data sensitivity of the i-th data asset sub-information, n is the total number of data asset sub-information in the data asset management database, i is the item index parameter of the data asset sub-information, R i is the data asset availability parameter of the i-th data asset sub-information, α is the data asset availability impact weight coefficient, D i is the leakage risk measurement parameter of the i-th data asset sub-information, β is the data asset leakage risk impact weight coefficient, γ is the control leakage risk impact exponential attenuation coefficient, C i is the data asset confidentiality parameter of the i-th data asset sub-information, δ is the data asset confidentiality impact weight coefficient, m is the total number of additional attribute information in the data asset sub-information, j is the item measurement parameter of the additional attribute information, P i,j is the specific value of the i-th data asset sub-information in the j-th additional attribute information, κ j is the influence weight coefficient of the jth additional attribute information, and η is the correction coefficient of the data sensitivity.

[0035] Furthermore, step S25 includes the following steps:

[0036] Step S251: classifying the data sensitivity level of each data asset sub-information in the data asset management database based on the data sensitivity level of each data asset sub-information to obtain the data sensitivity level of each data asset sub-information;

[0037] Step S252: performing access utility weighted analysis on each data asset sub-information corresponding to the data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity level of each data asset sub-information to obtain a data access utility weighted value set for each data asset sub-information;

[0038] Step S253: performing data asset access utility evaluation and analysis on each data asset sub-information corresponding to the data asset management database according to the data access utility weighted value set of each data asset sub-information to obtain the data access utility of each data asset sub-information.

[0039] Furthermore, step S3 includes the following steps:

[0040] Step S31: Performing a management role access business requirement analysis on each data asset sub-information in the data asset management database to obtain the database management role access business requirement for each data asset sub-information;

[0041] Step S32: performing data usage constraint analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency of each piece of data asset sub-information to obtain data usage constraint conditions for each piece of data asset sub-information;

[0042] Step S33: performing access utility restriction analysis on each data asset sub-information in the data asset management database based on the data access utility of each data asset sub-information, and obtaining a data access utility restriction range for each data asset sub-information;

[0043] Step S34: Dynamically control the access rights of the database management role to each data asset sub-information based on the data usage constraints and data access utility limitation range of each data asset sub-information to obtain dynamic access control rights for the data asset management database.

[0044] Furthermore, step S4 includes the following steps:

[0045] Step S41: performing management role authority matching and mapping on the dynamic access control authority of the data asset management database to obtain the data asset management role dynamic access matching and mapping control authority;

[0046] Step S42: Based on the data asset management role dynamic access matching mapping control authority, the corresponding data asset information in the data asset management database is analyzed for full life cycle stage authority characteristics to obtain a data asset full life cycle stage management role authority characteristic set;

[0047] Step S43: Formulate full life cycle management rules for the data asset full life cycle management role authority feature set to generate data asset information full life cycle management rules;

[0048] Step S44: performing automated management processing on the corresponding data asset information in the data asset management database based on the data asset information full life cycle management rules to obtain the data asset information full life cycle automatic management results;

[0049] Step S45: Record the management report feedback of the automatic management results of the data asset information throughout its life cycle to generate a data asset information management report.

[0050] Furthermore, the present invention also provides a data asset management system for executing the above-mentioned data asset management method, the data asset management system comprising:

[0051] The data asset information structured storage module is used to collect data asset information sets from databases, file systems, and API data sources in real time, and divide the data asset information sets into tag structure sub-columns to obtain data asset information tag description structured sub-columns; create time series synchronization processing for each data asset information tag description structured sub-column and integrate and store them in the data asset management database;

[0052] The usage frequency and access utility analysis module is used to perform usage frequency and sensitivity evaluation analysis on each data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each data asset sub-information; based on the data usage frequency and data sensitivity of each data asset sub-information, the data asset access utility evaluation analysis is performed on each data asset sub-information in the data asset management database to obtain the data access utility of each data asset sub-information;

[0053] A database dynamic access permission control module is used to dynamically control the access permission of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information, thereby obtaining dynamic access control permissions for the data asset management database;

[0054] The data asset full life cycle management module is used to perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database, so as to generate data asset information full life cycle management rules; based on the data asset information full life cycle management rules, the corresponding data asset information in the data asset management database is automatically managed and report feedback records are recorded to generate data asset information management reports.

[0055] Beneficial effects of the present invention:

[0056] 1. The data asset management method proposed in the present invention has the beneficial effect of ensuring the real-time and accuracy of data assets by collecting a subset of information related to data assets from an enterprise database data source. Through this method, the latest database content can be obtained, ensuring that the data reflected in the management system is the latest and most timely, which can reduce the risk of information lag and make data asset management more reliable and efficient. By collecting a subset of information related to data assets from a file system data source in real time, data assets stored in the file system can be effectively monitored and managed. As one of the main forms of data storage, the file system contains a large number of files and documents. These files often carry important business information and historical records. Real-time collection can ensure that the management of data assets in the file system is up-to-date and promptly reflects the addition, modification or deletion of files, thereby providing basic data guarantee for subsequent processing. By collecting a subset of information related to data assets from an API data source in real time, data from external systems or third-party services can be obtained. This method enables effective data integration with other systems or platforms, enhances system interoperability and data richness, and real-time collection of API data not only provides the latest external data, but also helps understand the impact of the external environment on its own data assets. At the same time, by processing the subsets of data asset information collected from different data sources into information sets, data asset information from multiple sources can be comprehensively summarized. This step, through the processing of information sets, generates a unified data asset information set, which is of great significance for the comprehensive management and analysis of data. It also structures the data asset information through the label structure sub-column division, assigning each data asset label to make its description clearer and more orderly. This structured description method helps to quickly identify and classify data assets and improve the efficiency of data retrieval. By creating time series synchronization processing for each data asset information label description structured sub-column and integrating and storing it in the data asset management database, it helps to maintain the temporal consistency and system consistency of data asset information. The time series synchronization processing ensures the consistency of data at different time points and avoids data inconsistencies caused by time differences. This step, by integrating and storing the structured label description information in the data asset management database, provides a reliable data foundation for subsequent data analysis, report generation, and decision support, thereby enabling the efficient processing of large-scale and dynamically changing data asset information, while supporting more accurate business decisions and strategic planning. Secondly, by conducting a usage frequency statistical analysis on each data asset sub-information in the data asset management database, this process aims to combine the number of data accesses with the time pattern to evaluate the usage frequency of data assets. This analysis can identify the actual usage of data assets and determine which data assets are frequently used and which are less frequently used.Through statistical analysis, managers can understand the actual needs of data and make more accurate resource allocation decisions. In addition, by conducting a corresponding sensitivity assessment on each data asset sub-information within the data asset management database, the sensitivity of each data asset is determined. This assessment takes into account factors such as data content, access rights, and potential risks, thereby determining the confidentiality and security requirements of the data. Through sensitivity assessment, it is possible to ensure that highly sensitive data assets receive appropriate protection measures to avoid damage caused by data leakage or misuse. Based on the data usage frequency and data sensitivity of each data asset sub-information within the data asset management database, a data asset access utility assessment analysis is also conducted. The purpose of this process is to comprehensively consider the frequency of data use and sensitivity and evaluate the overall utility of data assets. This assessment can identify the most valuable data assets in the management and use process, thereby improving the overall effectiveness of data asset management. Then, by dynamically controlling the access rights of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information, the access rights of the database management role can be dynamically controlled by comprehensively considering the data usage frequency and data access utility. This process allows real-time adjustment and optimization of data access rights based on actual business needs and data access conditions. The dynamic access rights control processing process can flexibly adjust access policies according to changes in data usage and utility to ensure data security and effectiveness. This method can effectively respond to changes in dynamic business environments, reduce the limitations of static permission configuration, and thus provide more accurate and timely data asset protection. Finally, by conducting full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database, the design and formulation of corresponding management rules can provide operational guidance and control standards for each stage of the data life cycle, ensuring the consistency and standardization of data asset management. By formulating clear management rules, data leakage or compliance risks caused by improper permission settings or non-standard operations can be avoided. This step helps to establish a systematic management framework, making data asset management more transparent and controllable, thereby improving the overall efficiency and effectiveness of data asset management, and providing guarantees for the compliance and security of the enterprise.By automating the management of the corresponding data asset information in the data asset management database based on the full life cycle management rules of data asset information, efficient and continuous monitoring and control of data assets can be achieved. Automated processing not only reduces human operational errors, but also responds to changes in data management needs in real time, improving management flexibility and accuracy. This step makes the management of data assets throughout the entire life cycle more efficient, automatically executing rules, updating and maintaining data status in real time, thereby reducing manual intervention and improving the reliability and efficiency of overall data asset management. This automated management result also facilitates comprehensive analysis and report generation of data assets. In addition, by reporting and recording the data asset management results obtained after previous automated management, systematic management reports can be formed. These reports not only record the results of automated management, but also provide a comprehensive evaluation of the management process and results. This process helps to comprehensively review all aspects of data asset management, identify existing problems and areas for improvement, and the generated reports can provide important reference for decision makers, help optimize data management strategies and processes, and thus improve the management effect of data assets and the scientific nature of business decisions.

[0057] 2. The data asset management system proposed in the present invention is composed of a data asset information structured storage module, a usage frequency and access utility analysis module, a database dynamic access permission control module and a data asset full life cycle management module. It can implement any data asset management method described in the present invention, and is used to combine the operations between computer programs running on each module to implement the data asset management method. The internal structure of the system cooperates with each other, which can greatly reduce duplication of work and manpower investment, and can quickly and effectively provide a more accurate and efficient data asset management process, thereby simplifying the operation process of the data asset management system. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments thereof made with reference to the following drawings:

[0059] Figure 1 Schematic diagram of the steps of the data asset management method of the present invention;

[0060] Figure 2 for Figure 1 Detailed step flow diagram of step S1;

[0061] Figure 3 for Figure 2 Detailed step flow chart of step S15. DETAILED DESCRIPTION

[0062] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0063] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0064] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0065] To achieve this, please refer to Figures 1 to 3 The present invention provides a data asset management method, which includes the following steps:

[0066] Step S1: Data asset information sets are collected in real time from enterprise databases, file systems, and API data sources, and the data asset information sets are divided into tag structure sub-columns to obtain data asset information tag description structured sub-columns; each data asset information tag description structured sub-column is created and synchronized, and integrated and stored in the data asset management database within the data processing unit;

[0067] Step S2: Performing a usage frequency and sensitivity assessment analysis on each piece of data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each piece of data asset sub-information; performing a data asset access utility assessment analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity of each piece of data asset sub-information to obtain the data access utility of each piece of data asset sub-information;

[0068] Step S3: Dynamically control the access rights of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information to obtain dynamic access control rights of the data asset management database;

[0069] Step S4: Perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database to generate full life cycle management rules for the data asset information; perform automated management and report feedback records for the corresponding data asset information in the data asset management database based on the full life cycle management rules for the data asset information to generate a data asset information management report.

[0070] In the embodiment of the present invention, please refer to Figure 1 FIG. 1 is a flow chart of the steps of the data asset management method of the present invention. In this example, the data asset management method includes the following steps:

[0071] Step S1: Data asset information sets are collected in real time from enterprise databases, file systems, and API data sources, and the data asset information sets are divided into tag structure sub-columns to obtain data asset information tag description structured sub-columns; each data asset information tag description structured sub-column is created and synchronized, and integrated and stored in the data asset management database within the data processing unit;

[0072] In an embodiment of the present invention, information data related to data assets is collected in real time from an enterprise database, and the relevant data asset information is extracted from the target database table using the SQL query language by configuring the database connection interface. This process obtains information related to data assets, such as the name of the database table, field type, number of records, etc., by executing specific SELECT statements. These query statements must accurately reflect the structure and content of the data assets, thereby obtaining a subset of the database source data asset information. By using a corresponding file system retrieval tool (such as Python's os module or Java's NIO), information data related to data assets is collected from the file system data source in real time, and the specified file directory is traversed according to the screening criteria of the file path and file type to identify qualified files. For each file, the file metadata (such as file name, creation date, size, etc.) and file content (such as the text or structured data of the file) are read to retrieve and collect the subset of the file system source data asset information. In addition, by collecting information data related to data assets from the corresponding API data source in real time, configuring the parameters of the API request, including the request method (such as GET or POST), request header and request body, and initiating the API request by using an HTTP client library (such as Python's Requests library), obtaining the returned JSON or XML data, parsing the response data to extract information related to the data asset, such as asset ID, attribute description, etc., thereby integrating and obtaining the API source data asset information subset. At the same time, by integrating the data asset information subsets of different data sources previously collected in real time from different data sources, the information subsets obtained from different data sources are standardized and formatted, and by using data integration tools (such as Apache NiFi or Talend) to merge data from different sources, and remove duplicate and redundant information, and by using data fusion algorithms (such as data matching and merging technology), the various information subsets are integrated to generate a unified data asset information set, thereby obtaining a data asset information set. We also define label classification standards based on the actual content and structure of data assets, and use data annotation tools (such as Labelbox or custom scripts) to assign labels to data asset information sets, divide data asset information into different label categories (such as data asset type, data format, storage location, creation time, and modification time, etc.), and subdivide the label structure into sub-columns of different label categories to ensure the accuracy and effectiveness of label division, thereby obtaining structured sub-columns describing data asset information labels.Then, by performing time series synchronization processing on the structured sub-columns described by the data asset information tags obtained previously according to the corresponding creation time, a time series data processing framework (such as Apache Kafka and Apache Flink) is established to achieve creation timestamp marking of each structured sub-column, and by using time series data synchronization algorithms (such as time window technology) to process data, the time series consistency and integrity of the structured sub-columns described by the data asset information tags are ensured. At the same time, the data after time series synchronization processing is integrated and stored in the data asset management database (such as PostgreSQL or MongoDB), and a data backup and recovery mechanism is set up to ensure long-term stable storage and access of the data.

[0073] Step S2: Performing a usage frequency and sensitivity assessment analysis on each piece of data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each piece of data asset sub-information; performing a data asset access utility assessment analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity of each piece of data asset sub-information to obtain the data access utility of each piece of data asset sub-information;

[0074] In an embodiment of the present invention, a detailed access log is generated in the data asset management database for each data asset sub-information access behavior using a logging tool (such as ELK Stack or Splunk). Each access event to the data asset sub-information is recorded to obtain data access log information for each data asset sub-information. The data access log information of each data asset sub-information is analyzed using a data analysis tool (such as Apache Hadoop or the Pandas library in Python). The access frequency of each data asset sub-information is quantitatively calculated by statistically analyzing the number of accesses to each data asset sub-information, thereby obtaining the data usage frequency of each data asset sub-information. At the same time, information sensitivity assessment calculations are performed in combination with relevant parameters of the data asset sub-information to quantitatively calculate the corresponding sensitivity value, thereby obtaining the data sensitivity degree of each data asset sub-information. Then, by combining the data usage frequency and data sensitivity of each data asset sub-information obtained from the previous statistical analysis, a weighted scoring method is used to evaluate the access utility of each data asset sub-information in the data asset management database, so as to calculate the scores of the two dimensions according to the usage frequency and sensitivity respectively, and combine the scores of the two dimensions to obtain the final access utility score through the weighted average method. This step is usually automatically calculated using data analysis tools or custom scripts. For example, the corresponding calculation formula can be: Where, E kScore the final access utility of the i-th data asset sub-information, F i is the data usage frequency of the i-th data asset sub-information, a is the usage frequency weight coefficient, S i is the data sensitivity of the i-th data asset sub-information, and b is the sensitivity weight coefficient, which ensures the accuracy and objectivity of the evaluation results, and finally obtains the data access utility of each data asset sub-information.

[0075] Step S3: Dynamically control the access rights of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information to obtain dynamic access control rights of the data asset management database;

[0076] In an embodiment of the present invention, a management role access business requirement analysis is performed on each data asset sub-information in the data asset management database, so as to collect the management role and corresponding business requirement of each data asset sub-information. This process is completed by accessing business requirement documents, interviewing management personnel, and analyzing existing permission settings. The management role information is extracted by using a permission management tool (such as LDAP or Active Directory), and the access requirement of each role to the data asset is recorded in detail in combination with a business requirement analysis tool (such as a business process modeling tool BPMN). In addition, a data usage constraint analysis is performed on the corresponding data asset sub-information in the data asset management database by combining the data usage frequency of each data asset sub-information obtained in the previous analysis. Based on the analysis results, data usage constraint conditions are defined, such as implementing stricter access control for frequently accessed sensitive data. The usage frequency data is combined with business requirements to set reasonable data usage restrictions. For example, if the access frequency of a data asset sub-information is high, it is necessary to limit the access time period or number of accesses to the data. In addition, the access utility restriction analysis is performed on the corresponding data asset sub-information in the data asset management database by combining the data access utility of each data asset sub-information obtained in the previous quantitative calculation. BI) generates an access utility chart to determine the utility range of data access. Then, by combining the data usage constraints and data access utility restriction range obtained from the previous analysis, the management role access business needs of each data asset sub-information obtained from the previous analysis are dynamically controlled to control access rights. By using a dynamic permission control system (such as an RBAC-based permission management system or an ABAC-based access control policy), the data usage constraints and access utility restriction range are combined with the business needs of the management role to create a dynamic access control policy. The implementation process includes configuring the policy rules of the permission management system and dynamically adjusting the role permissions to meet the requirements of data usage and access utility. For example, for a highly sensitive and high access utility data asset, permissions will be automatically restricted to only necessary management roles, and stricter access auditing and control measures will be implemented to ultimately obtain dynamic access control permissions for the data asset management database.

[0077] Step S4: Perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database to generate full life cycle management rules for the data asset information; perform automated management and report feedback records for the corresponding data asset information in the data asset management database based on the full life cycle management rules for the data asset information to generate a data asset information management report.

[0078] In an embodiment of the present invention, by matching and mapping the dynamic access control permissions in the data asset management database obtained by the previous analysis with the permissions of the management role, the access permission configuration of each management role is extracted from the permission management system, and these permissions are mapped to the corresponding data asset information, and a permission mapping tool (such as a role mapping module in an RBAC system) is used for mapping, the permissions of each management role are matched with the access requirements of the data asset, and the dynamic access permissions of each role to a specific data asset are determined. The dynamic access permissions of each role to a specific data asset are analyzed for the characteristics of the full life cycle permissions of the corresponding data asset information in the data asset management database by combining the dynamic access control permissions obtained by the previous mapping assignment, so as to perform the full life cycle permissions on the data asset at each life cycle stage (such as creation, storage, use) through a life cycle management tool (such as an enterprise data management system). , sharing, and destruction), record the management roles and their permission characteristics involved in each stage, and analyze the permission requirements and restrictions of the roles at different life cycle stages. In addition, full life cycle management rules are formulated based on the permission requirements and restrictions of the roles at different life cycle stages obtained in the previous analysis. By using a rule engine (such as Drools) or a policy formulation tool to create full life cycle management rules for data asset information, the permission characteristics obtained in the previous analysis are converted into management rules. For example, access control policies and operation requirements are defined for different life cycle stages of data. The formulated rules include how to implement permission control, monitoring and audit requirements at different stages of the data to ensure that the management of data assets throughout the life cycle complies with the set policies and security standards, thereby formulating and generating corresponding full life cycle management rules for data asset information. Then, by combining the full life cycle management rules for data asset information obtained in the previous analysis, the corresponding data asset information in the data asset management database is automatically managed and processed, so that the management rules are applied to data assets by utilizing an automated management system (such as a data governance platform or an automated workflow management system), and data access control, permission allocation, data auditing and compliance checks are automatically performed according to the rules to obtain full life cycle automatic management results. Finally, by using report generation tools (such as Microsoft Power BI or Tableau), feedback records of management reports are made on the full life cycle automatic management results obtained after previous automated management, and report generation tools (such as Microsoft Power BI or Tableau) are used to summarize the automatic management processing results and generate detailed management reports. The report content includes the implementation status of automated management, the effectiveness of authority control, the management effect of the life cycle stage, etc. Feedback records are generated regularly through the reporting system and provided to management for decision-making reference. The report also includes performance indicators and improvement suggestions for data asset management to help optimize the management process of data assets. Finally, the records are generated to generate corresponding data asset information management reports.

[0079] Furthermore, step S1 includes the following steps:

[0080] Step S11: obtaining a database source data asset information subset by collecting a data asset-related information subset from an enterprise database data source in real time;

[0081] Step S12: obtaining a file system source data asset information subset by collecting a data asset-related information subset from the file system data source in real time;

[0082] Step S13: obtaining an API source data asset information subset by collecting a data asset-related information subset from the API data source in real time;

[0083] Step S14: performing information set processing on the database source data asset information subset, the file system source data asset information subset, and the API source data asset information subset to obtain a data asset information set; dividing the data asset information set into label structure sub-columns to obtain data asset information label description structured sub-columns;

[0084] Step S15: Create a time series synchronization process for each data asset information tag description structured sub-column and integrate and store it in the data asset management database.

[0085] As an embodiment of the present invention, refer to Figure 2 As shown, Figure 1 Detailed step flow diagram of step S1 in FIG. 1 , in this embodiment, step S1 includes the following steps:

[0086] Step S11: obtaining a database source data asset information subset by collecting a data asset-related information subset from an enterprise database data source in real time;

[0087] In an embodiment of the present invention, information data related to data assets is collected in real time in an enterprise database, and the database connection interface is configured to extract relevant data asset information from the target database table using the SQL query language. This process obtains information related to data assets, such as the name of the database table, field type, number of records, etc., by executing specific SELECT statements. These query statements must accurately reflect the structure and content of the data assets and are executed regularly through scheduling tools (such as Apache Airflow) to ensure the real-time and accuracy of data collection, and ultimately obtain a subset of database source data asset information.

[0088] Step S12: obtaining a file system source data asset information subset by collecting a data asset-related information subset from the file system data source in real time;

[0089] In an embodiment of the present invention, information data related to data assets is collected in real time from a file system data source by using a corresponding file system retrieval tool (such as Python's os module or Java's NIO), so as to traverse a specified file directory according to the screening criteria of file path and file type, identify qualified files, and for each file, by reading the file's metadata (such as file name, creation date, size, etc.) and file content (such as the file's text or structured data), finally retrieve and collect a subset of the file system source data asset information.

[0090] Step S13: obtaining an API source data asset information subset by collecting a data asset-related information subset from the API data source in real time;

[0091] In an embodiment of the present invention, information data related to data assets is collected in real time from the corresponding API data source, and the parameters of the API request are configured, including the request method (such as GET or POST), request header and request body, and an API request is initiated by using an HTTP client library (such as Python's Requests library), and the returned JSON or XML data is obtained. The response data is parsed to extract information related to the data assets, such as asset ID, attribute description, etc. To ensure the real-time nature of data collection, a scheduled task or event triggering mechanism should be set to automatically trigger the API call, and finally the API source data asset information subset is integrated.

[0092] Step S14: performing information set processing on the database source data asset information subset, the file system source data asset information subset, and the API source data asset information subset to obtain a data asset information set; dividing the data asset information set into label structure sub-columns to obtain data asset information label description structured sub-columns;

[0093] In an embodiment of the present invention, a database source data asset information subset, a file system source data asset information subset, and an API source data asset information subset previously collected in real time from different data sources are subjected to data integration processing to standardize and format the information subsets obtained from different data sources, and data integration tools (such as Apache NiFi or Talend) are used to merge data from different sources and remove duplicate and redundant information. Through data fusion algorithms (such as data matching and merging technology), the various information subsets are integrated to generate a unified data asset information set. This process needs to consider the compatibility and consistency of the data to ensure the accuracy and completeness of the final information set, thereby obtaining a data asset information set. At the same time, by defining a label classification standard based on the actual content and structure of the data asset, and by using a data annotation tool (such as Labelbox or a custom script) to assign labels to the data asset information set, the data asset information is divided into different label categories (such as data asset type, data format, storage location, creation time, and modification time, etc.), and by analyzing the attributes and relevance of the data, the label structure is subdivided into sub-columns of different label categories to ensure the accuracy and effectiveness of the label division, and finally a data asset information label description structured sub-column is obtained.

[0094] Step S15: Create a time series synchronization process for each data asset information tag description structured sub-column and integrate and store it in the data asset management database.

[0095] In an embodiment of the present invention, each structured subcolumn described by the data asset information tag obtained by the previous division is synchronized according to the corresponding creation time, so as to establish a time series data processing framework (such as Apache Kafka and Apache Flink) to realize the creation timestamp marking of each structured subcolumn, and process the data by using a time series data synchronization algorithm (such as time window technology) to ensure the time series consistency and integrity of the structured subcolumns described by the data asset information tag. At the same time, the data after time series synchronization processing is integrated and stored in a data asset management database (such as PostgreSQL or MongoDB), and a data backup and recovery mechanism is set to ensure long-term stable storage and access of the data.

[0096] Furthermore, the division of the data asset information set into tag structure sub-columns in step S14 includes the following steps:

[0097] Performing information content semantic analysis on each piece of data asset information in the data asset information set to obtain content semantic data corresponding to the data asset content of each piece of data asset information;

[0098] In an embodiment of the present invention, natural language processing (NLP) technology is used to perform information content semantic recognition and analysis on each data asset information in a previously integrated data asset information set, so as to input the data asset information into a semantic analysis model. The model includes modules such as word segmentation, part-of-speech tagging, and named entity recognition. For example, pre-trained language models such as BERT or GPT-3 are used to extract keywords, entities, and the relationships between them in the data asset content through these models, thereby obtaining content semantic data for each data asset information. These content semantic data can reflect the actual semantic meaning of the data asset, forming semantic labels containing topics, contexts, and related concepts, and ultimately obtaining content semantic data for the data asset content corresponding to each data asset information.

[0099] Preferably, data format attribute mining and analysis is performed on each data asset information in the data asset information set to obtain the data format attribute of the data asset content corresponding to each data asset information;

[0100] In an embodiment of the present invention, data format attributes of each data asset information in a data asset information set are mined and analyzed, so that during the operation, the information content of each data asset is formatted and its data structure and format characteristics are identified. The data is read and parsed by using a data format detection tool (such as the pandas library in Python or Apache Avro) to parse and obtain the format attributes of the data asset information, such as data type, field length, data structure, etc., and record them as data format attributes, and finally obtain the data format attributes of the data asset content corresponding to each data asset information.

[0101] Preferably, based on the data format attributes of the data asset content corresponding to each data asset information, the content semantic data of the data asset content corresponding to each data asset information is subjected to data asset content attribute classification and labeling processing to obtain various data content attribute labels corresponding to the data asset content, wherein the various data content attribute labels include creation time, data type, storage location, data usage, and modification time;

[0102] In an embodiment of the present invention, the content attributes of the content semantic data of the data asset content corresponding to each data asset information are classified and labeled by combining the data format attributes of the data asset content corresponding to each data asset information obtained by previous mining and analysis, so as to classify and label the data asset content by combining the semantic data obtained by previous analysis with the format attributes obtained by mining and analysis, and analyze the data asset content by using a label classification algorithm (such as a decision tree, random forest and other machine learning algorithms) to identify the attributes such as the creation time, data type, storage location, data usage and modification time of the data asset content. For example, the data asset is trained using a classification model to obtain label categories and corresponding attribute information, and the data content is labeled and classified according to these attributes, and finally various data content attribute labels corresponding to the data asset content are obtained, wherein the various data content attribute labels include creation time, data type, storage location, data usage and modification time.

[0103] Preferably, each piece of data asset information in the data asset information set is divided into attribute label descriptions according to various data content attribute labels corresponding to the data asset content, to obtain a data asset information content attribute label description sub-column;

[0104] In an embodiment of the present invention, attribute label descriptions are divided for each corresponding data asset information in the data asset information set by combining various data content attribute labels corresponding to the data asset content obtained in the previous analysis, so that the previously generated attribute labels are applied to the data asset information set, and each data asset is divided into different label description sub-columns according to its information content. During the implementation process, attribute labels are assigned to each data asset through a label description generation tool (such as Excel's filtering function or a custom script), and the data asset information under the same attribute label is divided into corresponding sub-columns. For example, corresponding label description sub-columns are generated for labels such as creation time and data type, so that the information content of each data asset can be organized and classified according to these labels, forming a detailed attribute label description, and finally obtaining a data asset information content attribute label description sub-column.

[0105] Preferably, each data asset content in the data asset information content attribute tag description sub-column is structurally converted to obtain a data asset information tag description structured sub-column.

[0106] In an embodiment of the present invention, a data conversion tool (such as the DataFrame structure of the pandas library in Python) is used to perform data structuring conversion processing on each data asset content in the attribute label description subcolumn of the data asset information content obtained previously through division, so as to convert the attribute label description subcolumn obtained previously through division into a structured data format, and by mapping the information in the attribute label description subcolumn according to the data field, a normalized data table is generated. For example, the label attribute fields such as creation time and data type are used as column headers, and the information of each data asset is filled in the corresponding column by row. This process can convert the original label description into a structured data format, and finally obtain a structured subcolumn of data asset information label description.

[0107] Furthermore, step S15 includes the following steps:

[0108] Step S151: Perform creation time label sub-column screening processing on each data asset information label description structured sub-column to obtain a data asset information creation time label structured sub-column and data asset information remaining label structured sub-columns;

[0109] Step S152: performing creation time sequence synchronization processing on the remaining tag structured sub-columns of the data asset information based on the data asset information creation time tag structured sub-column, so as to obtain various data asset information tag structured sub-columns under the same creation time sequence dimension;

[0110] Step S153: Designing a data integration storage format for each data asset information tag structured sub-column under the same creation time sequence dimension through the data asset management database within the data processing unit to obtain a data asset information tag sub-column integration storage format specification;

[0111] Step S154: Integrate and store each data asset information tag structured sub-column under the same creation time sequence dimension into the data asset management database in the data processing unit according to the data asset information tag sub-column integration storage format specification.

[0112] As an embodiment of the present invention, refer to Figure 3 As shown, Figure 2 Detailed step flow diagram of step S15 in the embodiment, step S15 includes the following steps:

[0113] Step S151: Perform creation time label sub-column screening processing on each data asset information label description structured sub-column to obtain a data asset information creation time label structured sub-column and data asset information remaining label structured sub-columns;

[0114] In an embodiment of the present invention, records containing creation time label descriptions are retrieved from each data asset information label description structured sub-column obtained after previous structured conversion to extract the creation time label sub-column. This can be accomplished through SQL queries or filtering functions in data analysis tools. For example, the query SELECT label_description,creation_time FROM data_assets WHERE label_description IS NOTNULL is executed to obtain all data asset information records with creation time label descriptions, thereby storing the creation time label sub-column separately from the remaining label sub-columns, and finally obtaining the data asset information creation time label structured sub-column and the data asset information remaining label structured sub-columns.

[0115] Step S152: performing creation time sequence synchronization processing on the remaining tag structured sub-columns of the data asset information based on the data asset information creation time tag structured sub-column, so as to obtain various data asset information tag structured sub-columns under the same creation time sequence dimension;

[0116] In an embodiment of the present invention, the remaining label structured sub-columns of the filtered data asset information are synchronized in time sequence by combining the previously extracted data asset information creation time label structured sub-column, including associating the creation time label with the record of each data asset information label, and then sorting them according to the creation time to ensure that all label information is aligned in the time dimension. This can be done by using the python code data_frame.sort_values(by='creation_time') to sort the remaining label structured sub-columns, and applying the sorted data to other label columns, thereby synchronizing under the same time sequence dimension, and finally obtaining each data asset information label structured sub-column under the same creation time sequence dimension.

[0117] Step S153: Designing a data integration storage format for each data asset information tag structured sub-column under the same creation time sequence dimension through the data asset management database within the data processing unit to obtain a data asset information tag sub-column integration storage format specification;

[0118] In an embodiment of the present invention, a data asset management database preset in a data processing unit is used to design a storage format after data integration for each data asset information label structured sub-column under the same creation time sequence dimension that has previously undergone time series synchronization processing, so as to define a storage format specification, including column name, data type and index setting. For example, a table structure can be created, which contains fields such as "creation time", "label 1", and "label 2". The data types of the fields are VARCHAR, DATE, etc., and indexes are set to improve query performance. The corresponding table structure is created through a database design tool or an SQL script, such as executing CREATE TABLE asset_data(creation_time DATE, label1 VARCHAR(255), label2 VARCHAR(255)) to ensure that the data asset information can be integrated and stored in accordance with the defined format. Finally, the data asset information label sub-column integrated storage format specification is designed.

[0119] Step S154: Integrate and store each data asset information tag structured sub-column under the same creation time sequence dimension into the data asset management database in the data processing unit according to the data asset information tag sub-column integration storage format specification.

[0120] In an embodiment of the present invention, by combining the data asset information label sub-column integration storage format specification obtained in the previous design, the various data asset information label structured sub-columns under the same creation time sequence dimension obtained after time series synchronization processing are integrated and stored in the data asset management database of the data processing unit. The specific operation includes importing the previously obtained synchronized data into the database, so as to insert the data into a pre-created table by using a data import tool or SQL command, for example, by executing INSERT INTO asset_data(creation_time,label1,label2) VALUES(data_creation_time,data_label1,data_label2), the data is inserted into the database table, ensuring that the data format complies with the design specifications during the insertion process, and performing necessary data verification and error handling to maintain the accuracy and consistency of the data.

[0121] Furthermore, step S2 includes the following steps:

[0122] Step S21: Perform access log recording processing on each data asset sub-information in the data asset management database to obtain data access log record information of each data asset sub-information;

[0123] In an embodiment of the present invention, a logging tool (such as ELK Stack or Splunk) is used to generate a detailed access log for the access behavior of each data asset sub-information in the data asset management database. Each access event to the data asset sub-information, including user ID, access time, access type (read, modify, delete, etc.), will be recorded. The log data is collected through a specific API interface and stored in the log database to ensure that each access to each data asset sub-information can be fully recorded to ensure the real-time and integrity of the access log, and finally the data access log record information of each data asset sub-information is recorded.

[0124] Step S22: Analyze the access count and time distribution of the data access log record information of each data asset sub-information to obtain the data access count and data access time distribution of each data asset sub-information;

[0125] In an embodiment of the present invention, data access log record information of each data asset sub-information in the log database is analyzed by using a data analysis tool (such as Apache Hadoop or the Pandas library in Python) to statistically analyze the number of accesses to each data asset sub-information and draw a time distribution graph. The number of accesses is calculated by querying the relevant fields in the log database, and the time distribution analysis is performed by counting the access timestamps to generate an access heat map. This step ensures the visualization of the data access frequency and its time distribution, and ultimately obtains the data access number and data access time distribution of each data asset sub-information.

[0126] Step S23: performing a usage frequency statistical analysis on the number of data accesses to each data asset sub-information based on the data access time distribution of each data asset sub-information to obtain the data usage frequency of each data asset sub-information;

[0127] In an embodiment of the present invention, a statistical analysis tool (such as the Scipy library of R language or Python) is used to perform a statistical calculation of the frequency of use of the data access times corresponding to each data asset sub-information in combination with the data access time distribution of each data asset sub-information obtained in the previous analysis, so as to quantitatively calculate the access frequency of each data asset sub-information. The access frequency statistics are performed by aggregating the time distribution data by hour, day or week to obtain the frequency distribution of access, and finally the data usage frequency of each data asset sub-information is obtained.

[0128] Step S24: Using the data asset information sensitivity calculation formula, perform information sensitivity evaluation calculation on each data asset sub-information in the data asset management database to obtain the data sensitivity level of each data asset sub-information;

[0129] In an embodiment of the present invention, a suitable data asset information sensitivity calculation formula is constructed by combining the quantity measurement parameters of data asset sub-information, data asset availability parameters, data asset availability impact weight coefficient, leakage risk measurement parameters, data asset leakage risk impact weight coefficient, leakage risk control impact exponential attenuation coefficient, data asset confidentiality parameters, data asset confidentiality impact weight coefficient, quantity measurement parameters of additional attribute information, specific values ​​on additional attribute information, impact weight coefficient of additional attribute information and related parameters to perform information sensitivity assessment calculation on each data asset sub-information in the data asset management database, so as to quantitatively calculate the corresponding sensitivity degree value, and finally obtain the data sensitivity degree of each data asset sub-information.

[0130] Step S25: Based on the data usage frequency and data sensitivity of each data asset sub-information, a data asset access utility evaluation analysis is performed on each data asset sub-information corresponding to the data asset management database to obtain the data access utility of each data asset sub-information.

[0131] In an embodiment of the present invention, a weighted scoring method is used to evaluate the access utility of each data asset sub-information corresponding to the data asset management database by combining the data usage frequency and data sensitivity of each data asset sub-information obtained from previous statistical analysis, so as to calculate the scores of the two dimensions according to the usage frequency and sensitivity respectively, and combine the scores of the two dimensions to obtain the final access utility score through the weighted average method. This step is usually automatically calculated using a data analysis tool or a custom script. For example, the corresponding calculation formula can be: Where, E k Score the final access utility of the i-th data asset sub-information, F i is the data usage frequency of the i-th data asset sub-information, a is the usage frequency weight coefficient, S i is the data sensitivity of the i-th data asset sub-information, and b is the sensitivity weight coefficient, which ensures the accuracy and objectivity of the evaluation results, and finally obtains the data access utility of each data asset sub-information.

[0132] Furthermore, the calculation formula for the data asset information sensitivity in step S24 is specifically:

[0133]

[0134] Where S i is the data sensitivity of the i-th data asset sub-information, n is the total number of data asset sub-information in the data asset management database, i is the item index parameter of the data asset sub-information, R iis the data asset availability parameter of the i-th data asset sub-information, α is the data asset availability impact weight coefficient, D i is the leakage risk measurement parameter of the i-th data asset sub-information, β is the data asset leakage risk impact weight coefficient, γ is the control leakage risk impact exponential attenuation coefficient, C i is the data asset confidentiality parameter of the i-th data asset sub-information, δ is the data asset confidentiality impact weight coefficient, m is the total number of additional attribute information in the data asset sub-information, j is the item measurement parameter of the additional attribute information, P i,j is the specific value of the i-th data asset sub-information in the j-th additional attribute information, κ j is the influence weight coefficient of the jth additional attribute information, and η is the correction coefficient of the data sensitivity.

[0135] The present invention obtains a data asset information sensitivity calculation formula by using a specific mathematical model and after verification, which is used to perform information sensitivity evaluation calculation on each data asset sub-information in the data asset management database. The data asset information sensitivity calculation formula comprehensively considers data availability, leakage risk, confidentiality and additional attributes, making the evaluation of data sensitivity more comprehensive and accurate. Through the exponential decay of leakage risk and logarithmic adjustment of confidentiality, the formula can dynamically reflect the changes in data sensitivity and adapt to different risk scenarios. Secondly, the weight coefficients in the calculation formula (such as α, β, δ, κ j ) allows the relative importance of influencing factors to be adjusted according to actual needs, thereby improving the flexibility and accuracy of sensitivity assessment. In addition, the formula also introduces a correction coefficient to adjust the sensitivity level according to actual conditions, thereby enhancing the applicability and reliability of the calculation formula. By evaluating data access frequency and sensitivity, it can effectively support data management decisions, optimize data access control and protection strategies, and thus improve the management efficiency of data assets. Therefore, this formula fully considers the data sensitivity level S of the i-th data asset sub-information i , the total number of data asset sub-information n in the data asset management database, the item index parameter i of the data asset sub-information, and the data asset availability parameter R of the i-th data asset sub-information i , data asset availability impact weight coefficient α, the leakage risk measurement parameter D of the i-th data asset sub-information i , data asset leakage risk impact weight coefficient β, control leakage risk impact exponential attenuation coefficient γ, data asset confidentiality parameter C of the i-th data asset sub-information i , data asset confidentiality impact weight coefficient δ, the total number of additional attribute information in the data asset sub-information m, the item measurement parameter j of the additional attribute information, the specific value P of the i-th data asset sub-information on the j-th additional attribute information i,j, the influence weight coefficient κ of the jth additional attribute information j , the correction coefficient η of data sensitivity, according to the data sensitivity S of the i-th data asset sub-information i The mutual correlation between the above parameters constitutes a functional relationship:

[0136]

[0137] This formula can realize the information sensitivity assessment calculation process of each data asset sub-information in the data asset management database. At the same time, by introducing the correction coefficient η of the data sensitivity level, it can be adjusted according to the errors occurring in the calculation process, thereby improving the accuracy and applicability of the data asset information sensitivity calculation formula.

[0138] Furthermore, step S25 includes the following steps:

[0139] Step S251: classifying the data sensitivity level of each data asset sub-information in the data asset management database based on the data sensitivity level of each data asset sub-information to obtain the data sensitivity level of each data asset sub-information;

[0140] In an embodiment of the present invention, the sensitivity level of each data asset sub-information corresponding to each data asset sub-information in the data asset management database is divided by combining the data sensitivity level of each data asset sub-information obtained by previous quantitative calculation, so as to be divided according to a pre-set sensitivity classification standard (such as high (for example, the data sensitivity level is 86-100), medium (for example, the data sensitivity level is 56-85), and low (for example, the data sensitivity level is 0-55)). The classification standard can be based on factors such as the data type and the privacy level of the personal or business information involved, and adopts specific methods, such as data classification tools or manual review methods, to ensure that each data asset sub-information is appropriately identified with a sensitivity level, and finally the data sensitivity level of each data asset sub-information is obtained.

[0141] Step S252: performing access utility weighted analysis on each data asset sub-information corresponding to the data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity level of each data asset sub-information to obtain a data access utility weighted value set for each data asset sub-information;

[0142] In an embodiment of the present invention, access utility weighted analysis is performed on the corresponding data asset sub-information in the data asset management database by combining the data usage frequency of each data asset sub-information obtained by previous statistical analysis and the data sensitivity level obtained by classification, so as to analyze and obtain the access utility weighted weight coefficient corresponding to the data sensitivity level and the usage frequency, thereby obtaining a corresponding weighted value set, including the sensitivity weight coefficient and the usage frequency weight coefficient, which can reflect the utility impact parameters of the usage frequency and sensitivity level corresponding to each data asset sub-information in actual access, and finally obtain the data access utility weighted value set for each data asset sub-information.

[0143] Step S253: performing data asset access utility evaluation and analysis on each data asset sub-information corresponding to the data asset management database according to the data access utility weighted value set of each data asset sub-information to obtain the data access utility of each data asset sub-information.

[0144] In an embodiment of the present invention, a weighted average calculation is performed based on the data access utility weighted value set of each data asset sub-information obtained from the previous analysis and combined with the data usage frequency and data sensitivity of the corresponding data asset sub-information in the data asset management database to quantitatively calculate the final access utility score. For example, the corresponding calculation formula can be: Where, E k Score the final access utility of the i-th data asset sub-information, F i is the data usage frequency of the i-th data asset sub-information, a is the usage frequency weight coefficient, S i is the data sensitivity of the i-th data asset sub-information, b is the sensitivity weight coefficient, and the data access utility of each data asset sub-information is finally quantitatively calculated.

[0145] Furthermore, step S3 includes the following steps:

[0146] Step S31: Performing a management role access business requirement analysis on each data asset sub-information in the data asset management database to obtain the database management role access business requirement for each data asset sub-information;

[0147] In an embodiment of the present invention, a business demand analysis of management role access is performed on each data asset sub-information in the data asset management database, so as to collect the management role and corresponding business demand of each data asset sub-information. This process is completed by accessing business demand documents, interviewing management personnel and analyzing existing permission settings. The management role information is extracted by using permission management tools (such as LDAP or Active Directory), and combined with business demand analysis tools (such as business process modeling tools BPMN) to record in detail the access requirements of each role to the data assets. For example, if a data asset sub-information involves financial data, the access requirements of the finance department role will be recorded in detail and associated with the data asset, and finally the database management role access business requirements of each data asset sub-information are obtained.

[0148] Step S32: performing data usage constraint analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency of each piece of data asset sub-information to obtain data usage constraint conditions for each piece of data asset sub-information;

[0149] In an embodiment of the present invention, a data usage constraint analysis is performed on the corresponding data asset sub-information in the data asset management database by combining the data usage frequency of each data asset sub-information obtained in the previous analysis, so as to process the access log of the data asset by using a data analysis platform (such as Apache Spark or Python's Pandas library), calculate the usage frequency of each data asset sub-information, and define data usage constraint conditions based on the analysis results, such as implementing stricter access control for frequently accessed sensitive data, combining usage frequency data with business needs, and setting reasonable data usage restrictions. For example, if the access frequency of a data asset sub-information is high, it is necessary to limit the access time period or number of accesses to the data to reduce the risk of data leakage, and finally obtain the data usage constraint conditions for each data asset sub-information.

[0150] Step S33: performing access utility restriction analysis on each data asset sub-information in the data asset management database based on the data access utility of each data asset sub-information, and obtaining a data access utility restriction range for each data asset sub-information;

[0151] In an embodiment of the present invention, an access utility restriction analysis is performed on the corresponding data asset sub-information in the data asset management database by combining the data access utility of each data asset sub-information obtained by previous quantitative calculation, so as to generate an access utility chart by using a data visualization tool (such as Tableau or Microsoft Power BI) to determine the utility range of data access, wherein the data access utility restriction range is set according to the access frequency and data sensitivity. For example, for high-utility and high-sensitivity data assets, strict access control restrictions are set, such as restricting access rights or strengthening identity authentication requirements, and finally the data access utility restriction range of each data asset sub-information is obtained.

[0152] Step S34: Dynamically control the access rights of the database management role to each data asset sub-information based on the data usage constraints and data access utility limitation range of each data asset sub-information to obtain dynamic access control rights for the data asset management database.

[0153] In an embodiment of the present invention, dynamic access rights are controlled based on the business needs of database management roles for access to each data asset sub-information previously analyzed by combining the data usage constraints and data access utility restriction range of each data asset sub-information previously analyzed. By utilizing a dynamic permission control system (such as an RBAC-based permission management system or an ABAC-based access control policy), data usage constraints and access utility restriction range are combined with the business needs of the management role to create a dynamic access control policy. The implementation process includes configuring the policy rules of the permission management system and dynamically adjusting the role permissions to meet the requirements of data usage and access utility. For example, for a data asset with high sensitivity and high access utility, permissions will be automatically restricted to only necessary management roles, and stricter access auditing and control measures will be implemented to ultimately obtain dynamic access control permissions for the data asset management database.

[0154] Furthermore, step S4 includes the following steps:

[0155] Step S41: performing management role authority matching and mapping on the dynamic access control authority of the data asset management database to obtain the data asset management role dynamic access matching and mapping control authority;

[0156] In an embodiment of the present invention, the dynamic access control permissions in the data asset management database obtained by previous analysis are matched with the permissions of the management roles, so as to extract the access permission configuration of each management role from the permission management system, and map these permissions to the corresponding data asset information. A permission mapping tool (such as the role mapping module in the RBAC system) is used for mapping, and the permissions of each management role are matched with the access requirements of the data assets. The dynamic access permissions of each role to specific data assets are determined through access control lists (ACLs) or permission configuration tables. The generated mapping results include the role and its corresponding data asset access permissions, and finally the dynamic access matching mapping control permissions of the data asset management role are obtained.

[0157] Step S42: Based on the data asset management role dynamic access matching mapping control authority, the corresponding data asset information in the data asset management database is analyzed for full life cycle stage authority characteristics to obtain a data asset full life cycle stage management role authority characteristic set;

[0158] In an embodiment of the present invention, a full life cycle permission feature analysis is performed on the corresponding data asset information in the data asset management database by combining the dynamic access matching mapping control permissions of the data asset management role obtained by previous mapping assignment, so as to analyze the various life cycle stages (such as creation, storage, use, sharing, and destruction) of data assets through life cycle management tools (such as enterprise data management systems), record the management roles and their permission features involved in each stage, and analyze the permission requirements and restrictions of roles in different life cycle stages. For example, in the data creation stage, a role with higher permissions is required to enter data, while in the data use stage, it is necessary to ensure that only specific roles can access it. The generated full life cycle stage management role permission feature set includes the permission requirements and feature descriptions of each life cycle stage, and finally a full life cycle stage management role permission feature set of data assets is obtained.

[0159] Step S43: Formulate full life cycle management rules for the data asset full life cycle management role authority feature set to generate data asset information full life cycle management rules;

[0160] In an embodiment of the present invention, full life cycle management rules are formulated for the set of management role permissions characteristics of the full life cycle stages of data assets obtained through previous analysis, so as to create full life cycle management rules for data asset information by using a rule engine (such as Drools) or a policy formulation tool. The permission feature set obtained through previous analysis is converted into management rules. For example, access control policies and operation requirements are defined for different life cycle stages of data. The formulated rules include how to implement permission control, monitoring and audit requirements at different stages of data to ensure that the management of data assets throughout the entire life cycle complies with the set policies and security standards, and finally the corresponding full life cycle management rules for data asset information are formulated and generated.

[0161] Step S44: performing automated management processing on the corresponding data asset information in the data asset management database based on the data asset information full life cycle management rules to obtain the data asset information full life cycle automatic management results;

[0162] In an embodiment of the present invention, the corresponding data asset information in the data asset management database is automatically managed and processed by combining the data asset information full life cycle management rules obtained from the previous analysis, so that the management rules can be applied to data assets by utilizing an automated management system (such as a data governance platform or an automated workflow management system), and data access control, permission allocation, data auditing and compliance checks will be automatically performed according to the rules. For example, when data enters a new life cycle stage, access rights are automatically adjusted, permission settings are updated, and necessary management operations are performed. The results include the automated management status and related logs of data asset information throughout its life cycle, and ultimately the automatic management results of data asset information throughout its life cycle are obtained.

[0163] Step S45: Record the management report feedback of the automatic management results of the data asset information throughout its life cycle to generate a data asset information management report.

[0164] In an embodiment of the present invention, a report generation tool (such as Microsoft Power BI or Tableau) is used to record the feedback of the management report on the automatic management results of the entire life cycle of data asset information obtained after previous automated management, so as to summarize the automatic management processing results using a report generation tool (such as Microsoft Power BI or Tableau) and generate a detailed management report. The report content includes the execution status of automated management, the effectiveness of authority control, the management effect of the life cycle stage, etc. Feedback records are generated regularly through the reporting system and provided to the management for decision-making reference. The report also includes performance indicators and improvement suggestions for data asset management to help optimize the management process of data assets, and finally the records are generated to generate corresponding data asset information management reports.

[0165] Furthermore, the present invention also provides a data asset management system for executing the above-mentioned data asset management method, the data asset management system comprising:

[0166] The data asset information structured storage module is used to collect data asset information sets from databases, file systems, and API data sources in real time, and divide the data asset information sets into tag structure sub-columns to obtain data asset information tag description structured sub-columns; create time series synchronization processing for each data asset information tag description structured sub-column and integrate and store them in the data asset management database;

[0167] The usage frequency and access utility analysis module is used to perform usage frequency and sensitivity evaluation analysis on each data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each data asset sub-information; based on the data usage frequency and data sensitivity of each data asset sub-information, the data asset access utility evaluation analysis is performed on each data asset sub-information in the data asset management database to obtain the data access utility of each data asset sub-information;

[0168] A database dynamic access permission control module is used to dynamically control the access permission of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information, thereby obtaining dynamic access control permissions for the data asset management database;

[0169] The data asset full life cycle management module is used to perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database, so as to generate data asset information full life cycle management rules; based on the data asset information full life cycle management rules, the corresponding data asset information in the data asset management database is automatically managed and report feedback records are recorded to generate data asset information management reports.

[0170] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0171] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A data asset management method, characterized in that: The following steps are involved: Step S1: Data asset information sets are collected in real time from enterprise databases, file systems, and API data sources, and the data asset information sets are divided into tag structure sub-columns to obtain data asset information tag description structured sub-columns; each data asset information tag description structured sub-column is created and synchronized, and integrated and stored in the data asset management database within the data processing unit; Step S2: Performing a usage frequency and sensitivity assessment analysis on each piece of data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each piece of data asset sub-information; performing a data asset access utility assessment analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity of each piece of data asset sub-information to obtain the data access utility of each piece of data asset sub-information; Step S3: Dynamically control the access rights of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information to obtain dynamic access control rights for the data asset management database; wherein step S3 includes the following steps: Step S31: Performing a management role access business requirement analysis on each data asset sub-information in the data asset management database to obtain the database management role access business requirement for each data asset sub-information; Step S32: performing data usage constraint analysis on each piece of data asset sub-information in the data asset management database based on the data usage frequency of each piece of data asset sub-information to obtain data usage constraint conditions for each piece of data asset sub-information; Step S33: performing access utility restriction analysis on each data asset sub-information in the data asset management database based on the data access utility of each data asset sub-information, and obtaining a data access utility restriction range for each data asset sub-information; Step S34: Dynamically control the access permissions of the database management role for each data asset sub-information based on the data usage constraints and data access utility restriction range of each data asset sub-information to obtain dynamic access control permissions for the data asset management database; Step S4: Perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database to generate full life cycle management rules for the data asset information; perform automated management and report feedback records for the corresponding data asset information in the data asset management database based on the full life cycle management rules for the data asset information to generate a data asset information management report.

2. The data asset management method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining a database source data asset information subset by collecting a data asset-related information subset from an enterprise database data source in real time; Step S12: obtaining a file system source data asset information subset by collecting a data asset-related information subset from the file system data source in real time; Step S13: obtaining an API source data asset information subset by collecting a data asset-related information subset from the API data source in real time; Step S14: performing information set processing on the database source data asset information subset, the file system source data asset information subset, and the API source data asset information subset to obtain a data asset information set; dividing the data asset information set into label structure sub-columns to obtain data asset information label description structured sub-columns; Step S15: Create a time series synchronization process for each data asset information tag description structured sub-column and integrate and store it in the data asset management database.

3. The data asset management method according to claim 2, characterized in that: The step S14 of dividing the data asset information set into sub-columns of the tag structure includes the following steps: Performing information content semantic analysis on each piece of data asset information in the data asset information set to obtain content semantic data corresponding to the data asset content of each piece of data asset information; Perform data format attribute mining and analysis on each data asset information in the data asset information set to obtain the data format attribute of each data asset information corresponding to the data asset content; Based on the data format attributes of the data asset content corresponding to each data asset information, the content semantic data of the data asset content corresponding to each data asset information is classified and labeled to obtain various data content attribute labels corresponding to the data asset content, wherein the various data content attribute labels include creation time, data type, storage location, data usage and modification time; Divide each data asset information in the data asset information set into an attribute label description according to various data content attribute labels corresponding to the data asset content, and obtain a data asset information content attribute label description sub-column; Perform structural conversion on each data asset content in the data asset information content attribute label description sub-column to obtain a data asset information label description structured sub-column.

4. The data asset management method according to claim 2, characterized in that: Step S15 includes the following steps: Step S151: Perform creation time label sub-column screening processing on each data asset information label description structured sub-column to obtain a data asset information creation time label structured sub-column and data asset information remaining label structured sub-columns; Step S152: performing creation time sequence synchronization processing on the remaining tag structured sub-columns of the data asset information based on the data asset information creation time tag structured sub-column, so as to obtain various data asset information tag structured sub-columns under the same creation time sequence dimension; Step S153: Designing a data integration storage format for each data asset information tag structured sub-column under the same creation time sequence dimension through the data asset management database within the data processing unit to obtain a data asset information tag sub-column integration storage format specification; Step S154: Integrate and store each data asset information tag structured sub-column under the same creation time sequence dimension into the data asset management database in the data processing unit according to the data asset information tag sub-column integration storage format specification.

5. The data asset management method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Perform access log recording processing on each data asset sub-information in the data asset management database to obtain data access log record information of each data asset sub-information; Step S22: Analyze the access count and time distribution of the data access log record information of each data asset sub-information to obtain the data access count and data access time distribution of each data asset sub-information; Step S23: performing a usage frequency statistical analysis on the number of data accesses to each data asset sub-information based on the data access time distribution of each data asset sub-information to obtain the data usage frequency of each data asset sub-information; Step S24: Using the data asset information sensitivity calculation formula, perform information sensitivity evaluation calculation on each data asset sub-information in the data asset management database to obtain the data sensitivity level of each data asset sub-information; Step S25: Based on the data usage frequency and data sensitivity of each data asset sub-information, a data asset access utility evaluation analysis is performed on each data asset sub-information corresponding to the data asset management database to obtain the data access utility of each data asset sub-information.

6. The data asset management method according to claim 5, characterized in that: The calculation formula for data asset information sensitivity in step S24 is specifically: ; Where, For the The data sensitivity of each data asset sub-information, The total number of data asset sub-information in the data asset management database. It is the item index parameter of the data asset sub-information. For the Data asset availability parameters for each data asset sub-information, is the weight coefficient affecting the availability of data assets, For the Leakage risk measurement parameters of data asset sub-information, is the weight coefficient of data asset leakage risk impact, In order to control the exponential attenuation coefficient of leakage risk, For the The data asset confidentiality parameter of each data asset sub-information, is the weight coefficient of data asset confidentiality impact, is the total number of additional attribute information in the data asset sub-information, is the item-level measurement parameter of the additional attribute information, For the The data asset sub-information is in The specific value of the additional attribute information, For the The influence weight coefficient of additional attribute information, is the correction factor for data sensitivity.

7. The data asset management method according to claim 5, characterized in that: Step S25 includes the following steps: Step S251: classifying the data sensitivity level of each data asset sub-information in the data asset management database based on the data sensitivity level of each data asset sub-information to obtain the data sensitivity level of each data asset sub-information; Step S252: performing access utility weighted analysis on each data asset sub-information corresponding to the data asset sub-information in the data asset management database based on the data usage frequency and data sensitivity level of each data asset sub-information to obtain a data access utility weighted value set for each data asset sub-information; Step S253: performing data asset access utility evaluation and analysis on each data asset sub-information corresponding to the data asset management database according to the data access utility weighted value set of each data asset sub-information to obtain the data access utility of each data asset sub-information.

8. The data asset management method according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: performing management role authority matching and mapping on the dynamic access control authority of the data asset management database to obtain the data asset management role dynamic access matching and mapping control authority; Step S42: Based on the data asset management role dynamic access matching mapping control authority, the corresponding data asset information in the data asset management database is analyzed for full life cycle stage authority characteristics to obtain a data asset full life cycle stage management role authority characteristic set; Step S43: Formulate full life cycle management rules for the data asset full life cycle management role authority feature set to generate data asset information full life cycle management rules; Step S44: performing automated management processing on the corresponding data asset information in the data asset management database based on the data asset information full life cycle management rules to obtain the data asset information full life cycle automatic management results; Step S45: Record the management report feedback of the automatic management results of the data asset information throughout its life cycle to generate a data asset information management report.

9. A data asset management system, characterized in that: For executing the data asset management method according to claim 1, the data asset management system comprises: The data asset information structured storage module is used to collect data asset information sets from databases, file systems, and API data sources in real time, and divide the data asset information sets into tag structure sub-columns to obtain data asset information tag description structured sub-columns; create time series synchronization processing for each data asset information tag description structured sub-column and integrate and store them in the data asset management database; The usage frequency and access utility analysis module is used to perform usage frequency and sensitivity evaluation analysis on each data asset sub-information in the data asset management database to obtain the data usage frequency and data sensitivity of each data asset sub-information; based on the data usage frequency and data sensitivity of each data asset sub-information, the data asset access utility evaluation analysis is performed on each data asset sub-information in the data asset management database to obtain the data access utility of each data asset sub-information; A database dynamic access permission control module is used to dynamically control the access permission of each data asset sub-information in the data asset management database based on the data usage frequency and data access utility of each data asset sub-information, thereby obtaining dynamic access control permissions for the data asset management database; The data asset full life cycle management module is used to perform full life cycle management design for the corresponding data asset information in the data asset management database based on the dynamic access control permissions of the data asset management database, so as to generate data asset information full life cycle management rules; based on the data asset information full life cycle management rules, the corresponding data asset information in the data asset management database is automatically managed and report feedback records are recorded to generate data asset information management reports.

Citation Information

Patent Citations

  • Access controlled distributed ledger system for asset management

    CN111868768A

  • Urban development situation analysis method, device and system based on big data

    CN114418275A