Data asset management platform
By designing a data asset management platform, the problems of data silos, inefficient development and governance, insufficient security management and high development thresholds are solved, and panoramic data views, real-time data collection, transparent governance and efficient security management are realized.
Patent Information
- Application Number
- CN202510108719.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology is difficult to form a panoramic corporate data view, resulting in serious data silos, low data development and governance efficiency, insufficient data security management, high development threshold, and opaque data processing chain.
A data asset management platform has been designed, including data source management module, data acquisition module, data cataloging module, data governance module, label management module and data warehouse. Through metadata architecture, breaks data silos, collects data in real time, provides data governance and security management functions, and lowers development thresholds.
It realizes a panoramic enterprise data view, improves data asset management efficiency, reduces data access difficulty, ensures data security, adapts to the needs of different developers, and provides a transparent data processing chain.
Smart Images

Figure CN120045622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data management technology, and more particularly to a data asset management platform. Background Art
[0002] In today's digital economy era, data has become a key production factor after land, labor, and capital. However, the core challenges faced by current enterprises are mainly reflected in two aspects: First, with the rapid growth of data volume, data is distributed in different systems and databases, and traditional management methods are difficult to form a panoramic enterprise data view, resulting in serious data island phenomena and difficulty in meeting modern business needs; Second, the existing big data technology has a high threshold, and developers need to be familiar with the underlying complex technical architecture, resulting in low efficiency of data development and governance and difficulty in quickly realizing data value. In addition, the functions of data security management, sensitive data identification, version tracking, etc. are limited in the existing platform and cannot meet the enterprise's requirements for efficient, secure, and visual data management. These problems have seriously hindered the transformation of data resources into data assets and the digital transformation of enterprises.
[0003] The Chinese patent application with publication number CN111586033A discloses an asset data middle platform for a data center, the Chinese patent application with publication number CN115344635A discloses a method for constructing a data asset platform based on data objects, and the Chinese patent application with publication number CN118350030A discloses an application assembly middle platform system based on a metadata asset platform.
[0004] The above prior arts have the following technical problems:
[0005] Data island problem. It is difficult for the prior art to form a panoramic enterprise data view, resulting in data being distributed in different systems and databases, forming data islands and unable to meet modern business needs.
[0006] Low efficiency of data development and governance. The existing big data technology has a high threshold, and developers need to be familiar with the underlying complex technical architecture, resulting in low efficiency of data development and governance and difficulty in quickly realizing data value.
[0007] Insufficient data security management. The existing platform has limited implementation of functions such as data security management, sensitive data identification, and version tracking, and cannot meet the enterprise's requirements for efficient, secure, and visual data management.
[0008] The data processing chain is opaque. The data processing chain of the prior art is not visual enough, resulting in an opaque development and governance process and difficulty in effective monitoring and management.
[0009] The development threshold is high. Some data development and governance platforms have relatively high skill requirements for developers. It is difficult for junior developers to get started quickly, while senior developers need to invest a lot of time and energy in the processing and development of underlying technologies. Summary of the Invention
[0010] In view of the deficiencies in the prior art, the present invention provides a data asset management platform.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] A data asset management platform, comprising:
[0013] A data source management module, used to manage the sources and acquisition methods of data assets;
[0014] A data collection module, used to collect data in real time from different data sources;
[0015] A data cataloging module, used to classify the collected data assets;
[0016] A data governance module, used for data quality management, data standardization, data analysis, data desensitization and sensitive data identification, and processing data through user-defined functions;
[0017] A label management module, used to formulate label templates for marking data;
[0018] A data warehouse, used to store the collected data.
[0019] To optimize the above technical solutions, the specific measures taken further include:
[0020] Furthermore, the sources of the data assets include databases, API data sources and CSV file data sources. The databases include relational databases, MPP databases and distributed databases. The management of the sources and acquisition methods of the data assets specifically is: allowing users to perform multi-level folder management on the data sources, providing functions for searching, filtering, editing, viewing details and viewing data source connection log information, and used for monitoring and managing the status of the data sources.
[0021] Furthermore, the real-time collection of data from different data sources specifically is: non-intrusively docking various data sources for real-time data collection, and providing functions of delay alarm, traffic control and developer API.
[0022] Furthermore, the classification of the collected data assets specifically is:
[0023] Register new data assets;
[0024] Publish the registered data assets to the data portal for other users to access; classify and display the overview information of the published data assets, including the source, type, and usage of the data.
[0025] Further, the rules for data desensitization are specifically as follows:
[0026] For names, remove some characters and retain the surname; for ID numbers, hide the middle part of the numbers; for mobile phone numbers, hide the middle four digits; for bank card numbers, hide the other numbers except the first three and the last four; for income amounts, retain the approximate range and remove the exact value.
[0027] Further, the specific rules for sensitive data identification include:
[0028] Search for preset sensitive keywords in the data, including "name", "ID number", and "bank card number". When these keywords are found, identify the relevant data as the corresponding sensitive data type according to the category to which the keyword belongs.
[0029] Use regular expressions to match the data format. The regular expression for the ID number is "^\\d{18}". When the data format conforms to the regular expression, identify the data as the corresponding sensitive data type and classify it according to the preset rules; through natural language processing technology, perform semantic understanding on the data content, analyze the meaning expressed by the data, and determine whether it involves sensitive content.
[0030] Combine the context environment where the data is located for identification and classification to determine whether it is sensitive data.
[0031] Further, the label template specifically includes:
[0032] Label name, label description, applicable data types, and examples.
[0033] Further, the data analysis includes: batch calculation, stream calculation, and BI analysis.
[0034] The beneficial effects of the present invention are as follows: First, the unified metadata architecture breaks the data silo problem, provides a global view, and improves the efficiency of data asset management; Second, the self-developed real-time acquisition framework significantly reduces the difficulty of data access and ensures data timeliness and accuracy; Third, the DAG task scheduling model makes the data governance process transparent and visual, greatly reducing the development threshold and meeting the needs of both junior and senior developers; Fourth, the sensitive data identification and desensitization module ensures data security and compliance from the source; Fifth, the scalable architecture of the platform supports docking with multiple data warehouses and storage engines, has strong compatibility, and is easy to deploy and migrate. Ultimately, this solution successfully realizes the value transformation of data from resources to assets, providing strong technical support for the digital transformation of enterprises. By building a one-stop data asset management platform for data development, data management, and data operation, it provides the key capabilities for the transformation of data resources into data assets, realizing the "visibility, retrievability, usability, and operability" of data assets. Based on the already-governed data resources, asset data is defined, discovered, and identified. By constructing an inventory directory of data assets, a set of context information is provided to fully release the potential value of data resources. At the same time, the asset data is published to the data portal with the asset directory as the carrier, enabling data consumers to see, find, and use the data, and establishing a feedback channel for data use, forming a virtuous feedback loop between the data supply side and the data consumption side, thus realizing the continuous operation of data assets. Brief Description of the Drawings
[0035] Figure 1 It is a panoramic view of the data asset management platform architecture;
[0036] Figure 2 It is a schematic diagram of metadata management. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0038] Embodiment 1
[0039] The present invention proposes a data asset management platform, and the overall architecture of the platform is as Figure 1 shown, and the metadata management in the platform is as Figure 2As shown, Metadata: is data about data, which describes the attributes, characteristics, and behaviors of data assets. Metadata provides a basic framework for data asset management, enabling data assets to be better understood, located, accessed, and used. Metadata management is the core of data development, governance, and management, which involves aspects such as data discovery, quality, security, and maintenance. Adopting a metadata-driven underlying service design architecture, with metadata as the core to build other data function modules, the underlying metadata of the product is divided into technical metadata, business metadata, management metadata, and service metadata, ensuring that metadata information is left throughout the data development and processing process, enabling quick and accurate positioning of the required data and related information, guaranteeing the management and application of the entire data life cycle, enriching the description dimension and accuracy of data, being the basis of data governance, and being a key part of data value mining.
[0040] The data asset management platform proposed by the present invention includes the following modules:
[0041] The data source management module, which is the basis of metadata management, is used to manage the sources and acquisition methods of data assets; the sources of data assets include databases, API data sources, and CSV file data sources to meet different data integration requirements. The databases include relational databases (such as MySQL, Oracle, SQL Server), MPP databases (such as Greenplum), and distributed databases (such as Hive). The specific management of the sources and acquisition methods of data assets is as follows: allowing users to perform multi-level folder management on data sources for easy organization and classification. Providing functions for searching, filtering, editing, viewing details, and viewing data source connection log information for monitoring and managing the status of data sources.
[0042] The data collection module is used to collect data in real time from different data sources; specifically: non-intrusively docking various data sources for real-time data collection, providing functions such as delay alarm, traffic control, and developer API to ensure the timeliness and accuracy of data extraction. The full name of API is Application Programming Interface, and its Chinese meaning is application programming interface. It is a set of predefined functions, protocols, and tools for building software applications. The data collection module can also clean, transform, and integrate data.
[0043] The data cataloging module is used to classify the collected data assets so that users can quickly find the required data. The specific classification of the collected data assets is as follows:
[0044] Register new data assets;
[0045] Publish the registered data assets to the data portal for other users to access; categorically display the overview information of the published data assets, including the data source, type, and usage.
[0046] A data governance module for data quality management, data standardization, data analysis, data masking, and sensitive data identification, as well as processing data through user-defined functions (UDFs); UDFs are a powerful tool that allows users to extend the functionality of a database or data processing system by handling complex data processing logic through custom functions.
[0047] Data quality management can ensure the accuracy, integrity, and consistency of data, and data standardization can establish data standards for the unified management and use of data.
[0048] Data masking can protect the privacy and security of data. The specific rules for data masking are as follows:
[0049] For names, remove some characters and keep the surname; for ID numbers, hide the middle part of the digits; for mobile phone numbers, hide the middle four digits; for bank card numbers, hide all digits except the first three and the last four; for income amounts, keep the approximate range and remove the exact value. The masking rules and masking methods are shown in Table 1:
[0050] Table 1
[0051] Sensitive data type Desensitization rule Desensitization method Name Remove some characters and keep the surname Replacement desensitization, such as "Zhang**” ID number Hide the middle part of the numbers Masking desensitization, such as "320**********1234” Mobile phone number Hide the middle four digits Masking desensitization, such as "138****1234” Bank card number Hide all numbers except the last four digits Masking desensitization, such as "622**********1234” Income amount Retain the approximate range and remove the exact value Replacement desensitization, such as "5000 - 10000 yuan”
[0052] The specific rules for sensitive data identification include:
[0053] Search for preset sensitive keywords in the data, including "name", "ID number", and "bank card number". When these keywords are found, based on the category to which the keyword belongs, identify the relevant data as the corresponding sensitive data type; for example, if the keyword "name" is found, identify the data as sensitive data of the personal basic information type.
[0054] Use regular expressions to perform format matching on the data. The regular expression for ID numbers is "^\\d{18}". When the data format conforms to the regular expression, identify the data as the corresponding sensitive data type and classify it according to the preset rules; through natural language processing technology, perform semantic understanding on the data content, analyze the meaning expressed by the data, and determine whether it involves sensitive content such as personal privacy, financial status, and health information; for example, a piece of data containing a patient's condition description and treatment plan can be identified as sensitive data of the medical and health data type through semantic analysis.
[0055] Identify and classify data in the context where it is located to determine whether it is sensitive data. For example, data such as "income" and "expenses" that appear in a financial statement can be identified as sensitive data of the financial data category based on the financial scenario of its context; data such as "latitude and longitude" and "city name" that appear in a geographic information system data can be identified as sensitive data of the geographic location information category based on its geographic context.
[0056] The rules for identifying sensitive data are shown in Table 2:
[0057] Table 2
[0058]
[0059]
[0060] Data analysis includes: batch computing, stream computing, and BI analysis.
[0061] The tag management module is used to formulate tag templates for marking data; enhance the discoverability and understandability of data through tags. The tag templates specifically include: tag name, tag description, applicable data types, and examples. It also allows users to create custom tags according to needs for personalized data organization and retrieval. The tag templates are shown in Table 3:
[0062] Table 3
[0063]
[0064] The data warehouse is used to store the collected data.
[0065] Industry models and standard codes: are models and codes designed to improve the usability and consistency of data. These include:
[0066] Industry models: Define data models and structures according to the characteristics of specific industries.
[0067] Standard codes: Provide standard data codes and terms for the unified understanding and use of data.
[0068] Through these metadata management functions, the data asset management platform can provide a comprehensive, structured, and efficient data management environment, enabling data assets to be effectively developed, governed, and utilized. Metadata is data that describes data. By structuring the description of data through metadata, data can be made easier to understand, search, and use. Therefore, scientific and standardized metadata management is the foundation of data development, governance, and management.
[0069] When the present invention is actually implemented, it features a no-code process operation combined with lightweight code development, which can flexibly adapt to both junior and senior developers. It provides most no-code data processing functions. For junior data developers, after clearly sorting out the business requirements of data development, they can quickly complete the entire data processing process without writing code. At the same time, a lightweight code development page is provided for data processing, processing, and calculation scenarios, enabling senior developers to handle more complex data and adapting to different development roles and capabilities.
[0070] The underlying storage platform is designed with a loose coupling structure, enabling flexible connection to common third-party big data platforms.
[0071] The product service and the underlying data technology platform adopt a common protocol and a loose coupling docking architecture, which can flexibly support multiple common Hadoop technology platforms including DDP, and adapt to various scenarios of big data platforms purchased from other manufacturers. It provides good compatibility and transition feasibility for projects that have been built using the Hadoop platform for some time. Compatibility tests have been completed, and it supports multiple big data platforms in the Hadoop system, including DDP, TDH, CDP, Huawei FusionInsightHD, Cloudera's CDH, Hortworks' HDP, and the Telecom Feilong platform.
[0072] A data-centric security management mechanism provides the ability to share data security governance.
[0073] The data security architecture is designed in accordance with the DSMM (Data Security Capability Maturity Model) to ensure the security of data in all aspects such as "storage, management, and usage". The platform has underlying data security technical guarantees at the system layer, interface layer, application layer, data layer, and platform facility layer, providing basic data security capabilities such as authentication access, permission control, classification, log auditing, monitoring and warning, and emergency recovery, covering the entire life cycle security of data including data collection security, data transmission security, data storage security, data processing security, data sharing security, and data destruction security, and effectively reducing data security risks by standardizing and restricting the entire data transfer process.
[0074] In the embodiments disclosed in the present application, a computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0075] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed in the present application can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0076] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A data asset management platform, characterized in that: include: Data source management module, used to manage the source and acquisition method of data assets; Data collection module, used to collect data from different data sources in real time; Data cataloging module, used to classify collected data assets; Data governance module, which is used for data quality management, data standardization, data analysis, data desensitization and sensitive data identification, as well as data processing through user-defined functions; The label management module is used to formulate label templates for labeling data; Data warehouse, used to store collected data.
2. The data asset management platform according to claim 1, characterized in that: The sources of the data assets include databases, API data sources and CSV file data sources. The databases include relational databases, MPP databases and distributed databases. The sources and acquisition methods of managing data assets are specifically as follows: allowing users to perform multi-level folder management on data sources, providing search, filtering, editing, detail viewing and data source connection log information viewing functions, which are used to monitor and manage the status of data sources.
3. The data asset management platform according to claim 1, characterized in that: The real-time data collection from different data sources specifically includes: non-invasively connecting to various data sources for real-time data collection, and providing delayed alarm, traffic control, and developer API functions.
4. The data asset management platform according to claim 1, characterized in that: The classification of the collected data assets is specifically as follows: Register new data assets; Publish registered data assets to the data portal for other users to access; categorize and display overview information of published data assets, including the source, type, and purpose of the data.
5. The data asset management platform according to claim 1, characterized in that: The specific rules for data desensitization are: For names, remove some characters and keep the surname; for ID numbers, hide the middle part of the numbers; for mobile phone numbers, hide the middle four digits; for bank card numbers, hide all digits except the first three and the last four digits; for income amounts, keep the approximate range and remove the exact value.
6. The data asset management platform according to claim 1, characterized in that: The specific rules for identifying sensitive data include: Search the data for preset sensitive keywords, including "name", "ID number" and "bank card number". When these keywords are found, identify the relevant data as the corresponding sensitive data type according to the category to which the keywords belong; Use regular expressions to match the data format. The regular expression for the ID card number is "^\d{18}". When the data format meets the regular expression, the data is identified as the corresponding sensitive data type and classified according to the preset rules. Use natural language processing technology to understand the semantics of the data content, analyze the meaning expressed by the data, and determine whether it involves sensitive content. Identify and classify the data based on its context to determine whether it is sensitive data.
7. The data asset management platform according to claim 1, characterized in that: The label template specifically includes: Tag name, tag description, applicable data type, and example.
8. The data asset management platform according to claim 1, characterized in that: The data analysis includes: batch computing, stream computing and BI analysis.
Citation Information
Patent Citations
Asset data center station of data center
CN111586033A
Data asset platform construction method based on data object
CN115344635A
Metadata asset platform-based application assembly platform system
CN118350030A