Cloud analysis data asset classification and level-to-level management system

Through the combination of deep learning and graph neural networks, intelligent management of multi-language, multi-source heterogeneous data is achieved, which solves the efficiency and accuracy problems of data classification and grading in existing technologies and provides an efficient and secure data management solution.

CN120763708AInactive Publication Date: 2025-10-10NINGXIA KAIXINTE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510981469.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack efficient algorithm support when processing multi-language and multi-domain data, and are unable to cope with complex, multi-source, and heterogeneous data environments, especially in terms of metadata extraction and classification and grading.

Method used

By adopting a metadata extraction algorithm based on deep learning and attention mechanism, combined with multi-language AI large-model translation and classification and grading algorithms of graph neural networks and multi-layer perceptrons, we build a multi-language, multi-source heterogeneous data intelligent processing platform to achieve automatic identification, classification and grading of data assets.

Benefits of technology

It improves the efficiency and security of data management, can accurately identify key metadata, ensure the consistency of translation results and the accuracy of classification and grading, provide data source analysis and security assessment functions, and build an integrated security management and control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763708A_ABST
    Figure CN120763708A_ABST
Patent Text Reader

Abstract

The invention provides a cloud analysis data asset classification and level-to-level management system, and relates to the technical field of asset management. The cloud analysis data asset classification and level-to-level management system comprises a data acquisition layer, the data acquisition layer is connected with a metadata extraction layer and a preprocessing layer, the preprocessing layer is connected with a classification and level-to-level engine layer, the classification and level-to-level engine layer is connected with a multi-language AI large model translation layer and then connected with a storage layer, and the storage layer is connected with an application layer. The metadata extraction layer is a metadata extraction algorithm based on deep learning and an attention mechanism and can automatically identify and extract key metadata fields from original data and keep metadata unchanged, and the preprocessing layer cleans, converts and standardizes the collected data. By combining metadata extraction, AI large model translation and classification and grading technologies, the automation level of data asset management is improved, and efficient management and protection of data assets are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of asset management, and in particular to a cloud analysis data asset classification and grading management system. Background Art

[0002] Existing technologies typically rely on manual labor and simple tools to achieve data asset classification and grading. When managing data assets, enterprises first conduct preliminary identification and sorting of data assets through manual inventory. To improve efficiency, they may use simple data management tools to assist with manual inventory and perform preliminary data classification and grading. Data classification and grading typically involves manual classification based on pre-defined classification standards and grading rules. Some tools may utilize simple rule engines to automatically classify and grade data based on preset rules, but these rules typically require manual setup and maintenance. The formulation and implementation of security policies are also largely manual. After data classification and grading is complete, enterprises manually formulate corresponding security policies based on the classification results. Security tools, such as data encryption and access control, are then used to enforce these policies.

[0003] With the advent of the big data era, the amount of data faced by businesses, institutions, and individuals is growing exponentially. How to effectively classify, grade, and manage this data has become a pressing issue. Traditional classification and grading methods rely on manual rules and simple machine learning models, making them inadequate for complex, multi-source, and heterogeneous data environments. Furthermore, while existing AI big data models possess powerful semantic understanding and translation capabilities, their application in metadata extraction and classification and grading remains limited, particularly when processing multilingual and multi-domain data, which lacks efficient algorithmic support. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a cloud analysis data asset classification and grading management system, which solves the problem of the existing technology lacking efficient algorithm support when processing multi-language and multi-domain data.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a cloud analysis data asset classification and grading management system, including a data acquisition layer, the data acquisition layer is connected to a metadata extraction layer and a preprocessing layer, the preprocessing layer is further connected to a classification and grading engine layer, the classification and grading engine layer is connected to a multi-language AI large model translation layer, and then connected to a storage layer, which is connected to an application layer;

[0006] The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism. It can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI large model translation layer uses a multilingual AI large model to perform semantic translation and understanding of the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural networks and multi-layer perceptrons to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.

[0007] The data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQL Server, PostgreSQL, as well as domestic databases such as Renmin University Jincang, DAMO, Nanda General, and Shenzhou General. The system utilizes data labeling technology, transforming the traditional "manual + tool" classification and grading model to an "AI + tool-in-one" model based on large AI models. This enables more efficient labeling and processing of classification and grading fields. The system features data source analysis, clearly demonstrating the source, flow, and complex relationships of data, assisting with data security assessments and decision-making. The system utilizes national cryptographic algorithm standards (national secret algorithms) for data encryption, ensuring the security and confidentiality of data during transmission. The system has built-in error handling and recovery mechanisms, automatically attempting to restore connections in the event of network anomalies or hardware failures, and recording detailed error logs to facilitate subsequent troubleshooting and maintenance. The system's functions encompass not only the identification and organization of data assets, as well as classification and grading, but also data flow detection, system vulnerability detection, API interface security testing, and data analysis, aiming to comprehensively improve the security and efficiency of data management. The system combines metadata extraction technology, translation and semantic understanding technology supported by large AI models, as well as graph neural network and multi-layer perceptron technologies to build an integrated platform for intelligent processing of multi-language, multi-source heterogeneous data, realizing the safe and efficient management of data assets.

[0008] The present invention provides a cloud analysis data asset classification and grading management system. It has the following beneficial effects:

[0009] The present invention provides a cloud analysis data asset classification and grading management system. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a schematic diagram of the system framework flow of the present invention;

[0011] Figure 2 This is a schematic diagram of the system framework algorithm flow of the present invention. DETAILED DESCRIPTION

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0013] like Figure 1 As shown, an embodiment of the present invention provides a cloud-based data asset classification and grading management system, including a data acquisition layer, the data acquisition layer is connected to a metadata extraction layer and a preprocessing layer, the preprocessing layer is further connected to a classification and grading engine layer, the classification and grading engine layer is connected to a multi-language AI large model translation layer, and then connected to a storage layer, and the storage layer is connected to an application layer;

[0014] The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism, which can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI large model translation layer uses a multilingual AI large model to perform semantic translation and understanding on the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural network and multi-layer perceptron to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.

[0015] Preferably, the data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQL Server, PostgreSQL, and domestic databases such as Renmin University of China Jincang, DAMO, Nanda Universal, and Shenzhou Universal.

[0016] Preferably, the system uses data labeling technology to convert the traditional "manual + tool" classification and grading model into an "AI + tool integrated" model based on the AI ​​big model, so that the system can label and process classification and grading fields more efficiently.

[0017] Preferably, the system has a data source analysis function, which can clearly display the source, flow and complex relationships of data, and assist in data security assessment and decision-making processes.

[0018] Preferably, the system uses the national cryptographic algorithm standard (national cryptographic algorithm) to encrypt data to ensure the security and confidentiality of data during transmission.

[0019] Preferably, the system has a built-in error handling and recovery mechanism, which can automatically attempt to restore the connection when encountering network anomalies or hardware failures, and record detailed error logs to facilitate subsequent troubleshooting and maintenance.

[0020] Preferably, the system functions not only cover the identification and sorting, classification and grading of data assets, but also include data flow detection, system vulnerability detection, API interface security detection, and data analysis, aiming to comprehensively improve the security and efficiency of data management.

[0021] Preferably, the system combines metadata extraction technology, translation and semantic understanding technology supported by AI large models, and graph neural network and multi-layer perceptron technologies to build an integrated platform for intelligent processing of multi-language, multi-source heterogeneous data, thereby achieving safe and efficient management of data assets.

[0022] Specifically, the cloud analysis data asset classification and hierarchical management system has the following advantages: Multi-database support: The system supports 32 mainstream database systems, such as MySQL, Oracle, SQLServer, PostgreSQL, etc., as well as various domestic databases, such as Renmin University of China Golden Warehouse, Dream, Nandatong Universal, and God Universal, etc. This extensive database support enables the system to adapt to various data environments, improving its applicability and flexibility. Efficiency: The system optimizes data reading and writing speed, reduces network delay and processing time, and improves overall application performance. This makes data asset management more efficient, especially in large-scale data scenarios. Security: Encryption technology is used to protect data transmission and storage security, while providing user authentication and permission management mechanisms to ensure that only authorized users can access sensitive data. Error handling and recovery: The system has a good error handling mechanism that can automatically attempt to reconnect when encountering network interruptions or other abnormal situations, and records error logs for problem troubleshooting. This improves the stability and reliability of the system. Intelligent data processing: It has the functions of automatic data asset discovery, data source identification, database information intelligent completion, and intelligent classification and grading. It uses the "cloud analysis big model" to use machine learning algorithms and rule engines to accurately classify and grade data. The system continuously learns to improve the accuracy and efficiency of classification, achieving self-optimization. Data blood source analysis: The system can clearly present the source and flow of data to assist data security evaluation decisions. This helps enterprises better understand and manage the entire life cycle of data, improving the overall level of data security. High availability: It adopts a multi-layer architecture, with a data collection layer supporting more than 50 databases through various interfaces and protocols, a preprocessing layer for data cleaning, conversion, and standardization, a classification and grading engine layer using the "cloud analysis big model" for intelligent classification and grading, and a storage layer using distributed storage technology to ensure reliable storage and efficient access of data. Integrated security management system: To address data leakage, data tampering, and illegal access, the system builds a security management system that integrates data asset identification and analysis, data classification and grading, data flow detection, vulnerability detection, data API detection, data analysis, access control, and identity authentication. This makes the system not only capable of managing data assets, but also provides comprehensive security protection. Through these advantages, the cloud analysis data asset classification and hierarchical management system can better address the challenges of data security, providing efficient, accurate, and secure solutions to help government and enterprise units effectively manage data assets.

[0023] Example 2

[0024] As Figure 2As shown, an embodiment of the present invention provides an algorithm for the Cloud Analysis Data Asset Classification and Grading Management System. Through innovative algorithm design, combined with the translation, classification, and grading capabilities of AI big models, it achieves efficient processing of multi-source heterogeneous data. The core of the invention lies in a unique metadata extraction and classification and grading algorithm that automatically extracts key metadata from raw data and uses AI big models to perform semantic translation and understanding, ultimately achieving intelligent classification and grading.

[0025] Algorithm Description:

[0026] Metadata extraction module:

[0027] This module uses a deep learning-based metadata extraction algorithm to automatically identify key metadata fields in raw data. The algorithm uses a pre-trained neural network model and contextual information to accurately locate the location and content of metadata. The specific steps are as follows:

[0028] 1. Obtain raw data through self-developed structured and unstructured data components, extract features through pre-trained deep learning models, and keep the metadata unchanged.

[0029] 2. Use the Attention Mechanism to identify key metadata fields in the data.

[0030] 3. Output the extracted metadata for subsequent modules to process.

[0031] Algorithm and formula:

[0032] 1. Data preprocessing:

[0033]

[0034] Algorithms: data cleaning, formatting, standardization, etc.

[0035] Input: Raw data (text, images, tables, etc.).

[0036] Output: preprocessed data

[0037] 2. Deep learning model feature extraction:

[0038]

[0039] Algorithm: Use pre-trained deep learning models (such as BERT, CNN, etc.) to extract features.

[0040] Input: preprocessed data

[0041] Output: Feature Representation dimension is where is the number of data samples, is the feature dimension.

[0042] Attention mechanism:

[0043]

[0044]

[0045] Algorithm: Attention Mechanism.

[0046] Input: Feature representation

[0047] Parameter: Learnable weight matrix

[0048] Output: Extracted metadata dimension is

[0049] AI Large Model Translation Module:

[0050] This module uses a multilingual AI large model to perform semantic translation and understanding on the extracted metadata. The algorithm achieves this through the following steps:

[0051] 1. Input the extracted metadata into the multilingual AI large model for semantic analysis and translation.

[0052] 2. Utilize the model's contextual understanding capabilities to ensure the accuracy and consistency of the translation results.

[0053] 3. Output the translated metadata for use by the classification and grading module.

[0054] Algorithm and Formula:

[0055] 1. Semantic Translation:

[0056]

[0057] Algorithm: Use a multilingual AI large model (such as GPT, BERT, etc.) for translation.

[0058] Input: Extracted metadata

[0059] Output: Translated metadata dimension is

[0060] 2. Contextual Understanding:

[0061]

[0062] Algorithm: Leverage the model's contextual information to optimize translation results.

[0063] Input: Translated metadata

[0064] Output: optimized translation metadata

[0065] Classification and Grading Module:

[0066] This module uses a classification and grading algorithm based on a graph neural network (GNN), which can intelligently classify and grade metadata based on its semantic information. The algorithm is unique in that it combines the semantic features of metadata with graph structure information. The specific steps are as follows:

[0067] 1. Construct the translated metadata into a graph structure, where nodes represent metadata fields and edges represent semantic relationships between fields.

[0068] 2. Use graph neural networks to extract features from graph structures and combine them with the semantic information of metadata to generate feature vectors for classification and grading.

[0069] 3. Classify and grade the feature vectors through a multi-layer perceptron (MLP) and output the final result.

[0070] Algorithm and formula:

[0071] 1. Graph structure construction:

[0072]

[0073] Algorithm: Build a graph structure, where nodes represent metadata fields and edges represent semantic relationships between fields.

[0074] Input: Translated metadata

[0075] Output: graph structure

[0076] 2. Graph Neural Network Feature Extraction:

[0077]

[0078] Algorithm: Use graph neural network (GNN) to extract feature representation of graph structure.

[0079] Input: graph structure

[0080] Output: Graph feature representation The dimension is in is the graph feature dimension.

[0081] 3. Classification and grading:

[0082]

[0083] Algorithm: Use multi-layer perceptron (MLP) for classification and grading.

[0084] Input: graph feature representation

[0085] Output:

[0086] Classification results The dimension is in is the number of categories.

[0087] Grading results The dimension is in is the level number.

[0088] Uniqueness of the patent:

[0089] Accuracy of metadata extraction

[0090] By combining deep learning with the attention mechanism, the algorithm can accurately identify key metadata fields in complex data, avoiding the problems of mis-extraction and missed extraction in traditional methods.

[0091] Intelligent multilingual translation

[0092] By utilizing a large multilingual AI model, the algorithm can achieve intelligent translation and understanding of multilingual metadata, ensuring the accuracy of classification and grading.

[0093] Graph structure optimization for classification and grading

[0094] By combining graph neural networks with multi-layer perceptrons, the algorithm can fully utilize the semantic relationships and graph structure information of metadata to achieve more accurate classification and grading.

[0095] Application Scenario

[0096] The present invention can be widely used in enterprise data management, financial risk control, medical data classification, government data governance and other fields, and is particularly suitable for the intelligent processing of multi-language, multi-source heterogeneous data.

[0097] Through innovative algorithm design, combined with metadata extraction, AI large model translation and graph neural network classification and grading technology, this invention provides an efficient and intelligent data classification and grading method and management system with broad application prospects and market value.

[0098] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. The Cloud Analysis Data Asset Classification and Hierarchy Management System includes a data collection layer, which features: The data collection layer is connected to the metadata extraction layer and the preprocessing layer, the preprocessing layer is further connected to the classification and grading engine layer, the classification and grading engine layer is connected to the multi-language AI large model translation layer, and then connected to the storage layer, and the storage layer is connected to the application layer; The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism, which can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI large model translation layer uses a multilingual AI large model to perform semantic translation and understanding on the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural network and multi-layer perceptron to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.

2. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQL Server, PostgreSQL, as well as domestic databases such as Renmin University of China Jincang, DAMO, Nanda Universal, and Shenzhou Universal.

3. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system uses data labeling technology to convert the traditional "manual + tool" classification and grading model into an "AI + tool integrated" model based on the AI ​​big model, enabling the system to label and process classification and grading fields more efficiently.

4. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system has the function of data blood source analysis, which can clearly display the source, flow and complex relationships of data, and assist in data security assessment and decision-making processes.

5. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system uses national cryptographic algorithm standards to encrypt data to ensure the security and confidentiality of data during transmission.

6. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system has a built-in error handling and recovery mechanism that can automatically attempt to restore the connection when encountering network anomalies or hardware failures, and record detailed error logs to facilitate subsequent troubleshooting and maintenance.

7. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system functions not only cover the identification, organization, classification and grading of data assets, but also include data flow detection, system vulnerability detection, API interface security detection and data analysis.

8. The Yunxi data asset classification and grading management system according to claim 1 is characterized by: The system combines metadata extraction technology, translation and semantic understanding technology supported by large AI models, as well as graph neural network and multi-layer perceptron technology.