Cloud analysis data asset classification and level-to-level management system
By adopting deep learning, AI big model translation and graph neural network technologies in the cloud analysis data asset classification and grading management system, the problem of the lack of efficient algorithm support in the existing technology when processing multilingual and multi-field data is solved, and the intelligent classification and grading of data assets is realized, and the security and efficiency of data management are improved.
Patent Information
- Application Number
- CN202510134690.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks efficient algorithm support when processing multilingual and multi-domain data, making it difficult to deal with complex, multi-source, and heterogeneous data environments.
A cloud analysis data asset classification and hierarchical management system was designed, and a metadata extraction algorithm based on deep learning and attention mechanism was adopted, combining multilingual AI big model translation and a classification and hierarchical algorithm based on graph neural networks and multi-layer perception machines to realize intelligent data classification and hierarchy.
The system can efficiently process multilingual and multi-source heterogeneous data, realize intelligent classification and grading of data assets, and improve the security and efficiency of data management.
Smart Images

Figure CN119989098A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of asset management, and in particular to a cloud analysis data asset classification and grading management system. Background Art
[0002] Existing technologies usually use manual labor and simple tools to achieve data asset classification and grading. In terms of data asset management, enterprises will first conduct preliminary identification and sorting of data assets through manual inventory. In order to improve efficiency, they may use some simple data management tools to assist manual inventory and conduct preliminary classification and grading of data. In the process of data classification and grading, data is usually classified and graded manually according to predetermined classification standards and grading rules through manual classification. Some tools may use simple rule engines to automatically classify and grade data according to preset rules, but these rules usually need to be set and maintained manually. The formulation and implementation of security policies are also mainly manual. After the data classification and grading is completed, the enterprise will manually formulate corresponding security policies based on the classification results. Then, some security tools are used to implement these policies, such as data encryption and access control.
[0003] With the advent of the big data era, the amount of data faced by enterprises, institutions and individuals is growing exponentially. How to effectively classify, grade and manage this data has become an urgent problem to be solved. Traditional classification and grading methods rely on manual rules and simple machine learning models, which are difficult to cope with complex, multi-source and heterogeneous data environments. In addition, although the existing AI big models have powerful semantic understanding and translation capabilities, they still have limitations in the application of metadata extraction and classification and grading, especially when processing multi-language and multi-domain data, there is a lack of efficient algorithm support. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a cloud analysis data asset classification and grading management system, which solves the problem of the prior art lacking efficient algorithm support when processing multi-language and multi-domain data.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a cloud analysis data asset classification and grading management system includes a data collection layer, the data collection layer is connected to a metadata extraction layer and a preprocessing layer, the preprocessing layer is further connected to a classification and grading engine layer, the classification and grading engine layer is connected to a multi-language AI large model translation layer, and then connected to a storage layer, and the storage layer is connected to an application layer;
[0006] The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism. It can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI big model translation layer uses the multilingual AI big model to perform semantic translation and understanding of the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural network and multi-layer perceptron to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.
[0007] The data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQLServer, PostgreSQL, and domestic databases such as Renmin University Jincang, DAMO, Nanda General, and Shenzhou General. The system uses data labeling technology to convert the traditional "manual + tool" classification and grading mode into an "AI + tool integration" mode based on the AI big model, so that the system can more efficiently label and process the classification and grading fields. The system has the function of data blood source analysis, which can clearly display the source, flow and complex relationship of data, and assist in data security assessment and decision-making process. The system uses the national cryptographic algorithm standard (national secret algorithm) for data encryption to ensure the security and confidentiality of data during transmission. The system has a built-in error handling and recovery mechanism, which can automatically try to restore the connection when encountering network anomalies and hardware failures, and record detailed error logs for later troubleshooting and maintenance. The system functions not only cover the identification and sorting of data assets, classification and grading, but also include data flow detection, system vulnerability detection, API interface security detection, and data analysis, aiming to comprehensively improve the security and efficiency of data management. The system combines metadata extraction technology, translation and semantic understanding technology supported by AI large models, as well as graph neural network and multi-layer perceptron technology to build an integrated platform for intelligent processing of multi-language, multi-source heterogeneous data, achieving safe and efficient management of data assets.
[0008] The present invention provides a cloud analysis data asset classification and grading management system. It has the following beneficial effects:
[0009] The present invention provides a cloud analysis data asset classification and grading management system. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a schematic diagram of the system framework flow of the present invention;
[0011] Figure 2 It is a schematic diagram of the system framework algorithm flow of the present invention. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] like Figure 1 As shown, an embodiment of the present invention provides a cloud analysis data asset classification and grading management system, including a data collection layer, the data collection layer is connected to a metadata extraction layer and a preprocessing layer, the preprocessing layer is further connected to a classification and grading engine layer, the classification and grading engine layer is connected to a multi-language AI large model translation layer, and then connected to a storage layer, and the storage layer is connected to an application layer;
[0014] The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism, which can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI large model translation layer uses the multilingual AI large model to perform semantic translation and understanding on the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural network and multi-layer perceptron to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.
[0015] Preferably, the data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQLServer, PostgreSQL, and domestic databases such as Renmin University of China Jincang, DAMO, Nanda General, and Shenzhou General.
[0016] Preferably, the system uses data labeling technology to convert the traditional "manual + tool" classification and grading model into an "AI + tool integration" model based on the AI big model, so that the system can label and process classification and grading fields more efficiently.
[0017] Preferably, the system has a data source analysis function, which can clearly display the source and flow of data and the complex relationship between them, and assist in data security assessment and decision-making process.
[0018] Preferably, the system uses the national cryptographic algorithm standard (national cryptographic algorithm) to encrypt data to ensure the security and confidentiality of data during transmission.
[0019] Preferably, the system has a built-in error handling and recovery mechanism, which can automatically attempt to restore the connection when encountering network anomalies or sudden hardware failures, and record detailed error logs to facilitate subsequent troubleshooting and maintenance.
[0020] Preferably, the system functions not only cover the identification, sorting, classification and grading of data assets, but also include data flow detection, system vulnerability detection, API interface security detection, and data analysis, aiming to comprehensively improve the security and efficiency of data management.
[0021] Preferably, the system combines metadata extraction technology, translation and semantic understanding technology supported by AI large models, and graph neural network and multi-layer perceptron technology to build an integrated platform for intelligent processing of multi-language, multi-source heterogeneous data, thereby achieving safe and efficient management of data assets.
[0022] Specifically, the advantages of the cloud analysis data asset classification and grading management system of the present invention include multi-database support: the system supports 32 mainstream database systems, such as MySQL, Oracle, SQL Server, PostgreSQL, etc., as well as a variety of domestic databases, such as Renmin University Jincang, Dameng, Nanjing University General, Shenzhou General, etc. This wide range of database support enables the system to adapt to a variety of data environments, improving its applicability and flexibility. Efficiency: The system optimizes the speed of data reading and writing, reduces network latency and processing time, and improves the overall performance of the application. This makes data asset management more efficient, especially in large-scale data scenarios. Security: Encryption technology is used to protect the security of data transmission and storage, and user authentication and permission management mechanisms are provided to ensure that only authorized users can access sensitive data. Error handling and recovery: The system has a good error handling mechanism. When encountering network interruptions or other abnormal situations, it can automatically try to reconnect and record error logs for troubleshooting. This improves the stability and reliability of the system. Intelligent data processing: It has the functions of automatic discovery of data assets, identification of data sources, intelligent completion of database information, and intelligent classification and grading. It uses the "cloud analysis big model" to accurately classify and grade data using machine learning algorithms and rule engines. The system continuously improves the accuracy and efficiency of classification through continuous learning and achieves self-optimization. Data blood source analysis: The system can clearly present the source and flow of data, and assist in data security assessment and decision-making. This helps enterprises better understand and manage the entire life cycle of data and improve the overall level of data security. High availability: It adopts a multi-layer architecture. The data acquisition layer supports more than 50 databases through multiple interfaces and protocols. The pre-processing layer cleans, converts and standardizes the data. The classification and grading engine layer uses the "Cloud Analysis Big Model" for intelligent classification and grading. The storage layer uses distributed storage technology to ensure reliable storage and efficient access to data. Integrated security management and control system: In response to problems such as data leakage, data tampering, and illegal data access, a security management and control system integrating data asset identification and sorting, data classification and grading, data traffic detection, vulnerability detection, data API detection, data analysis, access control, and identity authentication is constructed. This enables the system to not only manage data assets, but also provide multi-faceted security protection. Through these advantages, the Cloud Analysis Data Asset Classification and Grading Management System can better cope with the challenges faced by data security, provide efficient, accurate and secure solutions, and help government enterprises and institutions achieve effective management of data assets.
[0023] Example 2
[0024] like Figure 2As shown, the embodiment of the present invention provides an algorithm for the cloud analysis data asset classification and grading management system. Through innovative algorithm design, combined with the translation, classification and grading capabilities of the AI big model, efficient processing of multi-source heterogeneous data is achieved. The core of the present invention lies in a unique metadata extraction and classification and grading algorithm, which can automatically extract key metadata from raw data, and perform semantic translation and understanding through the AI big model, ultimately achieving intelligent classification and grading.
[0025] Algorithm Description:
[0026] Metadata extraction module:
[0027] This module uses a deep learning-based metadata extraction algorithm that can automatically identify key metadata fields in raw data. The algorithm uses a pre-trained neural network model combined with contextual information to accurately locate the location and content of metadata. The specific steps are as follows:
[0028] 1. Obtain original data through self-developed structured and unstructured data components, extract features through pre-trained deep learning models, and keep metadata unchanged.
[0029] 2. Use the Attention Mechanism to identify key metadata fields in the data.
[0030] 3. Output the extracted metadata for subsequent modules to process.
[0031] Algorithms and formulas:
[0032] 1. Data preprocessing:
[0033]
[0034] Algorithms: data cleaning, formatting, standardization, etc.
[0035] Input: Raw data (text, images, tables, etc.).
[0036] Output: Preprocessed data
[0037] 2. Deep learning model feature extraction:
[0038]
[0039] Algorithm: Use pre-trained deep learning models (such as BERT, CNN, etc.) to extract features.
[0040] Input: preprocessed data
[0041] Output: Feature Representation The dimension is in is the number of data samples, is the characteristic dimension.
[0042] Attention Mechanism:
[0043]
[0044]
[0045] Algorithm: Attention Mechanism.
[0046] Input: Feature Representation
[0047] Parameters: Learnable weight matrix
[0048] Output: Extracted metadata The dimension is
[0049] AI large model translation module:
[0050] This module uses a multilingual AI model to semantically translate and understand the extracted metadata. The algorithm is implemented through the following steps:
[0051] 1. Input the extracted metadata into the multilingual AI big model for semantic analysis and translation.
[0052] 2. Use the model’s contextual understanding capabilities to ensure the accuracy and consistency of translation results.
[0053] 3. Output the translated metadata for use by the classification and grading modules.
[0054] Algorithms and formulas:
[0055] 1. Semantic Translation:
[0056]
[0057] Algorithm: Use multilingual AI large models (such as GPT, BERT, etc.) for translation.
[0058] Input: Extracted metadata
[0059] Output: Translated metadata The dimension is
[0060] 2. Contextual understanding:
[0061]
[0062] Algorithm: Leverage the model’s contextual information to optimize translation results.
[0063] Input: Translated metadata
[0064] Output: optimized translation metadata
[0065] Classification and grading module:
[0066] This module uses a classification and grading algorithm based on a graph neural network (GNN), which can perform intelligent classification and grading based on the semantic information of metadata. The uniqueness of the algorithm lies in its combination of the semantic features of metadata and graph structure information. The specific steps are as follows:
[0067] 1. Construct the translated metadata into a graph structure, where nodes represent metadata fields and edges represent semantic relationships between fields.
[0068] 2. Use graph neural networks to extract features from graph structures and combine them with the semantic information of metadata to generate feature vectors for classification and grading.
[0069] 3. Classify and grade the feature vectors through a multi-layer perceptron (MLP) and output the final result.
[0070] Algorithms and formulas:
[0071] 1. Graph structure construction:
[0072]
[0073] Algorithm: Build a graph structure, where nodes represent metadata fields and edges represent semantic relationships between fields.
[0074] Input: Translated metadata
[0075] Output: Graph structure
[0076] 2. Graph Neural Network Feature Extraction:
[0077]
[0078] Algorithm: Use graph neural network (GNN) to extract feature representation of graph structure.
[0079] Input: Graph structure
[0080] Output: Graph feature representation The dimension is in It is the graph feature dimension.
[0081] 3. Classification and grading:
[0082]
[0083] Algorithm: Use multi-layer perceptron (MLP) for classification and grading.
[0084] Input: Graph feature representation
[0085] Output:
[0086] Classification results The dimension is in is the number of categories.
[0087] Grading results The dimension is in is the number of levels.
[0088] Uniqueness of patents:
[0089] Accuracy of metadata extraction
[0090] Through the combination of deep learning and attention mechanism, the algorithm can accurately identify key metadata fields in complex data, avoiding the problems of mis-extraction and missed extraction in traditional methods.
[0091] Intelligent multi-language translation
[0092] By utilizing a large multilingual AI model, the algorithm can achieve intelligent translation and understanding of multilingual metadata, ensuring the accuracy of classification and grading.
[0093] Graph structure optimization for classification and grading
[0094] Through the combination of graph neural network and multi-layer perceptron, the algorithm can make full use of the semantic relationship and graph structure information of metadata to achieve more accurate classification and grading.
[0095] Application Scenario
[0096] The present invention can be widely used in the fields of enterprise data management, financial risk control, medical data classification, government data governance, etc., and is particularly suitable for the intelligent processing of multi-language, multi-source heterogeneous data.
[0097] Through innovative algorithm design, combined with metadata extraction, AI large model translation and graph neural network classification and grading technology, the present invention provides an efficient and intelligent data classification and grading method and management system with broad application prospects and market value.
[0098] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. The Cloud Analysis data asset classification and grading management system includes a data collection layer, which is characterized by: The data collection layer is connected to the metadata extraction layer and the preprocessing layer, the preprocessing layer is further connected to the classification and grading engine layer, the classification and grading engine layer is connected to the multi-language AI large model translation layer, and then connected to the storage layer, and the storage layer is connected to the application layer; The metadata extraction layer is a metadata extraction algorithm based on deep learning and attention mechanism, which can automatically identify and extract key metadata fields from the original data and keep the metadata unchanged. The preprocessing layer cleans, converts and standardizes the collected data. The multilingual AI large model translation layer uses the multilingual AI large model to perform semantic translation and understanding on the extracted metadata to ensure the accuracy and consistency of the translation results. The classification and grading engine layer adopts an algorithm based on graph neural network and multi-layer perceptron to perform intelligent classification and grading according to the semantic information of metadata. The storage layer is a storage method based on distributed storage technology, and the application layer is a user interface that supports data query, analysis, and policy management functions.
2. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The data collection layer supports 32 mainstream database systems, including but not limited to MySQL, Oracle, SQLServer, PostgreSQL, and domestic databases such as Renmin University of China Jincang, DAMO, Nanda General, and Shenzhou General.
3. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system uses data labeling technology to convert the traditional "manual + tool" classification and grading model into an "AI + tool integrated" model based on the AI big model, enabling the system to label and process classification and grading fields more efficiently.
4. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system has the function of data blood source analysis, which can clearly display the source, flow and complex relationship of data, and assist data security assessment and decision-making process.
5. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system adopts national cryptographic algorithm standards to encrypt data to ensure the security and confidentiality of data during transmission.
6. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system has a built-in error handling and recovery mechanism, which can automatically attempt to restore the connection when encountering network anomalies or hardware failures, and record detailed error logs to facilitate subsequent troubleshooting and maintenance.
7. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system functions not only cover the identification, sorting, classification and grading of data assets, but also include data flow detection, system vulnerability detection, API interface security detection and data analysis.
8. The cloud analysis data asset classification and grading management system according to claim 1 is characterized by: The system combines metadata extraction technology, translation and semantic understanding technology supported by large AI models, and graph neural network and multi-layer perceptron technology.
Citation Information
Cited By
System and method for realizing data grading and classification based on metadata
CN120654064A