Intelligent data asset management system

The intelligent data asset management system solves the problems of data silos and scattered tags in enterprise data management, realizes unified management and efficient utilization of multimodal data, and improves the retrieval efficiency of data assets and business decision support capabilities.

CN121233654BActive Publication Date: 2026-07-24CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD
Filing Date
2025-09-28
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Enterprise data asset management suffers from data silos, inconsistent quality, and lack of standardized practices. Traditional software struggles to effectively utilize unstructured data, and the scattered labels across different data modalities make it difficult to form a unified business perspective, hindering the extraction of data value.

Method used

An intelligent data asset management system is adopted, including a data access layer, a data processing layer, an intelligent analysis layer, and an application service layer. Through format standardization, modal feature extraction, intelligent tag recommendation, multi-dimensional tag generation, and knowledge graph construction, unified management and intelligent question answering of data assets are achieved.

Benefits of technology

Break down the format barriers between structured, semi-structured, and unstructured data, generate accurate multimodal data asset tags, improve data asset retrieval efficiency and cross-modal correlation analysis capabilities, and achieve dynamic control of data quality and efficient service for business decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233654B_ABST
    Figure CN121233654B_ABST
Patent Text Reader

Abstract

The application relates to an intelligent data asset management system. The system comprises a data access layer, a data processing layer, an intelligent analysis layer and an application service layer; the intelligent analysis layer comprises an intelligent tag recommendation unit, an intelligent cataloging unit, an intelligent management unit and an intelligent housekeeper unit; the data access layer is used for receiving multi-modal data assets and performing format standardization conversion to obtain standardized multi-modal data assets; the data processing layer is used for preprocessing the standardized multi-modal data assets and extracting modal features of each modal data; the intelligent housekeeper unit is used for inputting user query into an intelligent question and answer large language model, performing intelligent question and answer according to an intelligent directory and a knowledge graph, and outputting answer content; and the application service layer is used for calling the output of the intelligent analysis layer to provide data asset query services, quality monitoring services and decision support services for users. The method can output professional, accurate and efficient data asset data analysis reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data asset management technology, and in particular to an intelligent data asset management system. Background Technology

[0002] With the deepening of digital transformation, data has become a core driver of corporate decision-making. As electronic information generated during the business operations and production process, data itself does not participate in product production, but it contains immense value. Data assets help companies with cost management, risk control, and organizational development; activating corporate data assets can also drive corporate value and enhance core competitiveness. As a new type of core productivity, the assetization of data has become an inevitable trend.

[0003] However, current enterprise data asset management generally suffers from numerous limitations. On the one hand, problems such as data silos, inconsistent quality, and lack of standardized practices are prominent. Traditional business intelligence (BI) software can handle structured data, but its utilization rate of unstructured data such as text and images is low, and data governance is difficult. On the other hand, data assets come from multiple sources and are highly interrelated to business operations. Traditional multimodal labeling technologies do not consider the inherent connections between different modalities of data in business scenarios. Labels generated from different modalities may belong to different dimensions, making it difficult to form a unified business perspective and thus hindering the extraction of data value. Therefore, an automated, intelligent, and convenient data asset management system is needed. Summary of the Invention

[0004] Therefore, it is necessary to provide an intelligent data asset management system to address the aforementioned technical problems.

[0005] An intelligent data asset management system, the system comprising: The system comprises a data access layer, a data processing layer, an intelligent analysis layer, and an application service layer; the intelligent analysis layer includes an intelligent tag recommendation unit, an intelligent cataloging unit, an intelligent management unit, and an intelligent concierge unit. The data access layer is used to receive multimodal data assets and perform format standardization conversion to obtain standardized multimodal data assets; the multimodal data assets include structured data, semi-structured data, and unstructured data; The data processing layer is used to preprocess the standardized multimodal data assets and extract the modal features of each modal data. The intelligent tag recommendation unit is used to embed the modal features and data metadata of the multimodal data assets into the tag generation prompt template, input it into the tag recommendation language model, output the corresponding multimodal data asset tags, and filter high-quality tag data to output to the intelligent cataloging unit; the tag generation prompt template includes metadata constraints, mapping examples of each modal feature and tag, and conflict resolution instructions for different modal tags; The intelligent cataloging unit is used to train a pre-built multi-class head classification model based on high-quality tag data, input the modal features of the multi-modal data assets to be classified into the trained multi-class head classification model, generate multi-dimensional tags, and construct an intelligent catalog and knowledge graph based on the multi-dimensional tag data; The intelligent management unit is used to generate quality rules and quality labels according to preset prompt templates, calculate the comprehensive confidence level of the labels, and feed back optimization instructions to the upstream unit based on the confidence level to generate quality reports; The intelligent butler unit is used to input user questions into the intelligent question-and-answer language model, perform intelligent question-and-answer based on the intelligent catalog and knowledge graph, and output the answer content. The application service layer is used to call the output of the intelligent analysis layer to provide users with data asset query services, quality monitoring services, and decision support services.

[0006] The aforementioned intelligent data asset management system, through the data access layer's standardized format conversion of multimodal data assets and the data processing layer's preprocessing and modal feature extraction, breaks down the format barriers between structured, semi-structured, and unstructured data, constructing a unified data management foundation. The intelligent tag recommendation unit embeds modal features and metadata into a prompt template containing metadata constraints, mapping examples, and conflict resolution instructions, inputting it into a tag recommendation language model to generate accurate multimodal data asset tags, solving the problems of scattered tag dimensions and modal fragmentation. The intelligent cataloging unit trains a multi-class head classification model with high-quality tags to generate multi-dimensional tags, thereby constructing an intelligent catalog and knowledge graph, improving data asset retrieval efficiency and cross-modal correlation analysis capabilities. The intelligent management unit generates quality rules, calculates tag comprehensive confidence, and provides feedback optimization instructions, enabling dynamic control and continuous optimization of data quality. The intelligent steward unit, combined with the intelligent catalog and knowledge graph, enables intelligent question answering, and, along with the query, monitoring, and decision-making services provided by the application service layer, lowers the barrier to entry for non-technical personnel, allowing the value of data assets to efficiently serve business decisions. Attached Figure Description

[0007] Figure 1 This is a structural block diagram of an intelligent data asset management system in one embodiment; Figure 2 Here is a flowchart for the automatic generation of quality reports in one embodiment. Detailed Implementation

[0008] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0009] In one embodiment, such as Figure 1 As shown, an intelligent data asset management system is provided, including a data access layer, a data processing layer, an intelligent analysis layer, and an application service layer; the intelligent analysis layer includes an intelligent tag recommendation unit, an intelligent cataloging unit, an intelligent management unit, and an intelligent steward unit; The data access layer is used to receive multimodal data assets and perform format standardization conversion to obtain standardized multimodal data assets; multimodal data assets include structured data, semi-structured data, and unstructured data; The data processing layer is used to preprocess standardized multimodal data assets and extract modal features of each modality. The intelligent tag recommendation unit is used to embed the modal features and metadata of multimodal data assets into the tag generation prompt template, then input it into the tag recommendation language model, output the corresponding multimodal data asset tags, and filter high-quality tag data to output to the intelligent cataloging unit; the tag generation prompt template includes metadata constraints, mapping examples of each modal feature and tag, and conflict resolution instructions for different modal tags; The intelligent cataloging unit is used to train a pre-built multi-class head classification model based on high-quality label data. The modal features of the multi-modal data assets to be classified are input into the trained multi-class head classification model to generate multi-dimensional labels. Intelligent catalogs and knowledge graphs are constructed based on the multi-dimensional label data. The intelligent management unit is used to generate quality rules and quality labels based on preset prompt templates, calculate the comprehensive confidence level of the labels, and send optimization instructions to the upstream unit based on the confidence level to generate quality reports; The intelligent butler unit is used to input user questions into the intelligent question-answering language model, perform intelligent question-answering based on the intelligent catalog and knowledge graph, and output the answer content; The application service layer is used to call the output of the intelligent analysis layer to provide users with data asset query services, quality monitoring services, and decision support services.

[0010] In the aforementioned intelligent data asset management system, the data access layer standardizes the format conversion of multimodal data assets, while the data processing layer performs preprocessing and modal feature extraction. This breaks down the format barriers between structured, semi-structured, and unstructured data, building a unified data management foundation. The intelligent tag recommendation unit embeds modal features and metadata into a prompt template containing metadata constraints, mapping examples, and conflict resolution instructions, inputting it into a tag recommendation language model. This generates accurate multimodal data asset tags, solving the problems of scattered tag dimensions and modal fragmentation. The intelligent cataloging unit trains a multi-class head classification model with high-quality tags to generate multi-dimensional tags, thereby constructing an intelligent catalog and knowledge graph, improving data asset retrieval efficiency and cross-modal correlation analysis capabilities. The intelligent management unit generates quality rules, calculates the comprehensive confidence score of tags, and provides feedback optimization instructions, enabling dynamic control and continuous optimization of data quality. The intelligent steward unit, combined with the intelligent catalog and knowledge graph, provides intelligent question answering, along with query, monitoring, and decision-making services provided by the application service layer. This lowers the barrier to entry for non-technical personnel, allowing the value of data assets to efficiently serve business decisions.

[0011] In one embodiment, embedding the modal features and metadata of a multimodal data asset into a tag generation prompt template and then inputting it into a tag recommendation language model to output the corresponding multimodal data asset tag includes: converting the modal features and metadata of the multimodal data asset into a format parsable by the tag recommendation language model; embedding the converted modal features and metadata into the tag generation prompt template; inputting the embedded tag generation prompt template into the tag recommendation language model to generate the original tag; and fusing the original tag according to the preset conflict handling rules in the tag generation prompt template to obtain the corresponding multimodal data asset tag.

[0012] In one embodiment, training a pre-built multi-class head classification model based on high-quality labeled data, and inputting the modal features of the multimodal data asset to be classified into the trained multi-class head classification model to generate multi-dimensional labels includes: constructing a multi-class head classification model; the multi-class head classification model includes a multimodal parallel encoding module, a feature fusion module, and a multi-class head output module; the multimodal parallel encoding module is used to branch and encode the modal features of the multimodal data asset to obtain the encoded features corresponding to each modality; the feature fusion module is used to fuse the encoded features of each modality to obtain multimodal fused features; the multi-class head output module is used to classify the multimodal fused features and output multi-dimensional labels; training the multi-class head classification model based on high-quality labeled data to obtain a trained multi-class head classification model; and inputting the modal features of the multimodal data asset to be classified into the trained multi-class head classification model to obtain the corresponding multi-dimensional labels.

[0013] In one embodiment, the intelligent tag recommendation unit includes a multimodal data unified tag module, an automated tag recommendation module, and a tag quality management module. The multimodal data unified tag module is used to construct a standardized tag system. The standardized tag system is used to standardize and unify the tag dimensions of multimodal data assets. The automated tag recommendation module is used to call the standardized tag system, input the modal features of each multimodal data asset and the pre-set tag generation prompt template into the tag recommendation language model to obtain multimodal data asset tags. The tag generation prompt template includes metadata constraints, tag generation examples corresponding to each modality of data, and tag conflict resolution instructions for different modalities. The metadata constraints include that all tags must be associated with multiple tag dimensions in the business rules, and the tag dimensions are standardized by the standardized tag system. The tag quality management module is used to perform quality detection on the multimodal data asset tags according to the pre-set tag quality assessment rules and filter out high-quality tag data.

[0014] In one embodiment, the intelligent cataloging unit includes an automatic classification module, an intelligent catalog construction module, and a graph-based management module. The automatic classification module is used to train a pre-built multi-class head classification model using high-quality tag data output by the intelligent tag recommendation unit, inputting the modal features of the data assets to be classified into the trained multi-class head classification model to generate multi-dimensional tags. The intelligent catalog construction module is used to construct a multi-level catalog based on the multi-dimensional tags and visualize it. When the tags are updated, the catalog is dynamically adjusted by calculating the similarity between the asset location vector and the catalog category, and the data association is synchronized to the graph-based management module. The graph-based management module is used to construct an asset graph based on the data association and the multi-modal data assets marked by multi-dimensional tags and visualize it.

[0015] In one embodiment, constructing an asset graph based on data associations and multimodal data assets labeled with multidimensional tags includes: defining node attributes of the asset graph based on the multimodal data assets labeled with multidimensional tags; defining edge attributes of the asset graph based on data associations; constructing a node set of the asset graph based on the node attributes and the multimodal data assets labeled with predicted tags; constructing an initial edge set of the asset graph based on the edge attributes and data associations; forming an initial structure of the asset graph based on the node set and the initial edge set; calculating the tag similarity between unrelated data assets in the initial structure to determine potential associations; updating the edge set of the asset graph based on the potential associations; and obtaining the asset graph based on the updated node set and edge set.

[0016] In one embodiment, the intelligent management unit includes a quality rule auto-configuration module, a quality report auto-generation module, and an improvement suggestion provision module. The quality rule auto-configuration module receives multimodal data asset tags, data metadata, and business domain relationships output by the intelligent tag recommendation unit, and combines them with preset prompt templates to introduce quality attributes such as completeness, timeliness, and downstream dependency risk, and binds them to business rules. It automatically generates quality rules, tag dispersion calculation parameters, and quality tag generation logic, calculates the comprehensive confidence level of tags, and feeds back optimization instructions to the upstream unit. The quality report auto-generation module calculates data asset quality scores, counts the amount of problem data, and analyzes quality trends based on quality rules, confidence level results, and the intelligent catalog, and automatically generates visualized quality reports. The improvement suggestion provision module receives quality-related parameters output by the quality rule auto-configuration module, problem data details output by the quality report auto-generation module, and modal characteristics of multimodal data assets. It generates customized improvement suggestions for different quality problems, synchronizes the generated improvement suggestions to the quality reports, and feeds back high-frequency problems to the upstream unit.

[0017] In one embodiment, the overall confidence level of the labels is:

[0018] in, For the first Each quality dimension score For the weight of the quality dimension, For the dispersion of multimodal data asset labels, It is the first The original labels generated by each modality For modal types.

[0019] In one embodiment, calculating the overall confidence score of a tag and feeding back optimization instructions to the upstream unit includes: determining the quality dimensions and weights of each dimension required for calculating the overall confidence score of the tag; calculating the scores of each quality dimension and the dispersion of the multimodal data asset tags; calculating the overall confidence score of the tag based on the quality dimension scores, quality dimension weights, and the dispersion of the multimodal data asset tags; obtaining the relationship between the overall confidence score of the tag and a threshold; generating tag generation optimization instructions and data flow control instructions when the overall confidence score of the tag is lower than the threshold; generating tag generation optimization instructions when the overall confidence score of the tag is higher than or equal to the threshold; the tag generation optimization instructions are used to adjust the prompt template parameters or modal feature weights of the tag recommendation large language model to optimize the generation quality of multimodal data asset tags; the data flow control instructions are used to restrict the flow range of low-confidence data assets in the intelligent catalog or mark low-confidence data assets as pending review; feeding back the tag generation optimization instructions to the intelligent tag recommendation unit and feeding back the data flow control instructions to the intelligent cataloging unit.

[0020] In one embodiment, the intelligent management unit includes a natural language interaction module, an intelligent retrieval and recommendation module, an MCP technology application module, and an intelligent question-answering module. The natural language interaction module receives spoken questions from business personnel via the intelligent question-answering interface, uses a large model to analyze semantics, associates data asset libraries with data tags, and generates visualizations including charts and alerts. The intelligent retrieval and recommendation module retrieves data asset libraries containing directories and knowledge graphs based on user questions, recommends relevant data assets, and provides data visualization and analysis results. The MCP technology application module quickly retrieves the required data from the system using Model Context Protocol technology, performs semantic understanding and content extraction, and answers user questions in natural language. The intelligent question-answering module understands the user's question intent based on the large model's natural language processing capabilities, integrates asset information retrieved from the directory with the relationship and attribute information of the knowledge graph, and generates answer content.

[0021] In one specific embodiment, such as Figure 1 As shown, the system adopts a front-end and back-end separation model, using the Spring + MySQL + Redis + Vue framework. The system includes a data access layer, a data processing layer, an intelligent analysis layer, and an application service layer, with the functions of each layer as follows: The data access layer is responsible for collecting metadata from various business systems and external data sources of the enterprise, and performing preliminary cleaning and format conversion to provide a unified data format for subsequent processing; The data processing layer includes functions such as data storage, data preprocessing, and feature extraction, providing structured input data for large AI models. Feature extraction is used to extract features from the preprocessed data assets (structured data / semi-structured data / unstructured data) according to their respective modalities, outputting structured modal features / semi-structured modal features / unstructured modal features. The intelligent analysis layer is the core of the system, including units such as intelligent cataloging of data assets, intelligent classification and tagging of data assets, intelligent management of data quality, and intelligent data asset steward, all of which are implemented based on a large AI model; the application service layer provides functions such as user interaction interface, data visualization, report generation, and intelligent Q&A, enabling users to easily use the services provided by the system.

[0022] Specifically, within the intelligent analysis layer, the data asset intelligent tag recommendation unit includes a multimodal data unified tagging module, an automated tag recommendation module, and a tag quality management module. Among these: The multimodal data unified tagging module constructs a standardized tagging system covering structured, semi-structured, and unstructured data (including dimensions such as technical features, business features, and security classifications). While transferring this tagging system framework to the automated tag recommendation module, it also synchronizes quality assessment-related tag dimensions (such as integrity requirements and timeliness standards) to the intelligent management unit, serving as the foundational dimension reference for its quality rule configuration. It supports unified tagging management for structured data (such as sales order tables containing fields like order_id and amount), semi-structured data (such as API logs "GET / api / orders?status=pending"), and unstructured data (such as data documents like "This table is updated daily at midnight via ETL for generating downstream visitor visitor metrics"), achieving standardization and normalization of data assets.

[0023] The automated label recommendation module calls the standardized label system output by the unified label module for multimodal data. Based on the trained label recommendation big language model, it extracts core features from the preprocessed multimodal data, loads industry-specific prompt templates (including metadata rules and feature-label mapping examples) to generate original labels, resolves semantic conflicts through fusion rules to form fused labels, and sends the label results to the label quality management module for internal label quality verification before distributing them to the multidimensional intelligent recognition module and intelligent management unit.

[0024] Specifically, firstly, the system receives structured, semi-structured, and unstructured data that have undergone preprocessing and initial feature extraction by the data processing layer. The extracted features are then transformed into an input format recognizable by the large model, ensuring that features from different modalities can be uniformly parsed. The large model loads industry-specific prompt templates, which include metadata rules, mapping examples between features and labels for each modality, and conflict resolution instructions for different modal labels. The large model performs semantic understanding and rule matching on the features of each modality, generating original labels that conform to the template rules. By comparing the original labels of different modalities, semantic association conflicts and semantic contradiction conflicts are identified. For association conflicts: fused labels are generated using the fusion rules in the template; for contradiction conflicts: authoritative data sources or high-confidence labels are prioritized, and the fused labels are correlated and verified with the standardized label system provided by the unified label module for multimodal data. Missing technical features (such as "access permissions - department level") and business features (such as "data dependency - visitor indicator table") are supplemented, forming complete labels. These labels are then synchronized to the label quality management module for quality inspection and sent as core data to the intelligent management unit to support subsequent quality assessment and rule configuration.

[0025] The system can understand data content and semantics, automatically extract key features, and generate corresponding tags. These tags not only reveal the technical characteristics of the data, such as access frequency and access permissions, but also indicate the business characteristics, such as data dependencies, data business types, data operation dimensions, and sensitivity levels. For tags from different dimensions, such as high-frequency access (technical dimension), order status query (business operation dimension), and metric dependencies (business-affected dimension), the system integrates them according to the following dedicated prompt template: prompt_template = { # Data Asset Metadata Constraints "Metadata Rules": "All tags must be associated with the following dimensions in the 'Data Security Technology Data Classification and Grading Rules' (which uses the grading standard to automatically identify unit dependencies):\n" - Data Categories (User Data / Business Data / Operations Management Data / System Operations and Maintenance Data)\n "- Business domain (such as sales data, risk control data)\n" - Sensitivity Level (PII / PCI / Public)\n "- Criticality of the link (core / auxiliary)", # Multimodal dynamic fusion "Structured Data Hint": "Example of generating tags based on table fields and access frequency:\n" 1. If it contains a customer_id and the access frequency is >1k / day → 'PII Data - High Frequency'\n 2. If it is a primary-foreign key relationship table → 'Core Link' / 'Secondary Link', API Log Notification: "Inferring Business Attributes from Request Path and Parameters:\n" 1. If the path contains / orders and the parameter contains status → 'Sales Order Status Query'\n 2. If the response time > 500ms → 'Performance Sensitive' Document Tip: "Extract downstream impact and update frequency:\n" 1. If 'risk control' is mentioned → 'compliance-related'\n 2. If 'Daily ETL' is mentioned, it should be replaced with 'T+1 data'. # Conflict resolution instructions "Merge Rule": "If the structured data tag is 'PII Data' and the document tag is 'Compliance Related', then merge them into 'Sensitive Data - Strong Compliance and Regulation'." } In metadata rules, based on the metadata and business rules of data objects, data security classification rules are automatically identified and data risk levels are configured to obtain sensitivity levels. Business rules are pre-set security classification criteria, which can be based on laws and regulations, industry standards, or internal company policies.

[0026] Fusion Tags Also through formula Perform calculations. Among them... It is a predefined set of business tags, such as "log error" and "field missing"; It is the first The original labels generated by each modality; These are custom modal weights, determined by data quality or business rules, such as log data. ; It is a business tag With the original label Semantic similarity is calculated using knowledge graphs or word vectors; It is a matching indicator function, when and The value is 1 if there is a business logic relationship, and 0 otherwise. In data governance, if... =“Log error” ="Field missing", , , , The fusion result .

[0027] Different customized prompt templates can be created based on different industry data. For banking data, additional annotations can be added for fund flow data and transaction risk levels; for e-commerce data, additional annotations can be added for product unit price and product popularity; for healthcare data, additional annotations can be added for clinical priority and abnormal imaging examination data. The application also allows setting different confidence weights for technical metadata (such as access frequency) and business metadata (such as document descriptions), prioritizing the labels from the authoritative data source in case of label conflicts.

[0028] The tag quality management module receives fused tags from the automated tag recommendation module. After detecting conflicts and errors, it synchronizes the tag quality assessment results (such as tag confidence scores and high-frequency error types) to the intelligent management unit to help it more accurately assess the reliability of tag data. At the same time, it feeds back user correction records to the automated tag recommendation module to optimize the model, and feeds back quality problems of the tag system to the automatic identification module of the grading standard to iterate the grading rules.

[0029] The intelligent tagging recommendation unit for data assets uses a large AI model to automatically classify and tag data assets, eliminating the tedious manual tagging process. By integrating data tags from different dimensions, it further systematically solves the core problems of scattered tag dimensions and insufficient business alignment in data assets, laying the foundation for value mining and business applications of data assets.

[0030] The intelligent cataloging unit for data assets includes an automatic classification module, an intelligent catalog construction module, and a graph-based management module. Among them: The automatic classification module, based on a deep learning convolutional neural network algorithm, enables the system to automatically classify data assets. It uses quality label filtering to filter training data, employing only data with high-quality labels for training, and injects business label information into the classification layer of the convolutional neural network for weight initialization. Automatic classification is then achieved through the trained convolutional neural network.

[0031] Specifically, firstly, modal features of each modality of the data asset to be classified are obtained from the data processing layer and used as input to the parallel encoding layer. A three-level classification network is constructed: "Multimodal Parallel Feature Input and Preliminary Encoding Layer—CNN Feature Fusion and High-Dimensional Feature Extraction Layer—Multi-Classification Head Output Layer." The parallel encoding layer processes the input structured data modal features (after field embedding and statistical feature encoding), semi-structured log modal features (after key parameter extraction and sequence encoding), and unstructured document modal features (after keyword embedding and semantic encoding) through three independent branches, concatenating them into a multimodal fusion feature matrix. The CNN feature extraction layer contains three convolutional blocks (paired with pooling layers), extracting high-dimensional abstract feature vectors based on a pre-trained CNN variant, and injecting business label association rules during convolutional kernel weight initialization, prioritizing business-related features. The multi-classification head output layer sets independent classification heads containing fully connected layers and Softmax activation functions for each label dimension, outputting candidate labels and feature fitting scores for each dimension in parallel, and selecting the highest-scoring candidate labels. High-level labels are concatenated into multi-dimensional predicted labels. During the training phase, a joint loss function is used for optimization. This function is a weighted sum of the standard classification loss, business rule constraint loss, and orthogonal regularization loss. The standard classification loss uses multi-task cross-entropy loss to ensure the basic classification accuracy of each dimension label. The business rule constraint loss introduces a penalty term based on business rules (e.g., if core business domain data is predicted as "public" sensitive level, or non-core link data is predicted as "core" link criticality, the loss value is increased) to ensure that the model output conforms to business logic. The orthogonal regularization loss introduces a penalty term based on the independence of feature representations learned by different classifiers (by calculating the inner product of feature vectors of different classifiers, the larger the inner product, the higher the penalty), to avoid feature redundancy affecting the model's generalization ability. At the same time, the Adam optimizer, batch size of 32, and early stopping strategy are used to iteratively optimize until the average accuracy of the test set meets the threshold requirements. Finally, automatic classification of multimodal data assets and generation of multi-dimensional predicted labels are achieved, while ensuring the unity of model convergence and business compliance.

[0032] The intelligent directory building module automatically creates a data asset directory, comprehensively considering data tags and characteristics from the tagging system to construct a multi-level directory structure, and dynamically adjusts the directory structure according to the characteristics of the data assets and business needs. This directory building is not static but dynamic, automatically updating the directory based on real-time changes in the data assets.

[0033] Specifically, the system receives multimodal data features output from the data processing layer and high-confidence candidate labels (predicted labels) output from the automatic classification module. It then integrates these two types of information to construct a multi-level directory structure of "business domain - data type - technical characteristics". When data asset characteristics or business needs change, the system receives updated label results from the automatic classification module and updated multimodal data features from the data processing layer in real time. It calculates the similarity between the asset's directory location vector (based on a comprehensive representation of features and labels) and existing directory categories in the directory tree. Based on a preset similarity threshold (e.g., if the similarity is ≥0.8, the asset belongs to an existing node; if it is <0.8, a new node is created), the system determines the directory affiliation of the data asset, thereby dynamically adjusting the directory hierarchy and data affiliation relationship. At the same time, the system synchronizes the data association relationships in the directory (e.g., "sales data - order table" and "high-frequency access - core link data") to the graph management module, providing a foundation for these association relationships.

[0034] The graph-based management module adopts a graph-based rather than a linear management approach to establish interconnected relationships between isolated data assets. This is similar to typical big data recommendation algorithms, which recommend potential associations by calculating the similarity of tag sets between assets, and displaying related or interesting data assets in a linked manner.

[0035] Specifically, the graph-based management module receives data relationships output by the intelligent directory construction module and predicted labels output by the automatic classification module. It first generates a standardized label set for each data asset, including "business domain, sensitivity level, technical characteristics, and criticality of the link." Using data assets as nodes and known relationships as initial edges, it constructs a basic asset graph. The core module calculates the similarity of label sets between assets using a weighted Jaccard similarity algorithm (weighted according to dimensions such as business domain and sensitivity level, with 1 point awarded for consistent labels). A similarity threshold (e.g., ≥0.5) is set to filter out "no known relationships but highly similar labels." Asset pairs are considered potential associations. Similar assets are recommended to users who are interested in the asset pair, and the strength of the association is marked. When new assets are added or existing asset tags are updated, the similarity is recalculated in real time, and the graph edge attributes are dynamically adjusted (potential associations can be upgraded to known associations). Finally, a visual graph is displayed—node colors distinguish business domains, size reflects asset importance, and the solid / phasic nature (known / potential associations) and thickness (similarity level) of edges represent the association type and strength. Users can interactively view tag sets, association details, and manually confirm potential associations, thus achieving potential association recommendation and display through tag set similarity calculation.

[0036] The intelligent asset cataloging system uses AI big data models to automatically classify and catalog data assets, effectively building asset catalogs based on the multi-dimensional characteristics of data assets, thereby significantly improving the retrieval efficiency and management effectiveness of data assets for enterprises.

[0037] The intelligent management unit includes a quality rule automatic configuration module, a quality report automatic generation module, and an improvement suggestion provision module. Among them: The automatic configuration module for quality rules is based on the metadata and sample data of data assets. By introducing quality attributes into the prompt template and binding business rules, the system can automatically configure data quality rules.

[0038] Specifically, the automatic configuration module for quality rules takes multimodal data asset tags (derived from the intelligent tag recommendation unit) and business relationships (derived from the intelligent cataloging unit) as core inputs, and combines them with preset prompt templates (including rules such as quality constraints and data type prompts) to construct a full-process processing logic of input-generation-synchronization-evaluation-feedback. First, the module receives multimodal data asset tags (covering business / technology / security / initial quality dimensions), data metadata, and business domain relationships. Based on this, it introduces quality attributes such as integrity, timeliness, and downstream dependency risk, and binds them to business rules (e.g., core business domain data integrity has a higher weight). It automatically generates differentiated quality rules (including quality dimension weights and tag dispersion calculation parameters) and quality tag generation logic (e.g., merging rules when multimodal quality tags conflict). Next, the quality rules are synchronized to the automatic quality report generation unit, providing a basis for data quality scoring. Then, the module calculates the overall confidence level using a tag confidence formula (combining a weighted average quality score and a tag dispersion correction term). If a "low confidence" is determined, a "requires manual review" tag and a data flow blocking instruction are generated and fed back to the hierarchical and tagging unit (to optimize the tagging algorithm) and the intelligent cataloging unit (to restrict the flow of low-quality data). Simultaneously, the quality tag generation logic is passed to the improvement suggestion providing unit to pinpoint the root cause of the problem. This ensures that quality rules align with business needs, and that multimodal data asset tags and data quality are improved synergistically.

[0039] For the grading and labeling unit, the "low confidence" conclusion and data flow blocking instruction fed back by the quality rule auto-configuration module serve as optimization inputs for this unit. The purpose is to allow it to adjust the label generation algorithm based on the issue of "insufficient label confidence"—for example, when the confidence of multimodal data asset labels is low due to high dispersion, this unit will combine feedback to optimize the label extraction logic (such as adjusting the weight allocation of different modal labels and optimizing label conflict judgment rules), thereby improving the accuracy of subsequent label generation, rather than directly using it as the base data for label generation. For the intelligent cataloging unit, the "low confidence" conclusion and blocking instruction fed back by the quality rule auto-configuration module serve as control inputs for this unit. The purpose is to allow it to identify low-quality data during the cataloging process and restrict its flow based on the data's business domain relationships—for example, if a core piece of risk control domain data is determined to be low confidence, this unit will mark it as "pending review" in the catalog and block its flow to downstream business systems (such as the risk control decision module) to prevent low-quality data from affecting business applications, rather than using it to adjust the cataloging logic itself.

[0040] Data quality is based on the label confidence formula. Perform calculations, where For the first Scores for each quality dimension, such as completeness. , calculated as ; The weights for the quality dimension are set by business requirements. The dispersion of multimodal data asset labels is the degree of dispersion; the higher the dispersion, the greater the label conflict. ,in It is the number of all possible unique label categories. Then it is the first Class tags in The frequency of occurrence in each modality. Data with a confidence level below the threshold will be automatically labeled "low confidence - requires manual review", and the flow of low-quality labeled data will be automatically blocked by the system.

[0041] As shown above, the system identifies data quality anomalies by using prompt templates and data characteristics and distribution, and generates corresponding labels to provide a basis for subsequent processing.

[0042] The automatic quality report generation module can automatically generate quality reports based on the data quality investigation results, including data quality scores, problem data statistics, quality trend analysis, etc., to intuitively display the data quality status.

[0043] Specifically, the system receives the "Comprehensive Confidence Result of Quality Rules, Data, and Tags" output by the automatic quality rule configuration unit, and the "Data Classification Directory (e.g., divided by business domain / technical characteristics)" output by the intelligent data asset cataloging unit. Based on the quality rules, it calculates the quality score of each data asset (including basic data quality score and tag quality score), counts the amount of problematic data in different business domains (e.g., the proportion of low-confidence data in the sales domain, the amount of tag conflict data in the risk control domain), analyzes recent quality trends (e.g., data confidence curve, tag dispersion decreasing trend), and automatically generates a visual quality report containing "business domain - quality indicators - problem statistics - trend charts." The "details of problematic data (e.g., missing field types, tag conflict types)" in the report are synchronized to the improvement suggestion providing unit to provide data support for generating targeted solutions. At the same time, the quality report is fed back to the intelligent management unit, which uses its interactive functions to display the data and tag quality status to business personnel, supporting quality decision-making.

[0044] The improvement suggestion module provides targeted improvement suggestions based on detected data quality issues, including methods such as data cleaning, data completion, and data verification, to help users improve data quality.

[0045] Specifically, the system receives the "quality dimension weights, tag dispersion entropy values, and quality tag generation logic" output by the automatic quality rule configuration unit, and the "detailed problem data" output by the automatic quality report generation unit. Combined with the "data modal characteristics (such as structured / semi-structured / unstructured)" transmitted by the intelligent data asset tag recommendation unit, customized improvement suggestions are generated for different quality issues. For example, for the "high tag dispersion" problem, it suggests "manually verifying the consistency of multimodal data asset tags and adjusting tag weight allocation"; for the "data integrity missing" problem, it suggests "completing missing fields through historical data or synchronously supplementing them through business system interfaces." The generated improvement suggestions are synchronized to the automatic quality report generation unit and added to the "optimization suggestions" column of the report for easy user access. Simultaneously, "high-frequency problem types (such as repeated tag conflicts in a certain business domain)" are fed back to the intelligent data asset tag recommendation unit to assist in optimizing tag generation rules (such as adjusting large model prompt templates), reducing quality problems from the source and forming a closed-loop quality governance system of "rule configuration - report statistics - suggestion optimization - rule iteration."

[0046] The intelligent butler unit includes a natural language interaction module, an intelligent search and recommendation module, an MCP technology application module, and an intelligent question-answering module. Among them: The natural language interaction module allows business personnel to ask questions directly in spoken language through an intelligent question-and-answer interface (such as "Which data assets have been of substandard quality in the past month?"). The big data model automatically parses the semantics, associates the data asset library and data tags, and generates visualization results (including charts, warning prompts, etc.).

[0047] The intelligent search and recommendation module, based on user queries, intelligently searches the data asset repository, recommends the most relevant data assets, and provides data visualization and analysis results. The data asset repository includes a catalog and a knowledge graph.

[0048] The MCP technology application module enables the smart home assistant to quickly obtain the necessary data from the system, perform semantic understanding and content extraction, and then answer the user's questions in natural language.

[0049] The intelligent question-answering module, based on the natural language processing capabilities of a large model, can understand the user's question intent, perform contextual understanding, and provide accurate and clear answers. It can integrate data asset information retrieved from the catalog with relevant knowledge obtained from the knowledge graph. If the data assets found in the catalog have more detailed relationship descriptions and attribute information in the knowledge graph, this information is added to enrich the answer.

[0050] In the smart butler unit, users can ask questions to the smart butler via voice and text. The smart butler obtains data from the system through model context protocol technology and answers the user in natural language, allowing non-technical personnel to analyze the data independently and thus unlock the value of the data.

[0051] Each unit in the aforementioned intelligent data asset management system can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the corresponding operations of each unit.

[0052] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0053] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An intelligent data asset management system, characterized in that, The system comprises a data access layer, a data processing layer, an intelligent analysis layer, and an application service layer; the intelligent analysis layer includes an intelligent tag recommendation unit, an intelligent cataloging unit, an intelligent management unit, and an intelligent butler unit. The data access layer is used to receive multimodal data assets and perform format standardization conversion to obtain standardized multimodal data assets; the multimodal data assets include structured data, semi-structured data, and unstructured data; The data processing layer is used to preprocess the standardized multimodal data assets and extract the modal features of each modal data. The intelligent tag recommendation unit is used to embed the modal features and metadata of multimodal data assets into a tag generation prompt template, input it into a tag recommendation language model, output the corresponding multimodal data asset tags, and filter high-quality tag data to output to the intelligent cataloging unit. The tag generation prompt template includes metadata constraints, mapping examples between each modal feature and the tag, and conflict resolution instructions for different modal tags. The high-quality tag data are multimodal data asset tags that have passed the verification of preset tag quality assessment rules. The intelligent cataloging unit is used to train a pre-built multi-class head classification model based on high-quality tag data, input the modal features of the multi-modal data assets to be classified into the trained multi-class head classification model, generate multi-dimensional tags, and construct an intelligent catalog and knowledge graph based on the multi-dimensional tag data; The intelligent management unit is used to generate quality rules and quality labels according to preset prompt templates, calculate the comprehensive confidence level of the labels, and feed back optimization instructions to the upstream unit based on the confidence level to generate quality reports; The intelligent butler unit is used to input user questions into the intelligent question-and-answer language model, perform intelligent question-and-answer based on the intelligent catalog and knowledge graph, and output the answer content. The application service layer is used to call the output of the intelligent analysis layer to provide users with data asset query services, quality monitoring services and decision support services. Calculating the overall confidence score of the tags and feeding back optimization instructions to the upstream unit includes: Determine the quality dimensions and weights of each dimension required for calculating the overall confidence score of the labels, calculate the scores of each quality dimension, and calculate the dispersion of the multimodal data asset labels. Based on the quality dimension scores, quality dimension weights, and multimodal data asset label dispersion, calculate the overall confidence score of the labels. The system obtains the relationship between the overall confidence level of the tags and a threshold. When the overall confidence level of the tags is lower than the threshold, it generates tag generation optimization instructions and data flow control instructions. When the overall confidence level of the tags is higher than or equal to the threshold, it generates tag generation optimization instructions. The tag generation optimization instructions are used to adjust the prompt template parameters or modal feature weights of the tag recommendation large language model to optimize the generation quality of multimodal data asset tags. The data flow control instructions are used to restrict the circulation range of low-confidence data assets in the intelligent catalog or mark low-confidence data assets as pending review. The label generation optimization instructions are fed back to the intelligent label recommendation unit, and the data flow control instructions are fed back to the intelligent cataloging unit.

2. The system according to claim 1, characterized in that, The modal features and metadata of multimodal data assets are embedded into tags to generate prompt templates. These templates are then input into a tag recommendation language model, which outputs corresponding multimodal data asset tags, including: Convert the modal features and metadata of multimodal data assets into a format that can be parsed by a tag recommendation language model; The transformed modal features and data metadata are embedded into a tag generation prompt template. The embedded tag generation prompt template is then input into a tag recommendation language model to generate original tags. The original tags are then fused according to the preset conflict handling rules in the tag generation prompt template to obtain the corresponding multimodal data asset tags.

3. The system according to claim 1, characterized in that, A pre-built multi-class head classification model is trained based on high-quality labeled data. The modal features of the multimodal data assets to be classified are input into the trained multi-class head classification model to generate multi-dimensional labels, including: A multi-class head classification model is constructed. This model includes a multimodal parallel encoding module, a feature fusion module, and a multi-class head output module. The multimodal parallel encoding module performs branch encoding on the modal features of the multimodal data asset to obtain the encoded features corresponding to each modality. The feature fusion module fuses the encoded features of each modality to obtain multimodal fused features. The multi-class head output module performs classification processing on the multimodal fused features and outputs multi-dimensional labels. The multi-class head classification model is trained based on high-quality labeled data to obtain a well-trained multi-class head classification model. The modal features of the multimodal data assets to be classified are input into the trained multi-class head classification model to obtain the corresponding multi-dimensional labels.

4. The system according to claim 1, characterized in that, The intelligent tag recommendation unit includes a multimodal data unified tag module, an automated tag recommendation module, and a tag quality management module; The unified labeling module for multimodal data is used to construct a standardized labeling system; the standardized labeling system is used to standardize and unify the labeling dimensions of multimodal data assets. The automated tag recommendation module is used to call the standardized tag system, input the modal features of the multimodal data assets and the pre-set tag generation prompt template into the tag recommendation language model, and obtain multimodal data asset tags. The tag generation prompt template includes metadata constraints, tag generation examples corresponding to each modality of data, and tag conflict resolution instructions for different modalities; the metadata constraints include that all tags must be associated with multiple tag dimensions in the business rules, and the tag dimensions are specified by the standardized tag system; The tag quality management module is used to perform quality detection on the multimodal data asset tags according to the pre-set tag quality assessment rules, and to filter out high-quality tag data.

5. The system according to claim 1, characterized in that, The intelligent cataloging unit includes an automatic classification module, an intelligent catalog construction module, and a graph-based management module; The automatic classification module is used to train a pre-built multi-class head classification model using high-quality label data output by the intelligent label recommendation unit. The modal features of the data asset to be classified are input into the trained multi-class head classification model to generate multi-dimensional labels. The intelligent directory building module is used to build multi-level directories based on multi-dimensional tags and display them visually. When the tags are updated, the directory is dynamically adjusted by calculating the similarity between the asset positioning vector and the directory category, and the data association is synchronized to the graph management module. The graph-based management module is used to construct an asset graph based on the data relationships and multimodal data assets marked with multidimensional tags, and then visualize and display it.

6. The system according to claim 5, characterized in that, The asset map is constructed based on the aforementioned data relationships and multimodal data assets labeled with multidimensional tags, including: Define the node attributes of the asset graph based on the multimodal data assets marked with multidimensional labels; Define the edge attributes of the asset graph based on the data relationships; Based on node attributes and multimodal data assets labeled by prediction tags, construct the node set of the asset graph; based on edge attributes and data associations, construct the initial edge set of the asset graph; and based on the node set and the initial edge set, form the initial structure of the asset graph. Calculate the label similarity between unrelated data assets in the initial structure, determine potential relationships, and update the edge set of the asset graph based on the potential relationships; The asset graph is obtained based on the updated set of nodes and edges.

7. The system according to claim 1, characterized in that, The intelligent management unit includes a quality rule automatic configuration module, a quality report automatic generation module, and an improvement suggestion provision module; The automatic quality rule configuration module receives multimodal data asset tags and data metadata output by the intelligent tag recommendation unit and business domain associations output by the intelligent cataloging unit. It then combines preset prompt templates to introduce quality attributes and bind business rules, automatically generating quality rules, tag dispersion calculation parameters, and quality tag generation logic. It calculates the comprehensive confidence level of the tags and feeds back optimization instructions to the upstream unit. The quality attributes include: completeness, timeliness, and downstream dependency risk. The automatic quality report generation module is used to calculate the data asset quality score, count the amount of problematic data and analyze quality trends based on quality rules, confidence results and intelligent catalog, and automatically generate a visual quality report. The improvement suggestion providing module is used to receive quality-related parameters output by the quality rule auto-configuration module, problem data details output by the quality report auto-generation module, and modal characteristics of multimodal data assets. It generates customized improvement suggestions for different quality problems, synchronizes the generated improvement suggestions to the quality report, and reports the problem of repeated label conflicts to the upstream unit.

8. The system according to claim 7, characterized in that, The overall confidence level of the labels is: in, For the first Each quality dimension score For the weight of the quality dimension, For the dispersion of multimodal data asset labels, It is the first The original labels generated by each modality For modal types.

9. The system according to claim 1, characterized in that, The intelligent butler unit includes a natural language interaction module, an intelligent search and recommendation module, an MCP technology application module, and an intelligent question-and-answer module; The natural language interaction module is used to receive spoken questions input by business personnel through the intelligent question-and-answer interface, use a large model to parse the semantics, associate the data asset library with data tags, and generate visualization results including charts and warning prompts. The intelligent retrieval and recommendation module is used to retrieve a data asset library containing a directory and knowledge graph based on user queries, recommend relevant data assets, and provide data visualization and analysis results. The MCP technology application module is used to quickly obtain the required data from the system through Model Context Protocol technology, perform semantic understanding and content extraction, and answer user questions in natural language. The intelligent question-answering module is used to understand the user's question intent based on the natural language processing capabilities of a large model, integrate the relationship and attribute information between the asset information retrieved from the directory and the knowledge graph, and generate answer content.

Citation Information

Patent Citations

  • Data asset management method based on knowledge graph

    CN119226353A

  • Intelligent file classification and retrieval method and system

    CN120086390A