Intelligent enterprise data asset analysis method and system based on AI identification

By using an AI-based multimodal fusion recognition model and dynamic data asset graph analysis, the problem of integrating and assessing multi-source heterogeneous data in enterprise data asset management has been solved, enabling real-time and accurate identification and efficient governance of data assets, and improving the real-time performance and accuracy of enterprise data management.

CN120975397AInactive Publication Date: 2025-11-18WUPO DIGITAL TECHNOLOGY (HANGZHOU) GROUP CO LTD
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202511121774.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in enterprise data asset management suffer from problems such as difficulty in integrating multi-source heterogeneous data, strong subjectivity in value assessment, and lagging risk identification. They are unable to meet the asset value-added needs in dynamic business scenarios, and lack cross-modal feature integration, resulting in insufficient real-time performance and accuracy of governance strategies.

Method used

An AI-based intelligent analysis method for enterprise data assets is adopted. By using a pre-trained multimodal fusion recognition model for joint feature extraction and semantic alignment, structured data asset recognition results are generated. A dynamic enterprise data asset map is constructed by combining real-time data access trajectories and permission metadata, and spatiotemporal evolution analysis is performed. A hierarchical topology map of data assets is generated using a self-organizing mapping network, and an executable data governance action sequence is generated through policy constraint reinforcement learning.

Benefits of technology

It improves the identification accuracy and real-time analysis capabilities of enterprise data assets, realizes standardized understanding and integration of multi-source heterogeneous data, ensures the real-time and accuracy of data asset management, and provides efficient data governance strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975397A_ABST
    Figure CN120975397A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise data asset intelligent analysis method and system based on AI recognition, and the method comprises the steps: receiving an enterprise multi-source heterogeneous data stream, carrying out the joint feature extraction and semantic alignment through a pre-trained multi-modal fusion recognition model, and generating a structured data asset recognition result; constructing a dynamic enterprise data asset atlas according to the structured data asset identification result in combination with the data access trajectory and authority metadata collected in real time; performing spatio-temporal evolution analysis on the dynamic enterprise data asset map, and extracting potential data value density features and risk exposure features; inputting the data value density features and the risk exposure features into a self-organizing mapping network to generate a data asset grading topological graph; and based on the data asset grading topological graph, through strategy constraint reinforcement learning, generating an executable data governance action sequence. According to the embodiment of the invention, the identification precision and real-time analysis capability of special assets of enterprises can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of AI technology, specifically a method and system for intelligent analysis of enterprise data assets based on AI recognition. Background Technology

[0002] Currently, the management of special corporate assets (such as intellectual property, goodwill, and data assets) generally faces challenges such as difficulties in integrating multi-source heterogeneous data, strong subjectivity in value assessment, and lagging risk identification. Traditional methods rely on manual classification and static rules, which are insufficient to meet the asset value-added needs in dynamic business scenarios. Although existing technologies attempt to introduce machine learning for data classification, cross-modal feature fusion is insufficient, and there is a lack of quantitative analysis of the spatiotemporal evolution of assets, resulting in insufficient real-time performance and accuracy of governance strategies. Summary of the Invention

[0003] The purpose of this invention is to provide an AI-based intelligent analysis method and system for enterprise data assets, in order to address the shortcomings of existing technologies and improve the accuracy of enterprise special asset identification and real-time analysis capabilities.

[0004] One embodiment of this application provides an AI-based intelligent analysis method for enterprise data assets, the method comprising:

[0005] The system receives multi-source heterogeneous data streams from enterprises, performs joint feature extraction and semantic alignment using a pre-trained multimodal fusion recognition model, and generates structured data asset recognition results. The multimodal fusion recognition model simultaneously processes text, image, table, and log data through a cross-modal attention mechanism to identify data types, content themes, and sensitive information tags.

[0006] Based on the structured data asset identification results, and combined with the real-time collected data access trajectory and permission metadata, a dynamic enterprise data asset graph is constructed. The nodes of the graph represent data entities, and the edge weights are dynamically calculated and generated based on data association, access frequency, and permission association.

[0007] Spatiotemporal evolution analysis is performed on the dynamic enterprise data asset map to extract potential data value density characteristics and risk exposure characteristics. The spatiotemporal evolution analysis uses graph neural network time series prediction to quantify the value decay curve of data assets and the probability of compliance risks.

[0008] The data value density features and risk exposure features are input into a self-organizing map network to generate a data asset hierarchical topology map. The self-organizing map network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning, forming a four-quadrant visual distribution. Each quadrant corresponds to a quantitative hierarchical label of value-risk.

[0009] Based on the data asset hierarchical topology map, an executable data governance action sequence is generated through policy constraint reinforcement learning, wherein the action sequence includes automatic archiving, encryption enhancement, access permission reconstruction, and compliance audit trigger instructions, and the action priority is dynamically allocated by the hierarchical matrix quadrant position.

[0010] Yet another embodiment of the present application provides an AI recognition-based enterprise data asset intelligent analysis system, which comprises:

[0011] A receiving module is configured to receive enterprise multi-source heterogeneous data streams, perform joint feature extraction and semantic alignment using a pre-trained multi-modal fusion recognition model, and generate a structured data asset recognition result, wherein the multi-modal fusion recognition model synchronously processes text, image, table, and log data through a cross-modal attention mechanism to identify data types, content themes, and sensitive information labels.

[0012] A construction module is configured to construct a dynamic enterprise data asset graph based on the structured data asset recognition result and in combination with real-time collected data access tracks and permission metadata, wherein the nodes of the graph represent data entities, and the edge weights are dynamically calculated and generated according to data correlation, access frequency, and permission correlation degree.

[0013] An extraction module is configured to perform spatio-temporal evolution analysis on the dynamic enterprise data asset graph to extract potential data value density features and risk exposure features, wherein the spatio-temporal evolution analysis is performed through graph neural network time series prediction to quantify the value decay curve and compliance risk probability of the data asset.

[0014] An input module is configured to input the data value density features and risk exposure features into a self-organizing mapping network to generate a data asset hierarchical topology map, wherein the self-organizing mapping network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning to form a four-quadrant visual distribution, and each quadrant corresponds to a quantitative hierarchical label of value-risk.

[0015] A generation module is configured to generate an executable data governance action sequence through policy constraint reinforcement learning based on the data asset hierarchical topology map, wherein the action sequence includes automatic archiving, encryption enhancement, access permission reconstruction, and compliance audit trigger instructions, and the action priority is dynamically allocated by the hierarchical matrix quadrant position.

[0016] Yet another embodiment of the present application provides a storage medium having a computer program stored therein, wherein the computer program is configured to execute the method described in any of the above embodiments when running.

[0017] Still another embodiment of the present application provides an electronic device comprising a memory having a computer program stored therein and a processor configured to execute the computer program to perform the method described in any of the above.

[0018] Compared with the prior art, the AI recognition-based enterprise data asset intelligent analysis method provided by the present application receives enterprise multi-source heterogeneous data streams, uses a pre-trained multi-modal fusion recognition model to perform joint feature extraction and semantic alignment, and generates a structured data asset recognition result; according to the structured data asset recognition result, in combination with real-time collected data access tracks and permission metadata, a dynamic enterprise data asset graph is constructed; the dynamic enterprise data asset graph is subjected to spatio-temporal evolution analysis, and potential data value density features and risk exposure features are extracted; the data value density features and the risk exposure features are input into a self-organizing mapping network to generate a data asset hierarchical topology graph; based on the data asset hierarchical topology graph, through policy constraint reinforcement learning, an executable data governance action sequence is generated, so as to improve the recognition accuracy and real-time analysis capability of enterprise special assets. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A hardware structure block diagram of a computer terminal of the AI recognition-based enterprise data asset intelligent analysis method provided by the embodiment of the present application is shown in the figure.

[0020] Figure 2 A flowchart of the AI recognition-based enterprise data asset intelligent analysis method provided by the embodiment of the present application is shown in the figure.

[0021] Figure 3 A structure diagram of the AI recognition-based enterprise data asset intelligent analysis system provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0022] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be explained as a limitation of the present application.

[0023] The embodiment of the present application first provides an AI recognition-based enterprise data asset intelligent analysis method, which can be applied to an electronic device such as a computer terminal, specifically, a general computer, etc.

[0024] The following will be described in detail taking a computer terminal as an example. Figure 1 A hardware structure block diagram of a computer terminal of the AI recognition-based enterprise data asset intelligent analysis method provided by the embodiment of the present application is shown in the figure. Figure 1 As shown in the figure, the computer device comprises a processor, a memory and a network interface connected through a system bus, wherein the memory can comprise a non-volatile storage medium and an internal memory.

[0025] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any one of the AI recognition-based enterprise data asset intelligent analysis methods. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, which, when executed by the processor, can cause the processor to perform any one of the AI recognition-based enterprise data asset intelligent analysis methods. The network interface is used for network communication, such as sending assigned tasks.

[0026] Referring to Figure 2 Embodiments of the present application provide an AI recognition-based enterprise data asset intelligent analysis method, which can include the following steps:

[0027] S201, receiving enterprise multi-source heterogeneous data streams, using a pre-trained multi-modal fusion recognition model for joint feature extraction and semantic alignment to generate structured data asset recognition results, wherein the multi-modal fusion recognition model synchronously processes text, image, table and log data through a cross-modal attention mechanism to identify data types, content themes and sensitive information labels;

[0028] Specifically, multi-modal data streams containing text, image, table and log can be received, and adaptive format parsers can be used to unify each modality data into a tensor sequence to output standardized multi-modal data tensors;

[0029] The system receives raw data streams from different business systems of an enterprise through a distributed message queue (such as Kafka, a high-throughput distributed stream processing platform). Text data may come from customer service tickets (in CSV format), contract documents (in PDF / Word format); image data includes product design drawings (in PNG / JPG format), scanned bills; table data covers financial Excel reports, database export CSV; log data comes from server logs (such as Nginx access logs), application operation logs (in JSON format). The adaptive format parser first identifies the data type (Text / Image / Table / Log) based on the meta-information of the data stream (such as the Content-Type field in the HTTP header, the file extension). For text data, the parser calls the OCR engine (Optical Character Recognition, such as Tesseract) to process scanned document images, and then uses a word segmentation tool (such as Jieba Chinese word segmentation) to split the text into a sequence of word tokens; for image data, use the OpenCV library (Open Source Computer Vision Library) for size normalization (unified to 224x224 pixels) and channel standardization (RGB three channels mean zero); table data is parsed by the Pandas library (Python data analysis tool) to map the row and column structure to key-value pairs; log data uses regular expression template matching to match structured fields such as timestamps and operation types. All parsed intermediate results are converted into numerical tensor sequences (Tensor Sequence, i.e. multi-dimensional array sequence): text is converted into word embedding index sequence (each word is mapped to a 300-dimensional vector index), image is converted into three-dimensional pixel tensor (dimension H, W, C), table is converted into two-dimensional feature matrix (row = record number, column = feature number), log is converted into timestamp-labeled event vector (dimension = event type number).

[0030] To achieve inter-modal comparability, the parser normalizes tensors of different modalities:

[0031] Dimension alignment: text sequences are unified to 512 tokens by padding or truncation; image tensors maintain fixed HxWxC dimensions; table feature matrices are unified to 128 columns by interpolation or dimension reduction; log event vectors are expanded to fixed-length vectors using one-hot encoding.

[0032] Numerical normalization: all numerical features (such as image pixel values 0-255) are scaled to the [-1, 1] interval (formula: (x-128) / 128), and categorical features (such as log operation types) are converted to embedding vectors.

[0033] Serialization Packaging: The processed single-modal tensors are packaged into standardized data units (SDUs) with modal labels, containing three core attributes:

[0034] modality_type (modal type: text / image / table / log); tensor_data (standardized tensor data); metadata (original data source information, timestamp, etc.).

[0035] The final output is a batch of standardized multimodal data tensors (SMDT) processed by time window (TW=5 minutes), with a four-dimensional tensor structure (dimension 1=Batch Size, BS=64; dimension 2=Modality Count, MC=4; dimension 3=SeqLength, SL; dimension 4=Feature Dim, FD).

[0036] Exception handling mechanism ensures robustness:

[0037] Format error: For documents that fail to parse (such as encrypted PDF), trigger retry mechanism (RM) or transfer to manual review queue (MRQ).

[0038] Data missing: KNN imputation (K-Nearest Neighbors Imputation, mean filling based on similar records) is used for table null values, and generative adversarial network (GAN) is used to generate alternative content for damaged images.

[0039] Resource scheduling: The parser dynamically adjusts the thread pool size (TPS) according to real-time load, and ensures that high-value data (such as text containing the keyword "contract") is processed first through priority queue (PQ). The processed SMDT is written to distributed storage (such as HDFS, Hadoop Distributed File System) for subsequent model calls.

[0040] Input the standardized multimodal data tensor into the pre-trained multimodal fusion recognition model, and use the gated cross-modal attention mechanism to calculate the interaction weight of text-image-table-log, and generate a joint feature embedding vector;

[0041] The pre-trained multi-modal fusion recognition model adopts a hierarchical encoder architecture:

[0042] Single-modal encoding layer:

[0043] Text encoder: BERT model (Bidirectional Encoder Representations from Transformers) extracts context features, outputting a sequence of 768-dimensional vectors. Image encoder: ResNet-50 (Residual Network 50 layers) convolutional neural network extracts visual features, outputting a 2048-dimensional feature map. Table encoder: Multilayer Perceptron (MLP) processes structured data, outputting a 256-dimensional vector. Log encoder: LSTM (Long Short-Term Memory) models temporal dependencies, outputting a 128-dimensional state vector. The outputs of each encoder are projected uniformly into a 512-dimensional common space (Projection Space, PS), forming a modality-aligned intermediate representation.

[0044] Gated Cross-modal Attention (GCA) enables multi-modal interaction:

[0045] Attention weight calculation: For any two modalities (e.g., text Text and image Image), calculate the Query-Key matching degree: Query vector (Q_text) = text feature matrix × weight matrix W_Q; Key vector (K_image) = image feature matrix × weight matrix W_K; Attention score (Attn_score = Softmax(Q_text · K_image^T / √d_k) (d_k=64 is the scaling factor).

[0046] Gating mechanism fusion: Introduce a learnable gating parameter (Gating Parameter, GP) to control information flow:

[0047] Gate value (Gate_value) = σ(W_g·[Q_text; Attn_score] + b_g), where σ = Sigmoid function; Weighted feature (Weighted_feature) = Gate_value × (Attn_score·V_image), where V_image = image Value vector.

[0048] Cross-modal aggregation: Parallel computing attention for all modal pairs such as text-image, text-table, image-log, and output aggregated multimodal feature block (MFB).

[0049] Finally, generate joint feature embedding vector (JFEV):

[0050] Feature concatenation: Concatenate the MFB of each modal pair in the channel dimension to form a high-dimensional mixed feature (dimension = number of modal pairs x 512).

[0051] Compression mapping: Reduce dimension to 1024 dimensions through fully connected layer (FC) and use layer normalization to stabilize training.

[0052] Nonlinear activation: Apply GeLU function (Gaussian Error Linear Unit) to enhance expression ability.

[0053] In the model pre-training stage, the Masked Multimodal Modeling task is used: randomly mask 15% of the input units (such as text tokens, image blocks), and require the model to reconstruct the masked content, so that the JFEV contains cross-modal semantic association.

[0054] Based on the joint feature embedding vector, through contrastive learning, the semantic space of different modalities is aligned, the semantic gap between modalities is eliminated, and the aligned semantic consistent feature matrix is output.

[0055] The InfoNCE loss (Noise Contrastive Estimation) is used in the contrastive learning framework:

[0056] Positive and negative sample construction:

[0057] Positive sample (Positive Pair): Different modal representations of the same data entity (such as "Product A" text description + design drawing). Negative sample (Negative Pair): Randomly sample modal representations of different entities (such as "Product A" text + "Product B" image). Similarity calculation: Define the cosine similarity (CS) function: sim(u,v) = (u ·v) / (||u||·||v||), where u and v are two modal JFEV vectors.

[0058] Loss function optimization: minimize the difference between positive sample similarity and negative sample similarity: Loss = -log[ exp(sim(u+,v+) / τ ) / Σ_{k=1}^K exp(sim(u-,v-) / τ ), τ=0.07 is the temperature parameter, K=256 is the number of negative samples.

[0059] Core technology of semantic space alignment:

[0060] Shared Projection Head: map 1024-dimensional JFEV input to low-dimensional alignment space (AS) through two-layer MLP (hidden layer dimension 2048, output layer 256-dimensional). Momentum Encoder: use exponential moving average (EMA) of main encoder parameters to update copy encoder to enhance training stability, formula: θ_m = m·θ_m + (1-m)·θ, m=0.99 is the momentum coefficient. Memory Bank: store AS vectors of historical samples (capacity=65536) to expand negative sample sources. Output Semantic-Consistent Feature Matrix (SCFM):

[0061] Alignment verification: calculate modality alignment score (MAS): MAS = average positive sample similarity - average negative sample similarity; when MAS > 0.85, determine alignment success. Feature reorganization: reorganize 256-dimensional AS vectors after alignment into a matrix according to the original batch (dimension=BS×256). Gap elimination effect: after alignment, the Euclidean distance between feature vectors of different modalities describing the same entity is reduced to less than 20% of the original distance (e.g. from >5.0 to <1.0), achieving cross-modal semantic unification.

[0062] Input the semantic-consistent feature matrix into the multilayer perceptron classifier head to simultaneously identify data types, content themes, and sensitive information labels, and integrate them into structured data asset identification results.

[0063] Multilayer Perceptron Classifier Head (MLP-Head) structure:

[0064] Input layer: receives 256-dimensional SCFM vectors. Hidden layers: two fully connected layers (FC1 dimension 128, FC2 dimension 64) with Dropout = 0.3 to prevent overfitting.

[0065] Parallel output layers: Data type classifier: outputs 5-dimensional Softmax probabilities (text / image / table / log / mixed). Content topic classifier: outputs 20-dimensional Sigmoid probabilities (topic labels like "finance," "customer," "R&D," etc.). Sensitive information detector: outputs 3-dimensional Sigmoid probabilities ("PII personal identity information," "PCI payment card information," "PHI health information").

[0066] Multi-task joint training strategy:

[0067] Loss function weighting: Total Loss (TL) = a·Loss_type + β·Loss_topic + γ·Loss_sensitivity, weight coefficients a = 0.3, β = 0.5, γ = 0.2 (adjustable according to business requirements). Class imbalance handling: for minority classes (e.g., "PHI"), use Focal Loss (FL): FL = -a_t(1-p_t)^γ log(p_t), where a_t = 0.75 is the class weight and γ = 2.0 is the difficult sample focusing parameter. Online Hard Example Mining (OHEM): select the top 30% samples with the highest loss value in each batch for backpropagation.

[0068] Generate Structured Data Asset Identification Result (SDAIR):

[0069] Label decoding:

[0070] Data type: take the label corresponding to the maximum Softmax probability (confidence threshold > 0.8). Content topic: take all labels with Sigmoid probability > 0.5 (multi-label output). Sensitive information: if the probability of any sensitive category is > 0.7, trigger the sensitive flag. Result packaging: output structured records in JSON-LD (Linked Data JSON) format:

[0071] {

[0072] "data_id": "DOC-20240501-001",

[0073] "data_type": ["text", "table"],

[0074] "content_topic": ["finance", "contract"],

[0075] "sensitivity_tags": ["PII", "PCI"],

[0076] "confidence_scores": [0.92, 0.88, 0.91]

[0077] }。

[0078] Quality monitoring: The recognition accuracy is evaluated by F1-score (harmonic mean of precision and recall), and the overall type F1 is required to be greater than 0.85. The results are written into the graph database Neo4j for downstream construction of data asset graph.

[0079] This step unifies the processing of different formats and sources of data within the enterprise through a pre-trained multi-modal fusion model, and uses cross-modal attention mechanism to capture the implicit associations between text, image, table and log data, and converts the original data into structured information with clear semantic labels. The model eliminates the semantic differences between different data modalities through contrastive learning, ensuring that the extracted features accurately reflect the essential properties of data assets, achieving standardized understanding and integration of multi-source heterogeneous data, and solving the recognition fragmentation problem caused by data format differences in traditional methods. The structured recognition results provide high-quality input for subsequent analysis, while the automatic labeling of sensitive information tags lays the foundation for data security governance.

[0080] S202, according to the structured data asset recognition result, combined with the real-time collected data access track and permission metadata, a dynamic enterprise data asset graph is constructed, wherein the nodes of the graph represent data entities, and the edge weights are dynamically calculated and generated according to the data correlation relationship, access frequency and permission correlation degree;

[0081] Specifically, the structured data asset recognition result can be parsed to extract data entities, and based on the content theme similarity between entities, an initial relationship edge is constructed, and a weighted initial relationship graph is output;

[0082] The system first receives the structured data asset identification results generated by the upstream multi-modal fusion recognition module. This result is a highly organized data set, where each record explicitly labels the identified data entities (Data Entity, DE) and their attributes, including data types (such as customer database tables, product design drawings, server log files), content topics (such as "financial reimbursement", "user portrait", "device monitoring"), and sensitive information labels (such as "PII - Personal Identity Information", "PCI - Payment Card Information"). The parsing process is performed by a special graph construction engine (Graph Construction Engine, GCE). GCE uses a rule-based and machine learning combined entity extractor (Entity Extractor, EE) to accurately identify and extract all valid data entity objects. For example, from the identification result, it may extract "CRM_Customer_Table" (customer relationship management system customer table), "Q3_Sales_Report.pdf" (third quarter sales report PDF file), "App_Server_Error_Log" (application server error log) and other specific entities. Each entity is assigned a unique identifier (Unique Identifier, UID) and carries its content topic label (Content Topic Label, CTL) and other metadata.

[0083] The core basis for constructing the initial relationship edge (IRE) is the content topic similarity (CTS) between data entities. The system uses a pre-trained topic embedding model (TEM), which is usually based on algorithms such as BERT or Doc2Vec, to convert the content topic label (CTL) and description text of each entity into a high-dimensional semantic vector (SV). When calculating the CTS of entity A and entity B, the cosine similarity (CS) algorithm is used to measure the cosine value of the included angle between their corresponding semantic vectors SvA and SvB. For example, two entities with topic labels "employee performance evaluation" and "salary structure" may have a CTS of 0.85 (close to 1 indicates high similarity), while the CTS of "employee performance evaluation" and "computer room temperature monitoring" may only be 0.1 (close to 0 indicates no correlation). The system sets a similarity threshold (ST, for example, 0.6), and when the CTS is greater than ST, an initial relationship edge is established between the corresponding two entity nodes. The initial weight of this edge (IEW) is directly set to the calculated CTS value (ranging from 0 to 1), which intuitively reflects the strength of the topic association. The output of this step is a weighted initial relationship graph (WIRG) with data entities as nodes and topic similarity as weights.

[0084] To improve the accuracy and efficiency of the initial graph, the system uses affinity propagation clustering (APC) for optimization. The APC algorithm can automatically determine "exemplars" and automatically aggregate data entities with high theme similarity into clusters. Within the same cluster, entities automatically establish fully connected edges (with weights equal to the average similarity within the cluster), and between different clusters, the representative points of the clusters establish connection edges according to the theme similarity of the entity. This effectively avoids the overhead of full connection similarity calculation on a super large graph, while ensuring the close connection of a strong correlation entity group. For example, all database tables, order processing logs, and shipping records related to the "customer order" theme are clustered into a cluster, and the edge weights between them are initialized to a high value (such as 0.9), while the edge weights between the representative points of this cluster and the representative points of the "supplier management" cluster are calculated according to the theme similarity (such as 0.7). The final output WIRG not only contains the binary relationships between entities, but also implicitly contains the community structure based on the theme, laying a good foundation for subsequent dynamic evolution.

[0085] The real-time collected data access trajectory is converted into a time series access frequency vector, and the time series access frequency vector is fused with the weighted initial relationship graph. The edge weights are updated by sliding window aggregation, and the dynamic access enhanced relationship graph is output.

[0086] The system captures data access traces stream (DATS) in real-time by deploying log probes (LP) at each data access entry of the enterprise (e.g. API gateway, database proxy, file server). Each trace record contains key information: access timestamp (TS), access subject (e.g. user ID, service account), access operation (e.g. SELECT read, UPDATE update, DOWNLOAD download) and target data entity UID (TDE_UID). These raw stream data are ingested into a stream processing engine (SPE, e.g. Apache Flink or Spark Streaming) in real-time. The core task of the SPE is to aggregate per data entity by pre-set time window (TW, e.g. 15 minutes, 1 hour, 1 day) and generate time-series access frequency vector (TSAFV). For example, for entity "CRM_Customer_Table", in the last 1 hour time window, a vector [read count = 120, update count = 5, download count = 0] can be generated, reflecting the access heat (AH) and pattern in this period.

[0087] The fusion process is to superimpose the dynamic access information represented by TSAFV on the static weighted initial relation graph (WIRG) to update the edge weight. The system adopts sliding window aggregation (SWA) strategy. For each edge (connecting entity A and entity B) in WIRG, the system retrieves the respective TSAFV of entity A and entity B in the current sliding window (CSW, e.g. the past 24 hours). The key operation is to calculate a joint access strength metric (JASM). A typical calculation method of JASM is: JASM = a * CoAccess_Freq + b * (Freq_A * Freq_B)^0.5. Wherein:

[0088] CoAccess_Freq (co-access frequency): statistics of the number of times that the same access subject (or the same session) accesses A and B in adjacent time (e.g. within 60 seconds) in CSW, reflecting the business relevance.

[0089] Freq_A and Freq_B: Total access times (or weighted sum, e.g. read weight 1, update weight 2) of entities A and B within CSW; a and b: harmonic coefficients (usually a>b, emphasizing direct collaboration), e.g. a=0.7, b=0.3.

[0090] JASM value is normalized to 0-1 range. This value is used to update original topic-similarity-based edge weight (IEW). A common update rule is: New_Edge_Weight = g * IEW + (1-g)* JASM. g is a decay factor (DF, e.g. 0.4) to balance the contribution of static topic association and dynamic access association. This process works on all edges in parallel.

[0091] System adopts incremental computation (IC) to optimize performance. Whenever a new time window (e.g. another 15 minutes) of TSAFV data arrives, engine only re-computes JASM and New_Edge_Weight of edges affected by new data (i.e. those connecting entities with access activities in this window), instead of full graph update. Meanwhile, engine maintains a heat decay model (HDM) to assign lower weight to more distant historical access data (e.g. using exponential decay: Weight_t = e^(-l* age), where l is decay rate, age is data age). This ensures that graph can sensitively reflect the latest access pattern changes. For example, if marketing department suddenly frequently collaborates on “promotion activity table” and “customer feedback table”, even if their initial topic similarity is not high, their edge weight will quickly increase. The final output, dynamic access-enhanced relation graph (DAERG), is a continuously evolving graph whose edge weight is a fusion of content topic static association and real-time access dynamic association.

[0092] Integrate permission metadata, calculate the permission association degree between data entities, and generate a permission association degree matrix;

[0093] The system extracts permission metadata (PMD) from the enterprise's Identity and Access Management (IAM) system, Permission Management Database (PMDB), or Access Control Lists (ACLs). The core information of PMD includes: permission subject (S, such as user, user group, role), permission object (O, i.e. data entity), operation permission (OP, such as read, write, delete), and possible permission conditions (C, such as time limit, IP limit). The system first performs permission object alignment (POA) to ensure that the objects in PMD (usually represented by resource identifiers Resource ID) are accurately matched with the data entity UIDs in the graph. For objects that cannot be automatically matched (such as newly created or unregistered entities), manual review or fuzzy matching algorithm (FMA) based on name and path is triggered for association.

[0094] The core goal of this step is to calculate the permission association degree (PAD) between data entities. The permission association degree measures the similarity or overlap of two data entities in permission configuration. The system uses a subject-based association calculation (SBAC) method:

[0095] For each data entity E, generate its permission vector (PV). The dimension of PV is the set of all possible permission subjects (Subject Set, SS) or the set of roles (Role Set, RS). The value of the vector element represents the highest permission level that the subject / role has for the entity E (such as 0=no permission, 1=read, 2=write, 3=manage).

[0096] For any two entities A and B, calculate the similarity of their permission vectors PVA and PVB. Common methods include:

[0097] Jaccard Similarity Coefficient (JSC): Focuses on the proportion of subjects that share permissions (non-zero elements). JSC = |PVA ∩ PVB| / |PVA ∪ PVB|. Cosine Similarity (CS): Considers the directional consistency of the permission level vectors. CS = (PVA · PVB) / (||PVA|| * ||PVB||). Inverse Weighted Euclidean Distance (IWED): PAD = 1 / (1 + d), where d is the weighted Euclidean distance between PVA and PVB, with weights assigned to give greater importance to high permission levels (e.g., administrator roles).

[0098] The system typically combines multiple methods, such as taking a weighted average of JSC and CS as the final PAD value.

[0099] All PAD values for entity pairs eventually form an N x N symmetric matrix (N is the total number of entities), i.e., the Permission Association Matrix (PAM). Matrix element PAM[i][j] represents the permission association degree between entity i and entity j, with a value range of [0, 1]. To optimize storage and calculation, the system uses Sparse Matrix (SM) technology, only storing non-zero (or greater than a certain threshold, such as 0.2) PAD values and their corresponding entity pair indexes. In addition, the system identifies and labels High-Permission Association Clusters (HPAC), i.e., groups of entities with high PAD values among each other (which can be found through community detection of PAM). This usually corresponds to specific business units or project teams within an enterprise that share core data sets. The generation of the Permission Association Matrix PAM is the third important dimension in addition to static permission configuration and dynamic access behavior, which reveals the inherent connection of data entities in the security management layer.

[0100] Combining the dynamic access enhancement relationship graph and the permission association matrix, use the entropy weight method to dynamically synthesize the comprehensive edge weight, output the normalized weight atlas;

[0101] By now, the system has three independent but related weight sources to depict the relationship between data entities:

[0102] W1: Initial edge weight based on topic similarity (from DAERG, combining static content association and dynamic access association).

[0103] W2: Edge weight based on relevance degree of authority (from PAM, i.e. PAD value).

[0104] To construct a single graph that comprehensively reflects the data association, W1 and W2 need to be synthesized into a comprehensive edge weight (CEW). The system uses the entropy weight method (EWM) to determine the objective weight (OW) of W1 and W2, avoiding subjective assignment bias. The core idea of the entropy weight method is that the smaller the entropy of an index, the more information it provides, and it should be given more weight in comprehensive evaluation.

[0105] The dynamic synthesis process of the entropy weight method is as follows:

[0106] Data matrix construction: For each edge (connecting entity i and j) in the graph, collect its two weight index values: Xij1 = W1(ij) (from DAERG), Xij2 = W2(ij) (from PAM). The index values of all edges form an Mx2 matrix (M is the total number of edges).

[0107] Standardization: Standardize each column (each index) (normalization processing). Common methods such as the range method (Range Method): Yijk = (Xijk - min(Xk)) / (max(Xk) - min(Xk)), where k = 1 or 2, indicating the index. Ensure that all Yijk are in the [0, 1] interval.

[0108] Calculate the proportion of each index value for each edge i: Pijk = Yijk / Σ(i=1 to M) Yijk.

[0109] Calculate the information entropy (IE) of the kth index: IEk = - (1 / ln(M)) * Σ(i=1 to M) (Pijk * ln(Pijk)). If Pijk = 0, define this term as 0.

[0110] Calculate the information utility value (IUV): IUVk = 1 - IEk. The smaller IEk, the larger IUVk, indicating that the information provided by the index is more valuable.

[0111] Calculate the index weight: The entropy weight of the kth index Wk = IUVk / Σ(k=1 to 2) IUVk.

[0112] Compute the comprehensive edge weight CEW: for each edge ij, its comprehensive edge weight CEWij = W1 * Yij1 + W2 * Yij2.

[0113] For example, the entropy weight of W1 (topic access weight) after calculation may be 0.65, and the entropy weight of W2 (permission weight) may be 0.35. This means that under the current data distribution, the topic access relevance provides more differentiated information, which dominates in the comprehensive weight.

[0114] The value range of the synthesized comprehensive edge weight CEWij is between [0, 1]. The system performs global normalization (GN) to linearly scale the CEW values of all edges, making their sum equal to 1 or the maximum value not exceeding a certain set value (such as 100), ensuring that the weights of the graphs constructed at different times are comparable. The final output is a normalized weight graph (NWG). In the NWG, the nodes are still data entities, and the weight of each edge is the normalized comprehensive weight CEW', which contains the relevance of data content topics, the synergy of real-time access behavior, and the similarity of permission configuration. This comprehensive weight can more comprehensively and objectively reveal the real association strength between enterprise data assets.

[0115] The normalized weight graph is input into the incremental graph learning engine, which optimizes node clustering through community detection algorithms and embeds a time trigger to respond to new data streams, outputting a dynamic enterprise data asset graph.

[0116] The incremental graph learning engine (IGLE) is the core component of building a dynamic graph. It receives the normalized weight graph (NWG) as the initial input. One of the core functions of IGLE is community detection (CD), which aims to find groups of nodes (communities) in the graph that are tightly connected (i.e., high edge weight), optimizing the node clustering structure. The system usually uses efficient and incremental updating algorithms such as the Louvain algorithm or its improved version. The Louvain algorithm is a hierarchical clustering method based on modularity optimization (MO). Modularity (Q) is a measure of the quality of community division, and the higher the value, the stronger the internal connection of the community than random connection. The Louvain algorithm iteratively performs two steps:

[0117] Local Optimization: Traverse each node, try to move it to the community where its neighbors belong, calculate the gain in modularity ΔQ. If ΔQ>0 and is maximal, move.

[0118] Community Folding: Fold nodes belonging to the same community into a new super node, the edge weight between the new nodes is the sum of all edge weights between the original nodes.

[0119] This process repeats until the modularity no longer improves significantly. The final output community structure more clearly reflects the natural grouping of data assets (e.g. "customer data domain", "supply chain data domain", "financial data domain").

[0120] The key feature of IGLE is incremental update (IU). Instead of re-running the full graph community detection algorithm (huge computational overhead) every time new data arrives, IGLE embeds a temporal trigger (TT) mechanism. The trigger listens to two main types of events:

[0121] Node / Edge Addition / Deletion Event: When new data entities are identified and added to the graph, or old entities are marked as invalid; or when real-time access streams or permission updates result in new edges or significant changes (change amount exceeds threshold ΔW_threshold) or even deletion of existing edge weights.

[0122] Timer Event: For example, trigger local optimization once every 1 hour or every day.

[0123] When an event is triggered, IGLE starts the incremental community detection algorithm (ICDA). For node / edge addition / deletion, the algorithm usually only recalculates the community affiliation of the affected nodes and their local neighborhood (Local Neighborhood, such as first or second neighbors), and evaluates its impact on the overall modularity, and if necessary, perform limited range of node movement or community split / merge. For timer triggers, it may perform batch local optimization on nodes connected by edges with recent significant weight changes. This greatly reduces the computational complexity and ensures the real-time performance of the graph update.

[0124] In addition to community structure, IGLE can also maintain and update other graph properties or embedded representations (e.g., node embeddings trained by incremental graph neural networks). The final output, the Dynamic Enterprise Data Asset Graph (DEDAG), is a continuously evolving knowledge graph. Its nodes represent data entities, rich in attributes (types, topics, sensitive labels, access heat, etc.); edges represent comprehensive association relationships, and weights dynamically reflect the strength of association; nodes form meaningful clusters through community detection; and the graph structure (nodes, edges, communities) can be automatically adjusted and optimized in quasi-real time as new inflow identification results, access trajectories, and permission metadata are recognized. This graph is a real-time, structured, and computable core representation of the enterprise data asset panorama, providing a solid foundation for subsequent value analysis, risk assessment, and governance decisions.

[0125] This step fuses static data identification results with dynamic behavior data, constructs a dynamic graph that reflects the real-time state of the enterprise data ecosystem by quantifying the multi-dimensional association strength between data entities (including content relevance, usage heat, and permission overlap), and dynamically calculates edge weights. The dynamic graph breaks through the static limitations of traditional data directories and visually presents the networked association characteristics of data assets. The edge weight calculated based on real-time behavior provides a quantitative basis for enterprise data flow and value transfer analysis.

[0126] S203, performing spatiotemporal evolution analysis on the dynamic enterprise data asset graph to extract potential data value density features and risk exposure features, wherein the spatiotemporal evolution analysis is performed by graph neural network time series prediction to quantify the value decay curve and compliance risk probability of data assets.

[0127] Specifically, based on the dynamic enterprise data asset graph, the graph can be divided into a sequence of continuous graph snapshots according to a time window, and a set of spatiotemporal evolution graph slices can be output.

[0128] The system periodically slices the dynamic enterprise data asset graph at preset time window lengths (Time Window Length, TWL, e.g., 30 days). At the end of each time window, the system captures a complete state snapshot of the graph, including all nodes (data entities), edges (association relationships), and real-time updated edge weights (Edge Weight, EW, a comprehensive value reflecting data association strength, access frequency, and permission correlation degree). For example, in the customer data governance scenario of a financial enterprise, the time window can be set to one natural month, and a snapshot containing nodes such as customer information table, transaction record library, risk assessment report, and their association relationships is generated at the end of each month. The snapshot sequence is strictly sorted by timestamp to form a set of spatio-temporal evolution graph slices (Spatio-Temporal Evolution Graph Slice Set, STEGSS). This set is essentially a sequence of graph versions stacked by time dimension, and each slice is marked with its valid time interval (e.g., 2023-Q1, 2023-Q2). To achieve efficient storage and retrieval, the system uses incremental snapshot technology (Incremental Snapshot Technology, IST) to record only the differences between adjacent slices (such as new nodes and weight changes), significantly reducing storage overhead.

[0129] The division of time windows needs to consider business rhythm and data life cycle. For example, a retail enterprise's sales data may use a weekly window (TWL=7 days) to adapt to the promotion cycle, while a research and development enterprise's patent data may use a quarterly window (TWL=90 days). The system has a built-in dynamic window adjustment algorithm (Dynamic Window Adjustment Algorithm, DWAA) that automatically optimizes TWL based on data update frequency: when the coefficient of variation (Coefficient of Variation, CV) of node attributes (such as access volume) exceeds a threshold (such as CV>0.5), the window is automatically shortened, otherwise it is lengthened. Each graph slice contains three types of core metadata:

[0130] Topology: connection relationship of nodes and edges; Dynamic weight: real-time edge weight (EW) synthesized by entropy weight method; Node attributes: such as data volume (Data Volume, DV), last access time (Last Access Time, LAT). Before outputting the slice set, it needs to pass the temporal consistency check (Temporal Consistency Check, TCC) to ensure that the node ID is stable and the weight change is continuous between adjacent slices, avoiding distortion of evolution analysis caused by data collection abnormalities.

[0131] To support efficient partitioning of large-scale graphs, the system adopts a distributed graph storage engine (DGSE), such as a sharded architecture based on Neo4j cluster or JanusGraph. The slicing operation is coordinated by a temporal slicing controller (TSC): the TSC listens to the graph update event stream, triggers a snapshot instruction when the TWL is reached, calls a graph state serializer (GSS) to convert the graph topology and weight matrix in memory into a binary slice file (SF) with timestamp metadata, and finally outputs the STEGSS to a time-series graph database (TSGD) for subsequent spatio-temporal analysis.

[0132] The input of the spatio-temporal graph convolutional network (STGCN) is a set of spatio-temporal graph snapshots (STEGSS), and the output is a value decay curve (VDC) prediction vector.

[0133] The spatio-temporal graph convolutional network (STGCN) is composed of alternating stacking of spatial convolutional layers and temporal convolutional layers. The spatial convolutional layer uses Chebyshev polynomial approximation (CPA) to handle non-Euclidean graph structures and calculate the degree of influence of each node on its neighbors. For example, in a supply chain data graph, when the edge weight (EW) between the "supplier inventory table" node (Node A) and the "production plan table" node (Node B) increases, the spatial convolution will increase the contribution weight of Node B to the feature update of Node A. The temporal convolutional layer uses one-dimensional causal convolution (1D-CC) to capture the temporal patterns of the historical slice sequence, ensuring that future predictions only rely on past information. The input of STGCN is the last K slices (e.g., K = 12 months) in STEGSS, and the output is the predicted state of the K+1 slice.

[0134] The quantification of the value decay curve (VDC) is achieved through a multi-task prediction head (MTPH). MTPH contains two parallel branches:

[0135] Node state regressor: predict the key attribute values of each node in the future slice, such as predicted access frequency (PAF), data age (DA);

[0136] Topology evolution simulator: predict the trend of edge weight change (such as the probability of weakening association).

[0137] Combining the two, the calculation formula of VDC is: Value Decay Rate = f(PAF decline slope, DA growth rate, correlation weight entropy increase).

[0138] For example, a certain customer portrait data currently has an average monthly access volume of 1000 times, and it is predicted that it will drop to 400 times in the next 3 months, and the associated edge weight of the active product will decay by 60%, so the value decay rate is as high as 70%. Finally, each node outputs a value decay prediction vector (VDPV), which contains a decay rate sequence for the next N time windows (such as N=4 quarters).

[0139] To train STGCN, supervised learning samples need to be constructed. The system extracts slice sequence-label pairs from historical STEGSS: the previous T slices are input, and the true state of the T+1 slice is the label. The model is optimized using mean squared error (MSE) and graph structure similarity loss (GSSL). To prevent overfitting, a spatio-temporal dropout mechanism (STD) is introduced to randomly mask some nodes or time steps. After the model is deployed, the prediction results are continuously updated through a sliding prediction window (SPW) to ensure that the VDPV reflects the latest data evolution trend in real time.

[0140] Combining the value decay prediction vector and historical compliance / risk records, the compliance risk probability of each data entity is calculated through a Bayesian risk model, and the risk exposure probability distribution is output;

[0141] The Bayesian risk model (BRM) takes the forward-predicted value decay features as observation evidence and historical compliance events as prior knowledge to calculate the posterior risk probability. The model input includes:

[0142] Value decay prediction vector (VDPV): reflects the trend of data utility decline; historical compliance / violation records set (HCVR): such as the number of data breach events, misuse of authority alerts, audit failure records; environmental risk factor (ERF): such as the current degree of regulatory rigor (such as the status of the implementation of the GDPR), the industry supervision intensity index. For example, if a certain employee's salary table has a history of violations due to unencrypted storage, and the current prediction shows a sharp decline in access (suggesting possible neglect), a high-risk determination is triggered.

[0143] The core of the BRM is the dynamic construction of the conditional probability table (CPT). The CPT defines the dependency relationship between key variables:

[0144] Prior probability P(R): based on HCVR statistics, the benchmark violation rate of each type of data (such as the violation rate of financial data = 0.05); likelihood P(VDPV|R): calculate the conditional probability of observing a specific VDPV feature when risk level R occurs, fit the historical data distribution by Gaussian Mixture Model (GMM); posterior probability P(R|VDPV): calculated according to Bayes' theorem, the formula is: P(high risk | VDPV) = P(VDPV | high risk) x P(high risk) / P(VDPV).

[0145] The system outputs a risk exposure probability distribution (REPD) for each node, including low risk (0-0.3), medium risk (0.3-0.7), and high risk (0.7-1.0) probability values.

[0146] Model calibration iteratively optimizes CPT parameters through the Expectation-Maximization Algorithm (EMA). For example, if a new compliance event is detected (such as a customer's data being accessed by unauthorized personnel due to incorrect permission configuration), the system will automatically: backtrack the VDPV features of the nodes involved in the event (such as a 40% decline in access volume in the 3 months before the incident); update the distribution parameters of P(VDPV|high risk); recalculate the REPD of all nodes in the graph.

[0147] When outputting the REPD, a risk confidence interval (RCI) is added to reflect the prediction uncertainty. Low confidence results (such as RCI width > 0.2) trigger the manual review process.

[0148] The value decay prediction vector and the risk exposure probability distribution are spliced into high-dimensional features, and the low-dimensional potential data value density features and risk exposure features are obtained by principal component analysis.

[0149] The value decay prediction vector (VDPV, dimension = 4, for example) and the risk exposure probability distribution (REPD, dimension = 3) of each data entity are spliced into a 7-dimensional raw feature vector (RFV). Since there is redundancy between dimensions (such as "access volume decline rate" and "value decay rate" are strongly correlated), the system uses principal component analysis (PCA) for feature compression and denoising. The core of PCA is the eigen decomposition of the covariance matrix (CM): calculate the eigenvalue (EV) and eigenvector (EVec) of CM, and select the principal component (PC) according to the descending order of EV.

[0150] The dimension reduction process consists of three steps:

[0151] Standardization: Z-score transformation (mean = 0, standard deviation = 1) is performed on each dimension of RFV to eliminate the influence of dimension; Principal component extraction: retain the principal components whose cumulative contribution rate (CCR) is greater than 85% (usually the first 2-3 PCs); Feature mapping: project the original RFV into the principal component space to generate a low-dimensional feature vector (LDFV).

[0152] For example, a certain RFV = [PAF decline = 0.6, DA increase = 0.8, high risk probability = 0.9,...] after PCA gets LDFV = [PC1 = 1.2, PC2 = -0.3], where PC1 represents "value-risk comprehensive intensity" and PC2 represents "decay speed and risk sensitivity balance".

[0153] The two key features finally output are:

[0154] Potential data value density feature (PDVDF): corresponds to the PC with the largest positive load in LDFV (such as PC1 > 0 indicates high value density), reflecting the current and future utility concentration of data;

[0155] Risk Exposure Feature (REF): Corresponding to the PC in LDFV that is strongly related to the risk probability (such as the negative part of PC1 or the special pattern of PC2), quantifying the exposure of data to compliance threats.

[0156] The system annotates the business meaning of the features after dimensionality reduction through the Feature Interpretation Module (FIM). For example, nodes with PDVDF>1.0 and REF<0.5 are labeled as "core high-value low-risk assets" and directly drive the generation of subsequent governance strategies.

[0157] This step uses a spatio-temporal graph neural network to sequence model the historical state of the graph, predicts the law of data value decay over time, and calculates the risk probability in combination with compliance event records. By analyzing the time-series change pattern of node and edge attributes, core features representing long-term data value and short-term risk are extracted, realizing dynamic evaluation and risk warning of data asset value, helping enterprises identify high-value data to be mined and high-risk data that need to be disposed of urgently. Spatio-temporal evolution analysis provides decision-making basis for data lifecycle management.

[0158] S204, input the data value density feature and the risk exposure feature into a self-organizing mapping network to generate a data asset hierarchical topology graph, wherein the self-organizing mapping network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning, forming a four-quadrant visual distribution, and each quadrant corresponds to a quantitative hierarchical label of value-risk;

[0159] Specifically, the data value density feature and the risk exposure feature can be combined into a high-dimensional feature vector, and Z-score standardization is performed to output a set of standardized feature vectors;

[0160] Feature vector merging and dimension processing

[0161] The system receives "Data Value Density Feature" (DVD_F) and "Risk Exposure Feature" (RE_F) from the spatio-temporal evolution analysis module. DVD_F is a numerical vector representing the potential economic utility of data assets (e.g., data usage frequency, business relevance, derived revenue prediction), and RE_F is a numerical vector quantifying compliance risks (e.g., violation probability, sensitive field density, access anomaly index). Before merging, the dimensions of the two feature vectors must be consistent - if DVD_F contains 5-dimensional indicators (e.g., timeliness value, scarcity weight, reuse potential, business criticality, derivative value coefficient), and RE_F contains 3-dimensional indicators (e.g., compliance deviation, leakage risk entropy, regulatory sensitivity), then the Feature Aligner (FA) automatically fills or projects to a unified dimension (e.g., 8 dimensions). The merging operation uses Vector Concatenation (VC) technology: sequentially concatenate the N elements of DVD_F with the M elements of RE_F to generate a High-Dimensional Feature Vector (HDFV) with a dimension of N+M (e.g., 8 dimensions). Each HDFV corresponds to an independent data entity (e.g., customer database table, production log set, design drawing library).

[0162] Z-score normalization process

[0163] Due to significant differences in the original feature dimensions (e.g., value density values range from 0 to 100, risk probability values range from 0 to 1), standardization is needed to eliminate scale effects. Z-score Normalization (ZN) algorithm is adopted:

[0164] Calculate mean and standard deviation: traverse the HDFV set of all data entities, and independently calculate the global mean (Mean, μ) and standard deviation (Standard Deviation, σ) for each feature dimension. For example, for the "timeliness value" dimension, calculate the global mean and standard deviation of this feature for all entities: and .

[0165] Standardize by dimension: for each entity's HDFV, perform conversion by dimension: standardized value = (original value - μ) / σ. For example, the original value of an entity's "leakage risk entropy" is 0.85, the global =0.62, =0.18, then the standardized value is (0.85-0.62) / 0.18≈1.28.

[0166] The process is implemented by a distributed statistics engine (DSE): the HDFV set is sharded to multiple computing nodes for parallel processing of μ and σ, and then broadcast to all nodes for standardization. The output result is a normalized feature vector set (NFVS) with a numerical distribution conforming to a Gaussian distribution with a mean of 0 and a standard deviation of 1, ensuring the stability of subsequent model training.

[0167] Outlier processing and boundary control

[0168] To prevent extreme values (such as an entity value density being abnormal) from distorting the standardization results, the system adds truncation rules (TR): if the standardized value of a feature dimension exceeds the preset range (such as [-3, 3]), it is forced to be truncated to the boundary value. For example, if a "derivative value coefficient" is standardized to 4.2, which exceeds +3, it is reset to 3.0. At the same time, the truncation event is recorded in the audit log (AL) for subsequent manual review. The final output NFVS serves as the ideal input source for the self-organizing map network.

[0169] Initialize the two-dimensional self-organizing map grid, input the normalized feature vector set into the self-organizing map network, and dynamically adjust the grid weights through unsupervised competitive learning, output the trained neuron weight matrix;

[0170] Grid initialization and topology

[0171] The self-organizing map (SOM) adopts a two-dimensional grid structure (GS), the size of which is dynamically set according to the size of the data entity (such as a 50x50 grid processing 100,000 entities). Each grid node is called a neuron (N), which contains a weight vector (WV) of the same dimension as the input features. During initialization, the principal component analysis initialization method (PCI) is used:

[0172] Perform principal component analysis (PCA) on the NFVS to extract the first two principal component directions (PC1, PC2).

[0173] The grid's horizontal and vertical axes are aligned with the PC1 and PC2 directions, respectively, and the neuron weights are linearly interpolated along the principal component directions. For example, the neuron WV at the upper left corner of the grid takes the minimum value of PC1 and the maximum value of PC2, and the neuron WV at the lower right corner takes the maximum value of PC1 and the minimum value of PC2. This method accelerates convergence and avoids local optimal traps caused by random initialization.

[0174] Unsupervised competitive learning mechanism

[0175] The training process is based on the Winner-Takes-All (WTA) principle:

[0176] Competitive phase: input a normalized feature vector (such as the NFV of a certain customer data entity), calculate its Euclidean distance (ED) with all neuron WV, and select the neuron with the smallest ED as the Best Matching Unit (BMU). For example, the BMU coordinates are (15, 32).

[0177] Collaborative phase: define a neighborhood function (NF) centered on the BMU, such as a Gaussian function. The weights of the neurons within the neighborhood are adjusted in the direction of the input vector, and the adjustment strength decays with the topological distance from the BMU. The neighborhood radius is initially large (covering 30% of the grid area), and shrinks exponentially to only the BMU itself.

[0178] Weight update formula: the weight of the i-th neuron at time t+1 is updated as: WV_i(t+1)=WV_i(t) +η(t)×NF(i, BMU, t)×[NFV - WV_i(t)]. The learning rate (Learning Rate, η) is initially 0.8 and linearly decreases to 0.01 with iterations.

[0179] Dynamic training and convergence control

[0180] The training is divided into two stages:

[0181] Coarse tuning phase: first 1000 iterations, neighborhood radius from initial value 10 to 1, learning rate from 0.8 to 0.2. Fast construction of global topology. Fine tuning phase: next 2000 iterations, neighborhood radius fixed at 1 (only BMU itself), learning rate from 0.2 to 0.01. Fine adjustment of weight vectors. Termination condition uses Quantization Error Threshold (QET): stop training when the average ED of all NFVs to the corresponding BMU changes less than 0.001 for 5 consecutive iterations. Final output Neuron Weight Matrix (NWM), dimensionality grid rows x columns x feature dimensionality (e.g. 50 x 50 x 8).

[0182] Based on the Neuron Weight Matrix, each data entity is mapped to a two-dimensional grid coordinate, and the grid is divided into four quadrants by K-means clustering, and a quadrant label mapping table is output, wherein the four quadrants include: high value-low risk, high value-high risk, low value-low risk, and low value-high risk.

[0183] Data entity grid coordinate mapping

[0184] Using the trained NWM, BMU Retrieval (BR) is performed on the standardized feature vector (NFV) of each data entity:

[0185] Calculate the Euclidean distance (ED) of NFV and all neuron weights in NWM. Select the neuron coordinates with the smallest ED as the mapped coordinates (MC) of the entity. For example, the NFV of a certain sales record set has the smallest ED with the neuron at grid (24, 17), so its MC is (24, 17). This process is optimized by Approximate Nearest Neighbor (ANN) algorithm: NWM is constructed as a K-Dimensional Tree (KDT) index, and the retrieval time is reduced from O(N²) to O(log N). The MC set of all entities forms a Grid Distribution Point Cloud (GDPC).

[0186] K-means four-quadrant division

[0187] To define the value-risk quadrants, clustering is performed on the neurons (non-data entities):

[0188] Feature extraction: Each neuron contains an 8-dimensional weight vector, from which the value-related dimensions (the first 5 dimensions corresponding to DVD_F) and the risk-related dimensions (the last 3 dimensions corresponding to RE_F) are separated. Calculate the Value Density Score (VDS) and Risk Exposure Score (RES) of the neuron: VDS = weighted sum of the first 5-dimensional weight vector (weights set by business experts); RES = maximum of the last 3-dimensional weight vector (highlighting the most significant risk).

[0189] Perform K-means clustering: Perform K-means clustering (KMC) on all neurons with (VDS, RES) as two-dimensional features, and the number of clusters K=4. The initial center selection uses the Max-Min Initialization (MMI) method: the first center selects the neuron with the maximum VDS and RES, and the subsequent centers select the neuron farthest from the existing centers.

[0190] Iterative optimization: Repeat the assignment of neurons to the nearest center and recalculate the center position until the center moves less than the threshold 0.01. Finally, obtain 4 cluster centers and their covered neuron sets.

[0191] Quadrant definition and label mapping

[0192] Define quadrants according to the (VDS, RES) coordinates of the cluster centers:

[0193] High value-low risk (HV-LR): VDS > total average and RES < total average; High value-high risk (HV-HR): VDS > total average and RES > total average; Low value-low risk (LV-LR): VDS < total average and RES < total average; Low value-high risk (LV-HR): VDS < total average and RES > total average. Generate a Quadrant Label Mapping Table (QLMT) to record the quadrant to which each neuron belongs (e.g., neuron (15,32) belongs to HV-HR). The quadrant of a data entity is determined by the quadrant of its MC neuron.

[0194] According to the Quadrant Label Mapping Table, calculate the value-risk average of data entities in each quadrant, and generate a labeled hierarchical grid containing quantitative hierarchical labels; input the labeled hierarchical grid into the rendering engine, add a heat map layer to represent the value density and risk exposure intensity, and generate a data asset hierarchical topology map with four-quadrant distribution.

[0195] Quantitative hierarchical label generation

[0196] Statistical indicators of entities in each quadrant are calculated based on QLMT:

[0197] For the HV-LR quadrant, the Mean Value Density (MVD) and Mean Risk Exposure (MRE) of all entities are calculated. The label is generated as: "High Value - Low Risk | Value Mean: 85±5 | Risk Mean: 0.15±0.03". Similarly, other quadrants are processed, and the label contains the Mean and Standard Deviation (SD) of key indicators, reflecting the stability of the group. Special entities (such as an entity with VDS>90 and RES<0.1) are added with a Diamond Label (DL), identifying strategic assets.

[0198] Labeling hierarchical grid construction

[0199] Quantitative labels are integrated with the neuron grid:

[0200] Neuron staining: Assign primary colors according to the quadrant - HV-LR is green (RGB: 0, 128, 0), HV-HR is yellow (RGB: 255, 255, 0), LV-LR is blue (RGB: 0, 0, 255), and LV-HR is red (RGB: 255, 0, 0). Label embedding: A semi-transparent text box is superimposed at the center of each quadrant, displaying the quantitative label of that quadrant. For example, in the HV-HR area, "High Value - High Risk | Value: 78±7 | Risk: 0.82±0.12" is displayed. Entity density visualization: Calculate the number of entities associated with each neuron and map it to the Radius Size (RS) of the grid point. For example, a neuron associated with 100 entities is displayed as a 10-pixel diameter circle, and a neuron associated with 1 entity is displayed as a 1-pixel point.

[0201] Thermal map layer rendering and output

[0202] Generate the final topology map through the Visualization Rendering Engine (VRE):

[0203] Value density heat map: with neuron VDS value as intensity, using Blue-Yellow-Red Gradient (BYRG), low value blue (VDS = 0), medium value yellow (VDS = 50), high value red (VDS = 100). Risk exposure heat map: with neuron RES value as intensity, using Opacity Gradient (OG), low risk opaque (RES = 0), high risk translucent (RES = 1), after superimposition, high risk area is dark color. Interactive function: support clicking neurons to view details (such as associated entity list, value / risk decomposition index), drag and rotate 3D view (if 2.5D rendering is enabled). Output is a vector format Data Asset Grading Topology Map (DAGTM), which can be directly embedded into enterprise data governance platform.

[0204] This step uses the dimension reduction capability of self-organizing mapping network to project complex value-risk characteristics to a two-dimensional plane, and automatically forms a four-quadrant distribution with semantic meaning through competitive learning. The boundary of each quadrant is dynamically determined by the clustering result of data characteristics, and a quantitative label is added to explain the strong and weak combination of value and risk, which converts abstract data characteristics into intuitive visual grading, enabling management personnel to quickly grasp the overall data asset distribution. The four-quadrant classification provides a clear framework for the development of differentiated governance strategies.

[0205] S205, based on the data asset grading topology map, generating an executable data governance action sequence through policy constraint reinforcement learning, wherein the action sequence includes automatic archiving, encryption enhancement, access permission reconstruction, and compliance audit trigger instructions, and the action priority is dynamically allocated by the grading matrix quadrant position.

[0206] Specifically, the four-quadrant position and quantitative grading label in the data asset grading topology map can be parsed and encoded into a state vector to obtain a grading state vector.

[0207] Data parsing and feature extraction

[0208] When the system receives the Data Asset Classification Topology Map, it first parses its core elements. The map is a two-dimensional grid structure, with the horizontal axis representing the Risk Exposure Feature (REF) and the vertical axis representing the Value Density Feature (VDF). The grid is divided into four quadrants: the first quadrant (High Value-Low Risk, HV-LR), the second quadrant (High Value-High Risk, HV-HR), the third quadrant (Low Value-Low Risk, LV-LR), and the fourth quadrant (Low Value-High Risk, LV-HR). Each data entity within each quadrant carries a Quantified Classification Label (QCL), such as:

[0209] Value Density Label: represented by a Value Score (VS) from 0 to 100, such as 95 points;

[0210] Risk Exposure Label: represented by a Risk Probability (RP) value, such as 0.85 (85% risk probability).

[0211] The system traverses the Grid Coordinate (GC) and label values of each data entity in the topology map, extracting the following features:

[0212] Quadrant ID (QID): an integer from 1 to 4;

[0213] Original numerical values of Value Score (VS) and Risk Probability (RP);

[0214] Euclidean Distance (ED) between the entity and the center point of the quadrant, used to measure the degree of deviation of its position within the quadrant.

[0215] State Vector Encoding Rules

[0216] Encode the above features into a fixed-dimensional Classification State Vector (CSV). The specific rules are as follows:

[0217] Quadrant feature: use One-Hot Encoding (OHE) to convert QID into a 4-dimensional vector. For example, QID=2 (HV-HR) is encoded as [0, 1, 0, 0].

[0218] Numerical features: Min-Max Normalization (MMN) is applied to VS and RP, scaling to the interval [0, 1]. For example, VS = 95 is normalized to 0.95, and RP = 0.85 remains unchanged.

[0219] Spatial features: Calculate the Relative Position Intensity (RPI) of entities within the quadrant, formula: RPI = 1 / (1 + ED). The higher this value, the closer the entity is to the core of the quadrant.

[0220] Finally, each data entity's CSV consists of 9 dimensions: [OHE_Q1, OHE_Q2, OHE_Q3, OHE_Q4, Norm_VS, Norm_RP, RPI, GC_X, GC_Y].

[0221] Among them, GC_X and GC_Y are grid coordinate values (such as [3, 5]), preserving the original position information to capture spatial distribution patterns.

[0222] Batch processing and vector storage

[0223] The system performs the above encoding on all entities in the topology graph, generating a set of hierarchical state vectors (CSV Set). This set is stored in the State Vector Database (SVDB) for subsequent reinforcement learning calls. At the same time, an Entity-Vector Mapping Index (EVMI) is established to ensure fast association between each data entity (such as "Customer Transaction Database Table") and its CSV.

[0224] Input the hierarchical state vector into the reinforcement learning framework, define the action space and constraint rules, and output the policy optimization framework;

[0225] Reinforcement learning framework initialization

[0226] Deep Reinforcement Learning Framework (DRLF) is adopted, its core components include: Agent: decision-making subject, responsible for selecting governance actions; Environment: simulates the state response of enterprise data systems; Reward Function (RF): quantifies the benefits and costs of actions.

[0227] The framework is based on Markov Decision Process (MDP) modeling, each decision cycle inputs CSV, and outputs action instructions.

[0228] Action Space Definition

[0229] Action Space (AS) contains four types of executable operations:

[0230] Auto-Archiving (AA): Migrate low-value data to cold storage, action coded as AA_Level (Archive Level, 1-3); Encryption Enhancement (EE): Enhance data encryption strength, action coded as EE_Type (e.g., AES-256 algorithm); Access Right Reconstruction (AR): Adjust user access rights, action coded as AR_Scope (permission range, such as department level / role level). Compliance Audit Trigger (CA): Start the audit process, action coded as CA_Urgency (urgency, high, medium, low).

[0231] Each action comes with parameters, for example, EE_Type optional values: 1=field-level encryption, 2=table-level encryption, 3=database-level encryption.

[0232] Constraint Rule Design

[0233] Constraint Rules (CR) ensure that actions comply with enterprise policies: Security constraints: High-risk data (RP>0.7) is prohibited from reducing encryption levels; Cost constraints: Archiving operations cannot trigger more than the budget threshold (e.g., $500) at a time; Dependency constraints: Permission reconstruction (AR) must be performed after encryption enhancement (EE) is completed; Compliance constraints: Financial data must meet GDPR provisions, and the audit (CA) trigger period cannot be less than 30 days.

[0234] Rules are coded into Rule Engine (RE), which pre-screens before action selection, eliminating rule-breaking action options.

[0235] The final output strategy optimization framework (Policy Optimization Framework, POF) includes the integrated logic of DRLF, AS, and CR.

[0236] In the policy optimization framework, a double deep Q network is used to train the agent, and a candidate action strategy set is output based on the trained agent;

[0237] Double Deep Q Network Architecture

[0238] Double Deep Q-Network (DDQN) is composed of two neural networks:

[0239] Main Network (MN): Real-time update parameters, output action value Q value; Target Network (TN): Regularly synchronize main network parameters, provide stable Q value estimation; Network input layer dimension = 9 (consistent with CSV dimension), hidden layer is 3 layers of full connection (neuron number: 128-64-32), output layer dimension = 12 (corresponding to 4 types of action × 3 parameter options).

[0240] Training process and reward mechanism

[0241] Reward function (RF) design example: successfully reduce high-risk data RP value: reward +10; high-value data (VS>80) is misfiled: penalty-20; violate cost constraints: penalty-15.

[0242] Experience Replay (ER): Store state transition records (CSV_t, Action, Reward, CSV_t+1), randomly extract batches for training to break data correlation.

[0243] Exploration strategy: Epsilon-Greedy Algorithm (EGA) is adopted, initial exploration rate ε=0.7, decay 0.05 per round of training, until ε=0.1.

[0244] Training continues until Q value converges (fluctuation amplitude <0.01) or reaches maximum iteration number (e.g. 10,000 times).

[0245] Candidate action policy generation

[0246] After training, the agent performs the following operations on each CSV:

[0247] Forward propagation: input CSV to main network to obtain Q values of 12 actions; constraint filtering: remove actions that violate CR through rule engine (RE); policy sorting: arrange the remaining actions in descending order of Q value to generate candidate action policy set (CAPS). For example, for a high-value-high-risk entity (QID=2), CAPS may be: [(EE_Type=3, Q=9.2), (CA_Urgency=high, Q=8.5), (AR_Scope=role level, Q=7.1) ].

[0248] The candidate action policy set is time-sequenced, a data governance action sequence is generated based on an action cost-benefit model, and a priority label dynamically assigned by a quadrant position is embedded, and finally an executable and prioritized data governance action sequence is obtained.

[0249] Time-dependent relationship modeling

[0250] The action sequence needs to meet two types of time logic: technical dependency: encryption (EE) must be completed before authority reconstruction (AR); business dependency: audit (CA) should be triggered after sensitive operations (such as EE / AR).

[0251] The system constructs an action dependency graph (ADG), the nodes are action types, and the edges represent execution order constraints (such as EE→AR). Through a topological sorting algorithm (TSA), the actions in the CAPS are globally sorted.

[0252] Cost-benefit model optimization

[0253] The cost-benefit model (CBM) quantifies the economy of each action: cost items (CI): computing resource consumption (such as CPU hours), storage fees, and manual labor hours; benefit items (BI): risk reduction (ΔRP), value retention rate (VS maintenance).

[0254] The net benefit (NB) formula is defined as: NB = α×ΔRP + β×VS - γ×CI. Where, α, β, γ are weight coefficients (such as α=0.6, β=0.3, γ=0.1), which are determined by fitting enterprise historical data.

[0255] On the basis of time sequencing, the action combination with the highest NB in the CAPS is selected to generate an initial sequence.

[0256] Priority Label (PL) is dynamically determined by Quadrant ID (QID): Quadrant 1 (HV-LR): Priority = 3 (lowest), action is mainly value maintenance (e.g. light archiving); Quadrant 2 (HV-HR): Priority = 1 (highest), forced immediate execution of encryption and audit; Quadrant 3 (LV-LR): Priority = 4, only needs periodic archiving; Quadrant 4 (LV-HR): Priority = 2, needs fast risk reduction (e.g. archiving or encryption). The final action sequence is arranged in ascending order of priority (PL = 1 is executed first), and the same priority is sorted by the NB value of CBM. For example, sequence example: 1. [Priority 1] Execute EE_Type = 3 (library-level AES-256 encryption) on "customer credit table"; 2. [Priority 1] Trigger CA_Urgency = high (urgent compliance audit); 3. [Priority 2] Archive "old version log backup" to AA_Level = 3 (deep cold storage); 4. [Priority 3] Set AR_Scope = department level (permission reconstruction) for "market analysis report". The output result is an executable instruction queue (Executable Action Sequence, EAS), which can be directly issued to the enterprise data governance platform.

[0257] Yet another embodiment of the application provides an AI recognition-based intelligent analysis system for enterprise data assets, which, as shown in Figure 3 , can include:

[0258] The receiving module 301 is configured to receive enterprise multi-source heterogeneous data streams, perform joint feature extraction and semantic alignment using a pre-trained multi-modal fusion recognition model, and generate a structured data asset recognition result, wherein the multi-modal fusion recognition model synchronously processes text, image, table and log data through a cross-modal attention mechanism to identify data types, content topics and sensitive information labels.

[0259] The construction module 302 is configured to construct a dynamic enterprise data asset graph based on the structured data asset recognition result and in combination with real-time collected data access tracks and permission metadata, wherein the nodes of the graph represent data entities, and the edge weights are dynamically calculated and generated according to data correlation, access frequency and permission correlation degree.

[0260] The extraction module 303 is configured to perform spatio-temporal evolution analysis on the dynamic enterprise data asset graph to extract potential data value density features and risk exposure features, wherein the spatio-temporal evolution analysis is performed through graph neural network time series prediction to quantify the value decay curve and compliance risk probability of the data asset.

[0261] The input module 304 is configured to input the data value density feature and the risk exposure feature into a self-organizing mapping network to generate a data asset hierarchical topology map, wherein the self-organizing mapping network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning to form a four-quadrant visualization distribution, and each quadrant corresponds to a quantitative hierarchical label of value-risk;

[0262] The generation module 305 is configured to generate an executable data governance action sequence based on the data asset hierarchical topology map through policy-constrained reinforcement learning, wherein the action sequence includes automatic archiving, encryption enhancement, access permission reconstruction, and compliance audit trigger instructions, and the action priority is dynamically allocated by the hierarchical matrix quadrant position.

[0263] The above detailed description of the embodiments shown in the drawings explains the structure, features and effects of the present application. The above description is only the preferred embodiments of the present application, but the present application is not limited by the drawings. Any changes or modifications made in accordance with the concept of the present application, or equivalent embodiments with equivalent changes, are still within the scope of the present application.

Claims

1. A method for intelligent analysis of enterprise data assets based on AI recognition, characterized in that, The method includes: The system receives multi-source heterogeneous data streams from enterprises, performs joint feature extraction and semantic alignment using a pre-trained multimodal fusion recognition model, and generates structured data asset recognition results. The multimodal fusion recognition model simultaneously processes text, image, table, and log data through a cross-modal attention mechanism to identify data types, content themes, and sensitive information tags. Based on the structured data asset identification results, and combined with the real-time collected data access trajectory and permission metadata, a dynamic enterprise data asset graph is constructed. The nodes of the graph represent data entities, and the edge weights are dynamically calculated and generated based on data association, access frequency, and permission association. Spatiotemporal evolution analysis is performed on the dynamic enterprise data asset map to extract potential data value density characteristics and risk exposure characteristics. The spatiotemporal evolution analysis uses graph neural network time series prediction to quantify the value decay curve of data assets and the probability of compliance risks. The data value density features and risk exposure features are input into a self-organizing map network to generate a data asset hierarchical topology map. The self-organizing map network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning, forming a four-quadrant visual distribution. Each quadrant corresponds to a quantitative hierarchical label of value-risk. Based on the data asset hierarchical topology, an executable data governance action sequence is generated through policy-constrained reinforcement learning. The action sequence includes instructions for automated archiving, encryption enhancement, access permission reconstruction, and compliance audit triggering, and the action priority is dynamically allocated by the quadrant position of the hierarchical matrix.

2. The method according to claim 1, characterized in that, The receiving enterprise's multi-source heterogeneous data streams are processed using a pre-trained multimodal fusion recognition model for joint feature extraction and semantic alignment to generate structured data asset recognition results. The multimodal fusion recognition model simultaneously processes text, image, table, and log data through a cross-modal attention mechanism to identify data types, content themes, and sensitive information tags, including: It receives multimodal data streams containing text, images, tables, and logs, unifies the data from each modality into a tensor sequence through an adaptive format parser, and outputs a standardized multimodal data tensor. Standardized multimodal data tensors are input into a pre-trained multimodal fusion recognition model. A gated cross-modal attention mechanism is used to calculate the interaction weights of text, image, table, and log, generating a joint feature embedding vector. Based on joint feature embedding vectors, contrastive learning is used to force semantic space alignment between different modalities, eliminate the semantic gap between modalities, and output an aligned semantically consistent feature matrix. The semantically consistent feature matrix is ​​input into the multilayer perceptron classification head to simultaneously identify data types, content themes, and sensitive information tags, and integrate them into structured data asset identification results.

3. The method according to claim 2, characterized in that, The process involves constructing a dynamic enterprise data asset graph based on the structured data asset identification results and real-time collected data access trajectories and permission metadata. Nodes in the graph represent data entities, and edge weights are dynamically calculated and generated based on data relationships, access frequency, and permission correlation. The structured data asset identification results are analyzed, data entities are extracted, and initial relationship edges are constructed based on the content theme similarity between entities, outputting a weighted initial relationship graph; The real-time collected data access trajectory stream is transformed into a time series access frequency vector, and the time series access frequency vector is fused with the weighted initial relationship graph. The edge weights are updated by aggregating through a sliding window, and a dynamic access enhancement relationship graph is output. Integrate permission metadata, calculate the permission association degree between data entities, and generate a permission association degree matrix; By combining the dynamic access enhancement relationship graph and the permission association matrix, the entropy weight method is used to dynamically synthesize the comprehensive edge weights and output the normalized weight graph. The normalized weighted graph is input into the incremental graph learning engine, the node clustering is optimized through the community detection algorithm, and a time-series trigger is embedded to respond to new data streams, outputting a dynamic enterprise data asset graph.

4. The method according to claim 3, characterized in that, The process involves performing spatiotemporal evolution analysis on the dynamic enterprise data asset map to extract potential data value density characteristics and risk exposure characteristics. This spatiotemporal evolution analysis utilizes graph neural network time-series prediction to quantify the value decay curve of data assets and the probability of compliance risks, including: Based on the dynamic enterprise data asset map, it is divided into a continuous map snapshot sequence according to the time window, and outputs a spatiotemporal evolution map slice set; The spatiotemporal evolution graph slice set is input into the spatiotemporal graph convolutional network to model the temporal dependencies of nodes and edges, predict future states, quantify the data value decay curve, and output the value decay prediction vector. By combining the value decay prediction vector and historical compliance / violation records, the compliance risk probability of each data entity is calculated using a Bayesian risk model, and the risk exposure probability distribution is output. The value decay prediction vector and the risk exposure probability distribution are concatenated into a high-dimensional feature. Principal component analysis is then used to reduce the dimensionality, resulting in low-dimensional potential data value density features and risk exposure features.

5. The method according to claim 4, characterized in that, The data value density features and risk exposure features are input into a self-organizing map network to generate a data asset hierarchical topology map. The self-organizing map network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning, forming a four-quadrant visual distribution. Each quadrant corresponds to a quantitative value-risk grading label, including: The data value density feature and risk exposure feature are merged into a high-dimensional feature vector and then standardized using Z-score to output a standardized feature vector set. Initialize a two-dimensional self-organizing map grid, input a standardized feature vector set into the self-organizing map network, dynamically adjust the grid weights through unsupervised competitive learning, and output the trained neuron weight matrix; Based on the neuron weight matrix, each data entity is mapped to a two-dimensional grid coordinate, and the grid is divided into four quadrants by K-means clustering, outputting a quadrant label mapping table, wherein the four quadrants include: high value-low risk, high value-high risk, low value-low risk, and low value-high risk. Based on the quadrant label mapping table, calculate the value-risk mean of data entities in each quadrant, and generate a labeled hierarchical grid containing quantitative hierarchical labels. The labeled hierarchical grid is input into the rendering engine, and a heat map is added to represent value density and risk exposure intensity, generating a hierarchical topology map of data assets with a four-quadrant distribution.

6. The method according to claim 5, characterized in that, Based on the data asset hierarchical topology map, an executable data governance action sequence is generated through policy-constrained reinforcement learning. This action sequence includes instructions for automated archiving, enhanced encryption, access permission reconstruction, and compliance audit triggering. The action priority is dynamically assigned based on the quadrant position in the hierarchical matrix, including: The four quadrant positions and quantitative classification labels in the data asset classification topology diagram are analyzed and encoded into state vectors to obtain classification state vectors. Input the hierarchical state vector into the reinforcement learning framework, define the action space and constraint rules, and output the policy optimization framework; In the policy optimization framework, a dual-deep Q-network is used to train the agent, and a set of candidate action policies is output based on the trained agent. The candidate action strategy set is sorted in time sequence, and a data governance action sequence is generated based on the action cost-benefit model. Priority labels dynamically assigned by quadrant position are embedded, and finally an executable, priority-based data governance action sequence is obtained.

7. An intelligent analysis system for enterprise data assets based on AI recognition, characterized in that, The system includes: The receiving module is used to receive multi-source heterogeneous data streams from enterprises, and use a pre-trained multimodal fusion recognition model to perform joint feature extraction and semantic alignment to generate structured data asset recognition results. The multimodal fusion recognition model processes text, image, table and log data simultaneously through a cross-modal attention mechanism to identify data types, content themes and sensitive information tags. The construction module is used to construct a dynamic enterprise data asset map based on the structured data asset identification results and combined with the real-time collected data access trajectory and permission metadata. The nodes of the map represent data entities, and the edge weights are dynamically calculated and generated based on data association, access frequency and permission association degree. The extraction module is used to perform spatiotemporal evolution analysis on the dynamic enterprise data asset map, and extract potential data value density characteristics and risk exposure characteristics. The spatiotemporal evolution analysis uses graph neural network time series prediction to quantify the value decay curve of data assets and the probability of compliance risks. The input module is used to input the data value density features and risk exposure features into the self-organizing map network to generate a data asset hierarchical topology map. The self-organizing map network maps high-dimensional features to a two-dimensional grid space through unsupervised competitive learning, forming a four-quadrant visual distribution. Each quadrant corresponds to a quantitative hierarchical label of value-risk. The generation module is used to generate an executable data governance action sequence based on the data asset hierarchical topology map through policy constraint reinforcement learning. The action sequence includes instructions for automated archiving, encryption enhancement, access permission reconstruction, and compliance audit triggering, and the action priority is dynamically allocated by the quadrant position of the hierarchical matrix.

8. The system according to claim 7, characterized in that, The receiving module is specifically used for: It receives multimodal data streams containing text, images, tables, and logs, unifies the data from each modality into a tensor sequence through an adaptive format parser, and outputs a standardized multimodal data tensor. Standardized multimodal data tensors are input into a pre-trained multimodal fusion recognition model. A gated cross-modal attention mechanism is used to calculate the interaction weights of text, image, table, and log, generating a joint feature embedding vector. Based on joint feature embedding vectors, contrastive learning is used to force semantic space alignment between different modalities, eliminate the semantic gap between modalities, and output an aligned semantically consistent feature matrix. The semantically consistent feature matrix is ​​input into the multilayer perceptron classification head to simultaneously identify data types, content themes, and sensitive information tags, and integrate them into structured data asset identification results.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-6.

Citation Information

Cited By

  • Multi-party collaborative enterprise data AI intelligent analysis and storage method

    CN121166040A

  • Logistics asset internet-of-things monitoring and processing system with unified multi-system data

    CN121234068A

  • A multi-system data unified logistics asset internet of things monitoring processing system

    CN121234068B

  • Data maintenance management method and system based on artificial intelligence

    CN121279622A

  • Classified hierarchical semantic enhanced dynamic graph community offset data access anomaly detection method and device

    CN121561369A