Enterprise data asset intelligent checking method and system based on knowledge graph

By using a knowledge graph-based approach, we can identify and track enterprise data assets, solving the problem of accurate identification and association of multi-source heterogeneous data, and achieving highly accurate and traceable data asset inventory.

CN122019554APending Publication Date: 2026-05-12QINGTIAN KUNDE TECHNICAL SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGTIAN KUNDE TECHNICAL SERVICE CO LTD
Filing Date
2026-01-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing enterprise data asset inventory methods are insufficient for accurately identifying, continuously tracking, and semantically associating multi-source, heterogeneous, and dynamically evolving data assets, resulting in omissions, duplications, and misjudgments in asset inventory results.

Method used

A knowledge graph-based approach is adopted to receive metadata, identify asset characteristics, determine similarity thresholds, use deep learning models for graph embedding and lifecycle tracking, construct change trajectories, and verify the initial inventory results.

Benefits of technology

It enables accurate identification and consistent correlation of enterprise data assets across time points, improves the accuracy, continuity and traceability of inventory, and provides highly reliable automated technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019554A_ABST
    Figure CN122019554A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of enterprise data management, in particular to an enterprise data asset intelligent checking method and system based on a knowledge graph, and the method comprises the steps: receiving to-be-checked data asset metadata, extracting asset features, and distributing a unique asset identifier according to a similarity threshold value; by comparing the asset feature similarity of adjacent time points, identifying the same asset across the time points and uniformly identifying the same asset, and generating an initial inventory result; processing the metadata by using a trained deep learning inventory model to obtain an initial knowledge node position and a graph embedding coordinate of the data assets in the knowledge graph, and performing life cycle tracking by combining an inventory tracking model to obtain a target tracking result; an asset change trajectory is constructed based on a graph embedding result, dynamic verification and correction are performed on an initial inventory result in combination with a tracking result, a final inventory result with high accuracy is output, and accurate identification, continuous tracking and automatic inventory of data assets are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise data management technology, and in particular to a method and system for intelligent inventory of enterprise data assets based on knowledge graphs. Background Technology

[0002] As enterprises deepen their digital transformation, data has become one of their core strategic assets. However, during the long-term process of IT construction, enterprises have accumulated a large amount of scattered, heterogeneous, and cross-system data resources, resulting in a wide variety of data assets with diverse sources and complex structures. Common problems include data silos, unclear assets, redundancy, and unclear responsibilities. Traditional data asset management methods mainly rely on manual inventory or simple statistics based on static metadata, making it difficult to achieve dynamic identification and accurate tracking of data assets throughout their entire lifecycle. Especially when dealing with frequently changing data tables, temporary datasets, or logically reused data interfaces, omissions, misjudgments, and duplicate counts are highly likely, seriously affecting the efficiency and accuracy of data governance.

[0003] In recent years, some enterprises have attempted to introduce automated tools or rule-based metadata analysis systems for data asset inventory. However, these methods generally lack a deep understanding of the semantic characteristics and evolutionary relationships of data assets, making it difficult to effectively identify data assets from the same source across time and systems, and also struggling to address misidentification caused by minor changes in data patterns. Furthermore, insufficient utilization of the correlations and contextual information between data assets results in an isolated and static inventory process, lacking contextual awareness.

[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this invention is to provide a knowledge graph-based intelligent inventory method and system for enterprise data assets. This aims to solve the technical problem that existing enterprise data asset inventory methods are unable to accurately identify, continuously track, and semantically associate multi-source heterogeneous and dynamically evolving data assets, resulting in omissions, duplications, and misjudgments in asset inventory results.

[0006] To achieve the above objectives, this invention provides a method for intelligent inventory of enterprise data assets based on knowledge graphs, the method comprising:

[0007] Receive data assets to be inventoried, identify the asset characteristics corresponding to the data assets to be inventoried included in the data assets to be inventoried ...

[0008] Determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, identify the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data assets to be inventoried across time points, unify the asset identifiers at multiple time points, and determine the number of all asset identifiers as the initial inventory result;

[0009] The metadata of the data assets to be inventoried is processed by a trained deep learning inventory model to obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried. The lifecycle of the data assets is then tracked based on a trained inventory tracking model to obtain the target tracking results.

[0010] Based on the initial knowledge node positions and the graph embedding coordinates, the change trajectories of each of the data assets to be inventoried are determined. Based on the change trajectories and the target tracking results, the initial inventory results are verified to obtain the inventory results.

[0011] Optionally, receiving the metadata of the assets to be inventoried includes:

[0012] Structured and unstructured data are acquired based on the enterprise's multi-source data acquisition interface. Semantic alignment and entity disambiguation are performed on the structured and unstructured data to obtain a snapshot of the data assets to be inventoried.

[0013] Metadata is extracted from the snapshot of the data assets to be inventoried at preset time intervals to obtain the initial metadata of the data assets to be inventoried.

[0014] Preprocessing operations are performed on the initial asset metadata to be inventoried to obtain the asset metadata to be inventoried. The preprocessing operations include schema unification, field normalization, and missing value filling.

[0015] Optionally, determining the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, and identifying the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data asset to be inventoried across time points, includes:

[0016] The asset features included in the metadata of multiple data assets to be inventoried at adjacent time points are time-series aligned, and a feature evolution window is established for the data assets to be inventoried corresponding to each asset feature in the metadata of the data assets to be inventoried at the current time.

[0017] Determine the actual similarity of the asset features at the corresponding time of the feature evolution window. When the actual similarity is not less than the preset similarity threshold, determine the data asset to be inventoried corresponding to the asset feature as the same data asset to be inventoried across time points.

[0018] When the actual similarity is less than the preset similarity threshold, the actual similarity between the asset feature included in the metadata of the data asset to be inventoried outside the current feature evolution window and the current asset feature is determined. When the actual similarity is not less than the preset similarity threshold, the data asset to be inventoried corresponding to the asset feature is determined as the same data asset to be inventoried across time points.

[0019] Optionally, identifying the asset characteristics corresponding to the data assets to be inventoried, as included in the metadata of each of the data assets to be inventoried, includes:

[0020] Extract local metadata features of the data assets to be inventoried, wherein the local metadata features include at least one of data pattern features, field constraint features, data format features, and data volume features;

[0021] Extract the global semantic features of the data assets to be inventoried, wherein the global semantic features include at least one of the following: business domain features, data sensitivity features, and ownership relationship features;

[0022] By fusing the local metadata features and the global semantic features, the initial asset features corresponding to the data assets to be inventoried are obtained. Vector embedding processing is then performed on the initial asset features to obtain the asset features.

[0023] Optionally, the step of verifying the initial inventory result based on the changed trajectory and the target tracking result to obtain the inventory result includes:

[0024] Construct a knowledge association graph model, in which each node in the knowledge association graph represents a data asset to be inventoried, and the edge weight represents the semantic association probability.

[0025] When the changed trajectory and the target tracking result meet the preset re-inspection conditions, the data assets to be inventoried are re-identified, and the inventory result is determined based on the re-identified data assets to be inventoried and the initial inventory result.

[0026] The preset re-detection conditions include at least one of the following: the actual similarity between asset features in the same change trajectory is greater than a preset mutation threshold; asset feature confusion exists in semantically overlapping areas between different change trajectories; and the deviation between the target tracking result and the change trajectory exceeds a preset fault tolerance threshold.

[0027] Optionally, after determining the actual similarity between any two asset features in the metadata of multiple asset data to be inventoried at adjacent time points, the method further includes:

[0028] Assign asset identifiers to the data assets to be inventoried that have an actual similarity less than the preset similarity threshold;

[0029] After determining the actual similarity of each of the data assets to be inventoried, the activity frequency of each asset identifier is determined. If the activity frequency is less than a preset frequency threshold, the asset identifier is deleted, and the number of all remaining asset identifiers is determined as the inventory result.

[0030] Optionally, the training process of the trained deep learning inventory model includes:

[0031] Construct a multi-level knowledge graph network, wherein the multi-level knowledge graph network includes a bottom-level network, a middle-level network, and a high-level network. The bottom-level network is used to extract field-level features, the middle-level network is used to capture entity relationship features, and the high-level network is used to integrate business semantic features.

[0032] The training metadata is input into the multi-level knowledge graph network, and the multi-level knowledge graph network is adjusted through temporal consistency constraints to obtain the initial deep learning inventory model.

[0033] The initial deep learning inventory model is jointly optimized and detected, and a relation loss function is added to obtain the trained deep learning inventory model.

[0034] Furthermore, to achieve the above objectives, the present invention also provides an intelligent inventory system for enterprise data assets based on knowledge graphs, the system comprising:

[0035] The feature identification module is used to receive the metadata of the data assets to be inventoried, identify the asset features corresponding to the data assets to be inventoried included in the metadata of the data assets to be inventoried, and assign different asset identifiers to the data assets to be inventoried where the asset similarity between any two asset features is less than a preset similarity threshold. The metadata of the data assets to be inventoried includes the data assets to be inventoried.

[0036] The cross-time matching module is used to determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, identify the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data assets to be inventoried across time points, unify the asset identifiers in multiple time points, and determine the number of all asset identifiers as the initial inventory result.

[0037] The model tracking module is used to process the metadata of the data assets to be inventoried through a trained deep learning inventory model, obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried, and perform lifecycle tracking of the data assets based on the trained inventory tracking model to obtain target tracking results.

[0038] The trajectory verification module is used to determine the change trajectory of each of the data assets to be inventoried based on the initial knowledge node position and the graph embedding coordinates, and to verify the initial inventory result based on the change trajectory and the target tracking result to obtain the inventory result.

[0039] Furthermore, to achieve the above objectives, the present invention also provides a knowledge graph-based intelligent inventory device for enterprise data assets. The device includes: a memory, a processor, and a knowledge graph-based intelligent inventory program for enterprise data assets stored in the memory and executable on the processor. The knowledge graph-based intelligent inventory program for enterprise data assets is configured to implement the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described above.

[0040] In addition, to achieve the above objectives, the present invention also provides a medium storing a knowledge graph-based intelligent inventory program for enterprise data assets, wherein when the knowledge graph-based intelligent inventory program for enterprise data assets is executed by a processor, it implements the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described above.

[0041] This invention provides an intelligent inventory method for enterprise data assets based on knowledge graphs. By integrating local and global features of metadata from multiple data assets and combining deep learning models with knowledge graph embedding technology, this method achieves accurate identification and consistent cross-time point association of enterprise data assets. It effectively solves the problems of misjudgment, duplicate counting, and omissions caused by minor changes in data patterns, naming differences, or dispersed sources in traditional inventory. By constructing asset change trajectories and introducing a lifecycle tracking mechanism, it can dynamically capture the evolution path of data assets and intelligently verify and correct initial inventory results using spatiotemporal and semantic consistency. This significantly improves the accuracy, continuity, and traceability of data asset inventory, providing highly reliable automated technical support for enterprise data governance, asset management, and compliance auditing. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating an embodiment of the intelligent inventory method for enterprise data assets based on knowledge graphs according to the present invention.

[0043] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0044] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0045] Reference Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the intelligent inventory method for enterprise data assets based on knowledge graphs according to the present invention.

[0046] In one embodiment, the knowledge graph-based intelligent inventory method for enterprise data assets includes:

[0047] Step S100: Receive the metadata of the data assets to be inventoried, identify the asset characteristics corresponding to the data assets to be inventoried included in the metadata of each data asset to be inventoried, and assign different asset identifiers to the data assets to be inventoried where the asset similarity between any two asset characteristics is less than a preset similarity threshold. The metadata of the data assets to be inventoried includes the data assets to be inventoried.

[0048] The metadata of the data assets to be inventoried can be structured data describing the technical attributes, business meaning, and source information of the data assets to be inventoried. It can be used as the input basis for the inventory method, providing the original feature information required for asset identification and association. In this embodiment, the metadata of the data assets to be inventoried can be obtained by collecting structured or semi-structured metadata information from databases, data lakes, API interfaces, or metadata repositories.

[0049] Asset features can be vectors or sets of attributes extracted from metadata to characterize the semantic and structural properties of data assets. They can be used to support asset similarity calculations and clustering, and serve as the basis for assigning asset identifiers. For example, asset features can be obtained by parsing metadata fields, extracting attributes such as table names, field lists, data types, annotations, and lineage relationships, and converting them into feature vectors. Furthermore, asset features can be implemented by extracting key attributes based on rule templates and concatenating them into feature vectors, or by using pre-trained language models to semantically encode metadata text to generate embedding vectors.

[0050] A preset similarity threshold can be used to determine whether two assets belong to the same logical entity. It can be used to control the granularity of asset clustering and achieve a balance between avoiding erroneous merging and excessive splitting.

[0051] Asset identifiers can be symbols or codes used to uniquely identify a logically independent data asset. They can be used to distinguish different assets during inventory checks, avoiding double counting or omissions, and maintaining asset identity consistency in cross-time point analysis. Furthermore, asset identifiers can work in conjunction with asset characteristics: determining whether to assign the same identifier based on feature similarity; and with change trajectories: assets under the same identifier form a continuous evolution chain in the trajectory.

[0052] The system receives metadata from assets awaiting inventory, which can be structured or semi-structured metadata information collected from databases, data lakes, API interfaces, or metadata repositories. Furthermore, receiving metadata from assets awaiting inventory can be achieved by batch-fetching metadata snapshots through standardized interface protocols, thus providing raw input for subsequent asset identification and modeling.

[0053] Identifying the asset characteristics corresponding to the data assets to be inventoried, as included in the metadata of each data asset to be inventoried, can be achieved by parsing metadata fields, extracting attributes such as table name, field list, data type, annotations, and lineage, and converting them into feature vectors. Furthermore, identifying the asset characteristics corresponding to the data assets to be inventoried, as included in the metadata of each data asset to be inventoried, can be achieved by extracting key attributes based on rule templates and concatenating them into feature vectors, or by using a pre-trained language model to semantically encode the metadata text to generate embedding vectors. This achieves the technical effect of transforming the raw metadata into a standardized representation that can be used for similarity calculation.

[0054] For assets in inventory data whose similarity between any two asset features is less than a preset similarity threshold, different asset identifiers are assigned. This can be achieved by calculating the similarity of all asset pairs; if the similarity is below the threshold, they are considered different assets and assigned a unique identifier. Furthermore, this operation can be implemented through pairwise similarity matrix construction and clustering segmentation algorithms, thereby achieving the technical effect of initially completing asset deduplication and identifier assignment, forming an initial asset set.

[0055] Step S200: Determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points. Determine the data assets to be inventoried that have an actual similarity of not less than a preset similarity threshold as the same data assets to be inventoried across time points. Unify the asset identifier across multiple time points and determine the number of all asset identifiers as the initial inventory result.

[0056] Among them, the metadata of multiple data assets to be inventoried at adjacent time points can be two or more sets of data asset metadata snapshots collected adjacently in the time series, which can be used to detect the evolution behavior of assets in a short period of time and support consistent association across time points.

[0057] Actual similarity can be a quantitative measure of the similarity between two data assets calculated based on asset characteristics. It can be used as an objective basis for determining whether they are the same asset and can be compared with a preset similarity threshold. In an exemplary embodiment, actual similarity can be obtained by calculating the structural similarity of a set of fields based on edit distance or Jaccard similarity coefficient, or by using the cosine similarity of graph embedding coordinates to measure semantic consistency.

[0058] The same data asset to be inventoried across time points can be data asset instances that appear at different time points but belong to the same entity semantically and logically. This can be used to ensure that inventory results are not counted repeatedly over time and to maintain the continuity of asset identity.

[0059] The initial inventory results can be preliminary asset counts based on the number of asset identifiers, and can be used as a baseline input for subsequent verification and correction.

[0060] Determining the actual similarity between any two asset features in the metadata of multiple assets to be inventoried at adjacent time points can be achieved by comparing cross-period asset features pairwise within a time window and calculating semantic or structural similarity scores. Furthermore, this operation can be implemented using a sliding time window mechanism combined with a feature vector similarity function, thereby achieving the technical effect of identifying potential homologous assets and providing a basis for cross-time point association.

[0061] Determining data assets to be inventoried that have an actual similarity of not less than a preset similarity threshold as the same data assets to be inventoried across different time points can be done by classifying asset pairs that meet the similarity criteria as the same logical entity. Furthermore, this operation can be implemented by setting a hard threshold or a soft clustering strategy, thereby achieving the technical effect of solving the problem of asset fragmentation caused by naming changes or minor structural adjustments.

[0062] Unifying asset identification across multiple points in time can be achieved by assigning the same identifier to different time instances identified as the same asset. Furthermore, this operation can be implemented by establishing a time-identifier mapping table and executing an identifier propagation algorithm, thereby ensuring the consistency of asset identity across time and avoiding duplicate counting.

[0063] Determining the quantity of all asset identifiers as the initial inventory result can be achieved by counting the total number of unique asset identifiers as the initial inventory output. Furthermore, this operation can be implemented by deduplicating and counting using a hash set, thereby achieving the technical effect of generating a baseline result that can be processed by subsequent verification processes.

[0064] Step S300: The metadata of the data assets to be inventoried is processed by the trained deep learning inventory model to obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried, and the lifecycle of the data assets is tracked based on the trained inventory tracking model to obtain the target tracking results.

[0065] A knowledge graph can be a knowledge base that represents entities and their semantic relationships using a graph structure. Nodes represent data assets or their attributes, and edges represent logical, lineage, or evolutionary relationships between assets. It can provide a unified semantic modeling framework for distributed, heterogeneous data asset metadata, supporting cross-system and cross-time point asset association and consistent reasoning. In this embodiment, the knowledge graph can be constructed by extracting entities, attributes, and relationships from multi-source metadata, followed by ontology alignment and graph pattern fusion. For example, the knowledge graph can include, but is not limited to, one or more of the following: data lineage graph, asset classification graph, and change evolution graph.

[0066] A deep learning inventory model can be a trained neural network model used to extract semantic features from the metadata of the data assets to be inventoried and generate graph embedding representations. It can be used to jointly encode the local structural features and global contextual semantics of data assets, improving the accuracy of asset similarity judgment. In one specific embodiment, the deep learning inventory model can be obtained through end-to-end training using graph neural networks or Transformer architectures based on historical inventory annotation data. Furthermore, the deep learning inventory model can include, but is not limited to, graph convolutional inventory models, graph attention inventory models, and sequence-graph hybrid inventory models.

[0067] An inventory tracking model can be a time-series modeling tool used to track changes in the state of data assets throughout their entire lifecycle. It can dynamically record the entire process of data assets from creation and modification to archiving or destruction, generating traceable tracking results. In this embodiment, the inventory tracking model can be obtained by combining graph embedding coordinates and timestamp information, and using recurrent neural networks or time-series graph networks to model the asset evolution path. Exemplary models may include, but are not limited to, state transition tracking models, event-driven tracking models, and version chain tracking models.

[0068] The initial knowledge node position can be the topological position of a data asset when it is first mapped in the knowledge graph. It can be used to reflect the contextual role of the asset in the overall semantic network and is used for subsequent trajectory construction.

[0069] Graph embedding coordinates can be a coordinate representation that maps knowledge nodes to a low-dimensional continuous vector space using knowledge graph embedding technology. This can be used to support efficient semantic similarity calculation and evolutionary pattern mining. Data assets can be data entities within an enterprise that have business value and can be identified and managed, such as tables, views, interfaces, or temporary datasets. They can be used as objects of inventory and constitute the basic units of enterprise data governance. Lifecycle tracking is a mechanism for monitoring and recording the state of data assets from creation to extinction. It can be used to provide a dynamic evolutionary perspective for inventory, enhancing the timeliness and completeness of the results. Target tracking results can be structured records of the lifecycle states of each data asset output by the inventory tracking model. These can be used to verify whether the initial inventory results conform to the actual evolutionary patterns of the assets.

[0070] Processing metadata of the data assets to be inventoried using a trained deep learning inventory model can be achieved by inputting the metadata into a trained neural network model for forward inference. Furthermore, this operation can be implemented by using a graph neural network to aggregate neighbor node information to generate node embeddings, or by employing a Transformer encoder to perform context-aware encoding on the metadata sequence, thereby achieving the technical effect of obtaining a high-dimensional feature representation that integrates local and global semantics.

[0071] Obtaining the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried can be achieved by mapping the model output to the node space and embedding vector space of the knowledge graph. Furthermore, this operation can be implemented using graph alignment algorithms or embedding projection layers, thereby achieving the technical effect of establishing the asset's location in the semantic network and supporting subsequent relational reasoning and trajectory construction.

[0072] Lifecycle tracking of data assets based on a trained inventory tracking model can be achieved by embedding coordinates and time-series information into an input graph, with the tracking model outputting the state evolution path. Furthermore, this operation can be implemented by using an LSTM network to model the temporal dependencies of asset states, or by using a temporal graph neural network to propagate updates across knowledge graph snapshot sequences, thereby achieving the technical effect of dynamically capturing the entire process of an asset from creation to change.

[0073] The target tracking results can be obtained by summarizing the asset state sequences output by the tracking model. Furthermore, this operation can be achieved by aligning structured log aggregation with the state sequences, thereby forming a structured asset evolution log for verifying the technical effectiveness of the initial inventory count.

[0074] Step S400: Determine the change trajectory of each data asset to be inventoried based on the initial knowledge node location and graph embedding coordinates; verify the initial inventory results based on the change trajectory and target tracking results; and obtain the inventory results.

[0075] The change trajectory can be an ordered path describing the structural, semantic, or contextual changes of a single data asset at multiple points in time. It can reflect the asset evolution process and provide a basis for verifying the spatiotemporal continuity of the initial inventory results. In this embodiment, the change trajectory can be derived based on the position and inference of the same asset identifier node in different time slices within a knowledge graph. For example, the change trajectory may include, but is not limited to, structural change trajectories, semantic drift trajectories, and context migration trajectories. The inventory result can be the final data asset inventory output after verification and correction, which can be used to provide a highly reliable asset list for enterprise data governance, compliance audits, etc.

[0076] Determining the change trajectory of each data asset to be inventoried based on the initial knowledge node position and graph embedding coordinates can track the knowledge node position shift and embedding vector changes of the same asset identifier at different points in time. Furthermore, this operation can construct trajectories through node path connectivity analysis or by fitting the evolution direction and rate based on the embedding vector difference sequence, thereby achieving the technical effect of constructing a visualized or computable path of asset evolution.

[0077] Verifying the initial inventory results based on changed trajectories and target tracking results can involve comparing the continuity of asset identifiers in the initial inventory with the consistency of the trajectories / tracking results to identify inconsistencies. Furthermore, this operation can be achieved by marking an asset identifier as omitted if it is interrupted in the trajectory but continues to exist as shown in the tracking results; or merging two identifiers as the same asset if their trajectories highly overlap and are semantically consistent. This effectively identifies misjudgments, omissions, or redundancies in the initial inventory.

[0078] The inventory results can be obtained by revising the initial inventory based on verification feedback and outputting the final asset list. Furthermore, this operation can be achieved through collaboration between a conflict resolution engine and a manual review interface, thereby achieving the technical effect of providing accurate, continuous, and traceable data asset inventory output.

[0079] Taking the asset inventory of a financial enterprise's data platform as an example, the intelligent inventory method for enterprise data assets based on knowledge graphs in this embodiment can be as follows: A bank needs to inventory all customer-related tables in its data platform during quarterly data governance. The system first receives metadata from hundreds of tables from the metadata management platform, including table names (such as cust_info_v2, customer_base_tmp), fields (id, name, phone), comments, and upstream lineage. Step S100 extracts the asset features of each table and calculates the similarity, finding that cust_info_v2 and the old version of cust_info are initially assigned different identifiers due to field fine-tuning. Step S200 compares the metadata of the previous month, calculates that the actual similarity is higher than the threshold, and unifies them into the same asset identifier. Step S300 maps each table to knowledge graph nodes through a deep learning inventory model and generates graph embedding coordinates; at the same time, the inventory tracking model confirms that the customer table is in an "active maintenance" state based on historical change logs. Step S400 constructs the change trajectory display field phone_type, which is newly added but whose core semantics remain unchanged. Combined with the tracking results, it is verified that the initial inventory should be merged. Finally, the corrected unique customer master data asset is output to avoid duplicate counting.

[0080] This embodiment provides an intelligent inventory method for enterprise data assets based on knowledge graphs. It receives metadata of the data assets to be inventoried and identifies asset characteristics, assigning different asset identifiers to assets with similarity below a threshold. By determining the actual similarity between metadata at adjacent time points, assets meeting the threshold condition are uniformly identified, generating initial inventory results. A deep learning inventory model generates initial knowledge node positions and graph embedding coordinates, and an inventory tracking model is used for lifecycle tracking to obtain target tracking results. The initial inventory results are verified and corrected by changing the trajectory and the target tracking results. This method maps dispersed and heterogeneous metadata to semantically related knowledge nodes, overcomes the semantic blind spots of static metadata, accurately clusters homogeneous assets to solve duplicate counting and omissions, and verifies the inventory results through dual constraints of spatiotemporal continuity and semantic consistency. This significantly improves inventory accuracy, continuity, and traceability, providing automated support for highly reliable data governance.

[0081] In one embodiment, the data metadata of the assets to be inventoried is received, including:

[0082] Structured and unstructured data are acquired based on enterprise multi-source data acquisition interfaces. Semantic alignment and entity disambiguation are performed on the structured and unstructured data to obtain a snapshot of the data assets to be inventoried.

[0083] Metadata is extracted from the snapshot of the data assets to be inventoried at preset time intervals to obtain the initial metadata of the data assets to be inventoried.

[0084] Preprocessing operations are performed on the initial asset metadata to be inventoried to obtain the asset metadata to be inventoried. The preprocessing operations include schema unification, field normalization, and missing value filling.

[0085] The enterprise multi-source data acquisition interface serves as a standardized access channel connecting various internal data systems (such as databases, log systems, APIs, and document repositories). It can be used to uniformly aggregate structured and unstructured data, resolving the issue of dispersed data sources. In this embodiment, the enterprise multi-source data acquisition interface can call standardized interfaces to synchronize raw data from sources such as databases, APIs, log systems, and document storage. For example, the enterprise multi-source data acquisition interface can achieve unified access to heterogeneous data sources, providing full input for semantic fusion.

[0086] Structured data can be data with a fixed format and well-defined fields, such as relational database tables. It can be used to provide directly parsable metadata fields, forming the basic source of asset characteristics. Unstructured data can be data without a fixed format or schema, such as text logs, configuration files, and interface documents. It can be used to extract implicit asset description information through semantic parsing, supplementing the deficiencies of structured metadata.

[0087] Semantic alignment is the process of mapping data elements from different systems, using different terminology but expressing the same business meaning, to a unified semantic space. It can be used to eliminate asset identity fragmentation caused by naming differences or terminological inconsistencies, improving the consistency of asset identification across systems. In an exemplary embodiment, semantic alignment can utilize ontology libraries, business dictionaries, or embedding models to identify cross-source synonyms and merge data records pointing to the same entity. Furthermore, semantic alignment can be achieved automatically through rule matching alignment based on a predefined business terminology list, or by using a pre-trained language model to calculate the semantic similarity of field descriptions. This allows for the construction of semantically consistent asset representations, eliminating identification barriers caused by differences in naming and expression. Exemplarily, semantic alignment can include, but is not limited to, one or more of ontology mapping alignment, dictionary rule alignment, and contextual embedding alignment.

[0088] Entity disambiguation is the process of identifying and merging data entities that refer to the same real-world object but have different representations. It can be used to prevent the same logical asset from being misclassified as multiple independent assets due to differences in representation, thus reducing duplicate counting. In one specific embodiment, entity disambiguation can utilize contextual information, data lineage, or business rules to determine whether different representations refer to the same entity. Furthermore, entity disambiguation can be implemented through context-based disambiguation, lineage-based disambiguation, and business rule-based disambiguation, thereby constructing semantically consistent asset representations and eliminating recognition barriers caused by differences in naming and representation. For example, entity disambiguation may include, but is not limited to, one or more of the following: context-based disambiguation, lineage-based disambiguation, and business rule-based disambiguation.

[0089] A snapshot of data assets to be inventoried can be a complete or representative capture of the status of multi-source data assets of an enterprise at a specific point in time. It includes a semantically unified view of both structured and unstructured data and can be used as a temporal reference source for metadata extraction, supporting the construction of asset observation sequences with temporal continuity. In this embodiment, the snapshot of data assets to be inventoried can be generated by integrating the raw data obtained from the enterprise's multi-source data acquisition interface after semantic alignment and entity disambiguation. Furthermore, the snapshot of data assets to be inventoried can integrate the semantically aligned and entity-disambiguated multi-source data into a single point-in-time view of the asset status, thereby forming a time slice with semantic unity and completeness, serving as the basis for metadata extraction.

[0090] The preset time interval can be a pre-defined time granularity parameter for periodically extracting metadata. It can be used to control the frequency of asset snapshot sampling, ensuring that the inventory process has time-series awareness and an evolution tracking foundation. In an exemplary embodiment, extracting metadata from the snapshot of the data assets to be inventoried at the preset time interval can be done by extracting table structure, fields, comments, and other metadata from the snapshot at a fixed period (such as daily or weekly), thereby generating a metadata sequence with time-series tags to support asset association across time points.

[0091] The initial metadata of the assets to be inventoried can be a set of raw metadata extracted directly from the snapshot of the assets to be inventoried at preset time intervals, without cleaning. It can be used as input for preprocessing operations, preserving the original asset status information. In this embodiment, obtaining the initial metadata of the assets to be inventoried can be achieved by temporarily storing the extraction results in their raw form without cleaning or standardization, thereby preserving the original asset status for subsequent preprocessing.

[0092] Preprocessing operations can be a set of data cleaning and transformation operations to standardize and improve the quality of initial data asset metadata to be inventoried. This can enhance the structural consistency, field completeness, and semantic clarity of the metadata, providing high-quality input for subsequent feature extraction and similarity calculation. In one specific embodiment, performing preprocessing operations on the initial data asset metadata to be inventoried can involve sequentially applying rules or models such as pattern unification, field normalization, and missing value imputation. Furthermore, preprocessing operations can standardize field names using regular expressions and mapping tables, or train an imputation model based on historical metadata distribution to predict missing field types or annotations, thereby improving metadata quality and ensuring the reliability of subsequent feature extraction. For example, preprocessing operations can include, but are not limited to, one or more of pattern unification, field normalization, and missing value imputation.

[0093] Schema unification can convert heterogeneous metadata schemas (such as field naming conventions and type definitions) from different systems into a unified standard format, which can be used to eliminate technical heterogeneity and make cross-system assets comparable. Field normalization can map semantically identical but differently named or formatted fields (such as 'cust_id' and 'customer number') to standard field names, which can be used to improve the consistency of asset features and support accurate similarity calculations. Missing value imputation can use rules or models to complete missing key attributes (such as annotations and data types) in metadata, which can be used to avoid asset feature distortion or inability to participate in clustering due to missing information.

[0094] Obtaining metadata for the assets to be inventoried can result in a preprocessed, structurally consistent, and semantically clear set of metadata, which can provide high-quality input for knowledge graph construction and deep learning models.

[0095] Taking the inventory of master data in a large manufacturing enterprise as an example, the intelligent inventory method for enterprise data assets based on knowledge graphs in this embodiment can be as follows: A manufacturing group needs to inventory its material master data distributed in ERP, MES, PLM, and manual Excel reports. The system synchronizes the material tables (structured) in OracleEBS, SAP annotation documents (unstructured), and local Excel lists through the enterprise multi-source data acquisition interface. After semantic alignment, 'Item_Code', 'Material Number', and 'PartNo.' are uniformly mapped to the standard field 'material_id'; entity disambiguation identifies that 10 variant names of the same material in different systems are actually the same entity. After integration, a snapshot of the data assets to be inventoried is formed at 2:00 AM every day. The system extracts metadata from it at preset time intervals every day to obtain an initial metadata dataset, in which some PLM systems have missing field annotations. The preprocessing stage performs unified execution mode (converting to ISO standard field names), field normalization (unifying the unit field to 'UOM'), and missing value imputation (inferring annotations based on similar materials), ultimately outputting high-quality asset metadata for inventory data, which can be used for subsequent knowledge graph modeling.

[0096] This embodiment provides a knowledge graph-based intelligent inventory method for enterprise data assets. It acquires structured and unstructured data through a multi-source data acquisition interface, performs semantic alignment and entity disambiguation on the structured and unstructured data to obtain a snapshot of the data assets to be inventoried. Metadata is extracted from the snapshot at preset time intervals to obtain initial metadata for the data assets to be inventoried. Preprocessing operations are performed on the initial metadata to obtain the final metadata for the data assets to be inventoried. These preprocessing operations include schema unification, field normalization, and missing value imputation. This method addresses data dispersion by unifying access to heterogeneous data sources and by using semantic alignment... By eliminating identification barriers caused by differences in naming and description through entity disambiguation, constructing a temporally continuous asset observation sequence by extracting metadata at preset time intervals, and improving the consistency and integrity of metadata structure through pattern unification, field normalization, and missing value filling, the system can ensure that the metadata input to knowledge graphs and deep learning models has the characteristics of high quality, semantic clarity, and temporal continuity. This supports accurate asset similarity calculation, cross-time point identification unification, and change trajectory construction, fundamentally alleviating the problems of misjudgment, duplication, and omission caused by data noise, semantic ambiguity, and temporal discontinuity in traditional methods, and significantly enhancing the robustness and credibility of inventory results.

[0097] In one embodiment, determining the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, and identifying the data assets to be inventoried with an actual similarity not less than a preset similarity threshold as the same data assets to be inventoried across time points, includes:

[0098] The asset features included in the metadata of multiple data assets to be inventoried at adjacent time points are time-series aligned, and a feature evolution window is established for the data assets to be inventoried corresponding to each asset feature in the metadata of the data assets to be inventoried at the current time.

[0099] Determine the actual similarity of asset features at the corresponding feature evolution window time. If the actual similarity is not less than the preset similarity threshold, the data assets to be inventoried corresponding to the asset features are determined as the same data assets to be inventoried across time points.

[0100] When the actual similarity is less than the preset similarity threshold, the actual similarity between the asset features included in the metadata of the data assets to be inventoried outside the current feature evolution window and the current asset features is determined. When the actual similarity is not less than the preset similarity threshold, the data assets to be inventoried corresponding to the asset features are determined as the same data assets to be inventoried across time points.

[0101] Temporal alignment can be a process of structurally mapping and synchronizing metadata of assets to be inventoried at different time points according to a unified time granularity or event sequence. This can be used to eliminate cross-period comparison biases caused by misaligned metadata collection times, inconsistent frequencies, or missing data, providing an aligned time benchmark for feature evolution analysis. In this embodiment, temporal alignment can sort and interpolate metadata snapshots based on timestamps or version numbers to establish a time index mapping table.

[0102] A feature evolution window can be a local temporal neighborhood constructed around a specific point in time, used to centrally observe the change patterns of specific asset features over a short period. It can be used to focus on the local dynamic behavior of asset features, improving the ability to perceive gradual changes (such as incremental field adjustments). In an exemplary embodiment, the feature evolution window can be centered on the current point in time, extending forward and / or backward over a fixed or adaptive time span to form a sliding or fixed window. Exemplarily, the feature evolution window can include, but is not limited to, one or more of forward evolution windows, backward evolution windows, and bidirectional symmetrical evolution windows.

[0103] The corresponding time point of the feature evolution window can be an instance of an asset feature that is time-aligned with the current asset feature within the feature evolution window. It can be used as a direct comparison object for calculating the actual similarity within the window, supporting preliminary cross-period association judgment.

[0104] The current asset characteristics can be the asset characteristic representation extracted from the asset metadata of the data to be inventoried at the current processing time point. It can be used as the central reference point for feature evolution window analysis and for comparison with other features inside and outside the window.

[0105] In addition to the current feature evolution window, the asset features included in the asset metadata of the data to be inventoried can be a set of asset features from historical or future time points outside the time range of the feature evolution window. This can be used to provide extended context when matching fails within the window, for secondary similarity verification, and to enhance robustness to discontinuous changes or delayed synchronization scenarios.

[0106] Time-series alignment of asset features included in the metadata of multiple assets to be inventoried at adjacent time points can be achieved by sorting, interpolating, or aligning metadata from multiple periods based on timestamps or version sequences, making each asset feature comparable on a unified timeline. Furthermore, this operation can be implemented by sorting and interpolating metadata snapshots based on timestamps or version numbers, thereby eliminating misjudgments caused by time asynchrony and ensuring that subsequent similarity calculations are based on the actual evolution path.

[0107] For each asset feature in the metadata of the assets to be inventoried, a feature evolution window is established. This can be achieved by dynamically defining a local time interval containing metadata from nearby time points for each current asset feature. Furthermore, this operation can be implemented by constructing a sliding window with a fixed time length (e.g., ±7 days) or by adaptively adjusting the window size based on the change frequency (using a narrower window for assets with high-frequency changes). This transforms global time series analysis into local evolutionary modeling, improving sensitivity to subtle and gradual changes.

[0108] Determining the actual similarity of asset features at the corresponding time points within the feature evolution window can be achieved by calculating the semantic or structural similarity between the current asset feature and the corresponding asset feature at the aligned time point within the window. Furthermore, this operation can be implemented using metrics such as vector embedding, edit distance, or field intersection ratio, thereby enabling a preliminary determination of whether assets maintain identity within their local temporal neighborhood.

[0109] When the actual similarity is not less than the preset similarity threshold, the data assets to be inventoried corresponding to the asset characteristics are identified as the same data assets to be inventoried across time points. This can be achieved by directly establishing a cross-period identity association and unifying the asset identifier if the similarity within the window meets the standard. Furthermore, this operation can be implemented by setting a threshold (such as 0.85) and merging asset IDs after successful matching, thereby efficiently completing the continuous identification of stable or slightly changing assets.

[0110] When the actual similarity is less than a preset similarity threshold, the actual similarity between the asset features included in the asset metadata of the data to be inventoried outside the current feature evolution window and the current asset features is determined. This can be achieved by expanding the search to historical or future metadata outside the window after a match fails within the window, to find potential matches. Furthermore, this operation can be implemented by traversing historical snapshots in reverse chronological order until the first high-similarity match is found, or by calculating the similarity of all time points outside the window in parallel and selecting the maximum value as the basis for secondary judgment. This can address the issue of connection interruption within the window caused by sudden renaming, structural reconstruction, or collection delays.

[0111] When the actual similarity is not less than a preset similarity threshold, the data assets to be inventoried corresponding to the asset characteristics are identified as the same data assets to be inventoried across time points. This can be achieved by establishing a cross-period identity association after finding matches that meet the threshold within an expanded search range. Furthermore, this operation can be implemented by backtracking and updating the asset lifecycle chain after a successful match outside the window, thereby improving the recall capability of "homogeneous but different" assets and reducing asset fragmentation caused by local mutations.

[0112] Taking the quarterly inventory of order data tables on an e-commerce platform as an example, the intelligent inventory method for enterprise data assets based on knowledge graphs in this embodiment can be as follows: When an e-commerce company is inventorying its main order table in Q2, it finds that the fields of the April version `order_main_v3` and the May version `order_core_final` are significantly different (e.g., `user_id` is renamed `customer_id`, and a `status` field is added). The actual similarity within the window (±15 days) is lower than the threshold. The system then initiates an out-of-window search and finds `order_main_v2` in the March snapshot, which is highly consistent with the current table in core fields (order ID, amount, time), and the similarity exceeds the threshold. Despite the drastic restructuring that occurred between April and May, the system still correctly identifies the three as the same logical asset through time-series alignment and out-of-window matching, avoiding duplicate counting, and assigning them a unified asset identifier to ensure the continuity of their lifecycle trajectories.

[0113] This embodiment provides an intelligent inventory method for enterprise data assets based on knowledge graphs. By performing temporal alignment on the asset features included in the metadata of multiple data assets to be inventoried at adjacent time points, a feature evolution window is established for each asset feature in the metadata of the data assets to be inventoried. The actual similarity of the asset features at the corresponding time of the feature evolution window is determined, and cross-period identity is confirmed when a threshold is met. If the threshold is not met within the window, the similarity is extended to features outside the window for secondary similarity verification and identity is determined again. By introducing a temporal alignment mechanism to eliminate the misalignment of the collection time, constructing a local feature evolution window to focus on dynamic changes, and combining a dual matching strategy inside and outside the window to deal with surface mutations, this method can achieve high-precision and high-coherence cross-time asset association in complex and frequently changing enterprise data environments, effectively improving the accuracy, continuity, and traceability of inventory results.

[0114] In one embodiment, identifying the asset characteristics corresponding to the data assets to be inventoried, as included in the metadata of each data asset to be inventoried, includes:

[0115] Extract local metadata features of the data assets to be inventoried. Local metadata features include at least one of the following: data pattern features, field constraint features, data format features, and data volume features.

[0116] Extract the global semantic features of the data assets to be inventoried. The global semantic features include at least one of the following: business domain characteristics, data sensitivity characteristics, and ownership relationship characteristics.

[0117] By integrating local metadata features and global semantic features, the initial asset features corresponding to the data assets to be inventoried are obtained. Vector embedding processing is then performed on the initial asset features to obtain the asset features.

[0118] Local metadata features can be a set of attributes extracted from the structured metadata of the data assets to be inventoried, reflecting their physical or logical structural characteristics. These attributes can be used to characterize the technical details of the data assets and support sensitive identification of structural changes such as minor schema changes and field additions / deletions. In an exemplary embodiment, local metadata features can be generated by reading column names, types, primary and foreign keys, and NOT NULL constraints from the database metadata table to form field constraint features, or by inferring the distribution pattern and scale from sampled data to form data volume features. Furthermore, local metadata features can include, but are not limited to, one or more of the following: data schema features, field constraint features, data format features, and data volume features.

[0119] Global semantic features can be high-level features describing the semantic roles and management attributes of data assets, obtained from governance or business contexts. They can be used to provide the business positioning and compliance attributes of assets within the enterprise data ecosystem, enhancing the ability to identify cross-system assets of the same origin. For example, global semantic features can be constructed by reading sensitivity level labels from a data classification and grading system as data sensitivity features, or by extracting department or responsible person information from a data lineage or responsibility matrix to construct ownership relationship features. In a specific embodiment, global semantic features may include, but are not limited to, one or more of the following: business domain characteristics, data sensitivity characteristics, and ownership relationship characteristics.

[0120] Initial asset features can be unembedded raw feature representations formed by fusing local metadata features and global semantic features. These features can be used to preserve both structural and semantic information, providing a complete context for subsequent embedding. It can be understood that initial asset features work in conjunction with local metadata features and global semantic features: these two are their sources of composition; and they work in conjunction with vector embedding processing: serving as its input.

[0121] Vector embedding can be a mathematical mapping process that transforms initial asset features into low-dimensional dense vector representations. This can be used to make asset features computable and support efficient similarity metrics and knowledge graph embedding. In one specific embodiment, vector embedding can employ a multilayer perceptron (MLP) for non-linear projection, or a pre-trained language model to encode concatenated feature text.

[0122] Extracting local metadata features of data assets to be inventoried can involve parsing and extracting technical attributes related to the data structure from the metadata. Furthermore, extracting local metadata features of data assets to be inventoried can be achieved by reading column names, types, primary and foreign keys, NOT NULL constraints, etc., from database metadata tables to generate field constraint features, or by inferring distribution patterns and scale from sampled data to form data volume characteristics. This allows for the acquisition of the asset's structural fingerprint, improving the accuracy of identifying pattern changes.

[0123] Extracting global semantic features of data assets to be inventoried can be achieved by obtaining the business and management attributes of the assets from business catalogs, data governance platforms, or permission systems. For example, extracting global semantic features of data assets to be inventoried can be done by reading sensitivity level tags from a data classification and grading system as data sensitivity features, or by extracting information on the department or person in charge from a data lineage or responsibility matrix to construct ownership relationship features. This introduces contextual semantics and solves the problem of business synonyms caused by relying solely on technical metadata.

[0124] By fusing local metadata features and global semantic features, the initial asset features corresponding to the data assets to be inventoried are obtained. This can be achieved by concatenating, weighting, or cross-combining the two types of features according to preset rules. In an exemplary embodiment, fusing local metadata features and global semantic features can be achieved by concatenating structured feature vectors with semantic label embedding vectors into a joint vector, or by dynamically weighting the importance of local and global features through an attention mechanism, thereby forming a composite feature representation that combines technical accuracy with business semantics.

[0125] Vector embedding is performed on the initial asset features to obtain asset features. This can be achieved by inputting the fused initial features into the embedding model and outputting a standardized vector representation. Furthermore, vector embedding can be performed on the initial asset features by using an autoencoder to compress the initial features into a fixed-dimensional embedding space, or by utilizing a contrastive learning framework to optimize the semantic similarity discriminative ability of the embedded vectors. This results in the generation of unified semantic vectors that can be used for similarity calculation and graph embedding, supporting accurate asset clustering.

[0126] Taking the integration of multi-source customer data assets in a retail enterprise as an example, the intelligent inventory method for enterprise data assets based on knowledge graphs in this embodiment can be as follows: A retail group needs to inventory customer-related tables from CRM, order system, and member mini-program. The system extracts local metadata features from each source: the CRM table contains the fields customer_id (primary key) and email (unique constraint), and the order system table contains cust_id and email_addr. The field names are different, but the data format is VARCHAR (255). At the same time, global semantic features are extracted: all three are marked as "customer domain", "personal sensitive information", and belong to "customer operations department". After fusion, initial asset features are formed, and high-dimensional vectors are generated through vector embedding. Although there are differences in field names, because the local structure (primary key + email) is highly consistent with the global semantics, the similarity of the embedded vectors exceeds the threshold, and they are judged as the same asset, avoiding duplicate counting caused by inconsistent naming.

[0127] This embodiment provides a knowledge graph-based intelligent inventory method for enterprise data assets. It extracts local metadata features and global semantic features from the data assets to be inventoried, fuses these features to obtain initial asset features, and performs vector embedding on these initial features to obtain the final asset features. By separately acquiring local metadata features reflecting the physical / logical structure of the data and global semantic features reflecting business intent and governance attributes, and fusing them into initial asset features that retain both information, and transforming them into a computable high-dimensional semantic vector, this method achieves both structural accuracy and context awareness in asset representation. It effectively distinguishes assets with the same semantics but different forms, avoiding misjudgments and duplicate counting, thereby significantly improving the robustness of asset identification and the semantic consistency of inventory results. This provides key technical support for building a high-fidelity data asset catalog.

[0128] In one embodiment, the initial inventory results are verified based on the changed trajectory and target tracking results to obtain the inventory results, including:

[0129] Construct a knowledge association graph model, where each node in the knowledge association graph represents a data asset to be inventoried, and the edge weight represents the semantic association probability.

[0130] The knowledge association graph model can be a model that models the semantic association strength between data assets to be inventoried using a graph structure. It can be used to explicitly express the contextual dependencies and semantic similarities between assets, providing a structured basis for re-detection triggering and confusion identification. In this embodiment, the knowledge association graph model can calculate the association score between asset pairs based on graph embedding coordinates or semantic features, and construct a weighted graph structure accordingly. For example, the knowledge association graph model can include, but is not limited to, one or more of the following: semantic proximity graph model, lineage-enhanced association graph model, and context co-occurrence association graph model.

[0131] A knowledge association graph can be a concrete graph instance of a knowledge association graph model, consisting of nodes and weighted edges. It can be used as a structural carrier for re-examination analysis and supports the visualization and quantitative reasoning of semantic associations.

[0132] A node can be a graph element in a knowledge association graph representing a single data asset to be inventoried. It can be used to carry asset identification and characteristic information and is the basic unit of association analysis.

[0133] Edge weights can be numerical values ​​attached to the edges connecting two nodes, representing the probability of semantic association between corresponding assets. They can be used to quantify the semantic similarity between assets and support confusion detection and clustering decisions.

[0134] Semantic association probability can be an estimate of the likelihood that two data assets belong to the same logical entity in terms of business or technical semantics. It can be used as the semantic basis for edge weights to determine whether assets should be associated or merged.

[0135] Constructing a knowledge association graph model can be based on the graph embedding coordinates or semantic features of the data assets to be inventoried, calculating the probability of association between each pair of assets, and building a weighted graph. Furthermore, the knowledge association graph model can be constructed by using cosine similarity or KL divergence to calculate the semantic distance between embedded vectors and converting it into probability weights, or by using a graph neural network to learn the potential association strength between nodes and output edge weights. This allows for the establishment of an explicit model of semantic dependencies between assets, providing a structural foundation for subsequent re-examination condition judgments.

[0136] Each node represents a data asset to be inventoried. This can be achieved by mapping each unique asset identifier to a node in the graph, thus aligning assets to entities in the graph structure.

[0137] Edge weights represent semantic association probabilities. They can be obtained by normalizing the semantic similarity of asset pairs to a probability value in the interval [0, 1] and assigning it to the corresponding edge, so that the graph structure has semantic interpretability and computability.

[0138] When the preset re-inspection conditions are met based on the changed trajectory and target tracking results, the data assets to be inventoried are re-identified, and the inventory results are determined based on the re-identified data assets to be inventoried and the initial inventory results.

[0139] The preset re-examination conditions can be a set of judgment rules used to determine whether the initial inventory results need to be re-identified. These rules can serve as a dynamic verification trigger mechanism to identify potential misjudgment scenarios and initiate a correction process. In an exemplary embodiment, the preset re-examination conditions may include, but are not limited to, one or more of the following: mutation detection conditions, trajectory confusion conditions, and tracking-trajectory deviation conditions.

[0140] The preset mutation threshold can be a critical value used to measure whether the change in the characteristics of adjacent assets in the same change trajectory constitutes an abnormal mutation. It can be used to identify discontinuous evolution behavior caused by field additions, deletions, type changes, etc., and prevent misjudgment as the same asset.

[0141] Semantic overlap regions can be areas in the semantic space where assets corresponding to different change trajectories have high similarity and overlap. They can be used to indicate that multiple assets may be incorrectly separated due to logical reuse or naming convergence and need to be merged.

[0142] Asset feature confusion can be a phenomenon where semantic overlap makes it difficult for the system to distinguish the identities of assets belonging to different change trajectories. It can be used as a re-examination trigger signal to indicate that asset identification needs to be re-clustered or reassigned.

[0143] Relationship bias can be the inconsistency between the asset evolution status reflected by the target tracking results and the path presented by the change trajectory. It can be used to reveal the conflict between life cycle tracking and map evolution, and to verify the authenticity of asset status.

[0144] The preset fault tolerance threshold can be the maximum deviation limit that allows a reasonable difference between the target tracking result and the changed trajectory. It can be used to filter out normal fluctuations and trigger a re-examination only when the deviation is significant, thus avoiding overcorrection.

[0145] When the changed trajectory and target tracking results meet the preset re-examination conditions, the data assets to be inventoried are re-identified. This can involve monitoring the consistency between the changed trajectory and the tracking results. If any re-examination condition is triggered, the process reverts to the feature extraction and similarity calculation stage. Furthermore, this operation can ensure global consistency by re-clustering only the local asset subsets involved in the triggering conditions, or by re-executing the S100-S200 process on all assets. This allows for dynamic correction of the initial inventory and improves the robustness of the results.

[0146] The inventory results are determined based on the re-identified assets to be inventoried and the initial inventory results. This can be achieved by merging or replacing the re-identified asset identifiers with the initial inventory results, thereby generating a final inventory output that has undergone closed-loop verification.

[0147] Determining whether the actual similarity between asset features within the same change trajectory exceeds a preset mutation threshold can be achieved by iterating through the asset features at adjacent time points within a single change trajectory, calculating their actual similarity, and comparing it to the mutation threshold. Furthermore, this determination can be made by calculating the Jaccard similarity change rate of the field set or by monitoring the Euclidean distance jump magnitude of the graph embedding coordinates, thereby identifying discontinuous evolution events such as table structure restructuring or logical replacement.

[0148] Determining asset feature confusion by identifying semantically overlapping areas between different change trajectories can be achieved by detecting whether the semantic similarity of the endpoints or intermediate nodes of multiple trajectories exceeds a confusion threshold. In one specific embodiment, this determination can be made by clustering the endpoint nodes of different trajectories; if the same cluster contains nodes from multiple trajectories, it is considered confusion. Alternatively, the minimum semantic distance between trajectories can be calculated; if it is below a threshold, it is marked as overlapping. This approach can uncover asset fragmentation issues caused by cross-system reuse or naming convergence.

[0149] Determining whether the deviation between the target tracking result and the changed trajectory exceeds a preset fault tolerance threshold can be achieved by comparing the state sequence output by the tracking model with the evolution path derived from the changed trajectory, quantifying the degree of inconsistency. For example, this determination can be made by calculating the edit distance between the trajectory node sequence and the tracking state sequence, or by assessing the offset between the two at the timestamps of key events (such as field changes), thereby revealing the conflict between model predictions and map evolution, and ensuring the consistency of the inventory logic.

[0150] For example, in the scenario of data asset governance on e-commerce platforms, the knowledge graph-based intelligent inventory method for enterprise data assets in this embodiment could be as follows: An e-commerce company discovers during its monthly inventory that user behavior log table v3 and the old version of the user behavior log table have been assigned different asset identifiers. The system constructs a knowledge association graph model, calculating the semantic association probability between the two to be 0.89, indicating a high edge weight. Further analysis of the change trajectory shows that within the same trajectory, the field from version v2 to v3 changes from (action_type, ts) to (event_type, timestamp), and the actual similarity drops sharply to 0.65, below the preset mutation threshold of 0.7, triggering a re-examination condition. Simultaneously, the target tracking model marks the asset as "continuously active," but the change trajectory is interrupted, and the relationship deviation exceeds the fault tolerance threshold. The system automatically re-identifies, combining the semantic association graph and deep learning features to confirm that the two are the same asset, merges the identifiers, corrects the initial inventory results, and avoids duplicate counting.

[0151] This embodiment provides an intelligent inventory method for enterprise data assets based on knowledge graphs. By constructing a knowledge association graph model and organizing the data assets to be inventoried as nodes, and quantifying their semantic association probability using edge weights, it introduces preset re-examination conditions as a dynamic verification trigger mechanism. When the conditions are met, the data assets to be inventoried are re-identified, and the initial inventory results are merged to generate the final inventory output. By detecting sudden changes in asset features within the same change trajectory, asset feature confusion caused by semantic overlap between different trajectories, and whether the deviation between the target tracking result and the change trajectory exceeds the fault tolerance threshold, it can achieve the technical effects of explicitly modeling the contextual dependencies between assets, effectively dealing with identity misjudgments caused by minor changes in data patterns or logical reuse, and realizing closed-loop correction of inventory results. This significantly improves the accuracy, interpretability, and audit credibility of data asset inventory.

[0152] In one embodiment, after determining the actual similarity between any two asset features in the metadata of multiple asset data to be inventoried at adjacent time points, the method further includes:

[0153] Assign asset identifiers to data assets to be inventoried that have an actual similarity of less than a preset similarity threshold.

[0154] This operation can be implemented by generating a unique identifier for assets that do not meet the cross-time point merging criteria. Furthermore, this operation can be achieved by assigning an independent asset identifier to each data asset that does not meet the similarity merging criteria, thereby ensuring that assets that do not meet the identity determination are counted independently, maintaining the integrity of the initial clustering.

[0155] After determining the actual similarity of each data asset to be inventoried in the metadata of all data assets to be inventoried, the activity frequency corresponding to each asset identifier is determined.

[0156] The activity frequency can be a quantitative indicator of the frequency or intensity of use of a certain asset identifier in the metadata of the data asset to be inventoried at multiple time points. It can be used to measure the actual activity level of the data asset and serve as a basis for determining whether it belongs to a valid business asset. In this embodiment, the activity frequency can be calculated by combining the number of times the asset identifier is observed within a continuous time window, or by combining auxiliary signals such as access logs and query frequency. For example, the activity frequency may include, but is not limited to, one or more of the following: frequency of occurrence within a time window, query call frequency, and downstream citation frequency.

[0157] If the activity frequency is less than the preset frequency threshold, delete the asset identifier.

[0158] The preset frequency threshold can be a critical value for determining whether an asset identifier belongs to a low-activity asset. It can be used to control the activity threshold for asset retention and filter temporary, obsolete, or experimental data assets. In an exemplary embodiment, the preset frequency threshold can be a short-term observation threshold, a long-term stable threshold, or a business scenario-adaptive threshold. Furthermore, the operation of deleting asset identifiers can be achieved by directly filtering low-frequency items from the asset identifier list or by marking them as "inactive" and excluding them during the final count. This can remove temporary tables, intermediate result sets, or obsolete assets, avoiding interference from "zombie assets" in the inventory results.

[0159] The quantity of all remaining assets identified is determined as the inventory result.

[0160] This operation can be performed by counting the set of asset identifiers after filtering by activity level. Furthermore, this operation can be achieved by performing a set cardinality check on the retained asset identifiers, thereby outputting a list of effective data assets focused on high business value and continuous use.

[0161] For example, in the scenario of quarterly inventory of an e-commerce platform's data warehouse, the knowledge graph-based intelligent inventory method for enterprise data assets in this embodiment could be as follows: When an e-commerce company is inventorying its DWD layer data assets, the system discovers a temporary table named user_behavior_temp. Its field structure is highly similar to the main user behavior table (actual similarity 0.82, higher than the preset similarity threshold of 0.8), but it only appears once in the past six metadata snapshots. After processing, this table is merged into the main asset identifier due to its high similarity. However, in the activity frequency assessment phase of this embodiment, this identifier is only active once in six time points, with an activity frequency of 1 / 6, lower than the preset frequency threshold (e.g., 0.3). Therefore, it is judged as a temporary intermediate table and its independent identifier is deleted (if not merged) or an abnormal fluctuation alarm for the main asset is triggered (if merged). The final inventory result only retains continuously active core assets, avoiding the inclusion of one-time ETL intermediate products in the total number of formal assets.

[0162] This embodiment provides an intelligent inventory method for enterprise data assets based on knowledge graphs. By assigning asset identifiers to data assets to be inventoried that have an actual similarity lower than a preset similarity threshold, determining the activity frequency of all asset identifiers, deleting asset identifiers when the activity frequency is lower than a preset frequency threshold, and determining the number of remaining asset identifiers as the inventory result, this method effectively identifies and removes temporary, experimental, or obsolete data assets by assigning independent identifiers to low-similarity assets to maintain cluster integrity, quantifying asset usage continuity by statistically analyzing activity frequency based on multi-time-point metadata, dynamically eliminating low-frequency assets based on preset frequency thresholds to exclude temporary or obsolete data, and counting the filtered asset identifiers to output a list of valid assets. This method avoids incorrectly including temporary, experimental, or obsolete data assets in the total number of valid assets, compensates for the potential for incorrect retention due to relying solely on static similarity thresholds, and enhances the focus of the inventory results on assets with real business value. The final inventory result not only reflects the existence of assets but also their activity level and lifecycle status within the enterprise data ecosystem, thereby significantly improving the accuracy, governance efficiency, and compliance audit credibility of data asset management.

[0163] In one embodiment, the training process of the trained deep learning inventory model includes:

[0164] Construct a multi-level knowledge graph network.

[0165] The multi-level knowledge graph network can be a hierarchical neural network structure used for feature extraction and fusion of data asset metadata at different semantic granularities. It comprises three processing layers: bottom, middle, and top, enabling multi-level feature modeling from field structure to business semantics, thus enhancing the richness and discriminative power of asset representation. In this embodiment, the multi-level knowledge graph network can be designed as a cascaded neural network architecture containing bottom, middle, and top layers, with each layer connecting to metadata inputs of different granularities. Furthermore, the construction of the multi-level knowledge graph network can be achieved by stacking graph convolutional networks to implement a three-layer structure, aggregating neighbors with different hop counts at each layer, or by using a multi-branch Transformer encoder to process fields, relationships, and business text in parallel before fusing the output. This allows for the establishment of an end-to-end feature learning path from technical details to business semantics.

[0166] The bottom-level network can be a sub-network in a multi-level knowledge graph network responsible for processing raw metadata field information. It can be used to extract fine-grained technical attributes such as field name, type, length, and value distribution to form basic structural features. In an exemplary embodiment, the bottom-level network can vectorize and encode the list of fields in the input metadata, preserving structural attributes. For example, the bottom-level network can be implemented by inputting field names into a multilayer perceptron (MLP) after concatenating them with word embeddings and types using a one-hot encoding, or by using a character-level convolutional neural network to extract field naming patterns as supplementary features, thereby generating basic representations that can distinguish different table structures. Furthermore, the bottom-level network can provide field-level input to the middle-level networks to support the construction of entity relationships.

[0167] The mid-level network can be a sub-network in a multi-level knowledge graph network used to model the relationships between data assets or across asset entities. It can capture structured relationships such as primary and foreign keys, lineage dependencies, and reused references, enhancing contextual understanding between assets. In one specific embodiment, the mid-level network can construct an asset relationship graph based on field-level features and learn relationship embeddings through a graph propagation mechanism. Furthermore, the mid-level network can leverage graph attention networks to weighted aggregate neighbor information on lineage edges, or construct subgraphs through relational path sampling and use GraphSAGE for inductive learning, thereby enabling the model to understand that assets do not exist in isolation but are situated within a complex dependency network. Furthermore, the mid-level network can receive field features output from lower-level networks and generate relationship embeddings for integration by higher-level networks.

[0168] The high-level network can be a top-level sub-network in a multi-level knowledge graph network used to integrate business semantic information. It can be used to integrate high-level semantics such as business terms, data classification tags, and usage scenarios, enabling the asset representation to have business interpretability. In this embodiment, the high-level network can be semantically aligned with the output of the mid-level network and external business ontology (such as a data dictionary or classification system). For example, the high-level network can be implemented by concatenating business tag embeddings and graph embeddings and then fusing them through cross-attention, or by using contrastive learning to align the semantic space of business terms and technical assets, thereby enabling the asset representation to have business readability and context adaptability. Furthermore, the high-level network can aggregate the relational features of the mid-level network and the external business ontology to output the final graph embedding coordinates.

[0169] Field-level features can be a set of features that describe the technical attributes of a single data field, such as field name, data type, whether it is empty, default value, etc. They can be used as the basic unit for asset structure identification and support similarity judgment at the table / view level.

[0170] Entity relationship characteristics can be features that reflect the logical or physical connections between data assets, such as primary and foreign key constraints, ETL lineage, API call dependencies, etc. They can be used to reveal the topological connections between assets and avoid misjudgments caused by isolated analysis.

[0171] Business semantic features can be high-order features that reflect the role and meaning of data assets in the business context, such as the subject domain, sensitivity level, business owner, etc. They can be used to improve the model's ability to identify homonymous or homonymous scenarios and enhance the business alignment of inventory results.

[0172] The training metadata is input into a multi-level knowledge graph network, and the multi-level knowledge graph network is adjusted through temporal consistency constraints to obtain the initial deep learning inventory model.

[0173] The training metadata can be a historical data asset metadata sample set used to train the deep learning inventory model. This sample set includes labeled or unlabeled snapshots from multiple time points and can serve as an input source for model parameter learning, driving the evolution of the feature extraction capabilities of the multi-level network. In this embodiment, inputting the training metadata into the multi-level knowledge graph network can involve organizing training samples by time slice or asset category and feeding them into the network for forward computation. This drives parameter updates at each level, gradually learning multi-level feature representations.

[0174] Temporal consistency constraints can be a regularization mechanism applied during model training, requiring the semantic stability of the embedded representations of the same data asset at different time points. This can be used to suppress embedding drift caused by minor pattern changes and ensure the consistency of asset identity across time points. In an exemplary embodiment, adjusting a multi-level knowledge graph network through temporal consistency constraints can involve adding a distance penalty term for the embeddings of the same asset across time points to the loss function and updating the network weights through backpropagation. Furthermore, temporal consistency constraints can be achieved by minimizing the L2 distance between embeddings at adjacent time points or maximizing the cosine similarity of embeddings across time periods to above a preset threshold, thereby ensuring that the model maintains identity stability for evolving assets. Temporal consistency constraints can include, but are not limited to, one or more of embedding distance constraints, identity preservation constraints, and trajectory smoothing constraints.

[0175] The initial deep learning inventory model can be an intermediate model version that has been trained forward through a multi-level knowledge graph network but has not yet been optimized by introducing a relation loss function. It can be used as a starting point for joint optimization and has basic multi-level feature extraction capabilities. In a specific embodiment, the initial deep learning inventory model can be obtained by completing preliminary training with temporal consistency constraints and then saving the model parameters, thereby obtaining asset representation capabilities with basic spatiotemporal consistency.

[0176] The initial deep learning inventory model is jointly optimized and detected, and a relation loss function is added to obtain the trained deep learning inventory model.

[0177] Joint optimization detection can be an evaluation and parameter tuning process that simultaneously optimizes multiple objectives (such as reconstruction loss, temporal consistency, and relation fidelity) in the later stages of model training. It can be used to balance learning objectives at different semantic levels and prevent features at one level from dominating the overall representation. In this embodiment, joint optimization detection on the initial deep learning inventory model can evaluate the model's performance in multiple dimensions such as structure reconstruction, relation prediction, and temporal stability on the validation set. This can identify the performance bottlenecks of the current model among multiple objectives and guide the direction of subsequent optimization.

[0178] A relation loss function is a loss term specifically designed to measure the modeling error of relationships between nodes in a knowledge graph. It is typically designed based on triples (head-relation-tail) and can be used to explicitly enhance the model's ability to capture semantic associations between assets, improving the relation fidelity of graph embeddings. For example, adding a relation loss function can be done by adding a relation modeling error term based on knowledge graph triples to the total loss, jointly optimizing it with the original loss. Furthermore, the relation loss function can be implemented by constructing positive and negative triples using negative sampling and calculating margin-based ranking loss, or by using mutual information maximization in knowledge graph embeddings as the unsupervised relation loss, thereby enhancing the model's accuracy in modeling semantic associations between assets. Relation loss functions can include, but are not limited to, one or more of TransE loss, RotatE loss, and graph-based relation loss.

[0179] The trained deep learning inventory model can be a final version with its parameters frozen after joint optimization, outputting a result usable for inference. This model can provide high-fidelity, high-consistency asset embedding generation capabilities to support subsequent inventory tasks. In one specific embodiment, the trained deep learning inventory model can be a final version with its parameters frozen after joint optimization, outputting a result usable for inference, thereby providing high-fidelity, high-consistency asset embedding generation capabilities to support subsequent inventory tasks.

[0180] Taking the inventory of customer data assets of a telecommunications operator as an example, the intelligent inventory method for enterprise data assets based on knowledge graphs in this embodiment can be as follows: A certain operator needs to inventory hundreds of tables scattered across its customer domain in the BSS / OSS system. During the training phase, the system collects metadata snapshots from the past 6 months as training metadata. In a multi-level knowledge graph network, the bottom layer extracts fields such as cust_id (string) and arpu (double); the middle layer identifies the lineage of these tables with the billing table and complaint table through cust_id; the top layer combines business tags such as "customer master data" and "high-value user tags" for semantic integration. Temporal consistency constraints are applied during training to ensure that even if a new field "five_g_flag" is added in a certain month, the embedding of the customer table is still highly consistent with the historical version. A relational loss function is introduced in the joint optimization phase so that the model can correctly judge that although the "customer basic information table" and the "customer profile wide table" have large structural differences, they are semantically closely related. The final trained deep learning inventory model can accurately map tables with different names (such as cust_base and customer_info_v3) but the same origin to the same asset identifier during inference, significantly improving the accuracy of inventory.

[0181] This embodiment provides a knowledge graph-based intelligent inventory method for enterprise data assets. It constructs a multi-level knowledge graph network, comprising a bottom-level network, a middle-level network, and a high-level network. The bottom-level network extracts field-level features, the middle-level network captures entity relationship features, and the high-level network integrates business semantic features. Training metadata is input into the multi-level knowledge graph network, and the network is adjusted through temporal consistency constraints to obtain an initial deep learning inventory model. The initial deep learning inventory model is then jointly optimized and tested, and a relationship loss function is added to obtain a trained deep learning model. The Xi Inventory Model extracts field-level, entity relationship, and business semantic features at different semantic granularities, introduces temporal consistency constraints to ensure the stability of asset identities across time points, and further strengthens relationship modeling through joint optimization and fusion of multiple objectives. This effectively overcomes the semantic blind spots caused by traditional methods that rely solely on shallow metadata, significantly enhancing the model's robustness to minor changes in data patterns, naming differences, and logical reuse scenarios. It provides a high-fidelity, high-consistency embedded representation foundation for accurately allocating asset identifiers, constructing change trajectories, and achieving lifecycle tracking, ultimately supporting the technical effect of high-accuracy and highly traceable intelligent inventory capabilities.

[0182] Furthermore, to achieve the above objectives, the present invention also provides an intelligent inventory system for enterprise data assets based on knowledge graphs, the system comprising:

[0183] The feature identification module is used to receive the metadata of the data assets to be inventoried, identify the asset features corresponding to the data assets to be inventoried included in the metadata of the data assets to be inventoried, and assign different asset identifiers to the data assets to be inventoried where the asset similarity between any two asset features is less than a preset similarity threshold. The metadata of the data assets to be inventoried includes the data assets to be inventoried.

[0184] The cross-time matching module is used to determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, identify the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data assets to be inventoried across time points, unify the asset identifiers in multiple time points, and determine the number of all asset identifiers as the initial inventory result.

[0185] The model tracking module is used to process the metadata of the data assets to be inventoried through a trained deep learning inventory model, obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried, and perform lifecycle tracking of the data assets based on the trained inventory tracking model to obtain target tracking results.

[0186] The trajectory verification module is used to determine the change trajectory of each of the data assets to be inventoried based on the initial knowledge node position and the graph embedding coordinates, and to verify the initial inventory result based on the change trajectory and the target tracking result to obtain the inventory result.

[0187] Other embodiments or specific implementations of the knowledge graph-based intelligent inventory system for enterprise data assets described in this invention can be found in the above-described method embodiments, and will not be repeated here.

[0188] Furthermore, to achieve the above objectives, the present invention also provides a knowledge graph-based intelligent inventory device for enterprise data assets. The device includes: a memory, a processor, and a knowledge graph-based intelligent inventory program for enterprise data assets stored in the memory and executable on the processor. The knowledge graph-based intelligent inventory program for enterprise data assets is configured to implement the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described above.

[0189] In addition, to achieve the above objectives, the present invention also provides a medium storing a knowledge graph-based intelligent inventory program for enterprise data assets, wherein when the knowledge graph-based intelligent inventory program for enterprise data assets is executed by a processor, it implements the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described above.

[0190] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A method for intelligent inventory of enterprise data assets based on knowledge graphs, characterized in that, The method includes: Receive data assets to be inventoried, identify the asset characteristics corresponding to the data assets to be inventoried included in the data assets to be inventoried ... Determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, identify the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data assets to be inventoried across time points, unify the asset identifiers at multiple time points, and determine the number of all asset identifiers as the initial inventory result; The metadata of the data assets to be inventoried is processed by a trained deep learning inventory model to obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried. The lifecycle of the data assets is then tracked based on a trained inventory tracking model to obtain the target tracking results. Based on the initial knowledge node positions and the graph embedding coordinates, the change trajectories of each of the data assets to be inventoried are determined. Based on the change trajectories and the target tracking results, the initial inventory results are verified to obtain the inventory results.

2. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, The receipt of asset metadata to be inventoried includes: Structured and unstructured data are acquired based on the enterprise's multi-source data acquisition interface. Semantic alignment and entity disambiguation are performed on the structured and unstructured data to obtain a snapshot of the data assets to be inventoried. Metadata is extracted from the snapshot of the data assets to be inventoried at preset time intervals to obtain the initial metadata of the data assets to be inventoried. Preprocessing operations are performed on the initial asset metadata to be inventoried to obtain the asset metadata to be inventoried. The preprocessing operations include schema unification, field normalization, and missing value filling.

3. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, The step of determining the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, and identifying the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data asset to be inventoried across time points, includes: The asset features included in the metadata of multiple data assets to be inventoried at adjacent time points are time-series aligned, and a feature evolution window is established for the data assets to be inventoried corresponding to each asset feature in the metadata of the data assets to be inventoried at the current time. Determine the actual similarity of the asset features at the corresponding time of the feature evolution window. When the actual similarity is not less than the preset similarity threshold, determine the data asset to be inventoried corresponding to the asset feature as the same data asset to be inventoried across time points. When the actual similarity is less than the preset similarity threshold, the actual similarity between the asset feature included in the metadata of the data asset to be inventoried outside the current feature evolution window and the current asset feature is determined. When the actual similarity is not less than the preset similarity threshold, the data asset to be inventoried corresponding to the asset feature is determined as the same data asset to be inventoried across time points.

4. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, The identification of asset characteristics corresponding to the data assets to be inventoried, included in the metadata of each of the data assets to be inventoried, includes: Extract local metadata features of the data assets to be inventoried, wherein the local metadata features include at least one of data pattern features, field constraint features, data format features, and data volume features; Extract the global semantic features of the data assets to be inventoried, wherein the global semantic features include at least one of the following: business domain features, data sensitivity features, and ownership relationship features; By fusing the local metadata features and the global semantic features, the initial asset features corresponding to the data assets to be inventoried are obtained. Vector embedding processing is then performed on the initial asset features to obtain the asset features.

5. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, The step of verifying the initial inventory results based on the changed trajectory and the target tracking results to obtain the inventory results includes: Construct a knowledge association graph model, in which each node in the knowledge association graph represents a data asset to be inventoried, and the edge weight represents the semantic association probability. When the changed trajectory and the target tracking result meet the preset re-inspection conditions, the data assets to be inventoried are re-identified, and the inventory result is determined based on the re-identified data assets to be inventoried and the initial inventory result. The preset re-detection conditions include at least one of the following: the actual similarity between asset features in the same change trajectory is greater than a preset mutation threshold; asset feature confusion exists in semantically overlapping areas between different change trajectories; and the deviation between the target tracking result and the change trajectory exceeds a preset fault tolerance threshold.

6. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, After determining the actual similarity between any two asset features in the metadata of multiple asset data to be inventoried at adjacent time points, the method further includes: Assign asset identifiers to the data assets to be inventoried that have an actual similarity less than the preset similarity threshold; After determining the actual similarity of each of the data assets to be inventoried, the activity frequency of each asset identifier is determined. If the activity frequency is less than a preset frequency threshold, the asset identifier is deleted, and the number of all remaining asset identifiers is determined as the inventory result.

7. The intelligent inventory method for enterprise data assets based on knowledge graphs as described in claim 1, characterized in that, The training process of the trained deep learning inventory model includes: Construct a multi-level knowledge graph network, wherein the multi-level knowledge graph network includes a bottom-level network, a middle-level network, and a high-level network. The bottom-level network is used to extract field-level features, the middle-level network is used to capture entity relationship features, and the high-level network is used to integrate business semantic features. The training metadata is input into the multi-level knowledge graph network, and the multi-level knowledge graph network is adjusted through temporal consistency constraints to obtain the initial deep learning inventory model. The initial deep learning inventory model is jointly optimized and detected, and a relation loss function is added to obtain the trained deep learning inventory model.

8. A knowledge graph-based intelligent inventory system for enterprise data assets, characterized in that, The system includes: The feature identification module is used to receive the metadata of the data assets to be inventoried, identify the asset features corresponding to the data assets to be inventoried included in the metadata of the data assets to be inventoried, and assign different asset identifiers to the data assets to be inventoried where the asset similarity between any two asset features is less than a preset similarity threshold. The metadata of the data assets to be inventoried includes the data assets to be inventoried. The cross-time matching module is used to determine the actual similarity between any two asset features in the metadata of multiple data assets to be inventoried at adjacent time points, identify the data assets to be inventoried with an actual similarity not less than the preset similarity threshold as the same data assets to be inventoried across time points, unify the asset identifiers in multiple time points, and determine the number of all asset identifiers as the initial inventory result. The model tracking module is used to process the metadata of the data assets to be inventoried through a trained deep learning inventory model, obtain the initial knowledge node positions and graph embedding coordinates of the data assets included in the data assets to be inventoried, and perform lifecycle tracking of the data assets based on the trained inventory tracking model to obtain target tracking results. The trajectory verification module is used to determine the change trajectory of each of the data assets to be inventoried based on the initial knowledge node position and the graph embedding coordinates, and to verify the initial inventory result based on the change trajectory and the target tracking result to obtain the inventory result.

9. A knowledge graph-based intelligent inventory device for enterprise data assets, characterized in that, The device includes: a memory, a processor, and a knowledge graph-based intelligent inventory program for enterprise data assets stored in the memory and executable on the processor, the knowledge graph-based intelligent inventory program for enterprise data assets being configured to implement the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described in any one of claims 1 to 7.

10. A medium, characterized in that, The medium stores a knowledge graph-based intelligent inventory program for enterprise data assets. When the knowledge graph-based intelligent inventory program is executed by a processor, it implements the steps of the knowledge graph-based intelligent inventory method for enterprise data assets as described in any one of claims 1 to 7.