Data asset management system and method for medical system
By generating structural and semantic fingerprints, a hierarchical asset model is constructed, which solves the problems of heterogeneity and cross-system association difficulties in the hospital's internal data management system, and realizes efficient and automated data management and rapid innovative applications.
Patent Information
- Application Number
- CN202511469826.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-11
AI Technical Summary
Due to heterogeneity and difficulties in cross-system connections, the hospital's internal data management system suffers from low data utilization efficiency, making it difficult to achieve rapid innovative applications and real-time access, and also resulting in low efficiency in data quality monitoring and compliance auditing.
By acquiring data source metadata from multiple business systems, a set of structural metadata is generated and synchronized using a consistent window approach. Structural and semantic fingerprints are calculated, clustering and merging are performed, an asset meta-model is constructed, a hierarchical asset model is generated, and data access configuration is output.
It enables accurate, efficient, and timely synchronization of data from multiple business systems, automatically identifies and uniformly manages semantically similar or duplicate data, improves the automation and accuracy of data asset management, reduces maintenance costs, and meets the needs of rapid innovation and real-time access.
Smart Images

Figure CN120932840A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data asset management technology, and more specifically, to a data asset management system and method for use in medical systems. Background Technology
[0002] Various hospital information systems (HIS), laboratory information systems (LIS), electronic medical record systems (EMR), operations management systems, and procurement management systems are widely deployed and operate long-term in hospitals. These systems continuously generate a large amount of structured, semi-structured, and unstructured data resources in daily medical activities, equipment maintenance, and logistics management. Currently, internal data management in hospitals mainly relies on reports or export tools provided by each business system itself, or on ETL (Extract-Transform-Load) methods to aggregate some data into a data warehouse for analysis. However, different systems are implemented by different vendors or technical architectures, resulting in inconsistent data table structures, field naming, business semantics, and update strategies, leading to heterogeneous data sources and difficulties in cross-system correlation.
[0003] To improve data utilization efficiency, some hospitals have built data catalogs, data lineage, and master data management (MDM) platforms, attempting to organize tables, fields, and their inter-reference relationships. However, these platforms mostly rely on manual registration and template filling, making it difficult to achieve full coverage, dynamic updates, and deep semantic understanding. They often only manage structured metadata and cannot provide consistent management of semi-structured documents and free text. Furthermore, existing systems often require manually writing SQL scripts or interface documentation when providing data interfaces or views, resulting in long development cycles and high maintenance costs, failing to meet hospitals' needs for rapid innovative applications and real-time data access from third-party systems. In addition, data quality monitoring, sensitive information control, and compliance auditing also largely depend on offline investigation and manual review, which is inefficient and difficult to achieve a closed-loop system.
[0004] In summary, there is an urgent need for a solution that can automate the inventory, model-based management, and continuously track data assets of various business systems. This solution should be able to uniformly and compatiblely handle heterogeneous data from multiple sources, and dynamically support the rapid opening and deep empowerment of data while ensuring security and compliance. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a data asset management system and method for medical systems.
[0006] Firstly, this application provides a method for data asset management in a medical system, including:
[0007] Obtain data source metadata from multiple hospital business systems, wherein the data source metadata includes: connection configuration, data tables, and field list of the target business database;
[0008] According to the preset collection strategy, the data source metadata in the multi-service system is synchronized to generate a set of structural metadata; the generation of the set of structural metadata also includes: extracting the structural definition of each target business database in a consistent window manner, and generating a set of structural metadata based on the structural changes generated during the incremental data definition language log or trigger event supplement window.
[0009] Based on the aforementioned structural metadata set, structural fingerprints and semantic fingerprints are generated for each data table, fusion similarity is calculated, and clustering and merging are performed under the adaptive threshold constraint of the business domain to generate asset entity nodes and construct an asset meta-model.
[0010] For the aforementioned asset meta-model, semantic parsing processing is performed to generate a semantically enhanced model;
[0011] The semantic enhancement model is mapped to a hierarchical asset model, which includes: a raw layer, a standard layer, and a business layer.
[0012] Based on the hierarchical asset model, a data access configuration is generated, and the corresponding data asset ledger is output.
[0013] As an optional implementation, obtaining the data source metadata of the hospital's multi-service system includes:
[0014] Perform a dependency scan on the system directory of the target business database to extract reference object information associated with each data table, wherein the reference object information includes: views, stored procedures, and triggers;
[0015] Record the referenced object information as dependency metadata;
[0016] The dependency metadata is merged with the connection configuration, data table, and field list to generate data source metadata.
[0017] As an optional implementation, the extraction of reference object information associated with each data table includes:
[0018] Query the system directory to obtain the dependency records between the target object and its first-level referenced objects;
[0019] For the first-level referenced object, continue to query the system directory recursively until all referenced objects are basic data tables, generating multi-level nested dependency records;
[0020] The multi-level nested dependency records are deduplicated and topologically sorted to form a complete directed acyclic dependency chain, and the dependency chain is written into the dependency relationship metadata.
[0021] As an optional implementation, forming a complete directed acyclic dependency chain includes:
[0022] Label each object in the dependency record with its business domain identifier and sensitivity level weight;
[0023] The dependency records are divided into domains according to the business domain identifier, and a subgraph is created for each business domain.
[0024] Within each subgraph, a weighted topological sort is performed first by hierarchy and then by weight to obtain the corresponding subgraph sorting result;
[0025] Based on the subgraph sorting result, a ternary layer sequence number is generated and written into the dependency metadata; wherein, the ternary layer sequence number includes: business domain number, topology layer number, and sensitivity level.
[0026] As an optional implementation, the generated structure metadata set includes:
[0027] Obtain the structure version identifier of each target business database and record it as the synchronization start marker;
[0028] Based on the synchronization start marker, open the same transaction consistency window, and extract the structure definition of each target business database within the consistency window to generate the first structure metadata subset;
[0029] After completing the extraction of the first structural metadata subset, based on the structural change logs or data definition language trigger events set for each target business database, the incremental structural changes generated before the consistency window is closed are extracted to generate the second structural metadata subset.
[0030] The first and second subsets of structure metadata are merged and deduplicated to obtain the final set of structure metadata.
[0031] As an optional implementation, the construction of the asset meta-model includes:
[0032] For each data table in the structure metadata set, a structure fingerprint and a semantic fingerprint are generated. The structure fingerprint is used to characterize the number of fields, primary key type, and field type distribution. The semantic fingerprint is based on a pre-trained language model to obtain semantic vectors of table name, field name, and field annotation.
[0033] For any two data tables, calculate the structural fingerprint distance and the semantic fingerprint cosine similarity, and obtain the comprehensive similarity score between the tables according to the preset weighting coefficient;
[0034] Based on the adaptive threshold of the business domain, clustering and merging are performed on data tables whose comprehensive similarity score is not lower than the adaptive threshold of the business domain to form asset entity nodes;
[0035] Establish a mapping relationship between the asset entity nodes and the corresponding original data tables, and write them into the asset meta-model.
[0036] As an optional implementation, the construction of the asset meta-model further includes:
[0037] A predetermined number of record samples are extracted from the business database corresponding to each data table in the structure metadata set to obtain a field value sample set;
[0038] For each numeric field in the sample set of field values, calculate the minimum, maximum, mean, and variance;
[0039] For each categorical field in the field value sample set, a preset number of categories are selected in descending order of category frequency, and their frequencies are counted. The proportion of null values is also counted to generate a data distribution fingerprint.
[0040] Calculate the similarity of the data distribution fingerprint between any two data tables to obtain the data distribution fingerprint similarity;
[0041] When calculating the overall similarity score between any two data tables, the data distribution fingerprint similarity, structural fingerprint distance, and semantic fingerprint cosine similarity are weighted and fused according to a preset distribution weight to adjust the overall similarity score between the tables.
[0042] Adaptive threshold clustering and merging are performed based on the adjusted inter-table comprehensive similarity score to form the asset entity nodes.
[0043] As an optional implementation, the generative semantic enhancement model includes:
[0044] For the asset entity node, its semantic fingerprint, structural fingerprint and data distribution fingerprint are extracted respectively. Based on the scenario weight determined by the business domain or field type, the semantic fingerprint, structural fingerprint and data distribution fingerprint are weighted and concatenated to generate a context representation vector for semantic disambiguation.
[0045] The context representation vector is matched with the concept vector in the medical terminology ontology to obtain a set of candidate semantic concepts;
[0046] From the set of candidate semantic concepts, candidate semantic concepts with a similarity of not less than the domain threshold are selected as the standard semantic labels for the asset entity nodes;
[0047] If the similarity of all candidate semantic concepts is lower than the domain threshold, a manual review tag is generated for the asset entity node.
[0048] The standard semantic tags or tags to be manually reviewed are written into the semantic enhancement model to complete the semantic enhancement of asset entity nodes.
[0049] As an optional implementation, mapping the semantic enhancement model to the hierarchical asset model includes:
[0050] Obtain the update timestamp of each original data table that constitutes each asset entity node in the most recent structural change log, and calculate the update frequency of each original data table accordingly.
[0051] When the variance of the update frequency of each original data table within the same asset entity node exceeds the preset tearing threshold, the asset entity node is split according to the update frequency to obtain real-time child nodes and regular child nodes.
[0052] Map the real-time child nodes to the business layer, map the regular child nodes to the standard layer, and establish a source layer reference pointer between the two;
[0053] In response to the update frequency variance being less than or equal to the preset tear threshold, the asset entity node is directly mapped to the business layer, standard layer, or original layer based on its average update frequency.
[0054] Secondly, this application provides a data asset management system for medical systems, including:
[0055] The acquisition module obtains data source metadata from the hospital's multi-business system, wherein the data source metadata includes: the connection configuration of the target business database, data tables, and field lists;
[0056] The synchronization module synchronizes the data source metadata in the multi-service system according to a preset acquisition strategy, and generates a set of structural metadata. The generation of the set of structural metadata also includes: extracting the structural definitions of each target business database in a consistent window manner, and generating a set of structural metadata based on the structural changes generated during the incremental data definition language log or trigger event supplement window.
[0057] The extraction module generates structural fingerprints and semantic fingerprints for each data table based on the structure metadata set, calculates the fusion similarity, performs clustering and merging under the adaptive threshold constraint of the business domain, generates asset entity nodes, and constructs an asset meta-model.
[0058] The enhancement module performs semantic parsing processing on the asset meta-model to generate a semantically enhanced model;
[0059] The mapping module maps the semantic enhancement model to a layered asset model, which includes: a raw layer, a standard layer, and a business layer.
[0060] The output module generates data access configurations based on the hierarchical asset model and outputs the corresponding data asset ledger.
[0061] Compared to existing technologies, this application achieves accurate, efficient, and timely synchronization of structural metadata for multiple hospital business systems through a pre-defined acquisition strategy and consistency window approach, avoiding structural inconsistencies or omissions that are easily caused by traditional full-scale scanning methods. Simultaneously, by fusing structural and semantic fingerprints and performing clustering, it can automatically identify and uniformly manage semantically similar or duplicated data tables, effectively reducing semantic differences and data redundancy among heterogeneous data assets and improving the automation and accuracy of data asset management. Furthermore, the establishment of a semantic enhancement model and a hierarchical asset model enables hospitals to flexibly adjust data governance levels and openness strategies based on data update frequency, sensitivity, and business importance, effectively improving the security and reusability of data assets, significantly shortening the data interface opening and management cycle, reducing maintenance costs, and meeting the hospital's actual needs for rapid data innovation and real-time data access. Attached Figure Description
[0062] Figure 1 A schematic diagram illustrating a data asset management method for a medical system provided in an embodiment of this application;
[0063] Figure 2 A flowchart illustrating a method for obtaining data source metadata from a hospital's multi-service system, as provided in this application embodiment;
[0064] Figure 3 A flowchart illustrating a method for extracting reference object information associated with each data table, provided in an embodiment of this application;
[0065] Figure 4 This is a schematic diagram of a data asset management system for a medical system provided in an embodiment of this application. Detailed Implementation
[0066] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0067] Reference Figure 1 The flowchart shown is a data asset management method for a medical system provided in an embodiment of this application, including:
[0068] S101: Obtain data source metadata of the hospital's multi-business system, wherein the data source metadata includes: connection configuration of the target business database, data tables and field list;
[0069] S102: According to the preset acquisition strategy, synchronize the data source metadata in the multi-service system to generate a set of structured metadata.
[0070] S103: The generated structure metadata set further includes: extracting the structure definition of each target business database in a consistent window manner, and generating a structure metadata set based on the structure changes generated during the incremental data definition language log or trigger event supplement window.
[0071] S104: Based on the aforementioned structural metadata set, generate structural fingerprints and semantic fingerprints for each data table, calculate the fusion similarity, and perform clustering and merging under the adaptive threshold constraint of the business domain to generate asset entity nodes and construct an asset meta-model.
[0072] S105: Perform semantic parsing processing on the asset meta-model to generate a semantically enhanced model;
[0073] S106: Map the semantic enhancement model to a layered asset model, wherein the layered asset model includes: a raw layer, a standard layer, and a business layer;
[0074] S107: Based on the hierarchical asset model, generate data access configuration and output the corresponding data asset ledger.
[0075] Regarding the above S101:
[0076] Data source metadata describes the structure and connection method of the source database, and includes at least:
[0077] Connection configuration: server address, port number, database type (e.g., MySQL, Oracle, SQL Server), access credentials;
[0078] List of data tables: table name, schema;
[0079] Field list: field name, data type, field length, whether nullable is allowed, primary key or foreign key identifier, etc.
[0080] The data source access module prioritizes batch retrieving table-field information through the database system's built-in data dictionary view; if the target business database does not have an open data dictionary, it then calls the JDBC interface DatabaseMetaData to obtain metadata, achieving zero-business-intrusion structure extraction.
[0081] For scenarios involving private databases from certain vendors or only offline backup files, this embodiment provides an "offline DDL script parsing" solution: first, the table creation script is exported periodically, and then the script parser extracts the table and field definitions, converts them into a JSON structure, and writes them into the metadata repository.
[0082] The data source access module supports batch registration and incremental registration. When going live for the first time, all business database connection attributes can be imported at once through the YAML configuration file. When the hospital adds new systems in the future, the connection information can be added in the management interface to automatically collect metadata.
[0083] In this way, the collected data source metadata is uniformly stored in the metadata repository (Meta-Store). The system assigns a unique identifier, source_id, to each record, which is referenced in subsequent structure synchronization, fingerprint calculation, and asset modeling stages. Through the above steps, efficient, comprehensive, and structured collection of metadata from heterogeneous data sources is achieved while ensuring database operational security.
[0084] Regarding S102 above:
[0085] Data Definition Language (DDL) refers to a subset of SQL used to create, modify, and delete schema objects in a database. Typical statements include CREATE, ALTER, DROP, TRUNCATE, RENAME, and COMMENT. DDL is used to define the structure of tables, views, indexes, stored procedures, etc., without directly manipulating the business data in the tables.
[0086] In practice, the platform performs structural synchronization of the data source metadata obtained by S101 according to the preset collection strategy, thereby generating a set of structural metadata. The collection strategy can be configured to either full synchronization or incremental synchronization, and supports both scheduled and real-time triggering methods.
[0087] It should be noted that full synchronization mode is typically used during initial deployment or cross-version rebuilds:
[0088] The platform sequentially traverses the system metadata views of each database according to the connection configuration of the data source; maps all the table structures obtained from the traversal to an internally unified format; writes the generated results to a temporary buffer, and after verification, writes them in batches to the structure metadata tables t_meta_table and t_meta_column to ensure the consistency of metadata.
[0089] Incremental synchronization mode is suitable for routine operation and maintenance scenarios, and the platform implements it based on a two-level triggering mechanism:
[0090] Scheduled triggering: For example, by default, an incremental check is performed once a day at 01:00 AM, comparing the last_ddl_time field in the system dictionary view with the platform's most recent synchronization time, and filtering out newly added or modified tables;
[0091] Real-time triggering: For example, when a DDL event (create, modify, delete table) is generated on the database side, the event is pushed to the message queue through a database trigger or binlog subscription, and the platform consumes it in real time and updates the corresponding table structure.
[0092] For incremental synchronization, a strategy of buffering first and then merging can be adopted. For example, the captured DDL events can be parsed as structural change records and written to the buffer table t_meta_change_log; periodic jobs can update the main structure table according to change_type (Add / Alter / Drop). If a field is deleted, a deletion mark is made in t_meta_column instead of physical deletion.
[0093] After the merge is complete, a new structure version number, schema_version, is generated and recorded in t_meta_source.
[0094] It is important to note that, in order to avoid the "structural tearing" problem caused by the asynchronous update time of cross-system structures, this embodiment generates a "snapshot identifier" before the synchronization task begins, such as Oracle SCN or MySQL GTID. All data sources read the system dictionary based on this snapshot identifier to ensure the consistency of the same batch of structural snapshots.
[0095] For example, the structure metadata set generated after synchronization includes three parts: table-level metadata: table name, system to which it belongs, estimated row count, etc.; field-level metadata: field name, data type, whether it is a primary key, sensitivity level; and change log: structure version number, change type, change time, and operator.
[0096] It is understandable that table names such as t_meta_table and t_meta_column are for illustrative purposes only and can be replaced with the actual database names.
[0097] Regarding the above S103:
[0098] S103 further improves the structural metadata based on S102, and adopts a consistent window extraction combined with incremental DDL supplementation method, which not only ensures the time consistency of different system structure snapshots, but also takes into account the timely update of real-time changes.
[0099] In practice, when the synchronization task starts, the platform first requests a globally consistent identifier from each target business database, such as Oracle's System Change Number (SCN) or MySQL's GTID starting point. Subsequently, each database reads the system dictionary view based on this identifier to obtain the table structure definitions "at the same time." By utilizing the consistency window, structural omissions or inconsistencies caused by online DDL operations in business systems can be avoided, thereby improving the integrity and reliability of the structure snapshot.
[0100] During the consistency window, if the business system performs operations such as creating, modifying, or deleting table fields, the corresponding DDL events will be captured by database triggers or binary logs and written to the message queue. Before closing the window, the synchronization task consumes the queue in real time, appending the captured structural changes to the buffer table `t_meta_change_log` and associating them with the window identifier. This supplementary recording method ensures that even if temporary DDL events occur during large-scale table structure reads, they can be included in the same batch of structure metadata, reducing the workload of manual error correction.
[0101] When the synchronization task is completed, the platform merges and deduplicates the "window extraction results" and "incremental DDL entry results," and generates a new structure version number, schema_version. If the same table has both modified and newly added fields within the window, the system will prioritize retaining the latest field definition to ensure the accuracy and continuity of the structure description.
[0102] By using a consistent window in conjunction with incremental DDL supplementation, compared with the traditional full scan solution of database-by-database and table-by-table, the time required for structure synchronization can be shortened and the risk of structure gaps caused by parallel DDL can be reduced. At the same time, it reduces the reliance of operation and maintenance personnel on manual correction of structure snapshots and improves the automation level of structure governance.
[0103] The merged structural metadata set is then written to the metadata repository, triggering the fingerprint generation and clustering operations in S104. At this point, S103 has completed the "consistent, complete, and timely" aggregation of structural information from multiple business systems, providing a data foundation for subsequent asset modeling.
[0104] Regarding S104 above:
[0105] In this implementation method, after the structural metadata aggregation is completed, the data tables need to be automatically grouped and an asset meta-model constructed. Step S104 specifically includes: generating structural fingerprints and semantic fingerprints; calculating fusion similarity; performing clustering and merging based on business domain adaptive thresholds; generating asset entity nodes and constructing the asset meta-model.
[0106] Specifically, the platform first generates a structural fingerprint for each data table in the structural metadata set. The structural fingerprint is a fixed-length vector composed of the following table structural features: total number of fields, proportion of each field's data type, number of primary keys and unique indexes, number of foreign key constraints, and number of large object fields (such as CLOB / BLOB). After normalization, the feature values are concatenated in a predetermined order to obtain a structural fingerprint that can distinguish the physical differences between the tables.
[0107] Subsequently, the platform generates semantic fingerprints for each data table. Specifically, the table name and field names are concatenated into short text, and the field annotation text is input into a domain-fine-tuned pre-trained language model (e.g., a BERT model based on the Transformer architecture) to extract semantic vectors. These semantic vectors are then concatenated and dimensionality-reduced to obtain semantic fingerprints with adjustable dimensions. The model structure and vector dimensions used in this embodiment are merely examples; those skilled in the art can choose other models and parameters with equivalent performance.
[0108] The platform calculates the similarity between the structural fingerprint and semantic fingerprint of any two data tables, using cosine distance for the former and cosine similarity for the latter; then, it follows the formula:
[0109] Fusion similarity Structural similarity Semantic similarity;
[0110] Perform weighted fusion, where the weights , The threshold is dynamically determined by the adaptive threshold of the business domain to which each table belongs, so as to take into account the importance of structure and semantics in different business scenarios.
[0111] Based on the above fusion similarity, the platform uses an improved K-Medoids clustering algorithm to divide the data table into clusters:
[0112] Initial cluster centers are selected based on the similarity matrix;
[0113] Iteratively adjust cluster members until similarity converges;
[0114] Fine-tuning or splitting elements within a cluster based on adaptive thresholds for the business domain;
[0115] Generate asset entity nodes for each cluster and set the cluster center table as the representative table.
[0116] Finally, an asset meta-model is constructed for each asset entity node, including: asset identifier, business domain attributes, list of source systems, list of component tables, fingerprint summary, adaptive threshold information, and construction time, etc., and the asset meta-model is written to the metadata repository. Through step S104, this embodiment realizes the automatic identification and merging of synonym tables and duplicate tables, improving the automation and accuracy of data asset modeling.
[0117] For example, in actual implementation, let the business domain be "inpatient medical orders," and there are three data tables: Table A (the main medical order table in the HIS system), Table B (the laboratory medical order record in the LIS system), and Table C (the medical order history table in the HIS system). For the above three data tables, the platform extracts the structural fingerprint and semantic fingerprint of each table in turn, calculates the structural similarity and semantic similarity between Table A and Table B, Table A and Table C, and Table B and Table C, and then calculates three sets of fused similarity.
[0118] For example, suppose the fusion similarity results are: Table A and Table C = 0.89, Table A and Table B = 0.57, and Table B and Table C = 0.55. Assuming the adaptive threshold for this business domain is set to 0.85, during the clustering process, the platform first selects Table A, which has the highest similarity, as the initial cluster center. Table C, whose similarity is greater than the threshold, is assigned to the same cluster. Table B, because its fusion similarity with the other two tables is less than 0.85, is placed in a separate cluster. After the clustering iteration converges, the platform assigns Table A and Table C to the same asset entity node, and Table B to another asset entity node, using Table A and Table B as the representative tables for their respective nodes.
[0119] Ultimately, the platform generates two asset entity nodes based on the clustering results, and uses the cluster center table of each node as the representative table for that node. Through the above implementation, the platform can automatically aggregate synonymous data tables across systems and data sources within the same business domain, providing fundamental support for the unified management and data governance of subsequent hierarchical asset models.
[0120] For example, in the actual clustering process, after obtaining the structural and semantic fingerprints of all data tables, the platform first selects an initial set of cluster centers based on the fusion similarity matrix. Taking a certain business domain as an example, assuming that tables A, B, C, and D participate in clustering, the platform initially uses table A as the first cluster center, and then calculates the fusion similarity between tables B, C, and D and table A in sequence. Tables with similarity not less than a preset business domain threshold are merged into the current cluster. For example, if the fusion similarity between table B and table A is 0.88, and between table C and table A is 0.83, both greater than the preset business domain threshold of 0.80, the platform initially assigns A, B, and C to the same cluster. The platform then calculates the mean or weighted mean of the structural and semantic fingerprints of all members in the cluster as new cluster center features, and recalculates the fusion similarity between all data tables and the new cluster center features, adjusting the cluster members according to the new results. If the fusion similarity between table D and the new cluster center is found to be 0.82, which is greater than the threshold, the platform includes D in the cluster. The cluster centers are then updated again, and the above process is repeated until the affiliation of all tables remains unchanged across both rounds of calculation, meaning the fusion similarity between all tables and cluster centers converges, at which point the clustering process terminates. This process ensures that the structural and semantic features of data tables within the same cluster are more similar, guaranteeing the cluster density of the final partitioning result.
[0121] For example, after clustering iteration convergence, the platform can further fine-tune merging or splitting operations on data tables within the same cluster according to the management requirements of the business domain. For instance, when the business domain has higher requirements for clustering granularity, the platform can set a local merging threshold or splitting threshold. For tables A, B, C, and D that have already been grouped into the same cluster, the platform can further calculate the pairwise fusion similarity between each member within the cluster. If it is found that the fusion similarity between table C and A, B, and D is less than the local threshold of 0.85, the platform can separate C to form a new cluster, achieving fine-grained splitting of the same asset entity node. Conversely, if the fusion similarity between the centers of two original small clusters is higher than the local merging threshold, the platform can merge the two small clusters into one asset entity node. The above fine-tuning merging and splitting operations can be flexibly adjusted according to the data governance specifications of different business domains in actual hospitals, further improving the business adaptability and usability of asset modeling results.
[0122] Through the above methods, the platform can obtain asset entity node partitioning results that reflect both the consistency of table structure and semantics and meet the management requirements of different business domains by dynamically adjusting thresholds and iteratively adjusting cluster members based on the comprehensive similarity of structural fingerprints and semantic fingerprints. This provides reliable data support for subsequent data governance, interface opening and asset lifecycle management.
[0123] Regarding the above S105:
[0124] In this embodiment, after automatically merging and constructing the asset meta-model, the platform further performs semantic parsing processing on each asset entity node to generate a semantically enhanced model. The goal of the semantic parsing processing is to automatically map each clustered asset entity node to a standardized semantic label in the medical business domain, facilitating subsequent asset stratification, data governance, and interface opening.
[0125] In its implementation, the platform first extracts the structural fingerprint, semantic fingerprint, and data distribution fingerprint for each asset entity node. Combining this with multi-dimensional information such as the node's source system, business domain attributes, and historical naming conventions, a complete node context feature vector is formed. Subsequently, the platform uses a pre-trained language model fine-tuned with a medical industry corpus to calculate the similarity between the node's context feature vector and standard concept vectors in a medical terminology ontology, obtaining a set of candidate semantic labels. For example, for a specific asset entity node, the platform can calculate the similarity between its fingerprint and ontology labels such as "Medical Order Master Table" and "Laboratory Request Form," selecting the label with the highest similarity exceeding a preset threshold as the node's standard semantic label.
[0126] It should be noted that if the similarity of all candidate tags is lower than the threshold set by the business domain, the platform will generate a mark for manual review for that node and include it in the manual review task pool. Data governance personnel will then conduct a secondary confirmation based on domain knowledge and actual business needs. The semantic tags corrected after manual review will be automatically included in the platform's ontology library, improving the accuracy of automatic matching for similar nodes in the future.
[0127] For example, the "pre-trained language model fine-tuned with medical industry corpus" in this embodiment can be built based on the Transformer-BERT framework. This model, based on a general large-scale pre-trained Chinese model, includes a 12-layer encoder, a hidden vector dimension of 768, and 12 self-attention heads. It then incorporates medical-specific corpus for two-stage fine-tuning. The first stage uses documents such as electronic medical records, test reports, medical orders, and medical journal articles as the training set, still employing a masked language modeling task to adapt the model to lexical context for three consecutive epochs. The second stage adds an entity-relation joint training task to the weights obtained in the first stage. Specifically, it combines medical entity recognition (diseases, symptoms, drugs, test items, etc.) with synonym discrimination (such as "ALT" and "alanine aminotransferase") into a multi-task loss function, which is then optimized through cross-entropy and contrastive learning. To improve the model's sensitivity to abbreviations of table and field names, short sentence pairs automatically generated by concatenating table and field names were added during the fine-tuning process. These were used as pseudo-sentence pair samples for training, thereby enhancing the model's ability to represent database naming conventions.
[0128] It should be noted that the above model layer number, hidden vector dimension, training steps and multi-task design are only exemplary implementations. Those skilled in the art can adopt deeper networks, different hidden vector dimensions or replace them with other types of language models (such as RoBERTa, Electra) according to computing power, corpus size or specific business needs, all of which fall within the protection scope of this invention.
[0129] In practical implementation, the platform's semantic parsing processing also supports compatibility with multiple versions of medical terminology ontology, including mainstream standards such as ICD-10, SNOMED-CT, and HL7, achieving semantic alignment across ontology and multiple versions for the same node. The platform can compare the label versions of historical interfaces with the latest semantic labels of existing assets, automatically recording the timestamps of label upgrades and changes, providing a foundation for subsequent data asset evolution management and interface traceability.
[0130] In this way, the platform can achieve automatic semantic normalization of asset entity nodes, business standard mapping, and version tracking, which greatly improves the semantic consistency of data assets among heterogeneous systems in the hospital and the ability to reuse data across systems, laying a semantic foundation for hierarchical asset modeling, intelligent permission management, and interface opening.
[0131] Regarding S106 and S107 above:
[0132] After completing the semantic enhancement model construction, the platform proceeds to map the semantic enhancement model to the layered asset model. The layered asset model, from bottom to top, includes the original layer, the standard layer, and the business layer, to meet the hierarchical requirements of traceability, governance, and application. The platform first reads the tags, sensitivity levels, update frequencies, and source system information of each asset entity node in the semantic enhancement model, and calculates the mapping priority of the nodes according to the mapping rules. The mapping rules use sensitivity level, business call frequency, and structural standardization as weighting factors to calculate a comprehensive mapping score; nodes with a comprehensive mapping score greater than the first threshold are mapped to the business layer, nodes with a score greater than the second threshold but not exceeding the first threshold are mapped to the standard layer, and the remaining nodes are retained in the original layer to ensure that highly sensitive data is fully governed while high-frequency data can be quickly made available.
[0133] For example, in the inpatient medical record business domain, a certain asset entity node has an update frequency of one record per minute, a sensitivity level of level two (including patient summary), and a structural standardization degree of 0.95. After comprehensive scoring, it exceeds the first threshold and is mapped to the business layer. Another historical examination record node has an update frequency of once a day, a sensitivity level of level three (including image file links), and a comprehensive score between the two thresholds, so it is mapped to the standard layer. After mapping, the platform generates source layer reference pointers and records semantic tag version numbers for nodes promoted to the business or standard layer, ensuring that any upper-level node can be traced back to the original structural definition and the first extraction snapshot.
[0134] After obtaining the hierarchical asset model, the platform generates data access configurations based on the node's hierarchy and sensitivity level. Specifically, it generates REST-based application interface descriptions for business layer nodes, read-only access definitions based on database views for standard layer nodes, and retains internal query scripts visible only to governance personnel for raw layer nodes.
[0135] The above interface description uses YAML format as an example, and includes asset identifier, interface path, parameter list, field anonymization strategy, and call frequency limit information. The platform also writes the interface version number and release time into the interface metadata table to facilitate subsequent version management and canary releases.
[0136] Simultaneously with the generation of the interface description, the platform outputs a data asset ledger according to a fixed template. This data asset ledger records at least the following fields: asset identifier, semantic tag, hierarchical level, sensitivity level, source system, latest structure version number, interface or view address, and last synchronization time. For business layer interfaces with automatic monitoring enabled, the data asset ledger also adds operational metrics such as the number of calls in the last seven days and average response latency for subsequent performance evaluation and capacity planning.
[0137] In this way, the hierarchical mapping and interface configuration of assets can be completed on the basis of unified semantics, which not only ensures that highly sensitive information is not transferred without authorization, but also enables high-value data to quickly support business innovation. At the same time, the output of the data asset ledger provides hospital management with a quantifiable and traceable view of data assets, which helps to continuously optimize data governance strategies.
[0138] As an optional implementation method, see [link to implementation details]. Figure 2 The flowchart below shows a method for obtaining data source metadata of a hospital's multi-service system, as provided in this application embodiment, including steps S201 to S203, wherein:
[0139] S201: Perform a dependency scan on the system directory of the target business database to extract reference object information associated with each data table, wherein the reference object information includes: views, stored procedures, and triggers;
[0140] S202: Record the referenced object information as dependency metadata;
[0141] S203: Merge the dependency metadata with the connection configuration, data table and field list to generate data source metadata.
[0142] In this implementation, the "system directory" refers to the metadata view or data dictionary built into the database management system (DBMS), used to record the names, definition scripts, and dependencies of objects such as tables, views, stored procedures, and triggers; "referenced object information" refers to the name of the accessed object that appears in the form of an SQL statement in the view definition, procedure body, or trigger body; and "dependency metadata" is a structured description of the lineage relationship between "referenced object and underlying data table", with recorded fields including referenced object name, object type, referenced table name, reference depth, and resolution timestamp.
[0143] During dependency scanning, the platform first reads the definition scripts for views, stored procedures, and triggers in the system directory. Using regular expressions combined with an abstract syntax tree parsing algorithm, the platform extracts object identifiers after key clauses such as FROM, JOIN, UPDATE, and MERGE in the scripts to determine the first-level direct reference relationship between the object and the data table. If the parsing result contains another view or procedure, the platform continues recursively into its definition script until it traces back to the lowest-level underlying data table. This process yields a complete multi-level reference chain.
[0144] To avoid duplicate recording of the same reference chain, the platform deduplicates the parsing results before writing dependency metadata: if the same "referenced object - referenced table" tuple already exists in the system directory, the platform only updates the most recent parsing timestamp; if it does not exist, a new dependency record is added, and the reference depth field is set to the current recursion level. This processing ensures that dependency metadata is compact and free of redundancy.
[0145] After the dependency metadata is generated, the platform merges this data with the connection configuration and the data table list using the same identifier source_id. On the one hand, statistical fields such as "number of times referenced by views" and "number of times called by procedures" are added to the data table entries to characterize the coupling degree of the data table in the business logic. On the other hand, if there is an unregistered basic table in the dependency chain, the platform triggers a supplementary recording mechanism to write the name, mode, and default field information of the table into the table list to ensure the complete closed loop of the lineage chain.
[0146] By introducing system directory dependency scanning and generating dependency metadata, the platform can grasp object-level lineage information during the data source acquisition phase. This information can provide a basis for incremental object detection during subsequent structural synchronization, and can also help identify logically closely related but structurally different views or processes during the asset modeling phase, thereby improving the accuracy and traceability of data asset merging.
[0147] As an optional implementation method, see [link to implementation details]. Figure 3 The flowchart below shows a method for extracting reference object information associated with each data table, as provided in an embodiment of this application, including steps S301 to S303, wherein:
[0148] S301: Query the system directory to obtain the dependency record between the target object and its first-level referenced object;
[0149] S302: For the first-level referenced object, continue to query the system directory recursively until all referenced objects are basic data tables, and generate multi-level nested dependency records;
[0150] S303: Deduplicate and topologically sort the multi-level nested dependency records to form a complete directed acyclic dependency chain, and write the dependency chain into the dependency metadata.
[0151] In its implementation, when extracting referenced object information, the platform not only records one level of dependency but also generates multi-level nested dependency records through recursive traversal. It then performs deduplication and topological sorting on the recorded results, ultimately obtaining a complete directed acyclic dependency chain. Here, "topological sorting" refers to linearly sorting all nodes in the graph while maintaining the pointing relationships of directed edges, ensuring that the starting point of any edge always precedes its ending point. "Directed acyclic dependency chain" refers to the set of paths in the dependency graph that do not contain cycles, providing a strict sequence for subsequent lineage visualization and influence analysis.
[0152] During the actual parsing process, the platform first queries the system directory to retrieve dependency records between the target object and the first-level referenced objects. For example, if the definition script of the view V_PRESCRIPTION contains FROM V_DRUGORDER, the platform records the first-level dependency relationship "V_PRESCRIPTION to V_DRUGORDER". Subsequently, the platform uses V_DRUGORDER as the new target object and continues to query its definition script, identifying table names such as FROM T_ORDER_MASTER and JOIN T_ORDER_DETAIL as the next-level referenced objects. The recursive process terminates when it encounters the underlying data table (i.e., an object of type TABLE in the system directory).
[0153] After the recursive query is completed, the platform aggregates the resulting object-to-object dependency tuples by depth, forming multi-level nested dependency records. If two different paths produce dependency tuples with identical content, the platform uses hash comparison to remove duplicates, avoiding duplicate records in the metadata. Taking V_PRESCRIPTION as an example, if another view V_DRUGORDER_HIST also references T_ORDER_MASTER, then "V_DRUGORDER_HIST to T_ORDER_MASTER" is considered a different record from the aforementioned tuple. However, if "V_PRESCRIPTION to T_ORDER_MASTER" already exists, it will not be written again.
[0154] To ensure the temporal consistency of dependency chains, the platform performs a topological sort on all dependency tuples related to the same target object after deduplication. The sorting priority is based on dependency depth from shallow to deep, and within the same depth, a stable sort is performed according to the lexicographical order of object names. The sorting result is converted into an index number and written to the `topo_rank` field in the dependency metadata table. Based on this field, downstream modules can directly traverse nodes in ascending order when displaying lineage graphs or performing influence range analysis, without needing to sort them again.
[0155] For example, after generating the dependency chain "V_PRESCRIPTION to V_DRUGORDER to T_ORDER_MASTER", the topology sorting results are as follows: Node 1 corresponds to V_PRESCRIPTION, Node 2 corresponds to V_DRUGORDER, and Node 3 corresponds to T_ORDER_MASTER, with corresponding totopo_ranks of 1, 2, and 3, respectively. If a new dependency "V_PRESCRIPTION to T_DRUG_PRICE" is subsequently added, the platform compares V_DRUGORDER and T_DRUG_PRICE, both with a depth of 2. Because the name comes first, T_DRUG_PRICE obtains a totopo_rank of 2, and the new chain is updated to 1, 2, 2, 3. This is presented as a branch in the visualization, ensuring that the link is acyclic while maintaining hierarchical accuracy.
[0156] Through the combined operations of recursive parsing, deduplication, and topological sorting described above, the platform can automatically generate complete directed acyclic dependency chains of overlay views, stored procedures, triggers, and underlying data tables, providing a verifiable, traceable, and hierarchical relationship foundation for subsequent data lineage tracing, change impact assessment, and hierarchical asset mapping.
[0157] As an optional implementation, forming a complete directed acyclic dependency chain includes:
[0158] Label each object in the dependency record with its business domain identifier and sensitivity level weight;
[0159] The dependency records are divided into domains according to the business domain identifier, and a subgraph is created for each business domain.
[0160] Within each subgraph, a weighted topological sort is performed first by hierarchy and then by weight to obtain the corresponding subgraph sorting result;
[0161] Based on the subgraph sorting result, a ternary layer sequence number is generated and written into the dependency metadata; wherein, the ternary layer sequence number includes: business domain number, topology layer number, and sensitivity level.
[0162] To further improve the usability of dependency chains in cross-business domain management and sensitive data protection scenarios, this implementation introduces business domain identifiers and sensitivity level weights on the basis of forming a complete directed acyclic dependency chain, and performs domain-based partitioning and weighted topological sorting of dependency records. The "business domain identifier" is used to distinguish functional modules within the hospital, such as outpatient EMR, inpatient EMR, laboratory, and imaging; the "sensitivity level weight" assigns weight values to objects containing fields such as personal identification information, medical information, and payment information according to data security level protection requirements, with higher weight values indicating more sensitive data.
[0163] Before writing dependency records, the platform first queries the business domain configuration table and marks each dependency record with a business domain identifier based on the system to which the object belongs, the table name prefix, or field characteristics. Simultaneously, it calculates the percentage of sensitive fields by referring to a sensitive field dictionary (such as name, ID number, medical record number), mapping the percentage range to discrete weights. For example, a percentage > 30% is assigned a weight of 3, between 10% and 30% is assigned a weight of 2, and less than 10% is assigned a weight of 1. After annotation, the dependency record is updated with the fields `biz_domain` and `sensitive_weight`.
[0164] Among them, "business domain identifier (biz_domain)" refers to the discrete code used to distinguish the functional domains within the hospital, such as outpatient 01, inpatient 02, laboratory 03, imaging 04, etc.
[0165] "Sensitivity weight" refers to the numerical weight assigned to objects containing sensitive fields such as personal identification information and medical information, based on data security level protection requirements. The higher the weight value, the more sensitive the data.
[0166] The platform then splits the dependency records into several subsets according to the business domain identifier, and creates a separate dependency subgraph for each business domain. Taking the "Laboratory Business Domain" as an example, this subgraph only contains laboratory-type views, stored procedures and their underlying tables; while the dependency chains of outpatient EMR and inpatient EMR are assigned to their respective subgraphs to avoid mixing of lineage graphs between different departments.
[0167] Within each subgraph, the platform performs a weighted topological sort, prioritizing hierarchy over weight. The sorting algorithm first ranks nodes based on their depth within the dependency chain (i.e., topological level), with shallower nodes appearing first. When two nodes are at the same level, their sensitivity levels are compared, with the heavier node appearing first. If the weights are still the same, the lexicographical order of the object names is used as the final comparison factor to ensure deterministic results. After sorting, a stable sequence is obtained under both hierarchy and weight dimensions, recorded in the `topo_rank` field.
[0168] The platform generates a ternary hierarchical number (D, L, W) for each dependency record based on the sorting result and writes it into the dependency metadata, where:
[0169] D is the business domain number, which can be automatically mapped by the business domain configuration table, such as inspection domain = 03 and image domain = 04.
[0170] L is the topological layer number, which is taken from the depth of the node in the subgraph;
[0171] W is the digital value for the sensitivity level, corresponding to sensitive_weight.
[0172] For example, the ternary layer number of the view V_LAB_RESULT in the inspection business domain can be (03,2,3), indicating that it is located in the inspection domain, the second layer, and the highest sensitivity level; the ternary layer number of the basic table T_LAB_PARAM in the same domain may be (03,3,1), indicating the third layer and the lowest sensitivity level.
[0173] By introducing a dual-dimensional sorting system of business domain identifiers and sensitivity level weights for the dependency chain, this implementation not only ensures the correctness of the topology but also provides an intuitively filterable and controllable "layer-sensitive" index at the metadata level. On one hand, the data governance module can quickly locate highly sensitive data links based on the ternary layer sequence number and automatically generate de-identification strategies; on the other hand, the interface opening module can use this to determine whether the caller has exceeded its authority to access cross-domain or highly sensitive assets, thereby achieving refined data security management and operation and maintenance auditing.
[0174] As an optional implementation, the generated structure metadata set includes:
[0175] Obtain the structure version identifier of each target business database and record it as the synchronization start marker;
[0176] Based on the synchronization start marker, open the same transaction consistency window, and extract the structure definition of each target business database within the consistency window to generate the first structure metadata subset;
[0177] After completing the extraction of the first structural metadata subset, based on the structural change logs or DDL trigger events set for each target business database, the incremental structural changes generated before the consistency window is closed are extracted to generate the second structural metadata subset.
[0178] The first and second subsets of structure metadata are merged and deduplicated to obtain the final set of structure metadata.
[0179] Furthermore, the generation of structural metadata employs a two-stage extraction strategy: "version number - consistency window - incremental addition." The structural version ID, referred to here, is a metadata timestamp that automatically increments after the database executes DDL operations, such as Oracle's System Change Number (SCN), MySQL's Global Transaction ID (GTID), or SQL Server's database version. The consistency window refers to a read-only transaction interval opened in each target business database based on the same version ID, used to ensure point-in-time consistency of multi-source structural snapshots; the "structural change log" includes data dictionary change tables, binlog DDL fragments, or event records written by DDL triggers.
[0180] The platform first reads the current structure version identifier from each target business database and takes the minimum value as the synchronization start flag. Taking Oracle as an example, if the SCNs of the three instances are 105000, 105008, and 105003 respectively, then the synchronization start flag is set to 105000. Subsequently, the platform opens read-only transactions in all target business databases with SCN=105000, thus establishing a consistency window.
[0181] Within the consistency window, the platform iterates through data dictionary views (such as DBA_TABLES, DBA_TAB_COLUMNS, or information_schema.columns), extracts table names, field names, data types, constraints, and other definitions, and writes them into the first structure metadata subset. During the extraction process, the platform only reads system pages and does not lock business tables, ensuring that online transactions are unaware of the process.
[0182] After the consistency window closes, the platform starts incremental monitoring. If the target business database has DDL triggers enabled, the triggers write the statement and timestamp to t_meta_change_log when performing CREATE / ALTER / DROP. If the target business database supports binary logs (binlog), the platform subscribes to QUERY_EVENT and filters out DDL statements. All records with timestamps earlier than the window closing time are written to the second structure metadata subset, forming the structure increment during the window period.
[0183] The platform then performs a full field comparison on the first and second structural metadata subsets using "object name + field name + schema" as the key: if the two subsets contain fields with the same name but different types, the new type from the second structural metadata subset is used to overwrite the existing fields; if a table missing from the first structural metadata subset appears in the second structural metadata subset, the platform adds it. After merging and deduplication, the final structural metadata set is obtained, and a new schema version id=105008 record is generated and added to t_meta_source for use in the next round of synchronization.
[0184] For example, in a single synchronization task, a total of 1,820 table structures and 52,970 fields were extracted from ten business databases; 12 DDL events were generated during the window period, and after incremental data entry, 3 new table structures and 48 new fields were added. Compared with the traditional serial fetching solution, this solution avoids the problem of "time point misalignment" in multi-source structures and compresses the structure snapshot update cycle.
[0185] In this way, without affecting online business, consistent structural snapshots and real-time incremental integration across databases are achieved, providing an accurate, complete, and synchronized structural metadata foundation for subsequent fingerprint generation, asset modeling, and lineage tracing.
[0186] As an optional implementation, the construction of the asset meta-model includes:
[0187] For each data table in the structure metadata set, a structure fingerprint and a semantic fingerprint are generated. The structure fingerprint is used to characterize the number of fields, primary key type, and field type distribution. The semantic fingerprint is based on a pre-trained language model to obtain semantic vectors of table name, field name, and field annotation.
[0188] For any two data tables, calculate the structural fingerprint distance and the semantic fingerprint cosine similarity, and obtain the comprehensive similarity score between the tables according to the preset weighting coefficient;
[0189] Based on the adaptive threshold of the business domain, clustering and merging are performed on data tables whose comprehensive similarity score is not lower than the adaptive threshold of the business domain to form asset entity nodes;
[0190] Establish a mapping relationship between the asset entity nodes and the corresponding original data tables, and write them into the asset meta-model.
[0191] In this optional implementation, the platform extracts two types of fingerprint features from each data table based on the final structural metadata set, and performs clustering and merging based on fingerprint similarity to generate an asset meta-model. The "structural fingerprint" describes the physical structural features of the data table, and is composed of normalized indicators such as the total number of fields, field type distribution, number of primary keys, number of foreign keys, and the count of large object fields, concatenated in a fixed order. The "semantic fingerprint," relying on a pre-trained language model, encodes the table name, field names, and field annotations into high-dimensional semantic vectors to characterize the business meaning of the data table. The dimension of the semantic fingerprint vector is not limited to the 256 or 768 dimensions in the example; those skilled in the art can adjust it according to computing power and accuracy requirements.
[0192] The platform calculates the structural fingerprint distance and semantic fingerprint cosine similarity for any two data tables: the structural fingerprint distance uses... Formal measurement measures structural differences; semantic fingerprint cosine similarity directly uses the cosine value. The platform presets weighting coefficients μ and ν, with μ+ν=1. μ is appropriately increased in business domains with highly standardized field types such as inspection and imaging, while ν is increased in text-based business domains such as medical records and nursing records, to obtain a comprehensive similarity score that is closer to the business context.
[0193] To improve clustering flexibility, the platform performs merging based on an adaptive threshold for the business domain. The adaptive threshold δ is automatically adjusted by the platform based on the historical clustering quality within the same domain. If the manual splitting rate was high in the previous round of clustering, the platform lowers δ; if the manual merging rate was high, δ is raised. Only when the combined similarity score of two tables is not lower than δ will the system classify them into the same asset entity node. This mechanism can maintain clustering accuracy while avoiding excessive fragmentation.
[0194] The platform employs an improved K-Medoids algorithm for clustering. First, using the comprehensive similarity matrix as input, a maximal density table is selected as the initial cluster centers. Then, the cluster members and centers are iteratively updated until the centers stabilize. After clustering, each cluster is defined as an asset entity node, recording its business domain, center table name, member table list, and adaptive threshold δ.
[0195] The platform then establishes asset mapping relationships. On the one hand, it writes a list of member tables into the metadata entries of asset entity nodes, forming a one-to-many mapping with the original table structure records. On the other hand, it adds a field `asset_id` and a membership score to each member table to represent its corresponding asset entity node and similarity weight. The asset meta-model stores fields including asset identifier, business domain, clustering time, central table, number of member tables, adaptive threshold δ, and average similarity.
[0196] For example, in the business domain of inspection, the structural fingerprint distance between tables LAB_ORDER_HIS and LAB_ORDER_CUR is 0.12, and the semantic fingerprint cosine similarity is 0.96; if μ=0.6 and ν=0.4, then the comprehensive similarity score is... If the adaptive threshold δ=0.85, the two tables meet the merging conditions and are grouped into the same asset entity node "Medical Orders"; the other table, LAB_PARAM, remains an independent node because its comprehensive similarity score of 0.72 is below the threshold. Through this process, the platform can automatically aggregate data tables with similar semantics or structures in heterogeneous systems without interfering with business data, providing a unified asset layer abstraction for subsequent data governance and interface encapsulation.
[0197] As an optional implementation, the construction of the asset meta-model further includes:
[0198] A predetermined number of record samples are extracted from the business database corresponding to each data table in the structure metadata set to obtain a field value sample set;
[0199] For each numeric field in the sample set of field values, calculate the minimum, maximum, mean, and variance;
[0200] For each categorical field in the field value sample set, a preset number of categories are selected in descending order of category frequency, and their frequencies are counted. The proportion of null values is also counted to generate a data distribution fingerprint.
[0201] Calculate the similarity of the data distribution fingerprint between any two data tables to obtain the data distribution fingerprint similarity;
[0202] When calculating the overall similarity score between any two data tables, the data distribution fingerprint similarity, structural fingerprint distance, and semantic fingerprint cosine similarity are weighted and fused according to a preset distribution weight to adjust the overall similarity score between the tables.
[0203] Adaptive threshold clustering and merging are performed based on the adjusted inter-table comprehensive similarity score to form the asset entity nodes.
[0204] In a further optional implementation, the platform introduces a "data distribution fingerprint" when generating the asset meta-model to compensate for the risk of mis-clustering that may arise from relying solely on structural and semantic fingerprints. Here, "data distribution fingerprint" is a quantitative description of the distribution characteristics of the actual field values in a data table; "numerical field" refers to a field whose data type is an integer, floating-point number, or date / time that can be converted to a timestamp; "categorical field" refers to a string or enumerated field with discrete values and a finite cardinality, such as department codes or test categories.
[0205] For example, the platform first samples records from 5% of the number of rows in each data table or a maximum of 10,000 rows (whichever is smaller), forming a field value sample set. For numerical fields, the platform calculates the minimum, maximum, mean, and variance within the sample set; for categorical fields, it selects the top-k categories (default k=5) from high to low frequency, records the frequency of each category, and calculates the proportion of null values in each field. All statistics are concatenated in field order to obtain the data distribution fingerprint vector. If a field is encrypted or anonymized and its original value cannot be read, the platform records a placeholder and skips it during similarity calculation.
[0206] Here, Top-k is a common information retrieval term, referring to selecting the top k results from all candidate results based on their similarity. In this embodiment, k=5 is defaulted, meaning that the 5 candidate semantic concepts with the highest similarity are returned.
[0207] The platform calculates the similarity of data distribution fingerprints for any two data tables. Numerical fields use standardized Euclidean distance, while categorical fields use the Jaccard similarity coefficient. The absolute difference is calculated separately for the proportion of null values. The multi-field similarities are averaged according to field weights to obtain the data distribution fingerprint similarity S_dist. Subsequently, the platform performs a three-factor fusion using the structural fingerprint distance D_struc, the semantic fingerprint cosine similarity S_sem, and the data distribution fingerprint similarity S_dist.
[0208] Adjusted inter-table comprehensive similarity score = γ×S_dist+α×(1-D_struc)+β×S_sem;
[0209] Where γ, α, and β are preset distribution weights, and γ + α + β = 1. For highly structured business domains such as laboratory tests and imaging with significant numerical differences, the platform increases γ; for text-dominated domains such as medical records and nursing records, β is increased.
[0210] After the adjusted inter-table comprehensive similarity score is calculated, the platform continues to use the aforementioned adaptive threshold mechanism, incorporating distribution fingerprint features into the clustering determination. When two tables are highly similar in structure and semantics but their numerical distributions deviate significantly, such as an adult test table and a pediatric test table, even if the comprehensive similarity score was originally slightly higher than the threshold δ, adding the data distribution fingerprint may lower it to below δ, thus preventing erroneous merging. Conversely, if the field names of two tables differ significantly but their data distributions are almost identical, such as a multilingual version table of the same system, the data distribution fingerprint can improve the comprehensive similarity score, facilitating correct merging. Here, δ is the business domain adaptive threshold, used to determine whether the adjusted inter-table comprehensive similarity score meets the clustering merging criteria.
[0211] For example, comparing the two tables LAB_RESULT_ADULT and LAB_RESULT_PED within the business domain: the structural fingerprint distance is 0.05, the semantic fingerprint cosine similarity is 0.91, and the data distribution fingerprint similarity is only 0.42 due to significant differences in the reference interval. Let γ=0.4, α=0.3, and β=0.3, then the adjusted inter-table comprehensive similarity score = 0.4×0.42+0.3×0.95+0.3×0.91≈0.726, which is below the threshold of 0.85. Based on this, the system assigns the two tables to different asset entity nodes. Conversely, the structural fingerprint distance between LAB_RESULT_ZH and LAB_RESULT_EN is 0.04, the semantic fingerprint cosine similarity is 0.70, and the data distribution fingerprint similarity is 0.92. The adjusted inter-table comprehensive similarity score is ≈0.866, which is above the threshold and is merged into the same asset entity node.
[0212] Among them, LAB_RESULT_ADULT: Adult Lab Result, recording laboratory results for inpatients or outpatients aged ≥18 years. LAB_RESULT_PED: Pediatric Lab Result, recording laboratory results for patients <18 years old. Due to differences in reference intervals and physiological ranges, there are significant differences in the numerical distribution between the two tables. LAB_RESULT_ZH: Chinese field naming version of the laboratory result table (…). Fields should use Pinyin or Chinese abbreviations. LAB_RESULT_EN: The English version of the test results table (…). The business content is consistent with the previous table, only the field naming language is different.
[0213] By introducing data distribution fingerprints and fusing them with three-factor weights, the platform effectively reduces the probability of false merging caused by differences in business scenarios while maintaining clustering accuracy, providing a more robust and detailed identification basis for data asset modeling across departments and populations in the medical system.
[0214] As an optional implementation, the generative semantic enhancement model includes:
[0215] For the asset entity node, its semantic fingerprint, structural fingerprint and data distribution fingerprint are extracted respectively. Based on the scenario weight determined by the business domain or field type, the semantic fingerprint, structural fingerprint and data distribution fingerprint are weighted and concatenated to generate a context representation vector for semantic disambiguation.
[0216] The context representation vector is matched with the concept vector in the medical terminology ontology to obtain a set of candidate semantic concepts;
[0217] From the set of candidate semantic concepts, candidate semantic concepts with a similarity of not less than the domain threshold are selected as the standard semantic labels for the asset entity nodes;
[0218] If the similarity of all candidate semantic concepts is lower than the domain threshold, a manual review marker is generated for the asset entity node.
[0219] The standard semantic tags or tags to be manually reviewed are written into the semantic enhancement model to complete the semantic enhancement of asset entity nodes.
[0220] In this optional implementation, the semantic enhancement model is generated based on four steps: "multi-fingerprint concatenation—ontology matching—threshold filtering—manual review." Here, the "context representation vector" is a high-dimensional vector obtained by concatenating the fingerprint features of asset entity nodes in the three dimensions of structure, semantics, and data distribution according to scenario weights; "scenario weights" refer to the proportional factors that weight the contribution of the three types of fingerprints based on the business domain or field type; and the "medical terminology ontology" refers to a multi-source knowledge graph that includes ICD-10, SNOMED-CT, HL7, and in-house self-built medical dictionaries, whose concept vectors are encoded using the same language model for comparison with node vectors in a unified vector space.
[0221] During implementation, the platform first extracts semantic fingerprints, structural fingerprints, and data distribution fingerprints for each asset entity node, and then calls the weight configuration table according to the business domain or field type to which the node belongs. For example, in the inspection business domain, the semantic fingerprint weight is configured as 0.5, the structural fingerprint as 0.3, and the data distribution fingerprint as 0.2; in the outpatient medical record domain, the semantic fingerprint weight is configured as 0.6, the structural fingerprint as 0.2, and the data distribution fingerprint as 0.2. The platform then linearly concatenates the three types of fingerprints according to these proportions and normalizes them to obtain the context representation vector of the node.
[0222] Subsequently, the platform performs cosine similarity matching between the context representation vector and the concept vectors in the medical terminology ontology, returning the top-k candidate semantic concepts with the highest similarity (default k=5). Taking the asset entity node "LAB_RESULT_ZH_EN" as an example, the set of candidate semantic concepts may include "test result table" and "laboratory report view," etc. In this implementation, the threshold G within the business domain is set to 0.85. The platform selects the candidate semantic concepts with a similarity of not less than G and the highest score as the standard semantic label for the node. Here, LAB_RESULT_ZH_EN is an exemplary asset entity node name, representing "test result table (mixed Chinese and English fields)"; it is only used to demonstrate the matching process, and the actual table name can be replaced according to the naming of each hospital's HIS or LIS.
[0223] If the similarity of all candidate semantic concepts is lower than G, the platform automatically generates a label for manual review for that node and pushes the node ID to the data governance workbench. Governance personnel can select or create new labels and submit them on the workbench. The platform will write the manual review results back to the medical terminology ontology and update the node records, enabling the model to learn and improve in subsequent matching.
[0224] Successfully matched standard semantic labels or tags awaiting manual review are written to the semantic enhancement model table t_asset_semantic, with fields including asset_id, std_label, label_source (AUTO / REVIEW), match_score, and timestamp. The semantic enhancement model is then used by the hierarchical mapping and permission management components to achieve unified semantic references for asset entity nodes across systems and levels.
[0225] Wherein, t_asset_semantic: the storage table name of the platform's semantic enhancement model, with the following example field meanings: asset_id is the unique identifier of the asset entity node; std_label is the standard semantic label obtained through matching or manually confirmed; label_source is the label source marker, AUTO indicates automatic matching, and REVIEW indicates manual review; match_score is the real value of similarity with the ontology concept; timestamp is the label writing or update time.
[0226] By combining multi-fingerprint context splicing with ontology matching, this implementation method introduces a traceable manual correction loop while maintaining automation. This not only improves the accuracy of tag matching but also ensures the continuous evolution of new business semantics, thereby significantly enhancing the understandability and reusability of heterogeneous data assets in the medical system under different scenarios.
[0227] As an optional implementation, mapping the semantic enhancement model to the hierarchical asset model includes:
[0228] Obtain the update timestamp of each original data table that constitutes each asset entity node in the most recent structural change log, and calculate the update frequency of each original data table accordingly.
[0229] When the variance of the update frequency of each original data table within the same asset entity node exceeds the preset tearing threshold, the asset entity node is split according to the update frequency to obtain real-time child nodes and regular child nodes.
[0230] Map the real-time child nodes to the business layer, map the regular child nodes to the standard layer, and establish a source layer reference pointer between the two;
[0231] In response to the update frequency variance being less than or equal to the preset tear threshold, the asset entity node is directly mapped to the business layer, standard layer, or original layer based on its average update frequency.
[0232] This optional implementation addresses the issue of excessively large differences in update speed among the original data tables within an asset entity node during the mapping of the semantic enhancement model to the hierarchical asset model. It proposes a splitting mechanism based on update frequency variance. The term "update frequency" refers to the average number of structural or data changes to the same table within a statistical period; "update frequency tearing" refers to a significant difference in update frequency among tables within the same node, exceeding the platform's set allowable threshold; and the "tear threshold" is a pre-configured upper limit in the platform's operation and maintenance strategy table, used to determine whether node splitting is necessary.
[0233] The platform first reads several recent structural change logs, recording the update timestamps of each original data table constituting a certain asset entity node before and after the most recent synchronization, and calculates the update frequency of each table within that period. Then, the platform calculates the variance of these frequency values to assess whether the update rhythm within the node is uniform.
[0234] When the calculated update frequency variance is higher than the tear threshold, the platform determines that the node has an update frequency tear and needs to be split.
[0235] For example, the splitting method is as follows: First, sort the member tables from high to low according to their update frequency. Then, assign the high-frequency tables, which account for approximately 80% of the total number of rows in the node, to the "real-time child nodes," and assign the remaining tables to the "regular child nodes." The real-time child nodes are directly mapped to the business layer, while the regular child nodes are mapped to the standard layer. A "source layer reference pointer" is generated synchronously between the two to record the backtracking rules of the real-time layer for the standard layer data.
[0236] If the update frequency variance does not exceed the tear threshold, the platform determines the target layer based on the average update frequency of the asset entity nodes: nodes with an average update frequency of more than ten times per hour are mapped to the business layer; those with an average update frequency between once and more than ten times per hour are mapped to the standard layer; and those with an average update frequency of less than once per hour remain in the original layer. This layering range can be flexibly adjusted by the operations and maintenance configuration.
[0237] For example, the asset entity node "ORDERS_ALL" contains three tables. One table is updated every minute, another approximately every fifteen minutes, and the last table is updated only once a day. The platform detects a significant difference in update frequency among the three tables, with the variance exceeding the tearing threshold. Therefore, the high-frequency table is separated into real-time child nodes mapped to the business layer, while the medium- and low-frequency tables are combined into regular child nodes mapped to the standard layer. Similarly, within the node "INVENTORY," the update frequencies of the tables are evenly distributed between two and three times per hour, with the variance less than the tearing threshold. The platform then maps it entirely to the standard layer without splitting it.
[0238] Here, ORDERS_ALL is an example asset entity node name, corresponding to a comprehensive data table set related to medical order operations in the Hospital Information System (HIS). Its member tables may include:
[0239] T_ORDER_RT: Real-time medical order table, updated frequently (approximately once per minute); T_ORDER_DAY: Daily summary medical order table, updated at a medium frequency (approximately once every fifteen minutes); T_ORDER_ARCH: Historical archived medical order table, updated infrequently (once a day).
[0240] INVENTORY is an example asset entity node name, corresponding to the inventory business domain of a pharmacy or consumables warehouse.
[0241] By updating the frequency variance and dynamically splitting the data, the platform can ensure the real-time availability of high-frequency data tables at the business layer, while avoiding low-frequency tables from excessively occupying real-time interface resources. This balances response speed and governance stability in the layered mapping of data assets.
[0242] Based on the same inventive concept, this application also provides a data asset management system for medical systems, which corresponds to the data asset management method for medical systems. Since the principle of the system in this application is similar to the data asset management method for medical systems described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0243] Reference Figure 4 The diagram shown is a schematic of a data asset management system for a medical system provided in an embodiment of this application. The system includes:
[0244] The acquisition module 10 acquires data source metadata of the hospital's multi-business system, wherein the data source metadata includes: connection configuration of the target business database, data tables and field list;
[0245] The synchronization module 20 synchronizes the data source metadata in the multi-service system according to the preset acquisition strategy and generates a set of structural metadata. The generation of the set of structural metadata also includes: extracting the structural definition of each target business database in a consistent window manner, and generating a set of structural metadata based on the structural changes generated during the incremental data definition language log or trigger event supplement window.
[0246] The extraction module 30 generates structural fingerprints and semantic fingerprints for each data table based on the structure metadata set, calculates the fusion similarity, performs clustering and merging under the adaptive threshold constraint of the business domain, generates asset entity nodes, and constructs an asset meta-model.
[0247] Enhancement module 40 performs semantic parsing processing on the asset meta-model to generate a semantically enhanced model;
[0248] The mapping module 50 maps the semantic enhancement model to a layered asset model, which includes: a raw layer, a standard layer, and a business layer.
[0249] The output module 60 generates data access configuration based on the hierarchical asset model and outputs the corresponding data asset ledger.
[0250] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
Claims
1. A data asset management method for medical systems, characterized in that, include: Obtain data source metadata from multiple hospital business systems, wherein the data source metadata includes: connection configuration, data tables, and field list of the target business database; According to the preset collection strategy, the data source metadata in the multi-service system is synchronized to generate a set of structural metadata; the generation of the set of structural metadata also includes: extracting the structural definition of each target business database in a consistent window manner, and generating a set of structural metadata based on the structural changes generated during the incremental data definition language log or trigger event supplement window. Based on the aforementioned structural metadata set, structural fingerprints and semantic fingerprints are generated for each data table, fusion similarity is calculated, and clustering and merging are performed under the adaptive threshold constraint of the business domain to generate asset entity nodes and construct an asset meta-model. For the aforementioned asset meta-model, semantic parsing processing is performed to generate a semantically enhanced model; The semantic enhancement model is mapped to a hierarchical asset model, which includes: a raw layer, a standard layer, and a business layer. Based on the hierarchical asset model, a data access configuration is generated, and the corresponding data asset ledger is output.
2. The data asset management method for a medical system according to claim 1, characterized in that, The data source metadata obtained from the hospital's multi-service system includes: Perform a dependency scan on the system directory of the target business database to extract reference object information associated with each data table, wherein the reference object information includes: views, stored procedures, and triggers; Record the referenced object information as dependency metadata; The dependency metadata is merged with the connection configuration, data table, and field list to generate data source metadata.
3. The data asset management method for a medical system according to claim 2, characterized in that, The extraction of reference object information associated with each data table includes: Query the system directory to obtain the dependency records between the target object and its first-level referenced objects; For the first-level referenced object, continue to query the system directory recursively until all referenced objects are basic data tables, generating multi-level nested dependency records; The multi-level nested dependency records are deduplicated and topologically sorted to form a complete directed acyclic dependency chain, and the dependency chain is written into the dependency relationship metadata.
4. The data asset management method for a medical system according to claim 3, characterized in that, The formation of a complete directed acyclic dependency chain includes: Label each object in the dependency record with its business domain identifier and sensitivity level weight; The dependency records are divided into domains according to the business domain identifier, and a subgraph is created for each business domain. Within each subgraph, a weighted topological sort is performed first by hierarchy and then by weight to obtain the corresponding subgraph sorting result; Based on the subgraph sorting result, a ternary layer sequence number is generated and written into the dependency metadata; wherein, the ternary layer sequence number includes: business domain number, topology layer number, and sensitivity level.
5. The data asset management method for a medical system according to claim 1, characterized in that, The generated structure metadata set includes: Obtain the structure version identifier of each target business database and record it as the synchronization start marker; Based on the synchronization start marker, open the same transaction consistency window, and extract the structure definition of each target business database within the consistency window to generate the first structure metadata subset; After completing the extraction of the first structural metadata subset, based on the structural change logs or data definition language trigger events set for each target business database, the incremental structural changes generated before the consistency window is closed are extracted to generate the second structural metadata subset. The first and second subsets of structure metadata are merged and deduplicated to obtain the final set of structure metadata.
6. The data asset management method for a medical system according to claim 5, characterized in that, The construction of the asset meta-model includes: For each data table in the structure metadata set, a structure fingerprint and a semantic fingerprint are generated. The structure fingerprint is used to characterize the number of fields, primary key type, and field type distribution. The semantic fingerprint is based on a pre-trained language model to obtain semantic vectors of table name, field name, and field annotation. For any two data tables, calculate the structural fingerprint distance and the semantic fingerprint cosine similarity, and obtain the comprehensive similarity score between the tables according to the preset weighting coefficient; Based on the adaptive threshold of the business domain, clustering and merging are performed on data tables whose comprehensive similarity score is not lower than the adaptive threshold of the business domain to form asset entity nodes; Establish a mapping relationship between the asset entity nodes and the corresponding original data tables, and write them into the asset meta-model.
7. The data asset management method for a medical system according to claim 6, characterized in that, The construction of the asset meta-model also includes: A predetermined number of record samples are extracted from the business database corresponding to each data table in the structure metadata set to obtain a field value sample set; For each numeric field in the sample set of field values, calculate the minimum, maximum, mean, and variance; For each categorical field in the field value sample set, a preset number of categories are selected in descending order of category frequency, and their frequencies are counted. The proportion of null values is also counted to generate a data distribution fingerprint. Calculate the similarity of the data distribution fingerprint between any two data tables to obtain the data distribution fingerprint similarity; When calculating the overall similarity score between any two data tables, the data distribution fingerprint similarity, structural fingerprint distance, and semantic fingerprint cosine similarity are weighted and fused according to a preset distribution weight to adjust the overall similarity score between the tables. Adaptive threshold clustering and merging are performed based on the adjusted inter-table comprehensive similarity score to form the asset entity nodes.
8. The data asset management method for a medical system according to claim 7, characterized in that, The generative semantic enhancement model includes: For the asset entity node, its semantic fingerprint, structural fingerprint and data distribution fingerprint are extracted respectively. Based on the scenario weight determined by the business domain or field type, the semantic fingerprint, structural fingerprint and data distribution fingerprint are weighted and concatenated to generate a context representation vector for semantic disambiguation. The context representation vector is matched with the concept vector in the medical terminology ontology to obtain a set of candidate semantic concepts; From the set of candidate semantic concepts, candidate semantic concepts with a similarity of not less than the domain threshold are selected as the standard semantic labels for the asset entity nodes; If the similarity of all candidate semantic concepts is lower than the domain threshold, a manual review marker is generated for the asset entity node. The standard semantic tags or tags to be manually reviewed are written into the semantic enhancement model to complete the semantic enhancement of asset entity nodes.
9. The data asset management method for a medical system according to claim 8, characterized in that, The step of mapping the semantic enhancement model to the hierarchical asset model includes: Obtain the update timestamp of each original data table that constitutes each asset entity node in the most recent structural change log, and calculate the update frequency of each original data table accordingly. When the variance of the update frequency of each original data table within the same asset entity node exceeds the preset tearing threshold, the asset entity node is split according to the update frequency to obtain real-time child nodes and regular child nodes. Map the real-time child nodes to the business layer, map the regular child nodes to the standard layer, and establish a source layer reference pointer between the two; In response to the update frequency variance being less than or equal to the preset tear threshold, the asset entity node is directly mapped to the business layer, standard layer, or original layer based on its average update frequency.
10. A data asset management system for a medical system, used to implement the data asset management method for a medical system according to any one of claims 1-9, characterized in that, include: The acquisition module obtains data source metadata from the hospital's multi-business system, wherein the data source metadata includes: the connection configuration of the target business database, data tables, and field lists; The synchronization module synchronizes the data source metadata in the multi-service system according to a preset acquisition strategy, and generates a set of structural metadata. The generation of the set of structural metadata also includes: extracting the structural definitions of each target business database in a consistent window manner, and generating a set of structural metadata based on the structural changes generated during the incremental data definition language log or trigger event supplement window. The extraction module generates structural fingerprints and semantic fingerprints for each data table based on the structure metadata set, calculates the fusion similarity, performs clustering and merging under the adaptive threshold constraint of the business domain, generates asset entity nodes, and constructs an asset meta-model. The enhancement module performs semantic parsing processing on the asset meta-model to generate a semantically enhanced model; The mapping module maps the semantic enhancement model to a layered asset model, which includes: a raw layer, a standard layer, and a business layer. The output module generates data access configurations based on the hierarchical asset model and outputs the corresponding data asset ledger.
Citation Information
Patent Citations
Hospital data analysis method, hospital data analysis platform, equipment and medium
CN111768850A
Data asset automatic generation and management system based on data operation technology
CN118941395A
Cross-domain data integration and fusion method based on large model, terminal and storage medium
CN119862531A
Digital management method and system for medical documents
CN120048463A
Data recommendation method and device based on space-time big data and medium
CN120632198A