A medical data resource intelligent inventory system

The intelligent inventory system for medical data resources has solved the problems of incomplete metadata collection and insufficient classification standards in medical data governance, and has achieved accurate collection and hierarchical classification of medical data, thereby improving the efficiency and accuracy of data governance.

CN122177388APending Publication Date: 2026-06-09JIANGSU CHUANGU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU CHUANGU TECHNOLOGY CO LTD
Filing Date
2026-03-06
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Current medical data governance suffers from problems such as incomplete metadata collection, insufficient processing of unstructured information, weak semantic understanding capabilities, vague classification standards, and excessive reliance on manual verification of results, which limit the efficiency, accuracy, and compliance management capabilities of data governance.

Method used

The intelligent inventory system for medical data resources is adopted, including a data access and acquisition module, an inventory analysis module, and a result confirmation module. Through connection testing, acquisition range detection, access permission assessment, semantic association analysis, data lineage analysis, and rule and intelligent recognition collaborative judgment, the system can achieve accurate acquisition, hierarchical classification, and traceable result confirmation of medical data.

Benefits of technology

It has achieved clear boundaries for medical data collection, accurate output of hierarchical classification results and traceability of the basis, and continuously improved the efficiency and accuracy of data governance by means of review, feedback and update technology to revise rules and models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177388A_ABST
    Figure CN122177388A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data governance technology, specifically to an intelligent inventory system for medical data resources. The system comprises: first, performing connection testing, collection scope detection, and access permission assessment on the target data source, and collecting medical metadata to form a metadata set; then, performing semantic association analysis and data lineage analysis on the metadata set, and generating classification results and corresponding evidence chains by combining data classification standards, data classification rules, and an intelligent recognition model; finally, outputting the data resource inventory results, and performing verification and feedback updates when inconsistencies are identified or when manual confirmation instructions are received. This invention can improve the completeness of metadata collection, classification accuracy, and result traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data governance technology, specifically to an intelligent inventory system for medical data resources. Background Technology

[0002] As healthcare informatization continues to deepen, medical institutions are gradually accumulating massive data assets scattered across various business systems. Existing methods largely rely on manual registration, general metadata management tools, or simple rule-based classification tools, which generally suffer from problems such as incomplete metadata collection, insufficient processing of unstructured information, weak semantic understanding capabilities, vague classification standards, and excessive reliance on manual verification of results. This limits the efficiency, accuracy, and compliance management capabilities of data governance. Summary of the Invention

[0003] This invention provides an intelligent inventory system for medical data resources, which addresses at least the issues of how to achieve accurate collection, reliable hierarchical classification, and traceable result confirmation of medical data in a multi-source heterogeneous data environment.

[0004] This invention provides an intelligent inventory system for medical data resources, the system comprising: The data access and acquisition module is used to perform connection tests, acquisition range detection, and access permission assessment on the target data source. Based on the metadata acquisition range obtained from the acquisition range detection, it acquires the medical metadata of the target data source and generates a metadata set and access permission assessment results. The inventory analysis module is used to perform semantic association analysis and data lineage analysis on the metadata set, and generate hierarchical classification results and corresponding evidence chains based on data classification standards, data classification rules and intelligent recognition models; The result confirmation module is used to output data resource inventory results based on the hierarchical classification results and the corresponding evidence chain. When the recognition results of the data hierarchical classification rules are inconsistent with the recognition results of the intelligent recognition model, or when a manual confirmation instruction is received, the hierarchical classification results are reviewed, and the data hierarchical classification rules and intelligent recognition model are updated according to the review results. The corresponding evidence chain is used to record the basis for hierarchical classification.

[0005] In one possible implementation, the data access acquisition module is used to compare the access permissions required for metadata acquisition with the current access permissions of the target data source, and generate an access permission assessment result. The access permission assessment result is used to indicate the difference between the access permissions required for metadata acquisition and the current access permissions of the target data source.

[0006] In one possible implementation, the data access acquisition module includes a connection testing unit, an acquisition range detection unit, an access permission evaluation unit, and a metadata acquisition unit. The connection testing unit is used to verify the connectivity status of the target data source and the availability status of the data interface. The acquisition range detection unit is used to determine the accessible data objects and corresponding metadata types in the target data source and generate the metadata acquisition range. The access permission evaluation unit is used to generate access permission evaluation results. The metadata acquisition unit is used to acquire medical metadata according to the metadata acquisition range and generate a metadata set.

[0007] In one possible implementation, the data access acquisition module is used to collect database table structure information, unstructured data document attribute information, medical image data tag information, data interface definition information, and file system file attribute information to generate a metadata set.

[0008] In one possible implementation, when medical image data label information is missing, the data access acquisition module is used to extract key image frames from the medical image data and perform image feature recognition on the key image frames to supplement anatomical location information and sequence type information. The anatomical location information and sequence type information are written into the metadata set.

[0009] In one possible implementation, the inventory analysis module is used to map field names, field descriptions, and data content descriptions in the metadata set to medical terms to generate semantic relationships; the inventory analysis module is also used to generate data lineage relationships based on data generation relationships, data transmission relationships, and data transformation relationships to achieve data lineage analysis.

[0010] In one possible implementation, the inventory analysis module is used to generate a corresponding evidence chain based on the rule hit information of semantic association, data lineage and data classification rules, and to determine the classification result based on the recognition result of the data classification rules and the recognition result of the intelligent recognition model.

[0011] In one possible implementation, when the identification results of the data classification rules are inconsistent with the identification results of the intelligent identification model, or when a manual confirmation instruction is received, the result confirmation module is used to generate a review task. The review task includes the target data resource identifier, the classification label to be reviewed, and the corresponding evidence chain.

[0012] In one possible implementation, the result confirmation module is used to receive the manual confirmation result corresponding to the review task and update the data classification rules based on the manual confirmation result; the result confirmation module is also used to generate labeled samples and update the intelligent recognition model based on the labeled samples.

[0013] In one possible implementation, the data resource inventory results include data source identifiers, data resource names, hierarchical classification labels, access permission status, and index information of the corresponding evidence chain.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By employing connection testing, collection range detection, access permission assessment, and multi-source metadata collection technologies, we have achieved clear boundary collection of medical metadata. Through semantic association analysis, data lineage analysis, and collaborative judgment technologies combining rules and intelligent recognition, we have achieved accurate output of hierarchical classification results and traceability of the basis. Through review, feedback, and update technologies, we have achieved continuous correction of rules and models. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the module composition of the system of the present invention; Figure 2 This is a schematic diagram of the module operation process of the present invention. Detailed Implementation

[0016] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0019] Intelligent inventory is not a static registration of data resources, but rather an automated identification and judgment process for data governance. Its core lies in the continuous processing of data access, metadata extraction, semantic association, lineage tracing, and classification judgment to uniformly identify the source, attributes, relationships, and management boundaries of data resources, resulting in verifiable and traceable inventory results. Its focus is not simply on summarizing the quantity of data, but on establishing a structured understanding and classification basis for data objects, providing a unified data foundation for subsequent data governance, catalog management, and rule enforcement. Based on this, this invention provides an intelligent inventory system for medical data resources, integrating the access, collection, analysis, judgment, and result confirmation of multi-source medical data.

[0020] like Figure 1 - Figure 2 As shown, a medical data resource intelligent inventory system includes: The data access and acquisition module is used to perform connection tests, acquisition range detection, and access permission assessment on the target data source. Based on the metadata acquisition range obtained from the acquisition range detection, it acquires the medical metadata of the target data source and generates a metadata set and access permission assessment results. In one embodiment, the data access and acquisition module is located at the entry point of the intelligent inventory system for medical data resources. It is used to confirm access to the target data source, determine the acquisition boundaries, and prepare for metadata storage. The data access and acquisition module first initiates a connection test based on the target data source's access protocol to confirm network connectivity, interface response, and basic authentication status. Then, it performs a acquisition range probe on the target data source, identifying accessible data objects and their corresponding metadata types to form a metadata acquisition range. After the acquisition range is defined, it performs an access permission assessment based on the current authorization status to obtain the access permission assessment result. Finally, it calls the corresponding acquisition adapter to read medical metadata according to the metadata acquisition range and organizes the acquired content into a unified metadata set. After this processing, the system can obtain stable, traceable, and clearly defined metadata input before entering the inventory analysis.

[0021] The data access and acquisition module is used to compare the access permissions required for metadata acquisition with the current access permissions of the target data source, and generate access permission assessment results. The access permission assessment results are used to indicate the difference between the access permissions required for metadata acquisition and the current access permissions of the target data source.

[0022] In one embodiment, the new limitation of access permission assessment is that it does not directly determine whether the target data source can be used for inventory based on whether the connection is successful. Instead, it compares the access permissions required for metadata collection with the current access permissions of the target data source item by item, and then outputs the access permission assessment result. After introducing this limitation, the criteria for determining access permission assessment are clearer. The data access and collection module can not only determine "whether a connection is possible", but also "which metadata can be collected under the existing authorization conditions, and which metadata cannot be collected temporarily due to insufficient permissions".

[0023] In practice, the data access and acquisition module pre-establishes a set of access permission templates corresponding to the metadata acquisition tasks. These templates include at least directory read permissions, object enumeration permissions, field description permissions, interface description read permissions, and file attribute read permissions. After the target data source completes the connection test, the data access and acquisition module extracts the current access permissions based on the information returned by the target data source. These current access permissions can be obtained by reading the system directory, calling the permission query interface, verifying the access token's permission range, or executing a read-only probe request. Subsequently, the data access and acquisition module compares the current access permissions with the access permission templates item by item, forming a permission difference list. This list records missing permission items, the scope of data objects corresponding to the missing permissions, the impact level of the missing permissions on the acquisition task, and suggested additional authorization types.

[0024] Access permission assessment results are compiled from a list of permission differences. These results can be stored in structured record form or accompanied by textual descriptions, indicating the differences between the access permissions required for metadata collection and the current access permissions of the target data source. When access permissions meet the collection requirements, the access permission assessment result can be marked as allowing continued full collection; when access permissions only meet partial collection requirements, the result can be marked as performing restricted collection at the current boundary, with the restricted portion included in subsequent prompts; when access permissions do not meet the most basic collection requirements, the result can be marked as pausing collection. This process ensures a direct correspondence between access permission assessment results and the subsequent scope of metadata collection, preventing collection failures, incomplete collection, or invalid requests caused by simply starting collection based on a successful connection.

[0025] The data access and acquisition module includes a connection testing unit, an acquisition range detection unit, an access permission evaluation unit, and a metadata acquisition unit. The connection testing unit is used to verify the connectivity status of the target data source and the availability status of the data interface. The acquisition range detection unit is used to determine the accessible data objects and corresponding metadata types in the target data source and generate the metadata acquisition range. The access permission evaluation unit is used to generate access permission evaluation results. The metadata acquisition unit is used to collect medical metadata according to the metadata acquisition range and generate a metadata set.

[0026] In one embodiment, the data access acquisition module is further implemented using a modular structure to establish a clear processing order and data transfer relationship between connection testing, acquisition range detection, access permission assessment, and metadata acquisition. This constraint increases the execution boundaries within the module, enabling those skilled in the art to deploy and implement it according to the unit's responsibilities.

[0027] Specifically, the data access and acquisition module includes a connection testing unit, an acquisition range detection unit, an access permission evaluation unit, and a metadata acquisition unit. The connection testing unit is responsible for receiving the target data source's address, port, access protocol, authentication information, and access type, initiating a connection request to the target data source, and verifying network connectivity and data interface availability. Network connectivity is used to confirm the target data source's reachability, and data interface availability is used to confirm whether the target data source can return basic query results according to the established protocol. After the connection test passes, the acquisition range detection unit initiates a detection operation based on the type of the target data source. For database-type target data sources, the acquisition range detection unit reads database-level, table-level, and field-level directory information; for file system-type target data sources, the acquisition range detection unit reads the directory structure and file list; for interface-type target data sources, the acquisition range detection unit reads the interface description, returned field examples, or interface metadata; for medical imaging target data sources, the acquisition range detection unit reads the image index and tag directory.

[0028] The data collection range detection unit determines the accessible data objects and their corresponding metadata types based on this information, and generates a metadata collection range. The access permission assessment unit receives the metadata collection range and performs a difference judgment based on the current access permissions, generating an access permission assessment result. The metadata collection unit then calls the appropriate collection method based on the metadata collection range and the access permission assessment result, performing metadata reading, field organization, and format normalization on the allowed data objects, ultimately generating a metadata set. Through this structure, the output of the previous unit becomes the input of the next unit, and there are no isolated actions between units, facilitating sequential control, logging, and exception rollback.

[0029] The data access and acquisition module is used to collect database table structure information, unstructured data document attribute information, medical image data tag information, data interface definition information, and file system file attribute information to generate a metadata set.

[0030] In one embodiment, the data access and acquisition module adds refined limitations to the metadata type during the metadata acquisition stage. Instead of simply collecting "data description information," it separately collects database table structure information, unstructured data document attribute information, medical image data tag information, data interface definition information, and file system file attribute information, and writes them uniformly into the metadata set. This limitation clarifies the boundaries of the metadata set, enabling it to directly support subsequent semantic association analysis and data lineage analysis.

[0031] In practical implementation, for database-type target data sources, the data access and acquisition module reads table names, field names, field types, length constraints, primary key relationships, foreign key relationships, index information, update time, etc., and organizes them into database table structure information. For unstructured data, the data access and acquisition module reads document names, document formats, storage paths, creation times, modification times, title information, summary information, keyword information, and directly obtainable document ownership information, and organizes them into unstructured data document attribute information. For medical imaging data, the data access and acquisition module prioritizes reading the tag items in the image header information, including examination identifiers, sequence identifiers, imaging modalities, examination sites, acquisition times, sequence descriptions, etc., and forms medical imaging data tag information. For data interfaces, the data access and acquisition module reads the interface address, request method, input parameter names, input parameter types, output field names, output field types, interface version, and call restriction information, and forms data interface definition information.

[0032] For file systems, the data acquisition module reads file names, directory locations, file sizes, file types, creation times, modification times, and access attributes to form file system attribute information. Before entering the metadata collection, the metadata from these different sources can be organized using a unified field template, retaining at least the source type, data object identifier, attribute name, attribute value, acquisition time, and acquisition status. This process preserves the differences between heterogeneous target data sources while maintaining a unified organizational structure, allowing for direct retrieval based on source type and attribute dimensions during subsequent processing without requiring extensive format conversions.

[0033] In the event of missing label information in medical image data, the data access and acquisition module is used to extract key image frames from the medical image data and perform image feature recognition on the key image frames to supplement anatomical location information and sequence type information. The anatomical location information and sequence type information are written into the metadata set.

[0034] In one embodiment, when medical image data label information is missing, the data access and acquisition module adds an image supplementation and recognition path to extract key image frames from the medical image data and supplement anatomical location information and sequence type information through image feature recognition. This limitation targets scenarios where medical image labels are incomplete, historical images are exported improperly, or labels are lost after cross-system migration. The aim is to ensure that medical image-related metadata retains basic usability before entering the metadata set, rather than altering the content of the image itself.

[0035] In practice, the data access and acquisition module first performs image slice reading or sequence reading on medical image data with missing labels, selecting key image frames based on image sequence length, time order, or slice position. Key image frames can be extracted preferentially from the first, middle, and last segments, or extracted at fixed intervals to ensure that the extracted images reflect the main structural features of the current sequence. After extraction, the data access and acquisition module performs basic preprocessing on the key image frames, including size unification, grayscale normalization, and invalid boundary cropping. The preprocessed key image frames are then input into the image feature recognition process. Image feature recognition can employ a lightweight convolutional neural network or a combination of preset image template matching and simple classification rules. The output of the recognition process is anatomical location information and sequence type information. The anatomical location information indicates the main human examination site corresponding to the image, and the sequence type information indicates the scanning sequence or image sequence category to which the image belongs.

[0036] To avoid erroneous supplementation, the data access and acquisition module can set a confidence level for the recognition results. When the recognition result reaches the preset confidence level, the anatomical location information and sequence type information are written into the metadata set, and the information is marked as originating from image supplementation recognition in the corresponding record. When the recognition result does not reach the preset confidence level, the data access and acquisition module can retain the missing state and record a confirmation mark in the metadata set. After this processing, even if the original medical image data label information is incomplete, the metadata set can still supplement it to form the basic descriptive information required for subsequent analysis, and the source of the supplementation is clear and will not be confused with the original label records.

[0037] The inventory analysis module is used to perform semantic association analysis and data lineage analysis on the metadata set, and generate hierarchical classification results and corresponding evidence chains based on data classification standards, data classification rules and intelligent recognition models; In one embodiment, the inventory analysis module, located after the data access and acquisition module, is used for unified analysis and classification of the metadata set. The inventory analysis module first performs semantic association analysis on the metadata set to establish the correspondence between data objects, field attributes, and their medical business meanings. Then, it performs data lineage analysis to identify upstream and downstream relationships of data objects during generation, transmission, and transformation. After the semantic association analysis and data lineage analysis results are generated, the inventory analysis module calls pre-set data classification standards, data classification rules, and intelligent recognition models to classify the metadata set, output classification results, and simultaneously generate an evidence chain corresponding to the classification process. The evidence chain records the basis for the classification judgment, facilitating subsequent result confirmation, audit traceability, and rule adjustment.

[0038] The inventory analysis module is used to map medical terms to field names, field descriptions, and data content descriptions in the metadata set, generating semantic relationships. The inventory analysis module is also used to generate data lineage relationships based on data generation relationships, data transmission relationships, and data transformation relationships, so as to realize data lineage analysis.

[0039] In one embodiment, the inventory analysis module further limits its analysis of the metadata set to two parallel processing paths: semantic association analysis and data lineage analysis. This limitation clarifies the input, intermediate relationships, and output paths of the inventory analysis, enabling the module to move beyond superficial matching of field names and generate structured association results that can be used for classification and determination.

[0040] In its implementation, the inventory analysis module first reads the field names, field descriptions, and data content descriptions from the metadata set. Field names provide technical naming information, field descriptions provide business explanations, and data content descriptions provide semantic clues about sample values, value ranges, document snippets, or fields returned by interfaces. The inventory analysis module then performs medical terminology mapping on the above content. Medical terminology mapping can be based on a pre-built medical terminology dictionary, a medical thesaurus, and a standard terminology encoding table. For field names containing abbreviations, aliases, or historical names, the inventory analysis module first performs standardization processing and then maps the standardized results to the corresponding disease, symptom, examination item, drug item, patient identity information type, or research identifier type. For field descriptions and data content descriptions containing natural language text, the inventory analysis module extracts medical entity words, modifiers, and context words, and cross-validates the extraction results with the field name mapping results to generate semantic associations.

[0041] Semantic relationships are used to characterize the correspondence between current data objects and medical business entities. For example, a field might correspond to a patient's identity, a field to a test result, or a document fragment to a clinical diagnosis and treatment scenario. In parallel with semantic relationship analysis, the inventory analysis module also performs data lineage analysis on the metadata set. Data lineage analysis is based on the source and flow relationships between data objects, focusing on identifying three types of relationships: data generation relationships, data transfer relationships, and data transformation relationships. Data generation relationships indicate which source data object generated the current data object; data transfer relationships indicate the transfer path of the current data object between systems; and data transformation relationships indicate the derivation relationships formed during the cleaning, mapping, summarization, or structural adjustment of the current data object. The inventory analysis module can identify these relationships based on interface call records, data synchronization configurations, field mapping configurations, inter-table relationships, and task execution records, and organize the identification results into data lineage relationships. Once the data lineage relationships are formed, the inventory analysis module can determine whether a data object originates from highly sensitive source data, whether it has undergone desensitization transformation, or whether it has been reused by multiple business systems, thus providing upstream and downstream basis for subsequent classification and judgment. Through the above processing, semantic association is used to express what the data "represents", and data lineage is used to express where the data "comes from, what processing it has undergone, and where it flows to". Together, they constitute the basic analysis results for hierarchical classification judgment.

[0042] The inventory analysis module is used to generate corresponding evidence chains based on the rule-hitting information of semantic association, data lineage, and data classification rules, and to determine the classification results based on the recognition results of the data classification rules and the recognition results of the intelligent recognition model.

[0043] In one embodiment, after completing semantic association analysis and data lineage analysis, the inventory analysis module further generates a corresponding evidence chain based on the rule-hitting information of semantic association, data lineage, and data classification rules, and determines the classification result accordingly. This constraint extends the classification judgment process from single rule matching to a closed-loop process of "rule judgment plus model judgment plus evidence retention," enabling the classification result not only to be output but also to explain its formation.

[0044] In practice, the inventory analysis module first performs rule matching on the metadata set according to data classification and grading rules. These rules can include sensitive information identification rules, clinical business importance rules, research sharing condition rules, and compliance restriction rules. After rule matching is complete, the inventory analysis module generates rule hit information. This information records at least the rule identifier, the hit field or data object, the hit keywords or structural features, the classification suggestion provided by the rule, and the rule priority. Subsequently, the module integrates semantic relationships, data lineage, and rule hit information to generate a corresponding evidence chain. This evidence chain can be organized according to the data object dimension, sequentially recording the terminology mapping source, upstream source path, key transformation nodes, hit rule, model output results, and reserved space for manual confirmation records. After assembling the evidence chain, the module obtains recognition results based on data classification and grading rules and recognition results based on the intelligent recognition model.

[0045] The intelligent recognition model is used to supplement scenarios where rule coverage is insufficient, focusing on identifying complex semantic combinations, cross-field correlation features, and multi-source joint features that are difficult for rules to fully cover. If the two types of recognition results are consistent, the inventory analysis module directly uses the consistent result as the hierarchical classification result; if the two types of recognition results are inconsistent, the inventory analysis module determines the result according to a preset adjudication order. The adjudication order can prioritize whether strong constraint rules are hit, then consider the completeness of the corresponding evidence chain, and then consider whether there is a highly sensitive source propagation relationship in the data lineage. After adjudication, the inventory analysis module outputs the final hierarchical classification result and binds the final result with the corresponding evidence chain for storage. After this processing, each hierarchical classification result can be traced back to terminology mapping, lineage path, rule hit, and model judgment. When performing manual review, rule revision, or model update, existing evidence can be directly called without re-executing the full analysis.

[0046] The result confirmation module is used to output data resource inventory results based on the hierarchical classification results and the corresponding evidence chain. When the recognition results of the data hierarchical classification rules are inconsistent with the recognition results of the intelligent recognition model, or when a manual confirmation instruction is received, the hierarchical classification results are reviewed, and the data hierarchical classification rules and intelligent recognition model are updated according to the review results. The corresponding evidence chain is used to record the basis for hierarchical classification.

[0047] In one embodiment, the result confirmation module is positioned after the inventory analysis module. It receives the hierarchical classification results and corresponding evidence chains, forming a directly usable data resource inventory result. The result confirmation module first reads the classification labels, judgment status, and object identifiers from the hierarchical classification results. Then, it reads the rule-hitting basis, semantic association basis, and data lineage basis from the corresponding evidence chain. After associating and organizing these two, it outputs the data resource inventory result. When the identification result of the data hierarchical classification rules is inconsistent with the identification result of the intelligent recognition model, or when an external manual confirmation instruction is issued, the result confirmation module transfers the current hierarchical classification result to the review process. After the review is completed, it adjusts the data hierarchical classification rules and the intelligent recognition model based on the manual confirmation result. The corresponding evidence chain is used in this process to record the hierarchical classification basis and to support review processing, result traceability, and subsequent updates.

[0048] In cases where the identification results of the data classification rules are inconsistent with the identification results of the intelligent identification model, or when a manual confirmation instruction is received, the result confirmation module is used to generate a review task. The review task includes the target data resource identifier, the classification label to be reviewed, and the corresponding evidence chain.

[0049] In one embodiment, the result confirmation module adds a review task generation path when entering the review process. This path transforms inconsistent results or manually specified results generated during the automatic identification phase into executable manual review objects. This limitation primarily addresses the issues of automatic identification results being generated but lacking a manual confirmation entry point, unclear review object boundaries, and scattered review criteria.

[0050] In practice, the result confirmation module continuously receives the rule recognition results and model recognition results output by the inventory analysis module and performs consistency checks on the two types of recognition results. Consistency checks can be performed sequentially using three methods: label value comparison, classification level comparison, and sensitivity level comparison. If the label values ​​and classification levels are consistent, the result confirmation module can directly proceed to the result output process; if the label values ​​are inconsistent, or the label values ​​are consistent but the classification levels differ, the result confirmation module marks the current data object as pending review. In addition to automatic triggering, the system management terminal, data governance personnel, or auditing processes can also send manual confirmation instructions to the result confirmation module. These manual confirmation instructions specify a particular data object, a specific data source range, or a batch of inventory results to enter the review process. Upon receiving either of these trigger signals, the result confirmation module generates a review task.

[0051] The review task includes at least a target data resource identifier, a classification label to be reviewed, and a corresponding chain of evidence. The target data resource identifier uniquely identifies the data object currently undergoing review, preventing object confusion during manual confirmation. The classification label records the classification candidates given in the current automatic identification phase. The corresponding chain of evidence demonstrates the basis for the classification result to the reviewers. To facilitate execution, the result confirmation module can also add generation time, trigger source, review status, and task priority to the review task. After generation, the result confirmation module writes the review task to the review task queue and marks the current data object's status in the data resource inventory results as pending confirmation. The review task queue can be grouped by data source, business type, or sensitivity level for batch processing during manual confirmation. This ensures the review process has clear triggering conditions, clear task carriers, and clear evidence input, avoiding reliance on verbal explanations or scattered records for manual confirmation and ensuring that subsequent confirmation results are accurately written back to the original inventory object.

[0052] The result confirmation module is used to receive the manual confirmation results corresponding to the review task and update the data classification rules based on the manual confirmation results; the result confirmation module is also used to generate labeled samples and update the intelligent recognition model based on the labeled samples.

[0053] In one embodiment, after the review task is generated, the result confirmation module further provides a linked processing path for receiving manual confirmation results and updating rules and models. This limitation mainly addresses the problem that after manual confirmation, only the current result is modified without feeding back into the automatic recognition mechanism, enabling the manual confirmation results to be continuously used to correct rules and supplement training samples.

[0054] In practice, the result confirmation module reads completed review tasks from the review task queue and receives the corresponding manual confirmation results. The manual confirmation results can include the final classification label, the confirmer's identifier, the confirmation time, and necessary confirmation instructions. The result confirmation module first binds and verifies the manual confirmation results with the target data resource identifier, confirming that the returned result is consistent with the original review task. After successful verification, the result confirmation module replaces the original classification label generated in the automatic identification stage with the manual confirmation result and synchronously updates the data resource inventory results of the current data object. Then, the result confirmation module updates the data classification rules based on the manual confirmation results. The update method can be adjusting the matching conditions of existing rules, adding new keyword matching items, correcting classification priorities, or writing repeatedly occurring judgment criteria from the manual confirmation into supplementary conditions in the rule base.

[0055] For confirmation results that are clearly exceptional, the result confirmation module can simply record them as manually exceptional samples without directly modifying the general rules. After completing the rule update, the result confirmation module continues to generate labeled samples. Labeled samples consist of metadata records of the target data resource, the final classification labels after manual confirmation, and the corresponding evidence chain summary used during review. The purpose of generating labeled samples is to solidify the reliable results after manual confirmation into data that can be directly used for subsequent model training. The result confirmation module writes the labeled samples to the sample storage area and sends them to the intelligent recognition model update process in predetermined batches. Model updates can use incremental training, eliminating the need to retrain all samples each time. After the update is completed, the result confirmation module records the number of samples involved in this update, the update time, and the updated model version to track the source of model changes later. Through the above processing, the manual confirmation results are no longer just a one-time manual correction, but can simultaneously affect both the rule side and the model side, allowing the automatic recognition capability to gradually approach actual business judgments as reviews accumulate.

[0056] The data resource inventory results include data source identifiers, data resource names, hierarchical classification labels, access permission status, and index information of the corresponding evidence chain.

[0057] In one embodiment, the data resource inventory results are output in a structured result record format, explicitly limited to including data source identifier, data resource name, hierarchical classification label, access permission status, and index information of the corresponding evidence chain. This limitation primarily addresses the issues of inconsistent output content of inventory results, difficulties in subsequent queries, and the difficulty in quickly locating the evidence chain.

[0058] In practical implementation, after automatic confirmation or manual review, the result confirmation module generates a data resource inventory result record for each data object. The data source identifier indicates the target data source from which the current data object originates, facilitating subsequent aggregation by system, department, or business scope. The data resource name indicates the name of the currently inventoried object, which can be a table name, field name, document name, image sequence name, interface name, or file name. The hierarchical classification label records the final confirmed classification result and is a core business field in the inventory result. The access permission status reflects the authorization status of the current target data source during the metadata collection phase. It can directly reference the status value from the access permission assessment result or be organized into explicit statuses such as allowed collection, partially allowed collection, or temporarily disallowed collection. The index information corresponding to the evidence chain establishes the correspondence between the inventory result and the storage location of the evidence chain. The result confirmation module does not need to repeatedly store all evidence content in each inventory result; instead, it uses the index information to point to the evidence chain record, thereby reducing redundant storage and improving query efficiency. The index information can use the evidence chain number, storage address identifier, or evidence record key value.

[0059] During output, the result confirmation module can write the above fields into the inventory result database in a unified format, or generate a catalog list, query list, or export file according to a preset template. After this processing, any data resource inventory result will simultaneously have source information, object information, classification information, permission information, and evidence location information. Subsequent catalog display, manual spot checks, compliance audits, or re-verification can all be carried out directly based on the unified result structure without the need to reorganize fields or re-retrieve classification criteria.

[0060] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0061] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A medical data resource intelligent inventory system, characterized in that, The system includes: The data access and acquisition module is used to perform connection tests, acquisition range detection, and access permission assessment on the target data source. Based on the metadata acquisition range obtained from the acquisition range detection, it acquires the medical metadata of the target data source and generates a metadata set and access permission assessment results. The inventory analysis module is used to perform semantic association analysis and data lineage analysis on the metadata set, and generate hierarchical classification results and corresponding evidence chains based on data classification standards, data classification rules and intelligent recognition models; The result confirmation module is used to output data resource inventory results based on the hierarchical classification results and the corresponding evidence chain. When the recognition results of the data hierarchical classification rules are inconsistent with the recognition results of the intelligent recognition model or when a manual confirmation instruction is received, the module reviews the hierarchical classification results and updates the data hierarchical classification rules and the intelligent recognition model according to the review results. The corresponding evidence chain is used to record the basis for hierarchical classification.

2. The system according to claim 1, characterized in that, The data access acquisition module is used to compare the access permissions required for metadata acquisition with the current access permissions of the target data source, and generate the access permission assessment result. The access permission assessment result is used to indicate the difference between the access permissions required for metadata acquisition and the current access permissions of the target data source.

3. The system according to claim 2, characterized in that, The data access and acquisition module includes a connection testing unit, an acquisition range detection unit, an access permission evaluation unit, and a metadata acquisition unit; The connectivity test unit is used to verify the connectivity status of the target data source and the availability of the data interface; The collection range detection unit is used to determine the accessible data objects and corresponding metadata types in the target data source, and to generate the metadata collection range; The access permission evaluation unit is used to generate the access permission evaluation result; The metadata collection unit is used to collect medical metadata according to the metadata collection scope and generate the metadata set.

4. The system according to claim 1, characterized in that, The data access and acquisition module is used to collect database table structure information, unstructured data document attribute information, medical image data tag information, data interface definition information, and file system file attribute information to generate the metadata set.

5. The system according to claim 4, characterized in that, In the event of missing label information in medical image data, the data access and acquisition module is used to extract key image frames from the medical image data and perform image feature recognition on the key image frames to supplement anatomical location information and sequence type information. The anatomical location information and the sequence type information are written into the metadata set.

6. The system according to claim 1, characterized in that, The inventory analysis module is used to map medical terms to the field names, field descriptions, and data content descriptions in the metadata collection, generating semantic relationships. The inventory analysis module is also used to generate data lineage relationships based on data generation relationships, data transmission relationships, and data transformation relationships, so as to realize the data lineage analysis.

7. The system according to claim 6, characterized in that, The inventory analysis module is used to generate the corresponding evidence chain based on the rule hit information of semantic association, data lineage and data classification rules, and to determine the classification result based on the recognition result of the data classification rules and the recognition result of the intelligent recognition model.

8. The system according to claim 1, characterized in that, If the identification result of the data classification rule is inconsistent with the identification result of the intelligent identification model, or if a manual confirmation instruction is received, the result confirmation module is used to generate a review task. The review task includes the target data resource identifier, the classification label to be reviewed, and the corresponding evidence chain.

9. The system according to claim 8, characterized in that, The result confirmation module is used to receive the manual confirmation results corresponding to the review task, and update the data classification rules based on the manual confirmation results; The result confirmation module is also used to generate labeled samples and update the intelligent recognition model based on the labeled samples.

10. The system according to claim 1, characterized in that, The data resource inventory results include data source identifier, data resource name, hierarchical classification label, access permission status, and index information of the corresponding evidence chain.