Clinical data asset management method and system based on data intelligent anonymization

By building a unified data asset management platform and intelligent anonymization algorithms, the problems of scattered storage and insufficient privacy protection of medical data have been solved, enabling efficient management and secure sharing of medical data and enhancing the value and application potential of the data.

CN121071916BActive Publication Date: 2026-03-27XUANWU HOSPITAL OF CAPITAL UNIV OF MEDICAL SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Medical data is scattered across different systems, lacking a unified management mechanism. Anonymization technology is not intelligent enough, privacy protection is incomplete, data quality management is insufficient, there is a lack of asset management thinking, cross-institutional collaboration is insufficient, existing technologies struggle to strike a balance between privacy protection and data availability, and security audit mechanisms are inadequate.

Method used

We will build a unified data asset management platform, adopt a machine learning-based intelligent anonymization algorithm, adaptively adjust the anonymization strategy, combine it with a multi-layered privacy protection system, provide a scientific privacy risk assessment method, improve data quality and security audit, establish a data asset management model, and support multi-center clinical research and cross-institutional collaboration.

Benefits of technology

It enables integrated management and comprehensive analysis of medical data, improves privacy protection and data availability, enhances data quality and security, and promotes the widespread application and value mining of medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071916B_ABST
    Figure CN121071916B_ABST
Patent Text Reader

Abstract

The application provides a clinical data asset management method and system based on data intelligent anonymization, and the method comprises the following steps: collecting, cleaning and integrating medical clinical data from multiple source heterogeneous medical information systems, and constructing a unified database providing data view function; according to the characteristics and use scenarios of the medical clinical data, the best anonymization strategy is adaptively selected to anonymize the medical clinical data, and the privacy protection and data availability are balanced; the privacy risk of the anonymized data is evaluated, the possibility of data re-identification is quantified by simulating the behavior of an attacker, and a scientific basis is provided for privacy protection strategy adjustment; a data asset directory is established and managed, medical data is regarded as important assets for classification, evaluation and management, and the discoverability and availability of data assets are realized. The application improves the integrity, accuracy and consistency of data, realizes all-round privacy protection, and maximizes the availability and analysis value of data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hospital data management, in particular, to the technical field of data asset management; specifically, it relates to a clinical data asset management method and system based on data intelligent anonymization. BACKGROUND

[0002] With the acceleration of medical informatization process, medical institutions have accumulated a large amount of clinical specialty data, which contains rich medical knowledge and potential value, and can be used in clinical decision making, medical research, drug development and other fields, and has great social and economic value.

[0003] However, medical data is highly sensitive, involving patient privacy and core assets of medical institutions, and how to realize safe sharing and value mining of data under the premise of protecting privacy has become a key challenge in the field of medical big data. Relevant policy documents clearly propose to strengthen the standardized management of medical and health data, and promote data security sharing and application innovation. This provides policy support for medical data asset management, and also puts forward higher requirements for data security and privacy protection.

[0004] Currently, the clinical data of medical institutions is usually scattered in different systems, such as electronic medical record system (EMR), laboratory information system (LIS), picture archiving and communication system (PACS), etc., forming a serious "data island" phenomenon. These systems lack effective data exchange and integration mechanism, leading to difficulty in data integration and full utilization. At the same time, medical data contains a large amount of sensitive information of patients, and faces serious privacy leakage risk in the process of data sharing and utilization.

[0005] Traditional hospital data management systems mainly focus on data storage and basic query functions. Such systems are usually maintained by the hospital information department, and the main functions include data backup, simple statistics and report generation, etc. For example, HIS (hospital information system) and EMR (electronic medical record system) commonly used in hospitals can store and manage patient basic information and medical records, but lack in-depth application of data asset management, and lack of deep data governance and privacy protection functions. Such systems mainly use simple de-identification processing in data anonymization, such as deleting patient name, ID number and other explicit identifiers, which cannot cope with complex re-identification attacks, and the privacy protection capability is limited. The main shortcomings of traditional hospital data management systems are as follows:

[0006] 1. The problem of data island is serious. Medical data is scattered in multiple systems, and lacks unified collection and management mechanism, making it difficult to realize data integration and comprehensive analysis. The data formats of each system are inconsistent, and lack of standardized data exchange interface, increasing the difficulty of data integration.

[0007] 2. The anonymization technology is not intelligent enough, and existing technologies mostly use static and rule-based anonymization methods, such as simple de-identification or fixed generalization rules, which cannot dynamically adjust the anonymization strategy according to data characteristics and usage scenarios, and it is difficult to achieve a good balance between privacy protection and data usability.

[0008] 3. The privacy protection is not comprehensive enough, most systems only focus on the deletion of direct identifiers, ignoring the re-identification risk caused by the combination of quasi-identifiers, and cannot cope with the increasingly complex privacy attack methods. At the same time, there is a lack of scientific evaluation mechanism for the effect of anonymization, making it difficult to quantify the effectiveness of privacy protection.

[0009] 4. The data quality management is insufficient, and existing systems have limited monitoring and management capabilities for data quality, lacking automated data cleaning and quality assessment tools, resulting in data errors, missing and inconsistencies, affecting the accuracy of subsequent analysis.

[0010] 5. Lack of asset management thinking, existing medical data management mostly stays at the level of storage and simple use, lacking the concept and tools of systematic management of data as assets, and cannot fully realize the potential value of data.

[0011] 6. The security audit mechanism is not perfect, the whole process monitoring of data access and use is insufficient, lacking fine-grained permission control and operation audit, there is a risk of data abuse and leakage.

[0012] 7. Cross-institutional collaboration is insufficient, lacking a safe and standardized data sharing mechanism, making it difficult to adopt multi-center clinical research and cross-institutional collaboration, restricting the wide application and value mining of medical data.

[0013] There are also general commercial data governance platforms on the market, such as IBM InfoSphere Information Governance Catalog, Informatica Data Governance, etc. These data governance platforms provide comprehensive data governance functions, including data quality management, metadata management, master data management, etc. However, such general platforms are not designed specifically for the medical field, lacking in-depth understanding and targeted processing of medical data characteristics. In terms of data anonymization, they usually provide basic desensitization functions, using rule-based data masking and replacement, but the anonymization algorithm is relatively simple, making it difficult to balance data usability and privacy protection.

[0014] There are also some open source tools such as ARX Data Anonymization Tool, etc., which provide the implementation of various privacy protection algorithms, including k-anonymity, l-diversity, etc. These tools focus on data anonymization processing, but usually only exist as independent components, lack deep integration with medical data management systems, and have high usage threshold. At the same time, such tools often lack complete data asset management functions and cannot meet the needs of medical institutions for data lifecycle management. SUMMARY

[0015] Therefore, the purpose of the present application is to provide a clinical data asset management method and system based on data intelligent anonymization, to build a unified data asset management platform, to establish a standardized data collection and integration mechanism, to break down data silos, to realize integrated management and comprehensive analysis of medical data, to provide complete data for medical decision-making and research; to develop intelligent anonymization algorithms based on machine learning, which can adaptively adjust the anonymization strategy according to the data characteristics, usage scenarios and privacy requirements, maximize the preservation of data usability and analysis value while protecting privacy; to build a multi-level privacy protection system that not only processes direct identifiers, but also assesses and processes quasi-identifier combinations, while providing a scientific privacy risk assessment method to quantify the anonymization effect and achieve comprehensive privacy protection; through intelligent data cleaning, standardization and quality evaluation mechanism, improve the integrity, accuracy and consistency of data, improve the quality and usability of data, and provide high-quality data foundation for subsequent analysis and application; regarding medical data as important assets, establishing a data asset management mode, through data classification, evaluation, cataloging and lifecycle management, realizing the maximization of data value, creating economic and social benefits for medical institutions; building a fine-grained permission control and whole-process operation audit mechanism, strengthening security audit and compliance management, ensuring the safe use and compliance management of data, preventing data misuse and leakage; establishing a safe and standardized data sharing mechanism, adopting multi-center clinical research and cross-institutional data collaboration, promoting the wide application and value mining of medical data.

[0016] The present application provides a clinical data asset management method based on data intelligent anonymization, comprising the following steps:

[0017] S1, collecting, cleaning and integrating medical clinical data from multiple heterogeneous medical information systems, and constructing a unified database providing data view function;

[0018] The S1 step comprises the following steps:

[0019] S11, based on the HL7 FHIR standard, develop multi-source data interfaces for hospital HIS, EMR, LIS, PACS and other systems, use RESTful API, WebService, database link and other ways to collect data; adopt plug-in design, quickly access new data sources by configuring new data source adapters without modifying the system core code. For traditional systems that do not support standard interfaces, provide ETL toolkits to realize data extraction through customized extraction scripts;

[0020] S12, based on the dual mechanism of timestamp and data version number, realize incremental data synchronization, extract only the newly added or modified data since the last synchronization, reduce network transmission and processing time. Use distributed message queue (such as Kafka) to process data flow to ensure the reliability and real-time performance of data collection. Through the configuration of the scheduling system, realize the data synchronization strategy of on-demand, timing or event triggering; meet the needs of different scenarios;

[0021] S13, clean and convert multi-source data from different medical systems, use regular expressions and natural language processing techniques to identify and fix format errors such as non-standard date formats, special characters, etc. Detect abnormal values based on business rules, mark data detection results that exceed the normal range, and identify potential data errors through machine learning algorithms; use fuzzy matching algorithm to handle the differences of the same patient information in different medical systems, build a unified patient index, and solve the problem of patient identity recognition;

[0022] Regular expression examples are as follows:

[0023] \d: match any digit;

[0024] \w: match letters, numbers and underscores;

[0025] \s: match white space characters;

[0026] []: character set, match any character in the square brackets;

[0027] (): grouping;

[0028] {n,m}: quantifier, match the preceding element n to m times;

[0029] Based on deep learning-based natural language processing technology to identify sensitive information in unstructured medical text, the specific steps are as follows, text preprocessing, word segmentation: split continuous text into independent words. Standardization: convert text to a unified format, such as lowercasing, special character processing.

[0030] Specifically, the present application adopts a multi-strategy fusion fuzzy matching algorithm to process the differences of the same patient information in different medical systems. Different fields have different discrimination abilities, and different weights are assigned. High-recognition fields (such as ID number): weight 0.6-0.8. Medium-recognition fields (such as name + date of birth): weight 0.3-0.5. Low-recognition fields (such as gender, blood type): weight 0.1-0.2. Self-developed weight algorithm, high confidence (> 90).

[0031] S14, standardize and convert medical clinical data from different medical systems, and uniformly convert disease codes in different medical systems into ICD-10 standard, drug codes into ATC classification system, and examination item codes into LOINC standard; develop a medical terminology mapping engine to automatically identify and standardize medical terms in unstructured text and improve the degree of data structuring; build a medical ontology database to establish a data view containing semantic associations between various medical concepts, realize data integration and query at the semantic level;

[0032] S15, automatically extract the structure information of the data source including tables, fields, data types, primary and foreign key relationships, etc. through the metadata API or table structure analysis tool of the medical ontology database. Analyze the data content using machine learning technology to automatically infer the semantic type and sensitivity level of the field, such as personal identifier, medical indicator, time and place, etc. Establish a metadata version control mechanism to record the history of metadata changes and realize the backtracking and comparison analysis of metadata;

[0033] S2, according to the characteristics and use scenarios of medical clinical data, adaptively select the best anonymization strategy, anonymize medical clinical data, balance privacy protection and data usability; evaluate the privacy risk of anonymized data, simulate the behavior of attackers, quantify the possibility of data re-identification, and provide scientific basis for privacy protection strategy adjustment;

[0034] S3, establish and manage data asset directory, classify, evaluate and manage medical data as important assets, realize the discoverability and availability of data assets.

[0035] Further, the method of S2 step according to the characteristics and use scenarios of medical clinical data, adaptively selecting the best anonymization strategy, anonymizing medical clinical data, and balancing privacy protection and data usability includes the following steps:

[0036] S201, identify sensitive information in medical clinical data, use custom recognition rules and regular expressions, identify direct identifiers such as name, ID number, phone number, etc. based on rule engine, extract sensitive information from unstructured text using natural language processing technology, such as patient name, address, phone number in medical records. Use graph pattern matching algorithm to identify quasi-identifier combinations, such as gender, age, treatment time, and zip code, which may lead to re-identification of attribute combinations;

[0037] Extract content from unstructured medical text (such as medical history records), and structure the identified test results for subsequent analysis.

[0038] Specifically, the graph pattern matching algorithm represents each attribute in the data set (such as age, gender, zip code, etc.) as a node in the graph, and establishes a connection between them when two attribute combinations may increase the risk of identification, for example: connecting "rare disease" and "hospital" in the graph, because the combination of these two may identify a specific patient. For example, a patient data set containing the following fields:

[0039] Basic information: gender, age, marital status;

[0040] Geographical information: residential area, hospital;

[0041] Medical information: diagnosis, surgery type, medication.

[0042] For example, the attribute graph constructed by the graph pattern matching algorithm shows that the combination of "residential area + rare disease + age" has high identification, because there are very few patients with rare diseases in a specific age and area.

[0043] S202, develop an adaptive anonymization strategy engine, automatically select the most appropriate anonymization algorithm and parameters according to data type, sensitivity level, usage scenario and privacy requirements; use a multi-dimensional k-anonymity method based on the Mondrian algorithm to anonymize medical clinical data, recursively divide the attribute space so that each equivalence class contains at least k records; use the information loss minimization criterion to select the optimal split attribute and split point, balancing privacy protection and data usability;

[0044] Specifically, different types of attributes are treated differently, direct identifiers can be deleted or replaced with hash, numerical attributes can be perturbed or generalized, time attributes can be date offset or granularity reduction, etc. Use multiple anonymization strength configurations, from low level (only process direct identifiers) to high level (meet differential privacy requirements), to meet the privacy protection needs of different scenarios.

[0045] Preferably, a weighting process based on attribute importance is adopted, with lower generalization degree for attributes that are more important for the study, maximizing the preservation of key information.

[0046] There are various anonymization algorithms, and the most suitable anonymization algorithm and parameters can be selected according to the data type, sensitivity level, usage scenario, and privacy requirements, for example:

[0047] Generalization, replacing precise values with ranges or more general categories. For example, replacing the exact age "37 years old" with the age range "35-40 years old". Applicable to both categorical attributes and numerical attributes, reducing precision while preserving data distribution.

[0048] Microagitation, grouping similar records and replacing original values with average values within the group. For example, replacing records with heights of 172 cm, 173 cm, and 175 cm with 173 cm.

[0049] Noise addition, adding random noise to the original data. For example, adding a random deviation of ±5 mmHg to the blood pressure value.

[0050] S203, on the basis of k-anonymity processing, an L-diversity enhancement method based on attribute reorganization and record merging technology is adopted, so that each equivalence class has at least l different values of sensitive attributes; for numerical sensitive attributes, an entropy-based L-diversity variant is implemented to ensure the uniformity of sensitive value distribution.

[0051] Preferably, an automatic detection mechanism is supported to identify attribute correlations that may cause L-diversity failure and provide repair suggestions.

[0052] Specifically, there are many ways to anonymize data depending on the application scenario, and k-anonymity is one of them. K-anonymity is a multi-dimensional algorithm that groups similar patients into the same anonymous group, ensuring that each anonymous group contains at least k patients, and patients in the same group show the same value or range in sensitive attributes, making it impossible to determine which specific record belongs to which real patient.

[0053] For example, there are 12 diabetic patient data, and 4-anonymity needs to be implemented (at least 4 people per group). Patients A-E: 35-52 years old, HbA1c value 6.8%-8.3%, patients F-L: 55-75 years old, HbA1c value 8.2%-9.6%.

[0054] Further, the method for evaluating the privacy risk of anonymized data in the S2 step by simulating the behavior of an attacker to quantify the possibility of data re-identification includes the following steps:

[0055] S211, construct a re-identification risk model, establish a risk assessment framework based on attacker knowledge model; consider the background knowledge and attack ability that the attacker may master, adopt a risk quantification method based on probability inference, calculate the probability of each record being successfully re-identified, and the average re-identification risk of the entire data set;

[0056] Preferably, different risk metrics are adopted, such as maximum risk, average risk and record proportion risk.

[0057] S212, develop multiple re-identification attack simulators, including linkage attack (using external data set), homologous attack (using other data of the same institution) and reasoning attack (based on internal association of data), perform systematic attack test on anonymized data through simulation attack test process, and evaluate the defense effect;

[0058] Preferably, a detailed attack test report is provided, including success rate, impact range and vulnerability analysis, guiding the optimization of privacy protection strategy.

[0059] S213, construct a risk-utility trade-off evaluation model, comprehensively consider the privacy risk and data availability, and find the best balance point; develop data utility quantification index, evaluate the availability of anonymized data from the dimensions of fidelity and analysis feasibility.

[0060] Preferably, an interactive risk-utility tuning tool is provided, allowing users to adjust anonymization parameters according to specific needs.

[0061] Further, the S3 step comprises the following steps:

[0062] S31, classify and label medical clinical data assets, establish a multi-dimensional asset classification system, the asset classification system includes: data theme (such as clinical diagnosis and treatment, test report, etc.), data type (structured, semi-structured, unstructured), data source; adopt automatic classification method based on machine learning, infer the most possible classification category according to data content and structure characteristics; adopt custom label and automatic label recommendation, realize asset label management, and improve the retrievability of assets;

[0063] S32, construct a data asset value evaluation model, score the asset value from the dimensions of data quality, business importance, use frequency and scarcity; through user feedback and statistical analysis of the actual use value of assets, establish a dynamically adjusted value evaluation mechanism;

[0064] Preferably, an asset value dashboard is developed to visually display the value distribution and ranking of various assets, guiding the priority allocation of resources;

[0065] S33, by ETL process analysis and system log mining, automatically constructing a data blood relationship graph, showing the source, flow and use path of the data, using an interactive blood analysis, the user can start from any data point, trace its source upwards and track its influence downwards;Blood influence analysis function is provided to evaluate the potential impact of source data changes on downstream applications to assist change management decisions.

[0066] Further, the clinical data asset management method based on data intelligent anonymization further comprises: monitoring and improving the data quality of medical clinical data, ensuring the integrity, accuracy and consistency of medical clinical data, and providing high-quality data basis for subsequent analysis and application;

[0067] The method for monitoring and improving the data quality of medical clinical data comprises the following steps:

[0068] S41, a comprehensive quality evaluation system including integrity, accuracy, consistency, timeliness and usability is constructed;Special quality indicators are designed for different data types for multi-dimensional quality evaluation, such as format compliance of structured data and semantic consistency of unstructured data;An automatic quality evaluation engine is developed to regularly scan the quality of data and generate quality scores and problem lists;

[0069] S42, based on historical data patterns and domain knowledge, develop missing value filling algorithms such as time series interpolation, similar case reference and other methods, apply machine learning techniques to identify and correct abnormal values and error data, and improve data accuracy through comparative analysis and rule verification;Develop a conflict resolution engine to automatically identify and handle data conflicts from different source systems, and select the optimal data version based on credibility scores;

[0070] Specifically, for variables that need to be repaired, the algorithm rules of the missing value filling algorithm can be developed separately, for example, if a patient's blood pressure values for a week are [120, 125,?, 130, 122], the missing value can be estimated as (125+130) / 2=127.5.

[0071] S43, implement closed-loop management of quality problems, establish a quality problem work order system, convert automatically detected quality problems into work orders, assign them to relevant responsible persons for processing, realize full-process tracking of problem processing, record problem description, processing scheme, execution result and verification feedback;Build a quality knowledge base to accumulate experience in handling common problems and provide references for similar problems in the future.

[0072] The application also provides a clinical data asset management system based on data intelligent anonymization, which is used to execute the clinical data asset management method based on data intelligent anonymization as described above, comprising:

[0073] Data acquisition and integration module: used for acquiring, cleaning and integrating clinical data from multiple source heterogeneous medical information systems, and constructing a unified data view;

[0074] Intelligent anonymization processing module: used for intelligent anonymization processing of clinical data, and self-adaptive selection of the best anonymization strategy according to data characteristics and use scenarios, balancing privacy protection and data availability;

[0075] Data asset directory module: used for establishing and managing a data asset directory, classifying, evaluating and managing medical data as important assets, and realizing data asset discoverability and availability.

[0076] Further, the intelligent anonymization processing module comprises:

[0077] Privacy risk assessment unit: used for assessing the privacy risk of anonymized data, quantifying the possibility of data re-identification by simulating the behavior of an attacker, and providing a scientific basis for privacy protection strategy adjustment.

[0078] Further, the clinical data asset management system based on data intelligent anonymization further comprises:

[0079] Data quality management module: used for monitoring and improving the data quality of medical clinical data, guaranteeing the integrity, accuracy and consistency of medical clinical data, and providing a high-quality data basis for subsequent analysis and application;

[0080] System management module: used for basic management functions of the platform, including user management, permission control, system configuration, log audit, etc., and ensuring the safe and stable operation of the system.

[0081] The application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the clinical data asset management method based on data intelligent anonymization.

[0082] The application also provides a computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the clinical data asset management method based on data intelligent anonymization.

[0083] Compared with the prior art, the application has the following advantages:

[0084] Compared with the traditional hospital data management system and the commercial data governance platform, the clinical data asset management method and system based on data intelligent anonymization can adaptively select the best anonymization strategy, dynamically adjust the anonymization parameters according to the data characteristics and use scenarios, and significantly improve the privacy protection effect and data usability; the present application is specially designed for medical data, fully considers the characteristics and application requirements of medical data, such as disease coding standardization, medical terminology mapping, clinical data quality management, etc., and is more professional, and is more suitable for the actual needs of medical institutions than the existing general data governance platform; compared with the existing hospital data management system, the present application provides stronger multi-source heterogeneous data integration capability, through standardized interfaces and ETL tool suites, the data integration capability is more prominent, can efficiently integrate medical data from different medical systems, and construct a complete data view; the privacy protection is more comprehensive; the present application not only provides a variety of advanced anonymization technologies, but also develops a scientific privacy risk assessment method, which can quantitatively evaluate the anonymization effect and predict the potential privacy leakage risk, and the privacy protection is more comprehensive and reliable than the prior art; the present application first introduces the asset management method into medical data management, regards the medical data as important assets for systematic management through data classification, evaluation, cataloging and life cycle management, greatly improves the value realization of data, and the method improves the innovation and applicability of medical data management, and has a wide application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0085] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not considered a limitation of the present application.

[0086] In the drawings:

[0087] Figure 1 is a flowchart of the clinical data asset management method based on data intelligent anonymization of the present application;

[0088] Figure 2 is a schematic diagram of the computer device of the embodiment of the present application. DETAILED DESCRIPTION

[0089] The exemplary embodiments will be described in detail hereinbelow with reference to the drawings. The following description is with reference to the drawings, wherein like numerals refer to like elements throughout. The embodiments described in the following exemplary embodiments are not meant to represent all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and articles of manufacture consistent with some aspects of the present disclosure as detailed in the appended claims.

[0090] The terminology used in the disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0091] It is to be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a particular order or hierarchy. These terms are used merely for the purpose of distinguishing between two or more information. For example, without departing from the scope of the present disclosure, a first information can be termed a second information, and similarly, a second information can be termed a first information. The word "if" as used herein means "when" or "upon" or "in response to the determination" depending on the context.

[0092] The embodiments of the present application are further described below.

[0093] The embodiments of the present application provide a clinical data asset management method based on data intelligent anonymization, referring to Figure 1 as shown, comprising the following steps:

[0094] S1, collecting, cleaning and integrating medical clinical data from multi-source heterogeneous medical information systems, and constructing a unified database providing data view function; the step S1 comprises the following steps:

[0095] S11, based on the HL7 FHIR standard, developing multi-source data interfaces of hospital HIS, EMR, LIS and PACS systems, collecting data in multiple ways such as RESTful API, WebService and database link; adopting plug-in design, through the configuration of new data source adapter, new data source can be quickly accessed (without modifying the system core code), for traditional systems that do not support standard interface, ETL tool kit is provided, and data extraction is realized through custom extraction script;

[0096] S12, based on the double mechanism of timestamp and data version number, incremental data synchronization is realized, only the data added or modified since the last synchronization is extracted, distributed message queue (Kafka) is used to process data stream, and through the configuration of scheduling system, on-demand, timing or event-triggered data synchronization strategy is realized;

[0097] S13, cleaning and converting multi-source data of different medical systems, applying regular expressions and natural language processing techniques to identify and repair format errors (non-standard date format, special characters, etc.), detecting abnormal values based on business rules, marking data outside the normal range, and identifying potential data errors through machine learning algorithms; using fuzzy matching algorithms to handle differences in the same patient information in different medical systems, building a unified patient index, and solving the problem of patient identification;

[0098] Examples of regular expressions in this embodiment are as follows:

[0099] \d: matches any digit;

[0100] \w: matches letters, numbers, and underscores;

[0101] \s: matches white space characters;

[0102] []: character set, matches any character within the square brackets;

[0103] (): grouping;

[0104] {n,m}: quantifier, matches the preceding element n to m times;

[0105] Natural language processing based on deep learning to identify sensitive information in unstructured medical text, the specific steps include: text preprocessing, word segmentation: continuous text is divided into independent words; normalization: convert text to a uniform format, such as lowercasing, special character processing.

[0106] In this embodiment: Mr. XX (male, 42 years old) visited a hospital on May 12, 2023, complaining of abdominal pain for 3 days, with a history of type 2 diabetes for 5 years, contact number: XXX, current address: XXX.

[0107] After processing, the structured storage is:

[0108] Name: Mr. XX (personal identifier);

[0109] Date: May 12, 2023 (date identifier);

[0110] Hospital name: Hospital X (institution identifier);

[0111] Phone number: XXX (contact identifier);

[0112] Address: XXX (geographical identifier);

[0113] The embodiment adopts a multi-strategy fusion fuzzy matching algorithm to process the differences of the same patient information in different medical systems. Different fields have different distinguishing abilities, and different weights are assigned. High-recognition field (such as ID card number): weight 0.6-0.8. Medium-recognition field (such as name + date of birth): weight 0.3-0.5. Low-recognition field (such as gender, blood type): weight 0.1-0.2. Self-developed weight algorithm, high confidence (> 90).

[0114] S14, standardize and convert the medical clinical data of different medical systems, and uniformly convert the disease codes in different medical systems into ICD-10 standard, uniformly convert the drug codes into ATC classification system, and uniformly convert the examination item codes into LOINC standard; develop a medical terminology mapping engine to automatically identify and standardize medical terms in unstructured text and improve the degree of data structuring; build a medical ontology database to establish a data view containing semantic associations between various medical concepts, realize data integration and query at the semantic level;

[0115] S15, through the metadata API or table structure analysis tool of the medical ontology database, automatically extract the structure information of the data source (including tables, fields, data types, primary and foreign key relationships, etc.), analyze the data content by using machine learning technology, automatically infer the semantic type and sensitive level of the field (personal identifier, medical indicator, time and place, etc.), establish a metadata version control mechanism, record the metadata change history, realize the backtracking and comparison analysis of the metadata;

[0116] S2, according to the characteristics and use scenarios of medical clinical data, adaptively select the best anonymization strategy, and anonymize the medical clinical data to balance privacy protection and data usability; evaluate the privacy risk of anonymized data, simulate the behavior of attackers, quantify the possibility of data re-identification, and provide scientific basis for privacy protection strategy adjustment;

[0117] The method for adaptively selecting the best anonymization strategy according to the characteristics and use scenarios of medical clinical data to balance privacy protection and data usability includes the following steps:

[0118] S201, identify the sensitive information of medical clinical data, adopt self-defined identification rules and regular expressions, identify direct identifiers (name, ID card number, phone number, etc.) based on a rule engine, extract sensitive information (patient name, address, phone number, etc. in medical records) from unstructured text by using natural language processing technology, and identify quasi-identifier combinations (gender, age, treatment time, postcode, etc. which may lead to re-identification) by using graph pattern matching algorithm;

[0119] Extracting content from unstructured medical text (such as medical records), structuring the identified test results for easy subsequent analysis.

[0120] In this embodiment, the original medical record text is "The patient's blood glucose test result is 7.8 mmol / L, ALT 45 U / L, total cholesterol 5.2 mmol / L", which can be identified and structured as:

[0121] Blood glucose: 7.8 mmol / L;

[0122] ALT: 45 U / L;

[0123] Cholesterol: 5.2 mmol / L;

[0124] The graph pattern matching algorithm represents each attribute in the dataset (such as age, gender, zip code, etc.) as a node in the graph, and establishes a connection between them when the combination of two attributes may increase the risk of identification, for example: connecting "rare disease" and "hospital" in the graph, because the combination of these two may identify a specific patient. If there is a patient data set containing the following fields:

[0125] Basic information: gender, age, marital status;

[0126] Geographical information: residential area, hospital;

[0127] Medical information: diagnosis, surgery type, medication.

[0128] In this embodiment, the attribute graph constructed by the graph pattern matching algorithm shows that the combination of "residential area + rare disease + age" has high identification, because there are very few patients with certain rare diseases in a specific age and area.

[0129] S202, develop an adaptive anonymization strategy engine, automatically select the most suitable anonymization algorithm and parameters according to data type, sensitivity level, usage scenario and privacy requirements; use the multi-dimensional k-anonymity method based on the Mondrian algorithm to anonymize medical clinical data, recursively divide the attribute space so that each equivalence class contains at least k records; use the information loss minimization criterion to select the optimal split attribute and split point, balancing privacy protection and data usability;

[0130] Different types of attributes are treated differently, direct identifiers can be deleted or replaced with hash, numerical attributes can be perturbed or generalized, time attributes can be date offset or granularity reduction, etc. Use multiple anonymization strength configurations, from low level (only process direct identifiers) to high level (meet differential privacy requirements), to meet the privacy protection needs of different scenarios.

[0131] This embodiment selects the most suitable anonymization algorithm and parameters according to data type, sensitivity level, usage scenario and privacy requirements, including:

[0132] Generalization, replacing precise values with ranges or more general categories, replacing the exact age "37 years old" with the age range "35-40 years old". Applicable to both categorical and numerical attributes, reducing precision while preserving data distribution.

[0133] Microagitation, grouping similar records and replacing original values with average values within the group. Records with heights of 172 cm, 173 cm and 175 cm are all replaced with 173 cm.

[0134] Noise addition, adding random noise to the original data, adding a random deviation of ±5 mmHg to the blood pressure value.

[0135] S203, on the basis of k-anonymity processing, using L-diversity enhancement method based on attribute reorganization and record merging technology, so that each equivalence class has at least l different values of sensitive attribute; for numerical sensitive attribute, realize L-diversity variant based on entropy, ensure the uniformity of sensitive value distribution.

[0136] There are 12 diabetic patient data in this embodiment, which need to realize 4-anonymity (at least 4 people in each group). Patients A-E: 35-52 years old, HbA1c value 6.8%-8.3%, patients F-L: 55-75 years old, HbA1c value 8.2%-9.6%.

[0137] To evaluate the privacy risk of anonymized data, the method of quantifying the possibility of data re-identification by simulating the behavior of attackers includes the following steps:

[0138] S211, construct re-identification risk model, establish risk assessment framework based on attacker knowledge model; considering the background knowledge and attack ability that attackers may master, using risk quantification method based on probability inference (using different risk measurement standards such as maximum risk, average risk and record proportion risk), calculating the probability of each record being successfully re-identified, and the average re-identification risk of the whole data set;

[0139] S212, develop a variety of re-identification attack simulators, including linkage attack (using external data set), homologous attack (using other data of the same institution) and reasoning attack (based on internal association of data), through simulation attack test process, perform systematic attack test on anonymized data, evaluate defense effect;

[0140] This embodiment provides detailed attack test report, including success rate, influence range and vulnerability analysis, guiding the optimization of privacy protection strategy.

[0141] S213, construct a risk-utility trade-off evaluation model, comprehensively consider privacy risk and data availability, and find the best balance point; develop data utility quantification indicators to evaluate the availability of anonymized data from the dimensions of fidelity and analysis feasibility.

[0142] The embodiment provides an interactive risk-utility tuning tool, which allows a user to adjust anonymization parameters according to specific needs.

[0143] S3, establish and manage a data asset directory, classify, evaluate and manage medical data as important assets, and realize the discoverability and availability of data assets.

[0144] The S3 step includes the following steps:

[0145] S31, classify and label medical clinical data assets, establish a multi-dimensional asset classification system, and the asset classification system includes: data theme (clinical diagnosis and treatment, test report, etc.), data type (structured, semi-structured, unstructured), data source; adopt an automatic classification method based on machine learning to infer the most possible classification category according to the data content and structure characteristics; adopt self-defined labels and automatic label recommendation to realize asset label management and improve the retrievability of assets;

[0146] S32, construct a data asset value evaluation model to score the asset value from the dimensions of data quality, business importance, use frequency and scarcity; through user feedback and statistical analysis of the actual use value of the assets, a dynamically adjusted value evaluation mechanism is established;

[0147] The embodiment develops an asset value dashboard to intuitively display the value distribution and ranking of various assets, and guide the priority allocation of resources;

[0148] S33, automatically construct a data blood relationship graph through ETL process analysis and system log mining, display the source, flow and use path of data, adopt interactive blood analysis, the user can start from any data point, trace its source upwards, and track its influence downwards; provide blood influence analysis function to evaluate the potential influence of source data change on downstream applications, and assist change management decision.

[0149] The clinical data asset management method based on data intelligent anonymization further includes: monitoring and improving the data quality of medical clinical data, ensuring the integrity, accuracy and consistency of medical clinical data, and providing a high-quality data basis for subsequent analysis and application;

[0150] The method for monitoring and improving the data quality of medical clinical data includes the following steps:

[0151] S41, a comprehensive quality evaluation system including integrity, accuracy, consistency, timeliness and availability dimensions is constructed; special quality indicators are designed for different data types for multi-dimensional quality evaluation, such as format compliance of structured data and semantic consistency of unstructured data; an automatic quality evaluation engine is developed to regularly scan the quality of data and generate quality scores and problem lists;

[0152] S42, based on historical data patterns and domain knowledge, missing value filling algorithms such as time series interpolation, similar case reference and other methods are developed, machine learning techniques are applied to identify and correct abnormal values and error data, and data accuracy is improved through comparative analysis and rule verification; a conflict resolution engine is developed to automatically identify and handle data conflicts from different source systems, and the optimal data version is selected according to the credibility score;

[0153] For variables that need to be repaired, the algorithm rules of the missing value filling algorithm are developed separately, in this embodiment, the blood pressure values of a patient in a week are [120, 125,?, 130, 122], and the missing value can be estimated as (125+130) / 2=127.5.

[0154] S43, quality problem closed-loop management is implemented, a quality problem work order system is established, automatically detected quality problems are converted into work orders and assigned to relevant responsible persons for processing, problem processing whole process tracking is realized, problem description, processing scheme, execution result and verification feedback are recorded, and a quality knowledge base is constructed to accumulate processing experience of common problems for future reference.

[0155] The embodiment of the application also provides a clinical data asset management system based on data intelligent anonymization, which is used for executing the clinical data asset management method based on data intelligent anonymization as described above, and includes:

[0156] A data acquisition and integration module is used for acquiring, cleaning and integrating clinical data from multiple source heterogeneous medical information systems, and constructing a unified data view.

[0157] An intelligent anonymization processing module is used for intelligent anonymization processing of clinical data, and the best anonymization strategy is adaptively selected according to data characteristics and use scenarios to balance privacy protection and data availability; the intelligent anonymization processing module includes:

[0158] A privacy risk assessment unit is used for assessing the privacy risk of anonymized data, quantifying the possibility of data re-identification by simulating the behavior of an attacker, and providing a scientific basis for privacy protection strategy adjustment.

[0159] Data quality management module: used for monitoring and improving the data quality of medical clinical data, ensuring the integrity, accuracy and consistency of medical clinical data, and providing high-quality data basis for subsequent analysis and application;

[0160] Data asset directory module: used for establishing and managing data asset directory, classifying, evaluating and managing medical data as important assets, realizing the discoverability and availability of data assets;

[0161] Privacy risk assessment module: used for assessing the privacy risk of anonymized data, quantifying the possibility of data re-identification by simulating the behavior of attackers, and providing a scientific basis for privacy protection strategy adjustment;

[0162] System management module: used for the basic management function of the platform, including user management, permission control, system configuration, log audit, etc., to ensure the safe and stable operation of the system.

[0163] The embodiment of the application also provides a computer device, Figure 2 is a structural schematic diagram of a computer device provided by the embodiment of the application; as shown in the figure Figure 2 The computer device includes an input system 23, an output system 24, a memory 22 and a processor 21; the memory 22 is used for storing one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the clinical data asset management method based on data intelligent anonymization provided by the above-mentioned embodiment; wherein the input system 23, the output system 24, the memory 22 and the processor 21 can be connected through a bus or other means, Figure 2 For example, the connection through the bus.

[0164] The memory 22 is a kind of readable and writable storage medium of computing device, which can be used to store software program, computer executable program, such as the program instruction corresponding to the clinical data asset management method based on data intelligent anonymization described in the embodiment of the application; the memory 22 can mainly include storage program area and storage data area, wherein the storage program area can store operating system, at least one application program required by function; the storage data area can store data created according to the use of equipment and the like; in addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, for example, at least one magnetic disk storage device, flash memory device or other non-volatile solid state storage device; in some examples, the memory 22 can further include a memory remotely arranged with respect to the processor 21, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0165] The input system 23 can be used to receive inputted digital or character information, and to generate key signal input related to user settings and function control of the device; the output system 24 can include a display device such as a display screen.

[0166] The processor 21 executes various function applications and data processing of the device by running software programs, instructions and modules stored in the memory 22, i.e. implements the above-mentioned clinical data asset management method based on data intelligent anonymization.

[0167] The computer device provided above can be used to execute the clinical data asset management method based on data intelligent anonymization provided in the above-mentioned embodiments, and has corresponding functions and advantages.

[0168] The embodiments of the present application also provide a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to execute the clinical data asset management method based on data intelligent anonymization provided in the above-mentioned embodiments. The storage medium is any various types of memory device or storage device, and includes: installation medium such as CD-ROM, floppy disk or tape system; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic medium (e.g. hard disk or optical storage); register or other similar types of memory elements, etc.; the storage medium can also include other types of memory or combinations thereof; in addition, the storage medium can be located in the first computer system where the program is executed, or can be located in a different second computer system, which is connected to the first computer system through a network (such as the Internet); the second computer system can provide program instructions to the first computer for execution. The storage medium includes two or more storage media which can reside in different locations (e.g. in different computer systems connected through a network). The storage medium can store program instructions (e.g. specifically implemented as a computer program) executable by one or more processors.

[0169] Of course, the storage medium containing computer executable instructions provided by the embodiments of the present application is not limited to the clinical data asset management method based on data intelligent anonymization described in the above embodiments, and can also execute the related operations in the clinical data asset management method based on data intelligent anonymization provided by any embodiments of the present application.

[0170] So far, the technical solutions of the present application have been described in combination with the preferred embodiments, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will all fall within the protection scope of the present application.

[0171] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application; the present application can have various changes and variations for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A clinical data asset management method based on intelligent data anonymization, characterized in that: Includes the following steps: S1. Collect, clean, and integrate medical clinical data from multi-source heterogeneous medical information systems to build a unified database that provides data view functionality; Step S1 includes the following steps: S11. Based on the HL7 FHIR standard, develop multi-source data interfaces for various hospital systems such as HIS, EMR, LIS, and PACS. Collect data using various methods including RESTful API, WebService, and database connections. Adopt a plug-in design, which can quickly connect to new data sources by configuring new data source adapters. For traditional systems that do not support standard interfaces, provide an ETL toolkit to extract data through customized extraction scripts. S12. Based on a dual mechanism of timestamp and data version number, incremental data synchronization is achieved. Only data added or modified since the last synchronization is extracted. A distributed message queue is used to process the data stream. Through a configurable scheduling system, on-demand, timed, or event-triggered data synchronization strategies are implemented. S13. Clean and transform multi-source data from different medical systems, apply regular expressions and natural language processing technology to identify and repair data with format errors, detect outliers based on business rules, mark data detection results that exceed the normal range, identify potential data errors through machine learning algorithms, and use fuzzy matching algorithms to handle differences in the same patient information in different medical systems, build a unified patient index, and solve the patient identification problem. S14. Standardize and convert medical clinical data from different medical systems, unifying disease codes from different medical systems to the ICD-10 standard, drug codes to the ATC classification system, and examination item codes to the LOINC standard; develop a medical terminology mapping engine to automatically identify and standardize medical terms in unstructured text, improving the structuring level of data; construct a medical ontology database, establish a data view containing semantic relationships between various medical concepts, and realize semantic-level data integration and querying; S15. Automatically extract the structural information of the data source through the metadata API or table structure analysis tool of the medical ontology database, use machine learning technology to analyze the data content, automatically infer the semantic type and sensitivity level of the fields, establish a metadata version control mechanism, record the metadata change history, and realize the retrospective and comparative analysis of metadata. S2. Based on the characteristics and usage scenarios of medical clinical data, adaptively select the best anonymization strategy to anonymize the medical clinical data, balancing privacy protection and data availability; assess the privacy risks of anonymized data, quantify the possibility of data re-identification by simulating attacker behavior, and provide a scientific basis for adjusting privacy protection strategies. S3. Establish and manage a data asset catalog, classify, evaluate and manage medical data as important assets, and realize the discoverability and usability of data assets; Step S3 includes the following steps: S31. Classify and label medical clinical data assets, and establish a multi-dimensional asset classification system, which includes: data theme, data type, and data source; adopt an automatic classification method based on machine learning to infer the most likely classification category based on data content and structural characteristics; and use custom tags and automatic tag recommendation to realize asset tag management and improve asset retrieval. S32. Construct a data asset value assessment model to score asset value from the dimensions of data quality, business importance, usage frequency, and scarcity; establish a dynamically adjusted value assessment mechanism through user feedback and statistical analysis of the actual use value of assets. S33. Through ETL process analysis and system log mining, automatically construct a data lineage diagram to display the source, flow and usage path of data. With interactive lineage analysis, users can start from any data point, trace its source upwards and track its impact downwards. Provide lineage impact analysis function to assess the potential impact of source data changes on downstream applications and assist in change management decisions.

2. The clinical data asset management method based on intelligent data anonymization according to claim 1, characterized in that, The method for adaptively selecting the optimal anonymization strategy based on the characteristics and usage scenarios of medical clinical data to anonymize the medical clinical data and balance privacy protection and data usability includes the following steps: S201. Identify sensitive information in medical clinical data by using custom identification rules and regular expressions, identifying direct identifiers based on a rule engine, extracting sensitive information from unstructured text using natural language processing technology, and identifying quasi-identifier combinations using a graph pattern matching algorithm. S202. Develop an adaptive anonymization strategy engine to automatically select the most suitable anonymization algorithm and parameters based on data type, sensitivity level, use case, and privacy requirements; adopt a multidimensional k-anonymization method based on the Mondrian algorithm to anonymize medical clinical data, and recursively partition the attribute space so that each equivalence class contains at least k records; adopt the information loss minimization criterion to select the optimal partitioning attribute and partitioning point to balance privacy protection and data availability. S203. Based on k-anonymization, an L-diversity enhancement method based on attribute recombination and record merging is adopted to ensure that the sensitive attribute in each equivalence class has at least l different values. For numerical sensitive attributes, an entropy-based L-diversity variant is implemented to ensure the uniformity of the sensitive value distribution.

3. The clinical data asset management method based on intelligent data anonymization according to claim 2, characterized in that, The method for assessing the privacy risks of anonymized data in step S2, by simulating attacker behavior and quantifying the likelihood of data re-identification, includes the following steps: S211. Construct a re-identification risk model and establish a risk assessment framework based on the attacker's knowledge model. Considering the background knowledge and attack capabilities that the attacker may possess, adopt a risk quantification method based on probability inference to calculate the probability of each record being successfully re-identified, as well as the average re-identification risk of the entire dataset. S212. Develop various re-identification attack simulators, including link attacks, same-origin attacks, and inference attacks. Through simulating the attack testing process, perform systematic attack tests on anonymized data and evaluate the defense effectiveness. S213. Construct a risk-utility trade-off assessment model to comprehensively consider privacy risks and data availability, and find the best balance point; develop quantitative indicators of data utility to assess the availability of anonymized data from the dimensions of fidelity and analytical feasibility.

4. The clinical data asset management method based on intelligent data anonymization according to claim 1, characterized in that, Also includes: Monitor and improve the quality of medical clinical data to ensure its integrity, accuracy, and consistency, providing a high-quality data foundation for subsequent analysis and application; The method for monitoring and improving the quality of medical clinical data includes the following steps: S41. Construct a comprehensive quality assessment system that includes dimensions of completeness, accuracy, consistency, timeliness, and availability; design specific quality indicators for different data types and conduct multi-dimensional quality assessments; develop an automated quality assessment engine to regularly scan data for quality and generate quality scores and problem lists. S42. Based on historical data patterns and domain knowledge, develop a missing value imputation algorithm, apply machine learning technology to identify and correct outliers and erroneous data, and improve data accuracy through comparative analysis and rule verification; develop a conflict resolution engine to automatically identify and handle data conflicts from different source systems, and select the optimal data version based on credibility scores. S43. Implement closed-loop management of quality issues, establish a quality issue work order system, convert automatically detected quality issues into work orders, assign them to relevant responsible persons for handling, realize full-process tracking of issue handling, record issue descriptions, handling solutions, execution results and verification feedback; and build a quality knowledge base.

5. A clinical data asset management system based on intelligent data anonymization, used to execute the clinical data asset management method based on intelligent data anonymization as described in any one of claims 1-4, characterized in that, include: Data acquisition and integration module: used to collect, clean, and integrate clinical data from multi-source heterogeneous medical information systems to build a unified data view; Intelligent anonymization module: Used to intelligently anonymize clinical data, adaptively selecting the best anonymization strategy based on data characteristics and usage scenarios, balancing privacy protection and data availability; Data Asset Catalog Module: Used to establish and manage data asset catalogs, classifying, evaluating and managing medical data as important assets, and realizing the discoverability and usability of data assets.

6. The clinical data asset management system based on intelligent data anonymization according to claim 5, characterized in that, The intelligent anonymization processing module includes: Privacy Risk Assessment Unit: Used to assess the privacy risks of anonymized data. By simulating attacker behavior, it quantifies the possibility of data re-identification and provides a scientific basis for adjusting privacy protection strategies.

7. The clinical data asset management system based on intelligent data anonymization according to claim 5, characterized in that, Also includes: Data Quality Management Module: Used to monitor and improve the data quality of medical clinical data, ensuring the integrity, accuracy and consistency of medical clinical data, and providing a high-quality data foundation for subsequent analysis and application; System Management Module: Used for basic management functions of the platform, including user management, access control, system configuration, and log auditing, to ensure the safe and stable operation of the system.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the clinical data asset management method based on data intelligent anonymization as described in any one of claims 1-4.

9. A computer device, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the clinical data asset management method based on intelligent data anonymization as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Data de-identification privacy protection method and system

    CN120068157A

  • Data treatment method, terminal equipment and storage medium

    CN120336294A