Medical data optimization method and device based on artificial intelligence
By employing semantic placeholder algorithms and local secure training sandbox technology, the security compliance and data silo issues in collaborative medical data development have been resolved. This has enabled high-fidelity desensitization and efficient model iteration, thereby improving the security and compliance of collaborative medical AI development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING GAZELLE YUNZHI TECH CO LTD
- Filing Date
- 2026-01-24
- Publication Date
- 2026-04-24
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing and artificial intelligence technology, and in particular relates to a method and apparatus for optimizing medical data based on artificial intelligence. Background Technology
[0002] In the field of collaborative research and development of medical artificial intelligence (AI), the conflict between data security, privacy protection, and research efficiency is becoming increasingly prominent in the collaboration between hospitals and external technology partners (such as AI companies and research institutions), becoming a key bottleneck for industry development. From a compliance perspective, original medical records, images, and other patient health information are highly sensitive data, and their use and sharing are subject to clear and stringent restrictions. Directly providing raw data to external parties for model training will face significant legal and compliance risks.
[0003] In R&D practice, external algorithm teams often find it difficult to legally and securely obtain high-quality data from real clinical scenarios. This leads to model training and iteration becoming detached from the actual application environment, which can easily result in inflated algorithm performance or mismatch with clinical scenarios. This seriously affects the final effectiveness and reliability of medical AI products, and the data silo phenomenon significantly restricts the R&D process.
[0004] Traditional desensitization techniques have obvious limitations. Although existing static desensitization methods can remove some direct personal identifiers, they often significantly reduce the training value of the desensitized data and may even introduce noise and bias, making it difficult to meet actual research and development needs.
[0005] Against this backdrop, how to break down data silos and support external teams in conducting effective model training, fine-tuning, and evaluation, while ensuring the security and controllability of original sensitive data and compliance requirements, has become a crucial issue that urgently needs to be addressed in the field of collaborative research and development of medical AI. Summary of the Invention
[0006] This invention provides an artificial intelligence-based medical data optimization method that aligns with the collaborative R&D needs of medical AI. While ensuring the security and compliance of original medical data, it preserves the medical semantic value of the data, facilitates efficient model iteration, and achieves full-process traceability, thereby improving R&D security and compliance. This artificial intelligence-based medical data optimization method includes:
[0007] Obtain raw medical data and training task configuration parameters for the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data;
[0008] The original medical data is classified into medical entities using a semantic placeholder algorithm to obtain de-identified medical data; the de-identified medical data includes etiology data, symptom data, and medical history data.
[0009] Repeat the following steps until the output accuracy of the medical data optimization model reaches the preset threshold:
[0010] Based on the training task configuration parameters, the desensitized medical data is input into the neural network model for training, resulting in a medical data optimization model;
[0011] Calculate the output accuracy of the medical data optimization model based on the anonymized medical data.
[0012] This invention provides an artificial intelligence-based medical data optimization device that meets the needs of collaborative R&D in medical AI. While ensuring the security and compliance of original medical data, it retains the medical semantic value of the data, facilitates efficient model iteration, and achieves full-process traceability, thereby improving R&D security and compliance. This artificial intelligence-based medical data optimization device includes:
[0013] The data acquisition module is used to acquire raw medical data and training task configuration parameters for the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data.
[0014] The desensitized data determination module is used to classify the original medical data into medical entities using a semantic placeholder algorithm to obtain desensitized medical data; the desensitized medical data includes etiology data, symptom data, and medical history data;
[0015] The repeat execution module is used to repeatedly execute the following steps until the output accuracy of the medical data optimization model reaches a preset threshold:
[0016] Based on the training task configuration parameters, the desensitized medical data is input into the neural network model for training, resulting in a medical data optimization model;
[0017] Calculate the output accuracy of the medical data optimization model based on the anonymized medical data.
[0018] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described artificial intelligence-based medical data optimization method.
[0019] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned artificial intelligence-based medical data optimization method.
[0020] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the aforementioned artificial intelligence-based medical data optimization method.
[0021] In this embodiment of the invention, raw medical data and training task configuration parameters for a medical data optimization model are obtained. The raw medical data includes patient information data, hospital information data, and medical data. A semantic placeholder algorithm is used to classify the raw medical data into medical entities, resulting in de-identified medical data. The de-identified medical data includes etiology data, symptom data, and medical history data. The following steps are repeated until the output accuracy of the medical data optimization model reaches a preset threshold: based on the training task configuration parameters, the de-identified medical data is input into a neural network model for training, resulting in a medical data optimization model. The output accuracy of the medical data optimization model is calculated based on the de-identified medical data. This embodiment of the invention aligns with the collaborative R&D needs of medical AI, ensuring the security and compliance of raw medical data while preserving the medical semantic value of the data, facilitating efficient model iteration, and achieving full-process traceability, thereby improving R&D security and compliance. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0023] Figure 1 This is a flowchart of an artificial intelligence-based medical data optimization method in an embodiment of the present invention;
[0024] Figure 2 This is a specific example diagram illustrating the desensitized medical data obtained in an embodiment of the present invention;
[0025] Figure 3 This is a specific example diagram of the medical data optimization model obtained in an embodiment of the present invention;
[0026] Figure 4 This is a specific example diagram illustrating the calculation of the output accuracy of a medical data optimization model in an embodiment of the present invention;
[0027] Figure 5 This is a structural example diagram of an artificial intelligence-based medical data optimization device in an embodiment of the present invention;
[0028] Figure 6 This is a specific example diagram illustrating the structure of the artificial intelligence-based medical data optimization device in an embodiment of the present invention;
[0029] Figure 7 This is a structural diagram of a computer device in an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0031] As mentioned earlier, in existing technologies, when medical AI is collaboratively developed, the original sensitive medical data cannot be directly shared due to strict regulations, resulting in a broken research and development loop and prominent data silos. Although traditional static desensitization technology can remove sensitive labels, it will destroy medical semantic logic and key statistical features, which will significantly reduce the value of desensitized data for AI model training. Furthermore, it lacks effective security collaboration mechanisms and reliable auditing methods, making it difficult to balance privacy protection and research and development efficiency.
[0032] To address this issue, the inventors discovered the need to construct a secure collaborative architecture where the data remains stationary while the model moves. This architecture must achieve high fidelity in de-identified data to safeguard its research and development value, while also preventing the leakage of original data through secure isolation and privacy enhancement technologies. Furthermore, it must establish an unalterable auditing system. Therefore, they proposed a secure collaborative R&D technology solution for medical AI based on privacy enhancement and high-fidelity de-identification, enabling secure and controllable collaborative R&D.
[0033] Figure 1 This is a flowchart of an artificial intelligence-based medical data optimization method in an embodiment of the present invention, such as... Figure 1 As shown, this AI-based medical data optimization method includes:
[0034] Step 101: Obtain the raw medical data and the training task configuration parameters for the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data.
[0035] Step 102: Using a semantic placeholder algorithm, classify the original medical data into medical entities to obtain de-identified medical data; the de-identified medical data includes etiology data, symptom data, and medical history data;
[0036] Repeat the following steps until the output accuracy of the medical data optimization model reaches the preset threshold:
[0037] Step 103: Based on the training task configuration parameters, input the desensitized medical data into the neural network model for training to obtain the medical data optimization model;
[0038] Step 104: Calculate the output accuracy of the medical data optimization model based on the de-identified medical data.
[0039] As shown in Figure 1, in this embodiment of the invention, the following steps are performed: obtaining original medical data and training task configuration parameters for a medical data optimization model; the original medical data includes patient information data, hospital information data, and medical data; using a semantic placeholder algorithm, the original medical data is classified into medical entities to obtain de-identified medical data; the de-identified medical data includes etiology data, symptom data, and medical history data; repeating the following steps until the output accuracy of the medical data optimization model reaches a preset threshold: based on the training task configuration parameters, the de-identified medical data is input into a neural network model for training to obtain the medical data optimization model; and the output accuracy of the medical data optimization model is calculated based on the de-identified medical data.
[0040] Compared with existing quality assessment algorithms, this approach uses fine-grained classification of medical entities and logically preserved semantic placeholders for desensitization. This removes sensitive information while maximizing the preservation of the research and development value of medical data. It solves the problems of traditional desensitization destroying data semantics and separating research and development from data, while completely blocking the risk of original data leakage and ensuring compliance and traceability. This achieves dual protection of privacy and research efficiency in collaborative research and development of medical AI.
[0041] In step 101, the raw medical data and the training task configuration parameters of the medical data optimization model are obtained; the raw medical data includes patient information data, hospital information data and medical data.
[0042] In this embodiment, the training task configuration parameters include model structure parameters, training target parameters, number of iterations, and performance evaluation metric types.
[0043] In a specific embodiment, obtaining the training task configuration parameters for the original medical data and the medical data optimization model includes:
[0044] Obtain raw medical data:
[0045] Raw medical data originates from data providers, such as internal clinical data resources within hospitals, specifically including:
[0046] Patient information data includes personal information related to the patient, such as name, ID number, medical record number, phone number, address, precise age, specific date, and occupation.
[0047] Hospital information data: including hospital name, departments, and other information related to the medical service provider;
[0048] Medical data includes core medical information reflecting the patient's condition and treatment process, such as disease name, symptom description, drug information, surgical procedure records, laboratory test indicators, anatomical location, and imaging diagnostic results.
[0049] All raw medical data is stored in a secure and controllable environment provided by the data provider, such as the hospital's intranet or a secure, certified private cloud, and remains within the security boundaries.
[0050] Get training task configuration parameters:
[0051] Training task configuration parameters are submitted by external collaborative R&D partners, such as AI companies and research institutions, through authorized interfaces. After being securely aggregated and verified by the communication gateway, they are transmitted to the system. Specifically, these parameters include:
[0052] Model structure parameters: Define the core architecture information of the neural network model, such as network layers, number of neurons, activation function type, and feature extraction module configuration;
[0053] Training objective parameters: Set the loss function type and optimizer selection for model training, such as SGD, Adam, learning rate, and other key training objective-related parameters;
[0054] Number of iterations: Specifies the maximum number of iterations during model training, serving as an important criterion for terminating model training;
[0055] Performance evaluation metric types: clearly defined macro-aggregate metrics used to measure the effectiveness of model training, including accuracy, recall, F1 score, AUC value, confusion matrix, loss curve, etc.
[0056] Figure 2 This is a specific example diagram illustrating the desensitized medical data obtained in an embodiment of the present invention, such as... Figure 2 As shown, the original medical data is classified into medical entities using a semantic placeholder algorithm to obtain anonymized medical data, which may include:
[0057] Step 201: Using a pre-defined medical domain named entity recognition model, the entities in the original medical data are classified into direct identifiers, quasi-identifiers, and core medical entities;
[0058] Step 202: Delete or randomly replace the direct identifiers, and generalize or offset the quasi-identifiers.
[0059] Step 203: Replace the medical core entities with semantically typed placeholders using a semantic placeholder algorithm to obtain desensitized medical data.
[0060] In a specific embodiment, de-identified medical data is obtained through a semantic placeholder algorithm, including:
[0061] Medical entity classification:
[0062] A pre-trained medical domain Named Entity Recognition (NER) model is used to perform deep recognition and fine classification of entities in raw medical data, including patient information, hospital information, and medical data, specifically dividing them into three categories:
[0063] Direct identifiers (Class I): include information directly related to the patient's personal identity, such as name, ID number, medical record number, telephone number, and address;
[0064] Quasi-identifiers / Sensitive attributes (Class II): These include sensitive attribute information that may indirectly relate to an individual, such as precise age, specific date, occupation, department, and hospital name.
[0065] Medical core entities (Class III): These include information that constitutes the core of medical semantic logic, such as disease names, symptom descriptions, drugs, surgical procedures, laboratory test indicators, and anatomical locations.
[0066] Sensitive entity handling:
[0067] Differentiated processing is performed on different types of entities after classification:
[0068] For direct identifiers (Class I): the approach is to completely delete or replace them with random identifiers that have no business meaning, thus completely severing their association with specific individuals;
[0069] Alignment identifiers / sensitive attributes (Class II): Perform generalization or offset processing, for example, generalize "52 years old" to "[AGE_GROUP_50-60]", and offset "2023-10-26" to "[DATE_MONTH_2023-10]", which hides specific sensitive information while retaining the general characteristics of the attribute.
[0070] Replacement of placeholders for core medical entities:
[0071] The system generates and maintains a unique dynamic desensitization mapping table for each original medical record. It uses a semantic placeholder algorithm to perform structured replacement of core medical entities (Class III): replacing various core medical entities with placeholders containing semantic type and relational information, while ensuring mapping consistency for the same entities within the same medical record. For example, "recurrent cough" is replaced with "[SYMPTOM_COUGH]", "chest CT" with "[EXAM_CT_CHEST]", and "right lower lobe mass" with "[FINDING_MASS_RIGHT_LOWER_LOBE]", fully preserving the clinical logic chain of "symptom-examination method-imaging findings". Ultimately, this results in desensitized medical data that balances privacy and semantic integrity.
[0072] Figure 3This is a specific example diagram of the medical data optimization model obtained in an embodiment of the present invention, such as... Figure 3 As shown, based on the training task configuration parameters, desensitized medical data is input into a neural network model for training, resulting in a medical data optimization model, which may include:
[0073] Step 301: Load the preset neural network model and training task configuration parameters into the local security training sandbox;
[0074] Step 302: Based on the training task configuration parameters, train the neural network model using desensitized medical data;
[0075] Step 303: After completing the preset number of training iterations, stop training to obtain the medical data optimization model.
[0076] In a specific embodiment, a medical data optimization model is obtained, including:
[0077] Load model and training task configuration parameters:
[0078] In a local secure training sandbox deployed within the data provider's secure and controllable internal environment, the preset neural network model is loaded. At the same time, the training task configuration parameters submitted by the external collaborative R&D party through the authorized interface and verified by the security aggregation and communication gateway are imported. These parameters include model structure parameters, training target parameters, number of iterations, performance evaluation index types, etc., providing basic configuration for model training.
[0079] Model trained based on desensitized medical data:
[0080] Strictly adhering to the principle of "data remains still, model moves," model training is conducted entirely within a local secure training sandbox, without transmitting any original medical data or reversible intermediate data outside the sandbox. Based on training objective parameters such as loss function type, optimizer, and learning rate configured in the training task, high-fidelity, anonymized medical data is used as training data to drive iterative training of the neural network model. If collaborative modes requiring the exchange of model parameters, such as federated learning, are adopted, random noise perturbations that meet differential privacy requirements are first applied to model gradients or weight updates generated during training within the sandbox to ensure that parameter updates do not leak individual privacy information. During training, data interaction is managed through a dedicated secure communication protocol, prohibiting any requests for reverse queries or locating original data, ensuring unidirectional data flow security.
[0081] Stop training and output the model:
[0082] The model training iteration count is continuously monitored. When the training rounds reach the preset number of iterations in the training task configuration parameters, model training is stopped. At this point, the model trained in the sandbox is the medical data optimization model. This model is trained on anonymized data that retains medical semantic logic and can adapt to the needs of actual clinical application scenarios.
[0083] Figure 4 This is a specific example diagram illustrating the calculation of the output accuracy of a medical data optimization model in an embodiment of the present invention, as shown below. Figure 4 As shown, calculating the output accuracy of the medical data optimization model based on anonymized medical data can include:
[0084] Step 401: Obtain the validation set data from the de-identified medical data;
[0085] Step 402: Input the validation set data into the medical data optimization model to obtain the model prediction results;
[0086] Step 403: Compare the model prediction results with the validation set data, and count the number of samples that were correctly predicted;
[0087] Step 404: Calculate the output accuracy of the medical data optimization model based on the ratio of the number of correctly predicted samples to the total number of samples in the validation set.
[0088] In a specific embodiment, calculating the output accuracy of the medical data optimization model includes:
[0089] Obtain validation set data:
[0090] A subset of high-fidelity, desensitized medical data is set as the validation set, which is strictly separated from the training set and retains the same medical semantic logic and clinical relevance features. It covers core medical information such as etiology, symptoms, and medical history, ensuring that the validation process can truly reflect the model's adaptability in clinical scenarios.
[0091] The model prediction results are obtained:
[0092] Within a local secure training sandbox, the validation set data is fully input into the pre-trained medical data optimization model. Based on the medical patterns and feature associations learned during training, the model infers clinical problems in the validation set data, such as disease diagnosis and efficacy prediction, and outputs the corresponding model prediction results. The entire process is executed within the sandbox, without disclosing any original validation set data or intermediate calculation results.
[0093] The number of samples that were correctly predicted:
[0094] Extract real clinical conclusions from the validation set data, such as diagnosed diseases and actual treatment effects, as benchmark answers. Compare the model's prediction results with these benchmark answers one by one. Based on preset judgment rules, such as the prediction of a disease matching the actual disease in a disease diagnosis task, the model is judged as correct. Count the number of samples where the model's prediction results match the benchmark answer, i.e., the number of correctly predicted samples.
[0095] Calculate the output accuracy:
[0096] Based on the number of correctly predicted samples and the total number of validation samples obtained from statistics, the output accuracy of the medical data optimization model is calculated.
[0097] At the same time, based on the performance evaluation index type specified in the training task configuration parameters, macro-aggregated indexes such as recall rate, F1 score, and AUC value are calculated simultaneously and used as the basis for evaluating model performance. All evaluation results are only fed back to external collaborators through secure aggregation and communication gateways, without involving the specific prediction details of individual samples.
[0098] In this embodiment, after the output accuracy of the medical data optimization model reaches a preset threshold, the following may be included:
[0099] Obtain operation log information for medical data optimization model training and write the operation log information to the blockchain;
[0100] Based on the operation log information in the blockchain, a compliance assessment report for medical data optimization is generated.
[0101] In a specific embodiment, after the output accuracy of the medical data optimization model reaches a preset threshold, the following steps are also included:
[0102] Obtain operation log information and write it to the blockchain:
[0103] The system automatically captures and extracts operation log information from the entire training process of the medical data optimization model through the audit and evidence storage module. Log elements include: timestamps of key operations, executing entities (e.g., external collaborating R&D party identifiers), system execution modules, operation types (e.g., data access requests, anonymization processing, sandbox training start / stop, external command reception, model performance evaluation, etc.), identifiers of the anonymized data resources involved, unique identifiers of the training tasks, and summary descriptions of data flowing outwards. Subsequently, the key content of the above operation log information, or its hash value, is written in real-time to a permissioned blockchain managed by the data provider or co-governed by multiple parties. Leveraging the distributed and immutable characteristics of blockchain, the authenticity and integrity of the operation logs are ensured, preventing unilateral modification or deletion.
[0104] Generate compliance assessment reports based on blockchain logs:
[0105] The system invokes the compliance report generation function of the audit and evidence storage module. Based on the complete operation logs already stored in the blockchain, and in accordance with the relevant privacy protection requirements of the medical industry, it automatically extracts key compliance information from the logs, such as authorized operation records, data anonymization compliance proof, evidence that the original data has not been leaked, and access control records, and organizes, verifies, and analyzes them. The final result is a standardized medical data optimization compliance assessment report, clearly presenting the compliance basis for the entire process of data use and model training, significantly reducing the cost of manual compilation and compliance proof, and providing verifiable compliance audit evidence for data providers and regulatory agencies.
[0106] The AI-based medical data optimization method of this invention has been verified to have the following beneficial effects:
[0107] 1. Through the architecture design of "data remains stationary while the model moves", the original sensitive medical data always remains in the local secure environment of the data provider. Combined with technologies such as medical entity classification and desensitization, differential privacy protection, etc., the risk of data leakage is completely eliminated. In particular, it removes the core legal obstacles to cross-border medical AI collaborative research and development.
[0108] 2. By adopting a dynamic semantic placeholder algorithm, the medical causal logic, clinical correlation characteristics and population statistical distribution in the medical records are fully preserved while removing privacy information, so that the de-identified data has a high semantic equivalence value and solves the problem of low R&D value of traditional de-identified data.
[0109] 3. By relying on the dual isolation of the local security training sandbox and the secure communication gateway, the entire process of data use is managed in a closed loop, ensuring that the hospital's original data assets are always within the security boundary and effectively avoiding the risk of data assets getting out of control during collaborative research and development.
[0110] This invention also provides an artificial intelligence-based medical data optimization device, as described in the following embodiments. Since the principle by which this device solves the problem is similar to that of the artificial intelligence-based medical data optimization method, the implementation of this device can refer to the implementation of the artificial intelligence-based medical data optimization method; repeated details will not be elaborated further.
[0111] Figure 5 This is a structural example diagram of an artificial intelligence-based medical data optimization device in an embodiment of the present invention, such as... Figure 5 As shown, the AI-based medical data optimization device includes:
[0112] The data acquisition module 501 is used to acquire raw medical data and training task configuration parameters of the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data.
[0113] The desensitized data determination module 502 is used to classify the original medical data into medical entities using a semantic placeholder algorithm to obtain desensitized medical data; the desensitized medical data includes etiology data, symptom data, and medical history data;
[0114] The repeat execution module 503 is used to repeatedly execute the following steps until the output accuracy of the medical data optimization model reaches a preset threshold:
[0115] Based on the training task configuration parameters, the desensitized medical data is input into the neural network model for training, resulting in a medical data optimization model;
[0116] Calculate the output accuracy of the medical data optimization model based on the anonymized medical data.
[0117] In one embodiment, the training task configuration parameters include model structure parameters, training target parameters, number of iterations, and performance evaluation metric type.
[0118] In one embodiment, the de-identified data determination module 502 is specifically used for:
[0119] Using a pre-defined medical domain named entity recognition model, entities in the original medical data are classified into direct identifiers, quasi-identifiers, and core medical entities;
[0120] Direct identifiers are deleted or randomly replaced, while quasi-identifiers are generalized or offset.
[0121] By using a semantic placeholder algorithm to replace semantically typed placeholders with medical core entities, desensitized medical data is obtained.
[0122] In one embodiment, based on training task configuration parameters, desensitized medical data is input into a neural network model for training to obtain a medical data optimization model, including:
[0123] Load the preset neural network model and training task configuration parameters into the local secure training sandbox;
[0124] Based on the training task configuration parameters, the neural network model is trained using desensitized medical data;
[0125] After completing the preset number of training iterations, training is stopped, and the optimized medical data model is obtained.
[0126] In one embodiment, the output accuracy of the medical data optimization model is calculated based on the anonymized medical data, including:
[0127] Obtain validation set data from de-identified medical data;
[0128] The validation set data is input into the medical data optimization model to obtain the model prediction results;
[0129] Compare the model's prediction results with the validation set data, and count the number of samples that were correctly predicted.
[0130] The output accuracy of the medical data optimization model is calculated based on the ratio of the number of correctly predicted samples to the total number of samples in the validation set.
[0131] Figure 6 This is a specific example diagram of the structure of the artificial intelligence-based medical data optimization device in an embodiment of the present invention, as shown below. Figure 6 As shown in one embodiment, Figure 5 The artificial intelligence-based medical data optimization device shown in the embodiment of the present invention may further include: a blockchain storage module 601.
[0132] In one embodiment, the blockchain storage module 601 is specifically used for:
[0133] After the output accuracy of the medical data optimization model reaches a preset threshold, the operation log information of the medical data optimization model training is obtained and written into the blockchain.
[0134] Based on the operation log information in the blockchain, a compliance assessment report for medical data optimization is generated.
[0135] Based on the aforementioned inventive concept, such as Figure 7 As shown, the present invention also proposes a computer device 700, including a memory 710, a processor 720, and a computer program 730 stored in the memory 710 and executable on the processor 720. When the processor 720 executes the computer program 730, it implements the aforementioned artificial intelligence-based medical data optimization method.
[0136] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned artificial intelligence-based medical data optimization method.
[0137] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the aforementioned artificial intelligence-based medical data optimization method.
[0138] In this embodiment of the invention, raw medical data and training task configuration parameters for a medical data optimization model are obtained. The raw medical data includes patient information data, hospital information data, and medical data. A semantic placeholder algorithm is used to classify the raw medical data into medical entities, resulting in de-identified medical data. The de-identified medical data includes etiology data, symptom data, and medical history data. The following steps are repeated until the output accuracy of the medical data optimization model reaches a preset threshold: based on the training task configuration parameters, the de-identified medical data is input into a neural network model for training, resulting in a medical data optimization model. The output accuracy of the medical data optimization model is calculated based on the de-identified medical data. This embodiment of the invention aligns with the collaborative R&D needs of medical AI, ensuring the security and compliance of raw medical data while preserving the medical semantic value of the data, facilitating efficient model iteration, and achieving full-process traceability, thereby improving R&D security and compliance.
[0139] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0143] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A medical data optimization method based on artificial intelligence, characterized in that, include: Obtain raw medical data and training task configuration parameters for the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data. The original medical data is classified into medical entities using a semantic placeholder algorithm to obtain de-identified medical data; the de-identified medical data includes etiology data, symptom data, and medical history data. Repeat the following steps until the output accuracy of the medical data optimization model reaches the preset threshold: Based on the training task configuration parameters, the desensitized medical data is input into the neural network model for training, resulting in a medical data optimization model; Calculate the output accuracy of the medical data optimization model based on the anonymized medical data.
2. The method as described in claim 1, characterized in that, The training task configuration parameters include model structure parameters, training objective parameters, number of iterations, and performance evaluation metric types.
3. The method as described in claim 1, characterized in that, The original medical data is classified into medical entities using a semantic placeholder algorithm to obtain anonymized medical data, including: Using a pre-defined medical domain named entity recognition model, entities in the original medical data are classified into direct identifiers, quasi-identifiers, and core medical entities; Direct identifiers are deleted or randomly replaced, while quasi-identifiers are generalized or offset. By using a semantic placeholder algorithm to replace semantically typed placeholders with medical core entities, desensitized medical data is obtained.
4. The method as described in claim 1, characterized in that, Based on the training task configuration parameters, anonymized medical data is input into a neural network model for training, resulting in a medical data optimization model, including: Load the preset neural network model and training task configuration parameters into the local secure training sandbox; Based on the training task configuration parameters, the neural network model is trained using desensitized medical data; After completing the preset number of training iterations, training is stopped, and the optimized medical data model is obtained.
5. The method as described in claim 1, characterized in that, Based on the anonymized medical data, calculate the output accuracy of the medical data optimization model, including: Obtain validation set data from de-identified medical data; The validation set data is input into the medical data optimization model to obtain the model prediction results; Compare the model's prediction results with the validation set data, and count the number of samples that were correctly predicted. The output accuracy of the medical data optimization model is calculated based on the ratio of the number of correctly predicted samples to the total number of samples in the validation set.
6. The method as described in claim 1, characterized in that, After the output accuracy of the medical data optimization model reaches a preset threshold, the following is included: Obtain operation log information for medical data optimization model training and write the operation log information to the blockchain; Based on the operation log information in the blockchain, a compliance assessment report for medical data optimization is generated.
7. A medical data optimization device based on artificial intelligence, characterized in that, include: The data acquisition module is used to acquire raw medical data and training task configuration parameters for the medical data optimization model; the raw medical data includes patient information data, hospital information data, and medical data. The desensitized data determination module is used to classify the original medical data into medical entities using a semantic placeholder algorithm to obtain desensitized medical data; the desensitized medical data includes etiology data, symptom data, and medical history data; The repeat execution module is used to repeatedly execute the following steps until the output accuracy of the medical data optimization model reaches a preset threshold: Based on the training task configuration parameters, the desensitized medical data is input into the neural network model for training, resulting in a medical data optimization model; Calculate the output accuracy of the medical data optimization model based on the anonymized medical data.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1-6.