Medical big data vectorization knowledge migration method for protecting medical privacy

By constructing a high-dimensional semantic vector database and embedding model, the problems of privacy leakage and cross-institutional knowledge fusion in medical data sharing are solved. This enables deep medical knowledge sharing across institutions and improves the performance of large language models, breaking down model barriers and promoting ecological collaboration in medical artificial intelligence.

CN120874111APending Publication Date: 2025-10-31ZHONGSHAN HOSPITAL FUDAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510977078.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies face challenges in medical data sharing and model training, including privacy risks, difficulties in cross-institutional knowledge integration, insufficient model generalization ability, significant differences in data formats, and a lack of unified semantic expression mechanisms. These issues limit the cross-institutional sharing and performance improvement of medical AI models.

Method used

By constructing a high-dimensional semantic vector database and embedding model, medical texts are desensitized and vectorized using natural language processing and deep learning techniques. Standardized API interfaces are provided, and a dual strategy is adopted to enhance training input, enabling fine-tuning of external large language models while ensuring privacy protection and knowledge sharing.

Benefits of technology

It enables cross-institutional sharing of deep medical knowledge without compromising personal privacy, enhances the medical understanding and reasoning capabilities of large language models, breaks down model barriers, reduces deployment costs, and promotes ecological collaboration in medical artificial intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874111A_ABST
    Figure CN120874111A_ABST
Patent Text Reader

Abstract

The invention provides a data vectorization knowledge migration method for protecting medical privacy, and relates to the technical field of cardiology medical treatment. According to the method, original data transmission is replaced by vectorization, so that sensitive information of a patient is thoroughly prevented from being leaked, and medical data privacy and compliance are guaranteed; according to the method, knowledge structures including medical logic, treatment paths, language expressions and the like implied in medical records are shared in an embedded form through the vector database and are called by multiple models, so that deep medical knowledge sharing is realized; according to the method, an integratable and accessible medical vector knowledge interface is provided for a general large model of each large manufacturer, the performance of the general large model in a medical task is improved, and the medical understanding and reasoning ability of the large language model are enhanced; according to the method, a cooperation mode of'data immobility and semantic circulation 'of knowledge between models is established, a closed barrier between current large medical models is broken, and ecological collaboration and model complementation of medical artificial intelligence are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical and artificial intelligence technologies, and specifically to a method for vectorized knowledge transfer of medical data that protects medical privacy. Background Technology

[0002] With the widespread application of artificial intelligence technology, especially Large Language Models (LLMs), in the medical field, the demand for model training, inference, and services based on medical record data is constantly growing. In order to realize more professional and generalizable large medical models, obtaining high-quality, structured, and semantically rich clinical data has become crucial. However, under the current technological conditions, the use and sharing of medical data still face significant challenges, especially in terms of compliance, security, and ethics.

[0003] Traditional methods for training medical AI models typically rely on closed-loop data development within a single institution, meaning that model training, evaluation, and application are confined to a single hospital or data platform. While this approach effectively ensures data remains within its domain, it lacks cross-institutional and cross-regional knowledge fusion, resulting in slow model performance improvement and insufficient generalization ability. Some technologies attempt to promote medical data sharing and collaborative modeling through data anonymization and joint modeling (such as federated learning), but these solutions still suffer from the following prominent problems in practice:

[0004] 1. Data anonymization cannot completely eliminate privacy risks: Even if explicit information such as names and ID numbers are removed, implicit information in clinical texts may still be used to infer the patient's identity, posing a risk of privacy leakage;

[0005] 2. Federated learning technology is complex, requires high computing power, and has a high deployment threshold: Although it avoids the direct sharing of raw data, there are high network, security, and resource coordination costs when deploying it between actual medical institutions;

[0006] 3. Model performance depends on the quality and quantity of local data: Due to significant differences in data formats and document structures among medical institutions, model transfer and generalization face the problem of "contextual inconsistency," which affects model performance;

[0007] 4. Lack of a unified, high-quality medical semantic expression mechanism: Currently, there is no standardized method that can preserve the deep semantic logic of the original data while protecting it, and achieve cross-model sharing;

[0008] Furthermore, as the application of general-purpose large language models (such as GPT-type models) in tasks such as medical question answering and case analysis deepens, how to enable them to better "understand medicine" has become an urgent technical need. The current mainstream approach is to improve the medical capabilities of the models through fine-tuning or knowledge injection. However, due to the closed and non-shareable nature of the data, these approaches are often limited to the data that each vendor can obtain, resulting in the coexistence of "data silos" and "model barriers," which prevent the formation of an ecosystem of collaboration.

[0009] Against this backdrop, the medical and health community urgently needs a protection solution that can ensure medical data stays within the institution and does not leak personal privacy, while also enabling deep knowledge sharing and capability enhancement across institutions and models. This solution should possess high security, strong semantic expression capabilities, good model compatibility, and low deployment costs, and be able to serve the continuous optimization and widespread application of various medical artificial intelligence models. Summary of the Invention

[0010] To address the aforementioned shortcomings of existing medical institution data processing methods involving medical privacy, we provide a vectorized knowledge transfer method for medical big data that can effectively protect personal medical privacy from disclosure. Specifically, this invention includes the following technical solutions;

[0011] The first aspect of this invention provides a data vectorization knowledge transfer method for protecting medical privacy, comprising the following steps:

[0012] S1. Corpus Acquisition and Preprocessing: Collect inpatient medical record data from the cardiology and cardiac surgery departments. The medical record data includes at least one of the following: admission record, discharge summary, surgical record, ward round record, examination report, nursing report, and laboratory report. Desensitize the medical record data to remove sensitive personal information and eliminate the risk of implicit identification.

[0013] The anonymized medical record data was processed by sentence segmentation, word segmentation, stop word removal, and standardization of professional terminology to construct a unified corpus.

[0014] S2. Vector Database Construction: Using natural language processing and deep learning technologies, the preprocessed medical text is transformed into high-dimensional semantic vectors at the word and sentence levels. The high-dimensional semantic vectors and their corresponding word or sentence indexes are stored in the vector database to construct an efficient near nearest neighbor index structure to support high-performance semantic retrieval.

[0015] S3. Embedded Model Generation: Based on the constructed vector database, an embedded model is further trained. The embedded model is used to provide word-level or sentence-level vector representations in external fine-tuning tasks. The embedded model is loaded by the external model during the call process as its word vector initialization or enhancement module.

[0016] S4. Fine-tuning and Enhancement: Provides a standardized API interface for external pre-trained large language models to access the embedded model and vector database through the API interface during the fine-tuning stage; during the fine-tuning process, a dual strategy is adopted to enhance the training input to generate new training samples, and the enhanced training samples are used to fine-tune the external pre-trained large language model to obtain a model adapted to downstream medical tasks.

[0017] As a preferred implementation, the dual-strategy augmentation training input in step S4 specifically includes:

[0018] With a probability of 70%, the vector representations of each word in the input text are fused, and the original word vectors of the external large model and the corresponding word vectors of the embedded model are merged in a weighted manner.

[0019] With a 30% probability, several words are randomly selected from the input sentence, and the most similar words are searched through the vector database and replaced with the original words to generate new training samples.

[0020] In one implementation, in step S1, the medical record data is real inpatient medical record data covering a ten-year period with a sample size of more than 30,000.

[0021] Optionally, in step S3, the preprocessed medical text is pre-trained using the Word2Vec model to obtain high-dimensional semantic vectors; the vector database uses the Faiss library to construct an efficient approximate nearest neighbor index structure.

[0022] Furthermore, in step S3, the embedding model is trained in a self-supervised manner, and the training tasks include semantic similarity judgment and context prediction.

[0023] In one implementation, in step S4, the external pre-trained large language model is at least one of GPT, GLM, and DeepSeek-Med;

[0024] The downstream medical tasks include semantic retrieval, medical question answering, intelligent decision support, and diagnostic reasoning related to heart diseases.

[0025] Preferably, in step S4, the weighting method is a weighted average, a gated, or a learning-based combination.

[0026] Furthermore, in step S4, when searching for the most similar word through the vector database, the search is based on cosine similarity.

[0027] A second aspect of this invention provides a data vectorization knowledge transfer system for protecting medical privacy, used to implement the data vectorization knowledge transfer method for protecting medical privacy as described above, comprising:

[0028] Corpus Acquisition and Preprocessing Module: This module collects inpatient medical record data from the cardiology and cardiac surgery departments and performs anonymization on the data. It then performs sentence segmentation, word segmentation, stop word removal, and professional terminology standardization on the anonymized medical record data to construct a unified corpus.

[0029] Vector Database and Embedding Model Building Module: Used to convert preprocessed medical text into high-dimensional semantic vectors and build a vector database, as well as train a medical-specific embedding model;

[0030] Model fine-tuning and enhancement module: used to integrate the embedded model with the external pre-trained large language model, and to enhance the training input and fine-tune the external pre-trained large language model using a dual strategy;

[0031] A standardized API interface module is provided for external pre-trained large language models to call the vector database and embedding model.

[0032] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the data vectorization knowledge transfer method for protecting medical privacy as described above.

[0033] Compared with the prior art, the present invention has at least the following beneficial effects:

[0034] 1. By replacing the original data transmission with vectorization, the leakage of sensitive personal information, especially patient information, can be completely avoided, thereby ensuring the privacy and compliance of medical data;

[0035] 2. By using a vector database, the medical logic, treatment pathways, language expressions, and other knowledge structures implicit in medical records are shared in an embedded form for use by multiple models, thereby achieving deep medical knowledge sharing;

[0036] 3. It can provide an integrable and accessible medical vector knowledge interface for general large models from various manufacturers, improve their performance in medical tasks, and enhance the medical understanding and reasoning ability of large language models;

[0037] 4. A collaborative model of "data not moving, semantic flow" between models has been created, breaking down the closed barriers between current large medical models and promoting ecological collaboration and model complementarity in medical artificial intelligence;

[0038] 5. This invention not only provides a new, safe, and efficient data support and knowledge injection path for medical large-scale models, but also offers a highly feasible solution for building national infrastructure for medical large-scale models and realizing in-depth mining of the value of medical data. Attached Figure Description

[0039] Figure 1The flowchart shows the data vectorization knowledge transfer method for protecting medical privacy provided by this invention.

[0040] Figure 2 This is an illustration of the overall architecture of the present invention. Detailed Implementation

[0041] This invention provides a technical method for enhancing the performance of externally pre-trained large language models by integrating a vector database and an embedding model. Its core features include: constructing a medical-specific semantic vector database and an embedding model based on a natural language corpus of more than 30,000 inpatient medical records from the Department of Cardiology and Department of Cardiac Surgery accumulated over ten years by Zhongshan Hospital affiliated with Fudan University; and achieving deep involvement of the database in the fine-tuning process of the external model through a dual strategy of probability control, thereby realizing the transfer of medical knowledge and the enhancement of the semantic capabilities of the large model without directly transmitting the original medical data.

[0042] This invention can solve a key problem in the development of existing medical artificial intelligence models: how to achieve cross-institutional sharing of medical knowledge and effective improvement of the performance of large language models without leaking individual / patient privacy or over-relying on raw data transmission.

[0043] To this end, we have developed a data sharing and model enhancement method based on a vector database. This method utilizes over 30,000 real medical records from cardiology and cardiac surgery departments. First, semantic features are extracted using deep learning algorithms and natural language processing techniques. Then, these features are converted into irreversible high-dimensional vector embeddings to construct a medical-specific semantic vector database. While ensuring data anonymization and irreversible restoration, this vector database and its corresponding embedding model parameters are made available to third-party large language model vendors, effectively enhancing the medical capabilities of the model and facilitating knowledge transfer.

[0044] The technical solution of the present invention will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them; and the structures shown in the accompanying drawings are merely illustrative and do not represent physical objects. It should be noted that all other embodiments obtained by those skilled in the art based on these embodiments of the present invention are within the scope of protection of the present invention. Furthermore, in the absence of conflict, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other.

[0045] It should be understood that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0046] Example 1

[0047] Please refer to Figure 1 A data vectorization knowledge transfer method for protecting medical privacy includes the following steps:

[0048] S1. Corpus Acquisition and Preprocessing: Collect inpatient medical record data from the cardiology and cardiac surgery departments. The medical record data includes at least one of the following: admission record, discharge summary, surgical record, ward round record, examination report, nursing report, and laboratory report. Desensitize the medical record data to remove sensitive personal information and eliminate the risk of implicit identification. Desensitization methods may include K-anonymity and differential privacy technology.

[0049] The anonymized medical record data is processed through sentence segmentation, word segmentation, stop word removal, and standardization of professional terminology to construct a unified corpus. The direct removal of sensitive personal information and the elimination of implicit identification risks (such as through the blurring of indirect identifiers such as dates and regions) fundamentally eliminate the possibility of data being traced back to patients. Preprocessing such as sentence segmentation and word segmentation transforms unstructured medical text into standardized structured data, while the standardization of professional terminology ensures the consistent expression of the same concept in the corpus and avoids semantic ambiguity from interfering with model learning.

[0050] S2. Vector Database Construction: Utilizing natural language processing and deep learning technologies, preprocessed medical text is transformed into high-dimensional semantic vectors at the word and sentence levels. These high-dimensional semantic vectors and their corresponding word or sentence indices are stored in a vector database. An efficient near-nearest neighbor index structure is constructed to support high-performance semantic retrieval. High-dimensional semantic vectors are abstract semantic mappings of the original text, containing no directly interpretable textual information and making it impossible to reverse-engineer the original content using mathematical methods, thus achieving privacy protection in terms of data form. The near-nearest neighbor index structure (such as the optimized index of the Faiss library) significantly reduces the computational complexity of high-dimensional vector retrieval, enabling the query response time of the large-scale vector database to be controlled at the millisecond level, providing performance support for real-time semantic matching in subsequent model fine-tuning.

[0051] S3. Embedding Model Generation: Based on the constructed vector database, an embedding model is further trained. The embedding model is used to provide word-level or sentence-level vector representations in external fine-tuning tasks. The embedding model is loaded by the external model during the call process as its word vector initialization or enhancement module.

[0052] S4. Fine-tuning Enhancement: A standardized API interface is provided, allowing external pre-trained large language models to access the embedded model and vector database during the fine-tuning phase. A dual strategy is employed to enhance the training input, generating new training samples. These enhanced training samples are then used to fine-tune the external pre-trained large language model, resulting in a model adapted to downstream medical tasks. The standardized API interface ensures that external large models with different architectures (such as GPT and GLM) can access resources in a unified manner, avoiding redundant development for interface adaptation. The dual enhancement strategy (vector fusion and similar word replacement) introduces semantic variations from the medical field, enriching the diversity of training samples and strengthening the model's understanding of multiple expressions of professional terminology, thereby improving the model's adaptability to medical tasks without accessing the original data.

[0053] In the design process of this invention, the training efficiency of large models and the security of data use are fully considered; all vectors are in an irreversible form and cannot restore the original text information, ensuring data privacy compliance; at the same time, the probability parameters of word vector fusion and replacement can be flexibly adjusted to adapt to different model structures and training objectives.

[0054] Through the above-mentioned technical features, this invention not only realizes the secure flow of medical knowledge in an embedded manner, but also establishes an efficient, controllable, and generalizable model enhancement path, enabling large external models to significantly improve their performance, semantic understanding, and clinical reliability in medical tasks without touching the original data.

[0055] Furthermore, the dual-strategy enhancement of training input specifically includes:

[0056] With a 70% probability, the vector representations of each word in the input text are fused, and the original word vectors of the external large model and the corresponding word vectors of the embedded model are merged in a weighted manner.

[0057] With a 30% probability, several words are randomly selected from the input sentence, and the most similar words are searched through the vector database and replaced with the original words to generate new training samples;

[0058] A high proportion (70%) of vector fusion ensures the effective injection of medical domain knowledge while avoiding over-coverage of the original capabilities of external models (such as general language understanding); a low proportion (30%) of similar word replacement introduces reasonable variants (such as replacing "coronary artery angiography" with "coronary angiography") to train the model to recognize synonyms of professional terms, reducing comprehension bias caused by differences in expression.

[0059] Furthermore, in step S1, the medical record data consists of real inpatient medical records covering a ten-year period with a sample size of over 30,000. The ten-year span of medical record data includes information on the evolution of medical technology (such as changes in terminology from traditional surgery to minimally invasive surgery), enabling the model to understand medical expressions at different times. The sample size of over 30,000 ensures the statistical significance of the data, reduces knowledge bias caused by small samples (such as insufficient semantic understanding of rare heart diseases), and improves the model's universality in real clinical scenarios.

[0060] Furthermore, in step S3, the preprocessed medical text is pre-trained using the Word2Vec model to obtain high-dimensional semantic vectors; the vector database uses the Faiss library to construct an efficient approximate nearest neighbor index structure; the Word2Vec model performs excellently in semantic association learning of medical texts and can effectively capture subtle semantic differences between professional terms (such as "angina pectoris" and "myocardial infarction"); the index structure of the Faiss library significantly improves vector retrieval efficiency, meets the real-time query requirements in large-scale data scenarios, and reduces the computational resource consumption of the system.

[0061] Furthermore, in step S3, the embedding model is trained in a self-supervised manner, and the training task includes semantic similarity. Furthermore, in step S4, the external pre-trained large language model is at least one of GPT, GLM, and DeepSeek-Med.

[0062] Downstream medical tasks include semantic retrieval, medical question answering, intelligent decision support, and diagnostic reasoning related to heart diseases.

[0063] Furthermore, in step S4, the weighting method is either weighted average, gating, or a learning-based combination. Weighted average is suitable for simple scenarios and is computationally efficient. Gating mechanisms (such as dynamically adjusting weights through the sigmoid function) can adaptively allocate the fusion ratio according to the characteristics of the input text (such as general statements vs. technical terms). Learning-based combination (such as learning weights through a neural network) can automatically optimize the fusion strategy during training and adapt to complex tasks (such as multi-factor association analysis in diagnostic reasoning). These three methods cover application needs from simple to complex, enhancing the practical value of the technology.

[0064] Furthermore, in step S4, when searching for the most similar words through the vector database, the search is based on cosine similarity. Cosine similarity can effectively measure the directional consistency between high-dimensional vectors and performs excellently in judging the semantic similarity of medical terms (e.g., the vector directions of "elevated myocardial enzymes" and "myocardial injury" are highly consistent), which improves the accuracy of similar word replacement and ensures the clinical rationality of the enhanced sample.

[0065] This invention also provides a method and system for achieving medical knowledge sharing and performance enhancement of large language models under the premise of strictly protecting medical data privacy. It is particularly suitable for artificial intelligence model training and application in the cardiovascular field. The method is based on more than 30,000 real inpatient medical records from cardiology and cardiac surgery, covering structured and unstructured text data such as admission records, discharge summaries, surgical records, and ward round records, which have high clinical representativeness and medical professional depth.

[0066] At this critical stage of AI model development, data remains the core resource for enhancing model capabilities. However, medical data involves highly sensitive personal health information, and its transmission, sharing, and cross-institutional use are subject to various restrictions, including policy, ethics, and compliance. Especially in the high-risk and highly complex clinical field of cardiovascular disease, data quality requirements are high, annotation costs are high, and privacy protection pressures are heavy. Traditional data sharing and model training methods are difficult to adapt to the needs of industry development.

[0067] To address this issue, we designed an innovative “vectorized sharing” method. This method uses natural language processing technology and deep learning algorithms to transform raw cardiovascular medical record data into high-dimensional semantic vectors and construct a high-performance vector database. All medical record texts are anonymized during the transformation process to ensure that the original data cannot be restored, while retaining deep knowledge information such as medical logic, semantic structure, and treatment patterns.

[0068] Based on this vector database, relevant model parameters (such as embedding models, knowledge encoders, etc.) can be made available to external medical large-scale language model vendors for them to call and integrate. Without transmitting any original medical data, third-party models can access the vectorized representation of medical record knowledge through the interface to achieve high-quality semantic retrieval, medical question answering, intelligent decision support and other functions. This mechanism effectively breaks the dilemma of "data cannot be moved and models are not accurate enough" in the training process of medical large-scale models, and provides a feasible solution of "data not leaving the hospital and knowledge being shared".

[0069] Reference Figure 2 , Figure 2This invention demonstrates the complete technical process of the vectorized knowledge transfer method for medical big data proposed in this paper. The system input is anonymized medical text data (such as medical records, examination reports, etc.), which undergoes operations such as sentence segmentation, word segmentation, and terminology standardization through a data preprocessing module. The preprocessed text is input into a Word2Vec model to construct and obtain a searchable word vector database and embedding model parameters. In the knowledge transfer stage, the system enhances the fine-tuning process of the external large language model through a dual strategy: 70% probability embedding fusion: weighted fusion of the original word vectors of the large model with the medical professional vectors of the embedding model to enhance medical semantic understanding; 30% probability semantic replacement: replacing keywords in the input text with similar words retrieved from the vector database to improve data diversity. The fine-tuned large language model (such as GPT, GLM, etc.) integrates classification and linear transformation modules at its output end, which can be adapted to downstream medical tasks (such as question answering, diagnostic reasoning). The entire process ensures that the original medical data "does not leave the domain," and only improves the model's capabilities through vectorized knowledge transfer, combining security, efficiency, and scalability.

[0070] As the application of large language models in the medical field continues to deepen, the professionalism and generalization ability of the models are increasingly becoming key factors restricting their clinical applicability. Especially in highly complex specialty scenarios such as cardiovascular diseases, the improvement of model performance urgently depends on high-quality, real, and long-term accumulated medical record data. However, due to the protection of patient privacy, legal and regulatory constraints, and the objective existence of data silos, the traditional method of training models by sharing original medical data is difficult to promote on a large scale in practice.

[0071] Based on the aforementioned background, this invention proposes an innovative solution that combines the semantic embedding principle in natural language processing with the evolutionary path of large model fine-tuning mechanisms. Semantic vector models (such as Word2Vec and BERT) have been widely used to represent the deep semantic structure of words or sentences in a language. Essentially, they map text into a high-dimensional vector space, making semantically similar words closer together in the space. This invention utilizes this principle to transform tens of thousands of high-quality medical records from cardiology and cardiac surgery departments accumulated over ten years in a hospital into a searchable, computable, and irreversible medical semantic vector database. It also trains an embedding model to preserve the structure and transferability of medical knowledge while protecting the original data.

[0072] During the fine-tuning phase, the word vector fusion mechanism and semantic substitution mechanism proposed in this invention are based on the theoretical foundation of data augmentation and knowledge transfer. The fusion mechanism enables external large models to more effectively receive supplementary expressions of medical semantics, while the semantic substitution mechanism improves the model's generalization ability and anti-interference ability for medical language through semantic preservation and surface perturbation. This "soft injection" approach not only conforms to the training rules of large models, but also reduces the integration cost between different model frameworks.

[0073] Therefore, this invention is reasonable both theoretically and practically, and can effectively release the value of medical knowledge without touching the original data, thus promoting the evolution of general large models into professional medical large models.

[0074] A key innovation of this invention lies in proposing a novel data utilization method and sharing concept, aiming to break through the core dilemma of "difficult to share and difficult to circulate" medical data. Without transmitting or disclosing the original medical data, by transforming large-scale real medical records into a semantic vector database and combining word vector fusion and semantic substitution fine-tuning mechanisms, medical big data can effectively participate in the fine-tuning process of external pre-trained large models in the form of embedded knowledge.

[0075] This approach not only protects individual / patient privacy and data security, but also opens up new paths for unlocking the value of medical data. It does not rely on traditional data exchange or federated learning frameworks, but instead establishes a semantic sharing mechanism where "data remains unchanged and knowledge is available." This mechanism has high universality and scalability, providing a safe, compliant, efficient and feasible solution for the future development of medical AI.

[0076] Therefore, this invention is not only a technological innovation, but also an important method for breaking down data silos and promoting the ecological collaboration of large medical models, which has significant strategic significance and promotional value.

[0077] The usage process of the data vectorization knowledge transfer method for protecting medical privacy provided by this invention is as follows:

[0078] This invention can be integrated into the fine-tuning process of various medical artificial intelligence models as a complete toolchain or plugin, and is particularly suitable for knowledge injection and performance enhancement of large pre-trained language models in medical scenarios.

[0079] First, users do not need to directly access the original medical data; they only need to obtain the medical vector database and its accompanying embedding model provided by this invention. The vector database is derived from more than 30,000 cardiology and cardiac surgery medical records, which have been anonymized, vectorized, and indexed. The embedding model can be used to generate word or sentence vectors with medical semantics.

[0080] When fine-tuning large pre-trained language models (such as GPT, GLM, DeepSeek, etc.), users can follow these steps:

[0081] 1. Load the embedded model of this invention as an auxiliary module and connect it in parallel with the word vector layer of the large model;

[0082] 2. Two strategies are used when constructing training samples:

[0083] One approach is to fuse each word vector in the training text with the original word vectors from the large model and the vectors in the embedding model (e.g., by weighted averaging) with a 70% probability, thereby injecting medical semantics.

[0084] Secondly, with a 30% probability, some keywords in the training text are replaced with the most similar words in the vector database to generate equivalent but more diverse medical expressions, thereby improving robustness.

[0085] 3. The fine-tuning training process remains unchanged, only introducing medical knowledge through the above-mentioned input enhancement mechanism, without modifying the underlying architecture of the large model;

[0086] 4. After fine-tuning, the model will have a stronger ability to understand medical semantics, making it suitable for tasks such as medical question answering, treatment suggestions, and case generation;

[0087] This approach allows external model developers to fully leverage deep medical semantic knowledge without accessing the original medical data, thereby enhancing the model's professional capabilities and application value in the medical field. This method boasts advantages such as clear interfaces, convenient deployment, and strong compatibility with existing fine-tuning systems, making it suitable for widespread use and integration by hospitals, health service organizations, research institutions, and large model companies.

[0088] Example 2

[0089] A data vectorization knowledge transfer system for protecting medical privacy, used to implement the data vectorization knowledge transfer method for protecting medical privacy as described above, includes:

[0090] Corpus Acquisition and Preprocessing Module: This module collects inpatient medical record data from the cardiology and cardiac surgery departments and performs anonymization on the data. It then performs sentence segmentation, word segmentation, stop word removal, and professional terminology standardization on the anonymized medical record data to construct a unified corpus.

[0091] Vector Database and Embedding Model Building Module: Used to convert preprocessed medical text into high-dimensional semantic vectors and build a vector database, as well as train a medical-specific embedding model;

[0092] Model fine-tuning enhancement module: used to integrate the embedded model with the external pre-trained large language model, and to enhance the training input and fine-tune the external pre-trained large language model using a dual strategy;

[0093] A standardized API interface module is provided for external pre-trained large language models to call vector databases and embedding models.

[0094] Example 3

[0095] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the data vectorization knowledge transfer method for protecting medical privacy as described above.

[0096] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.

Claims

1. A data vectorization knowledge transfer method for protecting medical privacy, characterized in that, Includes the following steps: S1. Corpus Acquisition and Preprocessing: Collect inpatient medical record data from the cardiology and cardiac surgery departments. The medical record data includes at least one of the following: admission record, discharge summary, surgical record, ward round record, examination report, nursing report, and laboratory report. Desensitize the medical record data to remove sensitive personal information and eliminate the risk of implicit identification. The anonymized medical record data was processed by sentence segmentation, word segmentation, stop word removal, and standardization of professional terminology to construct a unified corpus. S2. Vector Database Construction: Using natural language processing and deep learning technologies, the preprocessed medical text is transformed into high-dimensional semantic vectors at the word and sentence levels. The high-dimensional semantic vectors and their corresponding word or sentence indexes are stored in the vector database to construct an efficient near nearest neighbor index structure to support high-performance semantic retrieval. S3. Embedded Model Generation: Based on the constructed vector database, an embedded model is further trained. The embedded model is used to provide word-level or sentence-level vector representations in external fine-tuning tasks. The embedded model is loaded by the external model during the call process as its word vector initialization or enhancement module. S4. Fine-tuning and Enhancement: Provides a standardized API interface for external pre-trained large language models to access the embedded model and vector database through the API interface during the fine-tuning stage; during the fine-tuning process, a dual strategy is adopted to enhance the training input to generate new training samples, and the enhanced training samples are used to fine-tune the external pre-trained large language model to obtain a model adapted to downstream medical tasks.

2. The data vectorization knowledge transfer method according to claim 1, characterized in that, The dual-strategy enhancement of training input specifically includes: With a probability of 70%, the vector representations of each word in the input text are fused, and the original word vectors of the external large model and the corresponding word vectors of the embedded model are merged in a weighted manner. With a 30% probability, several words are randomly selected from the input sentence, and the most similar words are searched through the vector database and replaced with the original words to generate new training samples.

3. The data vectorization knowledge transfer method according to claim 1, characterized in that, In step S1, the medical record data is real inpatient medical record data covering a ten-year period with a sample size of more than 30,000.

4. The data vectorization knowledge transfer method according to claim 1, characterized in that, In step S3, the Word2Vec model is used to pre-train the preprocessed medical text to obtain high-dimensional semantic vectors; the vector database uses the Faiss library to construct an efficient approximate nearest neighbor index structure.

5. The data vectorization knowledge transfer method according to claim 1, characterized in that, In step S3, the embedding model is trained in a self-supervised manner, and the training tasks include semantic similarity judgment and context prediction.

6. The data vectorization knowledge transfer method according to claim 1, characterized in that, In step S4, the external pre-trained large language model is at least one of GPT, GLM, and DeepSeek-Med; The downstream medical tasks include semantic retrieval, medical question answering, intelligent decision support, and diagnostic reasoning related to heart diseases.

7. The data vectorization knowledge transfer method according to claim 2, characterized in that, In step S4, the weighting method is a weighted average, a gated, or a learning combination.

8. The data vectorization knowledge transfer method according to claim 1, characterized in that, In step S4, when searching for the most similar word through the vector database, the search is based on cosine similarity.

9. A data vectorization knowledge transfer system for protecting medical privacy, used to implement the data vectorization knowledge transfer method according to any one of claims 1-8, characterized in that, include: Corpus Acquisition and Preprocessing Module: Used to collect inpatient medical record data from the cardiology and cardiac surgery departments and to perform anonymization processing on it; The anonymized medical record data was processed by sentence segmentation, word segmentation, stop word removal, and standardization of professional terminology to construct a unified corpus. Vector Database and Embedding Model Building Module: Used to convert preprocessed medical text into high-dimensional semantic vectors and build a vector database, as well as train a medical-specific embedding model; Model fine-tuning and enhancement module: used to integrate the embedded model with the external pre-trained large language model, and to enhance the training input and fine-tune the external pre-trained large language model using a dual strategy; A standardized API interface module is provided for external pre-trained large language models to call the vector database and embedding model.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the data vectorization knowledge transfer method according to any one of claims 1-8.