A FLAT-based electronic medical record data desensitization method and system

CN115438379BActive Publication Date: 2026-09-25SHAN DONG MSUN HEALTH TECH GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211116144.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2026-09-25
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

[0010]因此,现有的电子病历数据脱敏方案存在数据样本局限性大、口语化实体识别困难、准确率不高、模型推理速度慢的问题,需要对其进行进一步的研究

Benefits of technology

本发明采用以FLAT-CRF模型为主的实体识别方案,利用FLAT技术将静态词向量和字向量融合起来,获得了比传统类主流模型准确率更高的准确率,同时未应用预训练模型做特征提取,保证了较好的推理速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438379B_ABST
    Figure CN115438379B_ABST
Patent Text Reader

Abstract

The application provides an electronic medical record data desensitization method and system based on FLAT, relates to the technical field of data desensitization, collects electronic medical record text data, carries out data generalization and knowledge embedding processing on the text data to obtain a text segment sequence sample set; the text segment sequence sample set is used for training an entity recognition model based on FLAT and CRF; the electronic medical record text to be desensitized is input into the trained entity recognition model to obtain sensitive entities and entity categories of the electronic medical record; the sensitive entities are subjected to specific desensitization processing according to the entity categories; the entity recognition scheme mainly using the FLAT-CRF model is adopted, the generalization mode of random replacement of the same type of entity is adopted for the labeled entity to carry out data enhancement, the representation of the entity is simultaneously added in the word vector and the word vector to carry out information embedding, the classified desensitization processing is carried out on the recognized entity, and the accuracy and reasoning speed of data desensitization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data anonymization technology, and particularly relates to a method and system for anonymizing electronic medical record data based on FLAT. Background Technology

[0002] With the widespread adoption of electronic medical information technology, electronic medical records have become an essential method for hospitals to record medical information. Analysis of data within electronic medical records is crucial for promoting intelligent healthcare services, improving service quality, and reducing response time. Due to patient privacy requirements, patient-related information must be anonymized before use, including information such as name, date, address, institution name, contact information, and various important serial numbers.

[0003] Due to the inherent structural and geographically specific characteristics of electronic medical records, data anonymization is a very challenging task.

[0004] (1) In electronic medical records, there is little structured text content, and its structure is also subject to random changes. It is often identified by human rules. In addition, the extraction of sensitive information from a large amount of unstructured text is the main difficulty.

[0005] The extraction of sensitive information from unstructured text falls under the field of named entity recognition (NER). Accuracy and inference speed are key aspects of NER applications. With the rapid development of neural networks, LSTM-CRF, GRU-CRF, and IDCNN-CRF models based on static word vectors have gradually become the mainstream frameworks for NER. With the development of dynamic word vector technology, pre-trained models based on the Transformer framework have become mainstream. After simple parameter tuning, they can achieve higher accuracy than previous network models, but their large number of parameters results in slower inference speeds.

[0006] (2) The sources of data are often concentrated in specific regions. For example, most of the data comes from a certain province. Therefore, the place names and organization names in the collected data have strong regional characteristics.

[0007] (3) The data is collected from the last 2 or 3 years, so the data only appears in the last 2 or 3 years.

[0008] (4) The frequency of surnames in the data varies greatly depending on the population they represent.

[0009] In addition, addresses in electronic medical records often use colloquial forms such as abbreviations and shortened forms. For example, Ningjin County, Xingtai City, Hebei Province is often written as "Hebei Xingtai Ningjin". Similarly, dates, addresses, organizational structures, etc., often show similar issues.

[0010] Therefore, existing electronic medical record data anonymization schemes suffer from problems such as limited data samples, difficulty in recognizing colloquial entities, low accuracy, and slow model inference speed, requiring further research. Summary of the Invention

[0011] To overcome the shortcomings of the prior art, this invention provides a FLAT-based method and system for desensitizing electronic medical record data. It adopts an entity recognition scheme based on the FLAT-CRF model, performs data augmentation by randomly replacing similar entities with other generalized methods for labeled entities, and adds entity representations to both character vectors and word vectors for information embedding. The identified entities are desensitized by replacing them with special characters, thereby improving the accuracy of data desensitization and the speed of inference.

[0012] To achieve the above objectives, the present invention provides the following technical solution: Collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data, and obtain a text fragment sequence sample set; The entity recognition model based on FLAT and CRF was trained using a set of text fragment sequence samples. The electronic medical record text to be desensitized is input into a trained entity recognition model for reasoning to obtain the sensitive entities and entity types in the electronic medical record. Based on the type of entity, specific desensitization processing is performed on sensitive entities.

[0013] Furthermore, the term "text fragment" is a collective term for characters and words.

[0014] Furthermore, the specific steps to obtain the sample set are as follows: The text is segmented into sentences based on special characters, punctuation marks, and the set maximum sentence length; Manually annotate entities and entity types in sentences, and perform data generalization processing on the annotated entities based on entity types; Construct character vectors and word vectors with added surnames and addresses as knowledge embedding representations; The extracted sentences are segmented into words to obtain a sequence of text fragments for each sentence. The text fragments and their position information together constitute the Flat-lattice data structure unit required by the model.

[0015] The characters and words in the text fragment sequence are vectorized into character vectors and word vectors, respectively, to obtain the text fragment sequence matrix for each sentence; Construct a relative position encoding matrix for the text fragment sequence matrix; Furthermore, the data generalization includes surname generalization, address generalization, organization name generalization, and date generalization.

[0016] Furthermore, the construction of character vectors and word vectors with added surnames and addresses using knowledge embedding representations specifically involves: Based on the character vector dictionary and word vector dictionary for social sciences, construct character vectors and word vectors; Add knowledge embeddings of surnames and addresses to the constructed character vectors and word vectors.

[0017] Furthermore, the relative position encoding matrix is ​​composed of the relative position encodings of each pair of text segments in the text segment sequence matrix, and the calculation method for the relative position encoding is as follows: Dense vectors are used to simulate the relative positional relationship between two different text segments, resulting in four distances between the beginning and end of the text, the end of the text, and the end of the text. The relative position codes of the text fragment sequence are obtained by concatenating the four distances and performing a nonlinear transformation.

[0018] Furthermore, the entity recognition model includes a multi-head self-attention layer, a feedforward network layer, and a CRF layer, with the following specific steps: In the multi-head self-attention layer, multi-attention position encoding is performed on the text segment sequence matrix and the corresponding relative position encoding matrix; based on the position encoding, multi-head self-attention mechanism is used to calculate the text segment matrix. In the feedforward network layer, residual connections and normalization are performed to obtain the encoded representation of the text fragments; At the CRF layer, the highest score of the text fragment is calculated to obtain the entity label.

[0019] A second aspect of the present invention provides an electronic medical record data anonymization system based on FLAT.

[0020] A FLAT-based electronic medical record data anonymization system includes a sample set construction module, a model training module, an entity recognition module, and an anonymization processing module. The sample set construction module is configured to: collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data, and obtain a text fragment sequence sample set; The model training module is configured to train the entity recognition model based on FLAT and CRF using a set of text fragment sequence samples. The entity recognition module is configured to input the electronic medical record text to be desensitized into the trained entity recognition model to obtain the sensitive entities and entity types of the electronic medical record. The desensitization module is configured to perform specific desensitization processing on sensitive entities based on the entity type.

[0021] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a FLAT-based electronic medical record data desensitization method as described in the first aspect of the present invention.

[0022] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a FLAT-based electronic medical record data desensitization method as described in the first aspect of the present invention.

[0023] The above one or more technical solutions have the following beneficial effects: This invention employs an entity recognition scheme based on the FLAT-CRF model. By using FLAT technology to fuse static word vectors and character vectors, it achieves a higher accuracy rate than traditional mainstream models. At the same time, it does not use a pre-trained model for feature extraction, thus ensuring better inference speed.

[0024] To meet the model's requirements for recognizing colloquial entities, this invention adds knowledge embeddings of surnames and addresses to character vectors and word vectors, enabling the model to automatically learn the colloquial representation structure of entities in the corpus and improve the model's accuracy.

[0025] To address the limitations of geographical and temporal sources in the corpus, data augmentation is achieved by using information such as surnames, addresses, and dates, and employing a generalization approach that randomly replaces similar entities. This reduces the limitations of the data samples, significantly mitigates model overfitting, and ensures the model's usability.

[0026] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0027] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0028] Figure 1 This is a flowchart of the method in the first embodiment.

[0029] Figure 2 This is a data structure diagram of Flat-Lattice in the first embodiment.

[0030] Figure 3 This is a structural diagram of the encoder in the first embodiment.

[0031] Figure 4 This is a system structure diagram of the second embodiment. Detailed Implementation

[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0033] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0034] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0035] Example 1 This embodiment discloses a FLAT-based method for de-identifying electronic medical record data; like Figure 1 As shown, a FLAT-based method for desensitizing electronic medical record data includes: Step S1: Collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data to obtain a text fragment sequence sample set; The text fragment sequence sample set consists of a text fragment sequence matrix and a relative position encoding matrix. The specific steps to obtain the text fragment sequence sample set are as follows: Step S1-1: Segment the text according to special characters, punctuation marks, and the set maximum sentence length; The collected electronic medical record text data is segmented into sentences based on special characters (\n, \r, etc.); the complex, varied and small amount of structured data in the electronic medical records is treated as unstructured data and the structural symbols (\n, \r, spaces, etc.) are treated as specific characters.

[0036] Sentences are truncated according to the set maximum sentence length. If an entity is marked in a sentence longer than the maximum sentence length, and there are entity marks outside the maximum sentence length, the sentence is truncated at the punctuation marks within the maximum sentence length, as follows: Rule 1: When a sentence has more than 90 characters, if it encounters a punctuation mark (",", ",", ".", ";", ";"), it will be truncated at the punctuation mark, making the sentence before the truncation mark a new sentence, and the remaining sentence will continue to be truncated according to the rules.

[0037] Rule 2: When a sentence has more than 120 characters, if a punctuation mark (" ", "、") is encountered, the sentence is truncated at the punctuation mark, making the sentence before the truncation mark a new sentence, and the remaining sentence continues to be truncated according to the rules.

[0038] Rule 3: If the number of characters in a sentence is greater than 150, and the character at position 150 is not the target entity, then the sentence is truncated at that position; if the character at position 150 is part of the target entity, then the sentence is split at the punctuation mark closest to the target entity.

[0039] For example, in the sentence “(2) The patient underwent an ultrasound examination on July 19, 2019, which revealed a liver mass. A contrast-enhanced CT scan suggested liver cancer. On July 24, 2019, the patient underwent a liver biopsy at Shandong Cancer Hospital. The pathology report showed intrahepatic cholangiocarcinoma (pathology number: 2019-511880). On July 29, 2019, and September 11, 2019, the patient received 'hepatic artery chemoembolization' at the interventional radiology department of the hospital. The specific medication was 'oxaliplatin 150mg + gemcitabine 1.4g + iodized oil 12ml and 15ml'. Two months ago, ascites formed. One month ago, the patient received abdominal catheterization for ascites drainage at Shandong Cancer Hospital. The patient then drained the ascites at home.”, the comma before “specific medication” conforms to rule 1, so a clause should be formed at that comma.

[0040] Step S1-2: Manually annotate the entities and entity types in the sentences, and perform data generalization processing on the annotated entities based on the entity types; After entity labeling is completed, data generalization is performed on the labeled entities based on entity type, namely last name, address, organization name, and date, to achieve data augmentation, including: Surname generalization 1) Prepare a relatively complete surname dictionary, extract the individual characters from the dictionary to form a surname dictionary; some words in the surname dictionary have a length of 1, but the length of each character in the surname dictionary is 1.

[0041] 2) For the names in the annotated sentences, extract the surname words through code, and then randomly select surnames from the surname dictionary to replace them; for example, in the sentence "Li Qiang", after extracting "Li", randomly select generalized words from the surname dictionary, such as "Wang", "Zhang" and "Ouyang", and combine them into "Wang Qiang", "Zhang Qiang" and "Ouyang Qiang".

[0042] Address and organization name generalization 1) Prepare a relatively complete address dictionary; according to the administrative divisions of the People's Republic of China, the address dictionary generally consists of five levels of units: provincial (province, municipality, autonomous region), prefecture (city), county (district), township (town, street), and village (village committee, neighborhood committee). Clean up the non-address representations, for example: ***town committee, remove "committee"; ***neighborhood committee, remove "neighborhood committee", etc. Here, only the provincial, prefecture, and county level dictionary representations are selected.

[0043] 2) For provincial, prefectural, and county-level dictionaries, extract the individual characters to form the corresponding provincial, prefectural, and county-level dictionaries; the character length of the provincial, prefectural, and county-level dictionaries is 1.

[0044] 3) For addresses and organization names in the labeled statements, extract the unit word representations at each level through code, and then randomly select addresses from the corresponding unit dictionary in the address dictionary for replacement; for example, "Zhao County, Shijiazhuang City, Hebei Province" in the sentence can be generalized to "Cao County, Jinan City, Shandong Province" and "Furong District, Changsha City, Hunan Province"; "Heze Maternal and Child Health Hospital" can be generalized to "Zhangye Maternal and Child Health Hospital" and "Cao County Maternal and Child Health Hospital".

[0045] Date generalization For dates in the tagged statements, the year is extracted using code, and then randomly replaced with the years from the last 5 years; the date format remains unchanged, such as Chinese characters and numbers. For example, "June 2021" in the sentence can be generalized to "June 2022", "June 2025", and "June 2026"; "July 2021" in the sentence can be generalized to "July 2023", "July 2025", etc.

[0046] Steps S1-3: Construct character vectors and word vectors with added surnames and addresses as knowledge embedding representations. The steps are as follows: 1) Prepare a dictionary of character vectors and word vectors for social sciences; since the entities to be identified fall within the scope of social sciences, using a social science vector dictionary can obtain more accurate representations. This embodiment uses 50-dimensional character vectors and word vectors. Here, the character vectors use... c In other words, word vectors are used w express.

[0047] ,in, k It indicates the position of a specific word in a sentence.

[0048] ,in, k It indicates the position of a specific word in a sentence.

[0049] 2) The character vector representation after adding knowledge embeddings of surname and address (provincial, prefectural, and county levels) is as follows:

[0050] in, The value is 0 / 1, representing a word k Has it appeared in a surname dictionary? The value is 0 / 1, representing a word k Has it appeared in the provincial address dictionary? The value is 0 / 1, indicating whether word k has appeared in the local address dictionary; The value is 0 / 1, indicating whether the word k has appeared in the county-level address dictionary.

[0051] 3) The word vector representation after adding knowledge embeddings for surname and address (provincial, prefectural, and county levels) is as follows:

[0052] in, The value is 0 / 1, representing a word (character). k Has it appeared in a dictionary of surnames? The value is 0 / 1, indicating whether the word k has appeared in the provincial address dictionary; The value is 0 / 1, indicating whether word k has appeared in the local address dictionary; The value is 0 / 1, representing words k Has it appeared in the county-level address dictionary?

[0053] Steps S1-4: Segment the extracted sentence into words, and vectorize the distribution of characters and words in the sentence to obtain the text fragment sequence matrix of the sentence; 1) The extracted sentence is segmented into words, and the sentence is represented by concatenating characters and words to obtain a text fragment representation of the sentence.

[0054]

[0055] in, The first sentence i Vector representation of each character, The first sentence j Vector representation of each word Indicates the first in the corpus u One sentence.

[0056] Note that here, words and phrases are uniformly described and represented using text fragments. Sentences can then be directly described using these text fragments, as shown in the following formula:

[0057] in, This represents the k-th text segment of the u-th sentence in the corpus.

[0058] 2) Sentence fragment sequence It can be expanded into a flat-lattice data structure, such as Figure 2 As shown, flat-lattice data is a collection of sequence fragments, each consisting of a token, a head, and a tail. The token represents a character or word in the text, while the head and tail represent the positions of the first and last characters in the token within the original sequence. It can be seen that when the token is a single character, its head and tail are the same.

[0059] Finally, each sentence can be vectorized into a sequence matrix of text fragments. , This refers to the dimension of the word vectors and character vectors in steps S1-4, which is 54 here.

[0060] Steps S1-5: Construct the relative position encoding matrix of the text segment sequence matrix; The relative position encoding matrix is ​​composed of the relative position encodings of each pair of text segments in the text segment sequence matrix. The calculation method for the relative position encoding is as follows: 1) In a lattice structure, two different text segments (words / phrases) and There are three types of relationships between them: intersection, containment, and disjointness. Dense vectors are used to simulate these relationships: Among them, head[ i ]、tail[ i ] respectively represent The head and tail positions, express head and The distance between their heads, the others , , It has a similar meaning.

[0061] 2) After concatenating the four distances, a nonlinear transformation is performed to obtain the relative position code of the text segment sequence. The specific formula is as follows:

[0062] in, These are learnable parameters. This indicates a splicing operation. The calculation formula is:

[0063] here,d express , , , ; k is the vector dimension of the text segment, and its value ranges from [0- ].

[0064] Step S2: Train the entity recognition model based on FLAT and CRF using the sentence sample set; The entity recognition model is composed of several stacked encoders, and the structure of each encoder layer is as follows: Figure 3 As shown, it mainly consists of a multi-head self-attention layer (Multih head self The encoder consists of an attention layer, a feed forward network layer, and residual connections and layer normalization throughout. To ensure inference efficiency, only one Transformer layer is used as the encoder.

[0065] The specific steps for each encoder layer are as follows: 1) In the multi-head self-attention layer, the relative position encoding is performed using the head and tail information inherent in the sequence segments in the flat-lattice data structure; based on the position encoding, the multi-head self-attention mechanism is calculated using the sequence vector; The training statements with a fixed batch size are input into the encoder, as shown in the following formula:

[0066] in, , , These are all trainable and learnable parameters. It is the matrix representation of the j-th sentence in the training statements. , It is the relative position encoding of sentence i and sentence j in step (6). It is a variant of the self-attention module, approximately equivalent to the following representation:

[0067] Will Replacing A in the formula below yields the calculation of the attention mechanism for batch statements:

[0068] In this model , , The inputs are the query vector, key vector, and value vector, respectively, containing... This indicates that this module is performing self-attention.

[0069] The different attention results obtained by the multi-head attention mechanism are concatenated to obtain the final output sequence vector.

[0070] 2) In the feedforward network layer, the output of the multi-head self-attention layer is residually connected and normalized to obtain the encoded representation of the text segment; 3) At the CRF layer, calculate the highest score of the text fragment to obtain the entity label.

[0071] Step S3: Input the electronic medical record text to be desensitized into the trained entity recognition model to obtain the sensitive entities and entity types of the electronic medical record; Step S4: Based on the entity type, perform specific desensitization processing on sensitive entities. This uses special character replacement, specifically: using "#person_name#" to replace a person's name; using "#date#" to replace a specific date; using "#location#" to replace an address; using "#organization#" to replace an organization name; using "#telephone#" to replace contact information; and using "#ID#" to replace various important numbers. The "#" symbol is a special character set to facilitate entity extraction.

[0072] The data used for anonymization in this embodiment comes from multiple general hospitals, and the labeled data was selected through keyword search. Emphasis was placed on the diversity of the selected labeled data; the number of entities in each category was determined based on the richness of the entity content. The number of entities in the selected samples is shown in Table 1. Table 1. Entity Count of Selected Samples

[0073] After data generalization, the resulting data structure is shown in Table 2: Table 2. Composition of Experimental Data

[0074] The effectiveness of several models in data anonymization was compared through experiments, and the specific data are shown in Table 3: Table 3. Accuracy Comparison of Different Schemes

[0075] As shown in Table 1, the FLAT model has a 3% accuracy improvement compared to the BILSTM model; after using FLAT as the baseline model, adding knowledge embeddings to character vectors and word vectors resulted in an additional 1% accuracy improvement.

[0076] After using data generalization and knowledge embedding, BILSTM+CRF also showed an accuracy improvement of about 4%, demonstrating the effectiveness of data generalization and knowledge embedding.

[0077] Example 2 This embodiment discloses an electronic medical record data anonymization system based on FLAT; like Figure 4 As shown, an electronic medical record data anonymization system based on FLAT includes a sample set construction module, a model training module, an entity recognition module, and an anonymization processing module. The sample set construction module is configured to: collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data, and obtain a text fragment sequence sample set; The model training module is configured to train the entity recognition model based on FLAT and CRF using a set of text fragment sequence samples. The entity recognition module is configured to input the electronic medical record text to be desensitized into the trained entity recognition model to obtain the sensitive entities and entity types of the electronic medical record. The desensitization module is configured to perform specific desensitization processing on sensitive entities based on the entity type.

[0078] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0079] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a FLAT-based electronic medical record data desensitization method as described in Embodiment 1 of this disclosure.

[0080] Example 4 The purpose of this embodiment is to provide an electronic device.

[0081] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in a FLAT-based electronic medical record data desensitization method as described in Embodiment 1 of this disclosure.

[0082] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for desensitizing electronic medical record data based on FLAT, characterized in that, include: Collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data to obtain a text fragment sequence sample set. The specific steps to obtain the sample set are as follows: The text is segmented into sentences based on special characters, punctuation marks, and a set maximum sentence length; entities and their types are manually labeled in the sentences, and data generalization is performed on the labeled entities based on their types; character vectors and word vectors with added surnames and addresses are constructed; the truncated sentences are segmented into words to obtain a sequence of text fragments for each sentence, and the text fragments and their positional information together constitute the Flat-lattice data structure unit required by the model; the characters and words in the text fragment sequence are vectorized into character vectors and word vectors to obtain a text fragment sequence matrix for each sentence; Construct a relative position encoding matrix for the text fragment sequence matrix; The entity recognition model based on FLAT and CRF was trained using a set of text fragment sequence samples. The electronic medical record text to be desensitized is input into a trained entity recognition model for reasoning to obtain the sensitive entities and entity types in the electronic medical record. Based on the type of entity, specific desensitization processing is performed on sensitive entities.

2. The method for desensitizing electronic medical record data based on FLAT as described in claim 1, characterized in that, The text fragments referred to are a collective term for characters and words.

3. The method for desensitizing electronic medical record data based on FLAT as described in claim 1, characterized in that, The data generalization includes surname generalization, address generalization, organization name generalization, and date generalization.

4. The method for de-identifying electronic medical record data based on FLAT as described in claim 1, characterized in that, The construction of character vectors and word vectors with added surnames and addresses using knowledge embedding representations is specifically as follows: Based on the character vector dictionary and word vector dictionary for social sciences, construct character vectors and word vectors; Add knowledge embeddings of surnames and addresses to the constructed character vectors and word vectors.

5. The method for desensitizing electronic medical record data based on FLAT as described in claim 1, characterized in that, The relative position encoding matrix is ​​composed of the relative position encodings of each pair of text segments in the text segment sequence matrix. The calculation method for the relative position encoding is as follows: Dense vectors are used to simulate the relative positional relationship between two different text segments, resulting in four distances between the beginning and end of the text, the end of the text, and the end of the text. The relative position codes of the text fragment sequence are obtained by concatenating the four distances and performing a nonlinear transformation.

6. The method for desensitizing electronic medical record data based on FLAT as described in claim 1, characterized in that, The entity recognition model includes a multi-head self-attention layer, a feedforward network layer, and a CRF layer. The specific steps are as follows: In the multi-head self-attention layer, multi-attention position encoding is performed on the text segment sequence matrix and the corresponding relative position encoding matrix; Based on positional encoding, a multi-head self-attention mechanism is used to compute the text fragment matrix. In the feedforward network layer, residual connections and normalization are performed to obtain the encoded representation of the text fragments; At the CRF layer, the highest score of the text fragment is calculated to obtain the entity label.

7. A FLAT-based electronic medical record data anonymization system, characterized in that, It includes a sample set construction module, a model training module, an entity recognition module, and a desensitization processing module; The sample set construction module is configured to: collect electronic medical record text data, perform data generalization and knowledge embedding processing on the text data, and obtain a text fragment sequence sample set. The specific steps for obtaining the sample set are as follows: The text is segmented into sentences based on special characters, punctuation marks, and the set maximum sentence length; The system manually annotates entities and their types in sentences, performs data generalization on the annotated entities based on their types, constructs character vectors and word vectors with added surnames and addresses as knowledge embeddings, segments the truncated sentences to obtain a sequence of text fragments for each sentence, and combines the text fragments with their positional information to form the Flat-lattice data structure unit required by the model. The characters and words in the text fragment sequence are then vectorized into character vectors and word vectors to obtain a text fragment sequence matrix for each sentence. Construct a relative position encoding matrix for the text fragment sequence matrix; The model training module is configured to train the entity recognition model based on FLAT and CRF using a set of text fragment sequence samples. The entity recognition module is configured to: input the electronic medical record text to be desensitized into the trained entity recognition model for reasoning, and obtain the sensitive entities and entity types of the electronic medical record; The desensitization module is configured to perform specific desensitization processing on sensitive entities based on the entity type.

8. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the FLAT-based electronic medical record data desensitization method as described in any one of claims 1-6.

9. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the FLAT-based electronic medical record data desensitization method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data enhancement method and device, electronic equipment and storage medium

    CN114595327A