Medical record automatic coding method and system

Through the automatic coding method of medical records, combined with natural language processing and cross attention modules, the inefficiency and high cost of traditional manual coding methods are solved, and the consistency and accuracy of medical record coding are improved.

CN120199397APending Publication Date: 2025-06-24PEOPLES HOSPITAL OF HENAN PROV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510266158.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The traditional artificial medical record coding methods have problems such as strong technicality, high professional requirements, many repetitive operations, high labor costs, low efficiency and error-prone. Especially in large public hospitals, coding is difficult and the quality of coding personnel is uneven, resulting in poor coding consistency and accuracy.

Method used

The automatic coding method of medical records is used to process the homepage of the medical record and other parts of the medical record using natural language processing, word coding, self-attention and cross-attention modules, extract keyword information and encode it, and obtain standardized diagnostic information through the mapping table to realize automatic coding of the medical record information.

Benefits of technology

It improves the consistency, accuracy and efficiency of medical record coding, reduces labor costs and error rates, reduces the coding work burden of doctors, and supports more accurate selection of standardized diagnostic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199397A_ABST
    Figure CN120199397A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical record automatic coding, and discloses a medical record automatic coding method comprising the following steps: S1, obtaining a to-be-coded medical record, and dividing the medical record into a medical record home page and other parts of medical records; s2, inputting the home page of the medical record and other parts of medical records into a preset coding model for processing to obtain standardized diagnosis information; and S3, querying a medical record coding result corresponding to the standardized diagnosis information in a preset mapping table, and taking the medical record coding result as a medical record code of the to-be-coded medical record. According to the automatic coding method for the medical records, full-information automatic coding of the medical records is achieved, meanwhile, the consistency, accuracy and efficiency of coding are greatly improved, and the error rate and the labor cost are effectively reduced. The invention further discloses a medical record automatic coding system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic medical record coding, and particularly relates to a method and system for automatic medical record coding. Background Art

[0002] Traditional disease classification coding work requires professionals to carefully read medical record information, which includes the front page of the medical record and other parts of the medical record, such as admission records, progress notes, discharge records, operation records, examinations and tests, and pathological diagnoses. Key information of the medical record is extracted from the medical record information, combined with international disease classification coding knowledge for diagnostic classification. At the same time, the diagnosis on the front page of the medical record is corrected, and the main diagnosis is selected, and finally manual coding is performed. For example, for a medical record to be coded of a lung cancer patient, through the lung cancer diagnosis in its front page of the medical record, and then by consulting the confirmation, treatment, and operations for lung cancer in the progress notes, combined with the pathological diagnosis results, the coding is finally completed. The traditional manual medical record coding method has problems such as strong technical requirements for the work, high professional requirements, a large number of repetitive operations, high labor costs, low efficiency, and easy errors. At present, there are many discharged medical records in large public hospitals, the disease types are complex, and the coding difficulty is relatively large. The traditional manual medical record coding method also has problems such as uneven qualities of coding personnel and inconsistent cognitive understanding standards, which further lead to coding errors. Therefore, there is an urgent need for a method and system for automatic medical record coding to solve the above problems. Summary of the Invention

[0003] In order to solve the technical problems of the traditional manual medical record coding method, which has strong technical requirements for the work, high professional requirements, a large number of repetitive operations, high labor costs, low efficiency, and easy errors, the present invention provides a method and system for automatic medical record coding, realizing automatic coding of all medical record information, replacing traditional manual medical record coding, reducing the workload of manual coding, and at the same time greatly improving the consistency, accuracy, and coding efficiency of coding, effectively reducing the error rate and labor costs.

[0004] The present invention provides a method for automatic medical record coding, including the following steps:

[0005] S1. Obtain the medical record to be coded, and divide the medical record into the front page of the medical record and other parts of the medical record;

[0006] S2. Input the front page of the medical record and other parts of the medical record into a preset coding model for processing to obtain standardized diagnostic information;

[0007] S3. Query the medical record coding result corresponding to the standardized diagnostic information in a preset mapping table as the medical record coding of the medical record to be coded;

[0008] Among them, the encoding model includes a natural language processing module, a word encoding module, a self-attention module, a cross-attention module, and a word decoding module; the natural language processing module processes the front page of the medical record and other parts of the medical record respectively to obtain the front page keyword information and other parts of the medical record keyword information; the word encoding module encodes the front page keyword information and other parts of the medical record keyword information respectively to obtain a first d-dimensional vector and a second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and the n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnosis information.

[0009] Further, the natural language processing module includes a first natural language processing sub-module and a second natural language processing sub-module; the first natural language processing sub-module first performs word segmentation on the front page of the medical record, and then extracts keywords from the word segmentation result of the front page of the medical record to obtain the front page keyword information; the second natural language processing sub-module first performs word segmentation on other parts of the medical record, and then extracts keywords from the word segmentation result of other parts of the medical record to obtain other parts of the medical record keyword information.

[0010] Further, according to the order of obtaining the front page keyword information, the first natural language processing sub-module performs structured arrangement on it; according to the order of obtaining other parts of the medical record keyword information, the second natural language processing sub-module performs structured arrangement on it. The structured arranged front page keyword information and other parts of the medical record keyword information have relatively strong semantic information in different positions, so it is easier to learn the correlation between data in different dimensions.

[0011] Further, the word encoding module includes a first word encoding sub-module and a second word encoding sub-module. The first word encoding sub-module encodes the structured arranged front page keyword information to obtain a first d-dimensional vector, and the second word encoding sub-module encodes the structured arranged other parts of the medical record keyword information to obtain a second d-dimensional vector.

[0012] Further, the self-attention module includes a first self-attention sub-module and a second self-attention sub-module. The first self-attention sub-module extracts features from the first d-dimensional vector to obtain m first K-dimensional feature vectors, and the second self-attention sub-module extracts features from the second d-dimensional vector to obtain n second K-dimensional feature vectors.

[0013] Further, the cross-attention module processes the first matrix X1 composed of m first K-dimensional feature vectors and the second matrix X2 composed of n second K-dimensional feature vectors to obtain an attention score vector.

[0014] Furthermore, the cross-attention module utilizes the linear transformation matrices W K and W V to calculate the key matrix K and the value matrix V respectively;

[0015] K = X1 × W K ;

[0016] V = X1 × W V ;

[0017] The cross-attention module utilizes the linear transformation matrix W Q to calculate the query matrix Q;

[0018] Q = X2 × W Q ;

[0019] The cross-attention module calculates the attention score vector attention(Q, K, V),

[0020]

[0021] where K T represents the transpose of the key vector K, and d k is the number of columns of the key matrix K.

[0022] Furthermore, the training process of the encoding model is as follows: Using the labeled medical records as the training data of the natural language processing module to obtain a trained natural language processing module; Based on the trained natural language processing module, using the labeled medical records as the input of the entire encoding model and the corresponding standardized diagnosis information as the output of the entire encoding model to train the entire encoding model to obtain a trained encoding model.

[0023] Compared with the prior art, the present invention has the following beneficial effects: It realizes the automatic encoding of all medical record information, replaces the traditional manual medical record encoding, reduces the workload of manual encoding, and at the same time greatly improves the consistency, accuracy and encoding efficiency of encoding, effectively reducing the error rate and labor cost. The automatic medical record encoding method of the present invention comprehensively considers the front page of the medical record and other parts of the medical record, extracts features from the front page of the medical record and other parts of the medical record respectively, and then uses the cross-attention mechanism to learn the feature correlation between the two, so as to realize the feature enhancement of the front page of the medical record, which can better express the information contained in the medical record and support more accurate selection of standardized diagnosis information and medical record encoding results.

[0024] The present invention also provides a medical record automatic encoding system, including:

[0025] A medical record acquisition unit for acquiring the medical record to be encoded and dividing the medical record into the front page of the medical record and other parts of the medical record;

[0026] A standardization unit for inputting the first page of a medical record and other parts of the medical record into a preset coding model for processing to obtain standardized diagnosis information;

[0027] A medical record coding unit for querying the medical record coding result corresponding to the standardized diagnosis information in a preset mapping table as the medical record coding of the medical record to be coded;

[0028] Among them, the coding model includes a natural language processing module, a word coding module, a self-attention module, a cross-attention module, and a word decoding module; the natural language processing module processes the first page of the medical record and other parts of the medical record respectively to obtain the first page keyword information and other parts of the medical record keyword information; the word coding module encodes the first page keyword information and other parts of the medical record keyword information respectively to obtain a first d-dimensional vector and a second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and the n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnosis information.

[0029] Compared with the prior art, the present invention has the following beneficial effects: realizing the automatic coding of all medical record information, replacing the traditional manual medical record coding, reducing the workload of manual coding, and at the same time greatly improving the coding consistency, accuracy and coding efficiency, effectively reducing the error rate and labor cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a schematic flow chart of the medical record automatic coding method of the present invention;

[0031] Figure 2 is a schematic block diagram of the coding model of the present invention;

[0032] Figure 3 is a structural block diagram of the medical record automatic coding system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0033] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0034] As Figures 1 - 2 shown, a medical record automatic coding method includes the following steps:

[0035] S1. Obtain the medical record to be coded, and divide the medical record into the first page of the medical record and other parts of the medical record; the first page of the medical record includes diagnosis information, surgical information, etc.; the other parts of the medical record include other medical record information such as daily course record information, discharge record information, inspection and test result information, and pathological diagnosis result information;

[0036] S2. Input the first page of the medical record and other parts of the medical record into a preset coding model for processing to obtain standardized diagnosis information;

[0037] S3. Query the medical record coding result corresponding to the standardized diagnosis information in a preset mapping table as the medical record coding of the medical record to be coded; wherein, the mapping table adopts the "Tenth Revision of the International Statistical Classification of Diseases and Related Health Problems" (ICD-10);

[0038] Among them, the coding model includes a natural language processing module, a word coding module, a self-attention module, a cross-attention module, and a word decoding module; the natural language processing module processes the first page of the medical record and other parts of the medical record respectively to obtain the keyword information of the first page and the keyword information of other parts of the medical record; the word coding module encodes the keyword information of the first page and the keyword information of other parts of the medical record respectively to obtain a first d-dimensional vector and a second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and the n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnosis information.

[0039] The content of the medical record information is rich, and the direct analysis cost is relatively high. At present, there is a lack of effective automated analysis means. The medical record automatic coding method of this embodiment, based on the first page of the medical record, also considers other parts of the medical record (that is, the daily course record information, the discharge record information, the examination and test result information, and the pathological diagnosis result information). Compared with the first page of the medical record, other parts of the medical record have richer medical record details, and these information can provide better supplements to the first page of the medical record. At the same time, the self-attention mechanism and cross-attention mechanism considering structured information are adopted, which improves the expression ability of the information provided, and enhances the feature representation ability of the first page of the medical record. The medical record automatic coding method of this embodiment comprehensively considers the first page of the medical record and other parts of the medical record, and proposes a feature extraction scheme for structured information. By extracting features from the first page of the medical record and other parts of the medical record respectively, and then using the cross-attention mechanism to learn the feature correlation between the two, the feature enhancement of the first page of the medical record can be realized, which can better express the information contained in the medical record and support more accurate selection of standardized diagnosis information and medical record coding results.

[0040] The automatic medical record coding method of this embodiment realizes automatic coding of all medical record information, replaces traditional manual medical record coding, reduces the workload of manual coding, and greatly improves the consistency, accuracy and efficiency of coding, effectively reducing the error rate and labor costs. The coding model of this embodiment, by comprehensively considering the attention distribution of the two input information sources of the medical record homepage and other parts of the medical record, integrates the ability of the two to capture correlation features in different subspaces, so that the coding model learns the richer feature expression required for coding, solves the problem of incomplete information from a single data source, and supports more accurate standardized diagnostic information selection and medical record coding results. By comparing the coding results of the medical record automatic coding method with the coding results of the traditional manual medical record coding method, it is also possible to promptly discover the coding problems of traditional manual (doctors), thereby improving the writing quality of the medical record homepage.

[0041] On the basis of the above embodiment, the natural language processing module includes a first natural language processing submodule and a second natural language processing submodule; the first natural language processing submodule first performs word segmentation processing on the medical record homepage, and then extracts keywords from the word segmentation results of the medical record homepage to obtain the keyword information of the homepage; the second natural language processing submodule first performs word segmentation processing on other parts of the medical record, and then extracts keywords from the word segmentation results of other parts of the medical record to obtain the keyword information of other parts of the medical record. Standardized diagnostic information should be a specific professional diagnostic term, but in the actual process, due to the influence of the doctor's input habits, the diagnostic information actually presented to the medical record homepage may be a non-standard diagnosis, or even include multiple words or sentences. To this end, it is necessary to perform word segmentation processing on the medical record homepage, disassemble multiple words or sentences, divide them into multiple professional terms or common words, and then extract keyword information from them. Keyword information is usually a professional term in the word segmentation result, which contains the most essential information of the medical record diagnosis. For example: the standard diagnosis is "lung malignancy", and the doctor is used to diagnose it as "lung cancer". The first natural language processing submodule can achieve word segmentation of the medical record homepage through natural language processing algorithms, such as methods based on statistical models, and extract keyword information containing the core content of the medical record, which is the basis for the next step of processing. Non-keywords are usually non-professional vocabulary or punctuation marks, such as ";" and ":". Non-keywords are usually easy to appear in other parts of the medical record. The second natural language processing submodule can achieve word segmentation of daily medical record information, discharge record information, examination and test result information, and pathological diagnosis result information through natural language processing algorithms, such as methods based on statistical models, and extract keyword information containing the core content of the medical record, which is the basis for the next step of processing. For example: the doctor's medical record homepage writes the diagnosis as "lung cancer", and the pathological information extracted is "left upper lobe, adenocarcinoma". The diagnosis of "lung cancer" on the medical record homepage is automatically encoded as "malignant tumor of the left upper lobe of the lung" through the keyword information on the homepage, and the pathological diagnosis of "malignant tumor" on the medical record homepage is filled in as "adenocarcinoma".

[0042] Based on the above embodiments, the first natural language processing sub-module performs structured arrangement on the obtained home page keyword information in the order of obtaining it; the second natural language processing sub-module performs structured arrangement on the obtained keyword information of other parts of the medical record in the order of the structure. Since the medical record information has very distinct structural characteristics, for example, the home page of the medical record usually consists of structural information such as the main diagnosis and the main operation, prior knowledge can be used to perform structured arrangement on the above-extracted home page keyword information, which can greatly improve the learning efficiency of subsequent feature extraction, cross-attention, etc. For example, the home page keyword information includes at least one keyword one obtained from the "main diagnosis information" and at least one keyword two obtained from the "main operation information", then all keyword ones and all keyword twos are arranged in the structural order of "main diagnosis information" and "main operation information", and all keyword ones and all keyword twos are arranged in the order of obtaining. The structured arranged home page keyword information and the keyword information of other parts of the medical record have relatively strong semantic information in different positions, making it easier to learn the correlation between data in different dimensions.

[0043] Based on the above embodiments, the word encoding module includes a first word encoding sub-module and a second word encoding sub-module. In a computer system, it is impossible to directly calculate words, which is inconvenient for subsequent model processing. Therefore, the first word encoding sub-module and the second word encoding sub-module are adopted. The first word encoding sub-module and the second word encoding sub-module usually adopt model structures related to deep learning networks; the first word encoding sub-module encodes the structured arranged home page keyword information to obtain a first d-dimensional vector, that is, in the form of a d-dimensional vector of (X1, X2,..., X d ), where d takes an integer value. Assuming that each medical record home page has m items of information, the word-encoded information of the medical record home page is a matrix with m rows and d columns, where each row is an item of information. Similarly, the second word encoding sub-module encodes the structured arranged keyword information of other parts of the medical record to obtain a second d-dimensional vector, that is, in the form of a d-dimensional vector of (X1, X2,..., X d ), where d takes an integer value.

[0044] Based on the above embodiments, the self-attention module includes a first self-attention sub-module and a second self-attention sub-module. The first self-attention sub-module extracts features from the first d-dimensional vector to obtain m first K-dimensional feature vectors, and the second self-attention sub-module extracts features from the second d-dimensional vector to obtain n second K-dimensional feature vectors. Specifically, based on the first d-dimensional vector and the second d-dimensional vector obtained by the first word encoding sub-module and the second word encoding sub-module, the d-dimensional features of each vocabulary only contain its own information expression and cannot express the information in the context of its medical record, which will result in insufficient feature expression ability. To improve the feature expression ability and extract the global features on which the diagnostic standardization information depends, a self-attention mechanism is introduced. In the first self-attention sub-module, each feature in the m items of information will learn the correlation expression with the other m-1 items of features, and through the correlation expression, calculate the correlation degree of each feature with the m items (including the feature itself) of features and the corresponding first k-dimensional feature vector, that is, based on the m first k-dimensional feature vectors of the first self-attention sub-module, where different dimensional distributions learn the correlation information between features from different angles. The first k-dimensional feature vector contains not only its own information but also the correlation information of other items. Such a feature representation has the ability to extract all diagnostic information from the medical record front page and can provide the accuracy and reliability of extracting diagnostic information. Similarly, it can be known that the second k-dimensional feature vector contains not only its own information but also the correlation information of other items.

[0045] Based on the above embodiments, the cross-attention module processes the first matrix X1 composed of m first K-dimensional feature vectors and the second matrix X2 composed of n second K-dimensional feature vectors to obtain an attention score vector. Further, the cross-attention module uses the linear transformation matrix W K and W V , and calculates the key matrix K and the value matrix V respectively;

[0046] K = X1 × W K ;

[0047] V = X1 × W V ;

[0048] The cross-attention module uses the linear transformation matrix W Q , and calculates the query matrix Q;

[0049] Q = X2 × W Q ;

[0050] The cross-attention module calculates the attention score vector attention(Q, K, V),

[0051]

[0052] Among them, K T represents the transpose of the key vector K, where d k is the number of columns of the key matrix K, and W K , W V and W Q are parameter matrices to be optimized and solved in the cross-attention module.

[0053] Through the first self-attention sub-module, the feature expression of each piece of information combined with the overall information of its medical record front page can be obtained, and the ability to extract diagnostic information from the medical record front page is possessed. However, in practice, doctors may write the main diagnosis incorrectly. In this case, more accurate expression information needs to be obtained from other parts of the medical record (i.e., daily progress record information, discharge record information, examination and test result information, and pathological diagnosis result information). For this reason, a cross-attention mechanism is introduced here. By calculating the correlation degree and the corresponding score vector between the features in the medical record front page and the features of other parts of the medical record, the feature expression combined with other parts of the medical record is obtained. The cross-attention mechanism enhances the feature representation ability of the medical record front page by considering other parts of the medical record. For example, the patient's previous diagnosis was pulmonary malignant tumor, and the main treatment during this hospitalization was chemotherapy. According to the disease classification and coding rules, the main diagnosis should be coded as "chemotherapy course for tumor". However, according to the doctor's writing habit, the main diagnosis information in the medical record front page is often incorrectly written as "lung cancer". Relying only on the self-attention mechanism, the main diagnosis information can be changed to "pulmonary malignant tumor", but it is still the incorrect main diagnosis information. At this time, by introducing the cross-attention module, it is learned that there are keywords such as "chemotherapy drugs" in the progress record. Combining the coding content of the medical record front page, the feature representation that fuses the medical record front page and other parts of the medical record can finally be obtained, so as to be correctly coded as "chemotherapy course for tumor".

[0054] Based on the above embodiments, the medical record coding results known from the example query mapping table are as follows:

[0055] Standardized diagnostic information Medical record coding result Malignant tumor of duodenum C17.0 Malignant tumor of jejunum C17.1 Malignant tumor of appendix C18.1 Secondary malignant tumor of lung C78.0 Multiple myeloma C90.0 Radiation therapy course Z51.0 For chemotherapy course of tumor Z51.1

[0056] Based on the above embodiments, the labeled medical records are used as the training data of the natural language processing module to obtain a trained natural language processing module; based on the trained natural language processing module, the labeled medical records are used as the input of the entire coding model, and the corresponding standardized diagnosis information is used as the output of the entire coding model to train the entire coding model to obtain a trained coding model. Further, the coding model in this embodiment is trained step by step. Training step 1: First, the natural language processing module is learned. At this time, the input-natural language processing part of the task is executed, and a trained natural language processing module is obtained through training. Training step 2: Training from the input end to the output end. On the basis of the training in training step 1, the entire coding model from the input end to the output end is learned. The advantage of step-by-step training is that a part (natural language processing module) of the entire coding model can be learned by using step-by-step training to accelerate the training and ultimately improve the training efficiency on the entire coding model. Since the natural language processing module has been trained in training step 1, only fine-tuning is required during the training in training step 2, and the focus of the training in training step 2 is the part from word encoding to the final output. This embodiment proposes a method for training a coding model step by step, which reduces the training difficulty of the coding model and improves the model training efficiency.

[0057] As Figure 3 shown, an embodiment of the present invention further provides a medical record automatic coding system, including:

[0058] A medical record acquisition unit, configured to acquire a medical record to be coded, and divide the medical record into a front page of the medical record and other parts of the medical record;

[0059] A standardization unit, configured to input the front page of the medical record and other parts of the medical record into a preset coding model for processing to obtain standardized diagnosis information;

[0060] A medical record coding unit, configured to query a medical record coding result corresponding to the standardized diagnosis information in a preset mapping table as the medical record coding of the medical record to be coded;

[0061] Among them, the encoding model includes a natural language processing module, a word encoding module, a self-attention module, a cross-attention module, and a word decoding module; the natural language processing module processes the front page of the medical record and other parts of the medical record respectively to obtain front page keyword information and other part medical record keyword information; the word encoding module encodes the front page keyword information and other part medical record keyword information respectively to obtain a first d-dimensional vector and a second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and the n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnosis information.

[0062] It should be noted that the medical record automatic encoding system provided by the embodiments of the present invention is for implementing the above method, and its functions can be specifically referred to the above method embodiments, which will not be elaborated here.

[0063] The above embodiments are only preferred embodiments of the present invention, which are only used to explain the present invention and do not limit the scope of implementation of the present invention. For those skilled in the art of this technology, of course, other implementation manners can be easily made by means of substitution or change according to the technical content disclosed in this specification. Therefore, all changes and improvements made in the principles and process conditions of the present invention should be included within the scope of the patent application of the present invention.

Claims

1. A method for automatic coding of medical records, characterized in that: The following steps are involved: S1. Obtain the medical record to be coded, and divide the medical record into the medical record front page and other partial medical records; S2, inputting the medical record front page and other parts of the medical record into a preset coding model for processing to obtain standardized diagnosis information; S3, searching a preset mapping table for a medical record coding result corresponding to the standardized diagnosis information as the medical record coding of the medical record to be coded; Among them, the encoding model includes a natural language processing module, a word encoding module, a self-attention module, a cross-attention module and a word decoding module; the natural language processing module processes the medical record homepage and other parts of the medical record respectively to obtain the homepage keyword information and other parts of the medical record keyword information; the word encoding module encodes the homepage keyword information and other parts of the medical record keyword information respectively to obtain the first d-dimensional vector and the second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnostic information.

2. The medical record automatic coding method according to claim 1, characterized in that: The natural language processing module includes a first natural language processing submodule and a second natural language processing submodule; the first natural language processing submodule first performs word segmentation processing on the medical record homepage, and then extracts keywords from the word segmentation results of the medical record homepage to obtain homepage keyword information; The second natural language processing submodule first performs word segmentation on the other parts of the medical records, and then extracts keywords from the word segmentation results of the other parts of the medical records to obtain keyword information of the other parts of the medical records.

3. The medical record automatic coding method according to claim 2, characterized in that: The first natural language processing submodule arranges the keyword information of the home page in a structured manner according to the order in which the keyword information of other medical records is obtained; the second natural language processing submodule arranges the keyword information of other medical records in a structured manner according to the order in which the keyword information is obtained.

4. The method for automatic medical record coding according to claim 3, characterized in that: The word encoding module includes a first word encoding submodule and a second word encoding submodule. The first word encoding submodule encodes the structured keyword information of the homepage to obtain a first d-dimensional vector, and the second word encoding submodule encodes the structured keyword information of other parts of the medical records to obtain a second d-dimensional vector.

5. The medical record automatic coding method according to claim 1, characterized in that: The self-attention module includes a first self-attention submodule and a second self-attention submodule. The first self-attention submodule performs feature extraction on the first d-dimensional vector to obtain m first K-dimensional feature vectors, and the second self-attention submodule performs feature extraction on the second d-dimensional vector to obtain n second K-dimensional feature vectors.

6. The method for automatic coding of medical records according to claim 5, characterized in that: The cross-attention module processes a first matrix X1 consisting of m first K-dimensional feature vectors and a second matrix X2 consisting of n second K-dimensional feature vectors to obtain an attention score vector.

7. The method for automatic coding of medical records according to claim 6, characterized in that: The criss-cross attention module uses the linear transformation matrix W K and W V , calculate the key matrix K and value matrix V respectively; K=X1×W K ; H=X1×W V ; The criss-cross attention module uses the linear transformation matrix W Q , calculate the query matrix Q; Q=X2×W Q ; The cross attention module calculates the attention score vector attention(Q,K,V), Among them, K T represents the transpose of the key vector K, where d k is the number of columns of the key matrix K.

8. The method for automatic medical record coding according to claim 1, characterized in that: The training process of the encoding model is as follows: The annotated medical records are used as training data for the natural language processing module to obtain a trained natural language processing module; based on the trained natural language processing module, the annotated medical records are used as the input of the entire encoding model, and the corresponding standardized diagnostic information is used as the output of the entire encoding model, the entire encoding model is trained to obtain a trained encoding model.

9. A medical record automatic coding system, characterized in that: include: A medical record acquisition unit, used for acquiring the medical record to be coded, and dividing the medical record into a medical record front page and other partial medical records; A standardization unit is used to input the medical record front page and other parts of the medical record into a preset coding model for processing to obtain standardized diagnosis information; A medical record coding unit, used to query a medical record coding result corresponding to the standardized diagnosis information in a preset mapping table as the medical record code of the medical record to be coded; Among them, the encoding model includes a natural language processing module, a word encoding module, a self-attention module, a cross-attention module and a word decoding module; the natural language processing module processes the medical record homepage and other parts of the medical record respectively to obtain the homepage keyword information and other parts of the medical record keyword information; the word encoding module encodes the homepage keyword information and other parts of the medical record keyword information respectively to obtain the first d-dimensional vector and the second d-dimensional vector; the self-attention module extracts features from the first d-dimensional vector and the second d-dimensional vector respectively to obtain m first K-dimensional feature vectors and n second K-dimensional feature vectors; the cross-attention module processes the m first K-dimensional feature vectors and n second K-dimensional feature vectors to obtain an attention score vector; the word decoding module decodes the attention score vector to obtain standardized diagnostic information.

Citation Information

Cited By

  • Tumor code detection and identification method and system based on artificial intelligence

    CN120745559A