A DRG Automatic Grouping Method and System Based on Semantic Information Fusion
By integrating semantic information on medical data, the problem of existing DRG systems ignoring the semantics of unstructured text data is solved, and more accurate and reliable DRG packets are achieved, which improves the efficiency of medical resources.
Patent Information
- Application Number
- CN202411296612.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-09-18
AI Technical Summary
When using medical data for grouping, existing DRG systems mainly rely on structured data, ignoring the semantic information in unstructured text data, resulting in the limitation of the accuracy and scientificity of the grouping.
A DRG automatic grouping method based on semantic information fusion is proposed. Through artificial intelligence technology, structured and unstructured information in medical data are semantic mining and fusion to generate more accurate and comprehensive DRG grouping information.
Through semantic information fusion, the accuracy and reliability of DRG packets can be improved, ensuring that the packet results are closer to actual medical needs, and improving the efficiency of medical resources.
Smart Images

Figure CN119092104B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of DRG automatic grouping, and particularly to a DRG automatic grouping method and system based on semantic information fusion. Background Art
[0002] In the medical industry, the Diagnosis Related Groups (DRG) system is a widely used patient classification tool for dividing hospital patients into several groups, each group containing clinically similar and resource consumption - similar cases. The emergence of the DRG system not only improves the management efficiency of medical services, but also provides a scientific basis for medical insurance reimbursement. Currently, many countries in the world have established their own DRG systems for medical expense settlement, hospital performance evaluation, and medical service quality control.
[0003] With the development of medical informatization, a large amount of medical data has been accumulated in hospital information systems. These data include structured personal information data and unstructured text data, such as doctors' diagnosis records, surgical records, and nursing records. These data contain rich medical knowledge and can provide more accurate and comprehensive information support for DRG grouping. However, when using these data for grouping, the existing DRG systems mainly rely on structured data and ignore the semantic information in unstructured text data, resulting in certain limitations in the accuracy and scientific nature of grouping. Summary of the Invention
[0004] In view of this, the present invention proposes a DRG automatic grouping method based on semantic information fusion, which makes full use of structured and unstructured information in medical data through artificial intelligence technology for semantic mining and fusion, improving the accuracy and reliability of DRG grouping.
[0005] To achieve the above - mentioned purpose, a DRG automatic grouping method based on semantic information fusion provided by the present invention includes the following steps:
[0006] S1: Collect patient personal information, medical record information, and diagnosis information to form electronic health record data. The electronic health record data contains structured data and unstructured data, where patient personal information is structured data, and patient medical records and diagnosis information are unstructured data;
[0007] S2: Extract semantic information from the patient's electronic health record data to obtain a segmented semantic information feature vector describing the patient information, where the multi - scale semantic fusion method is the implementation method of the semantic information extraction;
[0008] S3: Perform semantic enhancement processing on the segmented semantic information feature vector and generate the DRG grouping information of the patient;
[0009] S4: Based on the DRG grouping information and the personal information of patients, perform clustering processing on different patients to obtain several patient clusters, and the medical resources required by the patients in each patient cluster are approximately the same, where the improved peak density estimation is the implementation method of the clustering processing.
[0010] As a further improvement method of the present invention:
[0011] Optionally, in the step S1, collecting the personal information, medical records and diagnosis information of patients to form electronic health record data includes:
[0012] Collect the personal information, medical record information and diagnosis information of N patients to be DRG grouped to form the electronic health record data of the patients, where the electronic health record data of N patients is:
[0013] { X n =( X n 1 , X n 2 , X n 3 )|n∈[1,N]} ;
[0014] ;
[0015] ;
[0016] ;
[0017] Where:
[0018] represents the electronic health record data of the nth patient;
[0019] represents the personal information in the electronic health record data and successively represents the gender, age, occupation, height and weight of the nth patient;
[0020] i represents the i-th segmented phrase in the medical record information ∈ [1, Sum n ] ;
[0021] represents the electronic health record data in the diagnostic information, represents the diagnostic information in the word segmentation phrases, represents the j-th word segmentation phrase in the diagnostic information where j represents the j-th word segmentation phrase in the diagnostic information represents the diagnostic information in the total number of word segmentation phrases.
[0022] Optionally, the semantic information extraction of the patient's electronic health record data in step S2 includes:
[0023] Performing semantic information extraction on the patient's electronic health record data to obtain a segmented semantic information feature vector describing the patient's information, where the electronic health record data The semantic information extraction process is:
[0024] S21: One-hot encoding all phrases in the electronic health record data to obtain the one-hot encoding vector corresponding to the phrase;
[0025] S22: Performing word vector representation on the one-hot encoding vectors describing different information respectively to obtain the word vector encoding sequences corresponding to different information, where the word vector encoding sequence of the e-th type of information is:
[0026] X n e = W e *[ C n e (1), C n e (2),..., C n e (S),..., C n e ( S e )],e ∈ {1,2,3} ;
[0027] ;
[0028] where:
[0029] represents the encoding matrix of the e-th type of information, * represents the convolution operation, where the 1st - 3rd types of information are personal information, medical record information, and diagnostic information in sequence;
[0030] The one-hot encoding vector corresponding to the S-th phrase in the information; S ∈[1, S e ] ;
[0031] The word vector encoding sequence representing the e-th type of information, where ,
[0032] The word vector corresponding to the S-th phrase in the information;
[0033] S23: Perform multi-scale convolution on the word vector encoding sequences of different information to obtain the feature information of the word vector encoding sequences at D scales;
[0034] S24: Extract semantic information from the multi-scale feature information of different information to obtain the semantic feature vectors representing the patient's personal information, medical record information, and diagnosis information in the electronic health record data, where the semantic information extraction formula is:
[0035] F n =[ F n 1 , F n 2 , F n 3 ] ,e ∈ {1,2,3} ;
[0036] F n e = ∑ h=1 H [ 1 1+exp(- weight e * X n e,h ) + X n e,h ] ;
[0037] Where:
[0038] represents the semantic feature vector of the e-th type of information;
[0039] successively represent the semantic feature vectors of the patient's personal information, medical record information, and diagnosis information;
[0040] represents the semantic mapping matrix of the e-th type of information;
[0041] Represents a segmented semantic information feature vector describing patient information.
[0042] Optionally, in step S23, multi-scale convolution is performed on the word vector encoding sequences of different information, including:
[0043] S231: Set the current convolution scale as h, the initial value of h is 0, and the maximum value is H;
[0044] S232: Generate the feature information of each type of information at scale h + 1, where the feature information of the e-th type of information at scale h + 1 is:
[0045] ;
[0046] ;
[0047] Where:
[0048] Represents the feature information of the e-th type of information at scale h + 1;
[0049] Represents the convolution matrix at scale h; in the embodiments of the present invention, the convolution matrix has a size of ;
[0050] S233: Let , and return to step S231 until the maximum convolution scale H is reached.
[0051] Optionally, in step S3, semantic enhancement processing is performed on the segmented semantic information feature vector, including:
[0052] Perform semantic enhancement processing on the segmented semantic information feature vector, where the semantic enhancement processing process of the segmented semantic information feature vector corresponding to the n-th patient is as follows:
[0053] S31: Generate the disease group information :
[0054] ;
[0055] A n 2 = exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 ] exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 +exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] ;
[0056] A n 3 = exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 +exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] ;
[0057] Where:
[0058] Represents the semantic feature vector of the enhancement coefficient, Represents the semantic feature vector of the enhancement coefficient;
[0059] Represents the semantic feature vector of the enhancement weight matrix, Represents the semantic feature vector of the enhancement weight matrix;
[0060] T represents transpose;
[0061] Represents the semantic feature vector of the length, Represents the semantic feature vector of the length;
[0062] S32: Generate the disease group information representing the patient's disease treatment method :
[0063] ;
[0064] Where:
[0065] Represents the L1 norm;
[0066] Represents the ReLU function;
[0067] S33: Generate the disease group information representing complications :
[0068] f n 3 = exp W 1 1 * F n 1 ∙ ( W 1 2 * F n 1 ) T Len n 1 ] F n 1 || f n 2 || ∙f n 2 1 + || f n 2 || 2 ;
[0069] Where:
[0070] Represents the semantic feature vector of the enhancement weight matrix;
[0071] Represents the semantic feature vector of the length;
[0072] S34: Use the disease group information that fuses multiple semantic feature vectors as the patient's semantic enhancement processing result f n =[ f n 1 , f n 2 , f n 3 ] ;
[0073] Generate the DRG grouping information of N patients according to the semantic enhancement processing result.
[0074] Optionally, generating the DRG grouping information of N patients according to the semantic enhancement processing result includes:
[0075] The generation method of the DRG grouping information of the nth patient is:
[0076] Generate a DRG code representing the major diagnosis category of the patient :
[0077] Code n 1 = argmax r∈ Code 1 exp ( Q 1 r ) T f n 1 ] ∑ r∈ Code 1 exp ( Q 1 r ) T f n 1 ] ;
[0078] Wherein:
[0079] A set of DRG codes representing the major diagnosis category; in the embodiments of the present invention, The DRG codes in are divided into A to Z;
[0080] Represents the set of DRG codes The DRG code in the set of DRG codes The corresponding coding matrix;
[0081] Generate a DRG code representing the disease treatment method of the patient :
[0082] Code n 2 = argmax r∈ Code 2 exp ( Q 2 r ) T f n 2 ] ∑ r∈ Code 2 exp ( Q 2 r ) T f n 2 ] ;
[0083] Wherein:
[0084] A set of DRG codes representing the disease treatment method; in the embodiments of the present invention, The DRG codes in are divided into A to Z, where A to J represent the surgical part, K to Q represent the non-operating room operation part, and R to Z represent the medical part;
[0085] Represents the set of DRG codes The DRG code in the set of DRG codes The corresponding coding matrix;
[0086] Generate DRG codes according to the processing sequence of disease treatment methods The corresponding sequence code ; In the embodiment of the present invention, the sequence code is from 1 to 9;
[0087] Generate DRG codes representing the complications of patients :
[0088] ;
[0089] Wherein:
[0090] represents the age of the nth patient;
[0091] represents the set of DRG codes of the complications of the patient; in the embodiment of the present invention, the DRG codes in are divided into 1, 3, and 5, where 1 represents severe complications, 3 represents general complications, and 5 represents no complications;
[0092] represents the set of DRG codes the coding matrix corresponding to the DRG code c in;
[0093] constitute the DRG grouping information of the nth patient:
[0094] ;
[0095] Wherein:
[0096] represents the DRG grouping information of the nth patient.
[0097] Optionally, in step S4, based on the DRG grouping information and the personal information of the patient, clustering processing is performed on different patients to obtain a number of patient clusters, including:
[0098] Based on the DRG grouping information and the personal information of the patient, clustering processing is performed on different patients to obtain a number of patient clusters. The patients in each patient cluster consume approximately the same medical resources. The improved peak density estimation is the implementation method of the clustering processing, and the clustering process is as follows:
[0099] S41: Extract the DRG codes representing the major categories of the main diagnoses of the patients in the DRG grouping information of N patients, and use the patients with the same DRG codes as patients in the same diagnosis area to obtain a set of patients in 26 diagnosis areas, where 26 is the number of types of major categories of the main diagnoses;
[0100] S42: Obtain the DRG grouping information and the personal information of any patient in each patient set, and calculate the information distance between this patient and any patient in the same patient set. The information distance between the nth patient and any patient in the same patient set is:
[0101] ;
[0102] ;
[0103] ;
[0104] where:
[0105] represents the information distance between the nth patient and the mth patient in the same patient set, represents the semantic feature vector of the personal information of the mth patient in the same patient set, represents the DRG code corresponding to the disease treatment method of the mth patient in the same patient set, represents the DRG code corresponding to the complication of the mth patient in the same patient set;
[0106] represents the L2 norm;
[0107] S43: Calculate the weighted information density of any patient in each patient set. The weighted information density of the nth patient is:
[0108] ;
[0109] where:
[0110] represents the weighted information density of the nth patient, represents the patient set where the nth patient is located, represents the patient set any patient in;
[0111] represents the patient set the number of patients in whose information distance from the nth patient is within a preset distance threshold, represents the patient set the set of neighboring patients whose information distance from the nth patient is within a preset distance threshold, and u represents any patient in the set of neighboring patients any patient in;
[0112] represents the preset density truncation distance;
[0113] S44: Select the K patients with the highest weighted information density as the initial clustering centers, and use the Kmean algorithm to perform iterative processing and clustering of the clustering centers for each patient set, and divide each patient set into K patient clustering clusters.
[0114] Optionally, the word segmentation process for medical record information and diagnosis information is as follows:
[0115] S11: Use a word segmentation tool to perform word segmentation on the medical record information and diagnosis information to obtain the initial word segmentation sequences of the medical record information and diagnosis information; in the embodiments of the present invention, the word segmentation tool used is the jieba word segmentation tool;
[0116] S12: Calculate the importance weights of each initial word segmentation result in the initial word segmentation sequences corresponding to the medical record information and diagnosis information respectively, where the initial word segmentation sequence corresponding to the medical record information the y-th initial word segmentation result in the importance information is:
[0117] I 1 (y) = ln[ E 1 1 (y)exp( E 2 1 (y))+ E 2 1 (y)exp( E 1 1 (y)) | E 1 1 (y)- E 2 1 (y)| ] ;
[0118] ;
[0119] ;
[0120] where:
[0121] represents the importance information of the y-th initial word segmentation result in the initial word segmentation sequence corresponding to the medical record information ;
[0122] represents the exponential function with the natural constant as the base;
[0123] represents the left adjacent information of the initial word segmentation result ; represents the right adjacent information of the initial word segmentation result ;
[0124] represents the probability that the initial word segmentation result and the initial word segmentation result appear in the same sentence in all medical record information;
[0125] represents the initial word segmentation result The probability of occurrence in all medical record information;
[0126] Among them, the initial word segmentation sequence corresponding to the diagnostic information The importance information of the y-th initial word segmentation result is:
[0127] I 2 (y) = ln[ E 1 2 (y)exp( E 2 2 (y))+ E 2 2 (y)exp( E 1 2 (y)) | E 1 2 (y)- E 2 2 (y)| ] ;
[0128] ;
[0129] ;
[0130] Among them:
[0131] represents the importance information of the y-th initial word segmentation result in the initial word segmentation sequence corresponding to the diagnostic information ; represents the left adjacent information of the initial word segmentation result
[0132] represents the right adjacent information of the initial word segmentation result ; represents the probability that the initial word segmentation result and the initial word segmentation result
[0133] appear in the same sentence in all diagnostic information; represents the probability of occurrence of the initial word segmentation result
[0134] in all diagnostic information;
[0135]
[0136] S13: Filter the initial word segmentation results with importance information lower than the preset threshold, and use the remaining initial word segmentation results as the word segmentation phrases of the medical record information and the diagnostic information, so as to obtain the word segmentation phrase sequence of the medical record information and the diagnostic information.
[0137] To solve the above problems, the present invention provides a DRG automatic grouping system based on semantic information fusion, and the system includes:
[0138] An information collection module, which is used to collect patient personal information, medical record information and diagnostic information to form electronic health record data;
[0138] An information extraction module, which is used to perform semantic information extraction on the electronic health record data of patients, obtain a segmented semantic information feature vector describing patient information, and perform semantic enhancement processing on the segmented semantic information feature vector;
[0139] An automatic grouping device, which is used to generate DRG grouping information of patients, perform clustering processing on different patients based on the DRG grouping information and patient personal information, and obtain several patient clustering clusters.
[0140] To solve the above problems, the present invention also provides an electronic device, and the electronic device includes:
[0141] A memory that stores at least one instruction;
[0142] A communication interface to enable communication of the electronic device; and
[0143] A processor that executes the instructions stored in the memory to implement the above-mentioned DRG automatic grouping method for semantic information fusion.
[0144] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned DRG automatic grouping method for semantic information fusion.
[0145] Compared with the prior art, the present invention proposes a DRG automatic grouping method based on semantic information fusion, and this technology has the following advantages:
[0146] First of all, this solution proposes a way of semantic extraction of patient information. By collecting the personal information, medical record information, and diagnosis information of patients as electronic health record data, and respectively performing multi-scale semantic feature information extraction, a segmented semantic information feature vector describing different patient information is formed, realizing multi-level and in-depth information extraction of the patient's personal situation, condition description, and diagnosis situation.
[0147] At the same time, this solution proposes a DRG grouping method. According to the multi-segment code composition form of DRG grouping, the segmented semantic information feature vector is spliced, enhanced, and norm-constrained to form the disease group information representing the major disease diagnosis categories, disease treatment methods, and complications of patients, and map different disease group information to the corresponding DRG codes to achieve the mapping of DRG codes, obtain the DRG grouping information of patients, and combine the DRG grouping information and patient personal information. Based on DRG and personal information, the information distance between different patients is calculated, and users with similar personal information and similar DRG are formed into patient clustering clusters. The medical resources consumed by patients in each patient clustering cluster are approximately the same, realizing the batch allocation of medical resources and improving the use efficiency of medical resources. Description of the Drawings
[0148] Figure 1 It is a schematic flowchart of a DRG automatic grouping method based on semantic information fusion provided by an embodiment of the present invention;
[0149] Figure 2 It is a functional module diagram of a DRG automatic grouping system with semantic information fusion provided by an embodiment of the present invention;
[0150] Figure 2 Among them: 100 is a DRG automatic grouping system with semantic information fusion, 101 is an information collection module, 102 is an information extraction module, and 103 is an automatic grouping device;
[0151] Figure 3 It is a schematic structural diagram of an electronic device for implementing the DRG automatic grouping method with semantic information fusion provided by an embodiment of the present invention.
[0152] Figure 3 Among them: 1 is an electronic device, 10 is a processor, 11 is a memory, 12 is a program, and 13 is a communication interface;
[0153] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0154] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0155] The embodiments of the present application provide a DRG automatic grouping method based on semantic information fusion. The execution subject of the DRG automatic grouping method with semantic information fusion includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the DRG automatic grouping method with semantic information fusion can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.
[0156] Embodiment 1:
[0157] S1: Collect the patient's personal information, medical record information and diagnosis information to form electronic health record data. The electronic health record data includes structured data and unstructured data, where the patient's personal information is structured data, and the patient's medical record and diagnosis information are unstructured data.
[0158] In the S1 step of collecting the patient's personal information, medical record and diagnosis information to form electronic health record data, it includes:
[0159] Collect the personal information, medical record information, and diagnosis information of N patients to be grouped by DRG to form the electronic health record data of the patients, where the electronic health record data of N patients is:
[0160] { X n =( X n 1 , X n 2 , X n 3 )|n∈[1,N]} ;
[0161] ;
[0162] ;
[0163] ;
[0164] Among them:
[0165] represents the electronic health record data of the nth patient;
[0166] represents the personal information in the electronic health record data in represents the gender, age, occupation, height, and weight of the nth patient in sequence;
[0167] represents the medical record information in the electronic health record data in represents the medical record information in word segmentation phrases, represents the ith word segmentation phrase in the medical record information in represents the medical record information in i ∈ [1, Sum n ] ;
[0168] represents the diagnosis information in the electronic health record data in represents the diagnosis information in word segmentation phrases, represents the jth word segmentation phrase in the diagnosis information in Indicates diagnostic information The total number of segmented phrase groups in
[0169] The segmentation process for medical record information and diagnostic information is as follows:
[0170] S11: Use a word segmentation tool to segment the medical record information and diagnostic information to obtain the initial segmented sequences of the medical record information and diagnostic information;
[0171] S12: Calculate the importance weights of each initial segmentation result in the initial segmented sequences corresponding to the medical record information and diagnostic information respectively. Among them, for the initial segmented sequence corresponding to the medical record information the y-th initial segmentation result in the importance information is:
[0172] I 1 (y) = ln[ E 1 1 (y)exp( E 2 1 (y))+ E 2 1 (y)exp( E 1 1 (y)) | E 1 1 (y)- E 2 1 (y)| ] ;
[0173] ;
[0174] ;
[0175] Among them:
[0176] Indicates the initial segmented sequence corresponding to the medical record information the y-th initial segmentation result in the importance information;
[0177] Indicates the exponential function with the natural constant as the base;
[0178] Indicates the left adjacent information of the initial segmentation result and Indicates the right adjacent information of the initial segmentation result ;
[0179] Indicates the probability that the initial segmentation result and the initial segmentation result appear in the same sentence in all medical record information;
[0180] Indicates the probability that the initial segmentation result appears in all medical record information;
[0181] Among them, for the initial segmented sequence corresponding to the diagnostic information the y-th initial segmentation result in The importance information is as follows:
[0182] I 2 (y) = ln[ E 1 2 (y)exp( E 2 2 (y))+ E 2 2 (y)exp( E 1 2 (y)) | E 1 2 (y)- E 2 2 (y)| ] ;
[0183] ;
[0184] ;
[0185] Wherein:
[0186] represents the importance information of the y-th initial word segmentation result in the initial word segmentation sequence corresponding to the diagnostic information in the middle; of the importance information;
[0187] represents the left adjacent information of the initial word segmentation result in the middle, represents the right adjacent information of the initial word segmentation result in the middle;
[0188] represents the probability that the initial word segmentation result and the initial word segmentation result appear in the same sentence in all diagnostic information;
[0189] represents the probability that the initial word segmentation result appears in all diagnostic information;
[0190] S13: Filter the initial word segmentation results with importance information lower than the preset threshold, and use the retained initial word segmentation results as the word segmentation phrases of the medical record information and the diagnostic information to obtain the word segmentation phrase sequence of the medical record information and the diagnostic information.
[0191] S2: Extract semantic information from the electronic health record data of the patient to obtain a segmented semantic information feature vector describing the patient information.
[0192] In the step S2 of extracting semantic information from the electronic health record data of the patient, it includes:
[0193] Extract semantic information from the electronic health record data of the patient to obtain a segmented semantic information feature vector describing the patient information, where the electronic health record data The semantic information extraction process is as follows:
[0194] S21: For the electronic health record data All phrases in
[0195] S22: Perform word vector representations on the one-hot encoded vectors describing different information respectively to obtain word vector encoding sequences corresponding to different information, where the word vector encoding sequence of the e-th type of information is:
[0196] X n e = W e *[ C n e (1), C n e (2),..., C n e (S),..., C n e ( S e )],e ∈ {1,2,3} ;
[0197] ;
[0198] where:
[0199] represents the encoding matrix of the e-th type of information, * represents the convolution operation, and the first to third types of information are personal information, medical record information, and diagnosis information in sequence;
[0200] represents the information in the one-hot encoded vector corresponding to the S-th phrase, S ∈[1, S e ] ;
[0201] represents the word vector encoding sequence of the e-th type of information, where ,
[0202] represents the information in the word vector corresponding to the S-th phrase;
[0203] S23: Perform multi-scale convolution on the word vector encoding sequences of different information to obtain feature information of the word vector encoding sequences at D scales;
[0204] S24: Extract semantic information from the multi-scale feature information of different information to obtain semantic feature vectors in the electronic health record data describing the patient's personal information, medical record information, and diagnosis information, where the semantic information extraction formula is:
[0205] F n =[ F n 1 , F n 2 , F n 3 ] ,e ∈ {1,2,3} ;
[0206] F n e = ∑ h=1 H [ 1 1+exp(- weight e * X n e,h ) + X n e,h ] ;
[0207] Wherein:
[0208] represents the semantic feature vector of the e-th kind of information;
[0209] successively represent the semantic feature vectors of the patient's personal information, medical record information, and diagnostic information;
[0210] represents the semantic mapping matrix of the e-th kind of information;
[0211] represents the segmented semantic information feature vector describing the patient information.
[0212] In the step S23, multi-scale convolution is performed on the word vector encoding sequences of different information, including:
[0213] S231: Set the current convolution scale as h, the initial value of h is 0, and the maximum value is H;
[0214] S232: Generate the feature information of each kind of information at scale h + 1, wherein the feature information of the e-th kind of information at scale h + 1 is:
[0215] ;
[0216]
[0217] Wherein:
[0218] represents the feature information of the e-th kind of information at scale h + 1;
[0219] represents the convolution matrix at scale h;
[0220] Let , return to step S231 until the maximum convolution scale H is reached.
[0221] S3: Semantically enhance the segmented semantic information feature vector and generate the DRG grouping information of the patient.
[0222] In the S3 step, the semantic enhancement process of the segmented semantic information feature vector includes:
[0223] Semantically enhance the segmented semantic information feature vector, where the semantic enhancement process of the segmented semantic information feature vector corresponding to the nth patient is as follows:
[0224] S31: Generate the disease group information representing the major disease diagnosis of the patient :
[0225] ;
[0226] A n 2 = exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 ] exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 +exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] ;
[0227] A n 3 = exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] exp W 2 1 * F n 2 ∙ ( W 2 2 * F n 2 ) T Len n 2 +exp W 3 1 * F n 3 ∙ ( W 3 2 * F n 3 ) T Len n 3 ] ;
[0228] where:
[0229] represents the enhancement coefficient of the semantic feature vector , represents the enhancement coefficient of the semantic feature vector ;
[0230] represents the enhancement weight matrix of the semantic feature vector , represents the enhancement weight matrix of the semantic feature vector ;
[0231] T represents transpose;
[0232] represents the length of the semantic feature vector , represents the length of the semantic feature vector ;
[0233] S32: Generate the disease group information representing the disease treatment method of the patient :
[0234] ;
[0235] Among them:
[0236] represents the L1 norm;
[0237] represents the ReLU function;
[0238] S33: Generate the disease group information representing complications :
[0239] f n 3 = exp W 1 1 * F n 1 ∙ ( W 1 2 * F n 1 ) T Len n 1 ] F n 1 || f n 2 || ∙f n 2 1 + || f n 2 || 2 ;
[0240] Among them:
[0241] represents the enhanced weight matrix of the semantic feature vector ;
[0242] represents the semantic feature vector length;
[0243] S34: Use the disease group information integrating multiple semantic feature vectors as the semantic enhancement processing result of the patient f n =[ f n 1 , f n 2 , f n 3 ] ;
[0244] Generate the DRG grouping information of N patients according to the semantic enhancement processing result.
[0245] The generating the DRG grouping information of N patients according to the semantic enhancement processing result includes:
[0246] The generating method of the DRG grouping information of the nth patient is:
[0247] Generate the DRG code representing the major diagnosis category of the patient :
[0248] Code n 1 = argmax r∈ Code 1 exp ( Q 1 r ) T f n 1 ] ∑ r∈ Code 1 exp ( Q 1 r ) T f n 1 ] ;
[0249] Among them:
[0250] The set of DRG codes representing the major diagnostic categories; in the embodiments of the present invention, the DRG codes in
[0251] represent the set of DRG codes in the DRG code set corresponding coding matrix;
[0252] Generate DRG codes characterizing the disease treatment methods of patients :
[0253] Code n 2 = argmax r∈ Code 2 exp ( Q 2 r ) T f n 2 ] ∑ r∈ Code 2 exp ( Q 2 r ) T f n 2 ] ;
[0254] Among them:
[0255] The set of DRG codes representing the disease treatment methods; in the embodiments of the present invention, the DRG codes in
[0256] represent the set of DRG codes in the DRG code set corresponding coding matrix;
[0257] Generate the sequence code corresponding to the DRG code according to the treatment order of the disease treatment method ; in the embodiments of the present invention, the sequence code is from 1 to 9; ;
[0258] Generate DRG codes characterizing the complications of patients :
[0259] ;
[0260] Among them:
[0261] represents the age of the nth patient;
[0262] The set of DRG codes representing the complications of patients; in the embodiments of the present invention, the DRG codes in
[0263] represent the set of DRG codes The coding matrix corresponding to the DRG code c;
[0264] Constitute the DRG grouping information of the nth patient:
[0265] ;
[0266] Wherein:
[0267] Represents the DRG grouping information of the nth patient.
[0268] S4: Based on the DRG grouping information and the patient's personal information, perform clustering on different patients to obtain several patient clusters, and the medical resources required by the patients in each patient cluster are approximately the same.
[0269] In the step S4, based on the DRG grouping information and the patient's personal information, perform clustering on different patients to obtain several patient clusters, including:
[0270] Based on the DRG grouping information and the patient's personal information, perform clustering on different patients to obtain several patient clusters, and the medical resources required by the patients in each patient cluster are approximately the same. Among them, the improved peak density estimation is the implementation method of the clustering process, and the clustering process is as follows:
[0271] S41: Extract the DRG codes representing the major diagnosis categories of the patients from the DRG grouping information of N patients, and use the patients with the same DRG code as the patients in the same diagnosis area to obtain the patient sets in 26 diagnosis areas, where 26 is the number of types of major diagnosis categories;
[0272] S42: Obtain the DRG grouping information and the patient's personal information of any patient in each patient set, and calculate the information distance between this patient and any patient in the same patient set. The information distance between the nth patient and any patient in the same patient set is:
[0273] ;
[0274] ;
[0275] ;
[0276] Wherein:
[0277] Represents the information distance between the nth patient and the mth patient in the same patient set, Represents the semantic feature vector of the personal information of the mth patient in the same patient set, Denotes the DRG code corresponding to the disease treatment method of the m-th patient in the patient set, Denotes the DRG code corresponding to the complication of the m-th patient in the patient set;
[0278] Denotes the L2 norm;
[0279] S43: Calculate the weighted information density of any patient in each patient set, where the weighted information density of the n-th patient is:
[0280] ;
[0281] Where:
[0282] Denotes the weighted information density of the n-th patient, Denotes the patient set where the n-th patient is located, Denotes the patient set Any patient in;
[0283] Denotes the patient set The number of patients in the patient set whose information distance from the n-th patient is within the preset distance threshold, Denotes the patient set The set of neighboring patients whose information distance from the n-th patient is within the preset distance threshold in the patient set, u denotes any patient in the set of neighboring patients in;
[0284] Denotes the preset density cut-off distance;
[0285] S44: Select the K patients with the highest weighted information density as the initial clustering centers, and use the Kmean algorithm to perform iterative processing and clustering processing of the clustering centers for each patient set, and divide each patient set into K patient clustering clusters.
[0286] Example 2:
[0287] As Figure 2 shown, it is the functional module diagram of the DRG automatic grouping system for semantic information fusion provided by an embodiment of the present invention, which can implement the DRG automatic grouping method for semantic information fusion in Embodiment 1.
[0288] The DRG automatic grouping system 100 for semantic information fusion according to the present invention can be installed in an electronic device. According to the functions achieved, the DRG automatic grouping system for semantic information fusion may include an information collection module 101, an information extraction module 102, and an automatic grouping device 103. The modules in the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0289] The information collection module 101 is used to collect patient personal information, medical record information, and diagnosis information to form electronic health record data;
[0290] The information extraction module 102 is used to extract semantic information from the electronic health record data of the patient, obtain a segmented semantic information feature vector describing the patient information, and perform semantic enhancement processing on the segmented semantic information feature vector;
[0291] The automatic grouping device 103 is used to generate DRG grouping information of the patient, and perform clustering processing on different patients based on the DRG grouping information and the patient personal information to obtain several patient clusters.
[0292] Specifically, each module in the DRG automatic grouping system 100 for semantic information fusion in the embodiment of the present invention adopts the same technical means as those Figure 1 described in the above-mentioned semantic information fusion DRG automatic grouping method, and can produce the same technical effects, which will not be elaborated here.
[0293] Embodiment 3:
[0294] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the DRG automatic grouping method for semantic information fusion provided by an embodiment of the present invention.
[0295] The electronic device 1 may include a processor 10, a memory 11, a communication interface 13, and a bus, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as program 12.
[0296] Among them, the memory 11 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code of the program 12, etc., but also to temporarily store the data that has been output or will be output.
[0297] In some embodiments, the processor 10 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing the programs or modules (such as the program 12 for realizing semantic information fusion DRG automatic grouping) stored in the memory 11, and calling the data stored in the memory 11, to execute various functions of the electronic device 1 and process data.
[0298] The communication interface 13 can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is usually used to establish a communication connection between the electronic device 1 and other electronic devices, and to realize the connection communication between the internal components of the electronic device.
[0299] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection communication between the memory 11 and at least one processor 10, etc.
[0300] Figure 3 Only an electronic device with components is shown. It can be understood by those skilled in the art that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0301] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0302] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0303] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0304] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, a device, an article or a method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such a process, a device, an article or a method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, device, article or method including the element.
[0305] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0306] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A DRG automatic grouping method based on semantic information fusion, characterized in that: The method comprises: S1: Collecting patient personal information, medical record information and diagnosis information to form electronic health record data, wherein the electronic health record data includes structured data and unstructured data, wherein the patient personal information is structured data, and the patient medical record information and diagnosis information are unstructured data; S2: Extract semantic information from the patient's electronic health record data to obtain a segmented semantic information feature vector describing the patient's information; S3: Perform semantic enhancement processing on the segmented semantic information feature vector and generate the patient's DRG grouping information; The step S3 performs semantic enhancement processing on the segmented semantic information feature vector, including: The segmented semantic information feature vector is semantically enhanced, where the segmented semantic information feature vector F corresponding to the nth patient is n The semantic enhancement process is as follows: S31: Generate disease group information representing the major categories of patient disease diagnosis in: Represents semantic feature vector The enhancement factor, Represents semantic feature vector The enhancement factor of Represents semantic feature vector The enhanced weight matrix, Represents semantic feature vector The enhanced weight matrix of T represents transpose; n∈[1,N]; Represents semantic feature vector Length, Represents semantic feature vector Length; S32: Generate disease group information that characterizes the patient's disease management method in: ||·|| represents the L1 norm; ReLU(·) represents the ReLU function; S33: Generate disease group information to characterize complications in: Represents semantic feature vector The enhanced weight matrix of Represents semantic feature vector Length; S34: The disease group information of multiple semantic feature vectors is integrated as the semantic enhancement processing result of the patient Semantic feature vectors representing patient personal information, medical record information, and diagnosis information in turn; According to the semantic enhancement processing results, DRG grouping information of N patients is generated, including: The DRG grouping information of the nth patient is generated as follows: Generate DRG codes representing the patient's main diagnosis category in: Code1 represents the DRG code set of the main diagnosis category; Represents the coding matrix corresponding to DRG code r in DRG code set Code1; Generate DRG codes that characterize how the patient's disease is managed in: Code2 represents the DRG code set for disease treatment methods; Represents the coding matrix corresponding to the DRG code r' in the DRG code set Code2; Generate DRG codes based on the treatment order of disease treatment methods Corresponding sequence code Generate DRG codes that characterize patient complications in: represents the age of the nth patient; Code4 represents the DRG code set for patient complications; Represents the coding matrix corresponding to DRG code c in DRG code set Code4; The DRG grouping information of the nth patient is: in: DRG n Indicates the DRG grouping information of the nth patient; S4: Based on the DRG grouping information and patient personal information, different patients are clustered and processed to obtain several patient clusters. The medical resources required by the patients in each patient cluster are similar; The clustering process is as follows: S41: extracting the DRG codes representing the major diagnosis categories of the patients from the DRG grouping information of N patients, and treating the patients with the same DRG codes as patients in the same diagnosis area, thereby obtaining a set of patients in 26 diagnosis areas, where 26 is the number of major diagnosis categories; S42: Obtain the DRG grouping information and patient personal information of any patient in each patient set, and calculate the information distance between the patient and any patient in the same patient set, where the information distance between the nth patient and any patient in the same patient set is: in: dist(n,D(m)) represents the information distance between the nth patient and the mth patient in the same patient set, F 1 (m) represents the semantic feature vector of the personal information of the mth patient in the same patient set, Code 2 (m) represents the DRG code corresponding to the disease treatment method of the mth patient in the same patient set, Code 4 (m) represents the DRG code corresponding to the complication of the mth patient in the same patient set; ||·||2 represents the L2 norm; S43: Calculate the weighted information density of any patient in each set of patients, where the weighted information density of the nth patient is: in: ρ n represents the weighted information density of the nth patient, D(n) represents the set of patients to which the nth patient belongs, and β represents any patient in the set of patients D(n); Count(n) represents the number of patients in the same patient set D(n) whose information distance to the nth patient is within the preset distance threshold, Ω n represents the set of neighboring patients in the same patient set D(n) whose information distance to the nth patient is within the preset distance threshold, and u represents the set of neighboring patients Ω n Any patient in dis represents the preset density cutoff distance; S44: Select K patients with the highest weighted information density as initial cluster centers, use Kmean algorithm to iteratively process the cluster centers and cluster each patient set, and divide each patient set into K patient clusters.
2. The DRG automatic grouping method based on semantic information fusion according to claim 1, characterized in that: In step S1, the patient's personal information, medical history and diagnosis information are collected to form electronic health record data, including: The personal information, medical record information and diagnosis information of N patients to be grouped into DRGs are collected and word segmented to form the electronic health record data of the patients. The electronic health record data of N patients are as follows: in: X n represents the electronic health record data of the nth patient; Represents electronic health record data X n Personal information in represents the gender, age, occupation, height and weight of the nth patient in turn; Represents electronic health record data X n Medical records information in Display medical record information Sum in n participle phrases, Display medical record information The i-th participle phrase in Sum n Display medical record information The total number of word groups in i∈[1,Sum n ]; Represents electronic health record data X n Diagnostic information in Indicates diagnostic information Num n participle phrases, Indicates diagnostic information The jth participle phrase in n Indicates diagnostic information The total number of participle phrases in .
3. The DRG automatic grouping method based on semantic information fusion according to claim 2, characterized in that: The S2 extracts semantic information from the patient's electronic health record data, including: Semantic information is extracted from the patient's electronic health record data to obtain a segmented semantic information feature vector describing the patient information, where the electronic health record data X n The semantic information extraction process is as follows: S21: Electronic health record data X n All phrases in are represented by one-hot encoding, and the one-hot encoding vector corresponding to the phrase is obtained; S22: Perform word vector representation on the one-hot encoding vectors describing different information respectively to obtain word vector encoding sequences corresponding to different information, where the word vector encoding sequence of the e-th information is: in: W e represents the encoding matrix of the e-th type of information, * represents the convolution operation, where the 1st to 3rd types of information are personal information, medical record information, and diagnosis information respectively; Display information The unique hot encoding vector corresponding to the Sth phrase in , S∈[1,S e ]; The word vector encoding sequence representing the e-th information, where Display information The word vector corresponding to the Sth word group in ; S23: Perform multi-scale convolution on word vector encoding sequences of different information to obtain feature information of the word vector encoding sequences at D scales; S24: Extract semantic information from multi-scale feature information of different information to obtain electronic health record data X n The semantic feature vectors representing the patient's personal information, medical record information, and diagnosis information are expressed in the formula for extracting semantic information: in: The semantic feature vector representing the e-th type of information; Represents the characteristic information of the e-th information at scale h; weight e The semantic mapping matrix representing the e-th type of information; F n Represents the segmented semantic information feature vector describing the patient information.
4. The DRG automatic grouping method based on semantic information fusion according to claim 3, characterized in that: In S23, multi-scale convolution is performed on word vector encoding sequences of different information, including: S231: Set the current convolution scale to h, where the initial value of h is 0 and the maximum value is H; S232: Generate feature information of each type of information at scale h+1, where the feature information of the e-th type of information at scale h+1 is: in: Represents the characteristic information of the e-th information at scale h+1; W h represents the convolution matrix at scale h; S233: Let h=h+1, and return to step S231 until the maximum convolution scale H is reached.
5. The DRG automatic grouping method based on semantic information fusion according to claim 2, characterized in that: The word segmentation process of medical record information and diagnosis information is as follows: S11: using a word segmentation tool to perform word segmentation processing on the medical record information and the diagnosis information to obtain an initial word segmentation sequence of the medical record information and the diagnosis information; S12: Calculate the importance weight of each initial word segmentation result in the initial word segmentation sequence corresponding to the medical record information and the diagnosis information, where the initial word segmentation sequence corresponding to the medical record information is X 1 The yth initial segmentation result X 1 The important information of (y) is: in: I 1 (y) represents the initial word segmentation sequence X corresponding to the medical record information 1 The yth initial segmentation result X 1 (y) information on the importance of exp(·) represents an exponential function with a natural constant as the base; Represents the initial word segmentation result X 1 The left neighbor information of (y), Represents the initial word segmentation result X 1 Right neighbor information of (y); P 1 (y,y-1) represents the initial word segmentation result X 1 (y) and the initial word segmentation result X 1 (y-1) The probability of appearing in the same sentence in all medical records; P 1 (y) represents the initial word segmentation result X 1 (y) probability of appearing in all medical records; The diagnostic information corresponds to the initial word segmentation sequence X 2 The yth initial segmentation result X 2 The important information of (y) is: in: I 2 (y) represents the initial word segmentation sequence X corresponding to the diagnostic information 2 The yth initial segmentation result X 2 (y) information on the importance of Represents the initial word segmentation result X 2 The left neighbor information of (y), Represents the initial word segmentation result X 2 Right neighbor information of (y); P 2 (y,y-1) represents the initial word segmentation result X 2 (y) and the initial word segmentation result X 2 (y-1) The probability of appearing in the same sentence in all diagnostic information; P 2 (y) represents the initial word segmentation result X 2 (y) probability of appearing in all diagnostic information; S13: Filtering the initial word segmentation results whose importance information is lower than a preset threshold, using the retained initial word segmentation results as the word segmentation phrases of the medical record information and the diagnosis information, and obtaining the word segmentation phrase sequence of the medical record information and the diagnosis information.
6. A DRG automatic grouping system based on semantic information fusion, characterized in that: The system comprises: Information collection module, used to collect patient personal information, medical history information and diagnostic information to form electronic health record data; An information extraction module is used to extract semantic information from the patient's electronic health record data, obtain a segmented semantic information feature vector describing the patient's information, and perform semantic enhancement processing on the segmented semantic information feature vector; An automatic grouping device is used to generate DRG grouping information of patients, and based on the DRG grouping information and patient personal information, cluster different patients to obtain a number of patient clusters, so as to realize a DRG automatic grouping method with semantic information fusion as described in any one of claims 1-5.
Citation Information
Patent Citations
DRG automatic grouping method and system based on semantic information fusion
CN116150698A
Semantic enhancement method by fusing medical term entity description information
CN117494724A