Data processing method and device, equipment and storage medium

By constructing the DIP grouping feature correlation coefficient set and weight matrix, and combining constraint rules to perform abnormal detection, the problem of insufficient accuracy of traditional DIP methods is solved, and the accuracy and management efficiency of medical data processing are improved.

CN120373793APending Publication Date: 2025-07-25TANGSHAN CAOFEIDIAN LIANCHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510778059.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional disease diagnosis-related grouping (DIP) methods are insufficiently accurate in medical data processing, resulting in insufficient scientific medical insurance payment and hospital management.

Method used

By obtaining the DIP grouping feature correlation coefficient set and target medical data, feature extraction and splicing are performed, weight coefficient matrix is constructed, weight allocation is performed, target disease grouping is determined, and abnormal detection is performed based on constraint rules.

Benefits of technology

It improves the accuracy of DIP grouping, helps hospitals to scientifically manage medical resources, promptly detect abnormal situations, and ensures the rational use of medical insurance funds and the quality of medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373793A_ABST
    Figure CN120373793A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, equipment and a storage medium, and belongs to the technical field of data processing, and the method comprises the steps: obtaining a DIP grouping feature correlation coefficient set and target medical data, carrying out the feature extraction of the target medical data, sequentially splicing the extracted features, and obtaining a first feature vector; the target medical data is to-be-processed medical data, and the DIP grouping feature correlation coefficient set comprises correlation coefficients between a plurality of DIP grouping features and disease groups; determining a weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; performing weight distribution on the first feature vector based on the weight coefficient matrix to obtain a weighted target feature vector, and determining a target disease category group based on the target feature vector; and performing anomaly detection on the target medical data based on the constraint rule corresponding to the target disease category group, and marking the detected abnormal data. The accuracy of medical data DIP grouping can be improved through data processing, and then the hospital management quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing. More specifically, it relates to a data processing method, apparatus, device, and storage medium. Background Art

[0002] In the field of medical data processing, Diagnosis Related Groups (DIP) is of great significance for medical insurance payment, hospital management, and medical quality assessment. However, with the advancement of medical informatization construction, medical data has shown an explosive growth trend. These medical data are characterized by large scale, wide sources, and complex structures; traditional DIP grouping methods have problems of insufficient accuracy. Summary of the Invention

[0003] Objective of the Application It is to provide a data processing method, apparatus, device, and storage medium, which improve the accuracy of DIP grouping of medical data through data processing, thereby improving the quality of hospital management.

[0004] In the first aspect of the embodiments of this application, a data processing method is provided, including: Obtain a DIP grouping feature correlation coefficient set and target medical data, extract features from the target medical data, and sequentially splice the extracted features to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP grouping feature correlation coefficient set includes the correlation coefficients between multiple DIP grouping features and disease groupings; the DIP grouping feature correlation coefficient set is calculated based on historical medical data; Determine the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain a weighted target feature vector, and determine the target disease grouping based on the target feature vector; Perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease grouping, and mark the detected abnormal data.

[0005] In the second aspect of the embodiments of this application, a data processing apparatus is provided, including: A data acquisition module, configured to obtain a DIP grouping feature correlation coefficient set and target medical data, extract features from the target medical data, and sequentially splice the extracted features to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP grouping feature correlation coefficient set includes the correlation coefficients between multiple DIP grouping features and disease groupings; the DIP grouping feature correlation coefficient set is calculated based on historical medical data; The DIP grouping module is used to determine the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain the weighted target feature vector, and determine the target disease group based on the target feature vector. The data anomaly detection module is used to perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease group, and mark the detected abnormal data.

[0006] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above data processing method are implemented.

[0007] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above data processing method are implemented.

[0008] The beneficial effects of the data processing method, device, equipment, and storage medium provided by the embodiments of the present application are as follows: The embodiments of the present application calculate the DIP grouping feature correlation coefficient set through historical medical data, which can accurately measure the correlation between each feature and the disease group. Based on this, the weight coefficient matrix of the first feature vector is determined and weight assignment is performed, so that the target feature vector can more accurately reflect the internal connection between the data and the disease group, thereby significantly improving the accuracy of DIP grouping, providing a more reasonable basis for medical insurance payment, and helping the hospital manage medical resources more scientifically.

[0009] The embodiments of the present application perform anomaly detection and marking on the target medical data based on the constraint rules corresponding to the target disease group, which helps to timely discover abnormal situations in the medical data, such as unreasonable diagnoses, abnormal treatment costs, etc. This not only improves the quality of medical data, but also provides supervision for the standardization of medical behaviors, further ensuring the reasonable use of medical insurance funds and the improvement of medical service quality. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 It is a schematic flowchart of the data processing method provided by an embodiment of the present application; Figure 2 Structural block diagram of a data processing device provided in an embodiment of the present application; Figure 3 Schematic block diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0012] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0013] To make the purpose, technical solution, and advantages of the present application clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data processing method provided in an embodiment of the present application. This method can be executed by an electronic device. Specifically, this method may include S101 to S103.

[0015] S101: Obtain a DIP grouping feature correlation coefficient set and target medical data, perform feature extraction on the target medical data, and sequentially splice the extracted features to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP grouping feature correlation coefficient set includes the correlation coefficients between multiple DIP grouping features and disease groupings; the DIP grouping feature correlation coefficient set is calculated based on historical medical data.

[0016] In this embodiment, the DIP grouping is a standardized classification system for disease cost accounting in the reform of medical insurance payment methods, which groups cases based on features such as diagnosis, treatment methods, and resource consumption. The DIP grouping feature correlation coefficient set may include the correlation coefficients between various DIP grouping features such as diagnosis codes, surgical operation codes, medication categories, length of hospital stay, age, and complication indicators and the disease grouping results. Among them, the correlation coefficient may be the Pearson correlation coefficient or the Spearman rank correlation coefficient, etc., which is used to reflect the influence intensity and direction of the feature on the grouping. The target medical data refers to single-case or batch medical data to be processed, such as patient electronic medical records, inspection reports, and expense lists. In this embodiment, features related to DIP grouping need to be extracted from the target medical data. Feature vector splicing refers to combining the extracted features into a one-dimensional vector in a certain order for subsequent modeling analysis.

[0017] In this embodiment, the correlation coefficients between each DIP grouping feature and the disease grouping are calculated using historical medical data to form a set of correlation coefficients. Similar features are extracted from the target medical data to be processed and spliced into a vector. Based on the set of correlation coefficients, the influence of each feature on the grouping is determined, thereby realizing the DIP grouping of the target medical data. The logic is to find the key features and their influencing degrees that affect the grouping by analyzing historical data, and apply this rule to new data to complete the grouping.

[0018] Exemplarily, this embodiment can obtain the DIP grouping feature correlation coefficient set and the target medical data from data sources such as the hospital information system and the medical insurance database. The target medical data covers patient electronic medical records, examination reports, expense lists, etc.

[0019] This embodiment can extract corresponding features from the target medical data according to the feature types involved in the DIP grouping feature correlation coefficient set, such as diagnosis codes, surgical operation codes, medication categories, etc. For example, diagnosis codes are extracted from the electronic medical record, and medication categories are extracted from the expense list.

[0020] This embodiment can splice the extracted features in sequence according to a preset order to form a first feature vector. For example, the diagnosis code is placed first, followed by the surgical operation code, and then the medication category, etc., to form a one-dimensional vector for subsequent analysis.

[0021] S102: Determine the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain the weighted target feature vector, and determine the target disease grouping based on the target feature vector.

[0022] In this embodiment, the weight coefficient matrix is a two-dimensional matrix. The rows represent DIP grouping features (such as diagnosis codes, surgical operations, age, etc.), and the columns represent disease grouping categories (such as ADRG core groups, MDC major diagnosis categories, etc.). The matrix elements are the weight values of each feature for the corresponding grouping. Among them, the weight values are converted based on the correlation coefficients. The target feature vector is the result of weighting the first feature vector according to the weight coefficient matrix, reflecting the difference in the contribution degree of the features to different groupings.

[0023] This embodiment constructs a weight system reflecting the importance of features through the correlation coefficients between features and groupings in historical medical data, performs differential weighting on the first feature vector of the new cases to be processed, and finally determines the most likely disease grouping through weighted scores.

[0024] In this embodiment, performing weight assignment on the first feature vector based on the weight coefficient matrix may specifically include: performing weight assignment on multiple feature elements in the first feature vector based on the weight coefficient matrix.

[0025] Exemplarily, in this embodiment, the DIP group feature correlation coefficient set can be traversed. For each DIP group feature, for different disease group categories (such as ADRG core groups, MDC major diagnostic categories), its weight value is determined according to the correlation coefficient conversion rule. For example, the higher the correlation, the greater the weight value.

[0026] In this embodiment, the weight values of each DIP group feature for each disease group category can be filled into a two-dimensional matrix in the way that the rows correspond to the DIP group features and the columns correspond to the disease group categories, to complete the construction of the weight coefficient matrix. In this embodiment, each eigenvalue in the first eigenvector can be taken out, and according to the DIP group feature corresponding to this eigenvalue, the corresponding row in the weight coefficient matrix can be found.

[0027] In this embodiment, for each disease group category, the eigenvalue in the first eigenvector is multiplied by the weight value corresponding to this row in the weight coefficient matrix to obtain the weighted value of this feature for each disease group category. In this embodiment, the weighted values of all features for the same disease group category can be added together to form a target eigenvector, and each element in the vector represents the weighted score of the corresponding disease group.

[0028] In this embodiment, the elements in the target eigenvector (i.e., the weighted scores of each disease group) can be compared, and the disease group corresponding to the element with the highest score is the target disease group.

[0029] S103: Perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease group, and mark the detected abnormal data.

[0030] In this embodiment, the target disease group is the specific DIP group determined, such as the ADRG group "FB1", the MDC major category "nervous system diseases", etc., and each group corresponds to a set of standardized constraint rules. Among them, the constraint rules can include: diagnosis rules, treatment rules, resource consumption rules, and logical rules, etc. Among them, the diagnosis rules can include: the main diagnosis code belongs to the disease category defined by the group. For example, the "diabetes with complications" group requires the main diagnosis to be E10 - E14 and there is at least 1 complication code. The treatment rules can include: the surgical / operation code matches the diagnosis. For example, the "hip replacement" group requires the operation code to include 81.51 - 81.53 and there is no contraindicated operation code such as 98.25. The resource consumption rules can include: the reasonable intervals of indicators such as the length of hospital stay, total cost, and drug proportion. For example, the mean length of hospital stay in the historical data of a certain group ± 2 times the standard deviation. The logical rules can include: the logical consistency between diagnosis and treatment. For example, in the "pneumonia" group, if ventilator treatment is used, the diagnosis related to acute respiratory failure needs to be marked. Abnormal data refers to data that violates any of the above rules, such as the main diagnosis code is not within the allowed range of the group, the cost exceeds 150% of the group budget upper limit, etc.

[0031] In this embodiment, the constraint rules corresponding to the grouped target disease types are transformed into computable verification conditions, and each feature of the target medical data is matched item by item. The abnormal items that do not conform to the grouped definition are identified through a rule engine or an algorithm model.

[0032] Exemplarily, for each grouped target disease type, this embodiment can respectively parse its corresponding diagnosis rules, treatment rules, resource consumption rules, and logical rules into specific computable verification conditions. For example, for the diagnosis rules of the "diabetes mellitus with complications" group, it is parsed into judging whether the main diagnosis code is within the range of E10 - E14 and whether there is at least one complication code.

[0033] This embodiment can transform the matching requirements of the operation codes in the treatment rules into judgment conditions for whether the operation codes are within the specified range and do not include prohibited operation codes. For the resource consumption rules, corresponding numerical comparison conditions are set according to the reasonable ranges of indicators such as length of hospital stay, total cost, and drug proportion. The logical rules are transformed into logical judgment conditions between diagnosis and treatment. For example, when using a ventilator for treatment in the "pneumonia" group, it is judged whether the relevant diagnosis of acute respiratory failure is marked.

[0034] This embodiment can extract features related to each rule from the target medical data, such as diagnosis codes, surgical operation codes, length of hospital stay, total cost, drug proportion, etc. This embodiment can match the extracted features with the transformed verification conditions item by item. For example, comparing the main diagnosis code with the coding range of the diagnosis rules; comparing the surgical operation code with the ranges of permitted and prohibited operation codes in the treatment rules; comparing the numerical values such as length of hospital stay, total cost, and drug proportion with the reasonable ranges set by the resource consumption rules; and checking the logical consistency between diagnosis and treatment according to the logical rules.

[0035] If a certain feature of the target medical data does not conform to the corresponding verification condition, this embodiment can determine that the data is abnormal data. This embodiment can mark the detected abnormal data and record information such as the type of abnormality (such as diagnosis abnormality, treatment abnormality, etc.), the specific rules involved, and the feature values that do not conform to the rules for subsequent analysis and processing.

[0036] As can be seen from the above, the embodiment of the present application calculates the DIP grouping feature correlation coefficient set through historical medical data, which can accurately measure the correlation between each feature and the disease type grouping. Based on this, determining the weight coefficient matrix of the first feature vector and performing weight assignment can enable the target feature vector to more accurately reflect the internal connection between the data and the disease type grouping, thereby significantly improving the accuracy of DIP grouping, providing a more reasonable basis for medical insurance payment, and helping the hospital to manage medical resources more scientifically.

[0037] The embodiments of this application perform anomaly detection and marking on target medical data based on the constraint rules corresponding to the target disease groups, which helps to promptly discover anomalies in medical data, such as unreasonable diagnoses and abnormal treatment costs. This not only improves the quality of medical data but also provides supervision for the standardization of medical behaviors, further ensuring the rational use of medical insurance funds and the improvement of the quality of medical services.

[0038] In one embodiment of this application, the historical medical data includes the medical data of multiple patients, and each patient's medical data corresponds to a patient identifier and a pre-divided DIP group. The DIP group feature correlation coefficient set is determined based on the historical medical data in the following manner: feature extraction is performed on the historical medical data, and the features with the same patient common identifier obtained by extraction are sequentially concatenated to obtain a first feature set; the first feature set is divided based on the pre-divided DIP groups to obtain multiple feature subsets, and each feature subset corresponds one-to-one to a DIP group; the first feature similarity of the features in each feature subset is calculated, and the second feature similarity of the features between all feature subsets is calculated; the DIP group feature correlation coefficient set is determined based on the first feature similarity and the second feature similarity.

[0039] In this embodiment, performing feature extraction on the historical medical data specifically includes: dividing the historical medical data into multiple medical data subsets based on the data type and data source, and performing feature extraction on the multiple medical data subsets respectively; the data type includes unstructured type, semi-structured type, and structured type; the data source includes patient source data and diagnosis and treatment source data; each medical data subset is labeled with a data type label and a data source label.

[0040] In this embodiment, the feature subset includes multiple second feature vectors, and each second feature vector corresponds to a patient common identifier; the second feature vector includes patient data features and diagnosis and treatment data features, and the patient data features include unstructured patient features, semi-structured patient features, and structured patient features; the diagnosis and treatment data features include unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features. Calculating the first feature similarity of the features in each feature subset specifically includes: performing standardization processing on the unstructured patient features, semi-structured patient features, structured patient features, unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features in each second feature vector respectively to obtain a standardized third feature vector; for each DIP group feature in the third feature vector, the first feature similarity of the DIP group feature is calculated based on the values of the DIP group feature in all third feature vectors; the first feature similarity is directly proportional to the correlation coefficient between the DIP group feature and the disease group.

[0041] In this embodiment, calculating the second feature similarity between features of all feature subsets specifically includes: calculating the statistical mean of all DIP grouped features in each feature subset; for each DIP grouped feature, calculating the second feature similarity based on the statistical mean of this DIP grouped feature in all feature subsets; the second feature similarity is inversely proportional to the correlation coefficient between this DIP grouped feature and the disease group.

[0042] In this embodiment, the patient public identifier refers to the primary key that uniquely identifies a patient and is used to integrate the medical data of the same patient across data sources. Structured data refers to standardized data that can be directly recorded in a database table, such as age, diagnostic code ICD-10, expense amount, etc. Semi-structured data refers to data with a certain format but not completely tabular, such as surgical records in XML format, doctor's order lists in JSON format, etc. Unstructured data refers to pure text or multimedia data, such as medical records, medical imaging report texts, etc. Patient source data refers to the personal information of a patient, which may include the patient's basic information, demographic characteristics, such as place of residence, medical insurance type, etc. Diagnostic and treatment source data may include diagnosis-related data and treatment-related data, such as surgical operation codes, medication records, examination reports, etc.

[0043] The second feature vector is a vector that integrates all-dimensional features of patient data, and its structure is: [unstructured patient features, semi-structured patient features, structured patient features, unstructured diagnostic and treatment features, semi-structured diagnostic and treatment features, structured diagnostic and treatment features]. The first feature similarity is used to measure the consistency of a certain feature within the same DIP group (a feature subset), and the higher the value, the more similar the feature is in the patient data within the group. For example, the "troponin" values of patients in the "acute myocardial infarction" group fluctuate little and are directly proportional to the group correlation. The second feature similarity is used to measure the similarity of a certain feature between different DIP groups, and the higher the value, the lower the discrimination degree of the feature between the groups. For example, "gender" is evenly distributed in most disease groups and is inversely proportional to the group correlation.

[0044] In this embodiment, by jointly measuring the within-group feature consistency (the first similarity) and the between-group feature difference (the second similarity), features with strong discrimination ability for DIP grouping are screened. If a certain feature shows high similarity among patients in the target group, then the contribution of this feature to the grouping is high, that is, the first similarity is high. If the distribution difference of a certain feature between different groups is small, then the discrimination ability of this feature for the grouping is weak, that is, the second similarity is high.

[0045] Exemplarily, in this embodiment, historical medical data can be divided into multiple medical data subsets based on data types (unstructured, semi-structured, structured) and data sources (patient source, diagnosis and treatment source), and each subset is labeled with a data type label and a data source label. For example, the patient's medical record can be divided into an unstructured patient source data subset, and the diagnosis code can be divided into a structured diagnosis and treatment source data subset.

[0046] This embodiment can perform feature extraction on each medical data subset respectively. For the structured data subset, data such as age, diagnosis code, and cost amount are directly extracted as features; for the semi-structured data subset, by parsing formats such as XML and JSON, key information in surgical records and doctor's order lists is extracted as features; for the unstructured data subset, natural language processing technology is used to extract key features from medical records and medical imaging report texts.

[0047] This embodiment can use the patient public identifier to sequentially splice the features obtained with the same patient public identifier to form a first feature set that integrates the full-dimensional features of the patient. For example, the features extracted from the age, diagnosis code, and medical record of the same patient are spliced together.

[0048] This embodiment can divide the first feature set into multiple feature subsets according to the pre-divided DIP groups, and each feature subset corresponds to a DIP group one by one. For example, the features of all patients belonging to the "pneumonia" DIP group are grouped into one feature subset.

[0049] This embodiment can perform standardization processing on various features in the second feature vector in each feature subset to obtain a standardized third feature vector. For each DIP group feature in the third feature vector, this embodiment can calculate the first feature similarity based on the values of this DIP group feature in all third feature vectors. For example, in the feature subset of the "pneumonia" group, calculate the similarity of the "body temperature" DIP group feature of all patients. The higher the first feature similarity, the more similar the feature is among the patients in the group, and it is in a direct proportional relationship with the correlation coefficient between this DIP group feature and the disease group.

[0050] This embodiment can first calculate the statistical mean of all DIP group features in each feature subset. For example, calculate the mean of "white blood cell count" in the feature subset of the "pneumonia" group features.

[0051] For each DIP group feature, this embodiment can calculate the second feature similarity based on the statistical mean of this DIP group feature in all feature subsets. The higher the second feature similarity, the lower the discrimination of this feature among different groups, and it is in an inverse proportional relationship with the correlation coefficient between this DIP group feature and the disease group.

[0052] This embodiment can comprehensively consider the first feature similarity and the second feature similarity to determine the DIP grouping feature correlation coefficient set. Through this joint measurement of intra-group feature consistency and inter-group feature difference, features with strong discrimination ability for DIP grouping are screened out to form the DIP grouping feature correlation coefficient set.

[0053] This embodiment divides data from multiple dimensions of data type and source and extracts features, comprehensively covering all aspects of patient medical data. By dividing the feature subsets and calculating the first feature similarity and the second feature similarity, features with strong discrimination ability for DIP grouping can be accurately screened out. This embodiment jointly measures the intra-group and inter-group feature similarities, making the screened features more accurately reflect the characteristics of different DIP groupings. The correlation coefficient set composed of these features provides a more reliable basis for determining weights and groupings based on features in the subsequent process, thereby improving the accuracy of DIP grouping and contributing to more reasonable medical insurance payment and more scientific allocation of medical resources.

[0054] This embodiment's meticulous screening and measurement method of features reduces the interference of redundant and irrelevant features, enhancing the stability of the DIP grouping model constructed based on these features. When facing complex and variable medical data, the model can run more robustly, reducing grouping deviations caused by data fluctuations or feature interferences.

[0055] In one embodiment of the present application, determining the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set includes: For each feature element in the first feature vector, match the feature element with multiple DIP grouping features in the DIP grouping feature correlation coefficient set, and use the correlation coefficient between the successfully matched DIP grouping feature and the disease group as the weight coefficient of the feature element; Normalize the weight coefficients corresponding to all feature elements to obtain the weight coefficient matrix.

[0056] In this embodiment, the feature element refers to the specific data item in the first feature vector, such as the patient's age, diagnosis code, length of hospital stay, etc. The weight coefficient is used to measure the importance of the feature element for different disease groups, which is transformed from the correlation coefficient between the DIP grouping feature and the disease group. The higher the correlation, the greater the weight. The weight coefficient matrix is a two-dimensional matrix, where the rows correspond to the feature elements of the first feature vector, the columns correspond to different disease groups, and the matrix elements are the weights of the feature elements for specific disease groups, which are used for subsequent weighted calculations.

[0057] In this embodiment, the relevance of features is transformed into weights, enabling features with a greater impact on grouping to take precedence in decision-making. First, the correlation coefficient of feature elements is found through matching to determine the initial weights. Then, through normalization, the sum of the weights of the same feature for each disease grouping is made 1, ensuring the comparability of the weights. Finally, a structured weight system is formed, providing a basis for weighted summation of target feature vectors and disease grouping decisions.

[0058] Exemplarily, this embodiment can traverse each feature element in the first feature vector, such as patient age, diagnosis code, length of hospital stay, etc. For each feature element, it is compared one by one with multiple DIP grouping features in the DIP grouping feature correlation coefficient set. If a matching DIP grouping feature is found, the correlation coefficient between this DIP grouping feature and the disease grouping is used as the initial weight coefficient of this feature element. For example, if the feature element of the diagnosis code matches a specific code in the DIP grouping features, the correlation coefficient between this code and the disease grouping is set as the weight coefficient of the diagnosis code feature element.

[0059] This embodiment can perform normalization processing on the initial weight coefficients determined for all feature elements. Taking the weight coefficients of each feature element for different disease groupings as a set of data, calculate the sum of this set of data, and then divide each weight coefficient by the sum, so that the sum of the weights of the same feature element for each disease grouping is 1. For example, if the weight coefficients of a certain feature element for three disease groupings are 0.4, 0.3, and 0.3 respectively, and the sum is 1, no adjustment is required; if they are 0.5, 0.3, and 0.2 respectively, after normalization, they become 0.5, 0.3, and 0.2 (the sum is 1).

[0060] This embodiment can arrange the normalized weight coefficients in order to form a two-dimensional matrix. The rows of the matrix correspond to the feature elements of the first feature vector, the columns correspond to different disease groupings, and each matrix element is the weight coefficient of the feature element for a specific disease grouping. Thus, the construction of the weight coefficient matrix is completed.

[0061] In this embodiment, by transforming feature relevance into weights, features with a greater impact on grouping are highlighted, and priority is given to them in disease grouping decisions, improving grouping accuracy. Normalization processing makes the weights comparable, enabling the impacts of different features on each disease grouping to be measured on the same scale. The structured weight system provides an accurate basis for weighting target feature vectors, making subsequent disease grouping decisions more scientific and reasonable, ultimately improving the quality and efficiency of DIP grouping, and providing reliable support for medical insurance payment and medical management.

[0062] In an embodiment of the present application, before obtaining the DIP grouping feature correlation coefficient set, it further includes: Obtaining historical multi-source medical data, performing data preprocessing on the historical multi-source medical data to obtain historical medical data; Data preprocessing includes missing value processing, duplicate value processing, and noisy data processing.

[0063] In this embodiment, historical multi-source medical data refers to past medical information collected from multiple channels, covering hospital electronic medical record systems, inspection and examination systems, medical insurance settlement systems, etc., and includes multi-dimensional data such as patient basic information, diagnosis records, treatment plans, and expense details. Data preprocessing refers to cleaning and transforming the original data to improve data quality and ensure the accuracy and reliability of subsequent analysis. Missing values refer to the lack of information in some fields in the dataset, such as key data like patient age and diagnosis codes not being recorded. Duplicate values refer to the repeated entry of a patient's medical records multiple times, or the same data appearing repeatedly in different tables. Noisy data refers to the presence of errors, anomalies, or interfering information in the data, such as ages that are clearly illogical or incorrect diagnosis code formats.

[0064] In this embodiment, the original medical data often has incomplete, duplicate, or incorrect situations, and directly using it for analysis will lead to result deviations. By handling missing values to complete key information, handling duplicate values to avoid data redundancy interference, and handling noisy data to eliminate incorrect content, the historical medical data can be made more accurate and standardized, providing a reliable data basis for calculating the DIP grouping feature correlation coefficient set later and ensuring the scientificity and effectiveness of the analysis results.

[0065] Exemplarily, the missing value processing of historical multi-source medical data can include: Numeric missing values: For numeric missing values such as patient age, first calculate the mean and median of this feature. If the data distribution is uniform, fill it with the mean; if there are outliers and the distribution is non-uniform, then use the median to fill. It is also possible to use the K-nearest neighbor algorithm to find similar samples based on other features and fill with the corresponding feature mean of these samples.

[0066] Non-numeric missing values: For non-numeric missing values such as diagnosis codes, count the mode of this feature and fill the missing value with the mode. Or build a classification model (such as a decision tree) and use other features to predict the missing value.

[0067] The duplicate value processing of historical multi-source medical data can include: Record-level duplicates: Identify and delete exactly duplicate medical records through unique identifiers (such as patient ID, medical record number). If there is no unique identifier, calculate the hash value for key information (such as patient basic information, diagnosis records), and consider records with the same hash value as duplicates and delete them.

[0068] Field-level duplicates: For the situation where the same data appears repeatedly in different tables, establish a data mapping relationship, merge duplicate fields, and retain the most accurate or latest data.

[0069] Processing of noisy data in historical multi-source medical data may include: Numerical noise: For numerical noise such as age, the 3σ principle is used to identify outliers, which are corrected to reasonable boundary values or replaced with linear regression predicted values.

[0070] Coding noise: For incorrect diagnostic coding formats, they are corrected according to standard coding rules, and the codes that do not conform to the rules are eliminated.

[0071] Through the preprocessing of historical multi-source medical data in this embodiment, problems such as incomplete, duplicate, and incorrect original data can be effectively solved. Key information is supplemented through missing value processing, redundancy is eliminated through duplicate value processing, and errors are corrected through noisy data processing, laying a reliable foundation for calculating the DIP grouping feature correlation coefficient set and improving the scientificity and effectiveness of DIP grouping analysis results.

[0072] Corresponding to the data processing method in the above embodiment, Figure 2 is a structural block diagram of a data processing device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown. Refer to Figure 2 , the data processing device 20 includes: a data acquisition module 21, a DIP grouping module 22, and a data anomaly detection module 23.

[0073] Among them, the data acquisition module 21 is used to acquire the DIP grouping feature correlation coefficient set and target medical data, extract features from the target medical data, and splice the extracted features in sequence to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP grouping feature correlation coefficient set includes the correlation coefficients between multiple DIP grouping features and disease groupings; the DIP grouping feature correlation coefficient set is calculated based on historical medical data; The DIP grouping module 22 is used to determine the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain a weighted target feature vector, and determine the target disease grouping based on the target feature vector; The data anomaly detection module 23 is used to perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease grouping, and mark the detected abnormal data.

[0074] In an embodiment of the present application, the historical medical data includes the medical data of multiple patients, and each patient's medical data corresponds to a patient identifier and a pre-divided DIP grouping; the data acquisition module 21 is specifically used to extract features from the historical medical data, and splice the features with the same patient public identifier obtained by extraction in sequence to obtain a first feature set; Divide the first feature set based on pre-divided DIP groups to obtain multiple feature subsets, with each feature subset corresponding to a DIP group one by one; Calculate the first feature similarity of the features in each feature subset, and calculate the second feature similarity of the features between all feature subsets; Determine the DIP group feature correlation coefficient set based on the first feature similarity and the second feature similarity.

[0075] In an embodiment of the present application, the data acquisition module 21 is further specifically configured to divide the historical medical data into multiple medical data subsets based on the data type and data source, and perform feature extraction on the multiple medical data subsets respectively; The data type includes unstructured type, semi-structured type, and structured type; The data source includes patient source data and diagnosis and treatment source data; Each medical data subset is labeled with a data type label and a data source label.

[0076] In an embodiment of the present application, the feature subset includes multiple second feature vectors, and each second feature vector corresponds to a patient public identifier; the second feature vector includes patient data features and diagnosis and treatment data features, and the patient data features include unstructured patient features, semi-structured patient features, and structured patient features; the diagnosis and treatment data features include unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features; the data acquisition module 21 is further specifically configured to perform standardization processing on the unstructured patient features, semi-structured patient features, structured patient features, unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features in each second feature vector respectively to obtain a standardized third feature vector; For each DIP group feature in the third feature vector, calculate the first feature similarity of the DIP group feature based on the values of the DIP group feature in all third feature vectors; the first feature similarity is directly proportional to the correlation coefficient between the DIP group feature and the disease group.

[0077] In an embodiment of the present application, the data acquisition module 21 is further specifically configured to calculate the statistical mean of all DIP group features in each feature subset; For each DIP group feature, calculate the second feature similarity based on the statistical mean of the DIP group feature in all feature subsets; the second feature similarity is inversely proportional to the correlation coefficient between the DIP group feature and the disease group.

[0078] In an embodiment of the present application, the data acquisition module 21 is further specifically configured to, for each feature element in the first feature vector, match the feature element with multiple DIP group feature correlation coefficients in the DIP group feature correlation coefficient set, and use the correlation coefficient between the successfully matched DIP group feature and the disease group as the weight coefficient of the feature element; Normalize the weight coefficients corresponding to all feature elements to obtain a weight coefficient matrix.

[0079] In an embodiment of the present application, before obtaining the DIP group feature correlation coefficient set, the data processing device 20 further includes: acquiring historical multi-source medical data, and performing data preprocessing on the historical multi-source medical data to obtain historical medical data; The data preprocessing includes missing value processing, duplicate value processing, and noise data processing.

[0080] Refer to Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided in an embodiment of the present application. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned device embodiments, such as Figure 2 the functions of the data acquisition module 21, the DIP grouping module 22, and the data anomaly detection module 23 shown.

[0081] It should be understood that in the embodiments of the present application, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0082] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0083] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may further include a non-volatile random access memory. For example, the memory 304 may further store information of medical data.

[0084] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present application may execute the implementation manners described in the embodiments of the data processing method provided in the embodiments of the present application, and may also execute the implementation manner of the electronic device 300 described in the embodiments of the present application, which will not be elaborated herein.

[0085] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0086] The computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may further include both an internal storage unit and an external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0087] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0088] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0089] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical, or other forms of connection.

[0090] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.

[0091] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0092] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, Including: Obtain the DIP grouping feature correlation coefficient set and the target medical data, perform feature extraction on the target medical data, and splice the extracted features in sequence to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP grouping feature correlation coefficient set includes the correlation coefficients between multiple DIP grouping features and disease groupings; the DIP grouping feature correlation coefficient set is calculated based on historical medical data; Determine the weight coefficient matrix corresponding to the first feature vector based on the DIP grouping feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain the weighted target feature vector, and determine the target disease grouping based on the target feature vector; Perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease grouping, and mark the detected abnormal data.

2. The data processing method according to claim 1, characterized in that The historical medical data includes the medical data of multiple patients, and the medical data of each patient corresponds to a patient identifier and a pre-divided DIP grouping; The DIP grouping feature correlation coefficient set is determined based on historical medical data in the following manner: Perform feature extraction on the historical medical data, and splice the extracted features with the same patient public identifier in sequence to obtain a first feature set; Divide the first feature set based on the pre-divided DIP groupings to obtain multiple feature subsets, and each feature subset corresponds to a DIP grouping one by one; Calculate the first feature similarity of the features in each feature subset, and calculate the second feature similarity of the features between all feature subsets; Determine the DIP grouping feature correlation coefficient set based on the first feature similarity and the second feature similarity.

3. The data processing method according to claim 2, characterized in that The performing feature extraction on the historical medical data includes: Divide the historical medical data into multiple medical data subsets based on the data type and data source, and perform feature extraction on the multiple medical data subsets respectively; The data types include unstructured type, semi-structured type, and structured type; The data sources include patient source data and diagnosis and treatment source data; Each medical data subset is labeled with a data type label and a data source label.

4. The data processing method according to claim 2, wherein The feature subset includes multiple second feature vectors, and each second feature vector corresponds to a patient public identifier; the second feature vector includes patient data features and diagnosis and treatment data features, and the patient data features include unstructured patient features, semi-structured patient features, and structured patient features; the diagnosis and treatment data features include unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features; The calculating the first feature similarity of the features in each feature subset includes: Perform standardization processing on the unstructured patient features, semi-structured patient features, structured patient features, the unstructured diagnosis and treatment features, semi-structured diagnosis and treatment features, and structured diagnosis and treatment features in each second feature vector respectively to obtain a standardized third feature vector; For each DIP group feature in the third feature vector, calculate the first feature similarity of the DIP group feature based on the values of the DIP group feature in all the third feature vectors; the first feature similarity is directly proportional to the correlation coefficient between the DIP group feature and the disease group.

5. The data processing method according to claim 4, wherein, The calculating the second feature similarity between features of all feature subsets includes: Calculate the statistical mean of all DIP group features in each feature subset; For each DIP group feature, calculate the second feature similarity based on the statistical mean of the DIP group feature in all feature subsets; the second feature similarity is inversely proportional to the correlation coefficient between the DIP group feature and the disease group.

6. The data processing method according to claim 1, wherein The determining the weight coefficient matrix corresponding to the first feature vector based on the DIP group feature correlation coefficient set includes: For each feature element in the first feature vector, match the feature element with multiple DIP group features in the DIP group feature correlation coefficient set, and use the correlation coefficient between the successfully matched DIP group feature and the disease group as the weight coefficient of the feature element; Normalize the weight coefficients corresponding to all feature elements to obtain the weight coefficient matrix.

7. The data processing method according to claim 1, wherein Before obtaining the DIP group feature correlation coefficient set, it further includes: Obtain historical multi-source medical data, and perform data preprocessing on the historical multi-source medical data to obtain historical medical data; The data preprocessing includes missing value processing, duplicate value processing, and noise data processing.

8. A data processing device, characterized in that, It includes: A data acquisition module, used to obtain a DIP group feature correlation coefficient set and target medical data, extract features from the target medical data, and splice the extracted features in sequence to obtain a first feature vector; the target medical data is the medical data to be processed, and the DIP group feature correlation coefficient set includes the correlation coefficients between multiple DIP group features and disease groups; the DIP group feature correlation coefficient set is calculated based on historical medical data; A DIP grouping module, used to determine the weight coefficient matrix corresponding to the first feature vector based on the DIP group feature correlation coefficient set; perform weight assignment on the first feature vector based on the weight coefficient matrix to obtain a weighted target feature vector, and determine a target disease group based on the target feature vector; A data anomaly detection module, used to perform anomaly detection on the target medical data based on the constraint rules corresponding to the target disease group, and mark the detected abnormal data.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.