Patient abnormal data detection method and system for medical information system

By further confirming and analyzing suspected abnormal data in the medical information system, combining the status of medication and the degree of impact, the problem of high error detection rate in the prior art is solved, and the accuracy and efficiency of data processing are improved.

CN120123948AInactive Publication Date: 2025-06-10BEIJING SHIKU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503711.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing abnormality detection algorithms are prone to misjudging changes in normal physiological data as abnormal data in medical information systems, resulting in false detection and affecting the data processing effect.

Method used

A method for detecting abnormal data of patients in medical information systems is proposed. By obtaining medical data sequences and timing sequences, the abnormality detection algorithm is used to screen out suspected abnormal data and conduct further confirmation and analysis. For a single suspected abnormal data, determine whether it is an abnormality caused by normal physiological behavior based on the data type and medication status; for multiple suspected abnormal data, a matrix to be analyzed, and the degree of influence between different types of data is analyzed using the maximum likelihood estimation algorithm to be used to screen out the real abnormal data.

Benefits of technology

Effectively prevent data misdetection, improve the data processing effect of medical information systems, and ensure the accuracy of abnormal data detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123948A_ABST
    Figure CN120123948A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of abnormal data detection, in particular to a patient abnormal data detection method and system for a medical information system. According to the method, after suspected abnormal data is screened out by using an existing anomaly detection algorithm, further confirmation analysis is performed on the suspected abnormal data. For the condition that only single suspected abnormal data appears at one detection moment, whether the abnormality is caused by normal physiological behaviors or not can be judged according to the data type in combination with the medicine taking state, and then an abnormality label is given. And for the condition that multiple pieces of suspected abnormal data appear at one detection moment, constructing a matrix to be analyzed, analyzing the influence degree among different types of suspected abnormal data by using a maximum likelihood estimation algorithm, screening out real abnormal data according to the influence condition, and endowing the real abnormal data with a label. According to the invention, through secondary detection of the suspected abnormal data, false detection of the data can be effectively prevented, and the data processing effect of the medical information system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal data detection, and particularly to a method and system for detecting abnormal patient data for a medical information system. Background Art

[0002] Medical information systems can store a large amount of patient data and provide functions such as data analysis and decision support. However, due to various accidental factors in the system, such as human operation errors and equipment failures, there are outliers in the patient data, and these outliers need to be detected in time to avoid affecting the normal use of the medical information system.

[0003] In the prior art, outlier detection algorithms such as the Isolation Forest algorithm and the Local Outlier Factor algorithm can analyze the numerical changes of patient data in time series to determine whether the data is abnormal data. However, for some data determined to be abnormal, it may itself be normal physiological data due to factors such as taking medicine and diseases, and the misjudgment of such abnormal data will further affect the data processing effect in the medical information system. Summary of the Invention

[0004] In order to solve the technical problem that the existing outlier detection algorithms will misdetect medical data, the purpose of the present invention is to provide a method and system for detecting abnormal patient data for a medical information system, and the specific technical solutions adopted are as follows: The present invention proposes a method for detecting abnormal patient data for a medical information system, and the method includes: Obtain the medical data sequence of the target patient at each detection moment, where each element in the medical data sequence represents a type of medical data, and the medical data includes basic data and complex data; obtain the data time series sequence of each type of medical data of the target patient in time series; Use an outlier detection algorithm to obtain the suspected abnormal data in the data time series sequence; regard the medical data sequence containing only one suspected abnormal data as the first sequence to be analyzed, and regard the other medical data sequences containing suspected abnormal data as the second sequences to be analyzed; If the suspected abnormal data in the first sequence to be analyzed is complex data, directly assign an abnormal label; if it is basic data, determine the abnormal index according to the patient's medication situation, and judge whether to assign an abnormal label according to the abnormal index; Arrange the second sequences to be analyzed and the medical data sequences within the time series range to obtain a matrix to be analyzed; for any two suspected abnormal data in the matrix to be analyzed, use the maximum likelihood estimation algorithm to determine the influence degree between the two suspected abnormal data, and judge the influence situation of the two suspected abnormal data; assign abnormal labels to the suspected abnormal data that do not affect each other in the second sequences to be analyzed.

[0005] Further, the anomaly detection algorithm is the Isolation Forest algorithm, which obtains the normal probability of each medical data based on the distribution position and depth of the medical data in the isolation tree, and filters out the suspected abnormal data according to the normal probability.

[0006] Further, the method for obtaining the normal probability includes: For each medical data, obtain the path length between the medical data and the root node in the isolation tree; obtain the average path length of other medical data within the preset numerical range of the medical data; use the ratio of the path length to the depth of the isolation tree as the position weight; Obtain the normal probability according to the difference between the path length and the average path length, and the position weight.

[0007] Further, the method for obtaining the anomaly index includes: If the patient is in a medication state at the detection time corresponding to the suspected abnormal data, set the anomaly index to a first preset value; otherwise, set it to a second preset value; where the first preset value is less than the second preset value.

[0008] Further, the determination of whether to assign an anomaly label according to the anomaly index includes: If the anomaly index is the second preset value, assign an anomaly label to the corresponding suspected abnormal data.

[0009] Further, the method for obtaining the influence degree includes: Use the maximum likelihood estimation algorithm to count the frequencies of two types of suspected abnormal data in each change situation, obtain the goodness-of-fit between the two medical data using the goodness-of-fit formula; obtain the maximum inference probability less than the goodness-of-fit in the chi-square test table data, and perform a negative correlation mapping to obtain the influence degree.

[0010] Further, the determination of the influence situation between two types of suspected medical data includes: If the influence degree is greater than the preset influence degree threshold, it is determined that the two types of suspected medical data influence each other.

[0011] Further, the filtering out of the suspected abnormal data according to the normal probability includes: Use the medical data with a normal probability less than the preset normal probability threshold as the suspected abnormal data.

[0012] Further, each column of the matrix to be analyzed represents a type of medical data, and each row represents the medical data sequence corresponding to a detection time.

[0013] The present invention also provides a detection system for abnormal patient data in a medical information system, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the detection methods for abnormal patient data in a medical information system are implemented.

[0014] The present invention has the following beneficial effects: After screening out suspected abnormal data using the existing anomaly detection algorithm, the present invention further performs confirmation analysis on the suspected abnormal data. For the case where only a single suspected abnormal data appears at a detection moment, since there is no influence from other types of medical data, simple analysis can be performed in this case. According to the data type and combined with the medication status, it can be determined whether the abnormality is caused by normal physiological behavior, and then an accurate abnormal label can be assigned. For the case where multiple suspected abnormal data appear at a detection moment, the present invention further constructs a matrix to be analyzed, and uses the maximum likelihood estimation algorithm to analyze the influence degree between different types of suspected abnormal data. According to the influence situation, the truly abnormal data is screened out and labeled. Through the secondary detection of suspected abnormal data, the present invention can effectively prevent misdetection of data and improve the data processing effect of the medical information system. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart of a detection method for abnormal patient data in a medical information system provided by an embodiment of the present invention. Detailed Embodiments

[0017] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific embodiments, structures, features, and effects of a detection method and system for abnormal patient data in a medical information system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0019] The following specifically describes the specific solutions of a method and system for detecting abnormal patient data in a medical information system provided by the present invention in conjunction with the accompanying drawings.

[0020] Please refer to Figure 1 , which shows a flowchart of a method for detecting abnormal patient data in a medical information system provided by an embodiment of the present invention. The method includes: Step S1: Obtain the medical data sequence of the target patient at each detection moment. Each element in the medical data sequence represents a type of medical data, and the medical data includes basic data and complex data; obtain the data time series sequence of each type of medical data of the target patient in time series.

[0021] In the medical information system, various types of medical data of patients can be stored and called for analysis at any time. Among them, medical data at least includes information such as vital signs, examination results, and medical orders. It should be noted that when the system stores data, it is necessary to perform dimensionless normalization on the data, convert structured data such as blood pressure and blood sugar into a unified unit, extract key indicators from data such as impact reports and medical order cases through natural language processing methods, and then clean and normalize the data, so that the medical information system can effectively analyze and store the medical data of patients. The preprocessing algorithms for data storage are all well-known technical means to those skilled in the art and will not be elaborated here.

[0022] In the embodiment of the present invention, the medical data sequence represents a set of multi-dimensional data obtained after a medical examination of the target patient, where each element in the sequence represents a medical data. The embodiment of the present invention divides medical data into basic data and complex data. Among them, basic data is medical data such as heart rate and body temperature that can directly and simply characterize the human body state; complex data is data such as blood sugar, insulin, and blood lipid that has obvious complex correlation changes. When the medical system stores data, type labels can be set artificially in advance to assign basic or complex labels to each medical data.

[0023] Furthermore, in order to extract abnormal data, it is necessary to perform temporal analysis on the medical data. Therefore, the data time series sequence of each type of medical data of the target patient in time series is obtained, that is, each element in the data time series sequence represents the corresponding medical data of the target patient at the corresponding moment.

[0024] Step S2: Use an anomaly detection algorithm to obtain suspected abnormal data in the data time series sequence; use the medical data sequence containing only one suspected abnormal data as the first sequence to be analyzed, and use other medical data sequences containing suspected abnormal data as the second sequences to be analyzed.

[0025] The anomaly detection algorithm can analyze the abnormal fluctuation data in the time series of data time series, and then determine the suspected abnormal data. Since this kind of abnormal fluctuation may be affected by the patient's own physiological data or related to the medication status, the suspected abnormal data cannot be directly identified as real abnormal data, and further analysis of the suspected abnormal data is needed to determine its real attributes.

[0026] It should be noted that the real abnormal data in the embodiments of the present invention refers to the data abnormality caused by the hardware factors or human factors of the medical system, rather than the physiological abnormal data of the patient himself. The embodiments of the present invention aim to further analyze the suspected abnormal data, screen out the abnormal fluctuation data caused by the physiological abnormalities of the patient himself, and determine the abnormal data generated by the system.

[0027] The embodiments of the present invention consider that for each detection moment, the number of suspected abnormal data included is different, and the analysis method should be different. If only one suspected abnormal data is included, it can be specifically analyzed whether the data is real abnormal data according to the medical data type of the data and the patient's medication status; if two or more suspected abnormal data are included, it is necessary to consider whether some suspected abnormal data are normal fluctuations caused by normal physiological activities, and further analysis of the changes between the suspected abnormal data is needed to avoid misidentification. Therefore, the embodiments of the present invention use the medical data sequence containing only one suspected abnormal data as the first sequence to be analyzed, and use other medical data sequences containing suspected abnormal data as the second sequence to be analyzed. In the subsequent steps, targeted analysis is performed on the first sequence to be analyzed and the second sequence to be analyzed.

[0028] Preferably, in the embodiments of the present invention, the anomaly detection algorithm is the isolation forest algorithm, and the normal probability of each medical data can be obtained by using the distribution position and depth of each medical data in the isolation tree, and the suspected abnormal data is screened out according to the normal probability. It should be noted that the isolation forest algorithm can be built into the medical information system, and the normal probability of each medical data can be directly obtained through a pre-trained isolation forest model.

[0029] Furthermore, the method for obtaining the normal probability includes: For each medical data, obtain the path length between the medical data and the root node in the isolation tree. The longer the path length, the deeper the depth corresponding to the medical data in the isolation tree, and the more the medical data belongs to normal data. Further, in order to characterize the depth feature of the medical data in the isolation tree, obtain the average path length of other medical data within the preset numerical range of the medical data, and obtain the difference between the medical data and the average path length. The larger the difference, the deeper the depth of the medical data relative to other medical data, and the greater the normal probability.

[0030] For this medical data, the ratio of the path length to the depth of the isolated tree is used as the position weight, that is, the position weight can be regarded as normalizing the path length.

[0031] According to the difference between the path length and the average path length, and the position weight, a normal probability is obtained. The greater the normal probability, the more the medical data belongs to normal medical data. In an embodiment of the present invention, after normalizing the product of the difference and the position weight, the normal probability is obtained. The normalization algorithm is a well-known technical means for those skilled in the art and will not be elaborated here.

[0032] In the embodiment of the present invention, the medical data with a normal probability less than the preset normal probability threshold is used as the suspected abnormal data. In the embodiment of the present invention, the normal probability threshold is set to 0.3.

[0033] Step S3: If the suspected abnormal data in the first sequence to be analyzed is complex data, directly assign an abnormal label; if it is basic data, determine the abnormal index according to the patient's medication situation, and determine whether to assign an abnormal label according to the abnormal index.

[0034] For the first sequence to be analyzed, it only contains one suspected abnormal data. First, it is necessary to analyze the data type of the suspected abnormal data. If it is complex data, because complex data has strong relevance, it is impossible to have only one suspected abnormal data under normal circumstances. Therefore, this suspected abnormal data must be data abnormality caused by the medical information system and needs to be assigned an abnormal label.

[0035] Furthermore, if the suspected abnormal data is basic data, because basic data is data that intuitively represents the patient's physiological state, it is necessary to consider the patient's corresponding state at the detection moment. If the patient is in a medication state, it can be considered that this fluctuation is caused by the drug effect and is not a real abnormal data. Therefore, for basic data, the abnormal index can be determined according to the patient's medication situation, and whether to assign an abnormal label is determined according to the abnormal index.

[0036] Preferably, in the embodiment of the present invention, if the patient is in a medication state at the detection moment corresponding to the suspected abnormal data, the abnormal index is set to the first preset value; otherwise, it is set to the second preset value; where the first preset value is less than the second preset value. In the embodiment of the present invention, the first preset value is set to 0, and the second preset value is set to 1. If the abnormal index is the second preset value, the corresponding suspected abnormal data is assigned an abnormal label.

[0037] Step S4: Arrange the second sequence to be analyzed and the medical data sequence within the time range to obtain a matrix to be analyzed; for any two types of suspected abnormal data in the matrix to be analyzed, use the maximum likelihood estimation algorithm to determine the degree of influence between the two types of suspected abnormal data, and judge the influence situation of the two types of suspected abnormal data; assign abnormal labels to the suspected abnormal data that do not influence each other in the second sequence to be analyzed.

[0038] In the embodiment of the present invention, the second sequence to be analyzed is further analyzed. Since the second sequence to be analyzed contains suspected abnormal data of multiple types of medical data, it is necessary to analyze the degree of influence between the suspected abnormal data to judge whether the suspected abnormal data is the normal data fluctuation caused by the patient's physiological activities. For any two types of suspected abnormal data, the maximum likelihood estimation algorithm can be used to determine the degree of influence between the two types of suspected abnormal data. The greater the degree of influence, the more the two types of suspected abnormal data belong to the data that influence each other, indicating that the abnormal fluctuation at this time belongs to the multiple physiological data fluctuations caused by normal physiological fluctuations. Therefore, the influence situation of the two types of suspected abnormal data can be judged, and then abnormal labels are assigned to the suspected abnormal data that do not influence each other in the second sequence to be analyzed.

[0039] Preferably, in the embodiment of the present invention, the method for obtaining the degree of influence includes: Use the maximum likelihood estimation algorithm to count the frequencies of the two types of suspected abnormal data in each change situation. The change situations include four types, namely the first type of suspected abnormal data increases and the second type of suspected abnormal data decreases; the first type of suspected abnormal data decreases and the second type of suspected abnormal data increases; the first type of suspected abnormal data increases and the second type of suspected abnormal data increases; the first type of suspected abnormal data decreases and the second type of suspected abnormal data decreases. By counting the data in the matrix to be analyzed, the frequencies in each change situation can be counted. The specific method is the conventional algorithm of the maximum likelihood estimation algorithm and will not be elaborated here.

[0040] Use the goodness-of-fit formula to obtain the goodness-of-fit between the two types of medical data. The goodness-of-fit formula includes: ; where is the goodness-of-fit, is the frequency in the i-th change situation, is the frequency corresponding to this change situation when the two types of suspected abnormal data have no relationship. Among them can be directly obtained according to the formula in the maximum likelihood estimation algorithm. Taking the situation where the first type of suspected abnormal data increases and the second type of suspected abnormal data increases as an example, ; where is the frequency that the first type of suspected abnormal data increases and the second type of suspected abnormal data increases; is the frequency that the first type of suspected abnormal data decreases and the second type of suspected abnormal data increases; It is the probability of the increase of the first type of suspected abnormal data, which can be obtained according to the proportion of the increased data. This formula is a conventional formula in the maximum likelihood estimation algorithm, and its principle will not be elaborated here.

[0041] Obtain the maximum inference probability less than the goodness-of-fit in the chi-square test table data, and perform a negative correlation mapping to obtain the influence degree. The chi-square test table is shown in Table 1: Table 1 Because the goodness-of-fit is mainly obtained by assuming the frequency when there is no relationship between two types of suspected abnormal data, after finding the maximum inference probability less than the goodness-of-fit in Table 1, a negative correlation mapping needs to be performed to obtain the influence degree between the two types of suspected abnormal data. The negative correlation mapping can directly use the positive integer 1 minus the maximum inference probability. If the influence degree is greater than the preset influence degree threshold, it is determined that the two types of suspected medical data affect each other. In the embodiments of the present invention, the preset influence degree threshold is set to 0.6.

[0042] After obtaining all the abnormal labels in the embodiments of the present invention, the abnormal data containing abnormal labels can be corrected by re-collecting and replacing, etc. It is also possible to further analyze the formation reasons of the abnormal data and optimize it to avoid making the same mistake again.

[0043] In summary, after using the existing anomaly detection algorithm to screen out the suspected abnormal data, the present invention further confirms and analyzes the suspected abnormal data. For the case where only a single suspected abnormal data appears at a detection moment, according to its data type and combined with the medication status, it can be determined whether the abnormality is caused by normal physiological behavior, and then an abnormal label is assigned. For the case where multiple suspected abnormal data appear at a detection moment, a matrix to be analyzed is constructed, and the maximum likelihood estimation algorithm is used to analyze the influence degree between different types of suspected abnormal data. According to the influence situation, the truly abnormal data is screened out and labeled. Through the secondary detection of the suspected abnormal data, the present invention can effectively prevent the misdetection of data and improve the data processing effect of the medical information system.

[0044] Based on the same inventive concept, the present invention also proposes a detection system for abnormal patient data for a medical information system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of any one of the methods for detecting abnormal patient data for a medical information system.

[0045] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0046] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

Claims

1. A method for detecting abnormal patient data for a medical information system, characterized in that: The method comprises: Obtain a medical data sequence of the target patient at each detection moment, wherein each element in the medical data sequence represents a type of medical data, and the medical data includes basic data and complex data; obtain a data time series sequence of each type of medical data of the target patient in time series; Using an anomaly detection algorithm to obtain suspected abnormal data in the data time series; using a medical data sequence containing only one suspected abnormal data as a first sequence to be analyzed, and using other medical data sequences containing suspected abnormal data as a second sequence to be analyzed; If the suspected abnormal data in the first sequence to be analyzed is complex data, an abnormal label is directly assigned; if it is basic data, an abnormal index is determined according to the patient's medication situation, and whether to assign an abnormal label is determined according to the abnormal index; The second sequence to be analyzed is arranged with the medical data sequence within the time series range to obtain the matrix to be analyzed; for any two suspected abnormal data in the matrix to be analyzed, the maximum likelihood estimation algorithm is used to determine the degree of influence between the two suspected abnormal data, and the influence of the two suspected abnormal data is judged; the suspected abnormal data that do not affect each other in the second sequence to be analyzed are assigned abnormal labels.

2. A method for detecting abnormal patient data for a medical information system according to claim 1, characterized in that: The anomaly detection algorithm is an isolation forest algorithm, which uses the distribution position and depth of each medical data in the isolation tree to obtain the normal probability of each medical data, and screens out the suspected abnormal data based on the normal probability.

3. A method for detecting abnormal patient data for a medical information system according to claim 2, characterized in that: The method for obtaining the normal probability includes: For each medical data, obtain the path length of the medical data from the root node in the isolated tree; obtain the average path length of other medical data within a preset numerical range of the medical data; and use the ratio of the path length to the depth of the isolated tree as the position weight; The normal probability is obtained according to the difference between the path length and the average path length, and the position weight.

4. A method for detecting abnormal patient data for a medical information system according to claim 1, characterized in that: The method for obtaining the abnormality index includes: If the patient is taking medicine at the detection time corresponding to the suspected abnormal data, the abnormality index is set to a first preset value; otherwise, it is set to a second preset value; wherein the first preset value is less than the second preset value.

5. A method for detecting abnormal patient data for a medical information system according to claim 4, characterized in that: The step of determining whether to assign an abnormal label according to the abnormal index includes: If the abnormality index is the second preset value, an abnormality label is assigned to the corresponding suspected abnormal data.

6. A method for detecting abnormal patient data for a medical information system according to claim 1, characterized in that: The method for obtaining the degree of influence includes: The maximum likelihood estimation algorithm is used to count the frequencies of the two suspected abnormal data in each change situation, and the fit between the two medical data is obtained using the fit formula; the maximum probability of inference establishment that is less than the fit is obtained in the chi-square test table data, and negative correlation mapping is performed to obtain the degree of influence.

7. A method for detecting abnormal patient data for a medical information system according to claim 6, characterized in that: Determine the impact of two types of suspected medical data, including: If the influence degree is greater than a preset influence degree threshold, it is determined that the two suspected medical data influence each other.

8. A method for detecting abnormal patient data for a medical information system according to claim 2, characterized in that: The screening out of the suspected abnormal data according to the normal probability includes: The medical data whose normal probability is less than a preset normal probability threshold is regarded as suspected abnormal data.

9. A method for detecting abnormal patient data for a medical information system according to claim 1, characterized in that: Each column of the matrix to be analyzed represents a type of medical data, and each row represents a medical data sequence corresponding to a detection moment.

10. A patient abnormal data detection system for a medical information system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a method for detecting abnormal patient data for a medical information system as described in any one of claims 1 to 9 are implemented.