A data processing apparatus and device
By receiving and processing vital signs data, laboratory data, and diagnostic and treatment data, extracting clinical feature matrices and inputting them into a prediction model, the problem of delayed risk warning in the medical field has been solved, enabling timely and accurate risk warning and intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANFENG CHAOSHENG MEDICAL TECHNOLOGY (LIAONING) CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-29
AI Technical Summary
Current technologies in the medical field suffer from delayed risk warnings, often only providing alerts after the risk has already occurred, thus missing the golden window for intervention.
By receiving vital sign data, laboratory data, and diagnostic data sent by edge medical devices, a clinical feature matrix is extracted and input into a pre-trained prediction model to obtain analysis results, thereby improving the accuracy and timeliness of risk warnings.
It enables timely and accurate risk warnings, allowing users to intervene promptly based on the analysis results and prevent risks from occurring.
Smart Images

Figure CN122117372A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a data processing device and apparatus. Background Technology
[0002] In the medical field, risk warnings are often based on single-dimensional data. For example, for laboratory data, simple threshold judgments are used to issue risk alerts. Relevant staff, upon receiving such a risk alert, can then focus on related information. However, this method of risk warning relying solely on laboratory data is delayed, often only providing warnings after the risk has already occurred, missing the crucial 0-24 hour window for intervention.
[0003] Therefore, how to issue timely risk warnings has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a data processing apparatus and device to solve the problem of delayed risk warning in the prior art.
[0005] This application provides a data processing apparatus, the apparatus comprising: The receiving module is used to receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; The data processing module is used to extract a clinical feature matrix from the medical data, the clinical feature matrix being used to describe whether the medical data contains risk factors; input the clinical feature matrix into a pre-trained prediction model to obtain analysis results, and output the analysis results so that users can perform analysis based on the analysis results.
[0006] This application also provides a data processing method, the method comprising: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; A clinical feature matrix is extracted from the medical data, and the clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, which are then output so that users can perform analysis based on these results.
[0007] This application also provides an electronic device, which includes a processor. When the processor executes a computer program stored in a memory, it performs the following functions: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; A clinical feature matrix is extracted from the medical data, and the clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, which are then output so that users can perform analysis based on these results.
[0008] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the data processing methods described above.
[0009] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the steps of any of the data processing methods described above.
[0010] In this embodiment, the receiving module of the data processing device can receive medical data of the target object sent by the edge medical device. Since the medical data includes vital sign data, laboratory data, and diagnostic data, the data processing module can extract a rich and multi-faceted clinical feature matrix describing the presence of risk factors from the medical data. By inputting this rich clinical feature matrix into a pre-trained prediction model, accurate analysis results can be obtained and output more promptly, improving the accuracy and timeliness of risk warnings. Users can then analyze the results and intervene in a timely manner to avoid risks. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A schematic diagram of a data processing apparatus provided in an embodiment of this application; Figure 2 A flowchart for determining a clinical feature matrix is provided as an embodiment of this application; Figure 3 This is a schematic diagram of multimodal data fusion provided in an embodiment of this application; Figure 4 A schematic diagram of a system deployment architecture provided in this application embodiment; Figure 5 A schematic diagram of a prediction model structure provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the display of analysis results provided in an embodiment of this application; Figure 7 This is a schematic diagram of a data processing method provided in an embodiment of this application; Figure 8 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art are within the scope of protection of this application.
[0014] This application provides a data processing apparatus and device. The data processing apparatus includes a receiving module and a data processing module. The receiving module is used to receive medical data of a target object sent by an edge medical device. The medical data includes vital sign data, laboratory data, and diagnostic data. The data processing module is used to extract a clinical feature matrix based on the medical data. The clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, and the analysis results are output so that users can perform analysis based on the analysis results.
[0015] Figure 1 This is a schematic diagram of a data processing device provided in an embodiment of this application, such as... Figure 1 As shown, the device includes: The receiving module 101 is used to receive medical data of the target object sent by the edge medical device, the medical data including vital sign data, laboratory data and diagnosis and treatment data; The data processing module 102 is used to extract a clinical feature matrix from the medical data, the clinical feature matrix being used to describe whether the medical data has risk factors; input the clinical feature matrix into a pre-trained prediction model to obtain analysis results, and output the analysis results so that the user can perform analysis based on the analysis results.
[0016] To enable timely and accurate risk warnings, in this embodiment, the receiving module 101 can receive medical data of the target object in real time. This medical data may include the target object's vital signs, laboratory data, and diagnostic data. The medical data may be retrieved from a database by the edge medical device and sent when data processing is required, or it may be obtained by the edge medical device through an input device and then sent, such as when medical personnel input medical data into the edge medical device via an input device.
[0017] In this embodiment, the edge medical device can retrieve medical data of a target object from a set storage location at set time intervals and send it to the receiving module 101 of the data processing device. Alternatively, the edge medical device can send the target object's medical data to the receiving module 101 after receiving a data processing instruction.
[0018] Vital signs data may include parameters such as heart rate, systolic / diastolic blood pressure, blood oxygen saturation, hourly urine output, and body temperature. For example, this vital signs data can be collected using a Philips IntelliVue MX800 monitor or a smart urine output monitor. Of course, those skilled in the art can also use other peripheral medical devices to obtain the relevant parameters.
[0019] Laboratory data may include parameters such as creatinine, blood urea nitrogen, serum potassium, serum sodium, estimated glomerular filtration rate (eGFR), and quantified urine protein. This laboratory data can be obtained from a Laboratory Information Management System (LIS).
[0020] Clinical data may include parameters such as: records of clinical procedures (type / dosage / time of contrast agent use, duration / volume of dehydration therapy), data on nephrotoxic drugs (type of drug, dosage, duration of administration), and history of underlying diseases (stage of chronic kidney disease, history of diabetes, etc.). This clinical data can be obtained from electronic medical record (EMR) systems and hospital information systems (HIS).
[0021] After acquiring the medical data, the data processing module 102 can extract features from the medical data to obtain a clinical feature matrix. This clinical feature matrix can be used to describe whether there are risk factors in the medical data. These risk factors can be understood as the causes that could lead to the occurrence of the risk currently being analyzed. For example, the risk currently being analyzed could be the probability of occurrence of a certain disease, such as a prediction of acute kidney injury (AKI); or it could be a prediction of a certain medical value.
[0022] In this embodiment of the application, when extracting the clinical feature matrix from medical data, the received medical data can be used to extract features based on a pre-trained feature extraction model to obtain the clinical feature matrix.
[0023] After obtaining the clinical feature matrix, it can be input into a pre-trained prediction model to obtain the analysis results. This prediction model can be based on a bidirectional LSTM. If the risk being analyzed is a prediction of a specific disease, then the analysis result represents the probability of that disease occurring within a given future timeframe.
[0024] It should be noted that the analysis results obtained in this application embodiment are not directly applied to humans, but are intended to provide reference for relevant personnel to prevent them from overlooking the existence of potential risks. In this application embodiment, after obtaining the analysis results output by the prediction model, the analysis results can be output so that users can view the analysis results and perform corresponding analysis based on them, thereby making timely response strategies to avoid the occurrence of risks.
[0025] In this embodiment, the receiving module of the data processing device can receive medical data of the target object sent by the edge medical device. Since the medical data includes vital sign data, laboratory data, and diagnostic data, the data processing module can extract a rich and multi-faceted clinical feature matrix describing the presence of risk factors from the medical data. By inputting this rich clinical feature matrix into a pre-trained prediction model, the analysis results can be obtained and output more promptly, improving the accuracy and timeliness of risk warning. Users can then analyze the results and intervene in a timely manner to avoid risks.
[0026] To further improve the accuracy of data processing, based on the above embodiments, in this embodiment, the data processing module 102 is specifically used to divide the medical data into multiple sub-data sets according to the acquisition time, set window duration, and set window sliding step size of each parameter in the medical data; for each risk factor, it determines whether the corresponding sub-data set involves the risk factor according to the parameters in each sub-data set and the judgment criteria of the risk factor; if it exists, it fills the first value into the preset position in the original clinical feature matrix; if it does not exist, it fills the second value into the preset position; the completed original clinical feature matrix is determined as the clinical feature matrix.
[0027] In this embodiment, when extracting a clinical feature matrix from medical data, the data processing module 102 can divide the parameters in the medical data into multiple sub-data sets based on the acquisition time, set window duration, and set window sliding step size of each parameter. The set window duration can be a pre-configured fixed value, such as 30 minutes or 1 hour, and the set window sliding step size can be 15 minutes or 10 minutes.
[0028] In one possible implementation, the duration of the set window can also be determined based on the target object's condition. In this embodiment, the target Acute Physiology and Chronic Health Evaluation (APACHE II) score can be determined based on the medical data. The process for determining APACHE II is prior art, and will not be described further in this embodiment.
[0029] After determining the target APACHE II score, the target sliding window strategy corresponding to the target acute physiological and chronic health scores can be determined based on the pre-configured correspondence between different scores and sliding window strategies. Each sliding window strategy clearly describes the time window length and the window sliding step size. After determining the target sliding window strategy, the time window length in the target sliding window strategy can be defined as the set window duration, and the window sliding step size in the target sliding window strategy can be defined as the set window sliding step size.
[0030] Specifically, when the target acute physiological and chronic health score is ≥25, the window duration can be set to 30 minutes and the window sliding step can be set to 15 minutes; when the target acute physiological and chronic health score is ≤24 and the target acute physiological and chronic health score is 1 hour and the window sliding step can be set to 30 minutes; when the target acute physiological and chronic health score is <15, the window duration can be set to 2 hours and the window sliding step can be set to 1 hour.
[0031] Since each risk has a corresponding trigger, in this embodiment of the application, the trigger corresponding to the risk can be configured according to the risk that needs to be analyzed. The trigger can be one or multiple.
[0032] After obtaining multiple sub-data sets, the data in each sub-data set is analyzed to determine whether the corresponding sub-data set involves a risk factor. In this embodiment, for each risk factor, it can be determined whether the corresponding sub-data set involves the risk factor based on the parameters in each sub-data set and the judgment criteria for the risk factor. If it exists, a first value can be filled into a preset position in the original clinical feature matrix; if it does not exist, a second value can be filled into a preset position. The first and second values are different; for example, the first value is 1 and the second value is 0. Of course, those skilled in the art can configure the first and second values as needed.
[0033] In this embodiment, the fully filled original clinical feature matrix can be determined as the clinical feature matrix. This clinical feature matrix can be understood as the time-series encoding of the received medical data.
[0034] Specifically, one-hot encoding can be used. When any sub-data set involves a certain risk factor, the corresponding position in the original clinical feature matrix can be marked as 1; otherwise, it can be marked as 0. In this embodiment, the original clinical feature matrix can be a matrix of dimension "number of sub-data sets × number of risk factors". Here, the number of sub-data sets is the number of time windows. That is, when the first sub-data set involves the first risk factor, the intersection of the first row and the first column in the original clinical feature matrix can be marked as 1; when the nth sub-data set involves the mth risk factor, the intersection of the nth row and the mth column in the original clinical feature matrix can be marked as 1. In other words, the complete original clinical feature matrix is composed of 0s and 1s.
[0035] To further improve the accuracy of data processing, based on the above embodiments, in this embodiment of the application, the apparatus further includes: The determination module 103 is used to determine the risk feature matrix based on the clinical feature matrix and the weight corresponding to each risk factor; The fusion module 104 is used to fuse the clinical feature matrix and the risk feature matrix to obtain a first fusion matrix, and to update the clinical feature matrix using the first fusion matrix.
[0036] Since different risk factors have different impacts on the risk being analyzed, in this embodiment, a corresponding weight can be pre-configured for each risk factor, and the risk feature matrix can be determined based on the clinical feature matrix and the weight corresponding to each risk factor.
[0037] In constructing the clinical feature matrix, rows represent sub-data sets, and columns represent each risk factor. Therefore, in this embodiment, for each row of the clinical feature matrix, the product of each value in that row and its corresponding weight can be calculated, and the sum of all products corresponding to that row can be used to determine the risk value for that row. Since each row corresponds to each sub-data set, in this embodiment, the risk value corresponding to any row can be understood as the risk value of the sub-data set corresponding to that row, or it can be called the single-time-window risk value. After determining the risk value corresponding to each row, the resulting matrix is the risk feature matrix, which is a matrix of "number of sub-data sets × 1". Each item in the risk feature matrix corresponds to the risk value of a sliding window, that is, the risk scores of each "snapshot" after the time series is divided by the sliding window are arranged in chronological order. When determining the risk value corresponding to any row, it can be determined based on the following formula: R t =∑(x t i×W i ), where R t x represents the risk value corresponding to the t-th row in the clinical feature matrix; t The t-th row in the clinical feature matrix represents the i-th risk factor; W represents the t-th risk factor. i This represents the weight corresponding to the i-th risk factor. This risk feature matrix can be understood as a risk level code corresponding to medical data.
[0038] After obtaining the risk feature matrix, the clinical feature matrix and the risk feature matrix can be fused to obtain a first fused matrix, which is then used to update the clinical feature matrix. For example, during fusion, the obtained risk feature matrix can be appended to the last column of the clinical feature matrix to obtain the first fused matrix.
[0039] In one possible implementation, the determining module 103 can determine the weight corresponding to each risk factor by combining the expert's evaluation of each risk factor. In this embodiment, a preset number of judgment matrices can be obtained. The preset number can be any value such as 3, 4, or 8, and can be configured as needed by those skilled in the art. This judgment matrix is used to describe the risk importance ratio between every two risk factors. Any judgment matrix is determined based on the risk importance ratio judged by a certain expert.
[0040] After obtaining the judgment matrix, the preset number of judgment matrices can be analyzed using the analytic hierarchy process (AHP) to determine the weight of each risk factor.
[0041] Specifically, eight experts can compare the risk importance of any two risk factors pairwise, filling in a judgment matrix to obtain eight judgment matrices. These eight experts can consist of five chief physicians of nephrology and three chief physicians of ICU. The rows and columns of this judgment matrix represent risk factors. The dimension of the judgment matrix is: number of risk factors × number of risk factors. In this embodiment, rows can be understood as "compared risk factors," and columns as "compared risk factors." The elements in this judgment matrix are the risk importance ratios between corresponding risk factors.
[0042] After obtaining the judgment matrix, the Analytic Hierarchy Process (AHP) can be used to calculate the weight corresponding to each risk factor. For ease of understanding, the processing logic of AHP is briefly explained below: For the judgment matrices of the eight experts, a method of "independent calculation followed by comprehensive weighting" is used. For each judgment matrix, its eigenvector is first calculated and normalized to obtain the weight allocation of that expert to the target node; then, the weight results of the eight experts are weighted and averaged to finally obtain the comprehensive weight of each risk factor.
[0043] The standard output of the AHP algorithm can calculate the consensus ratio (CR). In this embodiment, the consensus ratio can be used to determine whether the opinions of various experts are consistent. In the AHP algorithm, for each judgment matrix, after calculating the consensus index (CI), and combining it with the average random consensus index (RI), CR = CI / RI can be obtained. When CR < 0.1, the logical consistency of the judgment matrix is considered acceptable. Since most mainstream AHP toolkits (such as Python's ahp library and MATLAB's AHP toolbox) can automatically calculate weights and CR values, this embodiment will not elaborate on the process of determining the weights corresponding to risk factors based on the analytic hierarchy process.
[0044] To further improve the accuracy of data processing, based on the above embodiments, in this embodiment, the data processing module 102 is further configured to obtain a pre-configured adjacency matrix and a risk trigger feature matrix. The adjacency matrix is used to describe the impact of the simultaneous existence of any two risk triggers on the occurrence of risk, and the risk trigger feature matrix is used to describe the information of each risk trigger about each preset attribute. The adjacency matrix and the risk trigger feature matrix are input into a pre-trained graph convolutional network to obtain a path association feature matrix of risk triggers. The fusion module 104 is specifically used to fuse the clinical feature matrix, the risk feature matrix, and the path association feature matrix to obtain the first fusion matrix.
[0045] In order to describe the impact of different risk factors on the occurrence of risk events, the path association feature matrix between risk factors can also be determined in the embodiments of this application.
[0046] In this embodiment of the application, a pre-configured adjacency matrix and risk factor feature matrix can be obtained.
[0047] The adjacency matrix can have the dimension of "risk trigger number × risk trigger number". This adjacency matrix describes the impact of the simultaneous presence of any two risk triggers on the occurrence of a risk. In this embodiment, historical data can be used to analyze the impact of the simultaneous occurrence of any two risk triggers on the occurrence of a risk. For example, by analyzing historical data, it can be determined whether the probability of a risk event occurring is high, moderate, or low when risk trigger A and risk trigger B occur simultaneously. In this embodiment, the occurrence of risk events can be classified into three levels: Level I, Level II, and Level III. Level I can represent a high probability of a risk event, Level II can represent a moderate probability of a risk event, and Level III can represent a low probability of a risk event. In this embodiment, the evidence support value can be configured as 0.9 for Level I, 0.7 for Level II, and 0.5 for Level III. If the risk level corresponding to the simultaneous occurrence of risk factor A and risk factor B is Level I, then the value at the intersection of the row and column represented by risk factors A and risk factor B in the adjacency matrix can be determined to be 0.9.
[0048] The risk trigger feature matrix describes the information of each risk trigger regarding each preset attribute. Assuming each risk trigger corresponds to 16 attributes, such as risk type and level of clinical evidence, then the dimension of the risk trigger feature matrix is "number of risk triggers × 16".
[0049] After obtaining the risk trigger feature matrix and the adjacency feature matrix, these matrices can be input into a pre-trained Graph Convolutional Network (GCN) to obtain the path association feature matrix of the risk triggers. In other words, this trained GCN is used to capture the clinical associations between risk triggers. In this embodiment, this path association feature matrix can be referred to as path association encoding.
[0050] Specifically, when using graph convolutional networks to determine the path association feature matrix, association features can be extracted through two layers of GCN (e.g., the input layer outputs a 33-dimensional feature matrix → the hidden layer outputs a 64-dimensional feature matrix → the output layer outputs a 32-dimensional path association feature matrix) to generate a 32-dimensional path association feature matrix.
[0051] After obtaining the path association feature matrix, the fusion module 104 can fuse the clinical feature matrix, risk feature matrix and risk feature matrix to obtain the first fusion matrix when performing fusion to determine the first fusion matrix.
[0052] In one possible implementation, after determining the clinical feature matrix, risk feature matrix, and risk feature matrix, coding standardization can be performed. In the embodiments of this application, Z-Score standardization can be based on the formula. The clinical feature matrix, risk feature matrix, and risk feature matrix are standardized. Here, X represents any one of the three matrices. This represents the mean value corresponding to X; This represents the standard deviation corresponding to X; and All of these were determined based on 8,000 sets of historical data.
[0053] In one possible implementation, after standardizing each matrix, a "numerical truncation" process can be performed. This operation aims to eliminate the interference of extreme outliers on model analysis and obtain a stable feature distribution. For example, values exceeding the range [-3, 3] can be truncated. That is, the elements in the standardized matrix are restricted to the interval [-3, 3]. If an element is less than -3, it is forcibly set to -3; if an element is greater than 3, it is forcibly set to 3; elements within this interval remain unchanged.
[0054] To further improve the accuracy of data fusion, based on the above embodiments, in this embodiment, the fusion module 104 is specifically used to process the clinical feature matrix, the risk feature matrix, and the path association feature matrix using a multi-head attention mechanism to determine the weights corresponding to the clinical feature matrix, the risk feature matrix, and the path association feature matrix respectively; and to determine the first fusion matrix based on the clinical feature matrix, the risk feature matrix, the path association feature matrix, and each weight.
[0055] Since feature matrices with different meanings have different impacts on the risk being analyzed, in this embodiment, the weights corresponding to different features can be determined, and fusion processing can be performed based on these weights.
[0056] In this embodiment, a multi-head attention mechanism can be used to process the clinical feature matrix, risk feature matrix, and path association feature matrix to determine the weights corresponding to the clinical feature matrix, risk feature matrix, and path association feature matrix, respectively.
[0057] Specifically, the number of heads in the multi-head attention mechanism can be set to 4, and the clinical feature matrix, risk feature matrix, and risk feature matrix can be processed by the multi-head attention mechanism to obtain a weight of 0.45 for the clinical feature matrix, a weight of 0.35 for the risk feature matrix, and a weight of 0.2 for the path association feature matrix.
[0058] In one possible implementation, temporal alignment can be performed before multimodal data fusion based on a multi-head attention mechanism.
[0059] Before performing time alignment, the Dynamic Time Warping (DTW) algorithm can be used to construct a distance matrix between the real-time data and the test data. The optimal alignment path is then found using the formula (D(i,j)=d(i,j)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))), with an alignment error ≤ 5 min. Here, i represents the time index of the real-time data; j represents the time index of the test data; d(i,j) represents the original distance between the i-th point of the real-time data and the j-th point of the test data; and D(i,j) represents the minimum cumulative distance from (1,1) to (i,j). After obtaining this optimal alignment path, the feature encodings of each dimension can be aligned based on this optimal alignment path. The specific steps are as follows: Path mapping: The optimal alignment path obtained by the DTW algorithm will establish a one-to-one correspondence between the real-time data time point (i) and the test data time point (j).
[0060] Dimensional alignment: For any feature matrix among the clinical feature matrix, risk feature matrix, and path association feature matrix, the i-th time point of real-time data can be matched with the j-th time point of test data according to the mapping relationship of the optimal alignment path.
[0061] Error control: During the alignment process, the alignment error must be ≤5 min to ensure the comparability of time series in a clinical sense.
[0062] In one possible implementation, after acquiring medical data, but before extracting the clinical feature matrix based on the medical data, outlier handling and missing value handling can be performed.
[0063] For outlier handling: Outliers can be identified based on clinical reference ranges (e.g., serum creatinine > 133 μmol / L, urine output < 10 ml / h or > 500 ml / h) combined with the 3σ criterion. Outliers are replaced by the median of the nearest neighbor window. The 3σ criterion can be understood as follows: if a parameter follows a normal distribution, 99.73% of the data should fall within the mean ± 3σ; values outside this range are considered outliers. Combining this with clinical reference ranges: First, use the clinical reference range (e.g., serum creatinine > 133 μmol / L) for initial screening of outliers, then use the 3σ criterion (based on historical data) for secondary verification. This approach aligns with clinical common sense and data distribution patterns. The specific steps are as follows: Preliminary screening: For each parameter in the medical data (such as serum creatinine and urine output), abnormal values are initially screened out by using clinical reference ranges (such as serum creatinine > 133 μmol / L, urine output < 10 ml / h or > 500 ml / h).
[0064] Secondary verification: The outliers initially selected are then verified again using the 3σ criterion to confirm whether they are true outliers.
[0065] Encoding update: After an outlier is confirmed, it is replaced by the middle value in the nearest window.
[0066] For handling missing values: tiered imputation can be performed, resulting in data integrity ≥98%. When the parameter missing duration is ≤1 hour, the missing value can be determined based on K-Nearest Neighbors Interpolation (KNN) with path similarity, where K=5; when 1 hour < parameter missing duration ≤3 hours, LSTM interpolation can be used to determine the missing value; when the parameter missing duration >3 hours, the parameter can be marked as invalid and removed.
[0067] LSTM interpolation can be understood as using an LSTM model to predict and complete time series data. The specific implementation is as follows: Model training: Train the LSTM model using complete historical time series data to learn the pattern of parameter changes over time.
[0068] Missing parameter prediction: When the duration of missing parameters is between 1 hour and 3 hours, valid historical data before the missing point can be input into the trained LSTM model to predict the value of the missing point.
[0069] Completion and Validation: After completion, ensure that the overall data integrity is ≥98% to meet the requirements of subsequent analysis.
[0070] The process of determining the clinical feature matrix will be explained below with reference to a specific embodiment. Figure 2 A flowchart for determining a clinical feature matrix is provided as an embodiment of this application, such as... Figure 2As shown, the process first involves a three-tiered risk trigger matching, which matches the acquired medical data with 33 pre-defined risk triggers to identify the clinical pathway elements currently involved in the patient's case. Specifically, these 33 risk triggers are divided into three layers: basic risk, treatment procedure layer, and real-time monitoring layer. These include the following risk triggers: Basic risk tier (10 categories): chronic kidney disease stage 3-5, diabetic nephropathy, hypertensive nephropathy, congestive heart failure, cirrhosis, sepsis, advanced age (≥75 years), malignant tumors, chronic liver disease, and history of immunosuppressant use; Diagnostic and therapeutic procedures (15 categories): use of hyperosmolar contrast agents (osmolarity ≥600mOsm / kgH2O), use of hypoosmolar contrast agents (280-320mOsm / kgH2O), use of nonsteroidal anti-inflammatory drugs (≥100mg / d), use of aminoglycoside antibiotics (treatment course ≥7d), dehydration therapy (24h negative balance ≥1000ml), etc. Real-time monitoring layer (8 categories): Serum creatinine increase ≥26.5 μmol / L within 48 hours, urine output <0.5 ml / (kg) h)(lasting for 6 hours), serum potassium ≥5.5mmol / L, etc.
[0071] After completing the risk trigger matching, multi-dimensional coding processing can be carried out: time series coding, risk level coding, and path association coding.
[0072] When determining the time series encoding, a sliding window algorithm can be used, dynamically adjusting the sliding window strategy based on the APACHE II score. Then, One-Hot encoding is used to generate a 33-dimensional clinical feature matrix (time window number × 33).
[0073] When determining the risk level code, the analytic hierarchy process (AHP) can be used to calculate the risk factor weights, construct a judgment matrix containing the opinions of 8 experts, and perform a consistency test (CR < 0.1). The single-window risk value is calculated using the formula Rt = ∑(xt, i × Wi), ultimately forming the risk feature matrix.
[0074] When determining path association encoding, a graph convolutional network can be used to determine a 32-dimensional path association feature matrix.
[0075] After obtaining the clinical feature matrix, risk feature matrix, and path association feature matrix, feature splicing and standardization can be performed. Specifically, the Z-Score standardization formula Z=(X -μ) / σ can be used to truncate values outside the range [-3,3], generating 128-dimensional clinical path embedding features, which are then input into the multimodal data fusion module.
[0076] A structured risk factor system consisting of a "basic risk layer, diagnosis and treatment operation layer, and real-time monitoring layer" is constructed. Through multi-dimensional processing such as time series coding, risk level coding, and path association coding, clinical pathway knowledge is transformed into embedded features of a unified dimension.
[0077] To further improve the accuracy of data processing, based on the above embodiments, in this embodiment, the data processing module 102 is further configured to: determine the occurrence label corresponding to each parameter in the medical data according to the target event occurrence judgment criteria corresponding to the parameter; determine a first marginal probability, a second marginal probability, and a joint probability based on the diagnostic data of historical objects and the parameter; wherein, the first marginal probability is the proportion of the parameter value present in the diagnostic data of the historical objects; the second marginal probability is the proportion of the event occurrence label corresponding to the parameter in the diagnostic data of the historical objects being the target event occurrence label; the joint probability is the proportion of the parameter value present in the diagnostic data of the historical objects and the event occurrence label corresponding to the parameter being the target event occurrence label; determine the entropy value corresponding to the parameter based on the first marginal probability, the second marginal probability, the joint probability, and the mutual information entropy algorithm; filter each parameter in the medical data according to the entropy value corresponding to each parameter and a preset threshold to obtain core parameters; fuse the clinical feature matrix and the core parameters to obtain a second fusion matrix, and update the clinical feature matrix using the second fusion matrix.
[0078] In this embodiment of the application, before inputting the clinical feature matrix into the pre-trained prediction model and obtaining the analysis results, each parameter in the medical data can be filtered based on the mutual information entropy algorithm.
[0079] Furthermore, for each parameter in the medical data, an occurrence label corresponding to that parameter can be determined based on the target event occurrence judgment criteria corresponding to that parameter.
[0080] For example, based on the diagnostic criteria of the 2021 Clinical Practice Guidelines for Acute Kidney Injury, if a parameter in the medical data meets the following criteria: serum creatinine increases by ≥26.5 μmol / L within 48 hours or increases by ≥1.5 times the baseline within 7 days, or urine output is <0.5 ml / (kg) If any of the following criteria is met (h) "lasting ≥6h…", then the occurrence label Y=1 for the corresponding parameter can be determined; otherwise, Y=0. For example, if the occurrence label Y=1 for parameter A, then it can be determined that acute kidney injury may have occurred when parameter A is the parameter value recorded in the medical data.
[0081] In this embodiment, a first marginal probability, a second marginal probability, and a joint probability can be determined based on the diagnostic data of historical objects and the parameter. The first marginal probability is the percentage of historical objects' diagnostic data where the parameter value exists; the second marginal probability is the percentage of historical objects' diagnostic data where the event occurrence label corresponding to the parameter is the target event occurrence label; and the joint probability is the percentage of historical objects' diagnostic data where both the parameter value and the event occurrence label corresponding to the parameter are the target event occurrence label.
[0082] The entropy value corresponding to this parameter is determined based on the first marginal probability, the second marginal probability, the joint probability, and the mutual information entropy algorithm. Specifically, it can be determined based on the following formula:
[0083] in, Let X represent the entropy value; Y represent the occurrence label corresponding to the parameter; p(x,y) represents the joint probability that parameter X takes the value x and occurrence label Y takes the value y, i.e., the joint probability; p(x) represents the marginal probability that parameter X takes the value x, i.e., the first marginal probability; p(y) represents the marginal probability that occurrence label Y takes the value y, i.e., the second marginal probability.
[0084] After obtaining the entropy value corresponding to each parameter, the core parameters can be obtained by filtering each parameter in the medical data based on the entropy value and a preset threshold. For example, the preset threshold can be any decimal such as 0.6, 0.5, or 0.8.
[0085] Specifically, after obtaining the entropy value for each parameter, each parameter can be sorted according to its entropy value, and then the 32 parameters with an entropy value ≥ 0.6 are selected as core parameters. These core parameters can be understood as having the strongest correlation with the risk being analyzed, thus ensuring the performance of subsequent model analysis while avoiding feature redundancy. In this embodiment, if the number of parameters with an entropy value ≥ 0.6 exceeds 32, the top 32 parameters can be selected as core parameters. If the number of parameters with an entropy value ≥ 0.6 does not exceed 32, all parameters that meet the threshold requirement are selected as core parameters, without forcing a number, ensuring that the selected parameters have sufficient correlation.
[0086] After determining the core parameters, the core feature matrix and the core parameters can be fused to obtain a second fused feature matrix, and the clinical feature matrix can be updated using the second fused matrix.
[0087] The following is combined with Figure 3 The process of multimodal data fusion is explained. Figure 3 This is a schematic diagram of multimodal data fusion provided in an embodiment of this application, such as... Figure 3 As shown, after obtaining the clinical feature matrix, risk feature matrix, and path association feature matrix processed as described above, a multi-head attention mechanism (number of heads = 4) can be used to assign weights. Specifically, the weight corresponding to the clinical feature matrix is 0.45, the weight corresponding to the risk feature matrix is 0.35, and the weight corresponding to the path association feature matrix is 0.2. Based on the clinical feature matrix, risk feature matrix, path association feature matrix, and each weight, a first fusion matrix is determined. This first fusion matrix is concatenated with 32 core features (total dimension 160), and then mapped through a fully connected layer (activation function LeakyReLU) to obtain a 256-dimensional second fusion matrix.
[0088] A multi-head attention weight allocation strategy is adopted to solve the problems of frequency differences, structural heterogeneity, and uneven reliability of multi-source data.
[0089] To improve the reliability of the analysis results, based on the above embodiments, in this embodiment of the application, the data processing module 102 is further used to process the clinical feature matrix based on the model result interpretation algorithm to obtain the contribution degree corresponding to each risk factor; and select and output the target risk factor according to the contribution degree corresponding to each risk factor.
[0090] In related technologies, the analysis models are mostly "black box" algorithms, which only output analysis results (such as risk level) without clarifying the risk causes (such as the synergistic risk of "contrast agent use + dehydration treatment"), and do not adapt to the differences in the basic risks of different patients (such as history of chronic kidney disease, advanced age). The warning threshold is "one-size-fits-all", and the user adoption rate is less than 50%.
[0091] In this embodiment of the application, the clinical feature matrix can also be processed by the model result interpretation algorithm to obtain the contribution of each risk factor, and the target risk factor can be selected and output based on the contribution of each risk factor.
[0092] Specifically, the feature contribution can be calculated using the SHAP (SHapley Additive exPlanations) value algorithm to output the risk factors of TopK.
[0093] To further improve the adoption rate of user analysis results, based on the above embodiments, this application embodiment can also automatically generate intervention suggestions based on the method of "model calculation + clinical rule mapping": After obtaining the analysis results and target risk factors, these quantitative results can be matched with the content in a pre-set knowledge base. This knowledge base contains standardized intervention protocols corresponding to different risk levels and risk factors. Successfully matched intervention protocols are then output. For example, this intervention protocol could be: medical staff adjust the fluid resuscitation volume to 80 ml / kg, increase the frequency of renal function monitoring to 6 hours / time, and the patient's serum creatinine drops to 148 μmol / L within 72 hours, with no acute kidney injury occurring, thus verifying the effectiveness of the protocol.
[0094] In one possible implementation, a large model can be introduced as an auxiliary tool, but the core logic is still based on the combination of clinical rules and model computation, rather than relying entirely on the generation of a large model.
[0095] The logic of electronic devices acquiring results: Electronic devices obtain patient baseline, risk window, intervention recommendations, and other results through the following steps: Patient baseline acquisition: Statistical features (mean, standard deviation, etc.) are extracted from raw monitoring data (such as heart rate, serum creatinine, urine output), and after preprocessing, a patient baseline feature vector is generated; Risk window determination: Time series data is segmented by sliding windows, a risk score is calculated for each window, and the time window corresponding to high risk is determined by combining the threshold judgment. Intervention suggestion generation: Input the patient's baseline characteristics, risk window and risk score into the preset clinical intervention rule engine, match the corresponding intervention plan and output suggestions.
[0096] The entire process is automated by electronic devices through algorithms, requiring no human intervention.
[0097] To facilitate the deployment of data processing across environments, in this embodiment, the data processing process and dependent environments described in the above embodiments can be packaged and integrated into a Docker container.
[0098] To improve data processing efficiency, in this embodiment, the data processing procedures described in the above embodiments can be imported into TensorRT, and the inference speed can be optimized through compression, layer fusion, precision calibration, etc., to achieve the requirement of latency ≤500ms.
[0099] The following explains the judgment criteria for each risk factor: Basic risk tier (10 categories, patient-inherent AKI susceptibility risk factors) (1) Chronic kidney disease stage 3-5 (based on KDIGO chronic kidney disease staging criteria) (2) Diabetic nephropathy (combined with type 2 diabetes and evidence of kidney damage) (3) Hypertensive nephropathy (long-term history of hypertension with abnormal renal function) (4) Congestive heart failure (New York Heart Classification III-IV) (5) Liver cirrhosis (Child-Pugh classification B or above) (6) Sepsis (meeting the Sepsis 3.0 diagnostic criteria, i.e., infection + sequential organ failure score ≥2) (7) Advanced age (age ≥ 75 years) (8) Malignant tumors (advanced stage or active tumors treated with chemotherapy / radiotherapy) (9) Chronic liver disease (non-cirrhotic chronic hepatitis, autoimmune liver disease, etc.) (10) History of immunosuppressant use (use of glucocorticoids ≥10mg / day or other immunosuppressive drugs within the past 3 months) Clinical Practice Layer (15 categories, high-risk AKI events during dynamic clinical practice) (11) Use of hypertonic contrast agents (contrast agent osmotic pressure ≥600mOsm / kg H2O) (12) Use of hypotonic contrast agents (contrast agent osmotic pressure 280-320 mOsm / kg H2O) (13) Use of nonsteroidal anti-inflammatory drugs (daily dose ≥100mg, continuous use ≥3d) (14) Use of aminoglycoside antibiotics (treatment course ≥7 days, or single dose ≥5 mg / kg) (15) Dehydration therapy (negative fluid balance ≥1000ml in 24 hours, lasting ≥12 hours) (16) Fluid resuscitation (24h fluid resuscitation volume ≥50ml / kg, used for septic shock and other scenarios) (17) Initiation of continuous renal replacement therapy (CRRT) (initiated due to acute kidney injury or renal failure) (18) Mechanical ventilation (PEEP ≥ 10 cmH2O, and lasting ≥ 24 h) (19) Use of vasoactive drugs (norepinephrine dose ≥0.1 μg / (kg) (min), lasting ≥6h) (20) Percutaneous coronary intervention (PCI) (with contrast agent used during the procedure and procedure duration ≥1 hour) (21) The operation lasts ≥3 hours (general anesthesia, and involves major surgery types such as abdominal and cardiovascular surgery) (22) Massive transfusion (≥10U of red blood cell suspension or ≥2000ml of whole blood within 24 hours) (23) Kidney biopsy (the kidney biopsy procedure should be completed within the last 7 days) (24) Intra-abdominal hypertension (intra-abdominal pressure ≥12mmHg and lasting ≥6h) (25) Multiple organ dysfunction syndrome (MODS) (with dysfunction of two or more organs simultaneously) Real-time monitoring layer (8 categories, dynamic indicators reflecting changes in renal function) (26) Serum creatinine increased by ≥26.5 μmol / L within 48 hours (based on laboratory test values, excluding non-renal factors). (27) Serum creatinine increases by ≥1.5 times the baseline value within 7 days (baseline value is the lowest value within 3 months prior to admission or the initial value at admission). (28) Urine output <0.5 ml / (kg) h) and lasting ≥6 hours (excluding mechanical factors such as urinary retention and catheter blockage) (29) Serum potassium ≥ 5.5 mmol / L (excluding non-renal hyperkalemia factors such as hemolysis and specimen error) (30) Estimated glomerular filtration rate (eGFR) <60 ml / (min) 1.73m² (Calculated using the CKD-EPI formula) (31) Metabolic acidosis (arterial blood pH < 7.35 and standard bicarbonate < 22 mmol / L) (32) Urine sodium ≥ 40 mmol / L (urine sodium test value, excluding the influence of diuretic use) (33) Urine specific gravity < 1.010 (routine urine test value, reflecting abnormal renal concentrating function) In one possible implementation, the data processing device provided in the embodiments of this application can be integrated with the "Intelligent Connected Critical Care System". Figure 4 This is a schematic diagram of a system deployment architecture provided for an embodiment of this application. Figure 4 As shown, a three-layer architecture is adopted: a data layer, an algorithm layer, and an application layer. The data layer can interface with multiple devices and systems via the HL7FHIR protocol. Acquired medical data is stored in a data buffer queue called Kafka. Then, the data in this queue is processed based on a data preprocessing stream, which may include time alignment, outlier handling, and missing value handling. After data preprocessing, data quality monitoring can be performed to ensure data integrity ≥98% and data acquisition latency ≤10s. If these data quality monitoring requirements are not met, an anomaly alarm can be triggered.
[0100] The algorithm layer integrates the processing procedures described in the above embodiments into a Docker container via Docker containerization. Specifically, it can integrate clinical pathway coding services, data fusion services, and early warning model inference services. To improve data processing efficiency, TensorRT can be used to accelerate inference, reducing latency to ≤500ms.
[0101] The analysis results output by the algorithm layer can be temporarily stored in Redis, allowing the application layer to visualize them in the newly added AKI early warning module of the intelligent critical care system. Specifically, this can include displaying risk levels, time windows, risk triggers, and historical trends.
[0102] In the medical field, the criteria for assessing different risks are iteratively updated. However, the risk assessment rules in related technologies are hard-coded and cannot be dynamically integrated with new standards. This results in technological iteration lagging behind reality, and a significant decrease in early warning accuracy after long-term use. Therefore, in this embodiment, a dynamic update mechanism is also configured, namely knowledge update and model update. Knowledge update is triggered when guidelines are updated or ≥2 new Level I evidence studies (sample size ≥1000 cases) are added. This triggers a simple update, recalculates weights based on the AHP algorithm, and updates the coding rules. Model update is triggered when ≥1000 new labeled data cases are added. This initiates incremental training, and if the area under the curve (AUC) on the validation set increases by ≥2%, the online model is replaced.
[0103] It can be directly integrated into existing intensive care systems without the need for additional hardware, thus reducing the cost of hospital IT transformation.
[0104] It should be noted that the analysis results obtained in this application embodiment do not directly affect the human body, but rather serve as a risk warning to remind relevant personnel to pay close attention to the occurrence of risk events.
[0105] In this embodiment, clinical pathway knowledge is transformed into feature vectors that can be recognized by AI models through multi-dimensional encoding, thereby achieving deep integration of clinical knowledge and algorithms.
[0106] In one possible implementation, the prediction model provided in this application is a bidirectional LSTM model that incorporates a clinical pathway embedding layer. Figure 5 This is a schematic diagram of a prediction model structure provided in an embodiment of this application. Figure 5As shown, the prediction model includes an input layer, a clinical pathway embedding layer, a bidirectional LSTM layer, and an output layer. The specific data processing flow is as follows: Input layer (outputs 256-dimensional features) → Clinical pathway embedding layer (i.e., a 128-dimensional fully connected layer) → Bidirectional LSTM layer (two layers, each with 64 neurons, dropout rate 0.2) → Output layer (using the Sigmoid activation function, outputting the AKI occurrence probability). Specifically, the warning result can be analyzed based on this occurrence probability, mainly including: risk level, time window, and risk trigger. The risk level can include: high risk: occurrence probability ≥ 70%, medium risk: occurrence probability 30%-69%, low risk: occurrence probability less than 30%. In this embodiment, the next 72 hours are divided into three time windows, and the window corresponding to the highest occurrence probability is taken as the time period of risk occurrence. For example, the 72 hours can be divided into three time windows: 0-24h, 24-48h, and 48-72h. Risk triggers are calculated based on the feature contribution of the SHAP algorithm, and the top 3 risk triggers are output.
[0107] In this embodiment, 8000 labeled data points can be used (3200 positive samples and 4800 negative samples, divided into training set / validation set / test set in a 7:1:2 ratio); the optimizer is AdamW (initial learning rate 1e-4, weight decay 1e-5), the loss function is weighted cross-entropy (positive sample weight 1.5); an early stopping mechanism is set (if the validation set AUC does not improve for 5 consecutive rounds, the training objective is AUC ≥ 0.9 and recall ≥ 0.85.
[0108] Tests revealed that when the data processing device provided in this application was tested using a test set, the AUC obtained based on the test set was 0.92, which is 10%-23% higher than that of related technologies (AUC 0.75-0.85), with a 40% reduction in false negative rate and a 35% reduction in false negative rate.
[0109] The data processing device provided in this application enables early warning 24 hours before a risk occurs, which is 200%-400% better than related technologies (lagging by 6-12 hours).
[0110] The algorithm outputs "risk level + time window + specific trigger", with a user adoption rate of ≥90%, solving the problem of implementing "black box algorithms".
[0111] The data processing procedure is illustrated below with a specific example. Assume that medical staff compile medical data for a target patient and send this data to the data processing device via a edge medical device. Assume the target patient is a 68-year-old male patient in the ICU of a hospital, weighing 65 kg, with an APACHE II score of 28. His admission diagnosis was "pulmonary infection, septic shock, stage 3 chronic kidney disease, and type 2 diabetes." After admission, he received "mechanical ventilation (PEEP = 12 cmH2O) + fluid resuscitation (60 ml / kg of fluid replacement every 24 hours) + cefoperazone / sulbactam for anti-infection treatment," and the risk of AKI within 72 hours of admission needed to be monitored.
[0112] Data collection: 24 hours after admission, serum creatinine was 156 μmol / L (baseline 112 μmol / L), and hourly urine output was 28 ml / h (0.43 ml / (kg)). h)), matching nodes include "stage 3 chronic kidney disease, type 2 diabetes, sepsis, mechanical ventilation, fluid resuscitation, elevated serum creatinine, and abnormal urine output"; Path coding: A 30-minute sliding window is used. The risk value of a certain window = sepsis (0.85) + mechanical ventilation (0.6) + abnormal urine output (0.6) = 2.05 (0.78 after standardization). GCN outputs 32-dimensional association features. Data fusion: Serum creatinine and heart rate data were aligned using DTW (error 3 min), missing values of 1.5 h urine output were filled in (LSTM interpolation), and a 256-dimensional feature vector was generated; Model inference: The probability of AKI occurring 48 hours after admission is 78% (high risk), with an expected time window of 48-72 hours. The top 3 causes are sepsis (38%), abnormal urine output (30%), and negative fluid balance (22%). Clinical intervention: It was recommended to adjust the fluid resuscitation volume to 80 ml / kg and increase the frequency of renal function monitoring to 6 hours / time. The patient's serum creatinine dropped to 148 μmol / L after 72 hours, and no AKI occurred, thus verifying the effectiveness of the treatment plan.
[0113] Figure 6 This is a schematic diagram illustrating the display of analysis results provided in an embodiment of this application. After obtaining the analysis results, corresponding information can be displayed in each designated area of the AKI intelligent early warning interface. For example... Figure 6 As shown, the interface includes a patient basic information section, an early warning status overview area, a risk factor analysis area, a historical trend review, clinical pathway node status, and recommended intervention measures.
[0114] The patient basic information section displays basic information about the target patient, such as: Bed: Bed 08; Name: Li XX; Gender: Male; Age: 68 years old; Diagnosis: Lung infection, septic shock, stage 3 chronic kidney disease, type 2 diabetes, APACHE II score: 28.
[0115] The alert status overview displays the analysis results output by the predictive model. For example, if the current AKI risk level has a 78% probability of being high risk, it will be displayed with a red warning indicator. The alert time window is 48-72 hours, and the expected occurrence time is after 14:30 on June 17, 2024.
[0116] The risk factor analysis area displays the main contributing factors to the current risk. For example, it shows the top 3 risk factors: 1. Sepsis (contribution 38%), abnormal urine output (contribution 30%), and negative fluid balance (contribution 22%). For user convenience, the SHAP value analysis chart can also be displayed, showing the distribution of feature importance.
[0117] Historical trend review is used to display trend graphs, such as an AKI risk probability trend graph, which is a prediction from the past 72 hours to the next 72 hours. Specifically, key events can be marked on the timeline, such as contrast agent use, dehydration treatment, etc.
[0118] The status of clinical pathway nodes is used to display the current risk factors involved, i.e., the activated clinical pathway nodes. Specifically, the basic risk layer can be marked as activated for chronic kidney disease stage 3 and type 2 diabetes, the treatment operation layer for sepsis, mechanical ventilation and fluid resuscitation, and the real-time monitoring layer for elevated serum creatinine and abnormal urine output.
[0119] Recommended interventions are used to indicate clinical intervention recommendations. Specifically, these recommendations can be based on KDIGO guidelines, such as immediately adjusting fluid resuscitation to 80 ml / kg, increasing the frequency of renal function monitoring to 6 hours / time, avoiding the use of nephrotoxic drugs, and considering preparation for renal replacement therapy. It should be noted that these clinical intervention recommendations are only intended to guide users, and their specific adoption is up to the user.
[0120] Based on the same inventive concept, embodiments of this application provide a data processing method. Figure 7 This is a schematic diagram of a data processing method provided in an embodiment of this application, such as... Figure 7 As shown, it includes the following steps: S701: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnosis and treatment data.
[0121] S702: Extract a clinical feature matrix from the medical data, the clinical feature matrix being used to describe whether the medical data contains risk factors.
[0122] S703: Input the clinical feature matrix into the pre-trained prediction model to obtain the analysis results, and output the analysis results so that the user can perform analysis based on the analysis results.
[0123] In one possible implementation, the step of extracting a clinical feature matrix from the medical data includes: Based on the acquisition time, set window duration, and set window sliding step size of each parameter in the medical data, the medical data is divided into multiple sub-data sets; For each risk factor, based on the parameters in each sub-data set and the judgment criteria of the risk factor, it is determined whether the corresponding sub-data set involves the risk factor. If it exists, the first value is filled into the preset position in the original clinical feature matrix; if it does not exist, the second value is filled into the preset position. The original clinical feature matrix, which is filled in completely, is determined as the clinical feature matrix.
[0124] In one possible implementation, the process of determining the set window duration and the set window sliding step size includes: Based on the aforementioned medical data, the target Acute Physiology and Chronic Health Assessment (APACHE II) score was determined. Based on the pre-configured correspondence between different scores and sliding window strategies, the target sliding window strategy corresponding to the target acute physiological and chronic health scores is determined. Each sliding window strategy includes the time window length and the window sliding step size. The time window length in the target sliding window strategy is determined as the set window duration, and the window sliding step size in the target sliding window strategy is determined as the set window sliding step size.
[0125] In one possible implementation, the method further includes: The risk feature matrix is determined based on the clinical feature matrix and the weight corresponding to each risk factor; The clinical feature matrix and the risk feature matrix are fused to obtain a first fusion matrix, and the clinical feature matrix is updated using the first fusion matrix.
[0126] In one possible implementation, the process of determining the weight corresponding to each risk factor includes: Obtain a preset number of judgment matrices, which are used to describe the risk importance ratio between every two risk factors. Different judgment matrices are determined based on the risk importance ratios judged by different experts. The analytic hierarchy process (AHP) is used to analyze the preset number of judgment matrices to determine the weight of each risk factor.
[0127] In one possible implementation, before fusing the clinical feature matrix and the risk feature matrix to obtain the first fusion matrix, the method further includes: Obtain a pre-configured adjacency matrix and a risk trigger feature matrix. The adjacency matrix is used to describe the impact of the simultaneous existence of any two risk triggers on the occurrence of risk. The risk trigger feature matrix is used to describe the information of each risk trigger about each preset attribute. The adjacency matrix and the risk factor feature matrix are input into a pre-trained graph convolutional network to obtain the path association feature matrix of the risk factors. The process of fusing the clinical feature matrix and the risk feature matrix to obtain a first fusion matrix includes: The clinical feature matrix, the risk feature matrix, and the path association feature matrix are fused to obtain the first fusion matrix.
[0128] In one possible implementation, fusing the clinical feature matrix, the risk feature matrix, and the path association feature matrix to obtain the first fusion matrix includes: A multi-head attention mechanism is used to process the clinical feature matrix, the risk feature matrix, and the path association feature matrix to determine the weights corresponding to the clinical feature matrix, the risk feature matrix, and the path association feature matrix, respectively. The first fusion matrix is determined based on the clinical feature matrix, the risk feature matrix, the path association feature matrix, and each weight.
[0129] In one possible implementation, after extracting the clinical feature matrix from the medical data and before inputting the clinical feature matrix into a pre-trained prediction model to obtain the analysis results, the method further includes: For each parameter in the medical data, the occurrence label corresponding to the parameter is determined according to the occurrence judgment criteria of the target event corresponding to the parameter; Based on the diagnostic data of historical objects and this parameter, a first marginal probability, a second marginal probability, and a joint probability are determined; wherein, the first marginal probability is the proportion of the historical object's diagnostic data where the parameter value exists; the second marginal probability is the proportion of the historical object's diagnostic data where the event occurrence label corresponding to this parameter is the target event occurrence label; and the joint probability is the proportion of the historical object's diagnostic data where the parameter value exists and the event occurrence label corresponding to this parameter is the target event occurrence label. The entropy value corresponding to this parameter is determined based on the first edge probability, the second edge probability, the joint probability, and the mutual information entropy algorithm. Based on the entropy value corresponding to each parameter and the preset threshold, each parameter in the medical data is filtered to obtain the core parameters; The clinical feature matrix and the core parameters are fused to obtain a second fusion matrix, and the clinical feature matrix is updated using the second fusion matrix.
[0130] In one possible implementation, the method further includes: The clinical feature matrix is processed based on the model result interpretation algorithm to obtain the contribution of each risk factor; Based on the contribution of each risk factor, select the target risk factor and output it.
[0131] Based on the same inventive concept, embodiments of this application provide an electronic device that can implement the steps of the data processing method described above. Figure 8 This application provides a schematic diagram of an electronic device structure, such as... Figure 8 As shown, it includes: processor 801, communication interface 802, memory 803 and communication bus 804, wherein processor 801, communication interface 802 and memory 803 communicate with each other through communication bus 804. The memory 803 stores a computer program. When the program is executed by the processor 801, the processor 801 performs the following steps: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; A clinical feature matrix is extracted from the medical data, and the clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, which are then output so that users can perform analysis based on these results.
[0132] In one possible implementation, the step of extracting a clinical feature matrix from the medical data includes: Based on the acquisition time, set window duration, and set window sliding step size of each parameter in the medical data, the medical data is divided into multiple sub-data sets; For each risk factor, based on the parameters in each sub-data set and the judgment criteria of the risk factor, it is determined whether the corresponding sub-data set involves the risk factor. If it exists, the first value is filled into the preset position in the original clinical feature matrix; if it does not exist, the second value is filled into the preset position. The original clinical feature matrix, which is filled in completely, is determined as the clinical feature matrix.
[0133] In one possible implementation, the process of determining the set window duration and the set window sliding step size includes: Based on the aforementioned medical data, the target Acute Physiology and Chronic Health Assessment (APACHE II) score was determined. Based on the pre-configured correspondence between different scores and sliding window strategies, the target sliding window strategy corresponding to the target acute physiological and chronic health scores is determined. Each sliding window strategy includes the time window length and the window sliding step size. The time window length in the target sliding window strategy is determined as the set window duration, and the window sliding step size in the target sliding window strategy is determined as the set window sliding step size.
[0134] In one possible implementation, the method further includes: The risk feature matrix is determined based on the clinical feature matrix and the weight corresponding to each risk factor; The clinical feature matrix and the risk feature matrix are fused to obtain a first fusion matrix, and the clinical feature matrix is updated using the first fusion matrix.
[0135] In one possible implementation, the process of determining the weight corresponding to each risk factor includes: Obtain a preset number of judgment matrices, which are used to describe the risk importance ratio between every two risk factors. Different judgment matrices are determined based on the risk importance ratios judged by different experts. The analytic hierarchy process (AHP) is used to analyze the preset number of judgment matrices to determine the weight of each risk factor.
[0136] In one possible implementation, before fusing the clinical feature matrix and the risk feature matrix to obtain the first fusion matrix, the method further includes: Obtain a pre-configured adjacency matrix and a risk trigger feature matrix. The adjacency matrix is used to describe the impact of the simultaneous existence of any two risk triggers on the occurrence of risk. The risk trigger feature matrix is used to describe the information of each risk trigger about each preset attribute. The adjacency matrix and the risk factor feature matrix are input into a pre-trained graph convolutional network to obtain the path association feature matrix of the risk factors. The process of fusing the clinical feature matrix and the risk feature matrix to obtain a first fusion matrix includes: The clinical feature matrix, the risk feature matrix, and the path association feature matrix are fused to obtain the first fusion matrix.
[0137] In one possible implementation, fusing the clinical feature matrix, the risk feature matrix, and the path association feature matrix to obtain the first fusion matrix includes: A multi-head attention mechanism is used to process the clinical feature matrix, the risk feature matrix, and the path association feature matrix to determine the weights corresponding to the clinical feature matrix, the risk feature matrix, and the path association feature matrix, respectively. The first fusion matrix is determined based on the clinical feature matrix, the risk feature matrix, the path association feature matrix, and each weight.
[0138] In one possible implementation, after extracting the clinical feature matrix from the medical data and before inputting the clinical feature matrix into a pre-trained prediction model to obtain the analysis results, the method further includes: For each parameter in the medical data, the occurrence label corresponding to the parameter is determined according to the occurrence judgment criteria of the target event corresponding to the parameter; Based on the diagnostic data of historical objects and this parameter, a first marginal probability, a second marginal probability, and a joint probability are determined; wherein, the first marginal probability is the proportion of the historical object's diagnostic data where the parameter value exists; the second marginal probability is the proportion of the historical object's diagnostic data where the event occurrence label corresponding to this parameter is the target event occurrence label; and the joint probability is the proportion of the historical object's diagnostic data where the parameter value exists and the event occurrence label corresponding to this parameter is the target event occurrence label. The entropy value corresponding to this parameter is determined based on the first edge probability, the second edge probability, the joint probability, and the mutual information entropy algorithm. Based on the entropy value corresponding to each parameter and the preset threshold, each parameter in the medical data is filtered to obtain the core parameters; The clinical feature matrix and the core parameters are fused to obtain a second fusion matrix, and the clinical feature matrix is updated using the second fusion matrix.
[0139] In one possible implementation, the method further includes: The clinical feature matrix is processed based on the model result interpretation algorithm to obtain the contribution of each risk factor; Based on the contribution of each risk factor, select the target risk factor and output it.
[0140] The communication bus mentioned in the aforementioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 702 is used for communication between the aforementioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.
[0141] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0142] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; A clinical feature matrix is extracted from the medical data, and the clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, which are then output so that users can perform analysis based on these results.
[0143] In one possible implementation, the step of extracting a clinical feature matrix from the medical data includes: Based on the acquisition time, set window duration, and set window sliding step size of each parameter in the medical data, the medical data is divided into multiple sub-data sets; For each risk factor, based on the parameters in each sub-data set and the judgment criteria of the risk factor, it is determined whether the corresponding sub-data set involves the risk factor. If it exists, the first value is filled into the preset position in the original clinical feature matrix; if it does not exist, the second value is filled into the preset position. The original clinical feature matrix, which is filled in completely, is determined as the clinical feature matrix.
[0144] In one possible implementation, the process of determining the set window duration and the set window sliding step size includes: Based on the aforementioned medical data, the target Acute Physiology and Chronic Health Assessment (APACHE II) score was determined. Based on the pre-configured correspondence between different scores and sliding window strategies, the target sliding window strategy corresponding to the target acute physiological and chronic health scores is determined. Each sliding window strategy includes the time window length and the window sliding step size. The time window length in the target sliding window strategy is determined as the set window duration, and the window sliding step size in the target sliding window strategy is determined as the set window sliding step size.
[0145] In one possible implementation, the method further includes: The risk feature matrix is determined based on the clinical feature matrix and the weight corresponding to each risk factor; The clinical feature matrix and the risk feature matrix are fused to obtain a first fusion matrix, and the clinical feature matrix is updated using the first fusion matrix.
[0146] In one possible implementation, the process of determining the weight corresponding to each risk factor includes: Obtain a preset number of judgment matrices, which are used to describe the risk importance ratio between every two risk factors. Different judgment matrices are determined based on the risk importance ratios judged by different experts. The analytic hierarchy process (AHP) is used to analyze the preset number of judgment matrices to determine the weight of each risk factor.
[0147] In one possible implementation, before fusing the clinical feature matrix and the risk feature matrix to obtain the first fusion matrix, the method further includes: Obtain a pre-configured adjacency matrix and a risk trigger feature matrix. The adjacency matrix is used to describe the impact of the simultaneous existence of any two risk triggers on the occurrence of risk. The risk trigger feature matrix is used to describe the information of each risk trigger about each preset attribute. The adjacency matrix and the risk factor feature matrix are input into a pre-trained graph convolutional network to obtain the path association feature matrix of the risk factors. The process of fusing the clinical feature matrix and the risk feature matrix to obtain a first fusion matrix includes: The clinical feature matrix, the risk feature matrix, and the path association feature matrix are fused to obtain the first fusion matrix.
[0148] In one possible implementation, fusing the clinical feature matrix, the risk feature matrix, and the path association feature matrix to obtain the first fusion matrix includes: A multi-head attention mechanism is used to process the clinical feature matrix, the risk feature matrix, and the path association feature matrix to determine the weights corresponding to the clinical feature matrix, the risk feature matrix, and the path association feature matrix, respectively. The first fusion matrix is determined based on the clinical feature matrix, the risk feature matrix, the path association feature matrix, and each weight.
[0149] In one possible implementation, after extracting the clinical feature matrix from the medical data and before inputting the clinical feature matrix into a pre-trained prediction model to obtain the analysis results, the method further includes: For each parameter in the medical data, the occurrence label corresponding to the parameter is determined according to the occurrence judgment criteria of the target event corresponding to the parameter; Based on the diagnostic data of historical objects and this parameter, a first marginal probability, a second marginal probability, and a joint probability are determined; wherein, the first marginal probability is the proportion of the historical object's diagnostic data where the parameter value exists; the second marginal probability is the proportion of the historical object's diagnostic data where the event occurrence label corresponding to this parameter is the target event occurrence label; and the joint probability is the proportion of the historical object's diagnostic data where the parameter value exists and the event occurrence label corresponding to this parameter is the target event occurrence label. The entropy value corresponding to this parameter is determined based on the first edge probability, the second edge probability, the joint probability, and the mutual information entropy algorithm. Based on the entropy value corresponding to each parameter and the preset threshold, each parameter in the medical data is filtered to obtain the core parameters; The clinical feature matrix and the core parameters are fused to obtain a second fusion matrix, and the clinical feature matrix is updated using the second fusion matrix.
[0150] In one possible implementation, the method further includes: The clinical feature matrix is processed based on the model result interpretation algorithm to obtain the contribution of each risk factor; Based on the contribution of each risk factor, select the target risk factor and output it.
[0151] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to execute any of the data processing methods described above. Implementation of the above-described computer program product can be found in the implementation of the method; repeated details will not be elaborated further.
[0152] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0153] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0154] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0155] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of user-operated steps to be executed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0156] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data processing apparatus, characterized in that, The device includes: The receiving module is used to receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; The data processing module is used to extract a clinical feature matrix from the medical data, the clinical feature matrix being used to describe whether the medical data contains risk factors; input the clinical feature matrix into a pre-trained prediction model to obtain analysis results, and output the analysis results so that users can perform analysis based on the analysis results.
2. The apparatus according to claim 1, characterized in that, The data processing module is specifically used to divide the medical data into multiple sub-data sets according to the acquisition time, set window duration, and set window sliding step size of each parameter in the medical data; for each risk factor, it determines whether the corresponding sub-data set involves the risk factor according to the parameters in each sub-data set and the judgment criteria of the risk factor; if it exists, it fills the first value into the preset position in the original clinical feature matrix; if it does not exist, it fills the second value into the preset position. The original clinical feature matrix, which is filled in completely, is determined as the clinical feature matrix.
3. The apparatus according to claim 2, characterized in that, The device further includes: The determination module is used to determine the target Acute Physiology and Chronic Health Score (APACHE II) based on the medical data; determine the target sliding window strategy corresponding to the target APACHE II score based on the pre-configured correspondence between different scores and sliding window strategies, wherein any sliding window strategy includes a time window length and a window sliding step size; determine the time window length in the target sliding window strategy as the set window duration, and determine the window sliding step size in the target sliding window strategy as the set window sliding step size.
4. The apparatus according to claim 2, characterized in that, The device further includes: The determination module is used to determine the risk feature matrix based on the clinical feature matrix and the weight corresponding to each risk factor; The fusion module is used to fuse the clinical feature matrix and the risk feature matrix to obtain a first fusion matrix, and to update the clinical feature matrix using the first fusion matrix.
5. The apparatus according to claim 4, characterized in that, The determining module is specifically used to obtain a preset number of judgment matrices. The judgment matrices are used to describe the risk importance ratio between every two risk factors. Different judgment matrices are determined based on the risk importance ratios judged by different experts. The analytic hierarchy process (AHP) is used to analyze the predetermined number of judgment matrices to determine the weight of each risk factor.
6. The apparatus according to claim 4, characterized in that, The data processing module is further configured to obtain a pre-configured adjacency matrix and a risk trigger feature matrix. The adjacency matrix describes the impact of the simultaneous existence of any two risk triggers on the occurrence of a risk. The risk trigger feature matrix describes the information of each risk trigger about each preset attribute. The adjacency matrix and the risk trigger feature matrix are input into a pre-trained graph convolutional network to obtain a path association feature matrix of the risk triggers. The fusion module is specifically used to fuse the clinical feature matrix, the risk feature matrix, and the path association feature matrix to obtain the first fusion matrix.
7. The apparatus according to claim 6, characterized in that, The fusion module is specifically used to process the clinical feature matrix, the risk feature matrix, and the path association feature matrix using a multi-head attention mechanism to determine the weights corresponding to the clinical feature matrix, the risk feature matrix, and the path association feature matrix, respectively. The first fusion matrix is determined based on the clinical feature matrix, the risk feature matrix, the path association feature matrix, and each weight.
8. The apparatus according to any one of claims 1-7, characterized in that, The data processing module is further configured to: determine the occurrence label corresponding to each parameter in the medical data according to the target event occurrence judgment criteria corresponding to the parameter; determine a first marginal probability, a second marginal probability, and a joint probability based on the diagnostic data of historical objects and the parameter; wherein, the first marginal probability is the proportion of the parameter value present in the diagnostic data of the historical objects; the second marginal probability is the proportion of the event occurrence label corresponding to the parameter in the diagnostic data of the historical objects being the target event occurrence label; the joint probability is the proportion of the parameter value present in the diagnostic data of the historical objects and the event occurrence label corresponding to the parameter being the target event occurrence label; determine the entropy value corresponding to the parameter based on the first marginal probability, the second marginal probability, the joint probability, and the mutual information entropy algorithm; filter each parameter in the medical data based on the entropy value corresponding to each parameter and a preset threshold to obtain core parameters; fuse the clinical feature matrix and the core parameters to obtain a second fusion matrix, and update the clinical feature matrix using the second fusion matrix.
9. The apparatus according to claim 1, characterized in that, The data processing module is also used to process the clinical feature matrix based on the model result interpretation algorithm to obtain the contribution degree corresponding to each risk factor; and to select and output the target risk factor according to the contribution degree corresponding to each risk factor.
10. An electronic device, characterized in that, The electronic device includes a processor, which performs the following functions when executing a computer program stored in a memory: Receive medical data of the target object sent by the edge medical device, the medical data including vital signs data, laboratory data and diagnostic data; A clinical feature matrix is extracted from the medical data, and the clinical feature matrix is used to describe whether the medical data has risk factors. The clinical feature matrix is input into a pre-trained prediction model to obtain analysis results, which are then output so that users can perform analysis based on these results.