Efficient perioperative period time series data recovery method and device based on Hash hidden feature learning
Through the combination of hash hidden feature learning and adaptive differential evolution algorithm, the high-dimensional and noise problems of perioperative physiological timing data are solved, and efficient and accurate data recovery and early warning model are achieved.
Patent Information
- Application Number
- CN202510462375.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing tensor hidden feature analysis method is difficult to efficiently process physiological timing data with high dimensionality, high noise, and data loss during the perioperative period, resulting in insufficient accuracy and efficiency of the early warning model and unable to meet the real-time requirements.
The end-to-end framework of hash hidden feature learning combined with the LF model is adopted, and the adaptive differential evolution algorithm is used for direct discrete optimization. The data is binary encoding through hashing operations, reducing storage and communication overhead, and efficient recovery of perioperative physiological timing data is achieved.
It effectively reduces the impact of mixed noise and high proportion of missing data, realizes efficient and accurate recovery of perioperative physiological timing data, and improves the accuracy and efficiency of the early warning model.
Smart Images

Figure CN120376020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an efficient recovery method and device for perioperative time-series data based on hash hidden feature learning, belonging to the fields of medical and health big data and data recovery, and is particularly applicable to the efficient recovery of perioperative time-series data based on hash hidden feature learning. Background Art
[0002] The incompleteness of perioperative physiological time-series data is mainly caused by the combined effects of dynamic acquisition limitations and technical barriers of multi-source heterogeneous data. On the one hand, there are significant differences in the sampling frequencies of different physiological indicators (such as electrocardiogram, blood pressure, and blood oxygen), making it difficult to synchronize and integrate high-frequency monitoring data and low-frequency recorded data. On the other hand, when multiple devices are used for collaborative monitoring, problems such as insufficient compatibility and heterogeneous data formats occur, leading to fragmentation of multi-source information. In addition, factors such as equipment operation interference, measurement technical errors, desensitization loss caused by data privacy protection, and human operation omissions further cause high-frequency time-series errors, mixed noise interference, and missing specific fields in the data. These factors make perioperative data exhibit characteristics such as high dimensionality, high noise, and time-series breaks. At the same time, due to its dynamic change characteristics, the data is discontinuous and heterogeneous in the time and space dimensions, forming a typical high-dimensional and incomplete (HDI) data structure, which greatly increases the difficulty of analysis.
[0003] The contradiction between the high fatality rate of perioperative cardiovascular critical adverse events and the lag of current early warning technologies highlights the urgency of exploring the potential value of incomplete data. Traditional methods are difficult to extract effective information from physiological time-series data with high noise and high missing rates, while artificial intelligence technologies (such as machine learning) can discover hidden risk signals through dynamic pattern recognition. For example, postoperative cardiovascular events can be predicted based on intraoperative monitoring data. The introduction of methods such as tensor hidden feature analysis (LF), by modeling high-dimensional incomplete data as low-rank tensors, can effectively extract the time-series correlations and potential pathological features of multiple physiological indicators, improving the robustness and accuracy of early warning models.
[0004] Although low-dimensional embedding methods based on tensor decomposition provide new ideas for processing HDI data, their practical applications still face multiple challenges. The exponential growth of medical data has led to a sharp increase in the computational cost of traditional continuous embedding spaces, and the real-time requirements of perioperative data further pose higher demands on algorithm efficiency. Breaking through these bottlenecks requires innovating lightweight tensor calculation frameworks and integrating domain knowledge to optimize feature extraction logic.
[0005] In summary, existing methods based on tensor hidden feature analysis still face great difficulties in dealing with high-dimensional, high-noise, and data-missing physiological time-series data during perioperative treatment processes. Summary of the Invention
[0006] In view of this, aiming at the problems of the existing technology, the present invention provides an efficient recovery method and device for perioperative time-series data based on hash hidden feature learning, intending to reduce data storage and communication overhead by introducing hash operations to represent data, and then combining the hash operations with an end-to-end framework of the LF model, and using an adaptive differential evolution algorithm for direct discrete optimization to avoid approximate loss and reduce information loss, and finally achieve efficient and accurate recovery of perioperative physiological time-series data.
[0007] To achieve the above object, an efficient recovery device for perioperative time-series data based on hash hidden feature learning consists of a server and N data acquisition devices, which communicate with each other through a network for data communication; the server includes: a server data storage module and a data recovery module; one of the data acquisition devices is a local computer, including: a local physiological time-series data receiving module and a local physiological time-series data storage module; where the value of N depends on the number of physiological time-series data acquisition devices;
[0008] The local physiological time-series data receiving module in the data acquisition device is connected to the physiological time-series data acquisition device through a network, and is used to receive the incomplete physiological time-series data fed back by the physiological time-series data acquisition device, and instruct the local physiological time-series data storage module to store the incomplete physiological time-series data;
[0009] The local physiological time-series data storage module in the data acquisition device is connected to the server data storage module in the server through a network, and sends the incomplete physiological time-series data to the server data storage module;
[0010] The server data storage module and the data recovery module are used to send the incomplete physiological time-series data of all data acquisition devices to the data recovery module for data recovery processing, and then save the recovered perioperative physiological time-series data;
[0011] The data recovery module uses the hash hidden feature learning method to recover the incomplete physiological time-series data.
[0012] Preferably, the server data storage module includes a cloud physiological time-series data storage unit, a data preprocessing unit, and a recovered data storage unit;
[0013] The cloud physiological time-series data storage unit in the server data storage module is used to receive and store the incomplete physiological time-series data sent by the local physiological time-series data storage module in the data acquisition device.
[0014] The data preprocessing unit, which is connected to the cloud physiological time-series data storage unit and the data recovery module, is used to preprocess the incomplete physiological time-series data of all data acquisition devices;
[0015] The described restored data storage unit is connected to the data recovery module and is used to store the restored perioperative physiological time series data.
[0016] The present invention also provides an efficient method for restoring perioperative time series data based on hash hidden feature learning, including the following steps:
[0017] S1: The server sends a data acquisition instruction to the data acquisition device, and the data acquisition device collects data from the physiological time series data acquisition device according to the set sampling frequency by using the local physiological time series data receiving module.
[0018] S2: The local physiological time series data storage module stores the collected data and uploads it to the server data storage module of the server.
[0019] S3: The server uses the server data storage module to store the collected data and preprocesses it into incomplete physiological time series data.
[0020] S4: The data recovery module of the server initializes the discrete hidden feature vector by binary quantization.
[0021] S5: The data recovery module of the server updates and learns the discrete hidden feature vector by using the hash hidden feature learning method.
[0022] S6: The data recovery module of the server determines whether the learned discrete hidden feature vector can complete the restoration. If it can, the restored perioperative physiological time series data is stored in the restored data storage unit of the server data storage module; otherwise, an error message is sent through the network.
[0023] Further, the specific step S1 is as follows: The server selects the corresponding data acquisition device from the data acquisition devices communicating with the server network by using the patient ID as an index and sends a data acquisition instruction; after receiving the data acquisition instruction, the data acquisition device collects data from the physiological time series data acquisition device according to the set sampling frequency by using the local physiological time series data receiving module.
[0024] The incomplete physiological time series data in step S3 is the HDI matrix where the patient ID set U and the physiological detection index set I, and the |·| operation represents the number of elements in the set; each element r in R u,i represents the monitoring value of the physiological detection index i∈I monitored by the patient with ID u∈U; the number |Λ| of all known element sets Λ in R is much smaller than the number |Γ| of the set Γ of unknown elements in R.
[0025] Further, the preprocessing process described in step S3 includes:
[0026] S301: Store the uploaded collected data as the HDI matrix R;
[0027] S302: Label the unknown elements and label the missing data in the sampled data with NAN;
[0028] S303: Normalize R according to the known elements.
[0029] The discrete latent feature matrix described in step S4 includes the patient latent feature matrix W and the physiological detection index latent feature matrix Q; where: W is of dimension |U|×D and is used to extract the patient latent features that the extraction device needs; Q is of dimension |I|×D and is used to extract the physiological detection index latent features that the extraction device needs; D is the dimension of the latent feature space and is the bit length of the artificially set hash code.
[0030] Furthermore, the binary quantization initialization process described in step S4 is: independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological detection index latent feature matrix Q until the filling is completed.
[0031] Furthermore, the hash latent feature learning method described in step S5 is specifically as follows:
[0032] S501: Set the maximum number of update learning times T, randomly initialize the hyperparameters α and β in the interval (0, 1), and initialize the current update learning times t = 0; where, T is a positive integer;
[0033] S502: Determine whether the current update learning times t is greater than or equal to T. If it reaches, go to step S6;
[0034] S503: Update the patient latent feature matrix W and the physiological detection index latent feature matrix Q;
[0035] S504: Calculate the loss function before and after the update according to the following formula
[0036]
[0037] where, MEAN(·) is the mean operation; NOT is the not operation; AND is the and operation; w u is the u-th row vector of W, q i is the i-th row vector of Q; w u,d is the element in the u-th row and d-th column of W, q i,d is the element in the i-th row and d-th column of Q;
[0038] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error. Otherwise, return to step S502 to perform the next update.
[0039] Furthermore, step S503 is specifically as follows:
[0040] S5031: Fix the implicit feature matrix Q of physiological detection indicators. Generate multiple samples of the patient implicit feature matrix W row by row using a random algorithm for the implicit feature matrix W of physiological detection indicators, and calculate the fitness corresponding to each sample according to the following formula
[0041]
[0042] S5032: Select the sample with the smallest fitness for each row and update the corresponding row vector of the patient implicit feature matrix W;
[0043] S5033: Fix the implicit feature matrix W of physiological detection indicators. Generate multiple samples of the implicit feature matrix Q of physiological detection indicators row by row using a random algorithm for the implicit feature matrix Q of physiological detection indicators, and calculate the adaptation value corresponding to each sample according to the following formula
[0044]
[0045] S5034: Select the sample with the smallest fitness for each row and update the corresponding row vector of the implicit feature matrix Q of physiological detection indicators.
[0046] Furthermore, the random algorithm described in step S5031 and step S5033 is the particle swarm algorithm.
[0047] Preferably, the random algorithm described in step S5031 and step S5033 is the adaptive differential evolution algorithm, and the adaptive differential evolution algorithm is specifically as follows:
[0048] 1) Initially, set the maximum number of update generations G and the current number of update generations g = 1, and randomly generate a binary discrete population X g ∈{+1, -1} NP×D ; where NP is the number of samples;
[0049] 2) Select a D-dimensional column vector from X g as a sample and calculate its corresponding fitness, and then record the sample with the smallest fitness as where k = 1,..., NP;
[0050] 3) Calculate the gene mutation coefficient F of the current update generation g ;
[0051]
[0052] Among them, rand1, rand2 and rand3 are random numbers between 0 and 1; at the initial moment F g A random number between 0 and 1.
[0053] 4) From X g Randomly select 4 different D-dimensional column vector samples from Calculate the mutant genes of sample numbers k=1,...,NP:
[0054]
[0055] 5) Binarize the mutant gene of the kth sample element by element as:
[0056]
[0057] in,
[0058] 6) The mutant gene of the kth sample will be element by element Perform crossover genetic operations to obtain crossover genes
[0059]
[0060] Among them, rand4 is a random number between 0 and 1, CR is a real number threshold between 0 and 1 set by humans, and d rand is a random index of [1,D], for the d-th element of ;
[0061] 7) Select the next generation of samples sample by sample:
[0062]
[0063] Where k = 1, ..., NP;
[0064] 8) Iterative update algebra g=g+1 and binary discrete population X g , until the maximum number of iterations G set manually.
[0065] Further, the perioperative physiological time series data after recovery described in step S6 in Denormalization is required to restore to the original dimensions.
[0066] Preferably, the criterion for determining whether recovery can be completed in step S6 is the calculated loss function Is it less than a manually set threshold? If so, recovery can be completed, otherwise not.
[0067] The effects of the present invention are as follows: The present invention provides an efficient recovery method and device for perioperative time-series data based on hash hidden feature learning. By introducing hash operations to binary-encode the data, it reduces the data storage and communication overhead. Then, combined with the end-to-end framework of the LF model, it can effectively handle physiological time-series data with mixed noise and a high proportion of missing values, and achieve the recovery of perioperative physiological time-series data. Furthermore, an adaptive differential evolution algorithm is used for direct discrete optimization to avoid approximate loss and reduce information loss, ultimately achieving efficient and accurate recovery of perioperative physiological time-series data. Description of the Drawings
[0068] To make the objectives and technical solutions of the present invention clearer, the following drawings are provided for illustration:
[0069] Figure 1 It is a flowchart of the efficient recovery method for perioperative time-series data based on hash hidden feature learning of the present invention
[0070] Figure 2 It is a framework diagram of the efficient recovery device for perioperative time-series data based on hash hidden feature learning in Embodiment 1 of the present invention;
[0071] Figure 3 It is a schematic diagram of the discrete hidden feature vector update learning in Embodiment 2 of the present invention. Detailed Embodiments
[0072] The present invention will be described in detail below in conjunction with the drawings and embodiments.
[0073] Example 1: There is a hospital that needs to perform real-time analysis of sepsis physiological time-series data for 4,574 perioperative patients in the hospital (once every 15 minutes). The time-series data of 31 features are collected using the physiological time-series data acquisition equipment set in the hospital. These features include: heart rate (HR, beats per minute), oxygen saturation (O2Sat, %), body temperature (Temp, °C), systolic blood pressure (SBP, mm Hg), mean arterial pressure (MAP, mm Hg), diastolic blood pressure (DBP, mm Hg), respiratory rate (Resp, breaths per minute), end-tidal carbon dioxide partial pressure (EtCO2, mm Hg), base excess (BaseExcess, mmol / L), bicarbonate (HCO3, mmol / L), fraction of inspired oxygen (FiO2, %), blood pH, arterial partial pressure of carbon dioxide (PaCO2, mm Hg), arterial oxygen saturation (SaO2, %), aspartate aminotransferase (AST, U / L), blood urea nitrogen (BUN, mg / dL), alkaline phosphatase (Alkalinephos, IU / L), blood calcium (Calcium, mg / dL), blood chloride (Chloride, mmol / L), creatinine (Creatinine, mg / dL), direct bilirubin (Bilirubin direct, mg / dL), blood glucose (Glucose, mg / dL), lactate (Lactate, mg / dL), blood magnesium (Magnesium, mmol / dL), blood phosphorus (Phosphate, mg / dL), blood potassium (Potassium, mmol / L), total bilirubin (Bilirubin total, mg / dL), troponin I (TroponinI, ng / mL), hematocrit (Hct, %), hemoglobin (Hgb, g / dL), partial thromboplastin time (PTT, seconds), white blood cell count (WBC, count×10 3 / μL), fibrinogen (Fibrinogen, mg / dL), and platelet count (Platelets, ×10 3 / μL). Due to problems such as insufficient compatibility and noise during multi-device collaborative monitoring, the proportion of known elements in the entire physiological time-series data is less than 5%. In order to use a deep learning model to extract features from physiological time-series data and analyze the sepsis status of patients, it is urgent to recover unknown data. For this reason, the present invention provides an "Efficient Recovery Device for Perioperative Time-Series Data Based on Hash Hidden Feature Learning", as shown in Figure 2 shown.
[0074] It consists of a server (1) and N data acquisition devices (2), which are connected to each other through a network for data communication; the server (1) includes: a server data storage module (11) and a data recovery module (12); one of the data acquisition devices (2) is a local computer, which includes: a local physiological time series data receiving module (21) and a local physiological time series data storage module (22); where the number of N depends on the number of physiological time series data acquisition devices;
[0075] The local physiological time series data receiving module (21) in the data acquisition device (2) is connected to the physiological time series data acquisition device through a network, and is used to receive the incomplete physiological time series data fed back by the physiological time series data acquisition device, and instruct the local physiological time series data storage module (22) to store the incomplete physiological time series data;
[0076] The local physiological time series data storage module (22) in the data acquisition device (2) is connected to the server data storage module (11) in the server (1) through a network, and sends the incomplete physiological time series data to the server data storage module (11);
[0077] The server data storage module (11) and the data recovery module (12) are used to send the incomplete physiological time series data of all data acquisition devices (2) to the data recovery module (12) for data recovery processing, and then save the recovered perioperative physiological time series data;
[0078] The data recovery module (12) uses the hash hidden feature learning method to recover the incomplete physiological time series data.
[0079] Preferably, the server data storage module (11) includes a cloud physiological time series data storage unit (111), a data preprocessing unit (112), and a recovered data storage unit (113);
[0080] The cloud physiological time series data storage unit (111) in the server data storage module (11) is used to receive and store the incomplete physiological time series data sent by the local physiological time series data storage module (22) in the data acquisition device (2).
[0081] The data preprocessing unit (112), which is connected to the cloud physiological time series data storage unit (111) and the data recovery module (12), is used to preprocess the incomplete physiological time series data of all data acquisition devices (2);
[0082] The recovered data storage unit (113), which is connected to the data recovery module (12), is used to store the recovered perioperative physiological time series data.
[0083] The present invention also provides a "method for efficiently recovering perioperative time-series data based on hash hidden feature learning", which combines Figure 1 , and includes the following steps:
[0084] S1: The server (1) sends a data collection instruction to the data collection device (2), and the data collection device (2) collects the data of the physiological time-series data collection device according to the set sampling frequency by using the local physiological time-series data receiving module (21).
[0085] Specifically: The server (1) selects the corresponding data collection device (2) from the data collection devices (2) that communicate with the server (1) network according to the patient ID as the index, and sends a data collection instruction; after receiving the data collection instruction, the data collection device (2) collects the data of the physiological time-series data collection device according to the set sampling frequency by using the local physiological time-series data receiving module (21).
[0086] S2: The local physiological time-series data storage module (21) stores the collected data and uploads it to the server data storage module (11) of the server (1).
[0087] S3: The server (1) uses the server data storage module (11) to store the collected data and preprocesses it into incomplete physiological time-series data by using the data preprocessing unit (112).
[0088] The incomplete physiological time-series data is an HDI matrix where the patient ID set U and the physiological detection index set I, and the |·| operation represents the number of elements in the set; each element r in R u,i represents the monitoring value of the physiological detection index i∈I for the patient ID u∈U; the number |Λ| of all known element sets Λ in R is much smaller than the number |Γ| of the set Γ of unknown elements in R. Among them, |U| = 4574 and |I| = 31.
[0089] The preprocessing process includes:
[0090] S301: Store the uploaded collected data as the HDI matrix R;
[0091] S302: Mark the unknown elements, and mark the missing sampling data with NAN;
[0092] S303: Normalize R according to the known elements.
[0093] S4: The data recovery module (12) of the server (1) initializes the discrete hidden feature vector by using binary quantization.
[0094] The discrete latent feature matrix described above includes a patient latent feature matrix W and a physiological detection index latent feature matrix Q; where: W is a |U|×D-dimensional matrix, used to extract the patient latent features that the extraction device needs to extract; Q is a |I|×D-dimensional matrix, used to extract the physiological detection index latent features that the extraction device needs to extract; D = 100 is the dimension of the latent feature space, which is the bit length of the artificially set hash code.
[0095] The binary quantization initialization process is as follows: Independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological detection index latent feature matrix Q until the filling is completed.
[0096] S5: The data recovery module (12) of the server (1) uses the hash latent feature learning method to update and learn the discrete latent feature vectors.
[0097] Combined with Table 1, the specific hash latent feature learning method is as follows:
[0098] S501: Set the maximum number of update and learning times T, randomly initialize the hyperparameters α and β in the interval (0, 1), and initialize the current update and learning times t = 0; where, T is a positive integer, and in this embodiment, T = 1000 is taken.
[0099] S502: Determine whether the current update and learning times t is greater than or equal to T. If it reaches, go to step S6.
[0100] S503: Update the patient latent feature matrix W and the physiological detection index latent feature matrix Q. Specifically:
[0101] S5031: Fix the physiological detection index latent feature matrix Q, and use the particle swarm algorithm to generate multiple samples of the patient latent feature matrix W row by row for the patient latent feature matrix W, and calculate the fitness corresponding to each sample according to the following formula
[0102]
[0103] S5032: Select the sample with the minimum fitness for each row and update the corresponding row vector of the patient latent feature matrix W;
[0104] S5033: Fix the patient latent feature matrix W, and use the particle swarm algorithm to generate multiple samples of the physiological detection index latent feature matrix Q row by row for the physiological detection index latent feature matrix Q, and calculate the fitness value corresponding to each sample according to the following formula
[0105]
[0106] S5034: Select the sample with the minimum fitness for each row Update the row vector corresponding to the physiological detection index hidden feature matrix Q with the smallest sample.
[0107] S504: Calculate the loss function before and after the update according to the following formula
[0108]
[0109] where MEAN(·) is the mean operation; NOT is the NOT operation; AND is the AND operation; w u is the u-th row vector of W, q i is the i-th row vector of Q; w u,d is the element in the u-th row and d-th column of W, q i,d is the element in the i-th row and d-th column of Q.
[0110] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error. Otherwise, return to step S502 to perform the next update.
[0111] Table 1 Update learning of discrete hidden feature vectors
[0112]
[0113] S6: The data recovery module (12) of the server (1) determines whether the learned discrete hidden feature vector can complete the recovery. If it can, store the recovered perioperative physiological time series data in the recovery data storage unit (113). Otherwise, send an error message through the network.
[0114] The recovered perioperative physiological time series data where needs to be inverse standardized to restore to the original dimension.
[0115] The criterion for determining whether the recovery can be completed is whether the calculated loss function is less than the artificially set threshold. If so, the recovery can be completed. Otherwise, it cannot.
[0116] Embodiment 2: For the scenario of Embodiment 1, in order to avoid approximate loss and reduce information loss, and reduce the search complexity of the solution space, the present invention provides an "adaptive differential evolution efficient recovery method for perioperative time series data based on hash hidden feature learning" on the basis of the device of Embodiment 1.
[0117] The specific steps are as follows:
[0118] S1: The server (1) sends a data acquisition instruction to the data acquisition device (2), and the data acquisition device (2) collects data from the physiological time series data acquisition device at the set sampling frequency using the local physiological time series data receiving module (21).
[0119] Specifically: the server (1) selects the corresponding data acquisition device (2) from the data acquisition devices (2) that communicate with the server (1) via network using the patient ID as an index, and sends a data acquisition instruction; after receiving the data acquisition instruction, the data acquisition device (2) acquires the data of the physiological time series data acquisition device at the set sampling frequency using the local physiological time series data receiving module (21).
[0120] S2: The local physiological time series data storage module (21) stores the acquired data and uploads it to the server data storage module (11) of the server (1).
[0121] S3: The server (1) stores the acquired data using the server data storage module (11) and preprocesses it into incomplete physiological time series data.
[0122] The incomplete physiological time series data is an HDI matrix where the patient ID set U and the physiological detection index set I, and the |·| operation represents the number of elements in the set; each element r in R u,i represents the monitoring value of the physiological detection index i∈I monitored by the patient with ID u∈U; the number of elements |Λ| in the set Λ of all known elements in R is much smaller than the number of elements |Γ| in the set Γ of unknown elements in R. Among them, |U| = 4574 and |I| = 31.
[0123] S4: The data recovery module (12) of the server (1) initializes the discrete hidden feature vector using binary quantization.
[0124] The discrete hidden feature matrix includes a patient hidden feature matrix W and a physiological detection index hidden feature matrix Q; where: W is a |U|×D-dimensional matrix, used to extract the patient hidden features that the extraction device needs; Q is a |I|×D-dimensional matrix, used to extract the physiological detection index hidden features that the extraction device needs; D = 100 is the dimension of the latent feature space and is the bit length of the artificially set hash code.
[0125] The binary quantization initialization process is: independently and randomly select an element from the set {+1, -1} to fill the patient hidden feature matrix W and the physiological detection index hidden feature matrix Q until the filling is completed.
[0126] S5: The data recovery module (12) of the server (1) updates and learns the discrete hidden feature vector using the hash hidden feature learning method.
[0127] Combined with Table 2, the hash hidden feature learning method is specifically:
[0128] S501: Set the maximum number of update learning times \(T\), randomly initialize the hyperparameters \(\alpha\) and \(\beta\) in the interval \((0, 1)\), and initialize the current update learning times \(t = 0\); where \(T\) is a positive integer.
[0129] S502: Determine whether the current update learning times \(t\) is greater than or equal to \(T\). If it reaches, go to step S6.
[0130] S503: Update the patient's latent feature matrix \(W\) and the physiological detection index latent feature matrix \(Q\).
[0131] S504: Calculate the loss function before and after the update according to the following formula
[0132]
[0133] where \(MEAN(\cdot)\) is the mean operation; \(NOT\) is the not operation; \(AND\) is the and operation; \(w\) u is the \(u\)-th row vector of \(W\), \(q\) i is the \(i\)-th row vector of \(Q\); \(w\) u,d is the element in the \(u\)-th row and \(d\)-th column of \(W\), \(q\) i,d is the element in the \(i\)-th row and \(d\)-th column of \(Q\).
[0134] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error. Otherwise, return to step S502 to perform the next update.
[0135] Table 2 Update learning of discrete latent feature vectors
[0136]
[0137] Combined Figure 3 , the random algorithm described in step S5031 and step S5033 is an adaptive differential evolution algorithm. The specific adaptive differential evolution algorithm is as follows:
[0138] 1) Initially, set the maximum number of update generations \(G\) and the current update generation \(g = 1\), and randomly generate a binary discrete population \(X\) g \(\in{+ 1, - 1\}\) NP×D ; where \(NP\) is the number of samples, \(G\) is a positive integer between 2 and 10. In this embodiment, \(NP = 100\) and \(G = 3\).
[0139] 2) Select a \(D\)-dimensional column vector from \(X\) g as a sample and calculate its corresponding fitness, and then record the smallest sample as where \(k = 1,\cdots,NP\);
[0140] 3) Calculate the gene mutation coefficient \(F\) of the current update generation g ;
[0141]
[0142] Among them, rand1, rand2, and rand3 are random numbers between 0 and 1; at the initial moment, F g is a random number between 0 and 1.
[0143] 4) Randomly select 4 distinct samples from X g and calculate the mutant genes of the sample numbers k = 1,..., NP:
[0144]
[0145] 5) Binaryize the mutant genes of the k-th sample element by element as:
[0146]
[0147] where d = 1,..., D, k = 1,..., NP;
[0148] 6) Perform crossover genetic operations on the mutant genes of the k-th sample element by element to obtain crossover genes
[0149]
[0150] where d = 1,..., D, k = 1,..., NP, rand4 is a random number between 0 and 1, CR is a real number threshold between 0 and 1 set by humans, d rand is a random index of [1, D], is
[0151] the d-th element of;
[0151] 7) Select the next-generation samples for each sample:
[0152]
[0153] where k = 1,..., NP;
[0154] 8) Iteratively update the generation number g = g + 1 and the binary discrete population X g , until the maximum number of iterations G = 10 set by humans.
[0155] The specific algorithm flow is shown in Table 3, where the calculation of the fitness needs to determine the corresponding calculation formula according to the object of the updated hidden feature matrix.
[0156] Table 3 Adaptive Differential Evolution Algorithm
[0157]
[0158] Among them, Θ represents the proportional relationship of complexity.
[0159] S6: The data recovery module (12) of the server (1) determines whether the well-learned discrete hidden feature vectors can complete the recovery. If so, the recovered perioperative physiological time-series data is stored in the recovery data storage unit (11). Otherwise, an error message is sent through the network.
[0160] S7: The server (1) uses the sepsis outcome prediction model to perform predictive analysis on the recovered perioperative physiological time-series data in the recovery data storage unit (11) to obtain the sepsis prediction result corresponding to the patient.
[0161] Furthermore, based on the open-source perioperative physiological time-series dataset, part of the data is artificially missing, and data recovery experiments are carried out according to this embodiment, and comparative experiments are carried out with the down-sampling method in the literature [1] of the prior art in five aspects: NPV (negative predictive value), PPV (positive predictive value), accuracy, sensitivity, and specificity. The predicted sepsis results of the patients are compared with the true labels. The sepsis outcome prediction models used by the two data recovery methods are the same, and the results are shown in Table 3. It can be seen from the comparison that the method of the present invention can help improve the prediction results of the model for sepsis prediction results when recovering data, and all indicators have been improved to a certain extent, which also proves the effectiveness and accuracy of the method of the present invention.
[0162] [1]. Qinhao Wu, Fei Ye, Qianqian Gu, etc., “A customised down-sampling machine learning approach for sepsis prediction,” International Journal of Medical Informatics, vol. 184, 2024.
[0163] https: / / doi.org / 10.1016 / j.ijmedinf.2024.105365 .
[0164] Table 3 Comparison experiment results
[0165] Reference [1] The method of the present invention NPV (Negative Predictive Value) 53.4% 53.8% PPV (Positive Predictive Value) 78.5% 80.2% Accuracy 72.9% 73.3% Sensitivity 41.9% 48.1% Specificity 82.8% 83.4%
[0166] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. An efficient recovery method for perioperative time-series data based on hash hidden feature learning, characterized in that The method includes the following steps: S1: The server (1) sends a data acquisition instruction to the data acquisition device (2), and the data acquisition device (2) acquires the data of the physiological time series data acquisition device at the set sampling frequency by using the local physiological time series data receiving module (21); S2: The local physiological time series data storage module (22) stores the acquired data and uploads it to the server data storage module (11) of the server (1); S3: The server (1) uses the server data storage module (11) to store the acquired data and preprocesses it into incomplete physiological time series data; S4: The data recovery module (12) of the server (1) initializes the discrete hidden feature vector by binary quantization; S5: The data recovery module (12) of the server (1) updates and learns the discrete hidden feature vector by using the hash hidden feature learning method; S6: The data recovery module (12) of the server (1) determines whether the learned discrete hidden feature vector can complete the recovery. If it can, the recovered perioperative physiological time series data is stored in the server data storage module (11). Otherwise, an error message is sent through the network; Among them, the incomplete physiological time-series data in step S3 is the HDI matrix Among them, the patient ID set U and the physiological detection index set I, the |·| operation represents the number of elements in the set; each element r in R u,i represents the monitoring value of the physiological detection index i∈I monitored by the patient with ID u∈U; the number |Λ| of all known element sets Λ in R is much smaller than the number |Γ| of the set of unknown elements Γ in R; The discrete hidden feature matrix described in step S4 includes a patient hidden feature matrix W and a physiological detection index hidden feature matrix Q; where: W is of dimension |U|×D and is used to extract the patient hidden features that the extraction device needs to extract; Q is of dimension |I|×D and is used to extract the physiological detection index hidden features that the extraction device needs to extract; D is the dimension of the latent feature space and is the length of the artificially set hash code; The hash hidden feature learning method described in step S5 is specifically as follows: S501: Set the maximum number of update learning times T, randomly initialize the hyperparameters α and β in the interval (0, 1), and initialize the current update learning times t = 0; where T is a positive integer; S502: Determine whether the current update learning times t is greater than or equal to T. If it reaches, go to step S6; S503: Update the patient hidden feature matrix W and the physiological detection index hidden feature matrix Q; S504: Calculate the loss function before and after update according to the following formula Among them, MEAN(·) is the mean operation; NOT is the NOT operation; AND is the AND operation; w u is the u-th row vector of W, q i is the i-th row vector of Q; w u,d is the element at the d-th column of the u-th row of W, q i,d is the element at the d-th column of the i-th row of Q; S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error. Otherwise, return to step S502 to perform the next update; The restored perioperative physiological time series data described in step S6 wherein it is necessary to reverse the standardization and restore it to the original dimension.
2. The efficient recovery method for perioperative time-series data based on hash hidden feature learning according to claim 1, wherein The specific content of step S1 is: The server (1) selects the corresponding data acquisition device (2) from the data acquisition devices (2) that communicate with the server (1) network by using the patient ID as an index and sends a data acquisition instruction; After receiving the data acquisition instruction, the data acquisition device (2) acquires the data of the physiological time series data acquisition device at the set sampling frequency by using the local physiological time series data receiving module (21).
3. The efficient restoration method for perioperative time-series data based on hash hidden feature learning according to claim 1, characterized in that, The preprocessing process described in step S3 includes: S301: Store the uploaded acquired data as an HDI matrix R; S302: Mark the unknown elements and mark the data with missing sampling data with NAN; S303: Normalize R according to the known elements.
4. The efficient recovery method for perioperative time-series data based on hash hidden feature learning according to claim 1, wherein The binary quantization initialization process described in step S4 is: Independently and randomly select an element from the set {+1, -1} to fill the patient hidden feature matrix W and the physiological detection index hidden feature matrix Q until the filling is completed.
5. The efficient restoration method for perioperative time-series data based on hash hidden feature learning according to claim 1, characterized in that The specific steps of step S503 are as follows: S5031: Fix the latent feature matrix Q of physiological detection indexes, generate multiple samples of the latent feature matrix W of patients row by row for the latent feature matrix W of physiological detection indexes using a random algorithm, and calculate the fitness corresponding to each sample according to the following formula S5032: Select the fitness of each row Select the sample with the smallest value, and update the corresponding row vector of the patient's latent feature matrix W; S5033: Fix the implicit feature matrix W of physiological detection indicators, generate multiple samples of the implicit feature matrix Q of physiological detection indicators row by row using a random algorithm for the implicit feature matrix Q of physiological detection indicators, and calculate the adaptation value corresponding to each sample according to the following formula S5034: Select the fitness of each row Select the sample with the smallest value, and update the corresponding row vector of the physiological detection index hidden feature matrix Q.
6. The method for efficiently recovering perioperative time-series data based on hash hidden feature learning according to claim 5, wherein The random algorithms described in steps S5031 and S5033 are particle swarm algorithms.
7. The method for efficient recovery of perioperative time-series data based on hash hidden feature learning according to claim 5, characterized in that The random algorithms described in steps S5031 and S5033 are adaptive differential evolution algorithms. The specific adaptive differential evolution algorithm is as follows: 1) Initially, set the maximum number of update generations G and the current update generation g = 1, and randomly generate a binary discrete population X g ∈ {+1, -1} NP×D ; where NP is the number of samples; 2) Select a D-dimensional column vector from X g as a sample and calculate its corresponding fitness, then record the sample with the smallest fitness as where k = 1, …, NP; 3) Calculate the gene mutation coefficient F of the current updated generation g ; where rand1, rand2, and rand3 are random numbers between 0 and 1; at the initial moment, F g is a random number between 0 and 1; 4) Select 4 mutually distinct D-dimensional column vector samples randomly from X g Calculate the mutant genes for sample numbers k = 1, …, NP: 5) Binaryize the mutated genes of the k-th sample element by element: Among them, 6) The mutated genes of the k-th sample will be subjected to cross-genetic operations element by element to obtain cross genes where rand4 is a random number between 0 and 1, CR is a real threshold value between 0 and 1 set by humans, and d rand is a random index in [1, D], and is the d-th element of 7) Select the next generation of samples sample by sample: where k = 1,..., NP; 8) Iteratively update the generation number g = g + 1 and the binary discrete population X g , until the maximum number of iterations G set manually.
8. The method for efficiently recovering perioperative time-series data based on hash hidden feature learning according to claim 1, wherein The criterion for determining whether recovery can be completed described in step S6 is the calculated loss function is less than the threshold set by humans. If so, recovery can be completed; otherwise, it cannot be completed.
9. Applied to the perioperative time-series data efficient recovery device based on hash hidden feature learning according to any one of claims 1 to 8, characterized in that It consists of a server (1) and N data acquisition devices (2), and they communicate with each other through a network for data communication; The server (1) includes: a server data storage module (11) and a data recovery module (12); one of the data acquisition devices (2) is a local computer, which includes: a local physiological time series data receiving module (21) and a local physiological time series data storage module (22); where the value of N depends on the number of physiological time series data acquisition devices; The local physiological time series data receiving module (21) in the data acquisition device (2) is connected to the physiological time series data acquisition device through a network, and is used to receive the incomplete physiological time series data fed back by the physiological time series data acquisition device, and instruct the local physiological time series data storage module (22) to store the incomplete physiological time series data; The local physiological time series data storage module (22) in the data acquisition device (2) is connected to the server data storage module (11) in the server (1) through a network, and sends the incomplete physiological time series data to the server data storage module (11); The server data storage module (11) and the data recovery module (12) are used to send the incomplete physiological time series data of all data acquisition devices (2) to the data recovery module (12) for data recovery processing, and then save the recovered perioperative physiological time series data; The data recovery module (12) uses the hash hidden feature learning method to recover the incomplete physiological time series data.
10. The perioperative time series data efficient recovery device based on hash hidden feature learning according to claim 9, characterized in that, The server data storage module (11) includes a cloud physiological time series data storage unit (111), a data preprocessing unit (112), and a recovered data storage unit (113); The cloud physiological time series data storage unit (111) in the server data storage module (11) is used to receive and store the incomplete physiological time series data sent by the local physiological time series data storage module (22) in the data acquisition device (2); The data preprocessing unit (112) is connected to the cloud physiological time series data storage unit (111) and the data recovery module (12), and is used to preprocess the incomplete physiological time series data of all data acquisition devices (2); The recovered data storage unit (113) is connected to the data recovery module (12), and is used to store the recovered perioperative physiological time series data.
Citation Information
Patent Citations
Cross-modal hash retrieval method based on semantic graph evolution
CN116680432A
Prediction method for single physiological index state in perioperative period
CN117219272A
Medical decision-oriented multi-modal data dynamic fusion and labeling method and system
CN119377894A
System, server and method for preventing suicide cross-reference to related applications
US20230138557A1