Hash hidden feature learning-based perioperative time-series data efficient recovery method and device

By employing hash latent feature learning and adaptive differential evolution algorithm, the problems of high dimensionality, high noise, and missing data in perioperative physiological time series data were solved, achieving efficient and accurate data recovery and improving the early warning model.

CN120376020BActive Publication Date: 2026-05-01SOUTHWEST UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHWEST UNIV
Filing Date
2025-04-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing tensor-based latent feature analysis methods struggle to efficiently process perioperative physiological time-series data that are high-dimensional, noisy, and lacking in data, especially in terms of computational cost and real-time requirements.

Method used

By employing hash latent feature learning combined with an end-to-end framework, data is represented through hash operations, and direct discretization optimization is performed using an adaptive differential evolution algorithm, reducing data storage and communication overhead and achieving efficient and accurate recovery of physiological time series data.

Benefits of technology

It effectively reduces data storage and communication overhead, can efficiently and accurately recover perioperative physiological time series data, improves the robustness and accuracy of the early warning model, and reduces computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376020B_ABST
    Figure CN120376020B_ABST
Patent Text Reader

Abstract

The application is a perioperative time series data efficient recovery method and device based on hash hidden feature learning, belonging to the field of medical health big data. The device contains a server and N data acquisition devices; the method contains the following steps: S1: collecting physiological time series data; S2: uploading to the server; S3: preprocessing into incomplete physiological time series data; S4: initializing a discrete hidden feature vector; S5: updating and learning the discrete hidden feature vector by using a hash hidden feature learning method; S6: judging whether the discrete hidden feature vector can complete recovery, if yes, storing the result, otherwise, sending an error information through the network. The method introduces hash operation to encode the data in binary, reduces the data storage and communication overhead, and combines the end-to-end framework of the LF model and the self-adaptive differential evolution algorithm, which can avoid approximate loss and reduce information loss, and realize efficient and accurate recovery of perioperative physiological time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Efficient Method and Apparatus for Perioperative Temporal Data Recovery Based on Hash Hidden Feature Learning Technical Field

[0001] This invention relates to a method and apparatus for efficient recovery of perioperative time-series data based on hash hidden feature learning, belonging to the fields of medical and health big data and data recovery, and is particularly suitable for efficient recovery of perioperative time-series data based on hash hidden feature learning. Background Technology

[0002] The incompleteness of perioperative physiological time-series data is mainly caused by the limitations of dynamic acquisition of multi-source heterogeneous data and technical barriers. On the one hand, the sampling frequencies of different physiological indicators (such as ECG, blood pressure, and blood oxygen) vary significantly, making it difficult to synchronously integrate high-frequency monitoring data with low-frequency recording data. On the other hand, insufficient compatibility and heterogeneous data formats exist when multiple devices are used for collaborative monitoring, leading to fragmentation of multi-source information. In addition, factors such as equipment operation interference, measurement technology errors, data anonymization due to data privacy protection, and human error further cause high-frequency time-series errors, mixed noise interference, and missing specific fields in the data. These factors make perioperative data exhibit high dimensionality, high noise, and temporal discontinuities. At the same time, due to its dynamic change characteristics, the data is discontinuous and heterogeneous in both time and space dimensions, forming a typical high-dimensional and incomplete (HDI) data structure, which greatly increases the difficulty of analysis.

[0003] The contradiction between the high mortality rate of perioperative cardiovascular adverse events and the lagging early warning technology highlights the urgency of tapping the potential value of incomplete data. Traditional methods struggle to extract effective information from highly noisy and missing physiological time-series data, while artificial intelligence technologies (such as machine learning) can discover hidden risk signals through dynamic pattern recognition, for example, predicting postoperative cardiovascular events based on intraoperative monitoring data. The introduction of methods such as tensor latent feature analysis (LF), by modeling high-dimensional incomplete data as low-rank tensors, can effectively extract the temporal correlations and potential pathological features of multiple physiological indicators, improving the robustness and accuracy of early warning models.

[0004] While tensor decomposition-based low-dimensional embedding methods offer a novel approach to processing HDI data, their practical application still faces multiple challenges. The exponential growth of medical data has led to a surge in the computational cost of traditional continuous embedding spaces, while the real-time requirements of perioperative data further demand higher algorithm efficiency. Overcoming these bottlenecks requires innovating lightweight tensor computation frameworks and integrating domain knowledge to optimize feature extraction logic.

[0005] In summary, existing methods based on tensor latent feature analysis still face significant challenges in processing high-dimensional, noisy, and missing physiological time-series data during the perioperative treatment process. Summary of the Invention

[0006] In view of this, and in response to the problems of the prior art, the present invention provides a method and apparatus for efficient recovery of perioperative time series data based on hash latent feature learning. The aim is to reduce data storage and communication overhead by introducing hash operations to represent the data, and then to combine the hash operations with the LF model in an end-to-end framework, and to use an adaptive differential evolution algorithm for direct discretization optimization, so as to avoid approximation loss and reduce information loss, and finally achieve efficient and accurate recovery of perioperative physiological time series data.

[0007] To achieve the above objectives, a highly efficient perioperative time-series data recovery device based on hash hidden feature learning is provided, comprising a server and N data acquisition devices, which communicate with each other via a network. The server includes a server data storage module and a data recovery module. Each data acquisition device is a local computer, comprising a local physiological time-series data receiving module and a local physiological time-series data storage module. The value of N depends on the number of physiological time-series data acquisition devices.

[0008] The local physiological time series data receiving module in the data acquisition device is connected to the physiological time series data acquisition device via a network. It is used to receive incomplete physiological time series data fed back by the physiological time series data acquisition device and instruct the local physiological time series data storage module to store the incomplete physiological time series data.

[0009] The local physiological time series data storage module in the data acquisition device is connected to the server data storage module in the server via a network, and sends incomplete physiological time series data to the server data storage module.

[0010] The server data storage module and data recovery module are used to send the incomplete physiological time series data from all data acquisition devices to the data recovery module for data recovery processing, and then save the recovered perioperative physiological time series data.

[0011] The data recovery module uses the hash latent feature learning method to recover incomplete physiological time series data.

[0012] Preferredly, the server data storage module includes a cloud-based physiological time-series data storage unit, a data preprocessing unit, and a recovery data storage unit;

[0013] The cloud-based physiological time-series data storage unit in the server data storage module is used to receive and store incomplete physiological time-series data sent by the local physiological time-series data storage module in the data acquisition device.

[0014] The data preprocessing unit is connected to the cloud-based physiological time series data storage unit and the data recovery module, and is used to preprocess the incomplete physiological time series data from all data acquisition devices.

[0015] The recovery data storage unit is connected to the data recovery module and is used to store perioperative physiological time series data after recovery.

[0016] This invention also provides an efficient method for recovering perioperative time-series data based on hash latent feature learning, comprising the following steps:

[0017] S1: The server sends a data acquisition command to the data acquisition device, and the data acquisition device uses the local physiological time series data receiving module to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0018] S2: The local physiological time series data storage module stores the collected data and uploads it to the server data storage module.

[0019] S3: The server uses the server data storage module to store the collected data and preprocesses it into incomplete physiological time series data;

[0020] S4: The server's data recovery module initializes the discrete latent feature vector using binary quantization;

[0021] S5: The server's data recovery module uses the hash latent feature learning method to update and learn the discrete latent feature vector;

[0022] S6: The server's data recovery module determines whether the learned discrete latent feature vector can be recovered. If it can, the recovered perioperative physiological time series data is stored in the recovery data storage unit of the server's data storage module. Otherwise, an error message is sent over the network.

[0023] Furthermore, step S1 specifically involves: the server selecting the corresponding data acquisition device from the data acquisition devices communicating with the server network using the patient ID as an index, and sending a data acquisition command; after receiving the data acquisition command, the data acquisition device uses the local physiological time series data receiving module to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0024] The incomplete physiological time series data in step S3 is the HDI matrix. Wherein, the patient ID set U and the physiological test index set I are, and the |·| operation represents the number of elements in the set; each element r in R u,i This indicates that the patient ID is u∈U and the monitored physiological test index is i∈I; the number of known elements in R, Λ|Λ|, is much smaller than the number of unknown elements in R, Γ|Γ|.

[0025] Furthermore, the preprocessing procedure described in step S3 includes:

[0026] S301: Store the uploaded collected data as an HDI matrix R;

[0027] S302: Label unknown elements by labeling missing data in the sampled data using NAN;

[0028] S303: Normalize R according to the known elements.

[0029] The discrete latent feature matrix mentioned in step S4 includes a patient latent feature matrix W and a physiological detection index latent feature matrix Q; where: W is |U|×D dimensional and is used to extract the patient latent features that the extraction device needs to extract; Q is |I|×D dimensional and is used to extract the physiological detection index latent features that the extraction device needs to extract; D is the dimension of the latent feature space and is the bit length of the hash code set by the user.

[0030] Furthermore, the binary quantization initialization process described in step S4 is as follows: independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological test index latent feature matrix Q until the filling is completed.

[0031] Furthermore, the hash latent feature learning method described in step S5 is specifically as follows:

[0032] S501: Set the maximum number of update learning times T, randomly initialize hyperparameters α and β in the interval (0, 1), and initialize the current number of update learning times t = 0; where T is a positive integer;

[0033] S502: Determine whether the current update learning count t is greater than or equal to T. If it is, proceed to step S6.

[0034] S503: Update the patient latent feature matrix W and the physiological test indicator latent feature matrix Q;

[0035] S504: Calculate the loss function before and after the update according to the following formula.

[0036]

[0037] Where MEAN(·) is the mean operation; NOT is the NOT operation; AND is the AND operation; w u Let q be the u-th row vector of W. i Let w be the i-th row vector of Q; u,d Let q be the element in row u and column d of W. i,d Let Q be the element in the i-th row and d-th column;

[0038] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error; otherwise, return to step S502 to execute the next update.

[0039] Furthermore, step S503 specifically includes:

[0040] S5031: Fix the latent feature matrix Q of the physiological test indicators, generate multiple samples of the latent feature matrix W of patients row by row using a random algorithm, and calculate the fitness of each sample according to the following formula.

[0041]

[0042] S5032: Select the fitness of each row The smallest sample is used to update the row vector corresponding to the patient's latent feature matrix W;

[0043] S5033: Fix the latent feature matrix W of the physiological test indicators, generate multiple samples of the latent feature matrix Q of the physiological test indicators row by row using a random algorithm, and calculate the fitness value corresponding to each sample according to the following formula.

[0044]

[0045] S5034: Select the fitness of each row The smallest sample is used to update the row vector corresponding to the latent feature matrix Q of the physiological detection index.

[0046] Furthermore, the random algorithm described in steps S5031 and S5033 is the particle swarm optimization algorithm.

[0047] Preferredly, the random algorithm described in steps S5031 and S5033 is an adaptive differential evolution algorithm, specifically as follows:

[0048] 1) Initially, set the maximum update generation G and the current update generation g = 1, and randomly generate the binary discrete population X. g ∈{+1, -1} NP×D Where NP is the number of samples;

[0049] 2) From X g Select D-dimensional column vectors as samples And calculate the corresponding fitness, then denote the smallest sample as... Where k = 1, ..., NP;

[0050] 3) Calculate the gene mutation coefficient F for the current generation. g ;

[0051]

[0052] Where rand1, rand2, and rand3 are random numbers between 0 and 1; the initial time F g A random number between 0 and 1.

[0053] 4) From X g Four distinct D-dimensional column vector samples are randomly selected from the data. Calculate the mutated genes for sample numbers k = 1, ..., NP:

[0054]

[0055] 5) Binary the mutated gene of the k-th sample element by element:

[0056]

[0057] in,

[0058] 6) Analyze the mutated genes of the k-th sample element by element. Cross-genes are obtained by performing cross-genetic manipulation.

[0059]

[0060] Where rand4 is a random number between 0 and 1, CR is a manually set real number threshold between 0 and 1, and d rand For random indices in [1, D], for The d-th element;

[0061] 7) Selecting the next generation of samples one by one:

[0062]

[0063] Where k = 1,...,NP;

[0064] 8) Iteratively update the algebra g = g + 1 and the binary discrete population X. g This continues until the maximum number of iterations, G, is set manually.

[0065] Furthermore, the perioperative physiological time-series data after recovery described in step S6 in It needs to be denormalized to restore the original dimensions.

[0066] Preferably, the criterion for determining whether recovery can be completed in step S6 is the calculated loss function. If the value is less than a manually set threshold, recovery can be completed; otherwise, it cannot.

[0067] The advantages of this invention are as follows: This invention provides a method and apparatus for efficient recovery of perioperative time-series data based on hash latent feature learning. By introducing hash operations to encode the data in binary form, data storage and communication overhead are reduced. Combined with the end-to-end framework of the LF model, it can effectively handle mixed noise and high proportion of missing physiological time-series data, achieving the recovery of perioperative physiological time-series data. Furthermore, an adaptive differential evolution algorithm is used for direct discrete optimization to avoid approximation loss and reduce information loss, ultimately achieving efficient and accurate recovery of perioperative physiological time-series data. Attached Figure Description

[0068] To make the objectives and technical solutions of this invention clearer, the following figures are provided for illustration:

[0069] Figure 1 is a flowchart of the efficient perioperative time-series data recovery method based on hash implicit feature learning according to the present invention.

[0070] Figure 2 is a framework diagram of the efficient perioperative time series data recovery device based on hash latent feature learning in Embodiment 1 of the present invention;

[0071] Figure 3 is a schematic diagram of discrete latent feature vector update learning in Embodiment 2 of the present invention. Detailed Implementation

[0072] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0073] Example 1: A hospital needs to perform real-time analysis of the physiological time-series data of 4574 perioperative patients with sepsis (analyzed every 15 minutes). The hospital uses physiological time-series data acquisition equipment to collect time-series data on 31 characteristics, including: heart rate (HR, beats per minute), oxygen saturation (O2Sat, %), body temperature (Temp, ℃), systolic blood pressure (SBP, mm Hg), mean arterial pressure (MAP, mm Hg), diastolic blood pressure (DBP, mm Hg), respiratory rate (Resp, breaths per minute), end-expiratory carbon dioxide partial pressure (EtCO2, mm Hg), base excess (BaseExcess, mmol / L), bicarbonate (HCO3, mmol / L), inhaled oxygen concentration (FiO2, %), blood pH, and arterial blood carbon dioxide partial pressure (PaCO2, mm Hg). Hg, arterial oxygen saturation (SaO2, %), aspartate aminotransferase (AST, U / L), blood urea nitrogen (BUN, mg / dL), alkaline phosphatase (Alkalinephos, IU / L), serum calcium (Calcium, mg / dL), serum chloride (Chloride, mmol / L), creatinine (Creatinine, mg / dL), direct bilirubin (Bilirubin direct, mg / dL), glucose (Glucose, mg / dL), lactate (Lactate, mg / dL), serum magnesium (Magnesium, mmol / dL), serum phosphorus (Phosphate, mg / dL), serum potassium (Potassium, mmol / L), total bilirubin (Bilirubin total, mg / dL), troponin I (Troponin I, ng / mL), hematocrit (Hct, %), hemoglobin (Hgb, g / dL), partial thromboplastin time (PTT, seconds), white blood cell count (WBC, count × 10⁻⁶). 3 / μL), fibrinogen (mg / dL) and platelet count (×10 3 / μL). Due to insufficient compatibility and noise issues in multi-device collaborative monitoring, known elements account for less than 5% of the total physiological time-series data. In order to utilize deep learning models to extract features from physiological time-series data and analyze the sepsis status of patients, it is urgent to recover unknown data. To this end, the present invention provides a "highly efficient recovery device for perioperative time-series data based on hash hidden feature learning", as shown in Figure 2.

[0074] It consists of a server (1) and N data acquisition devices (2), which communicate with each other through a network connection; the server (1) includes: a server data storage module (11) and a data recovery module (12); each of the data acquisition devices (2) is a local computer, which includes: a local physiological time series data receiving module (21) and a local physiological time series data storage module (22); where the number of N depends on the number of physiological time series data acquisition devices;

[0075] The local physiological time series data receiving module (21) in the data acquisition device (2) is connected to the physiological time series data acquisition device via a network. It is used to receive incomplete physiological time series data fed back by the physiological time series data acquisition device and instruct the local physiological time series data storage module (22) to store the incomplete physiological time series data.

[0076] The local physiological time series data storage module (22) in the data acquisition device (2) is connected to the server data storage module (11) in the server (1) via the network, and sends incomplete physiological time series data to the server data storage module (11);

[0077] The server data storage module (11) and data recovery module (12) are used to send the incomplete physiological time series data of all data acquisition devices (2) to the data recovery module (12) for data recovery processing, and then save the recovered perioperative physiological time series data.

[0078] The data recovery module (12) uses the hash latent feature learning method to recover incomplete physiological time series data.

[0079] Preferredly, the server data storage module (11) includes a cloud-based physiological time-series data storage unit (111), a data preprocessing unit (112), and a recovery data storage unit (113);

[0080] The cloud-based physiological time series data storage unit (111) in the server data storage module (11) is used to receive and store incomplete physiological time series data sent by the local physiological time series data storage module (22) in the data acquisition device (2).

[0081] The data preprocessing unit (112) is connected to the cloud physiological time series data storage unit (111) and the data recovery module (12) and is used to preprocess the incomplete physiological time series data of all data acquisition devices (2);

[0082] The recovery data storage unit (113) is connected to the data recovery module (12) and is used to store the perioperative physiological time series data after recovery.

[0083] This invention also provides an "efficient method for recovering perioperative time-series data based on hash latent feature learning," which, as shown in Figure 1, includes the following steps:

[0084] S1: The server (1) sends a data acquisition instruction to the data acquisition device (2), and the data acquisition device (2) uses the local physiological time series data receiving module (21) to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0085] Specifically: The server (1) selects the corresponding data acquisition device (2) from the data acquisition devices (2) that communicate with the server (1) via the network using the patient ID as an index, and sends a data acquisition instruction; After receiving the data acquisition instruction, the data acquisition device (2) uses the local physiological time series data receiving module (21) to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0086] S2: The local physiological time series data storage module (21) stores the collected data and uploads it to the server data storage module (11) of the server (1).

[0087] S3: The server (1) uses the server data storage module (11) to store the collected data and uses the data preprocessing unit (112) to preprocess it into incomplete physiological time series data.

[0088] The incomplete physiological time-series data mentioned above is the HDI matrix. Wherein, the patient ID set U and the physiological test index set I are, and the |·| operation represents the number of elements in the set; each element r in R u,i This represents the monitoring values ​​of physiological indicators for patients with ID u∈U and monitoring values ​​for indicators i∈I; the number of known elements in R, denoted as |Λ|, is much smaller than the number of unknown elements in R, denoted as |Γ|. Where |U|=4574, |I|=31.

[0089] The preprocessing process includes:

[0090] S301: Store the uploaded collected data as an HDI matrix R;

[0091] S302: Label unknown elements by labeling missing data in the sampled data using NAN;

[0092] S303: Normalize R according to the known elements.

[0093] S4: The data recovery module (12) of the server (1) initializes the discrete latent feature vector using binary quantization.

[0094] The discrete latent feature matrix includes a patient latent feature matrix W and a physiological detection index latent feature matrix Q; where: W is |U|×D dimensional and is used to extract the patient latent features that the extraction device needs to extract; Q is |I|×D dimensional and is used to extract the physiological detection index latent features that the extraction device needs to extract; D=100 is the dimension of the latent feature space and is the bit length of the hash code set by the user.

[0095] The binary quantization initialization process is as follows: independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological test index latent feature matrix Q until the filling is complete.

[0096] S5: The data recovery module (12) of the server (1) uses the hash latent feature learning method to update and learn the discrete latent feature vector.

[0097] Referring to Table 1, the hash latent feature learning method is specifically as follows:

[0098] S501: Set the maximum number of update learning times T, randomly initialize hyperparameters α and β in the interval (0, 1), and initialize the current number of update learning times t = 0; where T is a positive integer, and in this embodiment, T = 1000.

[0099] S502: Determine whether the current number of learning updates t is greater than or equal to T. If it is, proceed to step S6.

[0100] S503: Update the patient latent feature matrix W and the physiological test indicator latent feature matrix Q. Specifically:

[0101] S5031: Fix the latent feature matrix Q of the physiological test indicators, and use the particle swarm optimization algorithm to generate multiple samples of the latent feature matrix W of patients row by row, and calculate the fitness of each sample according to the following formula.

[0102]

[0103] S5032: Select the fitness of each row The smallest sample is used to update the row vector corresponding to the patient's latent feature matrix W;

[0104] S5033: Fix the latent feature matrix W of the physiological detection index, generate multiple samples of the latent feature matrix Q of the physiological detection index row by row using the particle swarm optimization algorithm, and calculate the fitness value corresponding to each sample according to the following formula.

[0105]

[0106] S5034: Select the fitness of each row The smallest sample is used to update the row vector corresponding to the latent feature matrix Q of the physiological detection index.

[0107] S504: Calculate the loss function before and after the update according to the following formula.

[0108]

[0109] Where MEAN(·) is the mean operation; NOT is the NOT operation; AND is the AND operation; w u Let q be the u-th row vector of W. i Let w be the i-th row vector of Q; u,d Let q be the element in row u and column d of W. i,d Let Q be the element in the i-th row and d-th column.

[0110] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error; otherwise, return to step S502 to execute the next update.

[0111] Table 1. Learning to update discrete latent feature vectors

[0112]

[0113] S6: The data recovery module (12) of the server (1) determines whether the learned discrete latent feature vector can be recovered. If it can, the recovered perioperative physiological time series data is stored in the recovery data storage unit (113). Otherwise, an error message is sent through the network.

[0114] The perioperative physiological time series data after recovery in It needs to be denormalized to restore the original dimensions.

[0115] The criterion for determining whether recovery can be completed is the calculated loss function. If the value is less than a manually set threshold, recovery can be completed; otherwise, it cannot.

[0116] Example 2: In order to avoid approximation loss and reduce information loss, and reduce the search complexity of the solution space, this invention provides an "efficient recovery method for perioperative time series data based on hash latent feature learning" to address the scenario in Example 1.

[0117] The specific steps are as follows:

[0118] S1: The server (1) sends a data acquisition instruction to the data acquisition device (2), and the data acquisition device (2) uses the local physiological time series data receiving module (21) to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0119] Specifically: The server (1) selects the corresponding data acquisition device (2) from the data acquisition devices (2) that communicate with the server (1) via the network using the patient ID as an index, and sends a data acquisition instruction; After receiving the data acquisition instruction, the data acquisition device (2) uses the local physiological time series data receiving module (21) to collect data from the physiological time series data acquisition device according to the set sampling frequency.

[0120] S2: The local physiological time series data storage module (21) stores the collected data and uploads it to the server data storage module (11) of the server (1).

[0121] S3: The server (1) uses the server data storage module (11) to store the collected data and preprocess it into incomplete physiological time series data.

[0122] The incomplete physiological time-series data mentioned above is the HDI matrix. Wherein, the patient ID set U and the physiological test index set I are, and the |·| operation represents the number of elements in the set; each element r in R u,i This represents the monitoring values ​​of physiological indicators for patients with ID u∈U and monitoring values ​​for indicators i∈I; the number of known elements in R, denoted as |Λ|, is much smaller than the number of unknown elements in R, denoted as |Γ|. Where |U|=4574, |I|=31.

[0123] S4: The data recovery module (12) of the server (1) initializes the discrete latent feature vector using binary quantization.

[0124] The discrete latent feature matrix includes a patient latent feature matrix W and a physiological detection index latent feature matrix Q; where: W is |U|×D dimensional and is used to extract the patient latent features that the extraction device needs to extract; Q is |I|×D dimensional and is used to extract the physiological detection index latent features that the extraction device needs to extract; D=100 is the dimension of the latent feature space and is the bit length of the hash code set by the user.

[0125] The binary quantization initialization process is as follows: independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological test index latent feature matrix Q until the filling is complete.

[0126] S5: The data recovery module (12) of the server (1) uses the hash latent feature learning method to update and learn the discrete latent feature vector.

[0127] Referring to Table 2, the hash latent feature learning method is specifically as follows:

[0128] S501: Set the maximum number of update learning times T, randomly initialize hyperparameters α and β in the interval (0, 1), and initialize the current number of update learning times t = 0; where T is a positive integer;

[0129] S502: Determine whether the current update learning count t is greater than or equal to T. If it is, proceed to step S6.

[0130] S503: Update the patient latent feature matrix W and the physiological test indicator latent feature matrix Q;

[0131] S504: Calculate the loss function before and after the update according to the following formula.

[0132]

[0133] Where MEAN(·) is the mean operation; NOT is the NOT operation; AND is the AND operation; w u Let q be the u-th row vector of W. i Let w be the i-th row vector of Q; u,d Let q be the element in row u and column d of W. i,d Let Q be the element in the i-th row and d-th column;

[0134] S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error; otherwise, return to step S502 to execute the next update.

[0135] Table 2. Learning to update discrete latent feature vectors

[0136]

[0137] Referring to Figure 3, the random algorithm described in steps S5031 and S5033 is an adaptive differential evolution algorithm, which specifically includes:

[0138] 1) Initially, set the maximum update generation G and the current update generation g = 1, and randomly generate the binary discrete population X. g ∈{+1, -1} NP×D Where NP is the number of samples, and G is a positive integer between 2 and 10. In this embodiment, NP = 100 and G = 3.

[0139] 2) From X g Select D-dimensional column vectors as samples And calculate the corresponding fitness, then denote the smallest sample as... Where k = 1, ..., NP;

[0140] 3) Calculate the gene mutation coefficient F for the current generation. g ;

[0141]

[0142] Where rand1, rand2, and rand3 are random numbers between 0 and 1; the initial time F g A random number between 0 and 1.

[0143] 4) From X g Four distinct samples were randomly selected from the data. Calculate the mutated genes for sample numbers k = 1, ..., NP:

[0144]

[0145] 5) Binary the mutated gene of the k-th sample element by element:

[0146]

[0147] Where d = 1,...,D, k = 1,...,NP;

[0148] 6) Analyze the mutated genes of the k-th sample element by element. Cross-genes are obtained by performing cross-genetic manipulation.

[0149]

[0150] Where d = 1, ..., D, k = 1, ..., NP, rand4 is a random number between 0 and 1, CR is a manually set real number threshold between 0 and 1, and d rand For random indices in [1, D], for The d-th element;

[0151] 7) Selecting the next generation of samples one by one:

[0152]

[0153] Where k = 1,...,NP;

[0154] 8) Iteratively update the algebra g = g + 1 and the binary discrete population X. gThis continues until the maximum number of iterations, G = 10, is set manually.

[0155] The specific algorithm flow is shown in Table 3. The calculation of fitness depends on the object of the updated latent feature matrix.

[0156] Table 3 Adaptive Differential Evolution Algorithm

[0157]

[0158] Here, Θ represents the direct proportionality of complexity.

[0159] S6: The data recovery module (12) of the server (1) determines whether the learned discrete latent feature vector can be recovered. If it can, the recovered perioperative physiological time series data is stored in the recovery data storage unit (11). Otherwise, an error message is sent through the network.

[0160] S7: The server (1) uses the sepsis outcome prediction model to predict and analyze the perioperative physiological time series data after recovery in the recovery data storage unit (11) to obtain the sepsis prediction result corresponding to the patient.

[0161] Furthermore, based on the open-source perioperative physiological time series dataset, some data was artificially missing, and a data recovery experiment was conducted according to this embodiment. The data was compared with the downsampling method in the prior art [1] in five aspects: NPV negative predictive value, PPV positive predictive value, accuracy, sensitivity, and specificity. The predicted sepsis results of patients were compared with the true labels. The sepsis result prediction models used by the two data recovery methods are the same. The results are shown in Table 3. As can be seen from the comparison, the method of the present invention can help improve the prediction results of the sepsis by recovering data. All indicators have been improved to a certain extent, which also proves the effectiveness and accuracy of the method of the present invention.

[0162] [1].Qinhao Wu, Fei Ye, Qianqian Gu, etc., "A customized down-samplingmachine learning approach for sepsis prediction," International Journal of Medical Informatics, vol.184, 2024.

[0163] https: / / doi.org / 10.1016 / j.ijmedinf.2024.105365 .

[0164] Table 3 Comparison of experimental results

[0165] Reference [1] The method of this invention has the following NPV negative predictive value: 53.4% ​​53.8%; PPV positive predictive value: 78.5% 80.2%; accuracy: 72.9% 73.3%; sensitivity: 41.9% 48.1%; specificity: 82.8% 83.4%. surface

[0166] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. An efficient method for recovering perioperative time-series data based on hash latent feature learning, characterized in that, The method includes the following steps: S1: The server (1) sends a data acquisition instruction to the data acquisition device (2), and the data acquisition device (2) acquires the data of the physiological time series data acquisition device using the local physiological time series data receiving module (21) according to the set sampling frequency; S2: The local physiological time series data storage module (22) stores the acquired data and uploads it to the server data storage module (11) of the server (1); S3: The server (1) stores the acquired data using the server data storage module (11) and preprocesses it into incomplete physiological time series data; S4: The data recovery module (12) of the server (1) initializes the discrete latent feature matrix using binary quantization; S5: The data recovery module (12) of the server (1) uses the hash latent feature learning method to update and learn the discrete latent feature matrix; S6: The data recovery module (12) of the server (1) determines whether the learned discrete latent feature matrix can be recovered. If it can, the recovered perioperative physiological time series data is stored in the server data storage module (11). Otherwise, an error message is sent through the network. Among them, the incomplete physiological time series data in step S3 is the HDI matrix. Among them, the patient ID set U and the physiological test index set I, The operation represents the number of elements in a set; each element in R... This represents the monitoring value of the patient ID u∈U and the monitoring physiological detection index i∈I; the number of all known elements in R, Λ|, is much smaller than the number of unknown elements in R, Γ|; the discrete latent feature matrix in step S4 includes the patient latent feature matrix W and the physiological detection index latent feature matrix Q; where: W is |U|×D dimensional, used to extract the patient latent features that the extraction device needs to extract; Q is |I|×D dimensional, used to extract the physiological detection index latent features that the extraction device needs to extract; D is the dimension of the latent feature space, which is the bit length of the hash code set by the user; the hash latent feature learning method in step S5 is specifically as follows: S501: Set the maximum number of update learning times T, randomly initialize the hyperparameters α and β in the interval (0,1), and initialize the current number of update learning times t= 0; where T is a positive integer; S502: Determine whether the current update learning count t is greater than or equal to T. If it is, proceed to step S6; S503: Update the patient latent feature matrix W and the physiological test index latent feature matrix Q; S504: Calculate the loss function before and after the update according to the following formula. : ;in, To perform the mean operation; This is a NOT operation; For operation; Let u be the row vector of W. Let Q be the i-th row vector; Let W be the element in the u-th row and d-th column. The element in the i-th row and d-th column of Q; S505: Determine whether the loss function before and after the update converges. If it does not converge, report an error; otherwise, return to step S502 to perform the next update; the perioperative physiological time series data after recovery described in step S6. ,in It needs to be denormalized to restore the original dimensions.

2. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 1, characterized in that, The specific steps of step S1 are as follows: the server (1) selects the corresponding data acquisition device (2) from the data acquisition devices (2) that communicate with the server (1) via the network using the patient ID as the index, and sends a data acquisition command; After receiving the data acquisition instruction, the data acquisition device (2) uses the local physiological time series data receiving module (21) to collect data from the physiological time series data acquisition device according to the set sampling frequency.

3. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 1, characterized in that, The preprocessing process described in step S3 includes: S301: storing the uploaded collected data as an HDI matrix R; S302: labeling unknown elements and labeling missing data in the sampled data using NAN; S303: normalizing R according to the known elements.

4. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 1, characterized in that, The binary quantization initialization process described in step S4 is as follows: independently and randomly select an element from the set {+1, -1} to fill the patient latent feature matrix W and the physiological test index latent feature matrix Q until the filling is complete.

5. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 1, characterized in that, Step S503 specifically comprises: S5031: Fixing the latent feature matrix Q of the physiological test indicators, generating multiple samples of the latent feature matrix W of patients row by row using a random algorithm, and calculating the fitness of each sample according to the following formula. ; S5032: Select the fitness of each row. The smallest sample is used to update the row vector corresponding to the patient's latent feature matrix W; S5033: With the physiological indicator latent feature matrix W fixed, multiple samples of the physiological indicator latent feature matrix Q are generated row by row using a random algorithm, and the fitness value corresponding to each sample is calculated according to the following formula. ; S5034: Select the fitness of each row The smallest sample is used to update the row vector corresponding to the latent feature matrix Q of the physiological detection index.

6. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 5, characterized in that, The random algorithm described in steps S5031 and S5033 is the particle swarm optimization algorithm.

7. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 5, characterized in that, The random algorithm described in steps S5031 and S5033 is an adaptive differential evolution algorithm. Specifically, the adaptive differential evolution algorithm is as follows: 1) Initially, the maximum update generation G and the current update generation g = 1 are set, and a binary discrete population X is randomly generated. g ∈{+1, -1} NP×D Where NP is the number of samples; 2) From Select D-dimensional column vectors as samples And calculate the corresponding fitness, then denote the smallest sample as . ;in 3) Calculate the gene mutation coefficient for the current generation. ; ;in, 、 and It is a random number between 0 and 1; at the initial time 4) From Four distinct D-dimensional column vector samples are randomly selected from the data. Calculate the sample number Mutated genes: ;5) The first element will be processed element by element. The mutated genes of each sample are binary converted as follows: ;in, ;6) The elements will be processed one by one. Mutated genes in each sample Cross-genes are obtained by performing cross-genetic manipulation. : ;in, It is a random number between 0 and 1. It is a real number threshold between 0 and 1 that is set manually. For random indices of [1, D], for The One element; 7) Selecting the next generation of samples one by one: ;in, ;8) Iteratively update the algebra g=g+1 and the binary discrete population. This continues until the maximum number of iterations, G, is set manually.

8. The efficient method for perioperative time-series data recovery based on hash latent feature learning according to claim 1, characterized in that, The criterion for determining whether recovery can be completed in step S6 is the calculated loss function. If the value is less than a manually set threshold, recovery can be completed; otherwise, it cannot.

9. A high-efficiency recovery device applied to the efficient perioperative time-series data recovery method based on hash implicit feature learning as described in any one of claims 1 to 8, characterized in that, It consists of a server (1) and N data acquisition devices (2), which communicate with each other via a network. The server (1) includes a server data storage module (11) and a data recovery module (12). Each data acquisition device (2) is a local computer, which includes a local physiological time series data receiving module (21) and a local physiological time series data storage module (22). The value of N depends on the number of physiological time series data acquisition devices. The local physiological time series data receiving module (21) in the data acquisition device (2) is connected to the physiological time series data acquisition device via a network. It is used to receive incomplete physiological time series data fed back by the physiological time series data acquisition device and to indicate the local physiological time series data. The data storage module (22) stores incomplete physiological time series data; the local physiological time series data storage module (22) in the data acquisition device (2) is connected to the server data storage module (11) in the server (1) via the network, and sends the incomplete physiological time series data to the server data storage module (11); the server data storage module (11) and the data recovery module (12) are used to send the incomplete physiological time series data of all data acquisition devices (2) to the data recovery module (12) for data recovery processing, and then save the recovered perioperative physiological time series data; the data recovery module (12) uses the hash hidden feature learning method to recover the incomplete physiological time series data.

10. The high-efficiency recovery device according to claim 9, characterized in that, The server data storage module (11) includes a cloud-based physiological time-series data storage unit (111), a data preprocessing unit (112), and a recovery data storage unit (113). The cloud-based physiological time-series data storage unit (111) in the server data storage module (11) is used to receive and store incomplete physiological time-series data sent by the local physiological time-series data storage module (22) in the data acquisition device (2). The data preprocessing unit (112) is connected to the cloud-based physiological time-series data storage unit (111) and the data recovery module (12) and is used to preprocess the incomplete physiological time-series data of all data acquisition devices (2). The recovery data storage unit (113) is connected to the data recovery module (12) and is used to store the recovered perioperative physiological time-series data.

Citation Information

Patent Citations

  • Cross-modal hash retrieval method based on semantic graph evolution

    CN116680432A

  • Medical decision-oriented multi-modal data dynamic fusion and labeling method and system

    CN119377894A