A health status prediction method based on multi-vital sign time series data
By dividing time-series data of multiple vital signs, feature extraction and frequency domain conversion, generating an enhanced feature matrix and inputting the Transformer module for prediction, the problem of insufficient prediction accuracy and generalization ability of the existing technology when processing multimodal and dynamic data is solved, and more efficient personalized health status prediction is achieved.
Patent Information
- Application Number
- CN202510036178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The prior art is difficult to achieve efficient personalized health status prediction when processing multimodal and dynamic vital sign data, especially in capturing fine-grained feature interactions between multimodal data and dynamic changes in individualized health status.
A health status prediction method based on multi-viral sign timing data is proposed. By obtaining standardized multi-modal sign timing data, time segment division, feature extraction and frequency domain conversion are performed, enhancement feature matrix is generated, and input it into the Transformer module for prediction.
This method can improve the adaptability and generalization of the model when dealing with multimodal and dynamic features, and improve the accuracy and real-time performance of health status prediction.
Smart Images

Figure CN119475100B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to health status prediction, and in particular to a health status prediction method based on multiple vital sign time series data. Background Art
[0002] With the rapid development of science and technology and the explosive growth of data, time series data prediction is increasingly used in finance, healthcare, intelligent monitoring and other fields. Especially in personalized health management, multimodal data prediction based on vital signs can provide real-time assessment and prediction of individual health status. However, due to the diversity and complexity of vital sign data, traditional time series prediction methods still face many challenges in dealing with multimodal collaboration, data missing and long-term dependency.
[0003] The current single-modality time series prediction model has great limitations when processing complex multimodal vital signs data. There are complex intercorrelations and nonlinear relationships between multimodal data. Traditional time series prediction methods are difficult to effectively capture the fine-grained feature interactions between different modalities, resulting in insufficient prediction accuracy and generalization ability of the model. In addition, vital signs data often have highly individualized characteristics, and existing models are difficult to predict and manage the dynamic health status of individuals in real time and accurately.
[0004] In recent years, time series prediction models based on deep learning have gradually emerged, especially models represented by deep learning models based on self-attention mechanisms (Transformer), which have shown great potential in capturing long-term dependencies and multimodal data fusion. However, the existing Transformer model still has deficiencies in processing multimodal and individualized data, making it difficult to achieve efficient personalized predictions in complex dynamic health scenarios. Summary of the invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a health status prediction method based on multi-vital sign time series data, which can improve the adaptability and generalization of the model in dealing with multi-modal and dynamic characteristics.
[0006] The present invention also provides a health status prediction system based on multiple vital signs time series data, a control device for executing the above-mentioned health status prediction method based on multiple vital signs time series data, and a computer-readable storage medium.
[0007] According to a first aspect of the present invention, a health status prediction method based on multiple vital signs time series data comprises:
[0008] Acquire standardized time series data of multiple vital signs in different modalities;
[0009] Input the plurality of vital sign time series data into the slice and block frequency domain time series processing module for preprocessing to obtain an enhanced feature matrix, wherein the preprocessing includes: dividing each of the vital sign time series data into time segments to obtain a plurality of segment data, extracting features from each of the segment data to obtain corresponding time domain segment features, performing frequency domain conversion on each of the time domain segment features to obtain corresponding frequency domain segment features, and then converting the frequency domain segment features back to the time domain to obtain the enhanced feature matrix;
[0010] The enhanced feature matrix is input into a Transformer module to obtain a health status prediction result, wherein the Transformer module includes an encoder module, a decoder module and an output layer module which are connected in sequence.
[0011] The health status prediction method based on multiple vital signs time series data according to the embodiment of the present invention has at least the following beneficial effects:
[0012] After obtaining standardized vital sign time series data of multiple different modalities, each vital sign time series data is divided into time segments to obtain multiple segment data. Each segment data can be processed independently, thereby enhancing the model's adaptability to distribution changes in different time periods. By extracting features from each segment data, the corresponding time domain segment features are obtained, and each time domain segment feature is converted to the frequency domain to obtain the corresponding frequency domain segment features. The frequency domain segment features are then converted back to the time domain. The frequency domain information can be used to further compress and filter the data, thereby reducing noise and redundant information while retaining key features. Finally, the enhanced feature matrix is input into the Transformer module to predict the health status prediction result, which can improve the adaptability and generalization of the model in dealing with multimodal and dynamic features.
[0013] According to some embodiments of the present invention, extracting features from each of the segment data to obtain corresponding time domain segment features includes:
[0014] Use a sliding window to segment each of the fragment data to obtain a local feature matrix corresponding to each segment;
[0015] Performing linear transformation and self-attention calculation on each of the local feature matrices to obtain a corresponding first attention result;
[0016] Aggregate the first attention results of all the local feature matrices corresponding to each of the fragment data to obtain the time domain fragment feature corresponding to each of the fragment data.
[0017] According to some embodiments of the present invention, performing frequency domain conversion on each of the time domain segment features to obtain corresponding frequency domain segment features includes:
[0018] Performing frequency domain conversion on each of the time domain segment features using fast Fourier transform to obtain a corresponding original frequency domain representation;
[0019] Performing interpolation processing on each of the original frequency domain representations to obtain a corresponding interpolated frequency domain representation;
[0020] Each of the interpolated frequency domain representations is input into a low-pass filter to obtain the corresponding frequency domain segment features.
[0021] According to some embodiments of the present invention, the Transformer module further includes an analog denoising module, and the step of inputting the enhanced feature matrix into the Transformer module to obtain a health status prediction result includes:
[0022] Inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix;
[0023] Inputting the initial output feature matrix into the analog denoising module to perform analog step-by-step denoising processing to obtain a final output feature matrix;
[0024] Inputting the final output feature matrix into the decoder module to obtain decoded output data;
[0025] The decoded output data is input into the output layer module to obtain the health status prediction result.
[0026] According to some embodiments of the present invention, the simulated stepwise denoising process comprises:
[0027] Get the preset noise matrix, denoising matrix and step threshold;
[0028] Adding the noise matrix to the initial output feature matrix to obtain a noisy feature matrix;
[0029] The current noisy feature matrix is denoised step by step using the denoising matrix until the number of denoising steps reaches the step threshold, thereby obtaining the final output feature matrix after denoising.
[0030] According to some embodiments of the present invention, inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix comprises:
[0031] Obtaining a mask matrix, wherein the mask matrix is equal to 1 when the signal at the target position belongs to the critical area, and is equal to 0 when the signal at the target position belongs to the non-critical area;
[0032] Using the mask matrix, the enhanced feature matrix is divided into a key area feature matrix and a non-key area feature matrix;
[0033] Performing linear transformation and self-attention calculation on the key area feature matrix and the non-key area feature matrix respectively, and obtaining a second attention result and a third attention result accordingly;
[0034] The second attention result and the third attention result are weightedly combined to obtain the initial output feature matrix.
[0035] According to some embodiments of the present invention, performing weighted combination on the second attention result and the third attention result to obtain the initial output feature matrix includes:
[0036] Acquire a pre-trained semantic anchor point set, wherein the semantic anchor point set includes a plurality of semantic anchor point vectors related to health status;
[0037] Performing a weighted combination of the second attention result and the third attention result to generate a fusion feature matrix;
[0038] Calculating the similarity between the features of each time step in the fusion feature matrix and all the semantic anchor point vectors in the semantic anchor point set;
[0039] Using cosine similarity, sorting multiple similarities corresponding to each time step and selecting the top multiple semantic anchor vectors as semantic cue embedding;
[0040] Concatenate and expand the features of each time step according to the semantic hint embedding to obtain an extended feature;
[0041] The expanded features are linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
[0042] According to some embodiments of the present invention, inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix comprises:
[0043] Performing linear transformation and multi-head attention calculation on the features of each time step of the enhanced feature matrix, respectively, to obtain multi-head attention results corresponding to each time step, wherein different heads have different linear transformation parameters;
[0044] Concatenate the multi-head attention results corresponding to each time step to obtain the multi-head attention output;
[0045] The multi-head attention output is linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
[0046] According to some embodiments of the present invention, performing a linear transformation on the multi-head attention output to map back to the same dimension to obtain the initial output feature matrix includes:
[0047] Perform a linear transformation on the multi-head attention output to map it back to the same dimension to obtain a multi-head attention transformation matrix;
[0048] Performing a residual connection between the multi-head attention transformation matrix and the enhanced feature matrix to obtain a first connection result;
[0049] Performing layer normalization on the first connection result to obtain a normalized result;
[0050] Inputting the normalized result into a feedforward network to obtain a feedforward output;
[0051] Performing a residual connection between the normalized result and the feedforward output to obtain a second connection result;
[0052] The second connection result is layer normalized to obtain the initial output feature matrix.
[0053] According to some embodiments of the present invention, dividing each of the vital sign time series data into time segments to obtain a plurality of segment data includes:
[0054] Dividing each of the vital sign time series data into time segments to obtain a plurality of original segments;
[0055] Local normalization processing is performed on each of the original segments to obtain a plurality of segment data.
[0056] According to a second aspect of the present invention, a health status prediction system based on multiple vital signs time series data comprises:
[0057] A time series data acquisition unit, used to acquire standardized time series data of vital signs in different modalities;
[0058] A preprocessing unit is used to input the plurality of vital sign time series data into the slice and block frequency domain time series processing module for preprocessing to obtain an enhanced feature matrix, wherein the preprocessing includes: dividing each of the vital sign time series data into time segments to obtain a plurality of segment data, extracting features from each of the segment data to obtain corresponding time domain segment features, performing frequency domain conversion on each of the time domain segment features to obtain corresponding frequency domain segment features, and then converting the frequency domain segment features back to the time domain to obtain the enhanced feature matrix;
[0059] The prediction unit is used to input the enhanced feature matrix into a Transformer module to obtain a health status prediction result. The Transformer module includes an encoder module, a decoder module and an output layer module connected in sequence.
[0060] The health status prediction system based on multiple vital signs time series data according to the embodiment of the present invention has at least the following beneficial effects:
[0061] After obtaining standardized vital sign time series data of multiple different modalities, each vital sign time series data is divided into time segments to obtain multiple segment data. Each segment data can be processed independently, thereby enhancing the model's adaptability to distribution changes in different time periods. By extracting features from each segment data, the corresponding time domain segment features are obtained, and each time domain segment feature is converted to the frequency domain to obtain the corresponding frequency domain segment features. The frequency domain segment features are then converted back to the time domain. The frequency domain information can be used to further compress and filter the data, thereby reducing noise and redundant information while retaining key features. Finally, the enhanced feature matrix is input into the Transformer module to predict the health status prediction result, which can improve the adaptability and generalization of the model in dealing with multimodal and dynamic features.
[0062] According to the third aspect of the present invention, the control device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the health status prediction method based on multiple vital signs time series data as described in the first aspect of the embodiment. Since the control device adopts all the technical solutions of the health status prediction method based on multiple vital signs time series data of the above embodiment, it has at least all the beneficial effects brought by the technical solutions of the above embodiment.
[0063] According to the computer-readable storage medium of the fourth aspect embodiment of the present invention, computer executable instructions are stored, and the computer executable instructions are used to execute the health status prediction method based on multiple vital signs time series data as described in the first aspect embodiment. Since the computer-readable storage medium adopts all the technical solutions of the health status prediction method based on multiple vital signs time series data of the above embodiment, it at least has all the beneficial effects brought by the technical solutions of the above embodiment.
[0064] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0066] Figure 1 It is a flowchart of a health status prediction method based on multiple vital signs time series data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0067] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0068] In the description of the present invention, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0069] In the description of the present invention, it should be understood that descriptions involving orientation, such as orientation or positional relationship indicated as up, down, etc., are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0070] In the description of the present invention, it should be noted that, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0071] The following will be combined Figure 1 A clear and complete description is given of the health status prediction method based on multiple vital signs time series data according to an embodiment of the present invention. Obviously, the embodiment described below is only a part of the embodiments of the present invention, not all of the embodiments.
[0072] refer to Figure 1 , Figure 1 It is a flowchart of a health status prediction method based on multiple vital signs time series data according to an embodiment of the present invention.
[0073] According to a first aspect of the present invention, a health status prediction method based on multiple vital signs time series data comprises:
[0074] Acquire standardized time series data of multiple vital signs in different modalities;
[0075] Input multiple vital sign time series data into the slice and block frequency domain time series processing module for preprocessing to obtain an enhanced feature matrix, the preprocessing includes: dividing each vital sign time series data into time segments to obtain multiple segment data, extracting features from each segment data to obtain corresponding time domain segment features, performing frequency domain conversion on each time domain segment feature to obtain corresponding frequency domain segment features, and then converting the frequency domain segment features back to the time domain to obtain an enhanced feature matrix;
[0076] The enhanced feature matrix is input into the Transformer module to obtain the health status prediction result. The Transformer module includes an encoder module, a decoder module and an output layer module connected in sequence.
[0077] In some embodiments, multiple different modalities of vital sign time series data include but are not limited to electrocardiogram data, electroencephalogram data, multi-channel post-meal blood glucose data, and personal health information data.
[0078] The ECG data is collected from multi-lead ECGs. The signal of each lead is a time series, and all the lead data at each time step can be represented by a matrix:
[0079] ;Formula (1)
[0080] in, T Indicates the number of time steps, rows n Represents the number of leads, each element Indicates that at time step t At that time, i The signal value of each lead.
[0081] For EEG data, similarly, n The EEG data of channels is also represented as a matrix:
[0082] ;Formula (2)
[0083] in, T Indicates the number of time steps, rows n Indicates the number of channels, each element Indicates that at time step t At that time, i The signal value of each channel.
[0084] For multi-channel postprandial blood glucose data, it is also represented as a matrix:
[0085] ;Formula (3)
[0086] in, T Indicates the number of time steps, rows n Indicates the number of channels, each element Indicates that at time step t At that time, i The signal value of each channel.
[0087] For personal health information data, such as static features such as height, weight, and lung capacity, this information is considered as an independent feature vector and integrated into the model through embedding. This static health information is represented as a vector:
[0088] ;Formula (4)
[0089] in, Indicates height, represents weight. Assume that features, then Just one Vector.
[0090] Normalization of the ECG data, EEG data, and postprandial blood glucose data is used to eliminate the feature scale differences between different modalities and channels. x The i channels, time steps t The signal value The standardized formula is:
[0091] ;Formula (5)
[0092] in, and are the mean and standard deviation of the channel respectively.
[0093] It can be understood that the sliced and blocked frequency domain time series processing module is a fragment-based comprehensive processing module for complex multimodal time series data. It integrates the steps of time slicing, local feature extraction, frequency domain processing and enhanced non-stationary adaptability, and is suitable for in-depth modeling and analysis of multi-fragment and multi-channel time series data.
[0094] Time segmentation is performed on the standardized vital sign time series data to adapt to the characteristics of non-stationary data. The purpose of time segmentation is to divide a long time series into multiple shorter, adjacent, non-overlapping segments. The data in each segment is considered to be independently processed, thereby enhancing the model's adaptability to distribution changes in different time periods.
[0095] Set a fixed time segment length L , indicating that each time segment contains L time step data. L The value of can be selected according to the characteristics and non-stationarity of the data. T If a data sequence of time steps T Cannot be L Divisible, the last deficiency can be ignored L Fragment of .
[0096] The standardized vital signs time series data is divided into multiple L The data matrix is represented as (each row represents a channel and each column represents a time step), and its size is ,in, n is the number of channels, T is the number of time steps.
[0097] Defining the fragment matrix for:
[0098] ;Formula (6)
[0099] in, Indicates that from the ( k -1) L +1 column to kL The submatrix of the column corresponds to a length of L The size of each fragment data is , indicating that it contains n channels, L time-step fragments.
[0100] In some embodiments of the present invention, each vital sign time series data is divided into time segments to obtain multiple segment data, including:
[0101] Divide each vital sign time series data into time segments to obtain multiple original segments;
[0102] Each original segment is locally normalized to obtain multiple segment data.
[0103] For each original fragment (Right now ), local normalization is performed on each channel:
[0104] = ;Formula (7)
[0105] in, For the original fragment k Middle i Channel, j Normalized vital sign time series data of time steps.
[0106] For the original fragment k Middle i The mean of the channels:
[0107] = ;Formula (8)
[0108] For the original fragment k Middle i Standard deviation of channels:
[0109] = ;Formula (9)
[0110] After completing the normalization within the fragment, the processed fragment matrix is obtained , contains the local normalized data (fragment data) of all original fragments:
[0111] ;Formula (10)
[0112] Local normalization can further reduce the mean and variance differences within the original segments, eliminate the non-stationarity between different time segments, and improve the adaptability of the model.
[0113] In some embodiments of the present invention, feature extraction is performed on each segment data to obtain corresponding time domain segment features, including:
[0114] Use a sliding window to segment each fragment data and obtain the local feature matrix corresponding to each segment;
[0115] Perform linear transformation and self-attention calculation on each local feature matrix to obtain the corresponding first attention result;
[0116] Aggregate the first attention results of all local feature matrices corresponding to each fragment data to obtain the time domain fragment features corresponding to each fragment data.
[0117] It can be understood that each sliding window generates a local feature matrix, which represents the time series dependency within the window, and each fragment data corresponds to multiple local feature matrices.
[0118] Set a sliding window size W , indicating that each window contains W The sliding window slides over each fragment of data to generate multiple local feature matrices. Starting from the first time step of each fragment of data, the window slides one step at a time to extract features within the window until the entire fragment of data is covered. The size is , n is the number of channels, L is the length of the fragment data. The local feature matrix generated by the sliding window is:
[0119] ;Formula (11)
[0120] in, For fragment data k From the first j Column to j + W - 1 The data submatrix of the column is of size .
[0121] This sliding window operation is performed on each fragment of data. Generate L - W + 1 A local feature matrix.
[0122] Taking each local feature matrix as an independent input segment, the self-attention mechanism models these local feature matrices to capture the temporal dependencies within each local feature matrix and the global dependencies between local feature matrices.
[0123] Perform a linear transformation on each local feature matrix to generate query (Q), key (K), and value (V) matrices:
[0124] , , ;Formula (12)
[0125] in, , , is the learnable parameter matrix.
[0126] Self-attention is calculated for each local feature matrix to capture the dependencies between time steps:
[0127] Attention ( , , )= Softmax ( ) ;Formula (13)
[0128] in, is the dimension of the key matrix, which is used to scale the attention scores to ensure numerical stability.
[0129] Aggregate the first attention results of all local feature matrices corresponding to each fragment data to obtain the time domain fragment feature corresponding to each fragment data, that is, the fragment k The global feature representation of is:
[0130] ;Formula (14)
[0131] In some embodiments of the present invention, each time domain segment feature is converted into a frequency domain to obtain a corresponding frequency domain segment feature, including:
[0132] Use fast Fourier transform to convert each time domain segment feature into frequency domain to obtain the corresponding original frequency domain representation;
[0133] Perform interpolation processing on each original frequency domain representation to obtain a corresponding interpolated frequency domain representation;
[0134] Each interpolated frequency domain representation is input into a low-pass filter to obtain the corresponding frequency domain segment features.
[0135] It can be understood that the fast Fourier transform is used to decompose the time series signal into a combination of different frequency components. The formula is as follows:
[0136] FFT( ); formula (15)
[0137] Frequency domain information is used to further compress and filter the data, thereby reducing noise and redundant information while retaining key features.
[0138] By interpolating, the frequency resolution can be increased in the frequency domain, thereby enhancing the prediction ability of the model. , the interpolated frequency domain signal size will be expanded to the original times. The interpolation formula is:
[0139] ;Formula (16)
[0140] The low-pass filter can remove frequency components higher than the specified cutoff frequency and retain only important low-frequency components to reduce high-frequency noise. The low-pass filter formula is:
[0141] ;Formula (17)
[0142] in, It is the mask matrix of the low-pass filter, which only retains the components with frequencies below the cutoff frequency and sets the components above the cutoff frequency to zero.
[0143] After completing the frequency domain processing, the frequency domain segment features are converted back to the time domain using the inverse Fourier transform to obtain the enhanced feature matrix. The formula is as follows:
[0144] ;Formula (18)
[0145] is the time domain signal processed by the model, representing the fragment k The enhanced features of each segment are stacked together in the order of the original time series to form a complete enhanced feature matrix, which is used to represent the final feature representation of the entire time series data. The formula is as follows:
[0146] H = Formula (19)
[0147] The size of the enhanced feature matrix is ,in, n is the number of channels, L is the length after interpolation. The enhanced feature matrix contains the data features of all modalities and is used as the input of the encoder module.
[0148] In some embodiments of the present invention, the enhanced feature matrix is input into the encoder module to obtain an initial output feature matrix, including:
[0149] Obtain a mask matrix, where the mask matrix is equal to 1 when the signal at the target position belongs to the critical area, and is equal to 0 when the signal at the target position belongs to the non-critical area;
[0150] Using a mask matrix, the enhanced feature matrix is divided into a key region feature matrix and a non-key region feature matrix;
[0151] Perform linear transformation and self-attention calculation on the key area feature matrix and the non-key area feature matrix respectively, and obtain the second attention result and the third attention result respectively;
[0152] The second attention result and the third attention result are weightedly combined to obtain the initial output feature matrix.
[0153] According to the region separation idea of the Residual Denoising Diffusion Model (RDDM), a mask matrix is defined , used to identify which time steps and channels belong to a specific "region of interest". The size of the mask matrix is the same as the enhanced feature matrix. The mask matrix is defined as: if at the target location ( i , t ) belongs to the critical area [ i , t ]=1, the signal at the target location belongs to the non-critical area, then [ i,t ]=0.
[0154] The formula of the key area feature matrix is:
[0155] = ;Formula (20)
[0156] in, Represents element-wise multiplication.
[0157] The formula of the non-critical area feature matrix is:
[0158] = ;Formula (21)
[0159] It should be noted that is with Matrices of the same size.
[0160] right and Perform a linear transformation and map them into query (Q), key (K), and value (V) matrices respectively for use in the subsequent attention mechanism.
[0161] Linear transformation of the key area feature matrix:
[0162] , , ;Formula (22)
[0163] in, , , is the learnable parameter matrix, and its dimensions are 、 、 , are the dimensions of the key and query matrices, is the dimension of the value matrix.
[0164] Linear transformation of feature matrix in non-critical area:
[0165] , , ; Formula (23)
[0166] in, , , is a learnable parameter matrix with the same dimension as the linear transformation matrix of the key region feature matrix.
[0167] In getting , , as well as , , After that, self-attention calculation is performed.
[0168] Self-attention calculation of key areas (second attention result):
[0169] =Softmax ( ) ;Formula (24)
[0170] Self-attention calculation of non-critical areas (third attention result):
[0171] = Softmax ( ) ;Formula (25)
[0172] Selective denoising of key areas can improve feature extraction quality and prediction accuracy.
[0173] In some embodiments of the present invention, the second attention result and the third attention result are weightedly combined to obtain an initial output feature matrix, including:
[0174] Obtain a pre-trained semantic anchor point set, where the semantic anchor point set includes multiple semantic anchor point vectors related to health status;
[0175] Perform a weighted combination of the second attention result and the third attention result to generate a fusion feature matrix;
[0176] Calculate the similarity between the features of each time step in the fusion feature matrix and all semantic anchor vectors in the semantic anchor set;
[0177] Use cosine similarity to sort multiple similarities corresponding to each time step and select the top multiple semantic anchor vectors as semantic cue embedding;
[0178] According to the semantic hint embedding, the features of each time step are concatenated and expanded to obtain the extended features;
[0179] The expanded features are linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
[0180] The second attention result and the third attention result are weighted combined to obtain the fusion feature matrix F :
[0181] +(1- ) ;Formula (26)
[0182] in, is a weight coefficient (which can be a hyperparameter or learned through training) used to control the weight distribution between key areas and non-key areas.
[0183] The final fusion feature matrix FIt is a feature matrix that integrates the information of key areas and non-key areas. It contains the characteristics of multimodal time series. It not only contains important local area information, but also integrates the global features of other non-key areas.
[0184] The fusion feature matrix can be enhanced by introducing a set of pre-trained semantic anchors F The semantic representation of , so that the model can capture more potential information related to health status when processing vital signs time series data, and realize dynamic monitoring and prediction of individual health status. Fusion feature matrix F The size is ,in n is the number of channels, T Introducing semantic anchor hints for the features of each time step can enhance the model's understanding of vital sign data, especially health-related semantic information. The semantic anchor vector represents a semantic information related to health status, such as "heart rate increase", "abnormal brain waves", etc. The dimension is d , that is, each .
[0185] for F Each column in (i.e., each time step t ), and calculate its similarity with all semantic anchor vectors:
[0186] = ;Formula (27)
[0187] For each time step, use cosine similarity to sort the multiple similarities corresponding to each time step and select the top similarities. K Semantic anchor vectors are embedded as semantic cues :
[0188] = Top - K ( ); formula (28)
[0189] The features of each time step are concatenated and expanded according to the semantic cue embedding to obtain the extended features, so that the extended feature representation of each time step contains the original time series features and semantic cue information.
[0190] The calculation formula of the extended feature is:
[0191] ;Formula (29)
[0192] Among them, the splicing operation makes The dimension becomes .
[0193] In order to integrate the original features and semantic hint information, the expanded features are linearly transformed to map back to the same dimension to ensure consistency with the input format of subsequent modules.
[0194] The formula for the initial output feature matrix is:
[0195] + ;Formula (30)
[0196] in, , , are all learnable parameters used to project the expanded features back to the original feature space.
[0197] The final initial output feature matrix It includes the multimodal features of time series data and integrates the most relevant semantic cues, providing richer semantic information for the next module, making the model more interpretable and predictive when processing vital sign time series data.
[0198] In some embodiments of the present invention, the enhanced feature matrix is input into the encoder module to obtain an initial output feature matrix, including:
[0199] Perform linear transformation and multi-head attention calculation on the features of each time step of the enhanced feature matrix to obtain the multi-head attention results corresponding to each time step, where different heads have different linear transformation parameters.
[0200] Concatenate the multi-head attention results corresponding to each time step to obtain the multi-head attention output;
[0201] The multi-head attention output is linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
[0202] In getting Then, for each time step, the feature A linear transformation is performed to map it into query, key, and value matrices for computing attention between time steps.
[0203] Define the linear transformation matrix , , , the dimensions are 、 、 , are the dimensions of the key and query matrices, is the dimension of the value matrix.
[0204] Compute the query, key, and value matrices:
[0205] Q = , K = , V = ;Formula (31)
[0206] in, Q The size is , K The size is , V The size is .
[0207] Calculate the self-attention score:
[0208] For the query and key matrices, an attention score is calculated between each time step, which measures the correlation between features at different time steps.
[0209] Attention_scores = ;Formula (32)
[0210] Here is a scaling factor used to prevent the inner product value from being too large, leading to numerical instability.
[0211] Attention_weights = Softmax ( ); formula (33)
[0212] in, Attention_weights The dimension is , represents the attention weight of each time step to other time steps.
[0213] Using the attention weight matrix Attention_weights Pair Matrix V Perform weighted summation to obtain a new feature representation for each time step.
[0214] Attention_output = Attention_weights V ;Formula (34)
[0215] in, Attention_output The size is , contains the representation of each time step in the global context.
[0216] Multi-head attention mechanism:
[0217] To capture different feature relationships, we use hEach head has its own linear transformation parameters, calculates the attention output independently, and then concatenates the outputs of each head to form the final multi-head attention output.
[0218] For i head, define its linear transformation matrix , , , and then calculate the query, key, and value matrices for the header respectively:
[0219] , , ;Formula (35)
[0220] Attention_ ;Formula (36)
[0221] Splicing h Output of each header:
[0222] MultiHead_output = Concat ( Attention_ Attention_ );
[0223] Formula (37)
[0224] By adding cross-modal attention allocation to the multi-head attention mechanism, increasing the multi-scale feature decomposition module, and optimizing the allocation and parallel design of computing resources, not only the fine-grained dependency modeling capabilities between different modalities are enhanced, but also the computational overhead is reduced and the real-time prediction efficiency is improved, thus providing a more efficient and accurate solution for personalized prediction of health status.
[0225] In some embodiments of the present invention, a linear transformation is performed on the multi-head attention output to map back to the same dimension to obtain an initial output feature matrix, including:
[0226] Perform a linear transformation on the multi-head attention output to map back to the same dimension and obtain the multi-head attention transformation matrix;
[0227] Perform a residual connection between the multi-head attention transformation matrix and the enhanced feature matrix to obtain a first connection result;
[0228] Perform layer normalization on the first connection result to obtain a normalized result;
[0229] Input the normalized result into the feedforward network to obtain the feedforward output;
[0230] Perform a residual connection between the normalized result and the feedforward output to obtain a second connection result;
[0231] The second connection result is layer normalized to obtain the initial output feature matrix.
[0232] It is understandable that residual connection can keep the gradient stable and enhance the training effect, layer normalization can ensure the numerical stability of the output, and the feedforward network can further extract complex features and increase the expressive power of the model. The specific principles and working processes of residual connection, layer normalization and feedforward network are prior arts known to those skilled in the art and will not be elaborated here.
[0233] The final initial output feature matrix It contains the global features of time series data and cross-modal contextual relationships, and is the input for subsequent prediction or classification tasks.
[0234] In some embodiments of the present invention, the Transformer module further includes an analog denoising module, which inputs the enhanced feature matrix into the Transformer module to obtain a health status prediction result, including:
[0235] Input the enhanced feature matrix into the encoder module to obtain the initial output feature matrix;
[0236] The initial output feature matrix is input into the simulation denoising module for simulation step-by-step denoising to obtain the final output feature matrix;
[0237] The final output feature matrix is input into the decoder module to obtain the decoded output data;
[0238] The decoded output data is input into the output layer module to obtain the health status prediction result.
[0239] In some embodiments of the invention, the simulated stepwise denoising process comprises:
[0240] Get the preset noise matrix, denoising matrix and step threshold;
[0241] Add the noise matrix to the initial output feature matrix to obtain the noisy feature matrix;
[0242] The denoising matrix is used to perform step-by-step denoising on the current noisy feature matrix until the number of denoising steps reaches the step threshold, and the final output feature matrix after denoising is obtained.
[0243] According to the inverse process of RDDM, the time series features can be further enhanced by simulating the generation method of stepwise denoising, thereby improving the model's cross-modal understanding ability of vital sign time series data, extracting high-fidelity signals, and realizing robust feature modeling in complex health data scenarios.
[0244] First, in the initial output feature matrix Add noise to simulate the characteristic matrix of the initial state. Noise matrix is a random matrix sampled from a standard normal distribution, and Same size.
[0245] ;Formula (38)
[0246] in, is the noise feature matrix, which will be used as the starting state of the inverse process.
[0247] Step-by-step denoising process:
[0248] The number of steps in the inverse process is defined as N , that is, through N The noise is gradually eliminated by the denoising steps. i =1 , 2 ,...,N In the example, based on the current noise feature matrix Perform denoising to generate the next step of denoising feature matrix.
[0249] Generate a denoising matrix:
[0250] In each step, we learn a denoising matrix To reduce the impact of noise on features. This denoising matrix is obtained through a learnable linear transformation:
[0251] = ;Formula (39)
[0252] in, , , are the first i step of learnable parameters.
[0253] Update the noisy feature matrix:
[0254] Using the Denoising Matrix Add noise to the current feature matrix Update to get the next step Here we introduce a scaling factor , used to control the denoising step.
[0255] = - ;Formula (40)
[0256] in, It is i The denoising strength coefficient of the step is usually set as a decreasing series during training, for example:
[0257] = ;Formula (41)
[0258] In passing N After the denoising step, the final output feature matrix is obtained, which is the feature matrix after denoising. .
[0259] The decoder module receives the output of the encoder module as input and generates query, key, and value matrices:
[0260] , , ;Formula (42)
[0261] in, , and is the decoder weight matrix.
[0262] The decoder uses a future time step mask in its attention mechanism to ensure that the model can only focus on the current and previous time steps, avoiding leakage of future information:
[0263] ;Formula (43)
[0264] in, is an upper triangular matrix used to mask out future time step information.
[0265] The decoder module combines the output of the encoder module through the encoder-decoder attention mechanism :
[0266] The output of the decoder module is matched with the output of the encoder module to obtain contextual information:
[0267] , , ;Formula (44)
[0268] Use the contextual information from the encoder module to calculate the attention scores and generate the output of the decoder module:
[0269] ;Formula (45)
[0270] This attention layer injects the global contextual information of the encoder module into the decoder module, enabling the decoder module to refer to the output of the encoder module when generating predictions.
[0271] In some embodiments, the output of the decoder module is residually connected to the input of the decoder module and layer normalized to form a more stable feature representation. , which contains a comprehensive representation of the time series information from the decoder’s self-attention and the encoder’s contextual information.
[0272] In some embodiments, the residual connection and layer normalization are processed Input to the feedforward network to further enrich the features. After the output of the feedforward network, residual connection and layer normalization are performed again to obtain the decoded output data, and then the decoded output data is input into the output layer module to obtain the health status prediction result.
[0273] The health status prediction results include classification task results and regression task results. The classification task is to use Softmax The function obtains the category probability distribution. The classification task results include but are not limited to the classification of the heart state of the electrocardiogram data or the classification of the emotional state of the electroencephalogram data. The regression task directly outputs the predicted value, that is, the regression task results include but are not limited to the continuous value prediction of the health indicator.
[0274] According to the health status prediction method based on multiple vital sign time series data of an embodiment of the present invention, after obtaining the standardized vital sign time series data of multiple different modalities, each vital sign time series data is divided into time segments to obtain multiple segment data, and each segment data can be processed independently, thereby enhancing the adaptability of the model to distribution changes in different time periods. By extracting features from each segment data, the corresponding time domain segment features are obtained, and each time domain segment feature is converted to the frequency domain to obtain the corresponding frequency domain segment features, and then the frequency domain segment features are converted back to the time domain, the frequency domain information can be used to further compress and filter the data, thereby reducing noise and redundant information while retaining key features. Finally, the enhanced feature matrix is input into the Transformer module to predict the health status prediction result, which can improve the adaptability and generalization of the model when dealing with multi-modal and dynamic features.
[0275] According to a health status prediction system based on multi-vital sign time series data according to an embodiment of the second aspect of the present invention, the system includes a time series data acquisition unit, a preprocessing unit and a prediction unit.
[0276] A time series data acquisition unit, used to acquire standardized time series data of vital signs in different modalities;
[0277] A preprocessing unit is used to input multiple vital sign time series data into the slice and block frequency domain time series processing module for preprocessing to obtain an enhanced feature matrix, wherein the preprocessing includes: dividing each vital sign time series data into time segments to obtain multiple segment data, extracting features from each segment data to obtain corresponding time domain segment features, performing frequency domain conversion on each time domain segment feature to obtain corresponding frequency domain segment features, and then converting the frequency domain segment features back to the time domain to obtain an enhanced feature matrix;
[0278] The prediction unit is used to input the enhanced feature matrix into the Transformer module to obtain the health status prediction result. The Transformer module includes an encoder module, a decoder module and an output layer module connected in sequence.
[0279] According to the health status prediction system based on multiple vital sign time series data of an embodiment of the present invention, after obtaining the standardized vital sign time series data of multiple different modalities, each vital sign time series data is divided into time segments to obtain multiple segment data, and each segment data can be processed independently, thereby enhancing the adaptability of the model to distribution changes in different time periods. By extracting features from each segment data, the corresponding time domain segment features are obtained, and each time domain segment feature is converted to the frequency domain to obtain the corresponding frequency domain segment features, and then the frequency domain segment features are converted back to the time domain, the frequency domain information can be used to further compress and filter the data, thereby reducing noise and redundant information while retaining key features. Finally, the enhanced feature matrix is input into the Transformer module to predict the health status prediction result, which can improve the adaptability and generalization of the model when dealing with multi-modal and dynamic features.
[0280] In addition, an embodiment of the present invention further provides a control device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor and the memory may be connected via a bus or other means.
[0281] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0282] The non-transient software program and instructions required to implement the health status prediction method based on multiple vital signs time series data of the above embodiment are stored in the memory, and when executed by the processor, the health status prediction method based on multiple vital signs time series data in the above embodiment is executed.
[0283] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0284] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor or controller, for example, by the processor of the above-mentioned embodiment, so that the above-mentioned processor can execute the health status prediction method based on multiple vital signs time series data in the above-mentioned embodiment.
[0285] It will be appreciated by those skilled in the art that all or some of the steps and systems in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0286] The embodiments of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in the relevant technical field without departing from the purpose of the present invention.
Claims
1. A health status prediction method based on multiple vital signs time series data, characterized in that: The method comprises: Acquire standardized time series data of multiple vital signs in different modalities; Input the plurality of vital sign time series data into the slice and block frequency domain time series processing module for preprocessing to obtain an enhanced feature matrix, wherein the preprocessing includes: dividing each of the vital sign time series data into time segments to obtain a plurality of segment data, extracting features from each of the segment data to obtain corresponding time domain segment features, performing frequency domain conversion on each of the time domain segment features to obtain corresponding frequency domain segment features, and then converting the frequency domain segment features back to the time domain to obtain the enhanced feature matrix; Inputting the enhanced feature matrix into a Transformer module to obtain a health status prediction result, wherein the Transformer module includes an encoder module, a decoder module, and an output layer module connected in sequence; The Transformer module further includes an analog denoising module, and the enhanced feature matrix is input into the Transformer module to obtain a health status prediction result, including: Inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix; The initial output feature matrix is input into the analog denoising module for analog step-by-step denoising processing to obtain a final output feature matrix, wherein the analog step-by-step denoising processing includes: obtaining a preset noise matrix, a denoising matrix and a step number threshold, adding the noise matrix to the initial output feature matrix to obtain a noisy feature matrix, and using the denoising matrix to perform step-by-step denoising processing on the current noisy feature matrix until the number of denoising steps reaches the step number threshold, thereby obtaining the final output feature matrix after denoising; Inputting the final output feature matrix into the decoder module to obtain decoded output data; The decoded output data is input into the output layer module to obtain the health status prediction result.
2. The health status prediction method based on multiple vital signs time series data according to claim 1 is characterized in that: The extracting features of each of the segment data to obtain corresponding time domain segment features includes: Use a sliding window to segment each of the fragment data to obtain a local feature matrix corresponding to each segment; Performing linear transformation and self-attention calculation on each of the local feature matrices to obtain a corresponding first attention result; Aggregate the first attention results of all the local feature matrices corresponding to each of the fragment data to obtain the time domain fragment feature corresponding to each of the fragment data.
3. The health status prediction method based on multiple vital signs time series data according to claim 1, characterized in that: The performing frequency domain conversion on each of the time domain segment features to obtain corresponding frequency domain segment features includes: Performing frequency domain conversion on each of the time domain segment features using fast Fourier transform to obtain a corresponding original frequency domain representation; Performing interpolation processing on each of the original frequency domain representations to obtain a corresponding interpolated frequency domain representation; Each of the interpolated frequency domain representations is input into a low-pass filter to obtain the corresponding frequency domain segment features.
4. The health status prediction method based on multiple vital signs time series data according to claim 1, characterized in that: The step of inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix comprises: Obtaining a mask matrix, wherein the mask matrix is equal to 1 when the signal at the target position belongs to the critical area, and is equal to 0 when the signal at the target position belongs to the non-critical area; Using the mask matrix, the enhanced feature matrix is divided into a key area feature matrix and a non-key area feature matrix; Performing linear transformation and self-attention calculation on the key area feature matrix and the non-key area feature matrix respectively, and obtaining a second attention result and a third attention result accordingly; The second attention result and the third attention result are weightedly combined to obtain the initial output feature matrix.
5. The health status prediction method based on multiple vital signs time series data according to claim 4 is characterized in that: The weighted combination of the second attention result and the third attention result to obtain the initial output feature matrix includes: Acquire a pre-trained semantic anchor point set, wherein the semantic anchor point set includes a plurality of semantic anchor point vectors related to health status; Performing a weighted combination of the second attention result and the third attention result to generate a fusion feature matrix; Calculating the similarity between the features of each time step in the fusion feature matrix and all the semantic anchor point vectors in the semantic anchor point set; Using cosine similarity, sorting multiple similarities corresponding to each time step and selecting the top multiple semantic anchor vectors as semantic cue embedding; Concatenate and expand the features of each time step according to the semantic hint embedding to obtain an extended feature; The expanded features are linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
6. The health status prediction method based on multiple vital signs time series data according to claim 1, characterized in that: The step of inputting the enhanced feature matrix into the encoder module to obtain an initial output feature matrix comprises: Performing linear transformation and multi-head attention calculation on the features of each time step of the enhanced feature matrix, respectively, to obtain multi-head attention results corresponding to each time step, wherein different heads have different linear transformation parameters; Concatenate the multi-head attention results corresponding to each time step to obtain the multi-head attention output; The multi-head attention output is linearly transformed to map back to the same dimension to obtain the initial output feature matrix.
7. The health status prediction method based on multiple vital signs time series data according to claim 6, characterized in that: The multi-head attention output is linearly transformed to map back to the same dimension to obtain the initial output feature matrix, including: Perform a linear transformation on the multi-head attention output to map it back to the same dimension to obtain a multi-head attention transformation matrix; Performing a residual connection between the multi-head attention transformation matrix and the enhanced feature matrix to obtain a first connection result; Performing layer normalization on the first connection result to obtain a normalized result; Inputting the normalized result into a feedforward network to obtain a feedforward output; Performing a residual connection between the normalized result and the feedforward output to obtain a second connection result; The second connection result is layer normalized to obtain the initial output feature matrix.
8. The health status prediction method based on multiple vital signs time series data according to claim 1, characterized in that: The step of dividing each of the vital sign time series data into time segments to obtain a plurality of segment data includes: Dividing each of the vital sign time series data into time segments to obtain a plurality of original segments; Local normalization processing is performed on each of the original segments to obtain a plurality of segment data.
Citation Information
Patent Citations
Intelligent bed management method and system
CN118136213A
Health monitoring method and system based on time sequence Transform and data fusion
CN119106392A