Private composite time series data generation

By using the Differential Privacy Generation Adversarial Network (DP-GAN) architecture, the generation of synthetic time series data that meets privacy metrics solves the privacy issues when sharing medical vertical time series data and maintains important characteristics of the data.

CN120226089APending Publication Date: 2025-06-27DEXCOM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380080222.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-25
Filing Date
2023-09-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Sharing patients’ medical longitudinal time series data faces serious legal and privacy issues, and a method to generate synthetic data to protect privacy while maintaining important features of the data.

Method used

Using a differential privacy generative adversarial network (DP-GAN) architecture, which includes a motif causal module and an autoencoder, generator and discriminator module, the generator module is trained to generate a series of synthetic time data that meets privacy metrics.

Benefits of technology

It realizes that while protecting patient privacy, the generated synthetic data can simulate important characteristics of the original time series data, and is suitable for practical applications that improve treatment development and technological advancements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226089A_ABST
    Figure CN120226089A_ABST
Patent Text Reader

Abstract

Methods and systems for generating synthetic data are provided. Longitudinal time series data is retrieved, and a neural network is trained based on the longitudinal time series data to generate composite time series data that satisfies a privacy metric. The longitudinal time series data is unmarked and univariate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 481,431, filed on January 25, 2023, which is assigned to the assignee of the present application and is hereby incorporated by reference in its entirety as if fully set forth herein and for all applicable purposes. Background of the Disclosure

[0003] The present disclosure relates to data processing systems. More particularly, the present disclosure relates to private synthetic time - series data generation for data processing systems.

[0004] Sharing medical longitudinal time - series data of patients enables improved treatment development and technological progress. For example, sharing measured analyte time - series data of patients can help understand associated disease mechanisms and the development of technologies for improving the quality of life of these patients. Not surprisingly, significant legal and privacy issues arise when sharing medical longitudinal time - series data of patients, such as those described by the Health Insurance Portability and Accountability Act of 1996 (referred to as HIPAA). Brief Description of the Drawings

[0005] Figure 1 A block diagram depicting an example system for generating synthetic data in accordance with an embodiment of the present disclosure.

[0006] Figure 2 A depiction of an example artificial neural network (ANN) in accordance with an embodiment of the present disclosure.

[0007] Figure 3A 、 Figure 3B 、 Figure 3C and Figure 3D Depictions of different views of an example recurrent neural network (RNN) in accordance with an embodiment of the present disclosure.

[0008] Figure 3E A depiction of an example data - flow diagram of a hidden recurrent module in accordance with an embodiment of the present disclosure.

[0009] Figure 4A A depiction of a view of an example long short - term memory (LSTM) network in accordance with an embodiment of the present disclosure.

[0010] Figure 4B and Figure 4C A depiction of an example data - flow diagram of an LSTM cell in accordance with an embodiment of the present disclosure.

[0011] Figure 5 A depiction of an example data - flow diagram of a differential privacy generative adversarial network (DP - GAN) in accordance with an embodiment of the present disclosure.

[0012] Figure 6 depicts an example loss function graph of the DP-GAN depicted in Figure 5 for training according to an embodiment of the present disclosure.

[0013] Figure 7 depicts an example data flow diagram of batch raw data for generating for training the DP-GAN depicted in Figure 5 according to an embodiment of the present disclosure.

[0014] Figure 8A and Figure 8B depicts an example data flow diagram for generating synthetic data by the DP-GAN depicted in Figure 5 according to an embodiment of the present disclosure.

[0015] Figure 9A depicts an example data flow diagram of a motif causality module according to an embodiment of the present disclosure.

[0016] Figure 9B depicts an example data flow diagram of a motif sequence block for generating for training the motif causality module depicted in Figure 9A according to an embodiment of the present disclosure.

[0017] Figure 10A depicts according to an embodiment of the present disclosure Figure 9A a data flow diagram of a motif network within the motif causality module depicted in

[0018] Figure 10B depicts an example data flow diagram of a neural network within the motif network for training the motif network depicted in Figure 10A according to an embodiment of the present disclosure.

[0019] Figure 11A depicts an example motif causality matrix according to an embodiment of the present disclosure.

[0020] Figure 11B depicts example motif time series data of two entries of the motif causality matrix according to an embodiment of the present disclosure.

[0021] Figure 12A depicts traditional time series data generation.

[0022] Figure 12B depicts motif causality time series data generation according to an embodiment of the present disclosure.

[0023] Figure 13 depicts a comparison of longitudinal time series data and synthetic time series data according to an embodiment of the present disclosure.

[0024] Figure 14 Depicted is a flow diagram representing functionality associated with generating synthetic data in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] A potential technical solution to the problem of sharing a patient's medical longitudinal time series data is to generate synthetic (fake) time series data based on the patient's original (real) time series data (such as, for example, the patient's measured glucose trajectory). However, the synthetic time series data must provide strong privacy guarantees and protect the privacy of the patient's medical longitudinal time series data while emulating certain important properties of the original time series data. Privacy guarantees refer to the degree to which sensitive data, such as a patient's medical data, is protected. The formal concept of strong privacy guarantees ensures that the probability of leaking sensitive data is very small (e.g., close to zero).

[0026] A variety of methods can be used to generate synthetic time series data, such as machine learning (ML) techniques, neural networks (NNs), artificial neural networks (ANNs), etc. These methods use training data (i.e., labeled data) that may include labels (which are the results or labeled parts of the trajectory that guides the generation of synthetic data), or additional information, such as multiple variables per time step (i.e., multivariate data), metadata, or auxiliary features (information calculated during model training). For example, a generative adversarial network (GAN) can be used to generate synthetic data based on original data. And, although GANs can be trained to generate synthetic time series data based on original time series data, these GANs do not inherently protect the privacy of the original time series data.

[0027] Synthetic time series data that protects the privacy of patients' medical longitudinal time series data can be publicly shared and integrated into many practical applications, such as, for example, blood glucose prediction, artificial pancreas systems, computer-based medical diagnostic methods, population-level medical research, etc.

[0028] Embodiments of the present disclosure advantageously provide a differentially private generative adversarial network (DP-GAN) architecture, which includes a motif causality module and an autoencoder, generator, and discriminator module. The autoencoder module includes an embedder module and a recovery module. Each module may include (in particular) one or more ANNs, such as RNNs, LSTM networks, etc., as described below.

[0029] In addition, embodiments of the present disclosure advantageously provide DP-GAN training methods that include raw data, motif data, and synthetic data processing techniques, integrated differential privacy metrics, and loss functions that characterize relationships between important motifs in raw time series data, as described below. A motif is a short ordered sequence of time steps from a time series (or trajectory) that characterizes important events in the time series data, such as peaks, valleys, etc. In the context of the present disclosure, motifs are not time-dependent and do not form recurring time patterns.

[0030] Importantly, certain embodiments of the present disclosure advantageously relate to training a DP-GAN using unlabeled and univariate raw data without any auxiliary (additional) information.

[0031] Figure 1 A block diagram of a system 100 for generating synthetic data in accordance with an embodiment of the present disclosure is depicted.

[0032] Generally speaking, system 100 includes a computer, server, etc. having one or more single-core or multi-core processors, specialized processors, etc., which are configured to train a neural network based on longitudinal time series data to generate synthetic time series data that satisfies privacy metrics.

[0033] More specifically, system 100 includes a computer 110 coupled to one or more networks 172, one or more I / O devices 182, and one or more displays 192. Computer 110 includes a bus 120 coupled to one or more processors 130, a storage element or memory 160, one or more communication interfaces 170, one or more I / O interfaces 180, and a display interface 190. In many embodiments, computer 110 also includes one or more specialized processors, such as, for example, a graphics processing unit (GPU) 140, a neural processing unit (NPU) 150, etc. Generally speaking, communication interface 170 is coupled to network 172 using a wired or wireless connection, I / O interface 180 is coupled to I / O device 182 using a wired or wireless connection, and display interface 190 is typically coupled to display 192 using a wired connection.

[0034] Bus 120 is a communication system that transfers data between processors 130, memory 160, communication interface 170, I / O interface 180, and display interface 190. In many embodiments, bus 120 also transfers data between these components and GPU 140 and / or NPU 150, as well as Figure 1 other components not depicted herein.

[0035] Processor 130 includes one or more general-purpose or special-purpose microprocessors that execute instructions to perform functions such as control, computing, input / output, etc. for computer 110. Each processor 130 may include a single integrated circuit such as a microprocessing device, or multiple integrated circuit devices and / or circuit boards that work together to achieve the appropriate functions. In addition, processor 130 may execute computer programs or modules stored in memory 160, such as operating system 162, software module 164, etc. For example, software module 164 may include a neural network, which includes one or more artificial neural networks (ANNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, convolutional neural networks (CNNs), etc.

[0036] Generally speaking, memory 160 stores data and instructions for processor 130 to execute. Memory 160 may include various non-transitory computer-readable media accessible by processor 130 and other components. In various embodiments, memory 160 may include volatile and non-volatile media, non-removable media, and / or removable media. For example, memory 160 may include any combination of random access memory (RAM), dynamic RAM (DRAM), static RAM (SRAM), read-only memory (ROM), flash memory, cache memory, and / or any other type of non-transitory computer-readable media.

[0037] Memory 160 includes various components for retrieving, presenting, modifying, and storing data 166. For example, memory 160 stores software module 164 that provides functions when executed by processor 130. Operating system 162 provides operating system functions for computer 110. Software module 164 provides various functions as described above. Data 166 may include data associated with operating system 162, software module 164, etc.

[0038] Communication interface 170 is configured to send data to and receive data from one or more networks 172 using one or more wired and / or wireless connections. Networks 172 may include one or more local area networks, wide area networks, the Internet, etc. that can execute various network protocols such as, for example, wired and / or wireless Ethernet, Bluetooth, etc. Networks 172 may also include various combinations of wired and / or wireless physical layers, such as, for example, copper wire or coaxial cable networks, fiber optic networks, Bluetooth wireless networks, WiFi wireless networks, CDMA, FDMA, and TDMA cellular wireless networks, etc.

[0039] The I / O interface 180 is configured to send and / or receive data to and / or from the I / O device 182. The I / O interface 180 enables connectivity between the processor 130, the memory 160, and the I / O device 182 by encoding data to be sent from the processor 130 or the memory 160 to the I / O device 182 and decoding data received from the I / O device 182 for the processor 130 or the memory 160. Generally, data can be sent via wired and / or wireless connections. For example, the I / O interface 180 may include one or more wired communication interfaces such as USB, Ethernet, etc., and / or one or more wireless communication interfaces coupled to one or more antennas such as WiFi, Bluetooth, cellular, etc.

[0040] Generally, the I / O device 182 provides input to and / or output from the computer 110. As described above, the I / O device 182 is operatively connected to the computer 110 using wired and / or wireless connections. The I / O device 182 may include a local processor coupled to a communication interface configured to communicate with the computer 110 using wired and / or wireless connections. For example, the I / O device 182 may include a keyboard, a mouse, a touchpad, a joystick, etc.

[0041] The display interface 190 is configured to send image data from the computer 110 to the monitor or display 192.

[0042] As described above, the software module 164 may include a neural network, which includes one or more ANNs, RNNs, LSTMs, etc.

[0043] The ANN uses a network of interconnected nodes to model the relationship between input data or signals and output data or signals, and this network of interconnected nodes is trained through a learning process. These nodes are arranged into various layers, including for example an input layer, one or more hidden layers, and an output layer. The input layer receives input data, such as for example image data, sensor time series data, etc., and the output layer generates output data, such as for example the probability that the image data contains known objects, medical conditions, etc. Each hidden layer provides at least a partial transformation of the input data to the output data. The DNN has multiple hidden layers in order to model the complex non-linear relationship between the input data and the output data.

[0044] In a fully connected feedforward ANN, each node is connected to all nodes in the previous layer and to all nodes in the subsequent layer. For example, each input layer node is connected to each hidden layer node, each hidden layer node is connected to each input layer node and each output layer node, and each output layer node is connected to each hidden layer node. Additional hidden layers are interconnected similarly. Each connection has a weight value, and each node has an activation function, such as for example a linear function, step function, sigmoid function, hyperbolic or tanh operation, rectified linear unit (ReLu) function, etc., which determines the output of the node based on the weighted sum of the inputs to the node. Input data propagates from the input layer nodes through the corresponding connection weights to the hidden layer nodes and then through the corresponding connection weights to the output layer nodes. For any given input, the sigmoid function and ReLu function output a number between 0 and 1, while the tanh operation outputs a number between -1 and 1.

[0045] More specifically, at each input node, the input data is provided to the activation function of that node, and then the output of that activation function is provided as the input data value to each hidden layer node. At each hidden layer node, the input data values received from each input layer node are multiplied by the corresponding connection weights, and the resulting products are summed or accumulated into an activation signal value, which is provided to the activation function of that node. Then, the output of that activation function is provided as the input data value to each output layer node. At each output layer node, the output data values received from each hidden layer node are multiplied by the corresponding connection weights, and the resulting products are summed or accumulated into an activation signal value, which is provided to the activation function of that node. Then, the output of that activation function is provided as the output data. Additional hidden layers can be configured similarly to process the data.

[0046] Figure 2 Depicted is an ANN 200 according to an embodiment of the present disclosure.

[0047] ANN 200 includes an input layer 210, one or more hidden layers, such as hidden layers 2101, 2202, ……, 220 N , and an output layer 230. The input layer 210 includes one or more input nodes, such as nodes I,1 , nodes I,2 , ……, nodes I,i . The hidden layer 2201 includes one or more hidden nodes, such as nodes 1,1 , nodes 1,2 , ……, nodes 1,j . The hidden layer 2202 includes one or more hidden nodes, such as nodes 2,1 , nodes 2,2 , ……, nodes 2,k。Hidden layer 220 N includes one or more hidden nodes, such as node N,1 , node N,2 , ……, node N,n . The output layer 230 includes one or more output nodes, such as node O,1 , node O,2 , ……, node O,o . In Figure 2 the depicted example, there are N hidden layers; the input layer 210 includes "i" nodes, the hidden layer 2301 includes "j" nodes, the hidden layer 2202 includes "k" nodes, the hidden layer 230 N includes "n" nodes, and the output layer 230 includes "o" nodes.

[0048] In some embodiments, N equals 3, "i" equals 3, "j", "k", and "n" equal 5 and "o" equals 3. The input nodes I,1 , node I,2 and node I,3 are each coupled to the hidden nodes 1,1 , node 1,2 , node 1,3 , node 1,4 and node 1,5 . The hidden nodes 1,1 , node 1,2 , node 1,3 , node 1,4 and node 1,5 are each coupled to the hidden nodes 2,1 , node 2,2 , node 2,3 , node 2,4 and node 2,5 . The hidden nodes 2,1 , node 2,2 , node 2,3 , node 2,4 and node 2,5 are each coupled to the hidden nodes 3,1 , node 3,2 , node 3,3 , node 3,4 and node 3,5 . The hidden nodes 3,1 , node 3,2 , node 3,3 , node 3,4 and node 3,5 are each coupled to the output nodes O,1 , node O,2 , node O,3 .

[0049] Many other variations of the input layer, hidden layer, and output layer are clearly possible, including hidden layers that are locally connected to each other (as opposed to fully connected).

[0050] Training an ANN involves optimizing the connection weights between nodes by minimizing the prediction error of the output data until the ANN achieves a specific accuracy level. One approach is backpropagation or backpropagation of error, which iteratively and recursively determines the gradient with respect to each weight (i.e., the partial derivative of the error function), and then adjusts each weight to improve the performance of the network.

[0051] A multi-layer perceptron (MLP) is a fully connected ANN with an input layer, an output layer, and one or more hidden layers. MLPs can be used to process time series data, such as natural language processing, machine translation, speech recognition, etc. Other ANNs include RNNs, LSTM networks, CNNs, etc.

[0052] Figure 3A Depicts a view of an RNN 300 according to an embodiment of the present disclosure.

[0053] Generally speaking, an RNN processes input sequence data and generates output sequence data, and can be used in many different applications, such as for example natural language processing applications (e.g., sentiment analysis, speech recognition, reading comprehension, summarization, and translation, etc.), image processing (e.g., image caption generation, video classification, etc.), etc. An RNN can be programmed to process many different types of input and output data, such as for example fixed input data and fixed output data for image classification, etc., fixed input data and sequential output data for image caption generation, etc., sequential input data and fixed output data for sentence "sentiment" classification, etc., sequential input data and sequential output data for machine translation, etc., synchronous sequential input data and sequential output data for video classification, etc.

[0054] The RNN 300 includes an input layer 310, one or more hidden layers such as a hidden recurrent layer 320, and an output layer 330. Generally speaking, an RNN can include one to four hidden recurrent layers; other numbers of hidden recurrent layers are also supported.

[0055] The input layer 310 includes one or more input nodes, such as nodes I,1 and nodes I,2 , which present the input data X as a sequence of input data values such as for example a sequence of letters, words, sentences, etc., a sequence of measured data values, a sequence of sensor data values, etc. to the hidden recurrent layer 320. Generally speaking, each sequence is a time step, and the input data is processed as a vector or matrix. The RNN 300 processes the input data values for each time step, and typically executes a loop to process the total number of time steps.

[0056] The hidden recurrent layer 320 is a fully-connected recurrent layer that includes hidden recurrent nodes, such as nodes R,1 , nodes R,2 , nodes R,3 , nodes R,4 , ……, nodes R,r . Each hidden recurrent node maintains or stores the state of the hidden state vector h of the layer, which is updated at each time step of the RNN 300. In other words, the hidden state vector h includes the states of each hidden recurrent node in the hidden recurrent layer 320. In many embodiments, the size of the hidden state vector h ranges from dozens or hundreds to thousands of elements, such as 64, 256, 4,096, etc. elements. In certain embodiments, the hidden state vector h may be subsampled to reduce processing requirements.

[0057] One or more additional, fully-connected hidden recurrent layers may follow the hidden recurrent layer 320. Each successive hidden recurrent layer includes hidden recurrent nodes and a corresponding hidden state vector h. The last hidden layer (e.g., Figure 3A the hidden recurrent layer 320 depicted in

[0058] ) presents the hidden state vector h to the output layer 330. O,1 . In certain embodiments, each output node provides an output, such as the probability of a predicted class score, word, sentence, etc., a predicted data value, a predicted correlation value, etc. A normalization function such as the Softmax function may be applied to the output of the output layer 330, or alternatively, by an additional fully-connected layer interposed between the last hidden layer and the output layer 330.

[0059] Figure 3B Depicts another view of the RNN 300 according to an embodiment of the present disclosure.

[0060] The input layer 310 is depicted as including a single element 310' of the input data X, the hidden recurrent layer 320 is depicted as including a single element, module, or unit 320' of the hidden state vector h, and the output layer 330 is depicted as including a single element 330' of the output data Y.

[0061] Figure 3C Depicts another view of the RNN 300 according to an embodiment of the present disclosure.

[0062] For Figure 3B the view of the RNN 300 depicted in t, the hidden recurrent module 320' includes a hidden state vector h t , and the output layer 330' includes output data Y t . In many embodiments, the input data X t is a vector having the same dimension as the hidden state vector h t .

[0063] Figure 3D depicts another view of the RNN 300 according to an embodiment of the present disclosure.

[0064] As noted above, the RNN 300 generally performs a loop such that the hidden recurrent module 320' can process the input data X and update the hidden state vector h at each time step. In Figure 3D the depicted view, the loop is "unrolled" and three time steps, namely t - 1, t, and t + 1, are shown. Thus, the RNN 300 can be regarded as a chain of repeating hidden recurrent modules or units 320'.

[0065] At time step t - 1, the input data X t-1 , the input hidden state vector h from the previous time step t-2 , the hidden state vector h t-1 , and the output data Y t-1 are shown. At time step t, the input data X t , the input hidden state vector h from the previous time step t-1 , the hidden state vector h t , and the output data Y t are shown. At time step t + 1, the input data X t+1 , the input hidden state vector h from the previous time step t , the hidden state vector h t+1 , and the output data Y t+1 are shown.

[0066] Generally speaking, the hidden state vector h can be updated by the following method t : applying an activation function f state to the sum of the product of the weight vector W t-1 multiplied by the hidden state vector h from the previous time step data and the product of the weight vector W t multiplied by the input data X c , as given by Equation 1.

[0067] h t = f c (W state · h t-1 + W data · X t ) Equation 1

[0068] Activation function f c can be a non - linear activation function, such as, for example, tanh(), ReLu, etc., which is applied to each element of the hidden state vector h. In some embodiments, a bias b can be added to the sum c before applying the activation function f c . The output data Y t is the product of the weight vector W output multiplied by the hidden state vector h t , as given in Equation 2.

[0069] Y t = W output · h t Equation 2

[0070] In some embodiments, an activation function f output can be applied to the product of the weight vector W t and the hidden state vector h o , such as, for example, tanh(), ReLu, etc., to generate the output data Y t , as given in Equation 3.

[0071] Y t = f o (W output · h t ) Equation 3

[0072] In some embodiments, a bias b can be added to the product o before applying the activation function f o .

[0073] Figure 3E FIG. 302 depicts a data - flow diagram of the hidden recurrent module 320' according to an embodiment of the present disclosure.

[0074] The hidden recurrent module 320' is shown at time step t. The hidden recurrent module 320' includes a tanh or sigmoid layer 322 that receives the hidden state vector h t-1 and the input data vector X t , applies a tanh operation to the sum of the product of the weight vector W state multiplied by the hidden state vector h from the previous time step t-1 and the product of the weight vector W data multiplied by the input data X t to generate the hidden state vector h t , as given in Equation 1. The hidden state vector h tIs output to the output layer 330 and provided to the next time step or stored for use in the next time step.

[0075] Similar to an ANN, training an RNN involves optimizing the weights by minimizing the prediction error of the output data until the RNN reaches a specific accuracy level. As noted above, backpropagation through time can be used to iteratively and recursively determine the gradient (i.e., the partial derivative of the error function) with respect to each weight, and then each weight is adjusted to improve the performance of the RNN. However, when the gradient of one or more of these weights becomes too small (i.e., when the gradient "vanishes"), these weights are not adjusted and training eventually stops. This problem is known as the vanishing gradient problem.

[0076] An LSTM network is a variant of an RNN that, among other advantages, solves the vanishing gradient problem by increasing the complexity of each hidden recurrent module or cell to generate and maintain more information than just the hidden state vector h (i.e., the cell state vector C). The LSTM network also avoids the long-term dependency problem of the RNN.

[0077] Figure 4A Depicts a view of an LSTM network 400 according to an embodiment of the present disclosure.

[0078] The LSTM network 400 generally also performs a loop such that the LSTM module or cell 420 can process each time step. In Figure 4A the view depicted, the loop is "unrolled" and three time steps are shown, similar to Figure 3D the view of the RNN 300 depicted in

[0079] At time step t-1, the input data X t-1 , the input hidden state vector h from the previous time step t-2 and the input cell state vector C t-2 , the hidden state vector h t-1 , the cell state vector C t-1 and the output data Y t-1 are shown. At time step t, the input data X t , the input hidden state vector h from the previous time step t-1 and the input cell state vector C t-1 , the hidden state vector h t , the cell state vector C t and the output data Y t are shown. At time step t+1, the input data X t+1 , the input hidden state vector h from the previous time step t and the input cell state vector Ct , the hidden state vector h t+1 , the cell state vector C t+1 and the output data Y t+1 .

[0080] Figure 4B FIG. 402 depicts a data flow diagram of the LSTM cell 420 according to an embodiment of the present disclosure.

[0081] The LSTM cell 420 is shown at time step t. The LSTM cell 420 receives or retrieves the cell state vector C from the previous time step t-1 and the hidden state vector h t-1 , receives the input data X at the current time step from the input layer 410 t , processes this data to generate the cell state vector C at the current time step t and the hidden state vector h t , sends the hidden state vector h at the current time step t to the output layer 430, and sends or stores the hidden state vector h t and the cell state vector C t for the next time step.

[0082] The LSTM cell 420 includes (in particular) cell storage (not shown for clarity), a forget gate 440, an input gate 450, an output gate 460, and a cell state update section 470. The LSTM cell 420 can be implemented by a software module, process, routine, etc., by a hardware component, circuit, etc., by a combination of a hardware component and a software component, etc.

[0083] The forget gate 440 determines which elements of the cell state vector C t-1 and the input data X t should be discarded (i.e., "forgotten") or retained (i.e., "remembered"). The input gate 450 generates new data to be added to the cell state vector C t-1 based on the hidden state vector h t-1 and the input data X t . The cell state update section 470 updates the cell state vector C t-1 based on the outputs of the forget gate 440 and the input gate 450 to generate the cell state vector C t-1 . The output gate 460 generates the hidden state vector h t based on the hidden state vector h t-1 , the input data X t and the updated cell state vector C t . t

[0084] Figure 4CDepicts the data flow diagram 404 of the LSTM cell 420 according to an embodiment of the present disclosure.

[0085] The LSTM cell 420 is shown at time step t. The hidden state vector h t-1 and the input data vector X t are provided to the forget gate 440, the input gate 450, and the output gate 460. In many embodiments, the concatenation operation 422 concatenates the hidden state vector h t-1 and the input data vector X t to form the concatenated input vector [h t-1 , X t , which is provided to the forget gate 440, the input gate 450, and the output gate 460. In other embodiments, the hidden state vector h t-1 and the input data vector X t are provided separately to the forget gate 440, the input gate 450, and the output gate 460.

[0086] The forget gate 440 includes a sigmoid layer 442 that receives the concatenated input vector [h t-1 , X t from the concatenation operation 422, applies the concatenated weight vector W t-1 to the concatenated input vector [h t to generate the weighted concatenated input vector W f · [h f , X t-1 , and applies the sigmoid function to the weighted concatenated input vector W t · [h f , X t-1 to generate the activation vector f t , as given by Equation 4. The concatenated weight vector W t is the concatenation of the weight vector of the hidden state vector h f and the weight vector of the input data vector X t-1 and the weight vector of the input data vector X t .

[0087] f t = σ (W f · [h t-1 , X t ) Equation 4

[0088] In some embodiments, a bias b f may be added to the weighted concatenated input vector W t-1 · [h t before applying the sigmoid function σ. The sigmoid layer 442 outputs the activation vector f f . tProvided to the element-wise multiplication operation 476 within the unit state update segment 470.

[0089] The input gate 450 includes a sigmoid layer 452, a tanh layer 454, and an element-wise multiplication operation 456. The sigmoid layer 452 receives the concatenated input vector [h t-1 ,X t from the concatenation operation 422, applies the concatenated weight vector W t-1 ,X t to the concatenated input vector [h i to generate the weighted concatenated input vector W i ·[h t-1 ,X t , and applies the sigmoid function to the weighted concatenated input vector W i ·[h t-1 ,X t to generate the activation vector i t , as given in Equation 5.

[0090] i t = σ (W i · [h t-1 ,X t ) Equation 5

[0091] In some embodiments, a bias b i can be added to the weighted concatenated input vector W t-1 ,X t before applying the sigmoid function σ. The sigmoid layer 452 provides the activation vector i i to the element-wise multiplication operation 456. t

[0092] The tanh layer 454 receives the concatenated input vector [h t-1 ,X t from the concatenation operation 422, applies the concatenated weight vector W t-1 ,X t to the concatenated input vector [h C to generate the weighted concatenated input vector W C ·[h t-1 ,X t , and applies the tanh operation to the weighted concatenated input vector W C ·[h t-1 ,X t to generate the activation vector as given in Equation 6.

[0093]

[0094] ​In some embodiments, the weighted cascade input vector W may be input to the weighted concatenation before applying the tanh operation C ·[h t-1 ,X t ]Add deviation b C The tanh layer 454 converts the activation vector is provided to the element-wise multiplication operation 456. The element-wise multiplication operation 456 activates the vector i t With the activation vector The multiplication is performed to generate an intermediate product which is provided to an element-wise addition operation 478 within the cell state update stage 470 .

[0095] Output gate 460 includes sigmoid layer 462, element-wise multiplication operation 466, and element-wise tanh operation 464. Sigmoid layer 462 receives the concatenated input vector [h t-1 ,X t ], to the cascade input vector [h t-1 ,X t ] Apply the cascade weight vector W o To generate the weighted cascade output vector W o ·[h t-1 ,X t ], and outputs the vector W to the weighted cascade o ·[h t-1 ,X t ] Apply the sigmoid function to generate the activation vector o t , as given in Equation 7.

[0096] o t = σ (W o · [h t-1 ,X t ]) Equation 7

[0097] In some embodiments, the input vector W may be fed to the weighted cascade before applying the sigmoid function σ o ·[h t-1 ,X t ]Add deviation b o The sigmoid layer 462 activates the vector o t Provided to element-wise multiplication operation 466.

[0098] The tanh operation 464 receives the cell state vector Ct, applies the tanh operation to the cell state vector Ct, and provides the result to the element-wise multiplication operation 466, which multiplies the output of the sigmoid layer 462 and the tanh operation 464 to generate a hidden state vector h t , as given in Equation 8.

[0099] h t = o t · tanh (C t ) Equation 8

[0100] The element-wise multiplication operation 476 within the cell state update section 470 receives the cell state vector C t-1 and multiplies the activation vector f t by the cell state vector C t-1 to generate an intermediate vector product, which is provided to the element-wise addition operation 478. The intermediate vector products generated by the element-wise multiplication operation 476 and the element-wise multiplication operation 456 are added together to generate the cell state vector C t as given by Equation 9.

[0101]

[0102] The hidden state vector h t is output to the output layer 430. The hidden state vector h t and the cell state vector C t are provided to the next time step or stored for use at the next time step.

[0103] Figure 5 Depicts a data flow diagram 502 of the DP-GAN 500 according to an embodiment of the present disclosure.

[0104] As noted above, GANs can be used to generate synthetic data based on raw data. Generally speaking, a GAN includes a generator neural network and a discriminator neural network. The generator neural network learns from the raw data and works to generate synthetic data. The discriminator neural network receives samples of both raw (real) data and synthetic (fake) data and "guesses" whether each sample is real or fake. The generator neural network and the discriminator neural network are trained adversarially, i.e., trained against each other. The generator neural network attempts to deceive the discriminator neural network into believing that the synthetic data is real, and the discriminator neural network attempts to become very good at guessing which samples are actually real or fake. When the training is successful, the generator neural network becomes very good at generating synthetic data that deceives the discriminator neural network into believing it is real.

[0105] Differential privacy is a formal privacy concept that limits the risk faced by anyone who provides data for subsequent processing. In a DP-GAN, noise is extracted from a carefully designed distribution and applied to the weights of the generator neural network and the discriminator neural network to protect the privacy of individuals associated with the data. From one perspective, the addition of noise prevents the generator and discriminator neural networks of the DP-GAN from memorizing or leaking any sensitive or personal data from the raw data.

[0106] The DP-GAN 500 includes a motif causality module 510, an autoencoder module 520, a generator module 530, a discriminator module 540, a preprocessor module 526, and a postprocessor module 528. The autoencoder module 520 includes an embedder module 522 and a recovery module 524. Each module may include, in particular, one or more neural networks, such as RNNs, LSTM networks, etc., which are implemented as software modules, processes, routines, etc. In many embodiments, these software components may be executed by the processor 130. In certain embodiments, at least a portion of these software components may be executed by the GPU 140 or the NPU 150. In other embodiments, these software components may be executed by a combination of the processor 130 and the GPU 140, the processor 130 and the NPU 150, or the processor 130, the GPU 140, and the NPU 150.

[0107] In many embodiments, the medical longitudinal time series data (trajectories) of a patient (human) population may be divided into two sets, namely set A and set B. Each set includes multiple trajectories of different patient (human) groups within the population. Set A may include the same number of trajectories as set B, set A may include fewer trajectories than set B, or set A may include more trajectories than set B. In certain embodiments, the medical longitudinal time series data may be completely divided into set A and set B, while in other embodiments, the medical longitudinal time series data may be partially divided into set A and set B based on a selection criterion. For example, the medical longitudinal time series data of patients may be selected based on quality. For example, the unselected medical longitudinal time series data of patients may exhibit undesirable characteristics, such as measurement noise, data loss, etc. Generally speaking, the medical longitudinal time series data of patients is bounded and includes at least 50 time steps. Some medical longitudinal time series data may include 100 or more time steps. For example, 24-hour continuous glucose monitoring (CGM) data has 288 time steps (i.e., 12 measurements per hour).

[0108] The original longitudinal time series data (set A) 550 is provided to the motif causality module 510 as the original data x m (data 553). The original longitudinal time series data (set B) 560 is provided to the preprocessor module 526 as the original data x (data 563).

[0109] The motif causality module 510 includes a data processing module (not shown for clarity), which processes the original data x mto generate a plurality of non-overlapping motif data partitions. Each motif data partition is provided to a different motif network 512. Each motif network 512 generates a motif causality matrix M i , and these motif causality matrices are aggregated into an aggregated motif causality matrix M A and provided to the generator module 530. During training, the motif networks 512 learn the relationships between these motifs and express these relationships in the causality matrices, which are aggregated into the aggregated motif causality matrix M A to protect patient privacy. The motif causality module 510 is discussed in more detail below.

[0110] The preprocessor module 526 preprocesses the raw data x to generate batched raw data x and provides the batched raw data x to the embedder module 522.

[0111] The embedder module 522 reduces the dimensionality of the batched raw data x to generate a set of embedded trajectories, i.e., embedded raw data x e (data 523). In many embodiments, the embedder module 522 includes a neural network having an input layer, at least one hidden layer such as, for example, an RNN layer (e.g., hidden recurrent layer 320, hidden recurrent module 320', etc.) or an LSTM layer (e.g., LSTM cell 420, etc.), and an output layer. In certain embodiments, the neural network can be a CNN. Other neural network architectures are also supported.

[0112] The generator module 530 generates embedded synthetic data e based on the embedded raw data x A and the aggregated motif causality matrix 513 (M ).

[0113] The postprocessor module 528 reconstructs the embedded synthetic data in the raw data space to generate synthetic data (data 573), which can be output as synthetic longitudinal time series data 570.

[0114] The recovery module 524 reconstructs the embedded raw data x in the raw data space e to generate recovered raw data (Data 525). In many embodiments, the recovery module 524 includes a neural network having an input layer, at least one hidden layer such as, for example, an RNN layer (e.g., hidden recurrent layer 320, hidden recurrent module 320', etc.) or an LSTM layer (e.g., LSTM cell 420, etc.), and an output layer. In certain embodiments, the neural network can be a CNN. Other neural network architectures are also supported.

[0115] The discriminator module 540 receives the embedded original data x e and guesses whether the embedded original data x e is real or fake. Similarly, the discriminator module 540 also receives the embedded synthetic data and guesses whether the embedded synthetic data is real or fake. These guesses can be output as the embedded original guess 543 and the embedded synthetic guess 545. In many embodiments, the discriminator module 540 includes a neural network having an input layer, at least one hidden layer such as, for example, an RNN layer (e.g., hidden recurrent layer 320, hidden recurrent module 320', etc.) or an LSTM layer (e.g., LSTM cell 420, etc.), and an output layer. Other neural network architectures are also supported.

[0116] During training, carefully calibrated noise (not shown for clarity) is added to the weights of the autoencoder module 520 (i.e., the embedder module 522 and the recovery module 524), the generator module 530, and the discriminator module 540 to ensure that each network maintains differential privacy (e.g., satisfies a privacy metric) to protect patient privacy. To generate synthetic data, the weight noise generator 580 generates weight noise (Z) 583, which is received as an input by the generator module 530 and passed through the recovery module 524 to the post-processor module 528, which outputs the final synthetic data. In many embodiments, the weight noise (Z) 583 is a random noise vector.

[0117] Using the embedded original data x e and the embedded synthetic data (instead of the original data x and the synthetic data ) to train the generator module 530 and the discriminator module 540. By reducing the dimension of the space in which the generator module 530 and the discriminator module 540 learn, these networks focus on and learn the most important parts or motifs of the trajectories.

[0118] Figure 6 Depicts a loss function graph 600 for training the Figure 5 DP-GAN 500 depicted in

[0119] In many embodiments, the modules of the DP-GAN 500 may be trained in a specific order. First, the motif causality module 510 is trained using the raw data x m (data 553) to generate an aggregated motif causality matrix 513 (M A ). Then (e.g., within each pass), the remaining modules of the DP-GAN 500 are trained in sequence, using the raw data x (data 563) and the aggregated motif causality matrix 513 (M A ) to generate synthetic data In some embodiments, the autoencoder module 520 is trained, then the generator module 530 and discriminator module 540 are trained adversarially, and then the embedding module 522 of the autoencoder module 520 is trained a second time.

[0120] In many embodiments, six loss functions are used to train the autoencoder module 520, generator module 530, and discriminator module 540, including a reconstruction loss (L R ) 610, a step loss (L S ) 620, a distribution loss (L D ) 630, a motif causality loss (L M ) 640, an adversarial loss false (L Af ) 650, and an adversarial loss true (L Ar ) 660. Other loss functions and subsets of these loss functions are also supported.

[0121] The reconstruction loss (L R ) 610 is the root mean square error (RMSE) between the raw data x and the recovered raw data . A "perfect" autoencoder perfectly reconstructs the raw data, such that

[0122] The step loss (L S ) 620 is the mean squared error (MSE) between a batch of the raw data x e embedded and a batch of the synthetic data embedded. The generator module 530 uses the step loss (L S ) 620 to compare and learn to correct the differences between the step data distributions. In other words, the generator module 530 learns to better generate batches of data for the next time step by looking at the differences between the next step it generates and the true next step.

[0123] The distribution loss (L D ) 630 is the moment loss between the distribution of the raw data x and the distribution of the synthetic data . The generator module 530 uses the distribution loss (L D)630 learns to generate different sets of trajectories instead of repeatedly generating the same type of trajectories.

[0124] The motif causality loss (L M )640 is the mean squared error (MSE) between the motif causality matrix M x calculated from the original data and the motif causality matrix calculated from the synthetic data. The generator module 530 calculates the motif causality matrix after the set of embedded synthetic data is returned by the recovery module 524 and the post-processor module 528 to generate synthetic data in the original space and uses the motif causality loss (L )640 to learn to generate synthetic data that produces a true causal matrix (thereby identifying appropriate causal relationships from motifs) and implicitly learns not to generate untrue motif sequences. M )640 to learn to generate synthetic data that produces a true causal matrix (thereby identifying appropriate causal relationships from motifs) and implicitly learns not to generate untrue motif sequences.

[0125] The adversarial loss fake (L Af )650 is the binary cross entropy (BCE) between the discriminator's guess of the synthetic data (i.e., the embedded synthetic guess 545) and the ground truth (i.e., the all-ones vector).

[0126] The adversarial loss real (L Ar )660 is the BCE between the discriminator's guess of the original data x (i.e., the embedded original guess 543) and the ground truth (i.e., the all-zeros vector).

[0127] The autoencoder module 520 is trained to minimize a weighted combination of the reconstruction loss (L R )610 and the step loss (L S )620 (α is a weight hyperparameter), as given in Equation 10, to avoid overspecialization.

[0128]

[0129] The generator module 530 is trained to minimize a weighted combination of the step loss (L S )620, the distribution loss (L D )630, the motif causality loss (L M )640, and the adversarial loss fake (L Af )650 (η is a weight hyperparameter), as given in Equation 11. The step loss (L S )620 enables the dual training of the autoencoder module 520 and the generator module 530.

[0130]

[0131] The discriminator module 540 is trained to minimize a weighted combination of the adversarial loss fake (L Af ) 650 and the adversarial loss real (L Ar ) 660, as given in Equation 12.

[0132]

[0133] In one embodiment, α is 0.1 and η is 10; other values are also supported. These training objectives and loss functions train the DP-GAN 500 to generate high-quality, long-time series synthetic data.

[0134] Figure 7 Depicts a data stream 700 for generating batch raw data for training the DP-GAN 500 depicted in Figure 5 in accordance with an embodiment of the present disclosure.

[0135] In Figure 7 the exemplary embodiment depicted, the raw data 710 includes 100 trajectories, namely trajectories 7121, ……, 712 100 , and each trajectory includes 288 data values (time steps). For example, trajectory 7121 includes the raw data X 1,1 , ……, X 1,288 and so on; trajectory 712 100 includes the raw data X 100,1 , ……, X 100,288 . In many embodiments, the preprocessor module 526 preprocesses the raw data x (data 563) to generate batch raw data x (as discussed above).

[0136] The preprocessor module 526 applies a sliding window (width 24, step size 1) to each trajectory in the raw data 710 to expand each trajectory into a batch data slice including 264 time chunks, each time chunk including 24 values (time steps). For trajectory 7121, the sliding window is applied to the first 24 data values, namely the data value sequence 7141, to generate the time chunk 7241 of the batch data slice 7221, which includes X 1,1 , X 1,2 , ……, X 1,23 , X 1,24 . Then, the sliding window is moved one data value position to the right and applied to the next 24 data values, namely the data value sequence 7142, to generate the time chunk 7242 of the batch data slice 7221, which includes X 1,2 , X 1,3 , ……, X 1,24 , X 1,25 . And so on. The data value sequence 714 263Time strip 724 for generating batch data slice 7221 263 , the time strip includes X 1,263 、X 1,264 、……、X 1,286 、X 1,287 , and data value sequence 714 264 Time strip 724 for generating batch data slice 7221 264 , the time strip includes X 1,264 、X 1,265 、……、X 1,287 、X 1,288 . The remaining trajectories are processed in a similar manner; finally, trajectory 712 100 generates batch data slice 722 100 . Batch raw data 720 includes all batch data slices 722 i , that is, batch data slice 7221, ……, batch data slice 722 100 . Other methods for generating batch raw data 720 are also supported.

[0137] The embedder module 522 then reduces the dimension of the batch raw data x to generate a set of embedded trajectories, i.e., embedded raw data x e (data 523). In an exemplary embodiment, the embedder module 522 can reduce the number of strips in each batch data slice from 264 to 128 to generate embedded raw data x e .

[0138] Figure 8A and Figure 8B depicts a data stream 800 for generating synthetic data 830 by the DP-GAN 500 depicted in Figure 5 .

[0139] In many embodiments, the post-processor module 528 reconstructs the embedded synthetic data (data 533) in the raw data space to generate synthetic data (data 573), and the synthetic data can be output as synthetic longitudinal time series data 570 (as described above).

[0140] In Figure 8A and Figure 8B shown exemplary embodiments, the embedded synthetic data 810 includes 100 trajectories, i.e., trajectories 8121, ……, 812 100 , each trajectory includes 128 time strips, and each time strip includes 24 data values (time steps). For example, trajectory 8121 includes time strips 8241, ……, time strip 824 128; The time bar 8241 includes the embedded synthetic data X h 1,1 ,..., X h 1,24 The time bar 824 128 includes the embedded synthetic data X h 128,1 ,..., X h 128,24 (In Figure 8A and Figure 8B , X h represents ).

[0141] The post-processor module 528 first serializes each track 812 of the embedded synthetic data 810 i into a single row of the reformed embedded synthetic data 820. For example, by first placing the time bar 8241 in the first row of the reformed embedded synthetic data 820, placing the time bar 8242 in the first row of the reformed embedded synthetic data 820 and after the time bar 8241, and so on, until the time bar 824 128 is placed in the first row of the reformed embedded synthetic data 820 and after the time bar 824 127 thereafter, the track 8121 is formed into a serialized track 8221 to complete the formation of the serialized track 8221. The indices of the elements of the serialized track 8221 are shown to change from the index based on the time bar / time step (e.g., X h 1,1 ,..., X h 1,24 , X h 128,1 ,..., X h 128,24 etc.) to the index based on the track / time step (e.g., X h 1,1 ,..., X h 1,3072 , X h 100,1 ,..., X h 100,3072 etc.). The remaining tracks 812 of the embedded synthetic data 810 i are serialized in a similar manner, ending when forming the serialized track 822 100 from the track 812 100 .

[0142] The post-processor module 528 then applies to each serialized track 822 of the reformed embedded synthetic data 820 iApply a reverse sliding window (i.e., a moving average) to reconstruct the embedded synthetic data in the original space of 100 trajectories, each trajectory having 288 data values (time steps). Generally speaking, the reverse sliding window averages the groups of data values in each serialized trajectory 822 i to generate each synthetic trajectory 832 i . For example, by applying the reverse sliding window to the data values X h 1,1 ,..., X h 1,3072 to generate the data values X h 1,1 ,..., X h 1,288 , the serialized trajectory 8221 is formed into the synthetic trajectory 8321. And so on. Finally, by applying the reverse sliding window to the data values X h 100,1 ,..., X h 100,3072 to generate the data values X h 100,1 ,..., X h 100,288 , the serialized trajectory 822 100 is formed into the synthetic trajectory 832 100 . The synthetic data 830 includes the synthetic trajectories 832 1, ,..., 832 100 . Other methods for generating the synthetic data 830 are also supported.

[0143] Figure 9A Depicts a data flow diagram 900 of the motif causality module 510 according to an embodiment of the present disclosure.

[0144] The motif causality module 510 includes a data processing module 910, multiple (N) motif networks 5121, 5122,..., 512 N and a motif causality matrix aggregation module 940. During the training of the causality module 510, the data processing module 910 generates multiple (N) non-overlapping motif data partitions 9201, 9202,..., 920 m from the original data x N . Each motif data partition 920 i includes data of different patients from the original longitudinal time series data (set A) 550. Each motif network 5121, 5122,..., 512 N receives different motif data partitions 920 i, that is, motif network 5121 receives motif data partition 9201, motif network 5122 receives motif data partition 9202, and so on.

[0145] Each motif network 512 i generates a motif causality matrix 930 i based on the corresponding motif data partition 920 i (M i ). That is, motif network 5121 generates motif causality matrix 9301 (M1) based on motif data partition 9201, motif network 5122 generates motif causality matrix 9302 (M2) based on motif data partition 9202, and so on. As pointed out above, each motif causality matrix 930 i (M i ) represents the relationships between motifs learned by motif network 512 i during training. Generally speaking, motif causality matrix 930 i (M i ) includes motif causality values c and has a width less than or equal to m and a height less than or equal to m (where m is the number of motifs to be analyzed in the data partition, as discussed below). Each causality factor c j,k represents the strength of the relationship between two motifs (e.g., motif j and motif k), and can have a value between 0 (i.e., indicating a weak relationship) and 1 (i.e., indicating a strong relationship). Other values are also supported.

[0146] Motif causality matrix aggregation module 940 aggregates motif causality matrices 9301 (M1), 9302 (M2), ……, 930 N (M N ) into an aggregated motif causality matrix 513 (M A ) to protect patient privacy, that is, to meet the privacy metric. Aggregated motif causality matrix 513 (M A ) is provided to the generator module 530 during the training of the generator module 530, so that the generator module 530 focuses on preserving important motifs (events) within the trajectory of the embedded original data x e .

[0147] Figure 9B Depicts a data flow diagram 902 for generating batch motif data 960 for training the motif causality module 510 depicted in Figure 9A according to an embodiment of the present disclosure.

[0148] After data processing module 910 generates each motif data partition 9201, 9202, ……, 920 N , data processing module 910 further processes each motif data partition 9201, 9202, ……, 920N to generate corresponding batch motif data 960.

[0149] In Figure 9B the exemplary embodiment depicted, the motif data partition 920 includes 100 trajectories, and each trajectory includes 288 data values (time steps). For example, the first trajectory includes the raw data X 1,1 , ……, X 1,288 , and so on; the last trajectory includes the raw data X 100,1 , ……, X 100,288 . Other numbers of trajectories and other numbers of data values (time steps) are also supported.

[0150] The motif data partition 920 can theoretically be divided into multiple motif blocks, one motif block for each motif to be analyzed. Three motif blocks are depicted, and each motif block includes 96 data values (time steps) of each trajectory, i.e., the motif block 921 of motif 1, the motif block 922 of motif 2, and the motif block 923 of motif 3. The motif block 921 includes the data values X 1,1 , ……, X 1,96 , ……, X 100,1 , ……, X 100,96 . The motif block 922 includes the data values X 1,97 , ……, X 1,192 , ……, X 100,97 , ……, X 100,192 . The motif block 923 includes the data values X 1,193 , ……, X 1,288 , ……, X 100,193 , ……, X 100,288 . Although the motif data partition 920 can be divided into 2 motif blocks, the motif data partition 920 is typically divided into 3 or more motif blocks.

[0151] The data processing module 910 divides the motif blocks 921, 922, and 933 into separate motif blocks for each trajectory, and then stacks these separate motif blocks into a motif block stack 950. For the first trajectory, the motif block 9211 includes the data values X 1,1 , ……, X 1,96 , the motif block 9221 includes the data values X 1,97 , ……, X 1,192 , the motif block 9231 includes the data values X 1,193 , ……, X 1,288 . And so on. For the last trajectory, the motif block 921 100 includes the data values X 100,1 , ……, X 100,96 , the motif block 922 100 includes the data values X 100,97 , ……, X100,192 , and motif block 923 100 includes data values X 100,193 、……、X 100,288 . Thus, the motif block stack 950 includes 300 motif blocks.

[0152] Then, the data processing module 910 applies a sliding window (width 24, step size 1) to each motif block in the motif block stack 950 to expand each motif block into a motif sequence block including 72 overlapping motif sequences, each motif sequence including 24 data values (time steps).

[0153] For motif block 9211, the sliding window is applied to the first 24 data values to generate motif sequence 9641 of motif sequence block 9621, which includes X 1,1 、X 1,2 、……、X 1,23 、X 1,24 . Then the sliding window is moved one data value position to the right and applied to the next 24 data values to generate motif sequence 9642 of motif sequence block 9621, which includes X 1,2 、X 1,3 、……、X 1,24 、X 1,25 . And so on. For example, motif sequence 964 71 includes X 1,71 、X 1,72 、……、X 1,94 、X 1,95 , motif sequence 964 72 includes X 1,72 、X 1,73 、……、X 1,95 、X 1,96 .

[0154] For motif block 9221, the sliding window is applied to the data values X 1,97 、……、X 1,192 to generate motif sequence block 9622 (not shown for clarity). For motif block 9231, the sliding window is applied to the data values X 1,193 、……、X 1,288 to generate motif sequence block 9623 (not shown for clarity). And so on for the remaining motif blocks. For example, for motif block 923 100 , the sliding window is applied to the data values X 100,193 、……、X 100,288 to generate motif sequence block 962 300 . Other methods for generating batch motif data 960 are also supported.

[0155] Figure 10ADepicts a motif network 512 according to an embodiment of the present disclosure i of the data flow diagram 1000.

[0156] The motif network 512 i includes a plurality (m) of neural networks 10101, 10102, ……, 1010 m and a weight combination module 1030. The number m is the number of motifs being analyzed, as described above. The neural network 10101 includes a weight matrix 10201 (W1), the neural network 10102 includes a weight matrix 10202 (W2), and so on. The neural network 1010 m includes a weight matrix 1020 m (W m ).

[0157] Use the motif data partition 920 i to train each neural network 1010 j , as described below. The weight combination module 1030 linearly combines the weight matrices 10201, 10202, ……, 1020 m to generate a motif causality matrix 930 i . Generally speaking, each weight matrix 1020 j includes weights w and has a width equal to the sliding window width (e.g., 24 time steps) and a height equal to m.

[0158] Figure 10B Depicts a data flow diagram 1002 for training the neural network 1010 within the motif network 512 depicted in Figure 10A according to an embodiment of the present disclosure i j .

[0159] In some embodiments, a loss module 1040 and a weight adjustment module 1050 may be provided for each neural network 1010 within the motif network 512 i . In other embodiments, a loss module 1040 and a weight adjustment module 1050 may be provided for the motif network 512 j and used to train each neural network 1010 i . j

[0160] In many embodiments, the neural network 1010 j includes an input layer, at least one hidden layer such as, for example, an RNN layer (e.g., a hidden recurrent layer 320, a hidden recurrent module 320', etc.) or an LSTM layer (e.g., an LSTM cell 420, etc.), and an output layer. In some embodiments, a convolutional layer may be before the output layer. Other neural network architectures are also supported. ​​

[0161] Generally speaking, neural network 1010 j is trained with respect to a specific "ground truth" motif sequence block 962 within batch motif data 960 j to learn the causal relationship between the ground truth motif sequence block 962 j and all other motif sequence blocks within batch motif data 960. More specifically, neural network 1010 j generates a predicted motif sequence block 1062 based on batch motif data 960. Loss module 1040 determines whether the weights of neural network 1010 j should be adjusted by comparing the predicted motif sequence block 1062 with the ground truth motif sequence block 962 using a loss function such as, for example, MSE, RMSE, etc. j (W j ).

[0162] Figure 11A Depicts a motif causality matrix 1100 according to an embodiment of the present disclosure.

[0163] The motif causality matrix 1100 is a 10×10 matrix presenting motif causality values for 100 pairs of motifs. The X-axis 1102 includes 10 motif bins, the Y-axis 1104 includes 10 motif bins, and the scale 1106 ranges from 0 (i.e., no causal relationship between motifs) to 1 (strong causal relationship between motifs). For example, motif causality element 1120 has a value of 0.382 and indicates a certain degree of causal relationship between motif 100 (i.e., bin 5 on the X-axis) and motif 281 (i.e., bin 7 on the Y-axis). Motif causality element 1120 has a value of 0.424 and indicates a slightly higher causal relationship between motif 140 (i.e., bin 6 on the X-axis) and motif 297 (i.e., bin 9 on the Y-axis).

[0164] Figure 11B Depicts motif comparisons 1120 and 1130 for motif causality elements 1108 and 1110, respectively, according to an embodiment of the present disclosure.

[0165] Motif comparison 1120 includes a graph 1122 depicting time series data 1124 of motif 100 (i.e., glucose values versus time), a graph 1126 depicting time series data 1128 of motif 281 (i.e., glucose values versus time), and a motif causality element 1108 having a value of 0.382. Similarly, motif comparison 1130 includes a graph 1132 depicting time series data 1134 of motif 140 (i.e., glucose values versus time), a graph 1136 depicting time series data 1138 of motif 297 (i.e., glucose values versus time), and a motif causality element 1110 having a value of 0.424.

[0166] For medical longitudinal time series data, the sequence of important motifs is more informative than each individual previous time step. Traditional time series data generation methods (such as autoregressive models) assume that the time series depends on all previous time steps within a window and generate the value of x at time t based on the sequence of previous values of x. Importantly, these methods only preserve temporal relationships (e.g., information from previous time steps within the window) while ignoring any other potential information relationships within the same time series (e.g., x) or between different time series.

[0167] Figure 12A Depicts traditional time series data generation 1200.

[0168] The value of x at time step t (i.e., x value 1206) depends on the value of x at time step t - 1 (i.e., x value 1201), the value at t - 2 (i.e., x value 1202), the value at t - 3 (i.e., x value 1203), the value at t - 4 (i.e., x value 1204), and the value at t - 5 (i.e., x value 1205). Although the previous values of x can be weighted in a linear combination, subjected to dropout, etc., traditional methods strongly depend on the window size and ignore long-term relationships between different time series.

[0169] Figure 12B Depicts motif-causal time series data generation 1210 according to an embodiment of the present disclosure.

[0170] As Figure 12B shown, the value of x4 at time step t (i.e., x4 value 1220) depends on the value of x1 at time t - 3 (i.e., x1 value 1213) and the value at time t - 5 (i.e., x1 value 1215), the value of x2 at time t - 1 (i.e., x2 value 1221) and the value at time t - 2 (i.e., x2 value 1222), and the value of x3 at time t - 4 (x1 value 1234). Similarly, the value of x5 at time step t (i.e., x5 value 1230) depends on the value of x1 at time t - 1 (i.e., x1 value 1211), the value of x2 at time t - 4 (i.e., x2 value 1224), and the value of x3 at time t - 2 (i.e., x3 value 1232) and the value at time t - 3 (i.e., x3 value 1233).

[0171] Advantageously, motif-causal time series data generation 1210 only uses previous lags with causal influence, finds relationships between motifs from different time series, and allows the DP-GAN 500 to learn relationships (patterns) between important event sequences contributing to the time series construction in the trajectory.

[0172] This is particularly advantageous for long time series because the network can easily become overwhelmed when trained to learn from each previous time step. Instead, by only saving relationships associated with important motif sequences, the DP-GAN 500 learns the actual time step sequence in the output trajectory more quickly.

[0173] For example, for glucose trajectories, for predicting the next glucose value at time t, large glucose peaks (e.g., hyperglycemic events) six or more time steps in the past (e.g., earlier than t - 6) are more informative than the five immediately preceding time steps (e.g., t - 1, t - 2, t - 3, t - 4, t - 5). This is due to the strong influence of the event (e.g., knowing that the glucose value will necessarily fall back from the peak, regardless of whether the previous glucose value dropped from 330 to 329 or from 290 to 289). Thus, patterns in these types of events can be exploited (e.g., if a large peak motif is observed, it is known that a decreasing slope motif will follow).

[0174] Figure 13 A comparison 1300 of longitudinal time series data and synthetic time series data in accordance with an embodiment of the present disclosure is depicted.

[0175] The longitudinal time series data 1310 includes glucose measurements (mg / dL) at 288 time steps. The synthetic time series data 1320 includes synthetic glucose values (mg / dL) at 288 time steps generated by the DP-GAN 500. As can be seen from samples of the synthetic trajectories, the patterns in the trajectories look very realistic (e.g., having a very similar overall structure to the real trajectories in terms of sequences of peaks, valleys, etc.).

[0176] Figure 14 A flowchart 1400 representing functions associated with generating synthetic data in accordance with an embodiment of the present disclosure is depicted.

[0177] At 1410, longitudinal time series data is received. In many embodiments, the longitudinal time series data is unlabeled and univariate.

[0178] As described above, the longitudinal time series data can be medical longitudinal time series data. Generally speaking, a patient's medical longitudinal time series data is bounded and includes at least 50 time steps. Some medical longitudinal time series data can include 100 or more time steps, such as, for example, 24-hour continuous glucose monitoring (CGM) data has 288 time steps (i.e., 12 measurements per hour).

[0179] At 1420, a neural network is trained based on the longitudinal time series data to generate synthetic time series data that satisfies a privacy metric.

[0180] In many embodiments, the neural network can be a DP-GAN, such as, for example, DP-GAN 500. Training of DP-GAN 500 is described above with reference to Figures 6 to 1 1C. Additionally, as described above, during training of the generator module 530, a privacy-preserving aggregated motif causality matrix 513 (M A ) is provided thereto so that the generator module 530 focuses on retaining important motifs (events) within the trajectory of the embedded original data x e . Noise can also be added to the weights of the embedder module 522, the recovery module 524, the generator module 530, and the discriminator module 540 to ensure that each network maintains differential privacy.

[0181] Many features and advantages of the present disclosure will be apparent from the detailed description, and thus, the appended claims are intended to cover all such features and advantages of the present disclosure that fall within the scope of the present disclosure. Additionally, since many modifications and variations will be readily apparent to those skilled in the art, it is not desired to limit the present disclosure to the exact construction and operation illustrated and described, and accordingly, all suitable modifications and equivalents may be considered to fall within the scope of the present disclosure.

Claims

1. A method for generating synthetic data, the method comprising: Retrieving unlabeled and univariate longitudinal time series data; And Training a neural network based on the longitudinal time series data to generate synthetic time series data that satisfies a privacy metric.

2. The method according to claim 1, wherein the longitudinal time series data includes at least 50 measured glucose levels for each of multiple individuals.

3. The method according to claim 1, wherein the privacy metric defines an upper limit on the amount of privacy loss allowed.

4. The method according to claim 3, wherein the neural network is a differentially private generative adversarial network (DP-GAN), and training the neural network includes: At a motif causality module: Receiving a first portion of the longitudinal time series data for a first group of people; And Generating an aggregated motif causality matrix based on the first portion of the longitudinal time series data, the aggregated motif causality matrix identifying causal relationships between motifs within the first portion of the longitudinal time series data.

5. The method according to claim 4, wherein the motif causality module includes a plurality of motif networks, and generating the aggregated motif causality matrix includes: Partitioning the first portion of the longitudinal time series data into data partitions, each data partition being associated with a different motif network and including a plurality of motifs, each motif being an ordered sequence of data values from the first portion of the longitudinal time series data; For each motif network, generating a motif causality matrix from the associated data partition; And Aggregating the motif causality matrices into the aggregated motif causality matrix based on the privacy metric.

6. The method according to claim 5, wherein each motif network includes a plurality of recurrent neural networks (RNNs), each RNN receiving motif data for a different motif from the associated data partition.

7. The method according to claim 4, wherein training the neural network includes: At an embedder module: Receiving a second portion of the longitudinal time series data for a second group of people different from the first group; Generating embedded time series data based on the second portion of the longitudinal time series data, the embedded time series data having a lower dimension than the second portion of the longitudinal time series data; At a generator module: Generating embedded synthetic time series data based on the aggregated motif causality matrix and the embedded time series data; At a recovery module: Generating recovered longitudinal time series data based on the embedded time series data; Generating synthetic time series data based on the embedded synthetic time series data, The synthetic time series data having the same dimension as the second portion of the longitudinal time series data; At a discriminator module: Determining whether each data value in the embedded time series data is real or synthetic, and Determining whether each data value in the embedded synthetic time series data is real or synthetic; And Train the embedder module, the recovery module, the generator module, and the discriminator module based on multiple loss functions to meet the performance metric and the privacy metric.

8. The method according to claim 7, wherein training the embedder module, the recovery module, the generator module, and the discriminator module comprises: Adding noise to weights associated with the embedder module, the recovery module, the generator module, and the discriminator module based on the privacy metric.

9. The method according to claim 8, wherein training the embedder module, the recovery module, the generator module, and the discriminator module comprises: Training the embedder module and the recovery module based on a reconstruction loss and a step loss; Training the generator module based on at least one of the step loss, a distribution loss, a motif loss, and a synthetic data adversarial loss; and Training the discriminator module based on the synthetic adversarial loss and an embedded data adversarial loss.

10. The method according to claim 9, wherein the motif loss is associated with a data sequence pattern within the second portion of the longitudinal time series data.

11. The method according to claim 7, wherein the embedder module, the recovery module, the generator module, and the discriminator module each comprise an RNN.

12. A system for generating synthetic data, the system comprising: A memory configured to store unlabeled and univariate longitudinal time series data; and At least one processor coupled to the memory and configured to: Train a neural network based on the longitudinal time series data to generate synthetic time series data that meets a privacy metric.

13. The system according to claim 12, wherein the privacy metric defines an upper limit on an amount of allowed privacy loss.

14. The system according to claim 13, wherein: The neural network is a differential privacy generative adversarial network (DP-GAN) comprising a motif causality module having a plurality of motif networks, an embedder module, a generator module, a discriminator module, and a recovery module; The motif causality module is trained based on a first portion of the longitudinal time series data for a first group of people; and The embedder module, the recovery module, the generator module, and the discriminator module are trained based on a second portion of the longitudinal time series data for a second group of people different from the first group of people.

15. The system according to claim 14, wherein: Each motif network comprises a plurality of recurrent neural networks (RNNs); and Each RNN receives motif data for different motifs from an associated data partition of the first portion of the longitudinal time series data.

16. The system according to claim 15, wherein: The motif causality module generates an aggregated motif causality matrix based on the first portion of the longitudinal time series data; and The aggregated motif causality matrix identifies causal relationships between motifs within the first portion of the longitudinal time series data.

17. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to: retrieve unlabeled and univariate longitudinal time series data; and train a neural network based on the longitudinal time series data to generate synthetic time series data that meets a privacy metric.

18. The non-transitory computer-readable medium of claim 17, wherein the privacy metric defines an upper bound on the amount of privacy loss allowed.

19. The non-transitory computer-readable medium of claim 18, wherein: the neural network is a differential privacy generative adversarial network (DP-GAN) including a motif causality module having a plurality of motif networks, an embedder module, a generator module, a discriminator module, and a recovery module; the motif causality module is trained based on a first portion of the longitudinal time series data for a first group of people; and the embedder module, the recovery module, the generator module, and the discriminator module are trained based on a second portion of the longitudinal time series data for a second group of people different from the first group of people.

20. The non-transitory computer-readable medium of claim 19, wherein: each motif network includes a plurality of recurrent neural networks (RNNs); and each RNN receives motif data for different motifs from an associated data partition of the first portion of the longitudinal time series data.

Citation Information

Patent Citations

  • 1.-2. Anti-splash nozzles for taps

    WO100288

  • A histochemical method to identify and predict disease progression of human papilloma virus-infected lesions

    WO2010003072A1