System and method for human activity recognition
Through multiple wearable sensors simultaneously collecting data and training neural networks, the problem of lack of labeled data and laboratory research limitations in the prior art is solved, and deep neural network training without manual labeling data is achieved, which improves the accuracy and computing power of human activity recognition.
Patent Information
- Application Number
- CN202380063540.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-18
- Filing Date
- 2023-08-17
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology is difficult to effectively identify human activities using deep neural networks, mainly due to the lack of labeled data and limitations of laboratory research, which makes it difficult to collect and label data, and it is difficult to accumulate enough data to train deep neural networks.
By simultaneously collecting data using multiple wearable sensors, the neural network is trained to understand human activity patterns, using paired data to avoid the need to manually label data, and optimizing feature generation of neural networks by comparing the matching or mismatch of feature vectors of neural networks.
It realizes training of deep neural networks without expert data marking and laboratory research, improves the computing power and accuracy of human activity recognition, and better recognizes related aspects of human activities.
Smart Images

Figure CN120051832A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 399,002, filed on Aug. 18, 2022, the entire content of which is incorporated herein by reference. Technical Field
[0003] The field of the present invention relates to training neural networks to understand human behavior patterns and / or using these trained neural networks in various applications. Background Art
[0004] Human activity recognition (HAR) is an area of machine learning (ML) research. The HAR field attempts to classify or identify human activities from available sensor data such as video signals or accelerometer signals from wearable sensors. HAR is valuable in many fields including healthcare, where the results can be used to monitor health states outside of acute care settings, or to measure the impact
[0005] Deep neural network architectures have been used in many machine learning domains. However, deep neural networks have not been as dominant in the HAR field as in other fields.
[0006] Deep neural networks require large amounts of data to be successfully trained. Ideally, this data is manually labeled by human experts. Especially in the field of image recognition, prior art models were created with cameras and human annotations to provide the necessary training data.
[0007] Compared to the amount of publicly available image, text, and audio data, the amount of publicly available wearable sensor data is almost zero. Additionally, unlike image or audio data, the labeling of wearable sensor data cannot be easily done by non-experts. For example, compared to labeling objects in an image or types of sounds, the data from a worn accelerometer cannot be directly recognized by the person responsible for labeling the data. This has led to the need to perform activity recognition studies in a laboratory environment where data is directly collected from human subjects, such that the researchers specify which activities the subjects must perform, so that the data can be a priori labeled as the specified activities, rather than attempting to recognize the raw accelerometer data.
[0008] Thus, laboratory studies typically result in artificial movements of human subjects, and it becomes very difficult to label human behavior. Human activities are hierarchical in nature. Any complex activity can always be decomposed into simpler movements, which in turn can be further decomposed or subdivided. For expediency, activity recognition research in the laboratory will pick some levels in the activity hierarchy, label those levels, and simply discard the activity labels of all other hierarchies because they are too difficult to capture. This results in almost every activity recognition study defining different activity categories for recognition, with different levels of overlap between other studies. Therefore, it is difficult to accumulate sufficient data to utilize the capabilities of deep neural network learning to improve activity recognition based on motion sensor data from wearable sensors.
[0009] There is a need for a method to enable the use of deep neural networks to learn activity recognition rather than through brute-force human labeling from clinical artificial activity studies. There is a need for a network architecture that exhibits improved computational capabilities to learn relevant aspects of human activities, thereby facilitating better recognition by focusing on relevant motion features at all levels of the motion hierarchy available in the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more fully understand the present disclosure, reference should be made to the following detailed description and the accompanying drawings, in which:
[0011] Figure 1A Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0012] Figure 1B Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0013] Figure 2 Flowcharts of methods according to various embodiments according to these teachings;
[0014] Figure 3 Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0015] Figure 4 Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0016] Figure 5 Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0017] Figure 6A Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0018] Figure 6B Illustrations of systems configured in accordance with various embodiments according to these teachings;
[0019] Figure 7 diagrams of systems configured according to various embodiments in accordance with these teachings; and
[0020] Figure 8 diagrams including an example of using these methods according to various embodiments in accordance with these teachings. Detailed Description
[0021] Advantageously, the methods provided herein allow for teaching a deep neural network to understand human behavior patterns without the need for expert data labeling or laboratory studies. Generally, the provided methods include a neural network that receives a time window (e.g., 1 minute) of wearable sensor signals (e.g., from an accelerometer, gyroscope, etc.) and produces a rich hierarchical description for each second of the input signal. The description cannot be directly understood by a human but can be readily used with a small amount of human-labeled data to produce any level of activity classification as needed. As described elsewhere herein, this neural network is trained and, once trained, is used in various applications to perform various actions.
[0022] In other aspects, the methods provided herein use a large amount of data from two or more simultaneously worn wearable sensors to train a neural network to understand human activities. Compared to collecting labeled laboratory data, it is relatively easy to provide a group of people with multiple wearable devices and have them collect data during their daily life activities. For example, a group of subjects can wear a smartwatch on their wrist and a monitoring patch on their chest. These two devices will collect accelerometer data simultaneously.
[0023] If a person collects data simultaneously from two or more locations on the body, then the signals must be from the same activity. This obviates the need to label data to train a deep neural network because the paired data is effectively "labeled" without the need to know the underlying activity class. The simultaneously collected data acquired in this manner must belong to the same activity class, regardless of what that activity class is, and if the paired data is artificially mixed, then such mismatched paired data cannot belong to exactly the same class.
[0024] In many of these embodiments, a first neural network is iteratively trained to obtain a first trained neural network, and a second neural network is iteratively trained to obtain a second trained neural network. Training includes receiving a first set of wearable sensor data at the first neural network. The first neural network responsively produces a first feature vector, and the first set of wearable sensor data describes a first physiological characteristic of a person. Any labels in the first set of wearable sensor data are ignored.
[0025] The training further includes receiving, at a second neural network, a second set of second wearable sensor data. Responsively, the second neural network generates a second feature vector. The second wearable sensor data describes a second physiological characteristic of a person. Any markings in the second set of wearable sensor data are ignored. At least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity that occurred for the same person at the same time, time period, or time frame.
[0026] The training further includes predicting, at a comparison neural network, whether a first feature vector from the first neural network and a second feature vector from the second neural network match or mismatch. The first feature vector and the second feature vector are determined to match when they are from the same person and are acquired at substantially the same time.
[0027] The training further includes backpropagating an error generated by a cost function to the first neural network and the second neural network, the cost function penalizing a failure to correctly determine whether the vectors match or mismatch. The backpropagation effectively and independently updates the parameters of the first neural network and the second neural network, optimizes the feature generation of the first neural network and the second neural network, and optimizes the prediction success of the comparison neural network.
[0028] Then, the first trained neural network is deployed. At least one current human subject is monitored to obtain current wearable sensor data, and the current wearable sensor data is applied to the first trained neural network to obtain a current feature vector.
[0029] A classifier is trained. At the trained classifier, the current feature vector is mapped to a classification representing one or more activity categories. Based on the classification, one or more actions are performed.
[0030] In one example, the action is to quantify, in a clinical trial, the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention. In another example, the action includes determining possible health alterations in a monitored human subject. In yet other examples, the action is to alert a clinician to investigate the health condition of the monitored human subject.
[0031] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is to control the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject. In yet another example, the action is to control the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
[0032] In other examples, the mapping utilizes known and marked vectors from other monitored human subjects.
[0033] In various aspects, the sensor is a wrist sensor or a chest sensor. Sensors are available and known in the art, which can be continuously worn by a person throughout daily life activities, record waveform data characterizing the person's movement, and typically also record data characterizing environmental conditions, cardiopulmonary function, temperature, etc. Such waveforms can include accelerometer waveforms, gyroscope waveforms, electrocardiogram (ECG) waveforms, photoplethysmogram (PPG) waveforms, impedance waveforms, electromyogram (EMG) waveforms, or electroencephalogram (EEG) waveforms. Such sensors can be watch devices or devices similar to band-aids. In other aspects, the user electronic device is a smartphone, a personal computer, a laptop computer, or a tablet computer. Other examples of sensors and user electronic devices are possible.
[0034] In other examples, the trained classifier includes a random forest or a third neural network. Other examples are possible.
[0035] In still other examples, the first neural network and the second neural network are trained at a central location. Training can also be performed at other locations.
[0036] In yet other aspects, the comparison network includes an integration of separate comparison networks with different time scales having a fixed input window size.
[0037] In other embodiments of these embodiments, the system for determining and implementing appropriate life improvement actions for a person includes a first neural network, a second neural network, and a comparison neural network. The comparison neural network is coupled to the first neural network and the second neural network.
[0038] The first neural network is configured to receive a first set of wearable sensor data and responsively generate a first feature vector. The first wearable sensor data describes a first physiological characteristic of the person. Any markings in the first set of wearable sensor data are ignored.
[0039] The second neural network is configured to receive a second set of second wearable sensor data. The second neural network responsively generates a second feature vector. The second wearable sensor data describes a second physiological characteristic of the person. Any markings in the second set of wearable sensor data are ignored.
[0040] At least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity occurring at the same time for the same person.
[0041] The comparison neural network is configured to predict whether the first feature vector from the first neural network and the second feature vector from the second neural network match or mismatch. When the first feature vector and the second feature vector are from the same person and are acquired at substantially the same time, it is determined that the first feature vector and the second feature vector match.
[0042] The error is backpropagated to the first neural network and the second neural network. The error is generated by a cost function that penalizes the failure to correctly determine whether a vector is a match or a mismatch. Backpropagation effectively updates the parameters of the first neural network and the second neural network independently to produce a first trained neural network and a second trained neural network.
[0043] The system further includes a trained classifier and a wearable sensor. The wearable sensor is worn by a current human subject.
[0044] The first trained neural network is deployed, and current wearable sensor data is obtained from the current human subject via the wearable sensor and applied to the first trained neural network to obtain a current feature vector. At the trained classifier, the current feature vector is mapped to a classification representing one or more activity categories.
[0045] One or more actions are performed based on the classification. In one example, the action is to quantify the health impact between at least one control group that receives a first intervention or no intervention and a test group that receives a second intervention in a clinical trial. In another example, the action includes determining possible health changes in a monitored human subject. In still other examples, the action is to alert a clinician to investigate the health condition of the monitored human subject.
[0046] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is to control the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject. In still another example, the action is to control the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
[0047] In other aspects, the mapping utilizes known and labeled vectors from other monitored human subjects.
[0048] In an example, the wearable sensor is a wrist sensor or a chest sensor. Other examples of sensors are possible.
[0049] In other examples, the first trained neural network is deployed at a central location. In other examples, the trained neural network can be deployed at a remote location.
[0050] In still other examples, the training occurs at a central location. The training can also occur at a remote location.
[0051] In various aspects, the trained classifier includes a random forest or a third neural network. Other examples are possible.
[0052] In still other embodiments of these embodiments, a system for training a neural network includes a first neural network; a second neural network; a comparison neural network coupled to the first neural network and the second neural network; and a control circuit coupled to the first neural network, the second neural network, and the comparison neural network.
[0053] The control circuit is configured to obtain a first set of samples of wearable sensor data from multiple individuals with a first type of wearable sensor. The control circuit is further configured to obtain a second set of samples of wearable sensor data from multiple individuals with a second type of wearable sensor. At least some of the samples match the samples of the first set to the same person and a substantially the same time window. The control circuit is further configured to train the first neural network to produce a first trained neural network by iteratively performing the following actions: inputting samples from the first set into the first neural network to generate features; inputting samples from the second set into the second neural network to generate features; inputting the features from the first neural network and the features from the second neural network together into the comparison neural network, which responds by predicting whether the input samples match or mismatch; and backpropagating the error generated by a cost function, which penalizes the failure to correctly determine whether the input samples match or mismatch, to the first neural network and the second neural network to independently change the neural parameters of the first neural network and the second neural network, optimizing the feature generation of the first neural network and the second neural network, and thus optimizing the prediction success of the comparison neural network.
[0054] In other aspects, the system further includes a trained classifier. Subsequently, the first trained neural network is deployed, and wearable sensor data from a monitored human subject is captured. The wearable sensor data is applied to the first neural network, and the first neural network responds by generating a set of features in response to a time window of the wearable sensor data.
[0055] The trained classifier maps the generated features to one or more activity categories in the time window. Based on the classification, one or more actions are performed. In one example, the action is to quantify the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention in a clinical trial. In another example, the action includes determining possible health changes in a monitored human subject. In still other examples, the action is to alert a clinician to investigate the health status of a monitored human subject.
[0056] In yet another example, the action includes selectively controlling the startup or deactivation of a device. In other examples, the action is an operation or setting for controlling a parameter of a medical device associated with treating or monitoring a monitored human subject. In still another example, the action is an operation or setting for controlling a parameter of a user electronic device associated with treating or monitoring a monitored human subject.
[0057] In other embodiments of these embodiments, a method of training a neural network includes obtaining a first set of samples of wearable sensor data from multiple people with a first wearable sensor. Obtaining a second set of samples of wearable sensor data from multiple people with a second wearable sensor. At least some of the samples are matched to the same person and substantially the same time window as the samples of the first set.
[0058] Training a first neural network to generate a first trained neural network by iteratively performing the following actions: inputting samples from the first set into the first neural network to generate features; inputting samples from the second set into a second neural network to generate features; inputting the features from the first neural network and the features from the second neural network together into a comparison network that predicts whether the input samples match or mismatch; backpropagating the error generated by a cost function that penalizes the failure to correctly determine whether the input samples match or mismatch to the first neural network and the second neural network to independently update the neural parameters of the first neural network and the second neural network, improving the feature generation of the first neural network and the second neural network, thereby improving the prediction success of the comparison neural network.
[0059] Deploying the first trained neural network. Capturing wearable sensor data from a monitored human subject.
[0060] The captured wearable sensor data is sent to the first trained neural network, and the first trained neural network generates a set of features in response to the time window of the captured wearable sensor data. A trained classifier maps the generated features to one or more activity categories in the time window.
[0061] Based on the classification, one or more actions are performed. In one example, the action is quantifying the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention in a clinical trial. In another example, the action includes determining possible health changes in a monitored human subject. In still other examples, the action is to alert a clinician to investigate the health status of a monitored human subject.
[0062] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is an operation or setting to control a parameter of a medical device associated with treating or monitoring a monitored human subject. In still another example, the action is an operation or setting to control a parameter of a user electronic device associated with treating or monitoring a monitored human subject.
[0063] In still other embodiments of these embodiments, wearable sensor data is captured from a monitored human subject. A first neural network generates a set of features in response to a time window of the wearable sensor data. A trained classifier maps the generated features to one or more activity categories in the time window. Based on the classification, one or more actions are performed. In one example, the action is to quantify the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention in a clinical trial. In another example, the action includes determining a possible health change in a monitored human subject. In still other examples, the action is to alert a clinician to investigate the health condition of a monitored human subject.
[0064] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is an operation or setting to control a parameter of a medical device associated with treating or monitoring a monitored human subject. In still another example, the action is an operation or setting to control a parameter of a user electronic device associated with treating or monitoring a monitored human subject.
[0065] The first neural network is created by: obtaining a first set of samples of wearable sensor data from multiple persons with a first type of wearable sensor; obtaining a second set of samples of wearable sensor data from multiple persons with a second type of wearable sensor, at least some of the samples being matched to the same person and substantially the same time window as the samples of the first set; training the first neural network by iteratively: inputting the samples from the first set into the first neural network to generate features; inputting the samples from the second set into the second neural network to generate features; inputting the features from the first neural network and the features from the second neural network together into a comparison network that predicts whether the input samples match or mismatch; backpropagating the error generated by a cost function that penalizes the failure to correctly determine whether the input samples match or mismatch to the first neural network and the second neural network to independently update the neural parameters of the first neural network and the second neural network, improving the feature generation by the first neural network and the second neural network, and thus improving the prediction success of the comparison neural network.
[0066] Now refer to Figure 1A, an example of a system 100 for training a neural network includes a first neural network 102, a second neural network 104, a comparison neural network 106 (coupled to the first neural network 102 and the second neural network 104), and a control circuit 108 (coupled to the first neural network 102, the second neural network 104, and the comparison neural network 106). As Figure 1A shown, these elements are used during the training phase according to the methods described herein.
[0067] The first neural network 102, the second neural network 104, and the comparison neural network 106 can be any kind of neural network or deep neural network, such as a convolutional neural network (CNN). Other examples of neural networks are possible. The first neural network 102 is configured to receive a first set of wearable sensor data and, in response, generate a first feature vector 107. The first wearable sensor data describes a first physiological characteristic of one or more persons 112 being trained. Any markings in the first set of wearable sensor data are ignored by the first neural network 102. "Marking" means any identifier used to identify the source, content, or other characteristics of the data. The person 112 being trained has wearable sensors 114 and 115. The wearable sensors 114 and 115 can be any type of sensor, such as a chest sensor or a wrist sensor. These sensors can obtain waveform data such as readings from an accelerometer or gyroscope, which characterize the movement experienced at the location where the sensor is worn. These sensors can also obtain waveform data such as readings from an electrocardiogram signal, an impedance signal, a photoplethysmogram signal, etc. These sensors can further record data such as skin and environmental temperature and can calculate derived vital sign characteristics such as heart rate or respiratory rate in the sensor firmware. By way of two examples, a watch sensor and a sticky torso patch sensor are known in the art to collect the above data. The neural networks 102, 104, and 106 can be formed and stored in a database or other electronic memory device. The database 116 is any type of electronic memory device that stores information electronically.
[0068] The second neural network 104 is configured to receive a second set of second wearable sensor data. Responsively, the second neural network 104 generates a second feature vector 109. As described above, the second wearable sensor data describes a second physiological characteristic of the person 112. Any markings in the second set of wearable sensor data are ignored by the second neural network 104. As previously mentioned, "marking" means any identifier used to identify the source, content, and / or other characteristics of the data. For the purpose of human activity recognition, a preferred embodiment uses a watch-type sensor with at least one continuous 3-axis accelerometer sampled at a frequency of 5 Hz or higher, and a sticky torso patch with at least one continuous 3-axis accelerometer sampled at a frequency of 5 Hz or higher. The first feature vector 107 and the second feature vector 109 represent features in the input signal, which are characteristics of, for example, running, standing, exercising, walking at a specific speed, climbing stairs, cycling, driving, playing a specific game, or engaging in a specific activity, to name just a few examples.
[0069] In various aspects, the first feature vector 107 and the second feature vector 109 are groupings of node output values having numerical values (e.g., real numbers or integers). These values are generated by an activation function of the neural network 102 or 104, which represents the activation level of the nodes at the output layer of the neural network 102 or 104. The activation function can take a functional form known in the art, such as sigmoid, arctangent, hyperbolic tangent, rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), etc. The number of such nodes (and thus the length of the vector) in any neural network can be on the order of dozens or hundreds of nodes. Each such node receives inputs from other nodes at a higher or upstream layer of the neural network, and these inputs represent the activation values of those nodes. The inbound activation values are multiplied by learned weights to obtain numbers that are summed at each node. In this way, the inbound activation values from any given upstream node can be amplified or attenuated by the weights on that connection, as it contributes to the sum of the inbound activations of the node. The sum for each node is also added to a bias value. As described elsewhere herein, the learned weights and biases are updated when the neural network is trained by backpropagation of error. Based on the task of measuring error, the adjustment of these weights and biases improves the performance of the neural networks 102 and 104.
[0070] The control circuit 108 trains the neural networks 102, 104, and 106 and applies inputs to these neural networks and causes the neural networks 102, 104, and 106 to produce outputs (e.g., vectors as described herein). It should be understood that the term "control circuit" as used herein generically refers to any microcontroller, computer, or processor-based device having a processor, memory, and programmable input / output peripherals, which is typically designed to manage the operation of other components and devices. It should further be understood to include common accessory devices, including memory, transceivers for communicating with other components and devices, etc. These architectural options are well known and understood in the art and do not require further description herein. The control circuit 108 can be configured (e.g., by using corresponding programming stored in memory, as will be well understood by those skilled in the art) to perform one or more of the steps, actions, and / or functions described herein.
[0071] At least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity that occurs at the same time from the same person 112. "Time" means a time period, time window, time frame, or instant (e.g., a specific time value).
[0072] The comparison neural network 106 is configured to predict whether the first feature vector 107 from the first neural network 102 and the second feature vector 109 from the second neural network 104 match or mismatch. When the first feature vector 107 and the second feature vector 109 are derived from input data from the same person and are acquired at substantially the same time, it is determined that the first feature vector 107 and the second feature vector 109 match. This determination can be made based on the timestamps of the collected data and knowledge of which human subjects the wearable sensor data are from, and avoids the dilemma of post hoc manual labeling by human experts with activity classification. By programmatically parsing the library of collected human subject wearable sensor data in an automated manner, match and mismatch data sets can be easily generated. The prediction made by the comparison neural network 106 can be made by comparing the vectors output by networks 102 and 104. In these aspects, the cross-entropy loss for determining whether the prediction of a match is correct is determined. This value can be calculated by an error function and is used to train the neural networks 102 and 104.
[0073] Generally, training the neural networks 102 and 104 involves two different phases. The first phase is the forward pass through the neural network 102 or 104, where the network parameters are frozen, an input is provided to the network 102 or 102, and the network 102 or 104 produces an estimate of the target output value. It should be understood that training changes the physical characteristics of the neural networks and classifiers described herein.
[0074] The second stage, referred to as backpropagation through neural network 102 or 104, involves control circuit 108 calculating an error amount regarding whether vectors 107 and 109 are correctly determined to match or mismatch. In an example, the partial derivative of each network parameter with respect to the error is calculated through the cost function for each layer in network 102 or 104. The neural network parameters of networks 102 and 104 are then updated by subtracting a small multiple of the partial derivative of each corresponding parameter. The cost function penalizes the failure to correctly determine whether the vectors match or mismatch. This backpropagation process effectively updates the parameters of the neural network independently to produce a first trained neural network 102 and a second trained neural network 104.
[0075] Training is performed on available training data batches that include match or mismatch data windows. The process is then repeated until network 102 or 104 has learned to accurately produce an estimate of the target output. Methods for accurately determining when training is complete are known in the art and are applicable here and typically involve achieving a desired performance level by comparing network 106, or when performance reaches a maximum steady state.
[0076] Once this training step is complete, the trained neural network 102 or 104 can be used to train classifier 120. In other words, the deep learning model that produces features is trained as discussed above. Once this training is complete, the trained networks 102 and 104 are frozen and used to generate features for new labeled input data, and these features and labels are used to train classifier 120.
[0077] In one specific example where a classifier can be trained to recognize walking behavior as opposed to all other activities, after training neural networks 102 and 104 to analyze 1 - second data windows, a small amount of labeled walking data is used to build a classifier using features from networks 102 and 104 by inputting 1 - second accelerometer data windows into one or the other of networks 102 and 104 corresponding to the type of sensor from which the accelerometer data originated, where the accelerometer data is considered to be captured during walking or during non - walking activities, and a resulting feature vector is generated for each such sample. These feature vectors are then used with a traditional random forest classifier to classify each second of accelerometer data as "walking" or "not walking".
[0078] Classifier 120 used herein can be any number of different types, including random forest, neural network, logistic regression, support vector machine, etc. Once trained, classifier 120 is used to estimate an activity label based on features (e.g., in the form of vectors) applied to classifier 120. The activity label represents the type of activity detected and, as discussed herein, this can be used to perform various actions.
[0079] Now refer to Figure 1B , and describe an example of using a trained neural network and a trained classifier (e.g., obtained from the method of Figure 1A ) during the monitoring phase. In one example, the process of Figure 1A (or a similar process) has been performed to generate the trained neural networks 102 and 104 and the trained classifier 120. Any one of these trained networks 102 and 104 is deployed, and data is applied to generate a feature vector in response.
[0080] In the example of Figure 1B , the trained first neural network 102 is deployed. In various aspects, the system 100 further includes a trained classifier 120 and a wearable sensor 122. The wearable sensor 122 is worn by the current human subject 124. In the example, the wearable sensor 122 can be a chest or wrist sensor to match the deployed network. Other examples of sensors are possible. It should be understood that the second trained neural network 104 can also be deployed, and the following description is also applicable.
[0081] In various aspects, the trained neural network 102 and the trained classifier 120 can be deployed at a central location 103. The central location 103 can be an enterprise, headquarters, home office, data center, or other physical location. In another example, the trained neural network 102 can be deployed at a remote location, such as at a hospital, doctor's office, or research institution. Other examples are possible. Data from the sensor can be uploaded to the location 103 via a mobile digital communication network or a local area network or other means. The upload can be performed periodically or continuously.
[0082] The first trained neural network 102 is deployed, and current wearable sensor data is obtained from the current human subject 124 via the wearable sensor 122 and applied to the first trained neural network 102 to obtain the current feature vector 128. At the trained classifier 120, the current feature vector 128 is applied to the trained classifier 120, and at the trained classifier 120, the current feature vector 128 is mapped to a classification representing one or more activity categories. In various aspects, the mapping utilizes known and labeled vectors from other monitored human subjects. In other aspects, the mapping includes comparing the current feature vector 128 with known and labeled vectors from other monitored human subjects and determining the activity category through a nearest neighbor calculation.
[0083] Perform one or more actions based on classification. In one example, the action is to quantify the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention in a clinical trial. In another example, the action includes determining possible health changes in a monitored human subject 124. In still other examples, the action is to alert a clinician to investigate the health status of the monitored human subject 124.
[0084] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is to control the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject 124. In yet another example, the action is to control the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject 124. For example, the frequency, speed, and / or range of treatment or monitoring can be changed. The settings on these devices (e.g., characteristics related to video or display quality or characteristics related to font size, color, or other characteristics on the screen of these devices) can also be controlled and changed, for example, when displaying an alert (e.g., based on the type of alert).
[0085] Now refer to Figure 2 , and describe a method for training and utilizing a trained neural network.
[0086] At step 202, iteratively train a first neural network to obtain a first trained neural network, and at step 204, iteratively train a second neural network to obtain a second trained neural network. Training includes receiving a first set of wearable sensor data at the first neural network. The first neural network responsively generates a first feature vector, and the first wearable sensor data describes a first physiological characteristic of a person. Ignore any labels in the first set of wearable sensor data. In various aspects, the sensor is a wrist sensor or a chest sensor.
[0087] Training further includes receiving a second set of second wearable sensor data at the second neural network. The second neural network responsively generates a second feature vector. The second wearable sensor data describes a second physiological characteristic of a person. Ignore any labels in the second set of wearable sensor data. At least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity occurring at the same time for the same person.
[0088] At step 206 and at the comparison neural network, it is predicted whether the first feature vector from the first neural network and the second feature vector from the second neural network match or mismatch. The first feature vector and the second feature vector are determined to match when they are determined to be from the same person and are acquired at substantially the same time or time period. In various aspects, the comparison network includes an ensemble of separate comparison networks with different time scales having a fixed input window size.
[0089] At step 208, the error is backpropagated to the first neural network and the second neural network, where the error is generated by a cost function that penalizes the failure to correctly determine whether the vectors match or mismatch. The backpropagation and application of the error to the neural networks effectively and independently update the parameters of the first neural network and the second neural network, optimize the feature generation of the first neural network and the second neural network, and optimize the prediction success of the comparison neural network.
[0090] In an example, the first neural network and the second neural network are trained using steps 202, 204, and 206 at a central location. These elements can also be deployed at separate and different locations.
[0091] At step 210, the first trained neural network is deployed. For example, the first neural network can be deployed at a central location such that it can be conveniently used to receive data.
[0092] At step 212, at least one current human subject is monitored to obtain current wearable sensor data, and the current wearable sensor data is applied to the first trained neural network to obtain a current feature vector.
[0093] At step 214 and at the trained classifier, the current feature vector is mapped to a classification representing one or more activity categories. The classifier itself can be trained. In an example, the mapping uses known and labeled vectors from other monitored human subjects and uses the k-nearest neighbor method to output the category of the nearest neighbor or neighbor centroid to the current feature vector. In various aspects, the trained classifier includes a random forest or a third neural network. Other examples are possible.
[0094] At step 216 and based on the classification, one or more actions are performed. The actions can take various different forms.
[0095] In one example, the action is to quantify the health impact between at least one control group receiving a first intervention or no intervention and a test group receiving a second intervention in a clinical trial. In another example, the action includes determining possible health changes in the monitored human subject. In still other examples, the action is to alert a clinician to investigate the health status of the monitored human subject.
[0096] In yet another example, the action includes selectively controlling the activation or deactivation of a device. In other examples, the action is an operation or setting to control a parameter of a medical device associated with treating or monitoring a monitored human subject. In yet another example, the action is an operation or setting to control a parameter of a user electronic device associated with treating or monitoring a monitored human subject. In an example, the user electronic device is a smart phone, a personal computer, a laptop computer, or a tablet computer. Other examples of sensors and user electronic devices are possible.
[0097] Now refer to Figure 3 , and describe an example of a method for training a neural network. In this example, the torso neural network 302 models human activities measured by sensors on a human torso (e.g., a chest sensor such as a chest patch), and the wrist neural network 304 models human activities measured by wrist sensors. A comparison neural network 306 is also provided. It should be understood that these are two examples of sensors and the locations of the sensors on the human body, and other types of sensors and / or locations on the human body are possible.
[0098] The example shows the matched accelerometer data 310 and the mismatched paired accelerometer data 320 for training the neural networks 302 and 304. The torso accelerometer data remains the same between both the matched data 310 and the mismatched data 320, but the wrist accelerometer data is changed to a selected segment of the wrist accelerometer signal from a different time point or time period than when the torso accelerometer segment was collected.
[0099] At step 322, depending on the data source, the accelerometer data paired from the matched and mismatched time periods is provided to one of two networks, namely the torso neural network 302 and the wrist neural network 304. In one example, the torso neural network 302 and the wrist neural network 304 are multi-layer convolutional neural networks, each of which generates a single vector representation of the input accelerometer segment. More specifically, the torso neural network 302 generates a vector 332 (a single vector representation of the chest or torso input accelerometer segment), and the wrist network generates a vector 334 (a single vector representation of the wrist input accelerometer segment). Using the mismatched data 320, the torso neural network 302 generates a vector 336 (a single vector representation of the torso input accelerometer segment), while the wrist neural network 304 generates a vector 338 (a single vector representation of the wrist input accelerometer segment selected from a different time point than when the torso data was collected).
[0100] The vectors 332 and 334 of the paired accelerometer segments generally look similar to each other (e.g., they may be the same or have slight deviations in amplitude, pattern, shape, size, or other characteristics), while the vectors 336 and 338 of the mismatched segments do not look very similar to each other (e.g., they are different and / or have large deviations in amplitude, pattern, shape, size, or other characteristics). The comparison network 306 receives the vector representations of the activities and attempts to correctly determine whether the vectors represent matched data or mismatched data. An error signal is then generated from the comparison network. The error signal indicates whether a correct or incorrect determination has been made regarding whether the data matches / mismatches. The error signal is then backpropagated to the networks 302 and 304, forcing the torso neural network 302 and the wrist neural network 304 to learn increasingly subtle descriptions of the accelerometer data so that the comparison network 306 can correctly classify them.
[0101] As a specific example, the matched data 310 is from a time period when the subject is walking, but the mismatched data 320 is from a different walking time period. The fact that the torso neural network 302 and the wrist neural network 304 learn that the data represents walking is not sufficient to distinguish the two activities. The torso neural network 302 and the wrist neural network 304 must provide even more sub-classifications to distinguish or classify the data. For example, to give two examples, the walking can be further classified by characteristics such as walking speed and / or gait style. In this way, the networks 302 and 304 can learn rich descriptions of the underlying activities without knowing the exact activity or the exact nature of the activity.
[0102] It can be seen that the present method collects data simultaneously from two or more locations on the body at the same time and indicates that the paired data comes from the same person and activity. Artificial data is introduced and distinguished from the determined activities because it is known that mismatched data cannot come from the same activity. These methods avoid the use of markers and the need to assign markers to the data during this training process of the deep neural network because the data collected simultaneously from the same person must come from the same activity category, regardless of what that category is.
[0103] Once the torso neural network 302 and the wrist neural network 304 are trained in this way, they can be used in various ways. For example, the learned features can be used together with a small amount of manually labeled activity data, and a conventional classifier can be trained to convert the features into a desired set of activity classifications.
[0104] The torso neural network 302 and the wrist neural network 304 can be fine-tuned using a small hand-labeled activity data set to obtain the desired set of activity classifications.
[0105] As described elsewhere herein, a trained torso or wrist neural network can be deployed during a monitoring phase, applying monitored human data to the trained neural network, and applying the synthetic vectors from the trained neural network to a trained classifier. To train the classifier, the vectors produced by the torso neural network 302 and the wrist neural network 304 can also be applied on-the-fly to a "one-shot" classifier.
[0106] For example, a clinician has a human subject in physical therapy perform a therapy while wearing a wearable device. Features from the therapy can be used to train a classifier for a particular action as performed by the subject. The clinician can then quantify the extent to which the subject adheres to their physical therapy regimen, thus training the classifier to associate the vectors with "good" or "poor" adherence.
[0107] In another specific example, a clinician can demonstrate the correct and incorrect ways of performing a particular activity during physical therapy, training a one-shot classifier or distance metric to identify "good" versus "poor" (or degrees thereof) of the activity. Later, when a human subject performs the activity on their own at home, the sensor signals are applied to the trained neural network, the output vectors of the trained neural network are applied to the classifier, and the classifier can classify the activity as being performed correctly or incorrectly. If the human subject performs the activity incorrectly, then the device can warn the human subject.
[0108] In still other examples, the sensor data is applied to a trained neural network, producing vectors representing characteristic attributes indicative of the effects of a drug (assuming the features capture these attributes). For example, the detection of a gait style can indicate certain effects of a drug taken by a human subject.
[0109] Now refer to Figure 4 , and an example of an apparatus for training a neural network is described. In this example, accelerometer inputs are used as inputs to the neural network. Additionally, additional signal inputs (e.g., gyroscopes) can be used. Although two neural networks are trained in this example, it should be understood that more than two neural networks can be used.
[0110] As described herein, these methods provide a comparison between two time windows of accelerometer data obtained from different devices. In various aspects, a multi-window time-scale vector is created. Generating multiple vectors of different time scales allows each vector to capture features related to its time scale, in various aspects, from features at the 1-second level to features at the 1-minute level. Features are meant to be general activity recognition features, and thus it is not known a priori what time scale is best for each activity recognition task. As an example, a feature vector created from a 4-second window may be suitable for performing running or walking detection, but may not be sufficient to determine a higher-level activity such as "playing basketball". In an alternative embodiment, if the desired feature window size is known prior to training the features, then it will not be necessary to use multiple window time scales, and only the desired feature window size will be needed.
[0111] In other examples, a convolutional network utilizing dilated convolutions can maintain the same input and output at each network layer, but the network receptive field (how long the network "looks" at a given network depth) doubles at each layer. In such a network, the outputs from each layer can be paired similarly.
[0112] Now turning to Figure 4 , the torso neural network 402 includes various blocks that perform convolutional and max pooling operations. The convolutional and max pooling operations approximately halve the temporal resolution of the data received at each of these blocks. In this way, the features generated at each layer become increasingly general and high-level, and less specific. The blocks that perform these functions include blocks 410, 411, 412, 413, 414, 415, 416, 417, 418, and 419.
[0113] The linear depth convolutional layers 420, 421, 422, 423, 424, 425, and 426 reduce the number of features (dimensions) of the data. In various aspects, the linear depth convolution simply performs a weighted average along the feature dimensions to produce (in one example) a 60×10 output feature vector. The linear depth convolutional layers 420, 421, 422, 423, 424, 425, and 426 are linear layers because learning occurs in the main path of the neural network, resulting in a dimensionality reduction of the information.
[0114] The wrist neural network 404 includes various blocks that perform convolution and max pooling operations. The blocks that perform these functions include blocks 430, 431, 432, 433, 434, 435, 436, 437, 438, and 439. The convolution and max pooling operations roughly halve the temporal resolution at each of these blocks. In this way, the features generated at each layer become increasingly general and high-level, and less specific. The linear depth convolution layers 440, 441, 442, 443, 444, 445, and 446 reduce the number of features (dimensions). The linear depth convolution performs a weighted average along the feature dimension to produce an output feature vector of 60×10 in one example. As mentioned before, the linear depth convolution layers 440, 441, 442, 443, 444, 445, and 446 are linear layers because learning occurs in the main path of the neural network, resulting in a reduction in the dimensionality of the information.
[0115] Data 450 flows through the torso neural network 402 in the direction indicated by the arrow labeled 451. Data 452 flows through the wrist neural network 404 in the direction indicated by the arrow labeled 453.
[0116] Generally, data 450 flows through network 402 and data 452 flows through network 404. The size of the data is reduced by each block of the corresponding neural network. As mentioned, some of the blocks flow to the linear depth convolution layers.
[0117] The linear depth convolution layers from each of the torso neural network and the wrist neural network are combined at the connectors 453, 454, 455, 456, 457, 458, and 459. The outputs of the connectors 453, 454, 455, 456, 457, 458, and 459 flow to the comparison neural network 460 that includes seven neural networks 461, 462, 463, 464, 465, 466, and 467. The chest and wrist features are linked together as shown. Once the features are combined, the features of 7 resolutions (from once per second to once per minute) are input into the 7 individual feed-forward comparison neural networks 461, 462, 463, 464, 465, 466, and 467 that form the overall comparison neural network 460.
[0118] The seven comparison neural networks include feedforward layers 470, 471, 472, 473, 474, 476, 477, 478, 479, 480, 481, 428, and 483, sigmoid layers 485, 486, 487, 488, 489, 490, and 491, and cross-entropy (also known as "x-entropy") layers 492, 493, 494, 495, 496, 497, and 498. At step 499, the data outputs through each of these are summed to a single loss. For example, data from linker 453 arrives at feedforward layer 470, feedforward layer 477, sigmoid layer 485, and x-entropy layer 492 to produce a loss on network 461, which is summed with the losses of the other networks 462, 463, 464, 465, 466, and 467 to produce the total loss 499.
[0119] The sigmoid activation blocks 485, 486, 487, 488, 489, 490, and 491 produce a "match" or "mismatch" probability for each timestamp in the feature vector (e.g., the probability of a feature that occurs once per second input to network 461 is 60×1, and the probability of a feature that occurs once every 30 seconds input to network 466 is 2×1).
[0120] Then networks 402 and 404 are trained to minimize the sum of the total cross-entropy losses at all time resolutions at step 499. This is backpropagated to the torso neural network 402 and the wrist neural network 404 to train these networks. In one example, the weights are adjusted within these networks.
[0121] In Figure 4 In an example of the operation of this system, a data sample 450 from the torso sensor includes a 60-second window of 3-axis accelerometer measurement waveform data. Similarly, a data sample 452 from the wrist sensor also includes a 60-second window of 3-axis accelerometer waveform data. The 3-axis accelerometer measures vibrations or movements in each of three orthogonal directions, commonly referred to as x, y, and z, relative to the circuit board plane of the sensor device (wrist or chest patch) on which the accelerometer is mounted. Thus, human walking movements are expressed in a complex manner in the orthogonal directions of the accelerometer unit from the wrist position or chest position or other positions where the sensor can be applied, but inherently encode relevant movements such as heel strikes, swings, and wobbles of typical human walking behavior. Non-walking behaviors will be expressed in some other complex manner in the orthogonal directions of the accelerometer unit from such positions and will be inherently encoded.
[0122] The accelerometer data for each window is resampled to 24 Hz (Hertz, samples per second), such that the input data has 1440 samples in the time dimension, and 3 accelerometer channels (e.g., 1440×3 input). Initially, 4 convolutional and max pooling layers are applied to the resampled accelerometer signal data to build features in the channel dimension and downsample to lose time resolution.
[0123] The wrist data 403 is applied to block 430 (producing a 1440×24 output of the convolutional layer, then a 480×24 output of the max pooling), then to block 431 (producing a 480×48 output of the convolutional layer, then a 240×48 output of the max pooling), then to block 432 (producing a 240×96 output of the convolutional layer, then a 120×96 output of the max pooling), then to block 433 (producing a 120×96 output of the convolutional layer, then a 60×96 output of the max pooling), then to block 434 (producing a 60×96 output of the convolutional layer, then a 30×96 output of the max pooling), then to block 435 (producing a 30×96 output of the convolutional layer, then a 15×96 output of the max pooling), then to block 436 (producing a 15×96 output of the convolutional layer, then an 8×96 output of the max pooling), then to block 437 (producing an 8×96 output of the convolutional layer, then a 4×96 output of the max pooling), then to block 438 (producing a 4×96 output of the convolutional layer, then a 2×96 output of the max pooling), and then to block 439 (producing a 2×96 output of the convolutional layer, then a 1×96 output of the max pooling).
[0124] At layer 433 (layer 4), the wrist input 1440×3 shaped data has been transformed to a size equal to 60×96; in other words, this is a 96-element vector of data per second in a 60-second input. It is at this point that the first linear depth convolutional filter 440 is applied to create features at the one-second resolution level. Linear depth convolution simply performs a weighted average along the feature dimension to produce a 60×10 output feature vector. It is a linear layer because learning occurs in the main path of the neural network; the linear layer is for dimensionality reduction of the existing information. The other layers 441, 442, 443, 444, 445, and 446 perform similar operations on the data siphoned from the convolutional / max pooling layers at 434, 435, 436, 437, 438, and 439. The same process is also applied to the torso data 401 through layers 410 - 419 and the linear depth layers 420 - 426.
[0125] The convolutional and max pooling layers roughly halve the temporal resolution along each layer. In this way, the features generated at each layer become increasingly "higher level" or less specific or less general. However, using linear depth convolutions at each step allows for the representation of features along the entire hierarchy. It is these features that the comparison network uses to determine matching / mismatching activities. At the highest level, there are 10 features per second in 60 seconds (60×10); at the lowest level, there are 30 features in the entire minute (1×30).
[0126] The chest and wrist features are linked together by linkers 453, 454, 455, 456, 457, 458, and 459. Once the features are combined, the features at 7 resolutions (from once per second to once per minute) are input into 7 separate feedforward neural networks 461, 462, 463, 464, 465, 466, and 467. As mentioned, the feedforward networks 461, 462, 463, 464, 465, 466, and 467 have two feedforward (hidden) layers, a sigmoid layer with a sigmoid activation function, and a cross-entropy layer that produces a "match" or "mismatch" probability for each timestamp in the feature vector (e.g., the probability for the once-per-second resolution feature in network 461 is 60×1, but the probability for the once-per-30-second resolution feature in network 466 is 2×1). The cross-entropy loss is averaged at each feedforward network, and the average cross-entropy losses over all feedforward networks are summed at 499.
[0127] Then networks 402 and 404 are trained to minimize the sum of the total cross-entropy losses over all temporal resolutions using the value 499. This is backpropagated to the torso neural network 402 and the wrist neural network 404 to train these networks. Through this backpropagation of the error, the two networks become better at generating features related to determining whether the activity represented by the raw data from the sensors is a match or a non-match. Training is performed on batches of available training data that include match or mismatch data windows until the comparison network 460 achieves a desired performance level in terms of the prediction error regarding matches, or the performance reaches a maximum steady state. The method for determining when training is complete is known in the art and applicable here.
[0128] Individually and otherwise, after the wrist neural network 404 and the torso neural network 402 are fully trained, a small amount of labeled data (e.g., labeled walking data) is used to construct a classifier by upsampling each time-resolution feature to once per second and using the features from the torso convolutional network 402 and the wrist convolutional network 404. For example, the feature values at layer 420 (which is 60×10) already have 60 groups of 10 values, and no upsampling is required there. At the feature values from layer 421 (30×10), each of the 30 groups of 10 values is upsampled by repeating the ten values adjacent to each value once. At the feature values at layer 424 (4×25), each of the 4 groups of 25 values is repeated more than 14 times adjacent to each other to produce 4 groups of 15 groups of 25 values. This is done for all time resolutions, so a 140-dimensional vector is produced for the data per second, where the 140 values come from 10, 10, 15, 20, 25, 30, and 30 feature values. These feature vectors are then used with a traditional random forest classifier to classify the accelerometer data for each second as "walking" or "not walking".
[0129] Now refer to Figure 5 , and describe another example of a method for training and using a neural network. The torso feature neural network 502, the wrist feature neural network 504, and the comparison neural network 506 are utilized. The neural networks 502, 504, and 506 are neural networks as described elsewhere herein. The method includes a first training mode, a second training mode, and a monitoring mode (which follows the first training mode and the second training mode in time).
[0130] In the first training mode, the unlabeled accelerometer signal 503 from a sensor deployed on a human torso (e.g., chest) is applied to the torso feature neural network 502. The unlabeled accelerometer signal 505 from a sensor deployed on a human wrist (the same person and the same time frame when "matched", but different time frames and / or different people when "not matched") is applied to the wrist feature neural network 504. The torso feature neural network 502 produces features 507, and the wrist feature neural network 504 produces feature vectors 509, which represent the features or characteristics of human activities in the corresponding time windows.
[0131] At a comparison neural network 506, feature vectors 507 and 509 are received. The comparison neural network 506 uses a cost function to determine whether the vectors match or do not match. The amount of match or mismatch is indicated by an error, which in the example is a numerical value. This value is applied to neural networks 502 and 504, which in the example adjust their parameters by subtracting small multiples of the partial derivatives of each respective parameter. At step 511, the cost function penalizes the failure to correctly determine whether the vectors match or do not match. This backpropagation process effectively updates the parameters of neural network 502 and neural network 504 independently, and results in a trained torso feature neural network 502 and a trained wrist feature neural network 504. The two networks become better at producing features related to determining whether an activity measured by the sensors matches or does not match.
[0132] The neural networks 502 and 504 that produce the feature vectors 507 and 509 are trained in a first training mode. Once the first training mode is complete, the trained torso feature neural network 502 and / or the trained wrist feature neural network 504 are used to generate feature vectors in a second training mode. In this example of the second training mode, the trained wrist feature network is used to produce feature vectors 522 from this trained network, and these feature vectors 522 are used as inputs to a classifier 520 and are learned by the classifier 520 via a supervised program in various aspects. In various aspects, the feature vectors 522 are used together with corresponding activity labels 524 to train a conventional classifier. This classifier 520 can be in any number of forms, including random forest, neural network, logistic regression, support vector machine, etc. The classifier 520 is trained to estimate an activity label 528.
[0133] It should be understood that the steps in the second mode effectively train the classifier 520 (to produce a completely separate classifier) while keeping the networks 502 and 504 frozen. It is also possible to perform fine-tuning on the trained networks 502 and 504 to estimate activity classification instead of training the classifier 520. One classifier can be used for wrist features and another classifier for torso features. It is also possible that feature data from more than one sensor can be used to train a classifier such as classifier 520, and thus feature data from both networks 502 and 504 can be used as inputs to such a classifier together with the activity label 524; this approach takes into account activity recognition during a monitoring mode (described below), in which a person will wear multiple sensors at the same time, for example, a chest patch sensor and a wrist sensor worn together.
[0134] In a monitoring mode, new accelerometer signals from different people are processed to obtain feature vectors, which are then further processed to estimate underlying activity categories. More specifically, a trained wrist feature neural network 504 is used in conjunction with a trained classifier 520 to obtain a classification of human activities. A new accelerometer signal 530 from a monitored human subject is applied to the trained wrist feature neural network 504 to produce features 532. The features 532 are applied to the trained classifier 520 to produce a classification 531 of the human activity in the signal 530. New accelerometer signals from torso sensors of different people may be applied to a trained torso network to obtain feature vectors, which are applied to the trained classifier 520 to obtain activity categories.
[0135] Now refer to Figure 6A , and describe an example of creating a feature vector by a trained network. Figure 6A Illustrates accelerometer data 602 applied to a neural network 604 (e.g., a wrist or torso neural network as described herein). Features 606 are produced and the features can be represented as a vector.
[0136] Features 606 are output (e.g., as a vector) on multiple time scales, thereby producing vectors 620, 622, 624, 626, 628, 630, and 632 such that the features capture both high-level and low-level information about the underlying activity. In this implementation, the features represented by the vectors 620, 622, 624, 626, 628, 630, and 632 at each time scale are not of the same shape and thus must be repeated to form a vector 634. The vector 634 is applied to a random forest 638 or other trained classifier to produce an activity classification 640.
[0137] Feature vector 620 represents 1 feature per second and has a depth of 10. Feature vector 622 represents 1 feature per two seconds and has a depth of 10. Feature vector 624 represents 1 feature per four seconds and has a depth of 15. Feature vector 626 represents 1 feature per 8 seconds and has a depth of 20. Feature vector 628 represents 1 feature per 15 seconds and has a depth of 25. Feature vector 630 represents 1 feature per 30 seconds and has a depth of 30. Feature vector 632 represents 1 feature per 60 seconds and has a depth of 30. The depths are added (10 + 10 + 15 + 20 + 25 + 30 + 30) to obtain a depth of 140 for vector 634, and each one-second time window has vector 634.
[0138] In the example, the features of the repeating vector 620 are repeated such that each feature occurs at an interval of once per second. For a feature that occurs once per second, this means no repetition occurs, while a feature that occurs once per minute is repeated 60 times by the feature vector 632. Overall, the resulting vector 634, and this is a 140-length feature vector within each 1 second of 60 seconds of accelerometer input.
[0139] Now refer to Figure 6B , a diagram showing the monitoring mode is described. This process can be seen in the features in the images of the following monitoring mode steps. In this example, the accelerometer data 602 is the accelerometer input for 60 seconds. This is applied to the neural network 604 to produce features in vector form. Seven outputs (vectors 620, 622, 624, 626, 628, 630, and 632) are produced, each output representing a different resolution of the features at different time scales. These can be combined into the output vector 634, which is applied to the random forest 638 to produce an activity class determination. That is, the vector 634 has a length of 140 and represents information at all time scales at a specific point in time to produce the activity classification 640. In this example, the activity classification is a binary number, where one value (e.g., 1) represents an activity (e.g., walking), and the other value (e.g., 0) represents another activity (e.g., resting). It should be understood that this is only one example of activity classification, and other examples with more than two levels (i.e., having any number of levels) are possible.
[0140] Now refer to Figure 7 , an example showing matching data and mismatching data is presented. The diagrams of matching and mismatching show data collected simultaneously from a wrist accelerometer and a torso accelerometer, and show how paired data segments are selected during the same time period (matching) or during different time periods (mismatching).
[0141] It can be seen that the matching data 702 includes torso sensor data 704 and wrist sensor data 706. The torso sensor data 704 and the wrist sensor data 706 are obtained from exactly the same time or exactly the same time frame or time period, and will indicate the same activity of a person.
[0142] On the other hand, the mismatched data 708 includes torso sensor data 710 and wrist sensor data 712. The torso sensor data 710 and the wrist sensor data 712 are not taken from exactly the same time or exactly the same time frame or time period, and do not necessarily indicate the same activity of a person. The matched and mismatched pairs include data for training the feature generation network; "matched" and "mismatched" are actually labels for this training. It is not necessary to know the category of the activity, even if the activities during the mismatched input are actually similar. The feature generation network is trained to optimize the chance that the comparison network can accurately predict a match. However, by generating features to do so, the generated features also become useful for improving activity classification in the classifier.
[0143] Now refer to Figure 8 , and describe an example of the application of these methods. In Figure 8 's example, the health impact between at least one control group that receives the first intervention or no intervention and a test group that receives the second intervention is quantified in a clinical trial.
[0144] A pharmaceutical company wishes to show that a treatment can improve the quality of life of people with heart failure. One group 802 receives treatment to address the disease, and another group 804 does not receive treatment. Both groups are monitored using a wrist-worn activity monitor or sensor 806. The accelerometer data 808 from the wrist device 806 is streamed to the cloud 810, where the data is processed according to the methods described herein, thereby producing a fine-grained measurement of the participants' activities. In these aspects, the cloud 810 can include a platform that has a control circuit, a memory, a trained neural network, and a trained classifier, which are collectively referred to as the model 812. The accelerometer data 808 is applied to the trained neural network to produce a feature vector, and the feature vector is applied to the trained classifier to produce an activity classification. The results can be depicted by a graph 814. Since the current method causes the creation of fine-grained activity classifications, much more information about the underlying activities is obtained compared to previous methods.
[0145] At the end of the study, the activity characteristics of each group are analyzed. The enhanced detection capabilities of the present method, such as differences in attributes like walking speed, gait confidence, and the increased incidence of sitting and standing, can be shown and illustrated between the groups, thereby indicating the new treatment efficacy.
[0146] Other examples of using these methods are possible and some examples are described below and utilize the above-described cloud-based platform. In an additional example, possible health changes in a monitored human subject are determined. For example, a patient is discharged from a hospital and enrolled in a cloud-based platform where they will receive a wrist-worn activity monitor or sensor. Accelerometer data from the wrist device is streamed to the cloud-based platform where the signal is automatically analyzed by the methods described herein to generate an activity classification. The discharging clinician expects the patient to slowly increase their activity level as they recover and monitors the activity classification generated by these methods on a daily basis.
[0147] In other examples, an alert can be sent to a clinician to investigate the health status of a monitored human subject. As an additional enhancement to the above use case (i.e., discharge), the clinician sets an alert level for a patient in recovery where they expect to be alerted if the patient does not walk for at least 20 minutes at least once a day within a week after discharge. If the clinician receives the alert, they will contact the patient and recommend a change in their medical care.
[0148] In yet other examples, these methods can be used to selectively control the activation or deactivation of a device. For example, operations or settings of parameters of a medical device associated with treating or monitoring a monitored human subject can be performed. In a specific example, a patient with Parkinson's disease is enrolled in a cloud-based platform and monitored using a wrist-worn activity monitor. Data from the accelerometer is streamed to the cloud-based platform where it is analyzed according to the present method to detect signs of tremors and gait pattern deterioration as indicated by the obtained activity classification. If the system detects signs of gait pattern deterioration, the patient's medication can be automatically adjusted to compensate. Medical monitoring devices can be adjusted or their parameters changed. These devices can be controlled via an electronic control signal created by a control circuit or other processing devices described herein.
[0149] In still other examples, operations or settings of parameters of a user electronic device associated with treating or monitoring a monitored human subject are controlled. For example, a patient undergoing physical rehabilitation is enrolled in the platform and monitored using a torso-worn activity monitor. Data from the accelerometer is streamed to the platform where the data is analyzed by a model which shows whether the patient is performing physical therapy activities correctly. Based on the analysis, additional instructions or encouragement are pushed to the patient's mobile phone so that they can perform the activities correctly. In various aspects, graphics or icons on the electronic device can be moved, controlled, and / or changed so as to present these instructions or encouragement to the patient in a specific pattern or manner.
[0150] This document describes the preferred embodiments of the present disclosure, including the best mode known to the inventors. It should be understood that the described embodiments are merely exemplary and should not be considered as limiting the scope of the appended claims.
Claims
1. A method for determining and implementing appropriate life improvement actions for a person, the method comprising: iteratively training a first neural network to obtain a first trained neural network and iteratively training a second neural network to obtain a second trained neural network by the following actions: receiving, at the first neural network, a first set of wearable sensor data, the first neural network responsively generating a first feature vector, the first wearable sensor data describing a first physiological characteristic of the person, wherein any markings in the first set of wearable sensor data are ignored; receiving, at the second neural network, a second set of second wearable sensor data, the second neural network responsively generating a second feature vector, the second wearable sensor data describing a second physiological characteristic of the person, wherein any markings in the second set of wearable sensor data are ignored; wherein at least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity occurring at the same time for the same person; and predicting, at a comparison neural network, whether the first feature vector from the first neural network and the second feature vector from the second neural network match or mismatch, determining that the first feature vector and the second feature vector match when the first feature vector and the second feature vector are from the same person and are acquired at substantially the same time; backpropagating an error generated by a cost function to the first neural network and the second neural network, the cost function penalizing a failure to correctly determine whether the vectors match or mismatch, the backpropagation effectively and independently updating the parameters of the first neural network and the second neural network, optimizing the feature generation of the first neural network and the second neural network, and optimizing the prediction success of the comparison neural network; deploying the first trained neural network; monitoring at least one current human subject to obtain current wearable sensor data and applying the current wearable sensor data to the first trained neural network to obtain a current feature vector; mapping, at a trained classifier, the current feature vector to a classification representing one or more activity categories; performing, based on the classification, one or more actions selected from the group consisting of: quantifying, in a clinical trial, the health impact between at least one control group receiving a first intervention and a test group receiving a second intervention; determining possible health changes in the monitored human subject; issuing an alert to a clinician to investigate the health condition of the monitored human subject; selectively controlling the activation or deactivation of a device; controlling the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject; controlling the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
2. The method according to claim 1, wherein the mapping utilizes known and marked vectors from other monitored human subjects.
3. The method according to claim 1, wherein the sensor is a wrist sensor or a chest sensor.
4. The method according to claim 1, wherein the user electronic device is a smartphone, a personal computer, a laptop computer, or a tablet computer.
5. The method according to claim 1, wherein the trained classifier includes a random forest or a third neural network.
6. The method according to claim 1, wherein the first neural network and the second neural network are trained at a central location.
7. The method according to claim 1, wherein the comparison network includes an ensemble of separate comparison networks of different time scales with a fixed input window size.
8. A system for determining and implementing appropriate life improvement actions for a person, the system comprising: a first neural network; a second neural network; a comparison neural network coupled to the first neural network and the second neural network; wherein the first neural network is configured to receive a first set of wearable sensor data, and the first neural network responsively generates a first feature vector, the first clinical data describing a first physiological characteristic of a person, wherein any markings in the first set of wearable sensor data are ignored; wherein the second neural network is configured to receive a second set of second wearable sensor data, and the second neural network responsively generates a second feature vector, the second wearable sensor data describing a second physiological characteristic of the person, wherein any markings in the second set of wearable sensor data are ignored; wherein at least some of the first set of wearable sensor data and the second set of wearable sensor data are obtained from the same human activity occurring at the same time from the same person; and wherein the comparison neural network is configured to predict whether the first feature vector from the first neural network and the second feature vector from the second neural network match or mismatch, and determine that the first feature vector and the second feature vector match when the first feature vector and the second feature vector are from the same person and are acquired at substantially the same time; wherein an error is backpropagated to the first neural network and the second neural network, the error being generated by a cost function that penalizes a failure to correctly determine whether the vectors match or mismatch, the backpropagation effectively and independently updating the parameters of the first neural network and the second neural network to produce a first trained neural network and a second trained neural network; wherein the system further includes a trained classifier and a wearable sensor worn by a current human subject; wherein the first trained neural network is deployed, and current wearable sensor data is obtained from the current human subject through the wearable sensor and applied to the first trained neural network to obtain a current feature vector; wherein at the trained classifier, the current feature vector is mapped to a classification representing one or more activity categories, and one or more actions are performed based on the classification, the actions being selected from the group consisting of: Quantifying health effects between at least one control group receiving a first intervention and a test group receiving a second intervention in a clinical trial; Determining possible health changes in the monitored human subject; Issuing an alert to a clinician to investigate the health condition of the monitored human subject; Selectively controlling the activation or deactivation of the device; Controlling the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject; Controlling the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
9. The system according to claim 8, wherein the mapping utilizes known and labeled vectors from other monitored human subjects.
10. The system according to claim 8, wherein the wearable sensor is a wrist sensor or a chest sensor.
11. The system according to claim 8, wherein the trained first neural network is deployed at a central location.
12. The system according to claim 8, wherein the training occurs at a central location.
13. The system according to claim 8, wherein the trained classifier comprises a random forest or a third neural network.
14. A system for training a neural network, the system comprising: a first neural network; a second neural network; a comparison neural network coupled to the first neural network and the second neural network; a control circuit coupled to the first neural network, the second neural network, and the comparison neural network, the control circuit being configured to: obtain a first set of samples of wearable sensor data from multiple individuals with a first type of wearable sensor; obtain a second set of samples of wearable sensor data from multiple individuals with a second type of wearable sensor, at least some of the samples being matched to the same individual and substantially the same time window as the samples of the first set; train the first neural network to produce a first trained neural network by iteratively performing the following actions: - inputting samples from the first set into the first neural network to generate features; - inputting samples from the second set into the second neural network to generate features; - inputting the features from the first neural network and the features from the second neural network together into the comparison neural network, the comparison neural network responsively predicting whether the input samples match or mismatch; - backpropagating an error generated by a cost function to the first neural network and the second neural network, the cost function penalizing a failure to correctly determine whether the input samples match or mismatch, to independently change the neural parameters of the first neural network and the second neural network, optimizing the feature generation of the first neural network and the second neural network, thereby optimizing the prediction success of the comparison neural network.
15. The system according to claim 14, further comprising a trained classifier, and: wherein the first trained neural network is subsequently deployed and wearable sensor data from a monitored human subject is captured; wherein the wearable sensor data is applied to the first neural network, and the first neural network responsively generates a set of features in response to a time window of the wearable sensor data, wherein the trained classifier maps the generated features to one or more activity categories in the time window; and wherein, based on the classification, one or more actions are performed, the actions being selected from the group consisting of: quantifying, in a clinical trial, the health impact between at least one control group receiving a first intervention and a test group receiving a second intervention; determining possible health alterations in the monitored human subject; issuing an alert to a clinician to investigate the health condition of the monitored human subject; selectively controlling the activation or deactivation of a device; controlling the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject; controlling the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
16. The system according to claim 15, wherein the sensor is a wrist sensor or a chest sensor.
17. The system according to claim 15, wherein the trained classifier comprises a random forest or a third neural network.
18. The system according to claim 15, wherein the training occurs at a central location.
19. A method of training a neural network, the method comprising: obtaining a first set of samples of wearable sensor data from multiple persons with a first wearable sensor; obtaining a second set of samples of wearable sensor data from multiple persons with a second wearable sensor, at least some of the samples being matched to the same person and substantially the same time window as the samples of the first set; training the first neural network by iteratively performing the following actions to produce a first trained neural network: inputting samples from the first set into the first neural network to generate features; inputting samples from the second set into a second neural network to generate features; inputting the features from the first neural network and the features from the second neural network together into a comparison network, the comparison network predicting whether the input samples match or mismatch; backpropagating an error generated by a cost function that penalizes failure to correctly determine whether the input samples match or mismatch to the first neural network and the second neural network to independently update the neural parameters of the first neural network and the second neural network, improving the feature generation of the first neural network and the second neural network, thereby improving the prediction success of the comparison neural network.
20. The method according to claim 18, further comprising: deploying the first trained neural network, capturing wearable sensor data from a monitored human subject; applying the captured wearable sensor data to the first trained neural network and generating, by the first trained neural network, a set of features in response to a time window of the captured wearable sensor data, and Map the generated features to one or more activity categories in the time window via a trained classifier; and Based on the classification, perform one or more actions selected from the group consisting of: Quantify the health impact between at least one control group receiving a first intervention and a test group receiving a second intervention in a clinical trial; Determine possible health alterations in the monitored human subject; Alert a clinician to investigate the health condition of the monitored human subject; Selectively control the activation or deactivation of a device; Control the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject; Control the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject.
21. The method according to claim 20, wherein the first wearable sensor and the second wearable sensor are wrist sensors or chest sensors.
22. The method according to claim 20, wherein the trained classifier comprises a random forest or a third neural network.
23. The method according to claim 20, wherein the training occurs at a central location.
24. A method, the method comprises: Capture wearable sensor data from a monitored human subject; Generate a set of features in response to a time window of the wearable sensor data via a first neural network, and Map the generated set of features to one or more activity categories in the time window via a trained classifier; and based on the activity categories, perform one or more actions selected from the group consisting of: Quantify the health impact between at least one control group receiving a first intervention and a test group receiving a second intervention in a clinical trial; Determine possible health alterations in the monitored human subject; Alert a clinician to investigate the health condition of the monitored human subject; Selectively control the activation or deactivation of a device; Control the operation or setting of parameters of a medical device associated with treating or monitoring the monitored human subject; Control the operation or setting of parameters of a user electronic device associated with treating or monitoring the monitored human subject; wherein the first neural network is created by: Obtain a first set of samples of wearable sensor data from multiple persons with a first type of wearable sensor; Obtain a second set of samples of wearable sensor data from multiple persons with a second type of wearable sensor, at least some of the samples being matched to the same person and substantially the same time window as the samples of the first set; Train the first neural network by iteratively performing the following actions: Input samples from the first set into the first neural network to generate features; Input samples from the second set into a second neural network to generate features; Input the features from the first neural network and the features from the second neural network together into a comparison network, which predicts whether the input samples match or mismatch; Backpropagate the error generated by the cost function to the first neural network and the second neural network, where the cost function penalizes failure to correctly determine whether the input sample is a match or a mismatch, to independently update the neural parameters of the first neural network and the second neural network, improve the feature generation of the first neural network and the second neural network, and thus improve the prediction success of the comparison neural network.
25. The method according to claim 24, wherein the wearable sensor data is obtained from a wrist sensor or a chest sensor.
26. The method according to claim 24, wherein the trained classifier comprises a random forest or a third neural network.
27. The method according to claim 24, wherein the training occurs at a central location.