Foundation Model Training for Physiological Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of digital-based physiological markers using deep neural networks for wearable devices is hindered by the lack of curated datasets with annotated medical labels, limiting the generalizability of learned models to population demographics.
Innovation Solution
Large-scale training of foundation models for physiological signals from wearable devices using self-supervised learning, incorporating stochastic user-level augmentation and momentum training with regularized normalized cross-entropy loss, enables the pre-training of models on unlabeled data, reducing the need for labeled data and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional supervised learning with annotated medical labels is used, then model accuracy for physiological state prediction can be improved, but the cost and time for data collection and annotation increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training foundation models on large-scale unlabeled physiological signal data before fine-tuning on smaller labeled datasets. This pre-training phase prepares the models with general physiological patterns, reducing the time and cost required for subsequent labeled data collection and annotation while maintaining high prediction accuracy.
Solution Approach 2:
The patent implements self-service through self-supervised learning approaches where models learn from unlabeled data without requiring manual annotation. The models automatically generate their own training signals from the raw physiological signals, eliminating the need for expensive and time-consuming medical expert annotation while still achieving high accuracy.
2Adaptability or versatility
If large-scale labeled datasets are collected for training, then model generalizability to population demographics improves, but the computational resources and cost required increase
Solution Approach 1:
The patent uses preliminary action by pre-training models on diverse unlabeled physiological data from multiple population demographics before fine-tuning. This preliminary exposure to varied data distributions enhances model generalizability across different populations while avoiding the computational burden of collecting and processing equally large labeled datasets for each demographic.
Solution Approach 2:
The patent applies universality by developing foundation models that learn universal physiological patterns from unlabeled data across different devices, populations, and signal types. These universal representations can be adapted to multiple downstream tasks and demographic groups without requiring separate large-scale labeled training for each, reducing overall computational resource requirements.
3Reliability
If deep neural networks are trained on wearable device data, then predictive capabilities for health conditions improve, but the lack of curated datasets limits model development
Solution Approach 1:
The patent implements self-service through self-supervised learning where models automatically learn from raw, uncurated physiological signals without requiring manual data curation or annotation. The models generate their own training objectives from the unlabeled data, eliminating the complex data curation process while maintaining reliable predictive capabilities for health conditions.
Solution Approach 2:
The patent applies preliminary action by pre-training foundation models on large volumes of uncurated physiological data to learn robust feature representations. This preliminary training on diverse uncurated data provides a strong foundation that can be subsequently fine-tuned for specific health condition predictions, reducing the need for extensive manual data curation while improving predictive reliability.
Data Source
AI summary
The subject technology provides for large-scale training of foundation models for physiological signals from wearable electronic devices. An apparatus receives receive input data having a plurality of physiological signal information segments associated with a user. The apparatus applies one or more augmentation functions to the plurality of physiological signal information segments to generate an augmented version of the plurality of physiological signal information segments. The apparatus trains a neural network to produce a trained machine learning model by generating, via an encoder, an embedding of the augmented version having a first number of dimensions in an embedding space. The apparatus maps, via a multilayer perceptron projection, the embedding into a representation having a second number of dimensions. The apparatus determines mutual information between a pair of representations of the augmented version. The apparatus can deploy the trained machine learning model to predict a physiological state of the user.


