Biomarker estimation
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-08-13
Smart Images

Figure EP2026051816_13082026_PF_FP_ABST
Abstract
Description
[0001] 2025PF00093
[0002] 1
[0003] BIOMARKER ESTIMATION
[0004] FIELD OF THE INVENTION
[0005] The present invention relates to a machine-learning method for estimating a value of a biomarker using biosignal measurements.
[0006] BACKGROUND OF THE INVENTION
[0007] Xiang Lan et al., (2025): GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images. 10.48550 / arXiv.2503.06073.
[0008] Zengding Liu et al., (2024): Large Language Models for Cuffless Blood Pressure Measurement From Wearable Biosignals. 10.48550 / arXiv.2406.18069.
[0009] Anand Chandrasekhar et al, (2020): PPG Sensor Contact Pressure Should be Taken into Account for Cuff-Less Blood Pressure Measurement. IEEE Transactions on Biomedical Engineering. PP.
[0010] 1-1. 10.1109 / TBME.2020.2976989.
[0011] Michael Moor et al, (2023): "Med-Flamingo: a Multimodal Medical Few-shot Learner", arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853.
[0012] It is known to use trained machine learning models to perform prediction of clinical indicators based on sensor data. This is typically achieved by building application-specific models trained for predicting one type of clinical indicator based on a supervised learning procedure utilizing a large volume of training data.
[0013] The process of training specific models for predicting clinical indicators is time consuming and complex. It requires a process of compiling and annotating appropriate training data in addition to the processes of training and validating the model. Furthermore, each model is limited in its applicability to predicting one clinical indicator, meaning that multiple models must be trained if it is desired to estimate values for a variety of different clinical indicators.
[0014] It would be of advantage to find a machine learning biomarker prediction method that is less time consuming and more flexible.
[0015] SUMMARY OF THE INVENTION
[0016] The invention is defined by the independent claims. The dependent claims define advantageous embodiments.
[0017] An aspect of the invention is a computer-implemented method for estimating a value of at least one biomarker for a subject based on inference from a biosignal for the subject generated by a biosignal sensor. The method comprises using a trained general purpose vision large language model to2025PF00093
[0018] 2
[0019] generate the estimated value of the at least one biomarker, wherein the general purpose vision large language model is operable to generate an output responsive to an input prompt comprising image data. The method further comprises obtaining a visual representation of a waveform of a time-series of values of the biosignal, and prompting the trained vision large language model with an input prompt comprising the visual representation of the waveform.
[0020] The trained general purpose vision large language model is operable to output an estimated value of the at least one biomarker responsive to an input prompt comprising a visual representation of a waveform of a time-series of values of the biosignal. Prompting the trained general purpose vision large language model with an input prompt comprising the visual representation of the waveform thus results in an output comprising an estimated value of the at least one biomarker.
[0021] Multi-modal large language models represent a significant advancement in artificial intelligence, extending the capabilities of traditional text-based language models to process and generate information across multiple modalities, such as text, images, and audio. These models are designed to understand and integrate information from various input types.
[0022] Multi-modal large language models are able to process visual information alongside text. This capability allows these models to analyze images, understand their content, and generate relevant textual descriptions or responses based on visual inputs. For instance, when presented with an image, these models can identify objects, describe scenes, and even answer questions about the visual content.
[0023] It is the recognition of the inventors that these visual reasoning capabilities of vision large language models may be utilized to provide a more efficient and flexible approach to biomarker value estimation. It has been found by the inventors that trained general purpose vision large language models, such as ChatGPT vision, can achieve reliable prediction of biomarkers based on graphical representation of correlated biosignal waveforms provided in an input prompt.
[0024] In prior art systems, such analysis is typically performed using rules-based signal processing or based on biomarker-specific machine learning models.
[0025] According to the invention, the trained large language model is a general-purpose vision large language model. A general -purpose vision large language model is an Al model designed to perform a broad, flexible range of visual tasks for example over any kind of visual input, combined with natural-language abilities. On the other hand, a specialized vision LLM is trained, further trained and / or fine-tuned to perform highly targeted tasks within a specific domain e.g., in the medical field. For achieving clinically acceptable accuracy, the intuitive choice would be a specialized medical model. Nevertheless, the inventors have surprisingly found out that a general-purpose vision large language model can also provide results of clinically acceptable accuracy. An advantage of using a general-purpose vision large language model instead of a specialized one is that no specialized training is required. Thus, downstream training effort and resources are lower and the costly steps traditionally required for training a specialized medical model can be avoided.2025PF00093
[0026] 3
[0027] Embodiments of the present invention provide the advantage that specific individual models do not need to be trained for each target biomarker. Reliable predictions for a range of biomarkers can be achieved using few-shot learning or even a zero-shot approach, using a large language model.
[0028] The biosignal may comprise measurements from one or more sensors configured for sensing one or more biological parameters of a subject, for example one or more physiological parameters. For example, the sensors may include electrocardiogram sensors, photoplethysmography devices, pressure sensors, or other physiological monitoring equipment. The physiological parameters may comprise electrical, optical, mechanical, or chemical signals that reflect biological processes occurring within the subject. The measurements may be acquired as time-series data that captures temporal variations in the physiological parameters being monitored.
[0029] The biomarker may be a biomarker known to be correlated with the biosignal.
[0030] For example, the biosignal and the at least one biomarker may be physiologically related such that variations in the biosignal correspond to, or are correlated with, changes in the biomarker value. The biosignal may be derived from measurements of a biological system or physiological process that influences or determines the biomarker being estimated. A correlation may exist between temporal patterns or characteristics of the biosignal and the quantitative value of the biomarker.
[0031] The visual representation of a waveform of a time-series of values of the biosignal may comprise an image of a waveform of a time-series of values of the biosignal. The visual representation of a waveform of a time-series of values of the biosignal may comprise one or a plurality of images of a waveform of a time-series of values of the biosignal. In some embodiments, the visual representation of a waveform of a time-series of values of the biosignal may comprise a series of images, for example a time series of images, for example a series of image frames, wherein each of the series of images or image frames represents a state of the waveform at a respective time point.
[0032] In some embodiments, the at least one biomarker comprises a hemodynamic parameter, such as an arterial blood pressure measure, and the biosignal comprises a pressure signal indicative of a pressure between a body part comprising an artery and a compressing surface applied to the body part over the artery. The method may comprise obtaining the pressure signal from a pressure sensor comprised by a hemodynamic measurement device comprising a body-mountable cuff, wherein the cuff comprises the compressing surface. In some embodiments, the cuff comprises an actuator for applying a controllable pressure to the body part and a shell portion between the actuator and the body part, with the pressure sensor disposed between the shell portion and the body part.
[0033] In accordance with this arrangement, the pressure signal obtained by the pressure sensor corresponds to a pulsatile arterial pressure waveform for the subject. The pressure signal may be a continuous pressure waveform which is indicative of an arterial pulsation signal for the subject, e.g. a pulse wave signal for the subject.
[0034] Using the obtained tissue pressure signal, various hemodynamic parameters may be inferred.2025PF00093
[0035] 4
[0036] Thus, in some embodiments, the at least one biomarker which is estimated using the large language model may comprise at least one hemodynamic parameter.
[0037] In some embodiments, the at least one biomarker which is estimated using the large language model may comprise a measure of arterial blood pressure. For example, in some embodiments, the at least one biomarker may comprise a measure of mean arterial pressure (MAP). Additionally or alternatively, the at least one biomarker may comprise a measure of systolic blood pressure (SBP).
[0038] Additionally or alternatively, the at least one biomarker may comprise a measure of diastolic blood pressure (DBP).
[0039] Additionally or alternatively, in some embodiments, the at least one biomarker which is estimated may comprise one or more of: heart rate, cardiac output, and stroke volume.
[0040] In some embodiments, the at least one biomarker is one or more vital signs.
[0041] In some embodiments, the at least one biosignal may comprise a photoplethysmogram (PPG) signal. In this case, the at least one biomarker may comprise one or more hemodynamic, cardiovascular or physiological parameters. For example, the at least one biomarker may comprise one or more hemodynamic parameters (e.g. heart rate, pulse rate variability, pulse wave velocity, pulse transit time), one or more blood or tissue oxygenation parameters (blood oxygen saturation (SpO2), perfusion index), one or more cardiovascular health indicators (e.g. arterial stiffness measures, vascular compliance), one or more respiratory parameters (e.g. respiration rate, derived from PPG variability), and / or one or more autonomic function indicators (e.g. heart rate variability).
[0042] In some embodiments, the input prompt to the large language model may further comprise a request for an estimated value of the at least one biomarker. In some embodiments, the large language model is a multi-modal large language model operable to generate an output responsive to an input prompt comprising one or both of image data and text, and the input prompt comprises a natural language text representation of the request.
[0043] The large language model can be prompted using a zero-shot approach or a few-shot approach. Neither approach requires fine-tuning of the model.
[0044] In accordance with the few-shot approach, the input prompt may comprise one or more training examples, each training example comprising an example visual representation of a waveform of the biosignal and a corresponding example value of the at least one biomarker. This few-shot learning approach can improve the model's performance at predicting the biomarker without requiring extensive retraining.
[0045] The method may comprise providing the one or more training examples within the same input prompt containing the visual representation of the biosignal waveform, or in a separate prompt but within the same context window or conversation context of the large language model as the prompt containing the visual representation of the biosignal waveform. More particularly, the training examples may be provided in the same input prompt containing the visual representation of the biosignal waveform or in a preceding prompt within the same context window of the large language model.2025PF00093
[0046] 5
[0047] In some embodiments, the one or more training examples may comprise a plurality of training examples, for example, three or more training examples, for example fifteen or more training examples.
[0048] In some embodiments, the input prompt may include an instruction to compare the visual representation of the biosignal waveform (which is to be analyzed) against each of the training examples. It has been found that this may improve model performance in estimating the biomarker.
[0049] The input prompt may include an instruction to determine which of the plurality of training examples the visual representation of the waveform (which is to-be-analyzed) is most similar to. The similarity assessment may be based on morphological similarity for example. It has been found that this too may improve model performance in estimating the biomarker.
[0050] In some embodiments, the input prompt may include an instruction to identify one or more waveform features comprised by the waveform. The input prompt may include an instruction to estimate the biomarker based on the identified one or more waveform features. The input prompt may include an instruction to count the number of a specified type of waveform feature within a specified temporal section of the visually represented waveform. The specified type of waveform feature may include any one or more of: waveform peaks, waveform troughs, zero-crossing points, and inflection points. For example, if the biomarker to be estimated comprises a heart rate of the subject, the input prompt may include an instruction to identify each peak of the waveform. It may include an instruction to count the number of peaks within a specified temporal section of the visually represented waveform.
[0051] In some embodiments, the input prompt may include an instruction to provide reasoning for the estimate of the biomarker. It has been found that this may improve model performance at estimating the biomarker.
[0052] In some embodiments, the method may comprise receiving a time series of values of the biosignal and generating the visual representation of the waveform of biosignal using the time series of values of the biosignal.
[0053] A further aspect of the invention is a computer program product comprising computer program code configured, when run on a processor, to cause the processor to perform a method in accordance with any embodiment described in this disclosure or in accordance with any claim.
[0054] A further aspect of the invention is a processing device comprising one or more processors configured to perform a method in accordance with any embodiment described in this disclosure or in accordance with any claim.
[0055] A further aspect of the invention is a system comprising a sensor operable to measure a time series of values of a biosignal for a subject, and the processing device described above, arranged to receive the time series of biosignal values from the biosignal sensor.
[0056] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment s) described hereinafter.2025PF00093
[0057] 6
[0058] BRIEF DESCRIPTION OF THE DRAWINGS
[0059] For a beter understanding of the invention, and to show more clearly how it may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0060] Fig. 1 outlines steps of an example method in accordance with one or more embodiments of the invention;
[0061] Fig. 2 is a block diagram of an example processing device and system in accordance with one or more embodiments of the invention; and
[0062] Fig. 3 schematically illustrates an example hemodynamic parameter measurement device comprising a tissue pressure sensor for acquiring a biosignal in the form of a tissue pressure signal.
[0063] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The invention will be described with reference to the Figures.
[0065] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems and methods of the present invention will become beter understood from the following description, appended claims, and accompanying drawings. It should be understood that the Figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the Figures to indicate the same or similar parts.
[0066] Embodiments of the invention provide a method for predicting values of at least one biomarker by prompting a trained large language model (LLM) with a visual representation of a waveform of values of a biosignal known to be correlated with the biomarker.
[0067] In the context of this disclosure, a biomarker may refer to a measurable biological indicator that provides information about a physiological or pathological process within a subject. A biomarker may comprise a quantitative parameter that can be objectively measured and evaluated as an indicator of normal biological processes, pathogenic processes, or pharmacologic responses to therapeutic intervention. The biomarker may include one or more hemodynamic parameters such as blood pressure measures, heart rate values, cardiac output measures, or stroke volume determinations. The biomarker may alternatively comprise one or more respiratory parameters, neurological indicators, or metabolic measures. The biomarker may be expressed as a numerical value with associated units of measurement corresponding to the specific physiological parameter being quantified. Alternatively, the biomarker may take the form of a classification or class label selected from a set of possible classifications.
[0068] In the context of this disclosure, a biosignal may refer to an electrical, optical, mechanical, or chemical signal measurable from a biological system and which is indicative of one or more physiological processes occurring within a subject. The biosignal may be acquired using one or more sensors positioned on or near the subject's body. The biosignal may comprise time-varying2025PF00093
[0069] 7
[0070] measurements that exhibit characteristic patterns or waveforms corresponding to underlying physiological activities. Non-limiting examples of biosignals may include pressure signals from tissue pressure sensors, photoplethysmography signals from optical sensors, electrocardiogram signals from electrical sensors, or respiratory signals from airflow sensors. The biosignal may be processed as a time series of digital values sampled at a predetermined frequency.
[0071] In the context of this disclosure, the term "large language model" may refer to a neural network model that comprises a transformer architecture with a plurality of attention layers. The model may comprise a parameter count of at least one billion trainable parameters. The model may comprise one or more multi-head self-attention mechanisms to process sequential input data. A self-attention mechanism may enable the model to compute attention weights between different positions in an input sequence. The attention weights may allow the model to capture long-range dependencies and contextual relationships within the input data.
[0072] The large language model may be a multimodal model operable to process multiple input modalities, for example including text data and image data, within a single inference session and / or context window. To implement the multimodal functionality, the model may comprise one or more encoding layers configured to convert visual input data into token representations compatible with the transformer architecture. The model may include one or more cross-attention mechanisms, wherein each cross-attention mechanism is configured to compute attention weights between elements of a first input modality and elements of a second input modality. This may thereby enable the model to establish correspondences between visual features and textual tokens for example.
[0073] The large language model may be pre-trained on one or more training datasets using selfsupervised learning objectives. The one or more training datasets may be large-scale datasets, for example the one or more datasets collectively comprising at least one billion tokens. The one or more datasets may comprise diverse, unlabeled text and image data from multiple domains rather than task-specific labeled training data. The self-supervised learning objectives may include next-token prediction. The pre-training may enable the model to develop general-purpose representations that can be applied to downstream tasks without task-specific retraining.
[0074] Traditionally, physiological waveforms are processed as time series data using signal processing techniques. Traditional signal analysis methods involve application of task-specific signal processing algorithms configured to estimate biomarkers based on signal data. Embodiments of the present invention are based on the recognition by the inventors that multimodal large language models (LLMs) can utilize the visual representation of these waveforms, such as images or videos, for biomarker estimation. This approach leverages the rich information contained in the visual representation, enabling inference of biomarkers using few-shot or even zero-shot learning.
[0075] Fig. 1 outlines steps of an example computer-implemented method 10 according to one or more embodiments. The steps will be recited in summary, before being explained further in the form of example embodiments.2025PF00093
[0076] 8
[0077] The method 10 is for estimating a value of at least one biomarker for a subject based on inference from a biosignal for the subject generated by a biosignal sensor. The method comprises using a trained large language model (LLM) to generate the estimated value of the at least one biomarker, wherein the large language model is operable to generate an output responsive to an input prompt comprising image data.
[0078] The method 10 comprises obtaining 12 a visual representation of a waveform of a timeseries of values of the biosignal, and prompting 14 the trained large language model (LLM) with an input prompt comprising the visual representation of the waveform. A biomarker value is output 16 responsive to the input prompt to the LLM.
[0079] The visual representation of a waveform of a time-series of values of the biosignal may comprise an image of a waveform of a time-series of values of the biosignal. The visual representation of a waveform of a time-series of values of the biosignal may comprise one or a plurality of images of a waveform of a time-series of values of the biosignal. In some embodiments, the visual representation of a waveform of a time-series of values of the biosignal may comprise a series of images, for example a time series of images, for example a series of image frames, wherein each of the series of images or image frames represents a state of the waveform at a respective time point.
[0080] The large language model may be a multimodal large language model.
[0081] Multi-modal large language models represent a class of foundation models that have been trained on diverse datasets spanning multiple data modalities including text, images, and other input types. These models are characterized by their general-purpose architecture and broad training objectives, distinguishing them from task-specific or domain-specific machine learning models.
[0082] As noted above, the method can also be embodied in hardware form, for example in the form of a processing device which is configured to carry out a method in accordance with any example or embodiment described in this document, or in accordance with any claim of this application.
[0083] To further aid understanding, Fig. 2 presents a schematic representation of an example processing device 32 configured to execute a method in accordance with one or more embodiments of the invention. The processing device is shown in the context of a system 30 which comprises the processing device. The processing device alone represents an aspect of the invention. The system 30 is another aspect of the invention.
[0084] The processing device 32 comprises one or more processors 36 configured to perform a method in accordance with that outlined above, or in accordance with any embodiment described in this document or any claim. In the illustrated example, the processing device further comprises an input / output or communication interface 34.
[0085] In the illustrated example of Fig. 2, the system 30 further comprises a biosignal sensor 42 operable to measure a time series of values of a biosignal for a subject. In some embodiments, this may comprise an arterial blood pressure measurement device. Additionally or alternatively, in some embodiments, this may comprise a PPG sensor.2025PF00093
[0086] 9
[0087] The system 30 may further comprise a memory 38 for storing computer program code (i.e. computer-executable code) which is configured for causing the one or more processors 36 of the processing unit 32 to perform the method as outlined above, or in accordance with any embodiment described in this disclosure, or in accordance with any claim.
[0088] As mentioned previously, the invention can also be embodied in software form. Thus another aspect of the invention is a computer program product comprising computer program code configured, when run on a processor, to cause the processor to perform a method in accordance with any embodiment of the invention described in this document, or in accordance with any claim.
[0089] The processing device is configured to communicate with a remote server which stores the LLM. The prompting of the LLM may comprise communicating with the remote server to transmit a message containing a data representation of the input prompt and receiving a return message from the remote server containing a data representation of the output of the LLM. Alternatively, the LLM may be stored locally on the memory 38.
[0090] In some embodiments, the system may comprise a patient monitoring system operatively coupled with the biosignal sensor 42. The patient monitoring system may be configured to compile, in real time, a data log of measured values of the biosignal generated by the biosignal sensor. The patient monitoring system may optionally include a display device, and the patient monitoring system may be configured to generate and display, using the display device, a visual representation of a waveform of a time-series of values of the biosignal, based on the data in the data log, and to update the displayed waveform at regular time intervals, in real time. The displayed waveform may therefore represent a time series of the most recently acquired time series of values of the biosignal generated by the biosignal sensor for the subject.
[0091] In some embodiments, the visual representation of a waveform of a time-series of values of the biosignal which is input to the trained LLM may be populated by the same set of biosignal values which is used to form the currently displayed version of the waveform displayed by the patient monitoring device. In other words, the method may in some embodiments comprise displaying a real-time biosignal waveform on a display device of the patient monitoring device, and may be configured to prompt the LLM using a visual representation of the same waveform (a waveform populated by the same biosignal data points) as the waveform being displayed on the display device. The visual representation of the waveform which is input to the LLM may be generated by accessing the biosignal data points stored in the data log of the patient monitoring device.
[0092] Regarding the biomarker, this may be a biomarker known to be correlated with the biosignal.
[0093] For example, the biosignal and the at least one biomarker may be physiologically related such that variations in the biosignal correspond to, or are correlated with, changes in the biomarker value. The biosignal may be derived from measurements of a biological system or physiological process that2025PF00093
[0094] 10
[0095] influences or determines the biomarker being estimated. A correlation may exist between temporal patterns or characteristics of the biosignal and the quantitative value of the biomarker.
[0096] The biosignal may comprise measurements from one or more sensors configured to detect physiological activities which are causally or functionally linked to the biomarker parameter. For example, the biosignal may be a signal indicative of a functioning or property or operation of at least one biological system or organ, and the biomarker may be an indicator related to a functioning of the same biological system or organ. For example, the biosignal may represent electrical, mechanical, optical, or pressure-based measurements from a cardiovascular system, and the biomarker may comprise a hemodynamic parameter derived from the same cardiovascular system.
[0097] The relationship between the biosignal and biomarker might be established through physiological principles, clinical validation studies, or empirical observations demonstrating predictive correlation between the measured signal characteristics and the target biomarker values.
[0098] In accordance with one example set of embodiments, the at least one biomarker which is estimated by the method may comprise an arterial blood pressure measure. This may for example include one or more of: mean arterial pressure (MAP), systolic blood pressure (SBP), and diastolic blood pressure (DBP).
[0099] In some examples, the biosignal may comprise an arterial blood pressure (ABP) waveform signal, or a signal indicative thereof.
[0100] By way of example, in some embodiments, the biosignal may comprise a pressure signal indicative of a pressure between a body part comprising an artery and a compressing surface applied to the body part over the artery. The method may comprise obtaining the pressure signal from a pressure sensor comprised by a hemodynamic measurement device comprising a body-mountable cuff, wherein the cuff comprises the compressing surface. For example, the cuff may comprise an actuator for applying a controllable pressure to the body part. For example, the cuff may include an inflatable portion having a controllable inflation level, and wherein a pressure applied to the body part is controllable by adjusting the inflation level. However, other implementations of the actuator are also possible, such as a retractable wire configuration.
[0101] In one example set of implementations, the body-mountable cuff may further comprise a shell portion between the actuator and the body part, and wherein the pressure sensor is disposed between the shell portion and the body part.
[0102] An example implementation of such a body-mountable cuff is schematically represented in Fig. 3.
[0103] The illustration shows a cross-sectional view of the cuff when applied to the upper arm 60 of a patient. An artery 58 of the patient is schematically shown. The apparatus has a design which differs from the most standard cuff-based blood pressure measurement apparatuses in that it includes a dedicated tissue pressure sensor 54 which is held against the tissue of the body part by the cuff when the cuff is inflated. The cuff includes a pneumatic actuator in the form of an inflatable bladder 52 for use in2025PF00093
[0104] 11
[0105] changing a pressure applied to the body part 10 by the cuff. The tissue pressure sensor 54 is arranged for sensing a pressure between a surface of the user’s body part and the cuff. Independently of the tissue pressure sensor, one or more operational parameters of the pneumatic actuator can be sampled. For example, a pressure inside the inflatable bladder can be sensed. An actuator activity signal indicative of a pumping power level or pumping rate of a pump of the pneumatic actuator can also be sampled.
[0106] Standard blood pressure measurement cuffs use the bladder air pressure as the sole measurement signal for sensing blood pressure. By additionally utilizing a tissue pressure sensor 54, the apparatus shown in Fig. 3 can achieve superior quality results and also enables the measurement of advanced hemodynamic parameters such as stroke volume, cardiac output and fluid responsiveness in addition to blood pressure. In particular, in standard blood pressure cuffs, damping air is assumed to destroy > 90% of the amplitude and contour of a tissue pressure pulse wave. By contrast, by using a dedicated tissue pressure sensor, the design illustrated in Fig. 3 allows recording high-fidelity (HiFi) arterial blood pressure and pulse waveform.
[0107] The tissue pressure sensor 54 may be implemented as, e.g. a fluid-filled pad which provides a conforming layer of fluid between the cuff bladder 52 and the surface of the user’s body. In this way, the integrated pneumatic actuator enables hydraulic coupling of the tissue pressure sensor 54 to the upper arm tissue. When blood pulses in the artery, this results in a pressure wave 56 which can be detected by the tissue pressure sensor 54. The tissue pressure sensor 54 can be connected to a pressure transducer via a fluid-filled conduit. The pressure transducer converts the pressure in the fluid into an electrical signal.
[0108] The operating principle is based on coupling of the pressure sensor to the body part (e.g. arm) in order to transcutaneously record tissue pressure pulse waves that result from arterial pulsation (e.g. in the brachial artery). In similar fashion to a conventional upper arm blood pressure cuff, the cuff compresses the upper arm using an integrated pneumatic actuator with rising clamping pressure.
[0109] In the example illustrated in Fig. 3, the cuff includes a shell portion 53, wherein the shell portion is arranged so as to be located between the pneumatic actuator 52 and the body part when the cuff is worn on the body part and arranged to surround the body part when the cuff is worn. The tissue pressure sensor 54 is arranged so as to be positioned between the shell 53 and the body part when the device is worn on the body part. The shell structure may be relatively stiff. Consequently, the measurement accuracy can be improved since the tissue pressure is measured against a relatively stiff support which prevents dampening of the amplitude and shape of the signal. In particular, if a tissue pressure sensor unit 54 is located at least partially between the shell portion 53 and the body part, the pulsatile component of the arterial pressure is sensed with high accuracy by the tissue pressure sensor unit.
[0110] The shell portion 53 may include overlapping sections, and wherein the overlapping sections of the shell portion may move or slide relatively to each other, thereby reducing the diameter of the shell portion responsive to increasing applied pressure by the pneumatic actuator.2025PF00093
[0111] 12
[0112] The design described above permits non-invasive hemodynamic monitoring and enables measurement of blood pressure, as well as cardiac output and other hemodynamic parameters.
[0113] For more expansive details about this example hemodynamic parameter measurement apparatus, reference is made to the document EP2953528 Al which describes the cuff design in more detail.
[0114] In one set of embodiments, the tissue pressure signal output by the tissue pressure sensor 54 of the above-described apparatus may be used as the biosignal for input to the LLM for predicting the at least one biomarker. The at least one biomarker may comprise a hemodynamic parameter, such as one or more of: mean arterial pressure (MAP), systolic blood pressure (SBP), diastolic blood pressure (DBP), cardiac output (CO), or stroke volume (SV).
[0115] The method may comprise plotting a waveform of measurements output by the tissue pressure sensor 54 as a function of time. This process may involve receiving a time series of digital values representing the tissue pressure signal from the tissue pressure sensor 54. The time series may be sampled at a predetermined sampling rate, for example between 100 Hz and 1000 Hz. The digital values may be processed to extract a segment corresponding to one or more cardiac cycles. The extracted segment may be plotted as a function of time on the x-axis and pressure amplitude on the y-axis to generate a waveform representation.
[0116] The method may further comprise generating a visual representation of the plotted waveform. This may be performed by utilizing graph plotting software or algorithms to render the waveform data as a visual image. The plotting process may involve scaling the axes appropriately to display the waveform features clearly. The visual representation may be generated in a standard image format such as PNG, JPEG, or bitmap. The image may include axis labels, gridlines, and scaling information to provide context for the waveform data. The resulting image may be formatted to optimize clarity and readability for processing by the large language model.
[0117] The method may comprise generating a prompt comprising the image and a request for inference of the biomarker, the request provided in text form. The text request may specify the particular biomarker or biomarkers to be estimated from the waveform. The text component may be formatted as a natural language query directed to the large language model.
[0118] Additionally or alternatively, in some embodiments, the biosignal may comprise a photoplethysmography (PPG) signal obtained from a PPG sensor.
[0119] A PPG sensor is configured to emit light at one or more wavelengths into tissue and detect the amount of light absorbed or reflected by the tissue. The detected light variations correspond to changes in blood volume within the tissue during cardiac cycles. The PPG signal may be acquired from various body locations, such as a fingertip, wrist, earlobe, or forehead.
[0120] In accordance with this set of embodiments, the target biomarker, which is inferred from the PPG signal, using the LLM, may comprise one or more hemodynamic, cardiovascular or2025PF00093
[0121] 13
[0122] physiological parameters. In one set of embodiments, the at least one biomarker may comprise a measure of heart rate.
[0123] The visual representation of the PPG waveform may be generated by plotting the PPG signal amplitude as a function of time over one or more cardiac cycles. The large language model may be prompted with this visual representation along with a text request to estimate the heart rate from the displayed PPG waveform.
[0124] In addition to or instead of predicting a heart rate from the PPG waveform, the LLM may be used to estimate a value of one or more other biomarkers. By way of non-limiting example, these may include one or more of: heart rate variability (HRV), pulse transit time (PTT), pulse wave velocity (PWV), pulse rate variability (PRV), blood oxygen saturation (SpO2), blood pressure (systolic, diastolic, mean arterial pressure), cardiac output (CO), and stroke volume (SV).
[0125] As discussed above, embodiments of the invention involve prompting 14 the trained large language model with an input prompt comprising the visual representation of the waveform.
[0126] By way of non-limiting example, in one example implementation, the inventors utilized GPT-4o with vision capabilities. In one example implementation, the standard GPT-4o model was utilized, without any fine-tuning or other modifications to the base model as provided by OpenAI.
[0127] However, it is also a possibility to fine-tune a multimodal LLM to better adapt to the target prediction domain. With sufficient data, the model can be retrained entirely. With more limited data, it can be fine-tuned using parameter-efficient fine-tuning (PEFT) methods such as LoRA (Low-Rank Adaptation).
[0128] GPT-4o employs a transformer-based neural network architecture that utilizes multi-head self-attention mechanisms to process sequential data across different modalities. The model comprises multiple transformer layers, each containing attention heads that can simultaneously attend to different aspects of the input data, enabling the model to capture relationships within and between text and image inputs. The architecture incorporates positional encoding schemes that allow the model to understand spatial relationships in visual data and temporal sequences in text data. Vision processing capabilities are integrated through specialized encoding layers that convert image data into token representations compatible with the transformer architecture. Cross-attention mechanisms allow the model to establish connections between visual and textual information, facilitating the integrated analysis of multi-modal inputs such as biosignal waveform images paired with natural language instructions for biomarker estimation.
[0129] However, the use of GPT-4o represents just one non-limiting example, and the embodiments of the invention may be implemented using any trained multi-modal large language model with vision capabilities. Further examples in this respect will be discussed later in this disclosure.
[0130] In accordance with one example implementation, a trained LLM is prompted using an input prompt which comprises the visual representation of the waveform of a time-series of values of the biosignal, in combination with a request for an estimated value of the at least one biomarker.2025PF00093
[0131] 14
[0132] By way of example, the large language model may comprise a multi-modal large language model operable to generate an output responsive to an input prompt comprising both image data and text, and wherein the input prompt comprises a natural language text representation of the request.
[0133] The model may be prompted using a zero-shot approach or a few-shot approach.
[0134] In accordance with a few-shot approach, the provided input prompt further includes one or more examples of biosignal waveforms and paired biomarker values. In other words, the method comprises prompting a multimodal LLM model with paired target biomarker values (e.g., mean blood pressure or heart rate) and the corresponding biosignal waveforms from which the biomarker values were derived (e.g. using standard signal processing techniques). The visual representation of the biosignal waveform, the request for estimation of the biomarker, and the set of one or more examples can all be provided in a single prompt. Alternatively, the examples may be provided in a separate preliminary prompt within the same context window in some examples. However, it has been found that the results are not different using multiple prompts compared to using one prompt.
[0135] In a zero-shot approach, the provided input prompt comprises the visual representation of the biosignal waveform, and the request for estimation of the biomarker without any examples.
[0136] By way of one example which is in accordance with the zero-shot approach, the input prompt to the LLM may include a single image of a waveform of a biosignal, and natural language text which instructs the model to:
[0137] - analyze the biosignal waveform, optionally further stating a temporal section of the waveform to consider, e.g. between 20 and 35 seconds;
[0138] - estimate a value of a target biomarker, e.g. heart rate; and
[0139] - briefly explain the reasoning.
[0140] In some embodiments, the input prompt may include an instruction to identify one or more waveform features comprised by the waveform. The input prompt may include an instruction to estimate the biomarker based on the identified one or more waveform features. The input prompt may include an instruction to count the number of a specified type of waveform feature within a specified temporal section of the visually represented waveform. The specified type of waveform feature may include any one or more of: waveform peaks, waveform troughs, zero-crossing points, and inflection points.
[0141] By way of example, the text part of two example input prompts, in accordance with the zero-shot approach are set out below. In both cases, the text prompt is submitted in combination with an image of the biosignal waveform. All images were passed as base64-encoded within the prompt. Both examples relate to an example implementation in which the biosignal comprises a tissue pressure signal output by a tissue pressure sensor 54, e.g. as comprised by a hemodynamic parameter measurement device in accordance with the example of Fig. 3 discussed above. However, it will be appreciated that the same principles may be applied to analysis of a different type of biosignal. Furthermore, in the following2025PF00093
[0142] 15
[0143] examples, the target biomarker is a heart rate value for the subject. However, it will be appreciated that the same principles may be applied to estimation of a different type of biomarker.
[0144] In the first example below, the prompt includes instructions indicative of one or more signal analysis steps to be performed as part of estimating the biomarker value. In particular, the prompt includes an instruction to count the number of distinct peaks and determine the heart rate (HR) based on this. Furthermore, in this example, the visual representation of the waveform comprises a visual representation of the AC component of the tissue pressure signal which has been extracted in a signal preprocessing step.
[0145] Example prompt 1 :
[0146] # prompt =
[0147] # I will provide the AC component of the Tissue Pressure (TP) signal, which represents the pulsatile part of the waveform.
[0148] # Please analyze the pattern of this AC waveform between 20 and 35 seconds, count the number of distinct peaks in that interval, and calculate the Heart Rate (HR).
[0149] # Provide the peak count, the calculated heart rate, and a brief explanation of your reasoning based on the waveform pattern.
[0150] In accordance with the following second example, the visual representation of the waveform comprises an image of the whole tissue pressure signal waveform (including the DC baseline component). Furthermore, in this example, no instructions are provided with regards to any particular signal processing steps to perform; the model is simply instructed to analyze the pattern of the waveform within a pre-defined temporal segment of the signal.
[0151] Example prompt 2:
[0152] prompt =
[0153] You are provided the full tissue pressure waveform (TP_Full waveforms) between 20 and 35 seconds.
[0154] Please analyze the pattern of this waveform between 20 and 35 seconds and calculate the Heart Rate (HR).
[0155] Provide the calculated heart rate, and a brief explanation of your reasoning based on the waveform pattern.2025PF00093
[0156] 16
[0157] By way of further example, the text part of one example input prompt, in accordance with a few-shot approach, is set out below.
[0158] In accordance with this example, the input prompt to the LLM includes multiple training examples, each training example comprising an example visual representation of a waveform of the biosignal and a corresponding example value of the at least one biomarker. In particular, the following example relates to an example implementation in which the biosignal comprises a tissue pressure (TP) signal output by a tissue pressure sensor 54, e.g. as comprised by a hemodynamic parameter measurement device in accordance with the example of Fig. 3 discussed above. However, it will be appreciated that the same principles may be applied to analysis of a different type of biosignal. Furthermore, in the following examples, the target biomarker is a heart rate value for the subject. However, it will be appreciated that the same principles may be applied to estimation of a different type of biomarker.
[0159] The following example prompt includes multiple training examples of waveform images with corresponding heart rate values (e.g., Measurement ID: 001, HR: 72 bpm). In addition, the prompt includes an image of a biosignal waveform to be analyzed and natural language text which instructs the model to:
[0160] - analyze the biosignal waveform, optionally further stating a temporal section of the waveform to consider, e.g. between 20 and 35 seconds;
[0161] - estimate a value of a target biomarker, e.g. heart rate;
[0162] - compare the waveform to be analyzed with the training examples; and - briefly explain the reasoning.
[0163] In particular, the input prompt includes an instruction to compare the visual representation of the biosignal waveform (which is to be analyzed) against each of the training examples. Furthermore, in this example, the input prompt includes an instruction to determine which of the plurality of training examples the visual representation of the waveform (which is to-be-analyzed) is most similar to. The similarity assessment may be based on morphological similarity for example. It has been found that this may improve model performance in estimating the biomarker.
[0164] Example prompt 3 :
[0165] def generate_prompt(training_set, test_set):
[0166] !!!!!!
[0167] Generates a structured prompt for training and test measurements,
[0168] including instructions to count peaks and calculate HR.
[0169] !!!!!!
[0170] prompt =
[0171] I will provide the AC component of the Tissue Pressure (TP) signal, which represents the pulsatile part of the waveform.2025PF00093
[0172] 17
[0173] For the following **training measurements**, I have included the AC waveform (from 20 to 35 seconds) along with the corresponding Heart Rate (HR) values:
[0174] # Add training examples with AC waveforms and HR (from pr_am)
[0175] for meas in training_set:
[0176] prompt += (
[0177] f 'Measurement ID: {measf'MeasID']}, Heart Rate: {meas['pr_am']} bpm\n"
[0178] f1[TP_AC_{meas['MeasID'] } ,png]\n"
[0179] )
[0180] # Add test set description
[0181] # prompt +=
[0182] # Now, for the **test measurement(s)** shown below, please analyze the AC waveform **between 30 and 40 seconds**, count the number of peaks, and calculate the Heart Rate (HR)
[0183] # Please return:
[0184] # 1. The number of peaks counted
[0185] # 2. The estimated heart rate (HR)
[0186] # 3. A brief explanation of your reasoning.
[0187] # Test Measurement(s):
[0188] prompt +=
[0189] Now, for the test measurement(s), analyze the waveform between 20 and 35 seconds.
[0190] Compare the shape and peak frequency to the training waveforms if helpful.
[0191] Please return:
[0192] 1. Number of peaks
[0193] 2. Estimated Heart Rate
[0194] 3. Which training image(s) it most resembles
[0195] 4. Reasoning2025PF00093
[0196] 18
[0197] for meas in test_set:
[0198] prompt += f - [TP_AC_{meas['MeasID']}.png]\n"
[0199] return prompt
[0200] As mentioned above, the method comprises obtaining a visual representation of a waveform of a time-series of values of the biosignal, and prompting the trained large language model with an input prompt comprising the visual representation of the waveform. The method may comprise receiving a time series of values of the biosignal and generating (e.g. by rendering) the visual representation of the waveform of the time series of values of the biosignal. In example implementations, it has been found that no signal pre-processing steps are necessary. The visual representation may be a representation of a waveform of the raw biosignal measurement data.
[0201] Results obtained from two example test implementations in accordance with embodiments of the invention will now be discussed.
[0202] In both test implementations, the OpenAI GPT-4o model having vision capability was utilized as the trained LLM.
[0203] In a first test implementation, the LLM was used to estimate values for mean arterial pressure (MAP) from a biosignal corresponding to a tissue pressure signal output by a tissue pressure sensor comprised by a hemodynamic parameter measurement apparatus such as that described with reference to Fig. 3 earlier.
[0204] For this example implementation, a subset of data was randomly selected from the human study described in the following paper: J Briegel et al., “Clinical evaluation of a high-fidelity upper arm cuff to measure arterial blood pressure during noncardiac surgery,” Anesthesiology, vol. 133, no. 5, pp.
[0205] 997-1006, 2020.
[0206] The subset of data consisted of 35 measurements (one per patient), with femoral artery catheter measurements used as the invasive blood pressure (IBP) reference. Cuff tissue pressure waveforms and reference IBP were recorded during non-cardiac surgery.
[0207] The target biomarker to be predicted in this example was mean arterial pressure (MAP). A test was performed for 20 patients in both zero-shot and few-shot learning scenarios. In the zero-shot case, the LLM was prompted to predict MAP without previously showing the model any example tissue pressure waveforms. In the few-shot case, the model was presented in the same context window with 15 patient examples, each including an example tissue pressure waveform and corresponding MAP value, before being prompted to predict MAP for 20 unseen patients. The estimated MAP values generated by the model in both scenarios were compared to the invasive reference MAP measurements.
[0208] The performance of the LLM in both the zero-shot and few-shot approaches were evaluated by computation of the correlation coefficient (r), bias, and precision errors between estimated values and reference invasive blood pressure measurements.2025PF00093
[0209] 19
[0210] Table 1 below shows these results.
[0211]
[0212] Table 1
[0213] The results in the table are indicative of good performance, especially considering the wide test set range of invasive MAP from 70 to 98 mmHg.
[0214] The results demonstrate a significant improvement in performance when using few-shot learning compared to zero-shot learning for mean arterial pressure (MAP) estimation. The correlation coefficient (r) increased substantially from 0.24 in the zero-shot scenario to 0.61 in the few-shot scenario, indicating a much stronger relationship between the LLM predictions and the invasive reference measurements.
[0215] The error metrics show marked improvement with few-shot learning. The bias error decreased from 10.1 mmHg to 4.9 mmHg, representing a reduction of over 50% in systematic error. Similarly, the precision error improved from 8.9 mmHg to 3.8 mmHg, indicating significantly reduced random variability in the predictions. These improvements demonstrate that providing the large language model with training examples substantially enhances its ability to accurately estimate blood pressure values from tissue pressure waveforms.
[0216] The few-shot results are particularly encouraging for clinical applications, as the bias and precision errors fall within acceptable ranges for many clinical monitoring scenarios. The correlation of 0.61 indicates a moderate to strong relationship between predicted and actual MAP values, suggesting the method has practical utility for non-invasive blood pressure monitoring.
[0217] While the zero-shot results show relatively modest performance with a correlation of 0.24 and error metrics of 10.1 mmHg bias and 8.9 mmHg precision, these results are nonetheless noteworthy given that the large language model achieved this performance without any domain-specific training examples. The ability to extract meaningful blood pressure estimates from tissue pressure waveforms using only general visual reasoning capabilities demonstrates the potential of the approach.
[0218] In a second test implementation, the LLM was used to estimate values for heart rate (HR) from a biosignal in the form of a PPG sensor signal.
[0219] A test was performed for 20 patients in both zero-shot and few-shot learning scenarios. In the zero-shot case, the LLM was prompted to predict HR without previously showing the model any example PPG signal waveforms. In the few-shot case, the model was presented in the same context window with 15 patient examples, each including an example PPG signal waveform and corresponding2025PF00093
[0220] 20
[0221] HR values, before being prompted to predict HR for 20 unseen patients. The estimated HR values generated by the model in both scenarios were compared to the reference HR measurements.
[0222] Table 2 below shows the correlation coefficient (r), and RMSE error between estimated HR values and reference values.
[0223]
[0224] Table 2
[0225] A root mean square error (RMSE) of 9 and a correlation of 0.83 was achieved using the few-shot approach.
[0226] Heart rate estimation from PPG signals shows marked improvement with few-shot learning compared to the zero-shot approach. The correlation coefficient increased from 0.49 to 0.82, representing a 67% improvement and achieving a strong positive correlation between predicted and actual heart rate values. The RMSE decreased modestly from 10.5 bpm to 9 bpm, indicating a 14% reduction in prediction error. The few-shot heart rate estimation results are particularly encouraging, as the 9 bpm RMSE and 0.82 correlation represent clinically acceptable accuracy for heart rate monitoring applications.
[0227] These results that further demonstrate scalability in accuracy performance without the need for fine-tuning of model parameters. These results demonstrate the effectiveness of the method for a range of different biosignals and biomarkers.
[0228] Table 3 below summarizes the correlation coefficients (r) and RMSE results for both of the test implementations:
[0229]
[0230] 2025PF00093
[0231] 21
[0232]
[0233] Table 3
[0234] The test implementations demonstrate the efficacy of the method for non-invasive prediction of biomarker values, such as non-invasive blood pressure estimation.
[0235] The test implementations did not utilize any fine tuning of the large language model (LLM). However, in further embodiments, fine-tuning techniques may be employed to adapt pre-trained models to specific biomarker prediction tasks through supervised learning on domain-specific datasets comprising biosignal waveforms and corresponding clinical measurements. Retrieval -Augmented Generation (RAG) approaches may additionally or alternatively be used to enhance model performance by incorporating external knowledge bases containing physiological relationships, clinical guidelines, or reference waveform databases during the inference process.
[0236] Applied for predicting blood pressure, this method could enhance the early detection and management of both hypertension and hypotension. By providing more accurate and timely predictions, clinicians can tailor interventions more effectively, potentially improving overall patient outcomes.
[0237] Although the example implementations related to prediction of blood pressure and heart rate, the same methodological approach can be used to predict additional hemodynamic parameters such as systolic blood pressure, diastolic blood pressure, cardiac output, stroke volume, or other cardiovascular biomarkers from the same or different types of biosignal signal inputs.
[0238] Furthermore, in some embodiments, the method may comprise prompting the LLM to estimate a plurality of biomarkers from a single visual representation of a biosignal waveform.
[0239] Although the above example implementation relates to use of GPT-4o with vision, more generally, any Multimodal LLM (Large Language Model with vision capabilities) can be used in the same manner.
[0240] Furthermore, although in the above-described example, the tissue pressure signal was obtained using a tissue pressure sensor comprised by a hemodynamic measurement cuff in accordance with Fig. 3, this is not essential. For example, it is also possible to measure a tissue pressure signal using a standard oscillometric cuff.
[0241] Embodiments of the invention involve the application of a multimodal large language model. Further implementation details relating to the large language model will now be outlined. Such a large language model may be substantially similar to the large language model behind ChatGPT.
[0242] A large language model, in general terms, may be understood as apparatus comprising a deep neural network and characterized by some or all of the following properties. The neural network may be instantiated with no fewer than 10A8 (one hundred million) independently trainable scalar2025PF00093
[0243] 22
[0244] parameters. With regards to the architectural topology, the network may comprise a plurality (e.g. L > 12) of stacked computational layers, each layer comprising, at a minimum: (i) a multi-head self-attention mechanism, and (ii) a position-wise, fully connected feed-forward subnetwork, said components being interconnected by residual connections and layer-normalization operations. Input and output symbols to and from the model have the form of discrete “tokens” obtained through a deterministic tokenizer mapping arbitrary natural -language text to fixed-size integer IDs. Positional information may be injected by means of either learned or analytic positional encodings. With regards to the training paradigm, the network may be trained, by stochastic gradient-based optimization, on a self-supervised predictive objective that minimizes the negative log-likelihood of target tokens conditioned on preceding and / or masked context tokens. The training corpus may contain at least 10A9 (one billion) distinct tokens sampled from one or more human languages. During deployment for inference, the network is operable to (i) compute conditional token-probability distributions and (ii) iteratively sample or select tokens so as to generate text sequences of unbounded length, without task-specific parameter modification. Owing to the scale and diversity of the training corpus, the network may manifest emergent zero-shot or few-shot proficiency across heterogeneous natural-language tasks, as evaluated without additional gradient updates.
[0245] Embodiments of the invention may employ, more particularly, a multimodal large language model. A multimodal large language model may be understood as an apparatus comprising a deep neural network and characterized by some or all of the following properties. The neural network may be instantiated with no fewer than 10A9 (one billion) independently trainable scalar parameters. With regards to the architectural topology, the neural network may comprise: a. a text encoder-decoder stack comprising a plurality (e.g. L t > 12) of layers, each layer comprising a multi -head self-attention mechanism and a position-wise feed-forward subnetwork, these components being interconnected by residual connections and layer-normalization operations; b. a visual encoder comprising a convolutional and / or vision-transformer backbone that converts an input raster image into a sequence of visual tokens or embeddings; c. a cross-modal fusion module implementing at least one of (i) cross-attention, (ii) coattention, or (iii) joint-projection operations so as to integrate information across the text and visual token sequences; d. a unified output decoder which shares parameters with the text encoder-decoder stack and operable to autoregressively generate text tokens conditioned on fused multi-modal representations. For the text modality, input and output symbols to and from the model have the form of discrete text tokens obtained through a deterministic tokenizer, and wherein positional information is injected by learned or analytic positional encodings. For the image modality, each input image may be partitioned into fixed-size patches or transformed by a learned tokenizer so as to yield a sequence of discrete visual tokens or continuous embeddings endowed with positional indices. With regards to the training paradigm, the network may be trained, by stochastic gradient-based optimization, on a composite self-supervised objective including, one or more of: (i) masked or autoregressive language modeling conditioned on visual context, (ii) masked or autoregressive image-token reconstruction conditioned on textual context;2025PF00093
[0246] 23
[0247] and (iii) a contrastive alignment loss encouraging semantically corresponding text-image pairs to occupy proximate regions in a shared embedding space. The training corpus may comprise (i) at least 10A7 aligned text-image pairs and (ii) at least 10A9 standalone text tokens and / or standalone images.
[0248] During deployment for inference, the neural network may be operable to (i) compute conditional probability distributions over text tokens given arbitrary combinations of textual and visual context tokens, and (ii) iteratively sample or select tokens to generate text sequences of unbounded length, without task-specific parameter modification. Owing to the scale and diversity of the training data, the network manifests emergent zero-shot or few-shot proficiency on heterogeneous vision-language tasks when evaluated without additional gradient updates. These may include, without limitation, image captioning, visual question answering, and multimodal dialogue.
[0249] In an example, a pre-trained publicly available model such as GPT-4 or GPT-4o with vision capabilities may be utilized.
[0250] Alternatively, a proprietary multimodal large language model may be trained with the above described characteristics, as described herein.
[0251] Training of the large language model may comprise any one or a combination of: unsupervised learning;
[0252] supervised learning; and
[0253] reinforcement learning.
[0254] In an example, training the large language model according to unsupervised learning comprises:
[0255] receiving an unlabeled data set comprising a plurality of instances, each instance including one or more features;
[0256] initializing the large language model with adjustable parameters; and
[0257] applying an unsupervised learning algorithm to the data set using the large language model, wherein the unsupervised learning algorithm modifies the adjustable parameters of the large language model based on relationships between the features in the instances without referring to any predetermined label or outcome and wherein the modification of the adjustable parameters is performed iteratively until a stopping criterion is met. The objective of the optimization algorithm tuning the parameters is to maximize the probability of predicting the next word in a sequence of inputs, which is iteratively done while incrementally crawling through text material, word for word.
[0258] Upon meeting the stopping criterion, the training according to unsupervised learning may be considered completed. The result at this point is then a trained large language model, wherein the trained large language model is configured to accept new instances of data and generate an output based on the adjusted parameters and the relationships learned during the unsupervised learning process, wherein the output provides an analysis or categorization of the new instance of data and generates outputs based on the most likely next word.2025PF00093
[0259] 24
[0260] Unsupervised learning typically comprises training on an unlabeled data set according to the above description with the aim of predicting the next most probable token in a series of tokens (e.g., a word in a series of words) given to the model. In other words, the trained large language model will predict in each a first word, given the first word a second word, given the first two words the third word, and so on, to generate full sentences, paragraphs and finally texts. By learning the inference of the next words of a sentence or conversation, a transformer architecture of a large language model (e.g., ChatGPT) attains general semantic knowledge.
[0261] For a large language model to be used to perform the steps of the method 10, it is thus advantageous to train the large language model on unlabeled data sets to learn general knowledge and the ability to paraphrase information in different ways, but it may also help to include data sets specific about one or more biosignals and correlated biomarkers which may be derived from analysis of the biosignals. These data sets may be attained, for example, from medical textbooks and / or medical reports used in, for example, training medicine students (e.g., at a university) or doctors (e.g., at a hospital). Additionally or alternatively, these data sets may be attained from historical patient measurement data, e.g. patient medical record databases. Additionally or alternatively, these data sets may be synthetically generated using one or more trained generative adversarial models.
[0262] Suitable unsupervised learning algorithms may be for example clustering, dimensionality reduction, anomaly detection. These algorithms are considered known to the skilled person.
[0263] Suitable data for unsupervised learning is readily available on the world wide web and is known by the skilled person in the art.
[0264] In an example, training the large language model according to supervised learning comprises:
[0265] receiving a labeled dataset comprising a plurality of instances, each instance comprising an input feature and an associated output feature.
[0266] initializing the large language model with adjustable parameters; and
[0267] applying a supervised learning algorithm to the labeled dataset using the large language model, wherein the supervised learning algorithm iteratively modifies the adjustable parameters of the large language model based on a comparison between the large language model prediction given the input feature and the associated output feature until a predetermined stopping criterion is satisfied.
[0268] Upon meeting the stopping criterion, the training according to supervised learning may be considered completed. The result at this point is then a trained large language model, wherein the trained large language model is configured to process new data instances and predict a corresponding output based on the learned relationships between the input features and the associated output features that make up the plurality of instances. The relationships are represented implicitly in the weights of the adjustable parameters.2025PF00093
[0269] 25
[0270] The supervised learning algorithm may be for example any one or a combination of: regression, classification, support vector machines, decision trees, neural networks, or diffusion models. These algorithms are considered state of the art and known to the skilled person.
[0271] While a large language model may be trained immediately following the supervised learning routine, it is typical to use the large language model trained according to unsupervised learning as a starting point to further train according to supervised learning. The step of supervised learning is thus commonly performed after the step of unsupervised learning.
[0272] In an example, the labelled data set comprises example visual representations (e.g. images) of biosignal waveforms paired with associated biomarker values, wherein, in each instance, the input feature comprises a visual representation of a biosignal waveform, in accordance with that employed in the method 10 described above, and the associated output features comprise the value of the at least one biomarker to be estimated in accordance with the method 10.
[0273] For example, the input features comprise: a visual representation of a waveform of a time-series of values of the biosignal for a subject, and the associated output comprises the value of the at least one correlated biomarker for the subject.
[0274] By way of one example, fortraining the large language model to estimate hemodynamic parameters from tissue pressure signals, the input features may comprise visual representations of tissue pressure waveforms captured during cardiac cycles. Each visual representation may comprise an image displaying a plotted waveform showing tissue pressure variations over time, typically spanning one or more heartbeats. The associated output features may comprise one or more corresponding hemodynamic parameter values such as one or more of: mean arterial pressure (MAP) measured in mmHg, systolic blood pressure (SBP) measured in mmHg, diastolic blood pressure (DBP) measured in mmHg, cardiac output (CO) measured in liters per minute, or stroke volume (SV) measured in milliliters. Additional output features may include pulse pressure, arterial compliance measures, or peripheral resistance values.
[0275] Additionally or alternatively, for training the large language model to estimate biomarkers from PPG sensor signals, the input features may comprise visual representations of PPG waveforms showing pulsatile patterns. Each visual representation may comprise an image displaying the PPG signal plotted against time, over one or more cardiac cycles. The associated output features may comprise one or more correlated biological parameters such as one or more of: heart rate measured in beats per minute, blood oxygen saturation (SpO2) percentage, perfusion index value, and pulse wave velocity. Additional output parameters may include heart rate variability metrics, respiratory rate estimates, or pulse transit time measurements.
[0276] The training data for tissue pressure signal analysis may be obtained through clinical studies involving patients undergoing hemodynamic monitoring procedures. Simultaneous recordings may be captured from both the tissue pressure sensor and reference measurement devices such as arterial catheters or validated non-invasive monitors. The tissue pressure waveforms may be recorded during various physiological states including rest, exercise, or medical interventions to capture a diverse range of2025PF00093
[0277] 26
[0278] hemodynamic conditions. Patient populations may include individuals with normal cardiovascular function as well as those with hypertension, hypotension, or other cardiovascular conditions to ensure robust model training across clinically relevant scenarios.
[0279] Clinical data collection protocols may involve standardized measurement procedures performed in hospital settings, outpatient clinics, or research facilities. Each measurement session may include multiple tissue pressure recordings paired with corresponding reference hemodynamic measurements obtained within the same time window. The data may be collected from diverse patient demographics including different age groups, genders, and clinical conditions to enhance model generalizability. Quality control measures may be implemented to ensure signal fidelity and measurement accuracy throughout the data collection process.
[0280] Training data for PPG signal analysis may be acquired through controlled studies using validated PPG sensors and reference monitoring equipment. Simultaneous recordings may be obtained from PPG devices and gold-standard measurement systems such as pulse oximeters, electrocardiograms, or arterial blood gas analyzers. The PPG signals may be captured from various anatomical locations including fingertips, wrists, or earlobes to account for site-specific variations in signal characteristics. Data collection may encompass different physiological conditions such as varying activity levels, breathing patterns, or ambient lighting conditions.
[0281] Large-scale datasets may be compiled from existing medical databases, clinical trials, or research repositories containing validated PPG measurements and corresponding biomarker values. Synthetic data generation techniques may be employed using physiological models or generative adversarial networks to augment the training dataset with additional examples. Cross-validation procedures may be implemented to ensure data quality and consistency across different measurement devices and clinical settings. The training datasets may be continuously expanded through ongoing clinical collaborations and real-world deployment feedback to improve model performance and robustness.
[0282] A large language model trained according to supervised learning on data as described above would thus naturally adhere to the training data examples it has been trained on and is thus fully enabled to perform method 10. However, to enhance the performance of the model in the case of unexpected and / or deviating input to the input on which the large language model was trained, the large language model may further be trained according to reinforcement learning.
[0283] Reinforcement learning is a type of learning in which a trained large language model is graded in use. For example, if the large language model performs as expected, it may be rewarded. If on the other hand the large language model performs against expectation and / or outside of its supposed bounds, it is penalized. This reward and penalty typically are performed via a pre-defined loss function. However, it may also be performed manually, in a process wherein testers (e.g., operator used for testing method 10 or very experienced subjects) may provide a score after each use. In this manner, the large2025PF00093
[0284] 27
[0285] language model algorithm will adjust the weights to better conform with the testers wishes and / or expectations.
[0286] In other words, feedback on the predicted biomarker values may be collected. Based on this feedback, the large language model may be fine-tuned to improve future performance, handling, and diagnostic value. The feedback may be collected in a testing phase or from users during deployment. The feedback may further be collected in free-text format (e.g., a tester or user may freely give feedback) or in a structured manner (e.g., a ranking system per instruction feature).
[0287] Based on the given feedback, new supervised training cases may be constructed to further improve the performance of method 10. Based on the given feedback, also additional algorithms may be put in place to restrict or influence the behavior of the trained large language model, thereby avoiding a further learning phase for the large language model. For example, a general adversarial network approach may be implemented as a second, independent large language model (e.g., a neural network, a large language model, etc.) to supervise the main model. This supervision large language model could be trained on validated clinical data and established physiological relationships to check the decisions of the main large language model. For example, the supervision large language model may in this manner approve or reject decisions of the main large language model. Thereby providing an extra safety net against undesired and / or unwanted outputs.
[0288] A typical reinforcement learning workflow comprises a first step of initializing the large language model with adjustable parameters.
[0289] In a second step, a reinforcement learning algorithm is applied, wherein the large language model interacts with an environment (e.g., the deployment or testing of the large language model), performs actions based on its current state and parameters (i.e., gives outputs, e.g., instructions, given inputs), and receives rewards or penalties based on the performed actions. For example, the large language model may be used to predict values of the at least one biomarker and the rewards or penalties based on the predictions may be based on a user defined performance or based on a pre-defined loss function.
[0290] In a third step, the model parameters are adjusted iteratively based on the received rewards or penalties until a predetermined stopping criterion is met.
[0291] Upon meeting the stopping criterion, the final large language model algorithm may be configured to make decisions in new states based on the adjusted parameters and learned rewards.
[0292] Suitable reinforcement learning algorithms may be any one or a combination of: Q-leaming, Deep Q Network, Policy Gradient, or Actor-Critic methods or proximal policy optimization.
[0293] For all types of learning, unsupervised, supervised and reinforcement, the adjustable parameters may be determined based on minimizing or maximizing an objective function derived from the received data set.
[0294] Although examples have been presented above relating to estimation of biomarkers from tissue pressure signals and from PPG signals, the trained large language model may be configured for2025PF00093
[0295] 28
[0296] application across a diverse range of biosignal monitoring scenarios beyond tissue pressure and PPG measurements.
[0297] By way of one non-limiting example, electrocardiogram (ECG) signals represent a significant application domain where the method may be employed to estimate cardiac biomarkers from visual waveform representations. The ECG waveform images may display characteristic P waves, QRS complexes, and T waves from which parameters such as heart rate, rhythm abnormalities, QT interval duration, and ST segment deviations may be inferred. Additional cardiac biomarkers including PR interval measurements, axis deviations, and arrhythmia classifications may be extracted through analysis of ECG waveform visualizations.
[0298] Standard oscillometric air cuffs may serve as another biosignal source for the trained model, wherein pressure oscillations detected during cuff deflation may be analyzed for blood pressure estimation. The oscillometric waveform patterns may be visualized as amplitude variations plotted against time, enabling extraction of systolic, diastolic, and mean arterial pressure values. The method may be particularly advantageous for oscillometric measurements in challenging patient populations where traditional algorithms may struggle, such as patients with arrhythmias, arterial stiffness, or weak pulse amplitudes.
[0299] Respiratory monitoring applications may utilize various biosignal types including respiratory effort belts, nasal airflow sensors, or thoracic impedance measurements. Visual representations of respiratory waveforms may enable estimation of respiratory rate, tidal volume, inspiratory-to-expiratory ratios, and breathing pattern irregularities. Capnography signals displaying carbon dioxide concentration over time may be analyzed to determine end-tidal CO2 levels, respiratory dead space, and ventilation efficiency parameters. The method may prove particularly valuable in sleep study applications where multiple respiratory parameters require simultaneous assessment from complex waveform patterns.
[0300] Neurological monitoring represents another application domain where the trained model may analyze electroencephalogram (EEG) signals for brain activity assessment. EEG waveform visualizations may enable detection of seizure activity, sleep stage classification, or cognitive load estimation through pattern recognition of characteristic frequency bands and amplitude variations.
[0301] Electromyography (EMG) signals from muscle activity monitoring may be processed to estimate muscle activation levels, fatigue indices, or movement intention parameters. Additional neurophysiological applications may include analysis of evoked potential waveforms for sensory pathway assessment or nerve conduction study results for peripheral nerve function evaluation.
[0302] Embodiments of the invention described above employ a processing device. The processing device may in general comprise a single processor or a plurality of processors. It may be located in a single containing device, structure or unit, or it may be distributed between a plurality of different devices, structures or units. Reference therefore to the processing device being adapted or configured to perform a particular step or task may correspond to that step or task being performed by any2025PF00093
[0303] 29
[0304] one or more of a plurality of processing components, either alone or in combination. The skilled person will understand how such a distributed processing device can be implemented. The processing device includes a communication module or input / output for receiving data and outputting data to further components.
[0305] The one or more processors of the processing device can be implemented in numerous ways, with software and / or hardware, to perform the various functions required. A processor typically employs one or more microprocessors that may be programmed using software (e.g., microcode) to perform the required functions. The processor may be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.
[0306] Examples of circuitry that may be employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).
[0307] In various implementations, the processor may be associated with one or more storage media such as volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. Various storage media may be fixed within a processor or controller or may be transportable, such that the one or more programs stored thereon can be loaded into a processor.
[0308] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0309] A single processor or other unit may fulfill the functions of several items recited in the claims.
[0310] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0311] A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0312] If the term "adapted to" is used in the claims or description, it is noted the term "adapted to" is intended to be equivalent to the term "configured to".
[0313] Any reference signs in the claims should not be construed as limiting the scope.
Claims
2025PF0009330CLAIMS:
1. A computer-implemented method (10) for estimating a value of at least one biomarker for a subject based on inference from a biosignal for the subject generated by a biosignal sensor, the method comprising:using a trained general purpose vision large language model to generate the estimated value of the at least one biomarker,wherein the general purpose vision large language model is operable to generate an output responsive to an input prompt comprising image data; andwherein the method comprises: obtaining (12) a visual representation of a waveform of a time-series of values of the biosignal, and prompting (14) the trained general purpose vision large language model with an input prompt comprising the visual representation of the waveform.
2. The method of claim 1,wherein the at least one biomarker comprises a hemodynamic parameter, for example an arterial blood pressure measure, andwherein the biosignal comprises a pressure signal indicative of a pressure between a body part comprising an artery and a compressing surface applied to the body part over the artery.
3. The method of claim 2, wherein the method comprises obtaining the pressure signal from a pressure sensor comprised by a hemodynamic measurement device comprising a body-mountable cuff, wherein the cuff comprises the compressing surface.
4. The method of claim 3,wherein the cuff comprises an actuator for applying a controllable pressure to the body part and a shell portion between the actuator and the body part, andwherein the pressure sensor is disposed between the shell portion and the body part.
5. The method of any of claims 2-4, wherein the at least one biomarker comprises a measure of mean arterial pressure.
6. The method of claim 1, wherein the biosignal comprises a PPG signal, and wherein the biomarker comprises at least one hemodynamic, cardiovascular or physiological parameter, for example heart rate.2025PF00093317. The method of any preceding claim, wherein the input prompt further comprises a request for an estimated value of the at least one biomarker.
8. The method of claim 7,wherein the trained general purpose vision large language model is a multi-modal large language model operable to generate an output responsive to an input prompt comprising one or both of image data and text; andwherein the input prompt comprises a natural language text representation of the request.
9. The method of any preceding claim, wherein the input prompt comprises one or more training examples, each training example comprising an example visual representation of a waveform of the biosignal and a corresponding example value of the at least one biomarker.
10. The method of claim 9, wherein the one or more training examples comprise a plurality of training examples, for example three or more training examples, for example fifteen or more training examples.
11. The method of any preceding claim, wherein the method comprises receiving a time series of values of the biosignal and generating the visual representation of the waveform of the time series of values of the biosignal.
12. A computer program product comprising computer program code configured, when run on a processor, to cause the processor to perform a method in accordance with any of claims 1-11.
13. A processing device comprising one or more processors configured to perform a method in accordance with any of claims 1-11.
14. A system comprising:a biosignal sensor operable to measure a time series of values of a biosignal for a subject; andthe processing device in accordance with claim 13, arranged to receive the time series of biosignal values from the biosignal sensor.