Method for predicting a future health status of a patient
The method enhances cardiac arrhythmia detection by denoising and classifying heart rhythm signals from consumer devices, enabling reliable prediction and monitoring of cardiac health conditions.
Patent Information
- Application Number
- PCT/EP2025/067144
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
Existing consumer-grade devices struggle to reliably detect cardiac arrhythmias due to noisy and infrequent heart rhythm signals, making it difficult for practitioners to differentiate between rhythmic sequences indicating cardiac disorders and recording artifacts, and preventing effective monitoring and prediction of patient health status.
A method for predicting future health status using consumer devices involves signal denoising, feature extraction, and classification into categories like 'normal sinus rhythm' and 'predetermined health status' through auto-regression models, leveraging noise type classification and multi-scale representations to optimize data sequences.
Enables robust classification and prediction of cardiac health conditions, allowing practitioners to interpret signals and establish a database for automated diagnosis and monitoring, even with consumer-grade devices.
Smart Images

Figure IMGF000018_0001 
Figure IMGF000018_0002 
Figure 00000030_0000
Description
Description Title of the invention: Method for predicting a patient's future health status
[0001] The invention relates to the field of screening for a patient's health condition that may involve a disorder of their heart rhythm, and in particular cardiovascular diseases.
[0002] Screening, monitoring, and preventing diseases that may involve a patient's heart rhythm requires, on the one hand, medical-grade devices, such as Holter monitors, capable of acquiring an electrocardiogram (ECG) over an extended period, and on the other hand, the intervention of a practitioner capable of analyzing the ECG thus acquired in order to diagnose the patient. This screening presents several challenges due to major difficulties, namely the availability of the necessary equipment, the patient's acceptability of the examination involving the placement of electrodes on the skin given its duration, and the availability of qualified practitioners, given that the analysis time is proportional to that of the examination.
[0003] In this context, various types of consumer-grade devices have been designed that can be worn easily by patients for extended periods. These devices are equipped with sensors capable of acquiring signals related to the patient's heart rhythm. Examples include smart bracelets or watches with optical sensors that can acquire a photoplethysmogram (PPG) signal, reflecting modulations due to the flow and reflux of blood in the venous system under the influence of the heart muscle. This allows for the deduction of the patient's heart rate and oxygen saturation. Other examples include smart watches capable of acquiring a single-lead ECG of the patient during short, one-off use, for example, 30 seconds, and belts equipped with a pair of electrodes capable of acquiring one over several hours.We can also mention phone or computer cameras capable of evaluating variations in the patient's skin color due to the flow and reflux of blood in the subcutaneous venous system under the effect of the heart muscle, in order to establish a remote photoplethysmogram, or even rPPG.
[0004] While these different types of signals are theoretically interpretable to assess the patient's health status, in practice there are several difficulties.
[0005] On the one hand, the signals collected by these various consumer devices are too infrequent, particularly with regard to single-lead point-by-point ECGs, and too noisy to reliably identify the patient's health status and detect a current or future cardiac arrhythmia. In particular, the noise from signal acquisition or measurement makes it difficult to distinguish between a portion of the signal indicating a cardiac arrhythmia and this measurement noise, which can be quite similar.
[0006] On the other hand, it is difficult for a practitioner to differentiate a random sequence in the signal from a rhythmic sequence indicating a cardiac disorder, such as atrial fibrillation. The biomarkers in the temporal rhythm signal indicating these cardiac disorders are indeed difficult to detect because they are masked by other signals. Rhythms that appear normal, particularly due to recording artifacts. This difficulty therefore increases the complexity of detecting cardiac disorders, especially at an early stage, and thus prevents reliable patient monitoring to prevent future events.
[0007] Furthermore, these various difficulties make it impossible to establish a reliable and comprehensive database of temporal rhythm signal sequences labeled with corresponding patient health statuses. Therefore, it is not currently possible to automate the diagnosis, monitoring, and prediction of patient health status through rhythm data collection in this way.
[0008] There is therefore a need for a method that provides practitioners with indications of a patient's possible future health status based on derived data, including heart rate resolution, from temporal signals, particularly those from consumer devices. This would allow practitioners to interpret these signals, diagnose, predict, and monitor the patient's condition, and propose appropriate treatments when necessary. There is also a need to establish a database of heart rhythm biomarkers labeled with a patient's possible or confirmed health status, designed to automate the prediction and monitoring of a patient's health.
[0009] The present invention falls within this context and aims to address one or both of these needs.
[0010] To this end, the invention relates to a method for predicting a patient's future health status from at least one signal containing information on the patient's cardiac activity, the method being implemented by computer and comprising the following steps: a. extraction of a data sequence relating to the heart rhythm from said signal; b. extraction of at least: i. a feature from a lexicographical representation of said data sequence; and / or ii. a feature from a representation of the density of the distribution of data from said data sequence around a centered value from said data sequence; c. classification of the signal into classes "normal sinus rhythm", "noise" and at least one class indicating a predetermined health status from said extracted feature(s); d.prediction, according to a given auto-regression model and from said signal class, of a probability of a future health state of said patient.
[0011] The invention aims to obtain, from an initial signal, a data sequence relating to the patient's heart rhythm. Heart rhythm is indeed an excellent indicator of the pathophysiological state of the heart and the body. In order to extract this data sequence, the signal can be denoised by exploiting... Characteristics relating to the noise contained in the signal, and in particular the classification of the noise type, i.e., whether a part of the signal contains impulsive noise or stationary noise. It is then possible to optimize the extraction of the data sequence from the signal by taking into account the unique characteristics of the heartbeats.
[0012] Once the data sequence has been extracted, it is then possible to extract various indicative characteristics, such as those relating to the data distribution and their correlations according to various transformations and kernels, particularly in lexicographic and data distribution contexts, or even in other contexts such as a multi-scale context, and especially according to various spectral categories, for example, by comparison with the spectrum of a random matrix. It has indeed been observed that certain health conditions alter the multifractality of heart rhythm, or that certain health conditions are characterized or preceded by chaotic or even random rhythmic sequences, or conversely, by frequent appearances or disappearances of patterns in the heart rhythm.The extracted characteristics thus make it possible to establish rhythmic signatures or imprints of health states, inscribed in the rhythmic heartbeat, and to differentiate these rhythmic signatures from each other as well as from measurement noise.
[0013] This allows the initial signal to be classified into different categories, indicating either a normal sinus rhythm, a signal with excessive noise, or a predetermined health condition. This classification can be performed robustly, even for signals obtained from consumer devices. A practitioner can then interpret the signal using this classification to make a diagnosis. Furthermore, it is possible to generate a database of rhythmic signatures, each labeled with its corresponding class.
[0014] Finally, the invention aims to project this current health state to predict the occurrence of a future health state, thereby enabling the practitioner to implement patient monitoring and, if necessary, a prevention plan. This probabilistic prediction could, for example, be implemented using a Bayesian autoregression model, particularly one of order greater than 1, which may or may not incorporate the patient's medical history to further strengthen the prediction.
[0015] In the present invention, and without limitation, the term "future health condition of a patient" means a future health condition involving the occurrence of cardiac rhythm disorders. This health condition may include an altered sinus rhythm, an arrhythmia such as a conduction abnormality, atrial or supraventricular arrhythmia, ventricular arrhythmia, cardiovascular disease, transient ischemic attack or stroke, heart failure, coronary ischemia, myocardial infarction, cardiac arrest, sudden death, diabetes, hypertension, stress, fatigue, depression, sleep apnea, cancer recurrence, or chronic obstructive pulmonary disease.
[0016] In the present invention, and without limitation, the term "signal containing information on the cardiac activity of said patient" means a signal such as a photoplethysmogram (PPG), remote photoplethysmogram (rPPG), electrocardiogram (ECG), electroencephalogram (EEG), electrogram (EGM), or any other signal whose characteristics, particularly temporal and / or frequency characteristics, directly or indirectly reflect, over any time scale, variations in the patient's heart rhythm. The signal may be acquired by one or more sensors that can be integrated into or connected to one or more devices such as a Holter monitor, wristband, smartwatch, belt, phone, tablet, camera, or computer.
[0017] If desired, the method according to the invention may include a preliminary step of acquiring said signal containing information on cardiac activity of said patient.
[0018] In the present invention, and without limitation, the term "heart rhythm data sequence" means a time sequence whose values or variations indicate a heart rhythm and / or heart rate variability. The data sequence may, for example, include a time sequence of the "R" peaks of the ECG, a time sequence of the intervals separating heartbeats such as the "RR" interval sequence, a tachogram, or any other representation of heart rate variability.
[0019] In the present invention, and without limitation, the term "predetermined health status classes" includes, but is not limited to, the following: abnormal sinus rhythm indicating diabetes, abnormal sinus rhythm indicating hypertension, abnormal sinus rhythm indicating heart failure, abnormal sinus rhythm indicating stress, abnormal sinus rhythm indicating fatigue, abnormal sinus rhythm indicating depression, abnormal sinus rhythm indicating cancer recurrence, abnormal sinus rhythm indicating chronic obstructive pulmonary disease, abnormal sinus rhythm indicating sleep apnea, rhythm indicating atrial fibrillation, rhythm indicating ventricular tachycardia, rhythm indicating frequent extrasystoles, rhythm indicating sinus tachycardia, rhythm indicating sinus bradycardia, rhythm indicating atrial flutter,rhythm indicating a conduction abnormality or any other rhythm class indicating a cardiac arrhythmia.
[0020] In the present invention, and without limitation, the term "computing unit" means one or more electronic and / or computer devices designed to implement, in an analog and / or digital manner, all or part of the steps of the process according to the invention. The computing unit may implement the steps of the process according to the invention as the signal is acquired or, alternatively, after a complete signal acquisition sequence. The computing unit may be equipped with one or more processors, arranged to execute instructions from one or more computer programs in order to implement the steps of the process according to the invention. The steps of the process according to the invention may be implemented centrally by a single computing unit or distributedly by several computing units. It may also be provided that all or part of the steps of the process according to the invention are implemented by a computing unit embedded in a device comprising the sensor(s) used to acquire the signal and / or by a computing unit located remotely from said device with which it can exchange data via wireless communication.
[0021] In one embodiment of the invention, the method includes a preliminary step of classifying the type of noise contained in the signal. The heart rhythm data sequence is then extracted from the signal using at least one filtering algorithm parameterized with the result of this noise type classification. These steps enable the signal to be denoised by exploiting characteristics of the noise contained in the signal, and in particular the noise type classification, namely whether a portion of the signal contains impulsive or steady-state noise. It is then possible to optimize the extraction of the data sequence from the signal by taking into account the unique characteristics of the heartbeats.
[0022] Advantageously, the noise type classification step includes a step for determining the temporal distribution of the signal's energy in at least one spectral band. The noise type contained in said signal is then classified into a stationary noise class and an impulsive noise class based on this distribution. It may be foreseen that the noise type will be classified into a plurality of noise classes, including a stationary or quasi-stationary noise class (such as white noise or pink noise), an impulsive noise class, and a non-stationary noise class (such as Brownian noise), based on this distribution.
[0023] In one embodiment of the invention, the step of determining the temporal distribution of the signal energy in at least one spectral band comprises determining a value representative of a proportion, or distribution, of the components of said signal of frequencies included in one or more given frequency bands with respect to the components of said signal of frequencies included in one or more other given frequency bands, the type of noise contained in said signal being classified according to at least said value.
[0024] For example, said representative value may be a ratio of the spectral energy of the components of said signal belonging to so-called low frequency ranges, for example frequency less than 0.5 Hz, and / or high frequency ranges, for example frequency greater than 4.5 Hz, to the spectral energy of the components of said signal belonging to a so-called normal frequency range, for example frequencies between 0.5 Hz and 4.5 Hz.
[0025] In another alternative or cumulative embodiment of the invention, the step of determining the temporal distribution of the signal energy in at least one spectral band includes determining a representative value of the energy diffusivity. between components of said signal of frequencies included in one or more given frequency bands and components of said signal of frequencies included in one or more other given frequency bands, the type of noise contained in said signal being classified according to at least said value.
[0026] For example, said representative value may be a ratio of the LO standard of the spectral energy of the components of said signal belonging to so-called low frequency ranges, for example frequency less than 0.5 Hz, and / or high frequency ranges, for example frequency greater than 4.5 Hz, to the LO standard of the spectral energy of the components of said signal belonging to a so-called normal frequency range, for example frequency between 0.5 Hz and 4.5 Hz.
[0027] In these embodiments, the components of said signal belonging to a given frequency range can be estimated by using indifferently a Fourier transform, a wavelet transform or any transform allowing a signal to be decomposed into a basis of kernels allowing this signal to be analyzed in the frequency domain or in the time-frequency domain.
[0028] In yet another alternative or cumulative embodiment of the invention, the step of determining the temporal distribution of the signal energy in at least one spectral band comprises the extraction of dispersion coefficients of the components of said signal of frequencies included in one or more given frequency bands and of the components of said signal of frequencies included in one or more other given frequency bands, the type of noise contained in said signal being classified according to said dispersion coefficients.
[0029] For example, these dispersion coefficients can be extracted using a cascade of wavelet decompositions of the signal, the characteristic being determined from the dispersion coefficients obtained at the end of each decomposition in the cascade. This cascade of decompositions, also called a Wavelet Scattering Transform (WST), performed on the signal and its magnitude, allows the dispersion coefficients to be obtained at the output of each decomposition, using a low-pass filter. The set of dispersion coefficients, also called a scatogram, forms a multi-scale representation in the time-frequency domain of the signal, allowing the signal structure to be analyzed at different time and frequency resolutions.
[0030] These different embodiments allow for obtaining different representative values of the signal structure, the distribution and / or diffusivity of its spectral energy, and / or its dispersion in the frequency domain and / or the time-frequency domain and / or the multi-scale time-frequency domain. Following the step of determining the temporal distribution of the signal's energy in at least one spectral band, these values can then be compared, alone or in combination, to one or more predetermined threshold values to classify the type of noise contained in sequences of said signal according to different time scales into said stationary noise class and said noise class impulsive.
[0031] It will be possible to foresee other variants of extraction of a value representative of a proportion, or a distribution, of the components of said signal of frequencies included in one or more given frequency bands with respect to the components of said signal of frequencies included in one or more other given frequency bands, and in particular variants based on Fourier transforms, wavelet transforms or even variational mode decompositions.
[0032] In yet another alternative or cumulative embodiment of the invention, said values representing the signal structure, the distribution and / or diffusivity of its spectral energy, and / or its dispersion in the frequency and / or time-frequency and / or time-frequency multiscale domains may be provided as input to a classification algorithm trained to classify the noise contained in a signal based on characteristics representing the distribution and / or diffusivity of spectral energy in frequency and / or time-frequency and / or time-frequency multiscale components of the signal. This algorithm may, for example, be a kernel classifier or a convolutional neural network, in particular a k-means or DBS-CAN classifier.
[0033] Advantageously, at the end of the noise type classification step, a predetermined value of a parameter, or a plurality of predetermined values forming a parameter vector, of the filtering algorithm used in the data sequence extraction step can be selected according to the noise class.
[0034] In one embodiment of the invention, the filtering algorithm parameterized using the result of said noise type classification includes a minimization of the differences between said extracted data sequence and said signal under a sparsity constraint of a maximum temporal gradient of said extracted data sequence.
[0035] This optimization ensures that the signal reconstructed by the parameterized filtering algorithm, which then forms the data sequence, is meaningfully represented by a reduced number of components, each exhibiting significant temporal variations. In other words, this reconstructed signal closely approximates a discrete signal. Thus, in one example, this reconstructed signal can represent only the R peaks of an ECG, or in another example, the systolic peaks of a PPG, thereby providing a discrete indication of the heart rhythm.
[0036] Advantageously, the filtering algorithm parameterized using the result of said noise type classification includes an edge-preserving filter parameterized using the result of said noise type classification. This type of filter ensures that strong signal variations are preserved during denoising, further improving the performance of minimizing discrepancies between said extracted data sequence and said signal under a sparsity constraint of a maximum temporal gradient.
[0037] Alternatively or cumulatively, the parameterized filtering algorithm may include a singular value decomposition and / or an independent component analysis.
[0038] In one embodiment of the invention, the method comprises, in addition to the step(s) of extracting said feature(s) from lexicographic and / or distribution representations, or even multi-scale representations, of said data sequence, a step of extracting at least one feature from a multi-scale representation of said data sequence. Said signal is thus classified into classes such as "normal sinus rhythm," "noise," and at least one class indicating a predetermined health status based on all the extracted features.
[0039] Where appropriate, it may be provided that the step of extracting at least one feature from a multi-scale representation of said data sequence includes at least one cascade of wavelet decompositions of said data sequence, said feature being determined from the dispersion coefficients obtained at the end of each decomposition of said cascade.
[0040] Alternatively or cumulatively, the step of extracting at least one feature from a multi-scale representation of said data sequence includes the extraction of a feature from the multifractal spectrum of said data sequence and / or from the linear and / or non-linear two-point correlation across the scales of said multi-scale representation.
[0041] In one example implementation, the step of extracting a feature from the multifractal spectrum of said data sequence includes the estimation of one or more multifractal cumulants of said data sequence. For example, said multifractal cumulants could be the cumulants CO, Cl, C2, and C3. Alternatively, the step of extracting a feature from the multifractal spectrum of said data sequence includes the estimation of the Hurst exponent, the fractal dimension of the data sequence support, the degree of fractionality of the data sequence, and / or its degree of multifractality.
[0042] Alternatively or cumulatively, the step of extracting at least one feature from a multiscale representation of said data sequence includes providing the multiscale representation, for example a scalogram or a local Fourier transform, including a Gabor transform, of said data sequence, to a machine learning algorithm trained to classify a data sequence from a multiscale representation of that data sequence.
[0043] These different characteristics quantify how moments in the data sequence, such as its mean and variance, vary with changes in scale and thus allow us to quantify how heart rate fluctuations are distributed and reproduced at different scales in the data sequence.
[0044] In one embodiment of the invention, the step of extracting a feature from a lexicographical representation of said data sequence includes a classification of each of the data of said sequence among a plurality of predetermined classes, a transformation of said data sequence from data classes to obtain a sequence of classes, a decomposition of said sequence of classes into a series of sub-sequences of classes, a determination of a lexicographic distribution of said series of sub-sequences of classes, said characteristic being determined from said lexicographic distribution.
[0045] Advantageously, when the data sequence represents a sequence of RR intervals, the classification step for each data point in said sequence involves classifying each data point into a plurality of interval length classes, including short interval, normal interval, and long interval classes. This classification could, for example, be implemented by thresholding each data point according to one or more given duration ranges. The data sequence, representing a sequence of RR intervals, can thus be transformed, or sampled, into a sequence of intervals, each of which can be short, normal, or long, this sequence forming said class sequence.
[0046] Advantageously, the decomposition of the said sequence of classes into a series of subsequences of classes may include the decomposition into subsequences of three or more successive classes. Each subsequence may, for example, comprise four or more successive intervals, each of which may be short, normal, or long. These subsequences form, in other words, patterns or words that may constitute part of a rhythmic signature of a given cardiac state.
[0047] Advantageously, the determination of a lexicographical distribution of the sequence of class subsequences may involve determining a distribution of the said sequence of class subsequences according to a dictionary of class subsequences. The said sequence of subsequences may, in particular, be transformed into a lexicographical order, in which the dictionary subsequences are ordered according to their frequency of occurrence in the said sequence of subsequences. The distribution formed by this lexicographical order thus indicates, for each dictionary subsequence, the number of occurrences of said subsequence in the sequence of subsequences, and thus provides an indication of the frequency or probability of occurrence of that subsequence in the said sequence.This lexicographic distribution is thus comparable to a law of random occurrences, such as Zipf's law, so that the distribution of the probabilities of occurrence of the subsequences of the dictionary in said sequence reflects a rhythmic signature.
[0048] Advantageously, said characteristic may be the lexicographic distribution. Alternatively, said characteristic may include one or more values relating to the lexicographic distribution, such as entropies, and / or dominant modes, namely sub-sequences of said dictionary most represented in said sequence, and / or moments of the sub-sequences in said sequence, and / or the asymptote of said distribution.
[0049] We can foresee other variants of lexicographical representations of the data sequence from which features can be extracted, such as a lexicographical transition graph where each node represents a data item, a data class or even a subsequence of classes and the edges represent the transitions between these data items, data classes or subsequences of classes in the data sequence, or such as a lexicographical recurrence matrix.
[0050] In one embodiment of the invention, the step of extracting a feature from a representation of the distribution density of the data of said data sequence around a centered value from said data sequence includes the estimation of a counting function of the data of the data sequence, said feature being determined from said counting function.
[0051] In this embodiment, the counting function of the data sequence is analyzed; that is, the evolution of the distribution of values within the data from a temporal window of the sequence around a centered value, such as an average value evaluated over this temporal window, with the size and / or position of the window. Comparing this counting function, through the value relative to this function, with a given statistical distribution, allows us to identify the distribution of values relative to the data in the sequence and to identify temporal trends in these data, and in particular specific signatures of variations, which can reflect a particular state of health. It is thus possible to characterize subclasses of processes that would be classified in the same way by a multi-scale analysis and / or by a lexicographic analysis.
[0052] For example, if the data sequence is a sequence of RR intervals, the mean interval value is determined over only one time window of the sequence. Each interval in this portion of the sequence is then placed on a unit circle, with one pole corresponding to the mean value. The circle thus represents a typical heart rate period within the time window. An interval value longer than the mean is placed clockwise past the pole, and an interval shorter than the mean is placed clockwise before the pole. The density of the distribution of these distributed intervals can then be analyzed—that is, the clustering of intervals around given points and the distances separating the interval values.By repeating this analysis by temporally moving the window along the data sequence and / or by varying the size of the window, and therefore the number of intervals considered, it is possible to measure the deviation of the evolution of the density of said distribution from a Poisson point process or a Tracy-Widom distribution, and in particular to analyze the variations and extreme values of said distribution with the number of intervals and to verify the law of evolution of this density with respect to a given statistical law, for example to the distribution of the eigenvalues of a random matrix of a specific class.
[0053] Advantageously, this characteristic may be the counting function of the data sequence. Alternatively, this characteristic may include one or more values relating to this counting function, and in particular extreme values and / or a diffusion law and / or measures of deviation of this function from a given statistical law.
[0054] We can foresee other variants of representation of the distribution density of the data of said data sequence around a centered value from said data sequence, from which characteristics can be extracted, such as a histogram or a density estimated by kernel.
[0055] In one embodiment of the invention, the extraction step comprises the extraction of at least one feature from a lexicographical representation of said data sequence and at least one feature from a representation of the distribution density of said data sequence around a centered value from said data sequence, said signal being classified among classes "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health from all the extracted features.
[0056] In one embodiment of the invention, the method comprises, in addition to the step or steps of extracting said feature(s) from the lexicographic and / or distribution, or even multi-scale, representations of said data sequence, the following steps: a. extraction of a statistical feature from said data sequence; b. extraction of a feature from the frequency spectrum of said data sequence; said signal being classified among said classes "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health from all the extracted features.
[0057] These additional features can be added to the other extracted features to further strengthen the signal classification and the prediction of the probability of occurrence of a future health condition.
[0058] Advantageously, the step of extracting a statistical characteristic from said data sequence includes estimating values relating to the statistical distribution of said data. These values may be the dominant modes, moments, and extreme values of said data.
[0059] Advantageously, the step of extracting a feature from the frequency spectrum of said data sequence includes calculating an energy spectral density of said data sequence and estimating values relative to one or more peak values of said energy spectral density. For example, said peaks could correspond to the physiological Fourier modes, namely low-frequency peaks corresponding to the sympathetic / adrenergic system and a high-frequency peak corresponding to the parasympathetic / vagal system. The values relating to these peaks may be the absolute values of the spectral energies corresponding to these peaks or ratios of these values to each other.
[0060] In one embodiment of the invention, the step of classifying the signal among the classes "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health, from said at least one extracted feature is implemented by means of at least one machine learning algorithm trained to classify a data sequence among said classes from features from lexicographic and / or distribution representations, or even multi-scale representations, of a data sequence.
[0061] This machine learning algorithm could, for example, be a support vector machine (SVM) trained on a training dataset to establish a decision function that assigns a new dataset to a class from among a set of classes based on the evaluation of a margin function between the dataset and a decision boundary, called the hyperplane. In other words, the dataset is thus placed, either directly or after processing by a kernel function, by the algorithm into a space divided into several subspaces by the hyperplane, each subspace corresponding to one of the classes in the set of classes. The position of the dataset in this space allows it to be assigned one of the classes.The algorithm's training is designed to identify the optimal hyperplane, that is, the one maximizing the margin function between the hyperplane and the training datasets closest to this hyperplane, called the support vector.
[0062] In other implementation examples, other types of machine learning algorithms capable of classifying a data sequence among said classes based on features from lexicographical and / or distributional representations, or even multi-scale representations, of a data sequence, which can be trained in a supervised, unsupervised, reinforcement manner, and in particular of the types deep neural networks, convolutional neural networks, recurrent neural networks, decision trees, data grouping and any combination of these elements.
[0063] In one embodiment of the invention, the prediction step, according to a given autoregression model and from said signal class, of a future health state of said patient, comprises a substep of generating, from said signal class, a potential future data sequence of said data sequence according to said autoregression model; and a substep of predicting, from said potential future data sequence and said signal class, a probability of a future health state of said patient.
[0064] This generation substep aims to synthesize a sequence of potential data by progressively propagating the data sequence obtained from said signal, and to predict, from this data sequence, the future health state of said patient.
[0065] Advantageously, said potential data sequence can be generated by means of a machine learning algorithm implementing a semantic auto-regression model applied to said data sequence, in particular to a sequence of classes obtained by transforming said data sequence from data classes, or even to a series of sub-sequences of classes obtained by decomposing said sequence of classes.
[0066] In one embodiment of the invention, said potential data sequence may be generated by means of a machine learning algorithm implementing an autoregression model trained to predict a data sequence from an input data sequence and which minimizes a cost function determined from a probability density of the data in the predicted sequence and a probability density of the data in the input sequence.
[0067] This model could, in particular, be a Gaussian process autoregression model. For example, this model could be trained unsupervised to infer or fit a Gaussian kernel of said model that minimizes a distance between a spectrum of a predicted data sequence and, therefore, a spectrum of said kernel and a spectrum of the training data provided as input.
[0068] In another embodiment of the invention, said potential data sequence may be generated by means of a diffusion model, in particular a probabilistic denoising diffusion model, for example implemented by an unsupervised trained neural network.
[0069] In one example implementation, the prediction substep can be carried out by a Bayesian-type classifier trained to estimate, from a data sequence and an initial health state, the probability of occurrence of each health state among a plurality of given health state classes. The class with the highest probability thus constitutes the patient's future health state.
[0070] Advantageously, the process includes a step of providing the patient's medical history, and the probability of the patient's future health condition is predicted according to the given autoregression model and based on the signal class and the patient's medical history. Information regarding the patient's medical history can thus be introduced into the Bayesian classifier to enhance the reliability of its predictions.
[0071] For example, the Bayesian classifier may include the same classifier used during the implementation of the signal classification step, with a layer of a Bayesian Vector Autoregressive (BVAR) model being added to said classifier to predict said probability of a future health condition from the patient's medical history and a class obtained by said classifier to which is provided as input said potential future data sequence.
[0072] The invention also relates to a system for predicting a patient's future health state, comprising at least one sensor capable of acquiring at least one signal containing information on the cardiac activity of said patient and a computing unit arranged to implement the method according to the invention.
[0073] It is conceivable that all the elements of the prediction system could be embedded in the same device, such as a laptop, a smartwatch or a smartphone, or conversely be distributed across different devices that are remote from each other, such as a mobile terminal and different computer servers, and connected to each other by means of wireless communication.
[0074] The invention also relates to a computer program product comprising instructions which, when the program is executed by a processor, lead the processor to implement the steps of the process according to the invention.
[0075] The invention also relates to a computer-readable storage medium comprising portions of code from a computer program intended to be executed by a processor to implement the steps of the process according to the invention.
[0076] The present invention is now described by means of purely illustrative and in no way limiting examples of the scope of the invention, and from the accompanying drawings, in which the various figures represent:
[0077] [Fig. 1] represents, schematically and partially, a system for predicting a patient's future health state according to one embodiment of the invention;
[0078] [Fig. 2] represents, schematically and partially, a method for predicting a future state of health of a patient according to an embodiment of the invention, implemented by the system of [Fig. 1];
[0079] [Fig. 3] represents, schematically and partially, an example of a signal containing information on a patient's cardiac activity used in the process of [Fig. 2];
[0080] [Fig. 4] represents, schematically and partially, a spectral energy density of the signal from [Fig. 3] and exploited during the implementation of a step in the process of [Fig. 2];
[0081] [Fig. 5] represents, schematically and partially, a sequence of heartbeats and a sequence of RR intervals extracted from the signal of [Fig. 3] during the implementation of a step in the process of [Fig. 2];
[0082] [Fig. 6] represents, schematically and partially, a step in the process of [Fig. 2] which aims to extract a feature from a multi-scale representation of the RR interval sequence of [Fig. 5];
[0083] [Fig. 7] represents, schematically and partially, a step in the process of [Fig. 2] which aims to extract a feature from a lexicographical representation of the RR interval sequence of [Fig. 5];
[0084] [Fig. 8] represents, schematically and partially, a step in the process of the [Fig. 2] whose object is the extraction of a feature from a representation of the density of the distribution of the RR interval sequence of [Fig. 5];
[0085] [Fig. 9] schematically and partially represents a support vector machine used during the implementation of a signal classification step in [Fig. 3]; and
[0086] [Fig. 10] represents, schematically and partially, a sequence of future RR intervals generated by an auto-regression model during the implementation of a prediction step of the process of [Fig. 2],
[0087] In the description that follows, identical elements, by structure or by function, appearing on different figures retain, unless otherwise specified, the same references.
[0088] The processes described below can also be implemented by software programs executable by a computer system. Furthermore, their implementation can be achieved through distributed processing and / or parallel processing, particularly for processing multiple data points simultaneously.
[0089] The figures described in this document are intended to provide a general understanding of the invention in various embodiments. These figures are not intended to serve as a complete description of all the elements and features of the devices, processors, and systems necessary for the invention. Many other embodiments of the invention, or combinations thereof, may be apparent to those skilled in the art upon reading this description, by combining the disclosed embodiments. Other embodiments may be derived from the description, so that structural and logical substitutions and changes may be made without departing from the scope of the present invention.
[0090] Furthermore, the description and figures should be considered illustrative rather than restrictive, and the appended claims are intended to cover all modifications, improvements, and other embodiments of the invention. Therefore, the scope of the following claims should be determined by the broadest possible interpretation of the claims and their equivalents and should not be restricted or limited by the preceding description.
[0091] Description of the implementation methods
[0092] Figure 1 shows a computer system 1 for predicting a patient's future health status.
[0093] This computer system 1 includes a computer terminal 2 comprising at least one sensor capable of acquiring at least one signal containing information on the patient's cardiac activity.
[0094] In the example described, computer terminal 2 is a laptop equipped with a camera and software for acquiring a video sequence V of the patient's face. This video sequence V can be processed to extract a remote photoplethysmogram or rPPG signal.
[0095] In unrepresented variants, terminal 2 may be another type of device This device can be handled and / or worn by the patient and is equipped with one or more sensors capable of acquiring, directly or indirectly, one or more electrophysiological signals whose characteristics, particularly temporal and / or frequency-related, directly or indirectly reflect, over any time scale, variations in the patient's heart rhythm. Terminal 2 could, for example, be a smartphone or tablet equipped with a camera to acquire video that can also be converted into an rPPG; a smart bracelet or watch; or a belt equipped with optical sensors to directly acquire a photoplethysmogram (PPG) signal, or electrodes to acquire an electrocardiogram (ECG), electroencephalogram (EEG), or electrogram (EGM) signal.
[0096] System 1 also includes a computing unit 3 capable of executing several steps of a process for predicting a future health state of the patient from the signal acquired by the sensor of terminal 2.
[0097] In the example described, the computing unit 3 is embedded in a remote server of the computer terminal 2. The computing unit 3 and the computer terminal 2 are thus each equipped with wireless communication means allowing them to exchange data through a wireless communication network.
[0098] The computing unit 3 is equipped with one or more processors, capable of executing instructions from one or more computer programs stored in a memory of the computing unit in order to implement the steps of the process of predicting a future health state of the patient.
[0099] In unrepresented variants, it may be possible for all these instructions to be executed by a single processor, or distributed across multiple processors on the server and / or across multiple processors on multiple servers, connected wirelessly and / or via wired communication. It may also be possible for the computing unit 3 to be integrated directly into the terminal 2.
[0100] Figure 2 shows a method for predicting a patient's future health status, implemented by system 1 of Figure 1.
[0101] The method includes a preliminary step E01 of acquiring a signal containing information about the patient's cardiac activity. In the context of [Fig. 1], this signal is formed by the video V of the patient's face. This video contains variations in the patient's skin color under the influence of blood flow and reflux in the subcutaneous venous system, and therefore reflects variations in heart rhythm.
[0102] In a step El, a photoplethysmogram type signal rPPGp is extracted from the video V. An example of a part of an rPPGp signal that can be extracted from a video V is shown in [Fig. 3].
[0103] This rPPGp signal can, for example, be extracted by segmenting a region of interest corresponding to a skin area in images of video V, using one of the methods the following or by a combination of several of these methods: colorimetric segmentation; segmentation by neural network, in particular of the ResNet type; extraction of region of interest by Haar cascade.
[0104] The average of the data from the region of interest segmented in each image of the video V thus translates the variations in light intensity and therefore in the color of the patient's skin and thus forms the rPPGp signal.
[0105] It should be noted that this step El can be implemented indifferently by the computing unit 3, as the images of the video V transmitted by the terminal 2 are received or after complete reception of the video V, or by a computing unit of the terminal 2 itself.
[0106] In a step E2, the type of noise contained in the rPPGp signal is classified temporally, so as to be able to identify whether the noise contained in a given sequence of the rPPGp signal is stationary noise or impulsive noise.
[0107] To this end, step E2 comprises several steps E21 and E22 for determining the temporal distribution of the energy of the rPPGp signal in at least one spectral band. The spectral energy density Es of the rPPGp signal from [Fig. 3] is shown in [Fig. 4], in which spectral bands are identified as low frequency Bf, whose frequency is less than 0.5 Hz, normal frequency Nf, whose frequency is between 0.5 Hz and 4.5 Hz, and high frequency Hf, whose frequency is greater than 4.5 Hz.
[0108] In a first step E21, the calculation unit 3 determines a value R representing the ratio of the spectral energy, for example measured by a weighted L2 standard, of the components of the rPPGp signal belonging to the low frequency ranges Bf and high frequency ranges Hf, to the spectral energy, for example also measured by the weighted L2 standard, of the components of the rPPGp signal belonging to the normal frequency range Nf.
[0109] [Math
[0110] In a second step E22, the computing unit 3 performs a time-frequency or time-scale decomposition of the rPPGp signal, by wavelet decomposition or a cascade of wavelet decompositions, and then determines one or more values No representing the ratio of the norm Lo of the spectral energy, by time scales, of the components of the rPPGp signal belonging to the low frequency ranges Bf and high frequency ranges Hf, to the norm Lo of the spectral energy of the components of the rPPGp signal belonging to the normal frequency range Nf.
[0111] [Math 2] / V0=
[0112] It should be noted that the components of the rPPGp signal can be estimated using, interchangeably, a Fourier transform, a wavelet transform, or any transform that allows decomposition of a signal into a basis of kernels, enabling analysis of this signal in the frequency domain or in the time-frequency or time-domain. ladders.
[0113] Alternatively or cumulatively, in other unrepresented steps, the computing unit 3 will be able to determine other values reflecting a temporal distribution of the energy of the rPPGp signal, such as dispersion coefficients extracted from a cascade of wavelet decompositions of this signal, also called WST (from the English "Wavelet Scattering Transform").
[0114] The different characteristics RA, NO (and others) obtained at the end of these steps E21, E22 (and others) are thus representative of the structure of the rPPGp signal, of the distribution and / or diffusivity of its spectral energy and / or of its dispersion in the frequency domain and / or time-frequency and / or multi-scale time-frequency domain and thus characterize the type of noise contained in this rPPGp signal.
[0115] In a step E23, the computing unit 3 can thus classify the type of noise contained in sequences of said rPPGp signal according to different time scales among a class of stationary noise and a class of impulse noise according to these values and determine a parameter vector X intended to parameterize a filter to denoise this rPPGp signal in order to extract a data sequence whose values or variations reflect the heart rate of the patient and / or a variability in the patient's heart rate.
[0116] It should be noted that the calculation unit 3 can classify the noise from the characteristics R and No, to then determine the values of the parameter vector X from this class, or directly determine the values of the parameter vector X from the characteristics RA and No.
[0117] In the example described, the calculation unit 3 compares each of the RA and No features to a predetermined threshold value to select a parameter vector X from a set of possible parameter vectors based on these comparisons.
[0118] In unrepresented variants, the computing unit 3 may determine the values of the parameter vector X using other methods, notably a classification algorithm trained to classify the noise contained in a signal based on characteristics representative of the spectral energy distribution and / or diffusivity in multiscale frequency and / or time-frequency and / or time-frequency components of the signal, or even based on the signal itself. This algorithm could, for example, be a k-means or DBSCAN type algorithm.
[0119] In a step E3, the computing unit 3 extracts a heart rate-related RR data sequence from the rPPGp signal by implementing a parameterized filtering algorithm using the parameter vector X.
[0120] In the example described, the filtering algorithm makes it possible to identify "R" peaks from the rPPGp signal, each peak corresponding to a heartbeat of the patient, the RR data sequence being formed by a temporal sequence of the duration of each interval separating two consecutive heartbeats. We have represented in the upper part of [Fig. 5] a sequence of R peaks identified from the rPPGp signal, and in the lower part a complete RR data sequence extracted from a complete sequence of R peaks of which the sequence of R peaks represented in the upper part is a part.
[0121] In order to identify the peak sequence R, the parameterized filtering algorithm implemented during step E3 is an edge-preserving filter parameterized using the parameter vector X and including a regularization intended to minimize the differences between said peak sequence R and the rPPGp signal under a sparsity constraint of a maximum time gradient of said peak sequence R.
[0122] The filter is thus intended to remove noise from the rPPGp signal by favoring strong variations in the rPPGp signal and ensuring that the sequence of peaks R can be meaningfully represented by only components approaching a Dirac peak, or in other words that the sequence of peaks R approaches a discrete signal as closely as possible.
[0123] It should be noted that other types of signals can be processed similarly to obtain a heart rhythm data sequence from a signal containing information about the patient's cardiac activity, adapting certain specific steps as needed. For example, for an ECG signal, the heart rhythm data sequence can be extracted by identifying certain wave segments, such as a QRS complex, and measuring the RR intervals, PP intervals, or QT intervals.
[0124] In a step E4, the computing unit 3 can then extract from the RR data sequence different features likely to define a signature or rhythmic fingerprint of a patient's health condition.
[0125] In a first step E41, the computing unit 3 determines one or more multi-scale representations of the RR data sequence represented by features Si, j.
[0126] Figure 6 illustrates an example of the implementation of step E41, in which the computing unit 3 performs a cascade of wavelet decompositions of the RR data sequence, also known as WST (Wavelet Scattering Transform). The Sy characteristics correspond to the dispersion coefficients obtained at the end of each decomposition of this cascade.
[0127] More specifically, in a SU substep, the computing unit 3 performs a first wavelet decomposition of the RR data sequence from first-rank wavelets selected from a wavelet bank.
[0128] In a substep S12, the calculation unit 3 determines the modulus of this first decomposition and then applies a scaling function to this modulus in a substep S13, such as an average function performed by a low-pass filter. Substep S13 thus yields a first series of coefficients Sj,i.
[0129] In a second iteration, the computing unit 3 then reiterates the wavelet decomposition of substep SU on the signal from the previous substep S12, in selecting higher rank wavelets from the data bank to change scale, redetermines the modulus in a new substep S12 and applies the scaling function in a new substep S13 to obtain a second set of coefficients Si, 2.
[0130] Steps S1, S12, and S13 are thus repeated in cascade. The set of coefficients Sij obtained at the end of steps S13, called dispersion coefficients Si, form a multi-scale representation in the time-frequency domain of the RR sequence. These dispersion coefficients Sij allow, among other things, the analysis of the structure of the RR sequence at different temporal and frequency resolutions, and in particular quantify how the moments of the RR sequence, such as its mean and variance, vary with changes of scale, and thus allow the quantification of how heart rate fluctuations are distributed and reproduced at different scales in the RR sequence.
[0131] The computing unit 3 will be able to determine other multi-scale representations of the RR sequence, such as a multifractal spectrum of this RR sequence, complete or estimated through certain characteristics such as multifractal cumulants, and in particular the cumulants Co, Ci, C2 and C3, the Hurst exponent, the fractal dimension of the support of the RR sequence, the degree of fractionality of the RR sequence and / or its degree of multifractality, or even the log of the variations of the RR sequence.
[0132] In a second step E42, the computing unit 3 determines one or more lexicographical representations of the RR data sequence represented by DL features.
[0133] Figure 7 shows an example of the implementation of step E42, in which the calculation unit 3 determines a lexicographic distribution Di of the sequence RR.
[0134] More specifically, in a first sub-step S21, the RR intervals of the R-R sequence are classified, by thresholding with regard to threshold values TS1 and TS2, among short interval classes "C", normal interval "N" and long interval "L".
[0135] The RR sequence can thus be transformed, or sampled, in a second sub-step S22, into an I CNL sequence of intervals “C”, “N” and “L”.
[0136] In a third sub-step S23, the sequence I CNL is decomposed into a series of MCNL sub-sequences, each composed, in the example described, of four successive intervals.
[0137] Each MCNL subsequence is, in other words, a pattern or word composed of the letters "C", "N" and "L". The set of all possible combinations of these letters "C", "N" and "L" thus forms a dictionary of subsequences, containing 256 different words in the example described.
[0138] In a fourth substep S24, the calculation unit 3 can then establish the lexicographic distribution DL of the RR sequence by counting, for each subsequence of the dictionary, the number of occurrences of that subsequence in the MCNL subsequence, thus giving an indication of the frequency or probability occurrence of this subsequence in the RR sequence. The dictionary subsequences can then be ordered according to their frequency of occurrence in the MCNL subsequence, to obtain the said lexicographic distribution DL of the RR sequence.
[0139] The computing unit 3 will also be able to characterize this DL lexicographic distribution by determining entropies, dominant modes and / or moments of the MCNL subsequence sequence, and / or the asymptote of said DL lexicographic distribution.
[0140] In a third step E43, the computing unit 3 determines one or more representations of the density distribution of the values of the intervals of the RR sequence represented by characteristics P(W).
[0141] Figure 8 shows an example of the implementation of step E43, in which the computing unit 3 determines a representation P(W) of the density of the distribution of the interval values of the RR sequence around its mean value.
[0142] More specifically, in a first step S31, the RR sequence is windowed by a window of size W, and the computation unit 3 determines the average value A of the RR sequence in this window W.
[0143] In a second step, S32, each value in the RR sequence is placed on a unit circle relative to a pole corresponding to its mean value, A. A time interval longer than the mean value A is thus placed clockwise past the pole, and a time interval shorter than the mean value A is placed clockwise past the pole. Groupings of values within these intervals can then be observed at specific points, RI to R4, either by sampling the RR sequence values at these points RI to R4 before step S32 or after the values have been placed on the unit circle.
[0144] The number of values clustering on these points RI to R4 can thus be counted in a step S33 to obtain a probability that an interval presents one of the values corresponding to these points RI to R4 and therefore estimate a distribution density of the values of the intervals around their average value in the window W.
[0145] By repeating this analysis, with the window W shifting temporally along the RR sequence, thus varying the intervals considered, and by normalizing the values of the intervals within this window to maintain a substantially constant average value, the computing unit can establish a counting function P(W) of the RR sequence. This function allows for the identification of the data distribution within the RR sequence and the identification of temporal trends in the RR intervals, particularly specific periods in which the heart rate deviates significantly from its average value according to a certain probability distribution, which may reflect a particular health condition.
[0146] In addition to establishing the function P(W), the computing unit 3 will be able to characterize this function P(W) with several values relating to this function, such as extreme values and / or its diffusion law and / or measures of deviations of this function with a given statistical law.
[0147] With further reference to [Fig. 2], calculation unit 3 will be able to, in addition to steps E41 at E43 allowing the extraction of multi-scale, lexicographic and distribution representations of said RR sequence, the computing unit will be able to extract other features or representations of the RR sequence, in steps not represented.
[0148] Calculation unit 3 will, for example, be able to extract features relating to the statistical distribution of the RR sequence, such as dominant modes, moments and extreme values of RR intervals.
[0149] The computing unit 3 will be able, for example, to extract features from the frequency spectrum of the RR sequence, such as the physiological Fourier modes, namely low frequency peaks corresponding to the sympathetic / adrenergic system and a high frequency peak corresponding to the parasympathetic / vagal system.
[0150] The computing unit 3 will be able, for example, to extract characteristics relating to the periodicity of the RR sequence, such as its level of periodicity evaluated through a singular value decomposition of the Toeplitz matrix of the RR sequence or through the spectrum of the covariance matrix of the RR sequence.
[0151] In step E5, the calculation unit 3 classifies, from the set of values extracted in step E4, the RR sequence into classes "normal sinus rhythm", "noise" and one or more rhythm classes indicating a predetermined state of health.
[0152] Without limitation, said classes may be one or more, or even all of the following: abnormal sinus rhythm indicating diabetes, abnormal sinus rhythm indicating hypertension, abnormal sinus rhythm indicating heart failure, abnormal sinus rhythm indicating stress, abnormal sinus rhythm indicating fatigue, abnormal sinus rhythm indicating depression, abnormal sinus rhythm indicating cancer recurrence, abnormal sinus rhythm indicating chronic obstructive pulmonary disease, abnormal sinus rhythm indicating sleep apnea, rhythm indicating atrial fibrillation, rhythm indicating ventricular tachycardia, rhythm indicating frequent extrasystoles, rhythm indicating sinus tachycardia, rhythm indicating sinus bradycardia, abnormal sinus rhythm indicating atrial flutter,Abnormal sinus rhythm indicating a conduction abnormality, or any other rhythm pattern indicating a heart rhythm disorder.
[0153] In the example described, this classification is carried out using at least one machine learning algorithm trained to classify a data sequence into said classes based on features from multi-scale, lexicographic and distributional representations of this data sequence.
[0154] Figure 9 illustrates an example of implementing step E5 using a support vector machine (SVM) trained to classify an RR sequence into five classes: "normal sinus rhythm," "noise," "rhythm indicating atrial fibrillation," "rhythm indicating ventricular tachycardia," and "other rhythm." This last class avoids an SVM convergence problem.
[0155] Previously, a dataset was established from a set of 500 signals from rPPG type and a set of 800 ID ECG signals from a plurality of patients with healthy status, atrial arrhythmia, or ventricular arrhythmia were used. RR interval sequences ranging from 10 seconds to 10 minutes in duration were extracted from these signals. Each RR sequence was manually annotated according to one of three classes: "normal sinus rhythm," "rhythm indicating atrial fibrillation," and "rhythm indicating ventricular tachycardia," "other rhythm," or "noise" if it was not possible to identify one of the other rhythms.
[0156] The Sy, DL, and P(W) features derived from multiscale, lexicographic, and distributional representations were extracted from each of these sequences according to steps E41 to E43. Each group of Sy, DL, and P(W) features was labeled with the class with which the sequence from which they were extracted was annotated. The set of annotated feature groups was separated into a training dataset and a validation dataset using a random sampling method.
[0157] A multi-class support vector machine was thus trained to classify each set of data in the training set among the four classes in order to identify support vectors, among these data in the training set, thus allowing to determine one or more optimal separation hyperplanes and to adjust weighting coefficients by a quadratic optimization method.
[0158] As a non-limiting example, the described example uses a multi-class linear SVM of the OVA or "one vs. all" type, composed of an elementary linear SVM (SVMi to SVMN) for each rhythm class (Ci to CN), trained to determine whether a given set of features belongs to that class or not. Each elementary linear SVM employs a Gaussian kernel function with a gamma bandwidth set between 1 and 10, a margin function regularization coefficient set less than 1, and is trained with the training dataset with a stopping tolerance set less than 0.001.
[0159] Each elementary linear SVM (SVMi to SVMN) of the multi-class linear SVM has its own hyperplane (HPI to HPN) and therefore its own decision function, which determines whether a given set of features belongs to the class Ci to CN associated with that elementary linear SVM. To classify the RR sequence, the feature set Sij, DL, and P(W) is provided to each elementary linear SVM (SVMi to SVMN), which predicts whether the RR sequence belongs to the associated class Ci to CN and provides an evaluation of the margin function fci to fcN. The final class Ck predicted by the multi-class linear SVM is the one, among all predicted classes, with the largest margin function. As an alternative to the "one vs all" type SVM, an OVA or "one vs one" type SVM can be used.
[0160] In other implementation examples, other types of machine learning algorithms can be used to classify a data sequence into these classes based on features derived from multi-scale, lexicographical representations. and the distribution of a data sequence, which can be trained in a supervised, unsupervised, self-supervised, and reinforcement manner, and in particular of the types deep neural networks, convolutional neural networks, recurrent neural networks, decision trees, data clustering, and any combination thereof.
[0161] It should be noted that at the end of step E5, a method is available for generating a database of rhythmic RR sequences, each labeled with the class predicted by the SVM. The extracted characteristics thus make it possible to establish rhythmic signatures or fingerprints of health states, inscribed in the heart rhythmic pulse, and to differentiate these rhythmic signatures from measurement noise.
[0162] In step E6, the computing unit 3 will predict a probability of a future health state of said patient, from the class Ck established at the end of step E5 and the RR sequence.
[0163] In a substep E61, the computing unit 3 synthesizes a sequence of potential future data R-RP by progressively propagating the data sequence RR.
[0164] We have represented in [Fig. 10] in the upper part a DGP machine learning algorithm implementing a deep model of autoregression of Gaussian process and in the lower part an example of R-RP sequence generated from the RR sequence of [Fig. 5] and by means of the DGP model.
[0165] The DGP model comprises a succession of GPi to GP layers m Each layer corresponds to a Gaussian process model, with the layers interconnected in pairs so that each output of each layer is connected to each input of each subsequent layer. Each GPi to GP layer m is an autoregression model capable of generating a data sequence from a sequence of input data, the model having been previously trained, in an unsupervised manner, to infer or fit a Gaussian kernel fi to f m said GPi to GP model mwhich minimizes a distance between a spectrum of a predicted data sequence and therefore a spectrum of said kernel and a spectrum of the training data provided as input.
[0166] In the example in [Fig. 10], the DGP model can have three interconnected layers. The first layer, GPi, has 200 inputs to receive the last 200 intervals of the RR sequence and predicts, according to its kernel fi, a data sequence that is provided as input to the second layer. The predicted data sequences are thus propagated from layer to layer until they reach the last layer, GPi. m which predicts the potential future R-RP sequence.
[0167] Alternatively, other Bayesian autoregression models, particularly those of order strictly greater than 1, can be used to generate the potential future sequence R-Rp.
[0168] In a substep E62, the computing unit predicts a probability Pk of a future health state of said patient, based on said potential future data sequence R-RP and the class Ck established at the end of step E5. In order to enhance the reliability of the predictions, the process includes a preliminary step E02 of providing a medical history of the patient, which is also taken into account in the prediction of the probability Pk.
[0169] This health condition may be a health condition involving an arrhythmia, an altered sinuous variation, such as a conduction abnormality, atrial and supraventricular arrhythmia, ventricular arrhythmia, cardiovascular disease, stroke, heart failure, cardiac arrest, sudden death, diabetes, hypertension, stress, fatigue, depression.
[0170] In the example described, the prediction of probability Pk can be implemented using the same SVM as that employed in step E5 to classify the potential future R-RP data sequence into classes of "normal sinus rhythm," "noise," and one or more rhythm classes indicating a predetermined health status. The class obtained as output from the SVM, along with the patient's medical history, is provided as input to a layer of a Bayesian autoregressive vector model to predict the probability Pk.
[0171] The preceding description clearly explains how the invention achieves its stated objectives, namely to provide a method for providing a practitioner with indications of a patient's possible future health status from data derived, at heartbeat resolution, from signals in particular from a consumer device in order to allow the practitioner to interpret these signals, to diagnose, predict and monitor the patient's condition and, where appropriate, to propose suitable treatments.
[0172] In any event, the invention cannot be limited to the embodiments specifically described in this document, and extends in particular to all equivalent means and to any technically operative combination of these means.
Claims
Demands
1. A method for predicting a patient's future health status from at least one signal (rPPGp) containing information on cardiac activity of said patient, the method being computer-implemented and comprising the following steps: a. (E3) extracting a data sequence (RR) relating to the heart rhythm from said signal; b. (E4, E41, E42, E43) extracting at least: i. a feature (DL) from a lexicographical representation of said data sequence; and / or ii. a feature (P(W)) from a representation of the data distribution density of said data sequence around a centered value from said data sequence; c. (E5) classifying the signal into classes (Ci, CN) "normal sinus rhythm", "noise" and at least one class indicating a predetermined health status from said extracted feature(s) (Sy, DL, P(W)); d.(E6) prediction, according to a given autoregression model (DGP) and from said class (Ck) of the signal, of a probability (Pk) of a future health state of said patient.
2. A method according to the preceding claim, characterized in that it comprises a preliminary step (E2) of classifying the type of noise contained in said signal, and in that said data sequence (RR) relating to the heart rhythm is extracted from said signal and from at least one filtering algorithm parameterized using the result of said noise type classification.
3. A method according to the preceding claim, characterized in that the noise type classification step (E2) comprises a step (E21, E22) of determining the temporal distribution of the energy (R, NO) of the signal (rPPGp) in at least one spectral band, the noise type contained in said signal being classified among a stationary noise class and an impulsive noise class according to said distribution.
4. A method according to any one of claims 2 or 3, characterized in that the filtering algorithm parameterized using the result (X) of said noise type classification includes a minimization of the differences between said extracted data sequence (RR) and said signal (rPPGp) under a sparsity constraint of a maximum temporal gradient of said extracted data sequence.
5. A method according to the preceding claim, characterized in that the filtering algorithm is parameterized using the result (X) of said classification of the type of noise includes an edge-preserving filter parameterized using the result of said noise type classification.
6. A method according to any one of the preceding claims, characterized in that it comprises, in addition to the extraction step(s) (E4, E41, E42, E43) of said feature(s) from the lexicographic and / or distribution representations of said data sequence (RR), an extraction step (E41) of at least one feature (Si, from a multi-scale representation of said data sequence (RR), in that said extraction step comprises at least one cascade of wavelet decompositions (Sil, S12, S13) of said data sequence, said feature being determined from the dispersion coefficients (Si, obtained at the end of each decomposition of said cascade, and in that said signal is classified among classes (Ci, CN) "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health from the set of extracted features (Sij, DL, P(W)).
7. A method according to any one of the preceding claims, characterized in that the extraction step (E42) of a feature (DL) from a lexicographic representation of said data sequence (RR) comprises a classification (S21) of each of the data of said sequence among a plurality of predetermined classes (C,N,L), a transformation (S22) of said data sequence from data classes to obtain a sequence of classes (ICNL), a decomposition (S23) of said sequence of classes into a series of sub-sequences of classes (MCNL), a determination (S24) of a lexicographic distribution of said series of sub-sequences of classes, said feature (DL) being determined from said lexicographic distribution.
8. A method according to any one of the preceding claims, characterized in that the extraction step (E43) of a feature (P(W)) from a representation of the distribution density of the data of said data sequence (RR) around a centered value from said data sequence comprises the estimation (S31, S32, S33) of a counting function (P(W)) of the data of the data sequence, said feature being determined from said counting function.
9. A method according to any one of the preceding claims, characterized in that the extraction step comprises extracting at least one feature (DL) from a lexicographical representation of said data sequence and at least one feature (P(W)) from a representation of the data distribution density of said data sequence around a centered value from said data sequence, and in that said signal is classified into classes (Ci, CN) "normal sinus rhythm", "noise", and at least one class indicating a predetermined health status based on all extracted characteristics (S L j, D L , P(W)).
10. A method according to any one of the preceding claims, characterized in that it comprises, in addition to the extraction step(s) (E4, E41, E42, E43) of said feature(s) from the lexicographic and / or distribution representations of said data sequence (RR), the following steps: a. extraction of a statistical feature from said data sequence; b. extraction of a feature from the frequency spectrum of said data sequence; said signal being classified among said classes (Ci, CN) "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health from the set of extracted features (Sy, DL, P(W)).
11. A method according to any one of the preceding claims, characterized in that the classification step (E5) of the signal among the classes (Ci, CN) "normal sinus rhythm", "noise" and at least one class indicating a predetermined state of health, from said at least one extracted feature (Sj, DL, P(W)) is implemented by means of at least one machine learning algorithm (SVM) trained to classify a data sequence (RR) among said classes from features from lexicographic representations and / or distribution of a data sequence.
12. A method according to any one of the preceding claims, characterized in that the prediction step (E6), according to a given auto-regression model (DGP) and from said class (Ck) of the signal (rPPGp), of a future health state of said patient, comprises a substep (E61) of generating, from said signal class, a future potential data sequence (R-RP) of said data sequence (RR) according to said auto-regression model; and a prediction substep (E62), from said future potential data sequence and said signal class, of a probability (Pk) of a future health state of said patient.
13. A method according to any one of the preceding claims, characterized in that it comprises a step of supplying (E02) a medical history of the patient and in that said probability (Pk) of a future health state of said patient is predicted, according to said given autoregression model (DGP) and from said class (Ck) of the signal (rPPGp) and the medical history of the patient.
14. A system for predicting (1) a patient's future health status, comprising at least one sensor (2) capable of acquiring at least one signal (rPPGp) containing information on cardiac activity of said patient and a computing unit (3) arranged to implement the process according to one of the preceding claims.