METHOD FOR ESTIMATING PHYSIOLOGICAL SIGNALS
A neural network-based method processes cardio-respiratory signals to estimate SAHS, addressing diagnostic complexity and cost issues, improving comfort and accessibility, and enhancing diagnostic capabilities.
Patent Information
- Application Number
- FR2021010882
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-10-14
AI Technical Summary
Current diagnostic methods for Sleep Apnea Hypopnea Syndrome (SAHS) are complex, costly, and uncomfortable, leading to underdiagnosis due to high equipment costs, logistical barriers, and patient discomfort, with existing digital technologies providing limited diagnostic value and material limitations.
A computer-implemented method using a non-linear regression model of the neural network type that processes cardio-respiratory signals, such as sounds and movements, to estimate physiological ventilatory polygraphy and electrocardiography signals, reducing the need for multiple sensors and improving diagnostic comfort and accessibility.
Enables reliable estimation of respiratory and cardiac signals using a single audio or movement sensor, simplifying the diagnostic process, reducing costs, and increasing the number of diagnosable patients by enhancing patient comfort and practicality.
Smart Images

Figure 00000026_0000 
Figure 00000026_0001 
Figure 00000027_0000
Abstract
Description
Title of the invention: METHOD FOR ESTIMATING PHYSIOLOGICAL SIGNALS FIELD OF THE INVENTION
[0001] The present invention relates to a computer-implemented system and method for estimating physiological signals of a subject. STATE OF THE ART
[0002] Sleep Apnea Hypopnea Syndrome (SAHS) is characterized by pauses (apneas) or breathing difficulties (hypopneas) during sleep. These respiratory events last at least 10 seconds and occur, depending on the severity of the SAHS, from 5 times per hour (mild SAHS) to more than 30 times per hour (severe SAHS). The causes of SAHS can be multiple, and several categories of respiratory events are classically defined by different clinical signs. SAHS can also be positional when the respiratory events mainly occur in the supine position. Two main types of SAHS are thus defined: obstructive sleep apnea syndrome (OSA) (Lévy, et al., 2015) and central sleep apnea syndrome (CSA) when at least 50% of respiratory events are of central origin (Eckert, Jordan, Merchia, & Malhotra, 2007).
[0003] SAHS has symptoms of excessive fatigue upon waking, daytime sleepiness or excessive snoring. It is associated with cardiovascular comorbidities such as heart failure, hypertension, stroke, cardiac arrhythmias or myocardial infarction (Javaheri, et al., 1998). Diagnosis is based on polysomnographic recording of respiratory activity (chest belt, abdominal belt, nasal flow, nasal pressure, snoring, oxygen saturation), cardiac activity (heart rate, photoplethysmogram, electrocardiogram), cerebral activity (EEG) and motor activity (actimeter, position sensor).In clinical practice in adults, a simplified outpatient polygraph recording is usually sufficient for the diagnosis of POSA and analyzes at least the following sensors: abdominal and thoracic belts, nasal cannula, microphone, pulse oximeter, position and oxygen saturation (Berry, et al., 2017).
[0004] OSA affects approximately 950 million people worldwide (Benjafield, et al., 2019). CSA and OSA are generally confused in these estimates. Risk factors have been identified such as obesity, age, sex (men are more affected), ethnicity, smoking, alcohol consumption, diabetes, respiratory tract abnormalities or neurological disorders. In France, we estimates that 24 million patients suffer from at least mild OSA, i.e. with an apnea-hypopnea index greater than 5 events per hour (Benjafield, et al., 2019). There are also strong correlations with other pathologies. Thus, 21% to 74% of patients with atrial fibrillation have POSA (Linz, et al., 2018); 29% to 67% of insomniac patients have POSA (Luyster, Buysse, & Strollo, 2010).
[0005] Treatments for SAHS exist and are of several types. However, screening and diagnosis are too rarely carried out and a large number of apneic patients are unaware of their condition. Thus, 80% of apneic patients are not diagnosed in the United States (Watson, 2016). This also represents a significant economic issue because an untreated apneic patient will cost three times more ($6,366) than a treated patient ($2,105) per year. The main obstacles to diagnosis are patient awareness, the lack of training of general practitioners in this pathology, the cost of diagnosis and treatment, and the immaturity of the healthcare system with regard to preventive medicine.
[0006] Regarding the cost of diagnosis, it has been considerably reduced by the systematic implementation of home respiratory polygraphy, which can be performed in private practice by doctors of different specialties (pulmonologist, ENT, cardiologist, psychiatrist) or by the general practitioner specializing in sleep (Safadi, Etzioni, Fliss, Pillar, & Shapira, 2014). Ventilatory polygraphy has a limited number of sensors, which allows the patient to equip themselves and perform the diagnostic examination on an outpatient basis at home. Although an overnight hospital stay is avoided and the cost is therefore lower in outpatient settings (Stewart, Penz, Fenton, & Skomro, 2017), the investment in equipment and the logistics of the examination (loan of equipment, recovery, disinfection, reading of tracings) remain a barrier to the number of diagnoses performed.Furthermore, the examination remains quite uncomfortable for the patient and he must go to his doctor several times during a single examination, which could be avoided. Finally, the average waiting time for a ventilatory polygraphy is shorter than for a polysomnography, but remains high in certain countries (10 months of delay between the first consultation and the diagnosis in the United Kingdom) which is detrimental for the patient (Flemons, Douglas, Kuna, Rodenstein, & Wheatley, 2004).
[0007] Over the past ten years, there has been increased innovation in screening and diagnostic devices driven by the constant increase in the use of digital technologies, the emergence of connected objects and artificial intelligence in health. There is a clear trend towards the simplification of sensors, to improve patient comfort and reduce the cost of equipment. These include measurement methods for pulse oximetry, tracheal sounds, snoring and respiratory sounds, ballistocardiogram, mandibular movements (Kelly, Strecker, & Blanchi, 2012) (Penzel, Schôbel, & Fietze, 2018). Electronic equipment is either purchased by patients (for consumer devices) or offered for rental: in the latter case, patients are invited to come to the office to collect the equipment and / or have it installed, or to receive the equipment by mail and install it themselves, or to welcome a service provider who can install the equipment. After examination, the patient must return the equipment or come directly to the office. Electronic equipment is connected, or simply equipped with a memory card that the doctor will connect to his PC to read the tracings and annotate respiratory events.Among the commonly available systems, only some connected objects have diagnostic value and simplify the respiratory polygraphy examination by reducing the number of sensors required; however, the number of examinations that can be carried out by these systems remains materially limited by the acquisition of these connected objects. Conversely, other systems, using for example only a microphone or a sonar process, have so far had no or little diagnostic value due to their limited functionalities and physiological measurements.
[0008] In this context, the present invention provides a solution for reducing the complexity of the acquisition system so as to drastically reduce the cost and logistics of the diagnostic examination, while being more comfortable and more practical for the patient. A greater number of patients could thus be diagnosed. SUMMARY
[0009] The present invention relates to a computer-implemented method for estimating physiological signals of a subject, said method comprising: - receiving at least one cardio-respiratory signal representative of sounds and / or movements originating from the rib cage and airways of the subject, the cardio-respiratory signal being acquired during the subject's sleep; - the use of a non-linear regression model of the neural network type configured and trained to receive, as input, at least one cardio-respiratory signal, and provide as output at least one estimate of a physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal; wherein the neural network comprises an encoder having one input and several outputs and a decoder having one output and several inputs, wherein at least two encoder outputs are directly connected to at least two respective inputs of the decoder, and at least two outputs of the encoder are connected to each other and wherein said neural network is configured to perform one-dimensional convolution operations; and - provide as output at least one estimate of a physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal.
[0010] The model used in the method of the invention advantageously makes it possible to obtain several signals having characteristic traces of the signals measured by the ventilatory polygraphy and / or electrocardiogram sensors by using as the only input a sound cardio-respiratory signal (i.e., acquired by a microphone) and / or a movement cardio-respiratory signal (i.e., acquired by a gyroscope and / or an accelerometer positioned in contact with the subject's thorax). This method therefore makes it possible to obtain ventilatory polygraphy or electrocardiogram (ECG) signals by means of only one audio sensor and / or one movement sensor instead of the set of sensors required in a standard ventilatory polygraphy or ECG acquisition (i.e., nasal cannula, thoracic belt, abdominal belt, electrodes, etc.).
[0011] In one embodiment, the neural network comprises a U-Net and the encoder comprises a ResNet. The U-shaped architecture of the U-Net advantageously makes it possible to connect several outputs of the encoder to several inputs of the decoder so as to extract information at several scales and then assemble them in the decoder. This gives a more reliable estimation of the physiological signals at the output.
[0012] In one embodiment, said physiological ventilatory polygraphy signal is chosen from: a nasal flow rate acquired via a nasal cannula, a nasal pressure acquired via a nasal cannula, a thoracic respiratory effort acquired via a thoracic belt or an abdominal respiratory effort acquired via an abdominal belt.
[0013] In one embodiment, said physiological electrocardiography signal is acquired via an electrocardiogram of at least one lead.
[0014] In one embodiment, the at least one cardio-respiratory signal is a cardio-respiratory movement signal obtained using a gyroscope and / or an accelerometer.
[0015] In one embodiment, the at least one cardio-respiratory signal is a cardio-respiratory sound signal; the method further comprising compressing the cardio-respiratory sound signal using a second convolutional neural network, prior to its application as input to the non-linear regression model.
[0016] In one embodiment, the method further comprises merging the compressed cardiorespiratory sound signal with the cardiorespiratory movement signal.
[0017] In one embodiment, the at least one cardiorespiratory signal is normalized before using the non-linear regression model.
[0018] In one embodiment, the neural network comprises an output layer having a linear activation function.
[0019] In one embodiment, the method further comprises detecting an abnormal episode associated with sleep apnea based on at least one ventilatory polygraphy signal delivered by the non-linear regression model.
[0020] In one embodiment, the method further comprises detecting an abnormal episode associated with a cardiac arrhythmia based on at least one cardiac physiological signal delivered by the non-linear regression model.
[0021] The present invention further relates to a device for estimating physiological signals of a subject, said device comprising: - at least one input configured to receive at least one cardio-respiratory signal representative of sounds and / or movements coming from the rib cage and the airways of the subject, the cardio-respiratory signal being acquired during the subject's sleep; - at least one processor configured to: estimate at least one physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal using a non-linear regression model of the neural network type configured and trained to receive, as input, the at least one cardio-respiratory signal, and provide as output the at least one estimate of a physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal; wherein the neural network comprises an encoder having an input and several outputs and a decoder having an output and several inputs, wherein each encoder output is directly connected to a respective input of the decoder, and at least two outputs of the encoder are connected to each other and wherein said neural network comprises one-dimensional convolution operations; - at least one output configured to provide at least one ventilatory polygraphy signal and / or a physiological electrocardiography signal.
[0022] In one embodiment, the system further comprises an accelerometer, a gyroscope and / or a microphone.
[0023] The present invention further relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the latter to implement the method for estimating physiological signals of a subject according to any one of the embodiments described above.
[0024] The present invention further relates to a non-transitory computer-readable recording medium comprising instructions which, when the program is executed by a computer, cause the computer to implement the method for estimating physiological signals of a subject according to any of the embodiments described above.
[0025] Such a non-transitory computer-readable recording medium may be, without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor device, or any suitable combination of the foregoing. It will be appreciated that the following, while providing more specific examples, is not that an illustrative and non-exhaustive list as easily appreciated by a person skilled in the art: a laptop floppy disk, a hard disk, a ROM, an EPROM (Erasable Programmable ROM) or a Flash memory, a portable CD-ROM (Compact-Disc ROM). DEFINITIONS
[0026] In the present invention, the terms below are defined as follows: the term "processor" should not be interpreted as being limited to computer hardware capable of executing software, and generally refers to a processing device, which may for example include a computer, a microprocessor, an integrated circuit or a programmable logic device (PLD). The processor may also include one or more graphics processing units (GPUs), whether exploited for computer graphics and image processing or other functions.In addition, the instructions and / or data for executing the associated and / or resulting functionalities may be stored on any medium readable by the processor such as, for example, an integrated circuit, a hard disk, a CD (Compact Disc), an optical disk such as a DVD (Digital Versatile Disc), a RAM (Random-Access Memory) or a ROM (Read-Only Memory). The instructions may in particular be stored in computer hardware, software, microprograms ("firmware") or in any combination thereof.
[0027] A "neural network" or "artificial neural network (ANN)" refers to a category of machine learning algorithm comprising nodes (called neurons), and connections between neurons modeled by "weights". For each neuron, an output is given as a function of an input or set of inputs by an "activation function". Neurons are generally organized into several "layers", so that neurons in one layer only connect to neurons in the immediately preceding and immediately following layers.
[0028] A "convolutional neural network" refers to a type of acyclic ("feed-forward") artificial neural network comprising multiple hidden layers. The hidden layers are generally convolutional layers followed by activation layers, some of which are followed by "pooling" layers.
[0029] “Inertial unit” relates to instruments comprising six sensors, of which three gyrometers measuring the three components of the angular velocity vector, and three accelerometers measuring the three components of the specific force vector. DESCRIPTION OF THE FIGURES
[0030] [Fig.l]: [Fig.l] is a functional diagram schematically representing a particular mode of a device for estimating respiratory signals and / or cardiac signals from cardio-respiratory signals, in accordance with the present disclosure.
[0031] [Fig.2]: [Fig.2] provides an example of cardio-respiratory signals received at input of the device and method of the present invention.
[0032] [Fig.3]: [Fig.3] is a functional diagram schematically representing the structure of the U-shaped architecture of the nonlinear regeneration model.
[0033] [Fig.4]: [Fig.4] is a functional diagram schematically representing the structure of the U-shaped architecture of the nonlinear regression model combined with the second neural network for sound signal compression.
[0034] [Fig.5]: [Fig.5] reports the respiratory physiological signals obtained from signals in [Fig.2].
[0035] [Fig.6]: [Fig.6] reports the cardiac physiological signals of type seismocardiogram (SCG) and gyrocardiogram (GCG) obtained from signals in [Fig.2],
[0036] [Fig.7]: [Fig.7] is a flowchart showing the successive steps carried out with the estimation device of [Fig.l].
[0037] [Fig.8]: [Fig.8] is a functional diagram schematically representing a particular mode of a regression model training device, in accordance with the present disclosure and adapted to the prediction device of [Fig.l].
[0038] [Fig.9]: [Fig.9] schematically represents a device integrating the functions of the estimation device of [Fig.l] and the training device of [Fig.8].
[0039] [Fig. 10]: [Fig. 10] shows a physiological signal of the respiratory effort type thoracic estimated by device 1 (curve Cl) superimposed on a signal acquired directly by a thoracic belt (curve C2) synchronously with the cardio-respiratory signal used in device 1 for the estimation.
[0040] [Fig. 11]: [Fig. 11] shows, from top to bottom, a first physiological signal of the thoracic respiratory effort type estimated by the device 1, a second physiological signal of the abdominal respiratory effort type estimated by the device 1, a superposition of the first signal with a signal acquired directly by a thoracic belt in a synchronous manner, and a superposition of the first signal with a signal acquired directly by an abdominal belt in a synchronous manner.
[0041] [Fig. 12]: [Fig. 12] shows, from top to bottom, a first physiological signal of the type a nasal flow estimated by the device 1, a second physiological signal of the type nasal pressure estimated by the device 1, a superposition of the first signal with a signal acquired directly by a nasal cannula synchronously and a superposition of the second signal with a signal acquired directly by a nasal cannula synchronously. DETAILED DESCRIPTION
[0042] The present description illustrates the principles of the present disclosure. It will therefore be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the disclosure and are included within its scope.
[0043] All examples and conditional language cited herein are intended for educational purposes to assist the reader in understanding the principles of the disclosure and the concepts contributed by the inventor to advance the state of the art, and should be construed as not being limited to these specifically cited examples and conditions.
[0044] Furthermore, all statements of principles, aspects and embodiments of the disclosure, as well as their specific examples, are intended to encompass their structural and functional equivalents. Furthermore, it is intended that these equivalents include both currently known equivalents and equivalents developed in the future, i.e., all developed elements that perform the same function, regardless of their structure.
[0045] Thus, for example, those skilled in the art will understand that the block diagrams presented herein may represent conceptual views of illustrative circuits implementing the principles of the disclosure. Similarly, it will be appreciated that all flowcharts, flow diagrams, and the like represent various processes that may be substantially represented on a computer-readable medium and thereby executed by a computer or processor, whether or not that computer or processor is explicitly shown.
[0046] The functions of the various elements illustrated in the figures may be performed by the use of dedicated computer hardware as well as computer hardware capable of executing software in association with appropriate software. When performed by a processor, the functions may be performed by a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which may be shared.
[0047] It is understood that the elements illustrated in the figures may be implemented in various forms of hardware, software, or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more suitably programmed general-purpose devices, which may include a processor, memory, and input / output interfaces.
[0048] The present disclosure will be described with reference to a particular functional embodiment of a device 1 for estimating physiological signals of a subject, as illustrated in [Fig.l].
[0049] The device 1 is in fact adapted to provide as output an estimation of at least one physiological ventilatory polygraphy signal and / or at least one physiological electrocardiography signal 31 from the analysis of the cardiorespiratory signals 21. The estimated physiological ventilatory polygraphy signal may be at least one of the following signals: a nasal flow rate acquired via a nasal cannula, a nasal pressure acquired via a nasal cannula, a thoracic respiratory effort acquired via a thoracic belt, an abdominal respiratory effort acquired via an abdominal belt and / or an oxygen saturation acquired by a pulse oximeter or saturometer. In one embodiment, the estimated physiological ventilatory polygraphy signal further comprises a photoplethysmographic signal and / or an electromyogram of chin or leg muscles.
[0050] The estimated electrocardiography physiological signal 31 may be a signal acquired via an electrocardiogram with at least one lead.
[0051] Indeed, the invention aims to use the information included in a cardio-respiratory signal to deduce therefrom an estimate of the respiratory signals and / or cardiac signals as they would have been acquired by dedicated sensors such as nasal cannulae, thoracic or abdominal belts, electrodes for electrocardiograms and others. The physiological signals 31 at the output of the method therefore have characteristic traces of ventilatory polygraphy or ECG sensors so that the doctor who examines them can interpret them in the same way as ECG or ventilatory polygraphy signals acquired directly by the dedicated sensors.
[0052] The cardio-respiratory signals 21 may be acquired by means of one or more sensors positioned near the subject, preferably in contact with the subject. In one embodiment, the cardio-respiratory signals 21 representative of the sounds and movements originating from the rib cage and the airways of the subject, which will be referred to in this disclosure as "cardio-respiratory movement signals", are acquired using at least one gyroscope and / or at least one accelerometer. In one embodiment, the cardio-respiratory movement signals 21 are acquired by an inertial unit. Said inertial unit may be included in an electronic telecommunications device, normally portable, such as a mobile telephone.
[0053] In one embodiment, the cardio-respiratory signals 21 representative of the sounds coming from the rib cage and the airways of the subject, which will be called in this disclosure "cardio-respiratory sound signals", are acquired using a microphone. The microphone may be an independent sensor positioned in contact with the subject's torso or neck (i.e., directly in contact with the skin or in contact with a thickness, such as the fabric of an undershirt, which is itself in direct contact with the subject's skin). Alternatively, the cardio-respiratory sound signal may be acquired using a microphone integrated into an electronic telecommunications device, normally portable, such as a mobile phone. Indeed, in one example, the subject may position a mobile phone on his chest using a fixing device and all cardio-respiratory signals are therefore acquired using the sensors included in said phone. The phone may further be combined with other sensors.For example, in addition to the signals acquired by the telephone, other cardio-respiratory sound signals may be acquired by a microphone in contact with the subject's neck or positioned near the subject's nose and mouth. The cardio-respiratory sound signals are received by file in MP3 (MPEG-1 / 2 Audio Layer III) or AAC ("Advanced audio Coding") format.
[0054] The sensors included in the same system, such as for example an integrated circuit of the inertial unit type, can be synchronized by the same internal clock so that the signals are time-stamped synchronously. Alternatively, in the case of signals recorded via independent sensors, such as the microphone and the inertial unit of a telephone, a system for synchronizing acquired signals can be set up so as to obtain time-stamped signals synchronously.
[0055] The cardio-respiratory signals 21 can be acquired continuously by the sensors, preferably while the subject is asleep. The cardio-respiratory signals 21 can therefore comprise data acquired over 5 to 12 hours, generally 8 hours. The sampling frequency of the sound signals is between 5 kHz and 15 kHz, and the sampling frequency of the movement signals is between 80 and 150 Hz. The graphs in [Fig.2] show an example of the signals received as input by the module 11, comprising an audio signal and six movement signals acquired by the six sensors of an inertial unit.
[0056] The estimation device 1 is associated with a training device 6 adapted to the adjustment of the parameters of the non-linear regression model of the device 1, represented in [Fig.8], which will be described later.
[0057] Although the presently described devices 1 and 6 are versatile and have several functions that can be performed alternatively or in any cumulative manner, other implementations within the scope of the present disclosure include devices having only portions of the functionalities described in the present disclosure.
[0058] Each of the devices 1 and 6 is advantageously an apparatus, or a physical part of an apparatus, designed, configured and / or adapted to perform the mentioned functions and produce the mentioned effects or results. In alternative implementations, any of the device 1 and the device 6 is realized as a set of devices or physical parts of devices, whether grouped in the same machine or in different, possibly remote, machines. The device 1 and / or the device 6 may for example have functions distributed on a cloud infrastructure and be available to users as a cloud-based service, or have remote functions accessible via an API.
[0059] The estimation device 1 and the training device 6 may be integrated into the same device or set of devices, and intended for the same users. In other implementations, the structure of the device 6 may be completely independent of the structure of the device 1, and may be provided to other users. For example, the device 1 may have a parameterized model available to the operators for the estimations of the cardiac and / or respiratory physiological signals, entirely defined from the training carried out upstream by other actors with the device 6.
[0060] In the following disclosures, the modules are to be understood as functional entities rather than as physically distinct hardware components. They may therefore be materialized either as grouped in a single tangible and concrete component, or distributed across several of these components. Similarly, each of these modules is possibly itself shared between at least two physical components. Furthermore, the modules are implemented in the form of hardware, software, firmware or any mixed form thereof. They are preferably incorporated into at least one processor of the device 1 or the device 6.
[0061] The device 1 comprises a module 11 for receiving at least one cardiorespiratory signal 21 as well as parameters of the non-linear regression model 20 stored in one or more local or remote database(s) 10. The latter may take the form of storage resources available from any type of suitable storage means, which may in particular be a RAM or an EEPROM (Electrically-Erasable Programmable Read-Only Memory) such as a Flash memory, possibly within an SSD (Solid-State Disk). In advantageous embodiments, the parameters of the model 20 have been previously generated by a system comprising the device 6 for learning. Alternatively, the parameters of the model 20 are received by a communication network.
[0062] The device 1 further optionally comprises a module 12 for preprocessing the cardio-respiratory signals 21. The module 12 may in particular be adapted to normalize the data of the cardio-respiratory signals received 21 for the purpose of efficient and reliable processing. It may transform the cardio-respiratory sound signals 21, for example by decompression of the acquired sound data.
[0063] In advantageous embodiments, the module 12 is configured to preprocess the cardio-respiratory signals 21 so that the signals are normalized. This can improve the efficiency of downstream processing by the device 1. Such normalization can be particularly useful when the signals being used come from different sources, possibly including different acquisition systems.
[0064] The signals may be normalized using Z-score normalization. The signals are separated into time windows of approximately 15 to 60 seconds for motion signals and approximately 1 to 60 seconds for sound signals. These lengths for the time windows are determined to be able to detect sleep apnea in the estimates of the physiological signals.
[0065] Normalization is advantageously applied to the signals 21 of the training data sets (e.g. by the device 6) in a similar manner. In particular, the device 1 can process signals originating from a given type of source, while its parameters result from learning based on signals obtained with different types of sources. Thanks to the standardization, the differences between the sources can then be neutralized or minimized. This can make the device 1 more efficient and more reliable.
[0066] In one embodiment, wherein the at least one cardio-respiratory signal comprises a cardio-respiratory sound signal, the module 12 is further configured to perform compression of the cardio-respiratory sound signal using a convolutional neural network. Said convolutional neural network may be a residual neural network (ResNet), which are neural networks using shortcuts (“skip connections”) to skip certain layers. Said shortcuts advantageously make it possible to avoid the problem of vanishing gradients. The ResNet models may be implemented with two- or three-layer skips that contain non-linearities (ReLU) and batch normalization between the two. The sound signals are acquired with a higher sampling frequency (i.e., large data files) to that of motion signals, so advantageously, the compression of sound signals allows to reduce the computational cost and reduces the RAM required at the model usage stage.
[0067] In this embodiment comprising both the sound signals and the cardio-respiratory motion signals, the compressed sound signals are merged with the motion signals. In one example, the network that performs the compression is a ResNet-18 with a one-dimensional convolutional operation, configured to compress packets of 8092 sound samples into 256 samples comprising 512 components (“features”). These 512 components are then concatenated with 6 components of the inertial unit (i.e., acceleration and angular velocity for each of the three axes).
[0068] The device 1 comprises, among other things, a module 13 configured to use a non-linear regression model on cardio-respiratory signals so as to obtain at least one estimate of a physiological signal representative of either the cardiac activity or the respiratory activity of the subject. Said non-linear regression model is in particular configured to receive as input the at least one cardio-respiratory signal, and provide as output at least one estimate of a physiological signal of ventilatory polygraphy and / or a physiological signal of electrocardiography 31. The non-linear regression model used is a neural network configured and trained to receive, as input, the at least one cardio-respiratory signal, and provide as output at least one estimate of a physiological signal 31. Said neural network is further configured to perform one-dimensional convolution operations.A convolution is a linear operation that involves multiplying a set of weights with the input. In the case of one-dimensional input, the multiplication is performed between a vector of input data (i.e. time series) and a one-dimensional array of weights, called a filter or kernel.
[0069] Advantageously, this neural network has a structure comprising an encoder and a decoder. The encoder is constructed with one input and several outputs and the decoder with one output and several inputs. Furthermore, at least two encoder outputs are directly connected to at least two respective inputs of the decoder, and at least two outputs of the encoder are connected to each other. This type of structure makes it possible to solve the problem of "vanishing gradient".
[0070] In one embodiment, the neural network comprises a fully convolutional neural network model, with an encoder and a decoder, called a U-Net. The encoder is an assembly of convolution layers, “max pooling” layers allowing to create a map of signal components and to reduce their size to reduce the number of network parameters. The decoder allows an expansion with a path symmetrical to the contraction path of the encoder, which gives the U-shape of the architecture. The U-net, advantageously, allows to connect several outputs of the encoder to several inputs of the decoder. Thus, it is possible to extract information at each scale (i.e., each block that precedes a "max pooling"), and to assemble them in the decoder. This is different from the case of the auto-encoder which only retains the information of the lowest scale. In this embodiment, the U-net encoder has a ResNet architecture.
[0071] [Fig.3] shows an example of the architecture of the neural network configured to receive cardio-respiratory movement signals as input, possibly after a pre-processing step performed by the module 12. This figure shows in particular the “skip connections” present in the encoder (left column) characteristic of a ResNet and the horizontal connections between the encoder and the decoder characteristic of a UNet (i.e., horizontal arrows between the left column and the right column representing the concatenation of components of an output of the encoder with those of another output of the decoder). [Fig.4] shows a second example of an architecture built to process both sound signals and movement signals. This architecture includes, in addition to the previous architecture, the ResNet-18 neural network for the compression of sound signals.The neural network for compression is located before the U-Net encoder (first column on the left of the path).
[0072] In one embodiment, the neural network comprises an output layer having a linear activation function. The linear activation function is the identity function, and its use is advantageously suited to regression problems because the expected output value belongs to the set of real numbers.
[0073] The module 13 can further be configured to perform a normalization of signals at the output of the regression model. As for the input signals of the regression model, also certain output signals can be normalized using a Z-score normalization. The normalized signals include in particular the physiological signals except oxygen saturation and heart rate, of which it is interesting to know the absolute amplitude.
[0074] Figures 5 and 6 give examples of respiratory and cardiac physiological signals respectively. These signals in Figures 5 and 6 are obtained by analyzing the cardio-respiratory signals represented in [Fig.2].
[0075] The device may further comprise a module 14 configured to detect an abnormal episode associated with sleep apnea. In one embodiment, this abnormal episode is detected by analyzing the at least one ventilatory polygraphy signal delivered by the non-linear regression model, by applying a segmentation model to it. In an alternative embodiment, the module 14 is configured to directly analyze the cardio-respiratory signals by applying a segmentation model.
[0076] The module 14 may further be configured to detect an abnormal episode associated with a cardiac arrhythmia. In one embodiment, this abnormal episode is detected by analyzing the at least one cardiac physiological signal delivered by the non-linear regression model. In an alternative embodiment, the module 14 is configured to directly analyze the cardio-respiratory signals by applying a segmentation model.
[0077] The device 1 interacts with a user interface 16, through which information can be entered and retrieved by a user. The user interface 16 comprises any suitable means for entering or retrieving data, information or instructions, including visual, tactile and / or audible capabilities which may include one or more of the following means well known to those skilled in the art: a screen, a keyboard, a trackball, a touch pad, a touch screen, a loudspeaker, a voice recognition system.
[0078] In its automatic actions, the device 1 can for example execute the following method ([Fig.7]): - receive at least one cardio-respiratory signal 21 representative of the sounds and / or movements coming from the rib cage and the airways of the subject (step 41); - the use of a non-linear regression model of the neural network type to obtain, from the at least one cardio-respiratory signal 21, at least one estimate of a physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal 31 (step 42); - providing as output at least one estimate of a physiological ventilatory polygraphy signal and / or a physiological electrocardiography signal 31 (step 43).
[0079] Further information will be given below on the device 6 for training. The device 6 may be integrated into the same apparatus as the estimation device 1 and share most of the functionalities of the device 1. More generally, the device 1 and the device 6 may have common processors configured to execute part of the functionalities of the device 1, whether localized or distributed. In alternative implementations, the device 1 and the device 6 may be completely disconnected, and possibly operated by separate actors (e.g., a service provider for the device 6 and clinic staff for the device 1).
[0080] In advantageous embodiments, the device 6 shares the same architecture with the device 1, as well as the same preprocessing of the data. For example, the preliminary normalization of the cardio-respiratory signals are identical. This is not necessary, however, and the device 6 may have an architecture different to some extent, in that it is able to provide model parameters suitable for correct estimations by device 1.
[0081] Similarly, the learning parameters 20 used by the device 1 may come from sources other than the device 6.
[0082] As shown in [Fig.8], the device 6 is adapted to recover initial learning parameters 721 defining an initial algorithmic model within the framework of a “machine learning” architecture, from one or more local or remote database(s) 60. The latter may take the form of storage resources available from any type of suitable storage means, which may in particular be a RAM or an EEPROM such as a Flash memory, possibly within an SSD. In advantageous embodiments, the initial training parameters 721 have been previously generated by the device 6 for training. Alternatively, the initial training parameters 721 are received from a communication network.
[0083] The device 6 is configured to determine modified learning parameters 722 defining a modified algorithmic model based on the input learning parameters 721, and to transmit these parameters 722 or record them, for example in the database(s) 60. This is achieved by using learning data sets 71 received by the device 6, allowing the learning parameters to be progressively adjusted so that an associated algorithmic model becomes more relevant to these learning data sets 71 (supervised learning).
[0084] More specifically, the training data sets 71 comprise input data including cardio-respiratory signals 711, and output data including the respiratory and cardiac physiological signals 712 corresponding respectively to the inputs, which were measured simultaneously. Furthermore, the device 6 is configured to iteratively produce results of the physiological signals evaluated from the input data 711, on the basis of training parameters defining a current algorithmic model, and to modify these training parameters so as to reduce the deviations between the evaluated respiratory physiological signals and the known respiratory physiological signals 712.
[0085] The modified learning parameters 722 can themselves be taken as initial learning parameters for subsequent learning operations.
[0086] Although the training aspects of the device 6 are currently being developed, it can be observed that the device 6 can be applied in the same manner to validation data sets and test data sets.
[0087] The learning device 6 comprises an execution stage 60 comprising a module 61 for receiving the learning data sets 71 and the learning parameters 721, a module 62 for preprocessing the cardio-respiratory signals 711 of the learning data sets 71, a module 63 for using the regression model to calculate the physiological signals and to output them. These modules 61, 62 and 63 will not be detailed here, because their description is similar to the above description of the respective modules 11, 12 and 13 of the prediction device 1. The machine learning-based models in the modules 62 and / or 63 are defined by the current values of the learning parameters.
[0088] In advantageous implementations, the modules 61, 62 and 63 of the device 6 are at least partially identical to the modules 11, 12 and 13 of the device 1. In particular, the same image normalization can be applied (modules 12 and 62) and / or the same neural network architecture can be used (modules 13 and 63). In particular, the execution step 60 can be an entity common to the device 1 and to the device 6. By "same network architecture" and "same methods" is meant subject to the possible exploitation of distinct learning parameters.
[0089] The device 6 further comprises a module 65 for comparing the estimates of the evaluated physiological signals with the known physiological signals 712 and for modifying the current training parameters in order to reduce the gap between them. The terms "comparisons" and "modifications" are used to simplify the expression, whereas in practice, the two operations can be closely combined in synthesis actions, for example by gradient descent.
[0090] The device 6 also comprises a module 66 for stopping the modifications of the learning parameters on the basis of the data sets 71. This module 66 can provide for stopping the associated iterations when the deviations between the evaluated physiological signals and the known physiological signals 712 become lower than a predefined threshold. It can, on the contrary, or cumulatively, provide for the end of these iterations when a predefined maximum number of iterations is reached.
[0091] In a basic learning version, the device 6 is configured to carry out learning exclusively directed towards the neural network with U-shaped architecture. In other modes, the device 6 comprises additional machine learning stages, which may in particular be intended to compress the sound signals 711. Then, the device 6 may also be configured to train these additional machine learning stages. In particular embodiments, the device 6 is configured to jointly train some or all of the additional learning stages with the neural network with U-shaped architecture of the module 63.
[0092] In implementations involving neural networks, the device 6 may be adapted to perform back-propagation through time (BPTT). According to this learning method, a gradient-based technique is applied at a cost with respect to all network parameters, as known to those skilled in the art. Stochastic gradient descent may also be used (i.e., batch gradient descent or mini-batch gradient descent).
[0093] For example, with the implementation described above involving the second compression neural network and the U-architecture neural network, both architectures can be trained jointly. At least one of the residual neural networks (ResNet) can be used, with transfer learning from a pre-trained model that used, for example, the known database. The associated parameters can be, for example, the following: layer_name: 18 epochs: 50 optimizer: Adam exponential decay for first moment estimation = 0.9 exponential decay for second moment estimation = 0.999 batch_dimension: 32 sound_signal_shape: (1, 81920) inertial_unit_signal_shape: (1, 5984)
[0094] The device 6 also comprises a module 67 making it possible to provide the modified training parameters 722 resulting from the training operations by the device 6. It further interacts with a user interface 68, for the description of which the reader may refer to the user interface 20 associated with the device 1.
[0095] The device 6 can for example execute the following process: 1) receiving the initial training parameters 721 and one or more of the training data sets 71, or a portion thereof, including inputs comprising the cardio-respiratory signals 711 and outputs comprising the known results of physiological signals 712; 2) preprocessing cardio-respiratory signals 711 for more efficient and / or more reliable processing; 3) run the regression model on the 711 cardiorespiratory signals after their preprocessing; 4) compare the estimated physiological signals to known physiological signals 712 and modify the current training parameters accordingly; 5) check the stopping criteria; 6) if the stopping criteria are not satisfied, return iteratively to step 3) with the modified values of the current learning parameters, and repeat the operations; 7) if the stopping criteria are met, keep the currently modified learning parameters as modified learning parameters 722 and provide them.
[0096] In alternative execution modes, the modified learning parameters 722 may be distinct from the currently modified learning parameters when the stopping criteria are met, and be calculated from these currently modified learning parameters as well as one or more previously obtained parameter values (e.g., by averaging).
[0097] A particular device 9, visible in [Fig.9], implements the device 1 as well as the device 6 described above. It corresponds for example to a workstation, a laptop, a tablet, a smartphone, or even a “head-mounted display” (HMD).
[0098] This apparatus 9 is suitable for estimating physiological signals and for training the corresponding model. It comprises the following elements, connected to each other by an address and data bus 95 which also carries a clock signal: - a microprocessor 91 (or CPU); - a graphics card 92 comprising several graphics processing units (or GPUs) 920 and a graphics random access memory (GRAM) 921; - a non-volatile memory of type ROM 96; - 97 RAM; - one or more input / output (I / O) devices 94 such as a keyboard, a mouse, a trackball, a webcam; other modes of entering commands such as voice recognition are also possible; - a power source 98; and - a radio frequency unit 99.
[0099] According to a variant, the power supply 98 is external to the device 9.
[0100] The apparatus 9 also comprises a screen-type display device 93 display device directly connected to the graphics card 92 to display synthesized images calculated and composed in the graphics card. The use of a dedicated bus to connect the display device 93 to the graphics card 92 offers the advantage of having much higher data transmission rates and therefore reducing the latency time for displaying the images composed by the graphics card.
[0101] According to a variant, a display device is external to the apparatus 9 and is connected to it by a cable or wirelessly to transmit the display signals. The apparatus 9, for example through the graphics card 92, comprises a transmission or connection interface adapted to transmit a display signal to a external display means such as for example an LCD or plasma screen or a video projector. In this regard, the radio frequency unit 99 can be used for wireless transmissions.
[0102] It should be noted that the word "register" used below in the description of the memories 97 and 921 can designate in each of the memories mentioned, a memory area of low capacity (some binary data) as well as a memory area of high capacity (making it possible to store an entire program or to calculate or display all or part of the data representative of the data). Similarly, the registers shown for the RAM 97 and the GRAM 921 can be arranged and constituted in any way, and each of them does not necessarily correspond to adjacent memory locations and can be distributed differently (which notably covers the case where a register comprises several smaller registers).
[0103] When powered up, the microprocessor 91 loads and executes the instructions of the program contained in the RAM 97.
[0104] As will be understood by those skilled in the art, the presence of the graphics card 92 is not mandatory, and can be replaced by complete processing by the central processing unit and / or simpler visualization implementations.
[0105] In alternative modes, the device 9 may comprise only the functionalities of the device 1, and not the learning capabilities of the device 6. Furthermore, the device 1 and / or the device 6 may be implemented differently than standalone software, and a device or set of devices comprising only parts of the device 9 may be operated via an API call or a cloud interface. EXAMPLES
[0106] The present invention will be better understood by reading the following examples which illustrate the invention in a non-limiting manner.
[0107] Example 1:
[0108] Materials and methods
[0109] 18 volunteers recruited for sleep apnea screening were used for this example.
[0110] For model training, the recordings were first shuffled. Then, since the recordings do not have the same durations (from ~4 to 12 hours), the 18 recordings were divided into three sets: a training dataset, a validation dataset, and a test dataset with an approximate ratio (33%,33%,33%), so that they have the same overall duration.
[0111] The device of the invention is used to receive several recordings. Each recording comprises 6 cardio-respiratory movement signals acquired by an inertial unit (i.e., signals obtained by a 3-axis accelerometer and a 3-axis gyroscope), sampled at 100 Hz and measured by the smartphone. Before any further processing, these recordings are manually synchronized with two signals from the Nox RIP belt: the thoracic and abdominal respiratory efforts. These signals were recorded by the device of the invention, and HSAT (Nox T3) or PSG (Nox Al) for one night for each volunteer.
[0112] The 6 cardio-respiratory movement signals obtained from the phone are used as inputs for two neural network-based models. Both models are trained to output a thoracic respiratory effort signal acquired via a chest belt for one, and an abdominal respiratory effort signal acquired via an abdominal belt for the other. Thus, the outputs always contain a single-channel signal.
[0113] The inputs (i.e., cardiorespiratory signals) and outputs (i.e., physiological signals) are normalized using sliding window z-score normalization. In this example, the sliding window has a duration of one minute.
[0114] Inputs and outputs are bandpass filtered between [0.05; 5] Hz (breathing frequencies) using an 8th order forward-backward Butterworth filter. Inputs and outputs are decimated to 20 Hz and are divided into 1 minute windows.
[0115] The neural network architecture for regression is a U-Net based on ResNetl8, with a final linear activation layer. The cost function is of type logcosh.
[0116] 10 training epochs were used. For regularization purposes, this number of epochs can be reduced by early stopping which is set when the validation loss does not decrease for at least 5 epochs.
[0117] Results
[0118] The loss functions for training and validation decrease rapidly. The error, for this example, is no more than 22% (in terms of logcosh). This result can be expected to improve with the use of data from a larger number of volunteers.
[0119] [Fig. 10] represents an example of a one-minute physiological signal estimated by the device of the invention for a volunteer, where the model obtains the worst predictions. An 11-second apnea occurs in the middle of this window. Even if the prediction is inaccurate just after the apnea, the general dynamics of the signal is well predicted.
[0120] Figures 11 and 12 show results obtained for one patient. The first and second signals at the top of [Fig. 11] respectively show the signal of the thoracic respiratory effort and the abdominal respiratory effort signal estimated by the method of the invention. The last line shows a superposition of these signals with, respectively, the signals acquired directly by a thoracic belt and an abdominal belt synchronously with the cardiorespiratory signals used for the estimation. Similarly, the first and second signals at the top of [Fig. 12] show respectively a nasal flow signal and a nasal pressure signal estimated by the method of the invention. The last line shows a superposition of these signals with, respectively, the signals acquired directly by a nasal cannula synchronously with the cardiorespiratory signals used for the estimation.
[0121] Considering the results obtained for all volunteers, the error rate is 20%. A larger database will facilitate generalization and improve the results.
Claims
Claims
1. A computer-implemented method for estimating physiological signals of a subject, said method comprising: - receiving (41) at least one cardio-respiratory signal (21) representative of sounds and / or movements originating from the rib cage and airways of the subject, the cardio-respiratory signal (21) being acquired during the subject's sleep; - using a non-linear regression model of the neural network type (43) configured and trained to receive, as input, the at least one cardio-respiratory signal, and to provide as output at least one estimate of a physiological ventilatory polygraphy signal (31);wherein said physiological ventilatory polygraphy signal is selected from: a nasal flow rate acquired via a nasal cannula, a nasal pressure acquired via a nasal cannula, a thoracic respiratory effort acquired via a thoracic belt or an abdominal respiratory effort acquired via an abdominal belt; wherein the neural network comprises an encoder having one input and several outputs and a decoder having one output and several inputs, wherein at least two encoder outputs are directly connected to at least two respective inputs of the decoder, and at least two outputs of the encoder are connected to each other and wherein said neural network is configured to perform one-dimensional convolution operations; and - provide as output the at least one estimate of a physiological ventilatory polygraphy signal (31).;
2. The method of claim 1, wherein the neural network comprises a U-Net whose encoder comprises a ResNet.
3. A method according to claim 1 or 2, wherein the non-linear regression model of the neural network type (43) is configured and trained to output at least one estimate of a physiological electrocardiography signal (31).
4. The method of claim 3, wherein said electrocardiography physiological signal is acquired through an electrocardiogram of at least one lead.
5. A method according to any one of claims 1 to 4, wherein the at least one cardiorespiratory signal comprises a cardiorespiratory movement signal obtained using a gyroscope and / or an accelerometer.
6. A method according to any one of claims 1 to 5, wherein the at least one cardio-respiratory signal comprises a cardio-respiratory sound signal; the method further comprising compressing the cardio-respiratory sound signal using a second convolutional neural network, prior to its application as input to the non-linear regression model.
7. The method of claims 5 and 6, further comprising merging the compressed cardiorespiratory sound signal with the cardiorespiratory movement signal.
8. A method according to any one of claims 1 to 7, wherein the at least one cardio-respiratory signal is normalized before using the non-linear regression model.
9. A method according to any one of claims 1 to 8, wherein the neural network comprises an output layer having a linear activation function.
10. Device (1) for estimating physiological signals of a subject, said device (1) comprising: - at least one input configured to receive at least one cardio-respiratory signal representative of sounds and / or movements coming from the rib cage and the airways of the subject (21), the cardio-respiratory signal (21) being acquired during the sleep of the subject; - at least one processor configured to: estimate at least one physiological signal of ventilatory polygraphy (31) using a non-linear regression model of the neural network type configured and trained to receive, as input, the at least one cardio-respiratory signal, and provide as output the at least one estimation of a physiological signal of ventilatory polygraphy;wherein said physiological ventilatory polygraphy signal is selected from: a nasal flow rate acquired via a nasal cannula, a nasal pressure acquired via a nasal cannula, a thoracic respiratory effort acquired via a thoracic belt or an abdominal respiratory effort acquired via an abdominal belt; wherein the neural network comprises an encoder having one input and several outputs and a decoder having one output and several inputs, wherein at least two encoder outputs are directly connected to at least two respective inputs of the decoder, and at least two outputs of the encoder are connected to each other and wherein said neural network is configured to perform one-dimensional convolution operations; - at least one output configured to provide the at least one ventilatory polygraphy signal (31).
11. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to implement the method for estimating physiological signals of a subject according to any one of claims 1 to Q.
12. id y. A computer-readable recording medium comprising instructions which, when the program is executed by a computer, cause the computer to implement the method for estimating physiological signals of a subject according to any one of claims 1 to 9.