A method and arrangement for automated speech-based health analysis

The bio-marker adaptation module addresses signal degradation issues by aligning feature representations with a reference path, enhancing health analysis accuracy and robustness in real-world conditions.

WO2025229258A1PCT designated stage Publication Date: 2025-11-06AALTO UNIV FOUND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/FI2025/050213
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-29
Filing Date
2025-04-29
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing speech-based health monitoring systems face challenges in accurately analyzing health conditions due to signal degradation caused by low-bitrate speech codecs, network conditions, and environmental noise, leading to performance reduction and mismatch issues.

Method used

A bio-marker adaptation module that processes speech signals to cancel distortions from background noise, recording equipment, transmission errors, and codec effects, aligning feature representations with a reference signal path using machine learning models trained on simulated distortions.

Benefits of technology

Enables robust health analysis in realistic environments by maintaining analysis performance, improving generalizability and accuracy by compensating for environmental and transmission-related distortions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FI2025050213_06112025_PF_FP_ABST
    Figure FI2025050213_06112025_PF_FP_ABST
Patent Text Reader

Abstract

A method for speech-based health analysis, comprising: receiving via a digital transmission channel an input signal related to a recorded speech signal, the input signal comprising at least one feature related to health information; processing the input signal by adapting the input signal for at least partly cancelling at least one effect on a signal path, wherein the signal path comprises path from speech signal output of a person to a device recording the speech signal and / or the digital transmission channel, wherein the at least one effect on the signal path comprises distortions caused to the input signal by at least one of the following: background noise, acoustic environment, recording equipment characteristics, transmission errors and / or error concealments, network handovers, codec, speech enhancement; and providing the processed signal to a health analysis, e.g. to a health analysis system; and an arrangement and a system thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A METHOD AND ARRANGEMENT FOR AUTOMATED SPEECH-BASED HEALTH ANALYSIS

[0002] TECHNICAL FIELD

[0003] The invention is related to a method and arrangement for health-analysis, the health analysis not limited to any particular pathology type. Specifically, the invention relates to a method and arrangement for automatic analysis of biomarkers of medical pathologies from speech signals, especially the invention claimed in the independent claims.

[0004] BACKGROUND

[0005] Speech analysis has emerged as a promising non-invasive tool in identifying pathologies. Speech production is closely linked to the physiological state, and changes in voice quality have been associated with various pathologies, including heart failure, Parkinson’s disease, Alzheimer’s disease, strokes, COVID-19, asthma, sleep apnea, multiple sclerosis (MS), amyotrophic lateral sclerosis (ALS), cerebral palsy, and various functional and organic diseases in laryngeal area including cancer, laryngitis, edema, nodules, polyps, ulcers, and lesions. Various mental disorders have been examined as well, including depression, post-traumatic stress disorder (PTSD), schizophrenia, anxiety, bipolar, bulimia, anorexia, and obsessive-compulsive disorder (OCD). Importantly, the automated detection of pathological markers is usually carried out in a data-driven manner by utilizing various methods of machine learning. Therefore, a typical approach consist of separate processes for training the detector by using a training data set, and for using the trained model for inference with new unseen data.

[0006] The widespread availability of smartphones, each equipped with microphones, provides a user-friendly platform for speech-based health monitoring. However, a crucial aspect of speech-based health monitoring is the ability to conduct the analysis in realistic practical environments, which often comprise speech transmission through bandwidth-constrained digital telephony channels. Significant challenges, however, arise when the input speech to a health monitoring system is transmitted over digital transmission channels. In speech transmission, the signal is always encoded using a low-bitrate speech codec. Encoding of speech results in two major degradations: reduction of speech bandwidth and generation of quantization noise. In speech encoding, transmission efficiency can be changed by varying the bit rate of the codec. Utilizing codecs that compress speech at a low bit rate can markedly alter the original signal, thereby affecting elements such as spectral characteristics, prosodic features, and formant information. These aspects are crucial as they can carry indicators related to health conditions. Consequently, differentiating between disordered and healthy voices can become more challenging. Furthermore, the codecs and bitrates used are not necessarily constant during a phone call, due to changing network conditions, including network handovers. This variation can easily lead to a codec mismatch between the data used for training a disorder detection model and the data ultimately used as input for detection in a realistic environment. Such a mismatch can also significantly reduce the performance of the model. Also, various other causes can distort the health-related information and cause mismatches in the realistic usageenvironment, including background noise, acoustic environment, recording device, transmission errors and general speech enhancements.

[0007] Thus, there is a need to conduct speech-based health analysis under such varying conditions, without degradation in the analysis performance, for example, detection accuracy of a symptom or disease.

[0008] SUMMARY

[0009] The following presents a simplified summary in order to provide basic understanding of some aspects of various invention embodiments. The summary is not an extensive overview of the invention. It is neither intended to identify key or critical elements of the invention nor to delineate the scope of the invention. The following summary merely presents some concepts of the invention in a simplified form as a prelude to a more detailed description of exemplifying embodiments of the invention.

[0010] According to the first aspect of the invention a method for speech-based health analysis is presented. The method comprising: receiving via a digital transmission channel an input signal related to a recorded speech signal, the input signal comprising at least one feature related to health information; processing the input signal by adapting the input signal for at least partly cancelling at least one effect on a signal path, wherein the signal path comprises path from speech signal output of a person to a device recording the speech signal and / or the digital transmission channel, wherein the at least one effect on the signal path comprises distortions caused to the input signal by at least one of the following: background noise, acoustic environment, recording equipment characteristics, transmission errors and / or error concealments, network handovers, codec, speech enhancement; and providing the processed signal to a health analysis, e.g. to a health analysis system, wherein the processed signal comprises a feature representation of the speech signal, the feature representation of the speech signal comprising the at least one feature related to health information adapted for the at least one effect on the signal path.

[0011] The at least partly cancellation of the at least one effect may comprise modification of the input signal so that the at least one feature related to the health information corresponds to a similar at least one feature related to the health information transmitted on a reference signal path.

[0012] The input signal related to the speech signal may comprise the speech signal and / or a feature representation of the speech signal.

[0013] The digital transmission channel may comprise a digital telephony channel and / or digital communication channel.

[0014] The processing may comprise using a machine learning model.

[0015] The machine learning model may be trained using data related to the at least one effect on the signal path.

[0016] The data related to the at least one effect on the signal path may be simulated using a channel simulator.

[0017] The data used for training the machine learning model may not include health related information used in the training.

[0018] The method may further comprise performing health analysis using the processed signal, e.g. in a health analysis system; and providing a health analysis result based at least in part on the performed health analysis.

[0019] The health analysis may comprise using a machine learning model for health analysis, the machine learning model for health analysis being trained using health related data.

[0020] The health analysis may be implemented longitudinally, e.g. over a period of time, utilizing one or multiple feature representations extracted from previous recording of the same person. According to the second aspect of the invention, an arrangement for a healthanalysis system is presented. The arrangement comprises at least one processing unit, such as a computer, wherein the arrangement is configured: to receive via a digital transmission channel an input signal related to a recorded speech signal, the input signal comprising at least one feature related to health information; to process the input signal by adapting for at least partly cancelling at least one effect on a signal path, wherein the signal path comprises path from speech signal output of a person to a device recording the speech signal and / or the digital transmission channel, wherein the at least one effect on the signal path comprises distortions caused to the input signal by at least one of the following: background noise, acoustic environment, recording equipment characteristics, transmission errors and / or error concealments, network handovers, codec, speech enhancement; and to provide the processed signal for health analysis, e.g. to a health analysis system, wherein the processed signal comprises a feature representation of the speech signal, the feature representation of the speech signal comprising the at least one feature related to health information adapted for the at least one effect on the signal path.

[0021] The arrangement may be configured to perform the method presented above.

[0022] According to third aspect of the invention a system for health analysis is presented, wherein the system comprises the arrangement presented above.

[0023] According to fourth aspect of the invention a computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method presented above, is presented.

[0024] According to fifth aspect of the invention a computer-readable medium comprising the computer program described above is presented.

[0025] Various exemplifying and non-limiting embodiments of the invention both as to constructions and to methods of operation, together with additional objects and advantages thereof, will be best understood from the following description of specific exemplifying and non-limiting embodiments when read in connection with the accompanying drawings.

[0026] The verbs “to comprise” and “to include” are used in this document as open limitations that neither exclude nor require the existence of unrecited features. The features recited in dependent claims are mutually freely combinable unless otherwise explicitly stated. Furthermore, it is to be understood that the use of “a” or “an”, i.e. a singular form, throughout this document does not exclude a plurality.

[0027] BRIEF DESCRIPTION OF FIGURES

[0028] The embodiments of the invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0029] Figure 1 presents schematically an example of an illustration of the parts of the system according to one embodiment of the invention.

[0030] Figure 2 illustrates schematically an example embodiment of the invention of different processes.

[0031] Figure 3 presents an example embodiment of the system within a digital telephony channel.

[0032] Figure 4 illustrates an example diagram of a biomarker adaptation network according to an embodiment of the invention.

[0033] Figure 5 illustrates a features modulator according to an example embodiment of the invention.

[0034] DESCRIPTION OF EXEMPLIFYING EMBODIMENTS

[0035] Herein, the technical details of a system that can be used for speech-based health-analysis in realistic environment of digital telephony channels is described. The solution can consist of a speech-based health-analysis system, together with a bio-marker adaptation module that conducts the cancellation of effects, e.g. all effects, the realistic environment has caused to the health-related information in the feature space of the health-analysis system.

[0036] Figure 1 depicts an example embodiment of a bio-marker adaptation network 105, BAN. The bio-marker adaptation network inputs a signal, the input signal 102, that may be a speech signal or a feature representation of the speech signal. The speech signal may be processed before a feature representation 104 of it is acquired. The speech signal may also be affected by different kind of effects before inputting the speech signal into the bio-marker adaptation network or extraction of the features 103. For example, the signal may be normalized before feature extraction 102. The speech signal may be travelling on a signal path that does not correspond to a reference signal path. The reference signal path is a signal path for a speech signal that different health analysis algorithms are trained to be used with.

[0037] The feature representation of the speech signal comprises features related to health information. The health information may be related to the pathologies described in the background. However, the invention is not related to only feature representation of the pathologies disclosed herein, but to feature representation of any pathology possible to distinguish from human speech.

[0038] The input signal may be distorted on the signal path due to different reasons. Some of the reasons comprise background noise, acoustic environment, recording equipment characteristics, transmission errors, transmission error concealments, network handovers, codec, and / or speech enhancement.

[0039] The background noise may be considered also as the ambient noise in the speech signal . The background noise may also comprise other type of analogue or digital noise not originally comprised in the speech.

[0040] The acoustic environment comprises different acoustical effects caused by the recording environment. One advantage of the invention is that the effect caused due to the acoustic environment may be adapted, e.g. cancelled, in a manner that the recording environment may be acoustically inferior to acoustic quiet room.

[0041] The characteristics of the recording equipment comprise the type and specifications of the recording equipment. The type of the recording equipment may be, for example, a phone, a microphone, a soundcard, or a computer, but is not limited to those. Also, the usage of the recording equipment may be comprised in the characteristics. For example, the position or orientation of the recording equipment relative to the human speech output, or the recording equipment settings may be taken into account.

[0042] Effect of the different speech or audio codecs or encoding are comprised in the codec. Different codecs may be AMR, Opus, MP3. However, this is not exhaustive list, and other codec are not excluded from the scope of the invention. In addition, the speech enhancement may comprise different audio enhancements performed to the input signal, such quality enhancements and / or noise suppressions.

[0043] The BAN adapts the input signal for at least partly cancelling at least one effect caused by the distortions due to the disclosed reasons. The adapting of the signal is done for compensating the at least one effect. By compensating, e.g. partly cancelling, for the at least one effect on the input signal related to a recorded speech signal, the input signal is adapted in a manner that modifies the input signal to correspond to a signal transmitted on a reference signal path.

[0044] The adaptation of the input signal to correspond to the signal on the reference signal path may correct the signal by at least partly cancelling the effects. It is also possible that after at least partly cancelling the effects, the reference signal path is simulated on the adapted input signal for achieving an input signal that corresponds to the signal transmitted on a reference signal path.

[0045] In the adaptation the information transmitted on the signal path may be converted to resemble the configurations of a reference signal path. The adaptation may be performed so that the health-related information in the speech signal is acquired so that it corresponds to a health-related information transmitted via a reference signal path.

[0046] The adaption of the input signal is performed to compensate for at least one effect. However, a choice of using at least two, three or more from the list is possible.

[0047] The BAN gives out a processed, e.g. adapted, feature representation 106 that may be used for health analysis 108. The processed feature representation may be processed further before the health analysis, e.g. the features may be normalized 107. The output of the health analysis may be a health analysis result 109.

[0048] According one example embodiment of the invention, the processing in BAN is performed using a machine learning model. The machine learning model may be trained using originally high-quality signals that are simulated to comprise at least one of the effects on one or multiple different signal paths. The simulation may be performed using a separate channel simulator. The machine learning model may also perform estimation of the configurations, e.g. at least one configuration, within the signal path. The configurations may also be acquired with other methods than with machine learning model.

[0049] The estimation of the configurations for the reference signal path may also performed using a machine learning model. The machine learning model for reference signal path may be same or different to the machine learning model used for the signal path estimation. The machine learning model used for the signal path estimation may take one, two, three, or more reference signals transmitted via the reference signal path as input.

[0050] According to a further example embodiment of the invention, the data used for training of the machine learning model does not comprise health related information that is used in the training. The data may comprise data points that are collected from people that have had bio-markers related to one or multiple pathologies in their speech signal, but according to the invention the health- related information is not used in the training. Thus, the adaptation performed with the BAN should not adapt or compensate for health-related information, conserving the health-related features as they are in the original speech signal.

[0051] According to an example embodiment of the invention, the health analysis may be performed using a separate machine learning model, i.e. a machine learning model for health analysis. The machine learning model for the health analysis may also be considered as a health analysis algorithm. The training performed of the machine learning model for health analysis may use health related information.

[0052] In Figure 2, there is presented an example embodiment, combining both the training of the BAN 203 and training of the health-analysis algorithm 201 , of the utilization of the proposed system within training and inference environments of a data-driven health-analysis frameworks, which involve distinct training and inference processes. During training, features are extracted, and the detection classifier is trained with the original high-quality signals. In the inference phase, the feature representation of the input signal, that was previously transmitted through a transmission channel, is adjusted by the bio-marker adaptation module, which processes both the signal and its extracted feature representation. The bio-marker adaptation module is also based on machinelearning models and undergoes a separate training procedure utilizing a channel simulator that generates realistically distorted versions of the training signals.

[0053] The inference process 202 may correspond to the steps presented in Fig. 1 .

[0054] In Figure 3, an example embodiment according to the invention relating to the recording, i.e. capturing, of the speech and an example of the transmission of the speech signal through one example signal path is presented. The effect on the signal path may comprise distortions caused by a capturing device 301 , e.g. a phone, effects due to transmission over network 302, and / or receiving device 303. Each of these distortion sources may comprise one or multiple different sources for the distortions. Possible sources of distortion in capturing device may be sound capture, analog-to-digital conversion, encoding, dynamic compression, and / or packetization. Possible sources of distortion in receiving device may be depacketization, buffering, decoding, and / or enhancement. After receiving the speech signal with a receiving phone, it may be transformed to a voice signal replayed to a recipient 304. The speech signal may also be, instead of directly processed in the system 306, especially in the BAN, routed via voicemail server 305, that may cause further distortions, having sources such as optional compression and / or storage, to the speech signal. The system may output the analysis result 307, similarly as depicted in Figs. 1 and 2.

[0055] An example embodiment of the implementation of the BAN illustrated in Figure 4 comprises a hidden representation extractor 408 and a feature modulator 409. A signal 401 , that may be a 16 kHz signal, may be trimmed 402 and normalized for volume 403 giving an input signal 404 for the hidden representation extractor 408. Features may also be extracted 405 of the trimmed signal, as well as Z- score normalized 406 the feature representation giving an input MFCC matrix 407 for the hidden representation extractor 408. The processing involves each input signal and its MFCC representation, and after Z-score denormalization 410 resulting in a modified version of the MFCC matrix as the output 411 .

[0056] An example embodiment the implementation of the feature modulator presented in Figure 5 comprises a fully-connected network (FCN) for frame interpolation 502, followed by a concatenation step 503, and another FCN for dimensionality reduction 504. A signal embedding 501 may be used as the input for the FCN for frame interpolator 502. An input MFCC matrix 503 may be input in the concatenation step 504 as well. The output of the FCN for the dimensionality reduction 505 may be an output MFCC matrix 506. An overview of an example embodiment of the invention is presented below.

[0057] The invention is a computerized system that is capable of conducting healthanalysis from a person’s voice in a realistic scenario where the person is speaking on any type of phone in an arbitrary environment, and the speech is transmitted through digital telephony channels prior to the analysis.

[0058] Such a realistic scenario implies that the characteristics of the transmission channel, phone, and the recording environment are unknown to the system, and all of them may vary over the duration of the transformed signal.

[0059] The invention consists of an automated speech-based health analysis system, including the extraction of features containing the health-related information, and conducting the health analysis based on the features. Optionally, health-analysis can be implemented longitudinally, utilizing also the feature representations that were extracted from any number of previous speech recordings from the same speaker.

[0060] The inventive step is the utilization of a biomarker-adaptation module, that modulates the feature representation to correspond to a scenario where the same source speech would have been recorded in, and not further transmitted from, an ideal noise-controlled laboratory environment with high bitrate by using professional recording equipment. The adaptation is conducted within the feature space used for the health-analysis, and the output is an adapted version of the feature representation.

[0061] The computerized system is integrated to the device that receives the speech after the transmission through telephony channels. In particular, after the speech is decoded back to a waveform, it is given as an input to the computerized system, by using any lossless means of transmission. The speaker is informed about the output of the system by utilizing any intermediate systems and processes.

[0062] Some benefits the invention, e.g. the example embodiments of the invention presented herein, will have over the background of the invention are presented below.

[0063] The bio-marker adaptation module ensures that the health-related information is consistently extracted from the speech, regardless of the environment, equipment, or channel -related configurations, making the process robust for dynamic changes that occur within a single phone-call and between phone-calls. Therefore, it allows speech-based health analysis, either longitudinal or cross- sectional, to be conducted in a realistic scenario of digital telephony, not requiring a strictly controlled laboratory environment or even consistent configuration within or between phone-calls.

[0064] Importantly, the module explicitly aims to modulate the extracted features, such that they align with the conditions where the model was trained. In other words, it cancels all distortions that were caused by the realistic usage environment to the health-related features. This is different to the existing approaches, which have simply included distorted data in the training corpus of the health-analysis algorithm.

[0065] There is one significant technical benefit to our modular approach: improved health analysis performance in realistic usage environments. This advantage stems from two related reasons. Firstly, dividing the task of health analysis in realistic environments into simpler sub-problems, health analysis and domain adaptation, reduces the overall complexity that the data-driven model needs to capture, thereby increasing the generalizability of the model. Secondly, this approach enables the feature distortion module to leverage non-health-related training datasets, which are typically much more extensive than the available health-related data.

[0066] An example embodiment of the system according to the invention is presented herein. First, the system components are presented and the positioning in a realistic speech transmission environment is described. The subsections provide implementation details for the system components.

[0067] The parts of the health-analysis system are shown in Figure 1. Signal normalization module (2) conducts operations to prepare the input signal (1 ) for the subsequent operations, for example, by scaling between values -1 and 1 to normalize the volume between different samples. Feature extraction module (3) extracts health-related features (4), typically feature matrices, from the normalized signal. Bio-marker adaptation module (5) is responsible for adapting the health-related information carried by the feature representation of the signal to correspond to a scenario in which the same source signal would have been processed in a laboratory environment. The adaptation considers the effects of the transmission channel between the two endpoint devices, including, for example, transmission errors, speech encodings, network handovers, and speech enhancements that are done by the endpoint devices. Also, it considers the effects of hardware and software equipment used to record the audio and perform analog to digital conversion, as well as the physical environment n of the recording device, as it can affect the health-related information in the feature representation in various ways, for example, due to varying acoustic properties and background noise in the environment. The bio-marker adaptation module outputs the modulated feature representation (6), which is then processed by a feature normalization module (7) that prepares the feature representation to be used with the health-analysis module, for example, by conducting z-score normalization. The normalized features are then analyzed by a health analysis module (8), for example, a machine learning classifier or regressor, that performs a health-related task based on the feature representation and outputs the analysis result (9), which can be, for example, a class index or a continuous output value. In the case of longitudinal analysis, the inverted feature representations can be stored to a database, from which some number of previous entries of the same speaker are obtained for the health-analysis.

[0068] The health-analysis module, as well as the bio-marker adaptation module, are based on a data-driven approach which involves training machine learning models by utilizing a training corpus. Figure 2 provides an illustration of the system’s usage in the training and inference environments of typical data-driven health-analysis systems, which consist of separate processes for training and inference. The approach includes a distinct training process for the models that are involved in the health-analysis, from those models that are used as a part of the bio-marker adaptation module. An important part of the system is therefore a channel simulator, that is utilized in the training process of the bio-marker adaptation models, to generate realistically distorted counterparts for the nontransmitted signals in the training corpus.

[0069] Figure 3 demonstrates the positioning of the claimed system during inference in a realistic speech transmission environment. The transmission is based on digital telephony channels employing any technologies (2G, 3G, etc.), including channels that utilize either circuit-switched or packet-switched technologies. Only the processing steps that are the most relevant for the illustration of the invention are shown. The input to the algorithm is transmitted from the speech processing pipelines by utilizing any transmission method, and the analysis result is informed to the speaker by using any system or process.

[0070] Some example embodiments according to the invention relating to the healthanalysis system are presented below.

[0071] The exact details of the health-analysis system are not important for the functioning of the proposed system, and they can consist of any methods, as long as the feature extraction step outputs numerical representations of the signal, for example, x e Rn, where n e N+, and x is the representation.

[0072] For example, signal normalization can include simple scaling between values -1 and 1 to normalize the volume between different samples. The features could be Mel Frequency Cepstral Coefficients (MFCCs), including the first 13 coefficients, with the exclusion of the Oth coefficient. Additionally, the delta and delta-delta coefficients can be included, resulting in 39-dimensional feature vectors. The extraction may utilize 128 mel filter banks, 2048 FFT-bins, and Hamming window function. The frame length can be set to 25 ms with a 10 ms shift between frames. Feature normalization can be based on z-score normalization. Assuming that a support vector machine (SVM) classifier is used, the MFCC matrices can be flattened for input, by computing a set of statistics over the MFCC frames. These stats could include the mean, standard deviation, minimum, maximum, median, kurtosis, skewness, and interquartile range. In the remaining sections of this document, these are the expected configurations.

[0073] Transmission channel simulator according to an example embodiment of the invention is presented herein.

[0074] This subsection describes the implementation details of the transmission channel simulator, which is utilized in the training of the bio-marker adaptation models. It aims to add realistic distortions to the signal, such that would occur in real usage scenario. It includes the addition of artificial background noise, artificial modeling of various types of acoustic environments, and artificial modeling of different types of recording devices. Channel related variations can be modeled by, for example, transmitting the signals through actual digital telephony networks by utilizing the same hardware / software configurations as in the intended use-case. These aspects can, of course, also be modeled artificially, for example by implementing channel simulation that is based on realistic set of speech encodings, and modeling of transmission errors, handovers and speech enhancements. Some examples of speech codecs that are widely used in digital telephony, and can therefore provide a realistic basis for channel simulation, include Adaptive Multi-Rate Narrowband (AMR-nb) and Adaptive Multi-Rate Wideband (AMR-wb).

[0075] Some example embodiments of the bio-marker adaptation are presented in the following paragraphs.

[0076] This section describes the implementation details of the bio-marker adaptation module, and the Figure 4 shows an example embodiment. The Bio-marker Adaptation Network (BAN) is the key component, responsible for the transformation of the signal. It consists of separate hidden representation extractor and feature modulator modules.

[0077] BAN assumes the signal is sampled at 16 kHz, the very first step involves resampling if needed. The input signal is then trimmed to ensure that the perceptive fields of the time-frames match between the MFCC extraction process and the BAN processing. In the BAN, the temporal alignment corresponds to 25 ms frames with a 20 ms shift. To align the frames of the BAN with those of the MFCCs, the input signal is trimmed to ensure both the BAN frames and MFCC frames align perfectly with the start and end of the signal.

[0078] Before the signal is processed by BAN, its amplitude is normalized to a range between -1 and 1. Additionally, the extracted MFCCs undergo Zscore normalization using the mean and standard deviation computed from the training data consisting of both non-distorted and distorted versions of all signals. Crucially, the outputs of the BAN are consistently denormalized using these same statistics. This step is important because the disorder detector, which will utilize the BAN outputs, employs its own normalization processes designed for inputs that are not pre-normalized.

[0079] The BAN is a deep learning architecture that takes as input raw speech signals and their corresponding feature representations. A key feature of this network is its ability to model the transformations between distorted and non-distorted versions of a speech signal, subsequently modulating the feature representations accordingly. The first part of the network is the hidden representation extractor. This example embodiment produces embeddings with 1024 dimensions. Upon obtaining the signal embeddings from an input, the feature modulator leverages these embeddings to adjust the initially extracted input signal features. The design of the feature modulator is depicted in Figure 5 The modulator begins by stretching the embeddings across the temporal dimension to synchronize their frame rates with the feature frames. For instance, the original frame shifts of 20 ms from the hidden representation extractor are modified to match the 10 ms frame shifts used by the MFCCs. This adjustment involves the addition of an extra frame between every two successive frames. More precisely, each pair of neighboring frames is merged and then fed into a three-layered fully connected network (FCN). The initial two layers of this network preserve the size of the merged vectors and utilize a leaky ReLU activation function for non-linearity, whereas the last layer applies a linear operation to reduce the dimensionality back to 1024. The resulting interpolated output is combined with the initially extracted features and further processed by another FCN. This second FCN contains two layers with leaky ReLU activation for non-linear processing and concludes with a linear layer that reduces dimensionality. The output from this layer approximates the 13-dimensional static MFCCs of the altered signal. From those, the dynamic coefficients and the temporal statistics are computed by the health-analysis system, as described in Section 4.1 .

[0080] An example embodiment according to the invention of model training is presented below.

[0081] The health-analysis model is trained separately from the bio-marker adaptation model, by utilizing laboratory quality recordings from speakers with medical conditions, diagnosed by medical doctors, and from healthy controls. Training, evaluation and model selection should be conducted by utilizing grid-search of hyper-parameters together with a nested cross-validation process, to avoid biasing to the training dataset.

[0082] The domain adaptation model is trained using general datasets unrelated to health. A crucial preprocessing step involves performing domain adaptation to align the environmental conditions of the two datasets, ensuring that the biomarker adaptation learns a useful mapping. The specifics of achieving this alignment are beyond the scope of this invention, as they depend on the chosen data and may involve modeling the acoustic environment, additive noise, or device characteristics. Typically, transmission-related effects are not considered in this task, as the data is not transmitted through telephony channels but simply recorded.

[0083] Firstly, the hidden representation extractor is trained in self-supervised manner, in isolation from other parts of the model, by utilizing speech that has not been distorted by the channel simulator. Initially, the model was pretrained on approximately 61 ,000 hours of audio. Secondly, the entire domain adaptation model is trained using simulated data. Each training batch contains both distorted and non-distorted versions of the same signals, with the model’s goal being to map distorted signals back to their non-distorted counterparts. The loss function should be designed to penalize differences between the feature representations of the non-distorted signal and the output feature representation from the model. Additionally, since some channel effects can be lossy and exact inversion may not exist, the model should be explicitly informed that approximate answers are sufficient to avoid overfitting to the ’easiest’ samples. One example of a loss-function can be written as follows. Consider P and Q to represent the predicted and desired matrices, each having dimensions n x m. Introduce the convolution operation conv(B) with a 2-width and 1 -height sliding window for any matrix B. Assume y is a chosen threshold. The loss function £ is then expressed as: conv(Q)iy| - y).

[0084] In this formulation, max(0, x) ensures values below 0 are reset to 0, and the summation extends across all matrix elements. The model can be optimized, for example, by using any gradient-based optimization method. The training is conducted for some number of epochs, by freezing the parameters of the hidden representation extractor for the initial epochs, to ensure the stability of the training process.

Claims

CLAIMS1 . A method for speech-based health analysis, the method comprising: receiving via a digital transmission channel an input signal related to a recorded speech signal, the input signal comprising at least one feature related to health information; processing the input signal by adapting the input signal for at least partly cancelling at least one effect on a signal path, wherein the signal path comprises path from speech signal output of a person to a device recording the speech signal and / or the digital transmission channel, wherein the at least one effect on the signal path comprises distortions caused to the input signal by at least one of the following: background noise, acoustic environment, recording equipment characteristics, transmission errors and / or error concealments, network handovers, codec, speech enhancement; and providing the processed signal to a health analysis, e.g. to a health analysis system, wherein the processed signal comprises a feature representation of the speech signal, the feature representation of the speech signal comprising the at least one feature related to health information adapted for the at least one effect on the signal path.

2. The method according to claim 1 , wherein the at least partly cancellation of the at least one effect comprises modification of the input signal so that the at least one feature related to the health information corresponds to a similar at least one feature related to the health information transmitted on a reference signal path.

3. The method according to any of claims 1 to 2, wherein the input signal related to the speech signal comprises the speech signal and / or a feature representation of the speech signal.

4. The method according to any of claims 1 to 3, wherein the digital transmission channel comprises a digital telephony channel and / or digital communication channel.

5. The method according to any of claims 1 to 4, wherein the processing comprises using a machine learning model.

6. The method according to claim 5 wherein the machine learning model is trained using data related to the at least one effect on the signal path.

7. The method according to claim 6, wherein the data related to the at least one effect on the signal path is simulated using a channel simulator.

8. The method according to any of claims 5 to 7, wherein the data used for training the machine learning model does not include health related information used in the training.

9. The method according to any of claims 1 to 8, the method further comprising: performing health analysis using the processed signal, e.g. in a health analysis system; and providing a health analysis result based at least in part on the performed health analysis.

10. The method according to claim 9, wherein the health analysis comprises using a machine learning model for health analysis, the machine learning model for health analysis being trained using health related data.11 . The method according to any of claims 1 to 10, wherein the health analysis is implemented longitudinally, e.g. over a period of time, utilizing one or multiple feature representations extracted from previous recording of the same person.

12. An arrangement for a health-analysis system, wherein the arrangement comprises at least one processing unit, such as a computer, wherein the arrangement is configured: to receive via a digital transmission channel an input signal related to a recorded speech signal, the input signal comprising at least one feature related to health information; to process the input signal by adapting for at least partly cancelling at least one effect on a signal path, wherein the signal path comprises path from speech signal output of a person to a device recording the speech signal and / or the digital transmission channel, wherein the at least one effect on the signal path comprises distortions caused to the input signal by at least one of the following: background noise, acoustic environment, recording equipment characteristics, transmission errors and / or error concealments, network handovers, codec, speech enhancement; and to provide the processed signal for health analysis, e.g. to a health analysis system, wherein the processed signal comprises a feature representation of the speech signal, the feature representation of the speech signal comprising the at least one feature related to health information adapted for the at least one effect on the signal path.

13. The arrangement according to claim 12, wherein the arrangement is configured to perform the method according to any of claims 2 to 11 .

14. A system for health analysis, wherein the system comprises an arrangement according to claim 12 or 13.

15. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any of claims 1 to 11.

16. A computer-readable medium comprising the computer program according to claim 15.

Citation Information

Patent Citations

  • Verbal periodic screening for heart disease

    US20190362740A1

  • Intelligent health monitoring

    WO2020102223A2