Model training and prediction method, device, equipment, medium and product
Patent Information
- Application Number
- CN202511447905.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-10-10
AI Technical Summary
[0003]然而,在临床设施之外的场景中,随着床边监测仪、可穿戴设备等便携式监测设备的日益丰富,这些设备虽能捕捉PSG技术所涵盖监测模态的部分子集,实现一定程度的睡眠监测,但也导致了不同监测设备之间、不同护理场景之下睡眠监测数据呈现出碎片化的格局
[0021] This disclosure provides a method, apparatus, device, medium, and product for model training and prediction. The method selected from multiple types of physiological signals to obtain feature vectors of the same predetermined dimension, which were then mapped to the same feature space for alignment. After processing any two types of physiological signals as described above, the trained model acquires generalization features for each type of physiological signal. This allows for fine-tuning of the trained model to obtain a specialized model capable of predicting different types of tasks based on any one type of physiological signal. Furthermore, the method, through a weighting mechanism, allows for better adjustment of model parameters using training loss, resulting in a model that better acquires generalization features for each type of physiological signal. The specialized model obtained by fine-tuning the model acquired using this method can predict user-related tasks in sleep scenarios based on any user's physiological signals, offering advantages such as reliability, convenience, efficiency, accuracy, and speed, enabling early detection of health risks.
Smart Images

Figure CN121188486B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of sleep monitoring, and more specifically, to a method for training a neural network model, a method for predicting a first task related to a user in a sleep setting, and corresponding devices, electronic devices, recording media, and computer program products. Background Technology
[0002] Sleep, as a core determinant of human health, plays a crucial role in shaping cognitive abilities, metabolic levels, cardiovascular function, and mental health. Sleep disorders are not only important signals of various diseases but also potential contributing factors to a wide range of illnesses. Given this importance, the field of clinical sleep assessment urgently needs precise and comprehensive monitoring technologies. Polysomnography (PSG) technology has thus become the gold standard for clinical sleep assessment. This technology is a multi-sensor recording system capable of simultaneously monitoring multiple key physiological indicators, including neurophysiological activity, ocular electrophysiological activity, muscle tone, dynamic changes in cardiopulmonary function, and blood oxygen saturation, providing crucial evidence for the diagnosis of clinical sleep-related disorders.
[0003] However, outside of clinical facilities, the increasing availability of portable monitoring devices such as bedside monitors and wearable devices, while capable of capturing a subset of the monitoring modalities covered by PSG technology and achieving a certain degree of sleep monitoring, has also led to a fragmented pattern of sleep monitoring data across different monitoring devices and in different nursing scenarios. More importantly, current sleep assessment models lack good adaptability and operational capabilities for this fragmented monitoring data pattern, failing to fully integrate and utilize fragmented data from different sources and modalities, thus making it difficult to guarantee the accuracy and reliability of sleep assessment results in non-clinical settings.
[0004] Given the shortcomings of existing technologies, namely the inability of current models to effectively address the fragmentation of monitoring data in non-clinical scenarios, a new model training method is urgently needed to meet the demand for accurate sleep assessment in practical applications. The model trained by this method can effectively adapt to and cope well with the aforementioned fragmented data pattern, thereby improving the accuracy and effectiveness of sleep assessment in non-clinical scenarios. Summary of the Invention
[0005] To address the aforementioned issues, this disclosure provides a novel model training method. This method enables efficient annotation and general modeling of real-world nocturnal biosignals by utilizing multiple types of physiological signals and employing cross-modal alignment. A dedicated model, fine-tuned using this method, can predict user-related tasks in sleep scenarios based on arbitrary physiological signals, offering advantages such as reliability, convenience, efficiency, accuracy, and speed, and enabling early detection of health risks.
[0006] This disclosure provides a method for training a neural network model, characterized in that the neural network model includes a first multilayer perceptron, a second multilayer perceptron, and a deep learning network based on a self-attention mechanism. The method includes: acquiring multiple types of physiological signals; selecting multiple physiological signals of a first type and a second type with predetermined durations from the multiple types of physiological signals; inputting the multiple physiological signals of the first type with predetermined durations into the first multilayer perceptron to obtain a first feature vector of a predetermined dimension for the multiple predetermined durations, and inputting the multiple physiological signals of the second type with predetermined durations into the second multilayer perceptron to obtain a second feature vector of the predetermined dimension for the multiple predetermined durations; and inputting the first feature vector of the predetermined dimension for the multiple predetermined durations into the second multilayer perceptron. The first fused feature vector is input into a self-attention-based deep learning network to obtain a first fused feature vector, and the second feature vectors of predetermined dimensions with predetermined durations are input into the self-attention-based deep learning network to obtain a second fused feature vector. The first fused feature vector and the second fused feature vector are mapped to the same feature space to obtain multiple third feature vectors corresponding to the first fused feature vector and multiple fourth feature vectors corresponding to the second fused feature vector. A first similarity between each third feature vector and each fourth feature vector is determined. Based on the first similarity, the training loss of the neural network model is calculated, and the parameters of the first multilayer perceptron, the second multilayer perceptron, and the self-attention-based deep learning network are adjusted based on the training loss.
[0007] According to an embodiment of this disclosure, the step of calculating the training loss of the neural network model based on the first similarity includes: determining a second similarity between feature vector pairs consisting of third and fourth feature vectors that satisfy predetermined conditions, wherein the predetermined conditions include that the third and fourth feature vectors in the feature vector pair are obtained based on different types of physiological signals of the same object for the same predetermined duration; and calculating the training loss of the neural network model based on the first similarity and the second similarity.
[0008] According to an embodiment of this disclosure, the step of calculating the training loss of the neural network model based on the first similarity and the second similarity includes: determining a first weight corresponding to the first similarity based on the first similarity, wherein the larger the first similarity, the larger the value of the first weight corresponding to the first similarity; and calculating the training loss of the neural network model based on the first similarity, the first weight, and the second similarity.
[0009] According to an embodiment of this disclosure, the method for calculating the training loss of the neural network model based on the first similarity, the first weight, and the second similarity includes: determining a first similarity adjustment value based on the first similarity and the value of the piecewise constant function corresponding to the first similarity; and calculating the training loss of the neural network model based on the product of the first similarity adjustment value and the first weight, and the second similarity; wherein, when the third feature vector and the fourth feature vector corresponding to the first similarity are obtained based on different types of physiological signals of the same object, the value of the piecewise constant function corresponding to the first similarity is a predetermined constant; otherwise, the value of the piecewise constant function corresponding to the first similarity is zero.
[0010] According to an embodiment of this disclosure, the sampling rate of the first type of physiological signal is different from the sampling rate of the second type of physiological signal.
[0011] According to embodiments of this disclosure, the various types of physiological signals include two or more of the following: electroencephalogram (EEG) physiological signals, electromyogram (EMG) physiological signals, electrooculogram (EOG) physiological signals, electrocardiogram (ECG) physiological signals, intercardiac interval (IBI) physiological signals, nasal airflow physiological signals, abdominal breathing band / thoracic breathing band (ABD / ThorBelt) physiological signals, respiratory movement physiological signals, and blood oxygen saturation (SpO2) physiological signals.
[0012] This disclosure provides a method for predicting a first task related to a user in a sleep scenario, characterized in that the method includes: acquiring at least one physiological signal of the user in a sleep scenario; and inputting the acquired at least one physiological signal into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task; wherein the first prediction model is obtained by fine-tuning a neural network model obtained according to any method for training a neural network model.
[0013] According to an embodiment of this disclosure, the first prediction model is obtained by fine-tuning the acquired neural network model through the following operations: obtaining the true value corresponding to the first task; setting a third multilayer perceptron corresponding to the first task after the deep learning network based on the self-attention mechanism in the acquired neural network model; obtaining the predicted value corresponding to the first task based on the physiological signal using the first multilayer perceptron, the second multilayer perceptron, the deep learning network based on the self-attention mechanism, and the third multilayer perceptron in the acquired neural network model; and adjusting the parameters of the first multilayer perceptron, the second multilayer perceptron, and the third multilayer perceptron corresponding to the first task based on the difference between the predicted value corresponding to the first task and the true value of the first task to obtain the first prediction model.
[0014] According to an embodiment of this disclosure, the first task is any one of the following tasks: a task related to blood oxygenation prediction, a task related to sleep staging, a task related to apnea-hypopnea index estimation, a task related to sleep apnea detection, a task related to demographics, a task related to heart failure, and a task related to gastroesophageal reflux.
[0015] According to an embodiment of this disclosure, the step of acquiring at least one physiological signal of a user in a sleep scenario includes: acquiring the at least one physiological signal through relevant sensor data, wherein the relevant sensor data includes data obtained based on at least one of the following sensors: piezoelectric sensor, piezoresistive sensor, fiber optic sensor, audio sensor, heart rate sensor, nasal pressure sensor, motion sensor, and accelerometer.
[0016] This disclosure provides an apparatus for training a neural network model, characterized in that the neural network model includes a first multilayer perceptron, a second multilayer perceptron, and a deep learning network based on a self-attention mechanism. The apparatus includes: an acquisition module configured to acquire multiple types of physiological signals; a selection module configured to select multiple physiological signals of a first type and a second type of physiological signal of predetermined duration from the multiple types of physiological signals; a first obtaining module configured to input the multiple physiological signals of the first type of physiological signal of predetermined duration into the first multilayer perceptron to obtain a first feature vector of predetermined dimension of the multiple predetermined durations, and to input the multiple physiological signals of the second type of physiological signal of predetermined duration into the second multilayer perceptron to obtain a second feature vector of predetermined dimension of the multiple predetermined durations; and a second obtaining module configured to input the multiple physiological signals of the second type of physiological signal of predetermined duration into the second multilayer perceptron to obtain a second feature vector of predetermined dimension of the multiple predetermined durations; and a second obtaining module configured to input the first feature vector of predetermined dimension of the multiple predetermined durations into the second multilayer perceptron. The system inputs feature vectors into a self-attention-based deep learning network to obtain a first fused feature vector, and inputs the multiple second feature vectors of predetermined dimensions for predetermined durations into the self-attention-based deep learning network to obtain a second fused feature vector; a third obtaining module is configured to map the first fused feature vector and the second fused feature vector to the same feature space to obtain multiple third feature vectors corresponding to the first fused feature vector and multiple fourth feature vectors corresponding to the second fused feature vector; a determining module is configured to determine a first similarity between each third feature vector and each fourth feature vector; and a parameter adjustment module is configured to calculate the training loss of the neural network model based on the first similarity, and adjust the parameters of the first multilayer perceptron, the second multilayer perceptron, and the self-attention-based deep learning network based on the training loss.
[0017] This disclosure provides an apparatus for predicting a first task related to a user in a sleep scenario, characterized in that the apparatus includes: an acquisition module configured to acquire at least one physiological signal of the user in a sleep scenario; and an acquisition module configured to input the acquired at least one physiological signal into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task; wherein the first prediction model is obtained by fine-tuning a neural network model obtained according to any method for training a neural network model.
[0018] This disclosure provides an electronic device, including: a processor and a memory, the memory storing computer-executable instructions that, when executed by the processor, cause the processor to perform any of the methods described above.
[0019] This disclosure provides a computer-readable recording medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, cause the processor to perform any of the methods described above.
[0020] This disclosure provides a computer program product including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, cause the processor to perform any of the methods described above.
[0021] This disclosure provides a method, apparatus, device, medium, and product for model training and prediction. The method selected from multiple types of physiological signals to obtain feature vectors of the same predetermined dimension, which were then mapped to the same feature space for alignment. After processing any two types of physiological signals as described above, the trained model acquires generalization features for each type of physiological signal. This allows for fine-tuning of the trained model to obtain a specialized model capable of predicting different types of tasks based on any one type of physiological signal. Furthermore, the method, through a weighting mechanism, allows for better adjustment of model parameters using training loss, resulting in a model that better acquires generalization features for each type of physiological signal. The specialized model obtained by fine-tuning the model acquired using this method can predict user-related tasks in sleep scenarios based on any user's physiological signals, offering advantages such as reliability, convenience, efficiency, accuracy, and speed, enabling early detection of health risks. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some exemplary embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0023] Figure 1 This is a flowchart illustrating a method 100 for training a neural network model according to an embodiment of the present disclosure;
[0024] Figure 2 This is a flowchart illustrating a method 200 for predicting a first task related to a user in a sleep scenario according to an embodiment of the present disclosure;
[0025] Figure 3 A schematic diagram of polysomnography (PSG) acquiring various physiological signals according to an embodiment of the present disclosure is shown;
[0026] Figure 4 A schematic diagram of a multimodal pre-training framework according to an embodiment of the present disclosure is shown;
[0027] Figure 5 A schematic diagram of t-SNE visualization of encoder embeddings comparing random initialization and pre-trained results according to an embodiment of the present disclosure is shown.
[0028] Figure 6 A schematic diagram of leave-one-out analysis for the SHHS sleep staging task according to an embodiment of the present disclosure is shown;
[0029] Figure 7 A schematic diagram is shown of ROC-AUC scores for a disease prediction task using different modality numbers (N) on the SHHS dataset according to embodiments of the present disclosure;
[0030] Figure 8 A schematic diagram of the Recall@1 retrieval accuracy matrix of the learned representation according to an embodiment of the present disclosure is shown;
[0031] Figure 9 A schematic diagram of a feature fusion technique based on a gating mechanism according to an embodiment of the present disclosure is shown;
[0032] Figure 10 A schematic diagram showing a comparison of sleep stage recognition performance at different model sizes on (a) SHHS and (b) WSC datasets according to embodiments of the present disclosure is illustrated.
[0033] Figure 11 This is a block diagram illustrating an apparatus 1100 for training a neural network model according to an embodiment of the present disclosure;
[0034] Figure 12 This is a block diagram illustrating an apparatus 1200 for predicting a first task related to a user in a sleep scenario, according to an embodiment of the present disclosure.
[0035] Figure 13 This is a block diagram illustrating an electronic device 1300 according to an embodiment of the present disclosure;
[0036] Figure 14 A schematic diagram of the recording medium 1400 according to this disclosure is shown. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0038] In this specification and accompanying drawings, substantially the same or similar steps and elements are indicated by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. Furthermore, in the description of this disclosure, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance or order.
[0039] In this specification and accompanying drawings, elements are described in singular or plural forms according to embodiments. However, the singular and plural forms have been suitably chosen for the presented cases merely for ease of explanation and are not intended to limit this disclosure. Thus, a singular form may include a plural form, and a plural form may include a singular form, unless the context clearly indicates otherwise.
[0040] As mentioned earlier, this fragmented pattern raises a core question: "Can a unified physiological representation model be constructed by cross-modal alignment of nocturnal biosignals, enabling it to achieve robust generalization across heterogeneous sensor clusters in the field of sleep medicine?"
[0041] Training physiological signals (such as pre-training) offers a promising paradigm—learning generalized representations from diverse biological signals with minimal supervision. However, real-world data presents stringent limitations: different acquisition points and devices employ varying sensor arrays, sampling rates may differ, and entire physiological signal acquisition channels are often missing. Furthermore, large-scale expert annotation is prohibitively expensive, making new training methods and the underlying frameworks derived from them both necessary and challenging.
[0042] To address one or more problems in existing technologies, this paper discloses a novel model training method. This method enables efficient annotation and general modeling of real-world nocturnal biosignals by utilizing multiple types of physiological signals and employing cross-modal alignment. A specialized model, fine-tuned using this method, can predict user-related tasks in sleep scenarios based on arbitrary physiological signals, offering advantages such as reliability, convenience, efficiency, accuracy, and speed, and enabling early detection of health risks.
[0043] The method provided in this disclosure will now be described in detail with reference to the accompanying drawings.
[0044] Figure 1 This is a flowchart illustrating a method 100 for training a neural network model according to an embodiment of the present disclosure.
[0045] As an example, the neural network model can be any suitable existing model, which can include various types such as deep neural networks (DNN), convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), transformers, long short-term memory networks (LSTM), residual networks (ResNet) or other neural networks, or combinations of the above neural network models. This application does not limit the type or number of neural network models.
[0046] As an example, a neural network model may include a first multilayer perceptron (MLP), a second multilayer perceptron, and a deep learning network based on a self-attention mechanism.
[0047] like Figure 1 As shown, the method 100 may include steps S110 to S170. For example, the execution subject of this method can be any suitable processing device. The processing device can be located in any electronic device, such as a mobile phone, tablet computer, portable computer, desktop computer, smart wearable device, smart home appliance, or smart vehicle terminal, etc., and the embodiments of this disclosure are not limited thereto.
[0048] Reference Figure 1 In step S110, various types of physiological signals can be acquired.
[0049] As an example, various types of physiological signals may include two or more of the following: electroencephalogram (EEG) physiological signals, electromyogram (EMG) physiological signals, electrooculogram (EOG) physiological signals, electrocardiogram (ECG) physiological signals, intercardiac interval (IBI) physiological signals, nasal airflow physiological signals, abdominal breathing band / thoracic breathing band (ABD / Thor Belt) physiological signals, respiratory movement (RESP) physiological signals, and blood oxygen saturation (SpO2) physiological signals.
[0050] In step S120, a first type of physiological signal and a second type of physiological signal of a predetermined duration can be selected from multiple types of physiological signals.
[0051] As an example, the predetermined duration can be any suitable and easily processed time length, such as 30 seconds. Multiple predetermined durations can form, for example, an entire night. A first type of physiological signal with multiple predetermined durations could be, for example, multiple 30-second EEG physiological signals. A second type of physiological signal with multiple predetermined durations could be, for example, multiple 30-second ECG physiological signals.
[0052] As can be seen, the method provided in this disclosure can process any two types of physiological signals from multiple types. This makes this disclosure a unified multimodal physiological signal pre-training method: the multimodal contrastive pre-training framework provided in this disclosure to date for building basic models of physiological signals. This framework synergistically integrates waveform and interval modal data to reveal comprehensive cross-modal physiological correlations.
[0053] In step S130, multiple physiological signals of a first type with a predetermined duration can be input into a first multilayer perceptron to obtain a first feature vector of a predetermined dimension with a predetermined duration, and multiple physiological signals of a second type with a predetermined duration can be input into a second multilayer perceptron to obtain a second feature vector of the predetermined dimension with a predetermined duration.
[0054] Each type of physiological signal has a corresponding multilayer sensor, such as EEG, EMG, EOG, ECG, IBI, nasal airflow, ABD / Thor belt, respiratory movement, and blood oxygen saturation (SpO2).
[0055] The sampling rates of physiological signals may differ; for example, the sampling rate of one type of physiological signal may differ from that of another type. By employing a multilayer perceptron corresponding to each type of physiological signal, it is possible to convert that type of physiological signal to the same predetermined dimension (e.g., 512). This achieves the alignment of physiological signals with different sampling rates.
[0056] The feature vectors obtained in the above manner (such as the first feature vector and the second feature vector) can characterize the local feature information of the physiological signal of the predetermined duration.
[0057] Continue to refer to Figure 1In step S140, multiple first feature vectors of predetermined dimensions and predetermined durations can be input into a deep learning network based on a self-attention mechanism to obtain a first fused feature vector, and multiple second feature vectors of predetermined dimensions and predetermined durations can be input into a deep learning network based on a self-attention mechanism to obtain a second fused feature vector.
[0058] As an example, the first feature vector and the second feature vector can be input into a self-attention-based deep learning network separately. Alternatively, the first feature vector and the second feature vector can be input into the self-attention-based deep learning network simultaneously. Whether input separately or simultaneously, the self-attention-based deep learning network can process the first feature vector and the second feature vector independently. The fused features can not only represent the local feature information of the physiological signal for the predetermined duration, but also the feature information of multiple predetermined durations (such as the entire night).
[0059] As an example, a deep learning network based on a self-attention mechanism can be any suitable network, such as RoFormer, which will be introduced later.
[0060] In step S150, the first fused feature vector and the second fused feature vector can be mapped to the same feature space to obtain multiple third feature vectors corresponding to the first fused feature vector and multiple fourth feature vectors corresponding to the second fused feature vector.
[0061] As an example, the above mapping can be implemented using a multilayer perceptron, so that the first fused feature vector and the second fused feature vector can be mapped to the same feature space, thereby facilitating related processing in the same feature space.
[0062] In step S150, a first similarity can be determined between each third feature vector and each fourth feature vector.
[0063] As an example, the first similarity between each third and fourth eigenvector can be determined using cosine similarity. Alternatively, other suitable similarity methods can be used to determine the first similarity between each third and fourth eigenvector.
[0064] In step S160, the training loss of the neural network model can be calculated based on the first similarity, and the parameters of the first multilayer perceptron, the second multilayer perceptron, and the deep learning network based on the self-attention mechanism can be adjusted based on the training loss.
[0065] According to an embodiment of this disclosure, calculating the training loss of the neural network model based on the first similarity may include: determining a second similarity between feature vector pairs composed of third and fourth feature vectors that satisfy predetermined conditions from among a plurality of third and fourth feature vectors; and calculating the training loss of the neural network model based on the first and second similarities. Figure 4 X in i It can indicate, for example, the third eigenvector, Y. i It can indicate, for example, the fourth eigenvector.
[0066] As an example, the predetermined condition could include that the third and fourth eigenvectors in the eigenvector pair were obtained based on different types of physiological signals from the same object at the same predetermined duration. That is, these two different types of physiological signals were obtained from the same object at the same predetermined duration (e.g., time). In this case, the eigenvector pair that satisfies the predetermined condition represents a positive sample pair. The eigenvector pair that does not satisfy the predetermined condition represents a negative sample pair.
[0067] According to embodiments of this disclosure, calculating the training loss of a neural network model based on a first similarity and a second similarity may include: determining a first weight corresponding to the first similarity based on the first similarity; and calculating the training loss of the neural network model based on the first similarity, the first weight, and the second similarity.
[0068] As an example, the higher the first similarity, the higher the value of the first weight corresponding to that first similarity. In this case, a predetermined method or any other suitable method can be used to determine the corresponding weight based on similarity; as long as the greater the similarity, the higher the value of the weight corresponding to that similarity.
[0069] According to embodiments of this disclosure, calculating the training loss of a neural network model based on a first similarity, a first weight, and a second similarity may include: determining a first similarity adjustment value based on the first similarity and the value of a piecewise constant function corresponding to the first similarity; and calculating the training loss of the neural network model based on the product of the first similarity adjustment value and the first weight, and the second similarity.
[0070] As an example, when the third and fourth feature vectors corresponding to the first similarity are obtained based on different types of physiological signals of the same object, the value of the piecewise constant function corresponding to the first similarity is a predetermined constant; otherwise, the value of the piecewise constant function corresponding to the first similarity is zero.
[0071] Based on the above combination Figure 1As described in the method 100 for training a neural network model, the method provided in this disclosure can select any two types of physiological signals from multiple types of physiological signals to obtain feature vectors of the same predetermined dimension, and then map them to the same feature space for alignment processing. After each pair of physiological signals from multiple types of physiological signals undergoes the above processing, the trained model can obtain the generalization features of each type of physiological signal. Therefore, when fine-tuning this trained model to obtain a specialized model, the specialized model can perform predictions for different types of tasks based on any type of physiological signal. Furthermore, based on the above combination... Figure 1 The described method also shows that, through the design of the weight mechanism, the training loss can be better used to adjust the parameters of the model, so that the trained model can better obtain the generalization features of each type of physiological signal.
[0072] In addition to providing the aforementioned method for training a neural network model, this disclosure also provides a method for predicting a first task related to a user in a sleep context. This will be explained below with reference to the accompanying drawings.
[0073] Figure 2 This is a flowchart illustrating a method 200 for predicting a first task related to a user in a sleep scenario according to an embodiment of the present disclosure.
[0074] As an example, the first task can be any of the following: a task related to blood oxygenation prediction, a task related to sleep staging, a task related to apnea-hypopnea index estimation, a task related to sleep apnea detection, a task related to demographics, a task related to heart failure, and a task related to gastroesophageal reflux.
[0075] like Figure 2 As shown, the method 200 may include steps S210 to S220. For example, the execution subject of this method can be any suitable processing device. The processing device can be located in any electronic device, such as a mobile phone, tablet computer, portable computer, desktop computer, smart wearable device, smart home appliance, or smart vehicle terminal, etc., and the embodiments of this disclosure are not limited thereto.
[0076] Reference Figure 2 In step S210, at least one physiological signal of the user in a sleep scenario can be acquired.
[0077] As an example, a sleep scenario can be any suitable scenario related to a user's sleep, such as a user's sleep scenario at night.
[0078] In a sleep scenario, at least one physiological signal from the user can be one or more of the following: electroencephalogram (EEG) physiological signal, electromyography (EMG) physiological signal, electrooculogram (EOG) physiological signal, electrocardiogram (ECG) physiological signal, intercardiac interval (IBI) physiological signal, nasal airflow physiological signal, abdominal breathing band / thoracic breathing band (ABD / Thor Belt) physiological signal, respiratory movement physiological signal, and blood oxygen saturation (SpO2) physiological signal.
[0079] According to embodiments of this disclosure, acquiring at least one physiological signal of a user in a sleep scenario may include: acquiring the physiological signal through relevant sensor data, wherein the relevant sensor data includes data obtained based on at least one of the following sensors: piezoelectric sensor, piezoresistive sensor, fiber optic sensor, audio sensor, heart rate sensor, nasal pressure sensor, motion sensor, and accelerometer. It should be noted that this sensor is merely an example, and any other suitable type of sensor may be used, as long as the sensor can acquire data related to any one or more of the following physiological signals: electroencephalogram (EEG) physiological signal, electromyography (EMG) physiological signal, electrooculogram (EOG) physiological signal, electrocardiogram (ECG) physiological signal, intercardiac interval (IBI) physiological signal, nasal airflow physiological signal, abdominal breathing band / thorax breathing band (ABD / Thor Belt) physiological signal, respiratory movement physiological signal, and blood oxygen saturation (SpO2) physiological signal.
[0080] In step S220, the acquired physiological signal can be input into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task.
[0081] As an example, the first prediction model could be based on the above combination Figure 1 The neural network model obtained by the described method is fine-tuned.
[0082] According to embodiments of this disclosure, the first prediction model is obtained by fine-tuning the acquired neural network model through the following operations:
[0083] First, obtain the actual value corresponding to the first task.
[0084] Secondly, a third multilayer perceptron corresponding to the first task is set after the deep learning network based on the self-attention mechanism in the obtained neural network model.
[0085] Then, based on physiological signals, the predicted value corresponding to the first task is obtained by using the first multilayer perceptron, the second multilayer perceptron, the deep learning network based on the self-attention mechanism, and the third multilayer perceptron in the obtained neural network model.
[0086] Finally, based on the difference between the predicted value corresponding to the first task and the actual value of the first task, the parameters of the first multilayer perceptron, the second multilayer perceptron, and the third multilayer perceptron corresponding to the first task are adjusted to obtain the first prediction model.
[0087] After the above fine-tuning, a specialized model for this first task can be obtained.
[0088] It is evident that the specialized model obtained by fine-tuning the model obtained using this method can predict user-related tasks in sleep scenarios based on any physiological signals of the user. It has the advantages of reliability, convenience, efficiency, accuracy, and speed, and can detect health risks at an early stage.
[0089] To make the methods provided in this disclosure clearer, examples will be given below. It should be noted that only some preferred examples are described below.
[0090] Note that concurrent nighttime signals represent multiple perspectives of the same underlying physiological state. These heterogeneous perspectives are then correctly aligned into a unified representation space (as described above). Figure 1 Using the same feature space (described above), downstream tasks can flexibly process arbitrary modal data without retraining a dedicated pipeline. This space must generate sufficiently robust modalities (as described above). Figure 1 The model describes different types of physiological signals and provides independent representations to ensure reliable inference even in the event of modality loss. This leads to the extension hypothesis: increasing modality diversity and model capacity enriches semantic coverage and normalizes modality-specific nuances. While the scalability principle may be extensively studied in language and vision domains, its application in the field of physiological signals remains largely unexplored. Therefore, this disclosure proposes and evaluates a framework that demonstrates predictable benefits from extending the PSG base model along the modality and parameter axes, particularly in cross-center generalization scenarios—scenarios where sensor configuration, subject (i.e., object) characteristics, and acquisition protocols are prevalent.
[0091] Figure 3 A schematic diagram of polysomnography (PSG) acquiring various physiological signals according to an embodiment of the present disclosure is shown. Figure 3Each modality is presented in 30-second intervals. High-sampling-rate electrophysiological channels include EEG, EMG, EOG, and ECG, while lower-sampling-rate cardiopulmonary and blood oxygenation channels cover nasal airflow, ABD / Thor bands, and SpO2. Although IBI and respiratory motion signals are not directly recorded by PSG, they can be derived from ECG and ABD / Thor bands, respectively, and can also be measured via wearable devices. These synchronously acquired nocturnal signals collectively provide complementary perspectives on underlying physiological states, highlighting the inherent multimodal complexity of sleep monitoring.
[0092] Regarding the physiological signals in polysomnography (PSG), PSG is a comprehensive nocturnal monitoring system that records various physiological signals during sleep, including brain activity, eye movements, muscle tone, heart rate, breathing patterns, and blood oxygen levels. Figure 3 The diagram illustrates a subject wearing a PSG device for nighttime sleep recording. This technology is widely used in clinical and research fields to objectively assess sleep stages and abnormalities, and is used to diagnose and study sleep disorders, including but not limited to sleep apnea, narcolepsy, and insomnia.
[0093] Previous research has only partially addressed some of the needs. Existing models are typically trained for specific downstream tasks, lacking the generality required as a base model and failing to support multi-task processing. Pre-training has shown potential on limited physiological datasets but has not yet scaled to full PSG sensor arrays. When multimodal data is involved, training objectives often focus on reconstruction rather than explicit cross-modal alignment. While reconstruction objectives can maintain the fidelity of modality-specific details, they cannot ensure that heterogeneous inputs map to a shared semantic manifold. Therefore, the inference process often assumes access to the same modality set as training, leading to a significant performance degradation in real-world sensor-deficient scenarios. Furthermore, systematic analyses of performance scaling with modality and parameters remain scarce.
[0094] This disclosure fills these gaps with a newly proposed model (hereinafter referred to as the sleep2vec model), which aligns heterogeneous nighttime signals to a unified embedding space. This framework collaboratively utilizes nine modalities of data: waveform channels including EEG, EOG, EMG, ECG, nasal airflow, ABD / Thor band, and SpO2; as well as interval-derived IBI and respiratory effect features. The data originates from physiological recordings of 42,249 nights. This is achieved by employing a context-aware InfoNCE objective function (as described above). Figure 1 The model (described by the training loss) can explicitly model physiological similarities (such as age, gender, and location of collection) to dynamically weight samples, thereby effectively distinguishing between difficult-to-classify and easy-to-classify negative samples and avoiding overfitting to subtle features of a specific dataset.
[0095] This public contribution includes one or more of the following:
[0096] (i) Unified multimodal physiological signal pre-training: This disclosure proposes the largest multimodal contrastive pre-training framework to date for building basic models of physiological signals. This framework synergistically integrates waveform and interval modal data to reveal comprehensive cross-modal physiological correlations.
[0097] (ii) Extension Law Research: This disclosure systematically explores the impact of multimodal diversity and parameter dimensionality on the extension of the PSG base model. It demonstrates that cross-cohort generalization capability can be predictably improved while minimizing task-specific labels.
[0098] (iii) Cross-modal training objective: This disclosure proposes InfoNCE based on demographics, age, collection site, and historical records (i.e., DASH-InfoNCE, which will be introduced later) – a context-based contrastive learning objective that conditionally weights negative samples using metadata about demographics, age, collection site, and historical records. This metadata-guided weighting mechanism can suppress queue-specific shortcuts and improve robustness and cross-site generalization under heterogeneous PSG sensor configurations.
[0099] (iv) Comprehensive downstream evaluation: This disclosure extensively tests sleep2vec on the SHHS and WSC datasets, covering tasks such as sleep staging, demographic prediction, and diagnostic results. This is the most comprehensive evaluation of the PSG base model to date.
[0100] Related work
[0101] Multimodal alignment enables flexible reasoning. Contrastive alignment maps heterogeneous inputs to a shared embedding space, thereby achieving robustness in zero-sample transfer, retrieval, and input permutation. In the vision-language domain, CLIP generalizes large-scale image-text alignment, while ImageBind achieves six-modal many-to-one binding. In particular, polysomnography (PSG) contains dozens of synchronous channels (EEG, EEG, EMG, ECG, nasal airflow, respiratory effort, blood oxygen saturation, etc.), but existing multimodal alignment research rarely goes beyond a single EEG channel or a small number of paired channels, and even less often deals with this extended hybrid transfer.
[0102] Self-supervised learning of sleep and PSG data. Self-supervised learning (SSL) of sleep data has evolved from early task-specific approaches (such as sleep staging using fixed PSG electrode configurations) to a broader pre-training framework. Recent research, while focusing on building foundational models, is often limited by: (i) pre-training strategies designed for a single downstream task or a limited set of labels; (ii) alignment limited to a selected subset of PSG channels, failing to achieve comprehensive multimodal integration; or (iii) prioritizing the fidelity of specific modal signals over explicit alignment of heterogeneous modalities when employing cross-modal generation methods. Therefore, cross-modal alignment covering the full PSG spectrum remains in its early stages of exploration, and current pre-training primarily focuses on fixed, small EEG configurations.
[0103] Expansion and generalization capabilities. In the language and vision domains, performance exhibits a predictable trend as model and data scale increases. Despite rapid research progress, systematic studies on the expansion patterns of physiological time series and polysomnography (PSG) remain scarce. Existing PSG self-organizing maps (SSLs) rarely explore modal diversity expansion or parameter expansion. It is noteworthy that a systematic modal diversity expansion map has not yet been established in sleep research; existing studies lack sufficient description of parameter / data expansion in PSG SSLs, resulting in the unclear mechanism of capability growth under the combined effect of model size and channel number.
[0104] The following is a description of the methods used in this disclosure.
[0105] About datasets and preprocessing
[0106] This disclosure utilizes publicly available PSG datasets for pre-training, including the Human Sleep Project (HSP) and four cohorts from the National Sleep Research Resource Center (NSRR): the Sleep Heart Health Study (SHHS), the Male Osteoporotic Fracture Study (MrOS), the Multi-Ethnic Atherosclerosis Study (MESA), and the Wisconsin Sleep Cohort Study (WSC). Specific descriptions of these datasets are provided below and will not be repeated here. These datasets collectively encompass multicenter, multi-device data collection from diverse demographics (age range: 1–109 years; recording span: 1995–present). Table 3, presented below, summarizes the five datasets involved. These five datasets were standardized and integrated into a unified corpus containing 42,249 overnight recordings from 30,852 participants. The standardized processing workflow minimized cohort-specific bias, ensuring symmetry in batch processing and evaluation.
[0107] like Figure 3As shown, a cross-cohort signal pool comprising nine PSG channels was established: consisting of two groups of signals distinguished by sampling rate—high-sampling-rate electrophysiological signals, including EEG, EOG, and EMG, uniformly resampled to 128 Hz; and low-sampling-rate physiological signals, including nasal airflow, abdominal / back band / chest band, blood oxygen saturation, IBI, and respiratory motion signals, uniformly resampled to 4 Hz. The high-sampling-rate signals underwent only minimal preprocessing to preserve the original signal characteristics, crucial for downstream physiological interpretation. Preprocessing included: resampling the signal in the time domain to the target frequency and standardizing the z-value using cohort-invariant statistics. The IBI channel was derived from ECG R-wave peak detection; the original heartbeat intervals were cleaned of outliers and artifacts and then converted to a continuous 4 Hz sequence via linear interpolation. Respiratory motion effects, reflecting respiratory cycles extracted from nasal airflow or abdominal band, were band-limited, standardized, and then resampled to 4 Hz. IBI and respiratory motion effects can be acquired not only through polysomnography (PSG), but also through simpler, less burdensome hardware, such as ballistic ECG pads or other non-contact sensors.
[0108] Participant information (including age, gender, and collection point) should be preserved where available to facilitate queue-aware analysis and difficulty estimation during pre-training. Participant-level data partitioning should be implemented to prevent data leakage between splits. A dedicated pre-training split set (N) should be used. pre-train =23,934 participants) were dedicated to basic model learning, and the downstream split set was divided in an 8:1:1 ratio (N train / N val / N test =8,792 / 1,102 / 1,116), ensuring no overlap among participants and complete consistency in modality coverage. In downstream evaluation, the sleep stage labels for SHHS and WSC, as well as the clinical diagnosis labels for SHHS, used the same participant-level segmentation scheme as in pre-training to ensure consistency and prevent data leakage.
[0109] Model Architecture
[0110] Figure 4 A schematic diagram of a multimodal pre-training framework according to an embodiment of the present disclosure is shown. (Refer to...) Figure 4 Each overnight polysomnography recording was divided into intra-subject (i.e., subject) segments (slices from different times within the same subject) and inter-subject segments (slices from different subjects). These segments were processed using a specific modality of MLP (as described above). Figure 1 The first MLP, second MLP, etc., described above, are processed independently by the tokenizer. This is achieved through a modality-independent RoFormer backbone network (as described above). Figure 1Before processing, each mask sequence in the self-attention-based deep learning network described above is preceded by a learnable [CLS] tag. The hidden states of the backbone network at each timestep are projected into a shared alignment space, thereby achieving cross-modal timestep alignment comparison.
[0111] A minimal multilayer perceptron (MLP) was implemented, consisting of two feedforward layers and a residual connection; it uses an initial linear transformation to reduce the time required for a 30-second (as described above) transition. Figure 1 The input segment (described as a predetermined duration) is mapped to an embedding vector of dimension D. This transformation projects the input to a 2D intermediate hidden representation layer, processed by the SiLU nonlinear activation function, and regularized with dropout at a probability of 0.1. A 30-second segment is selected to conform to the standard segment length recommended by the American Academy of Sleep Medicine (AASM) guidelines for polysomnography analysis. This hidden representation is then linearly transformed into the final embedding space (D). In parallel processing, the residual linear transformation directly maps the input to the output embedding dimension, thereby enhancing gradient flow and training stability. The final embedding vector is normalized using LayerNorm. Cross-modal sampling rate differences are addressed using a modality-specific MLP: the MLP processing results for the 30-second segment are converted into units with the same embedding dimension. Each MLP directly processes the original sampling rate data, generating cross-modal time-aligned embedding vectors for processing by the modality-independent main network.
[0112] like Figure 4 As shown, a simple yet effective sampling strategy is employed to ensure optimization stability as the number of modalities increases during pre-training: two modalities (m) are randomly selected from each mini-batch. a m b ), extract one instance for each modality (e.g. Figure 4 (As shown). Independent time-step masks are then applied to paired instances to enhance robustness and suppress shortcut learning. Each 30-second token has a 15% probability of being replaced by a learnable modality-specific mask token, and alignment is then performed only on these mask segments. A learnable [CLS] token is added to the beginning of the sequence, and the final input is processed through a modality-independent RoFormer backbone network. The backbone network outputs the hidden state at each time step and generates a global nighttime representation at the [CLS] position. These hidden states are mapped to a 128-dimensional shared alignment space via a shared three-layer MLP projector, thereby applying a cross-modal contrastive loss at each time step.
[0113] During the fine-tuning phase, both the masking mechanism and the contrastive learning projection head were removed, while the modality configuration remained fixed based on the downstream task. Task-specific heads directly applied to the backbone features: sequence-level tasks (such as sleep staging) used the hidden state at each time step, while aggregation tasks (such as gender, age, or clinical diagnosis) relied on the global nighttime representation at the [CLS] position. When multiple modalities existed during the inference phase, their representations were aggregated using simple fusion strategies (such as averaging, concatenation, or small gating modules). Details of the fusion methods used for specific tasks can be found in the experimental results section.
[0114] Cross-modal alignment target: DASH-INFONCE
[0115] During pre-training, each mini-batch contains B pairs of paired segments, each segment pair being L time steps in length. For segment indexing... Time Index and modality ,remember This is the corresponding d-dimensional embedding vector. Given... As its Norm-normalized product, by Given the cosine similarity (as described above) Figure 1 The first similarity (as described) is:
[0116]
[0117] In equation (1), Represents the dot product operation. Index mapping. Specify mode m a Mode m of the middle anchor point b The number of paired segments is usually determined by batch alignment. The demographic and metadata collection for fragment i is represented as follows: ,in and u i These represent age, gender, collection location, and subject-night identifier, respectively. These variables are used only for subsequent weight calculation and modulation and are never used as labels in the learning objective.
[0118] Basic formula: Time-domain InfoNCE
[0119] When the temperature τ > 0, make m a With m b The aligned baseline time step InfoNCE loss is:
[0120]
[0121] The objective is to promote similarity between paired cross-modal embeddings (i, π(i), t) that exceeds the similarity with all concurrent options (i, j, t) (where j ≠ π(i)). The temperature coefficient τ controls the degree of concentration of the induced softmax distribution.
[0122] In formula (2), the numerator of the expression enclosed in log indicates the similarity of the positive samples (as described above). Figure 1 The second similarity (described) indicates the similarity of all samples in the denominator.
[0123] The DASH-InfoNCE loss letter proposed in this disclosure number
[0124] This section proposes a novel DASH-InfoNCE loss function (as described above) Figure 1 The training loss described reshapes the negative sample set in two ways: (i) metadata-driven sample weighting, and (ii) margin-based pseudo-negative sample adjustment (i.e., negative samples from the same subject - night). For anchor (i, t), define:
[0125]
[0126] in To meet The segmented weights of the constraints, where γ ≥ 0 is the modulation intensity, is a predetermined value; Used to reduce the effective logit value of a specified spurious negative sample before softmax. Binary indicator variable. Determine which samples receive amplitude modulation, and follow The convention ensures that positive samples are not penalized. Optional factors. Encoding time-specific signals, in the implementation below, this disclosure sets ψ to a fixed interval value, and... This option is incorporated. Compared to equation (2), the numerator remains unchanged, while the denominator is changed by ω. i,j The probability quality is focused on negative samples with similar demographic characteristics (and which are more difficult to identify), while the competition intensity of negative samples from the same subject on the same night is weakened by the subtraction magnitude term γψ.
[0127] Sample weighting mechanism
[0128] set up It is a non-negative symmetric kernel function, and its value varies with age difference. Decrease. This disclosure further defines the gender similarity factor. And obtaining site similarity factors ,in and , where the value is based on gi Is it equal to g? j and c i Is it equal to c? j To be chosen.
[0129] Given a pseudo-negative sample indicator Unnormalized weights are defined as follows: ,in The normalized weights are calculated using the following formula:
[0130]
[0131] This weighting scheme assigns higher weights to negative samples that are highly matched in age, gender, and acquisition location. The constant ε ensures that negative samples from the same subject at night maintain a non-zero weight, stabilizing the denominator in equation (3) when highly matched demographic negative samples are scarce. Equation (4) indicates that the more similar the samples are, the greater their corresponding weight values. The similarity of the subjects can be determined by their age, gender, and acquisition location; for example, if the subjects corresponding to two physiological signals have the same age, gender, and acquisition location, then their similarity is very high.
[0132] Pseudo-negative modulation
[0133] This disclosure modulates only negative content extracted from the same subject at night. And instantiate using only the magnitude constraint of ψ:
[0134]
[0135] By combining equations (5) and (3), a fixed amplitude γm is subtracted from the logit values of negative samples from the same subject and the same night before softmax maximization, where m is a predetermined value. This reduces the tendency to over-penalize semantically similar negative samples from the same subject and the same night, while ensuring its presence in the denominator through equation (4). Equation (5) characterizes the above combination. Figure 1 The described piecewise constant function.
[0136] Final target loss function
[0137] The goal of DASH-InfoNCE is to average the loss per anchor across instances and time, as in equation (3):
[0138]
[0139] The mean over t is forced to align at each time step. Note that each part of equations (3)-(6) depends only on demographics and collected metadata (a i g i ci ) and the identifier u through Equation 4 i Downstream task labels are not used during pre-training.
[0140] Feature Fusion combine
[0141] In multimodal physiological tasks, the way modal-specific features are fused / aggregated directly impacts performance. Simple connection strategies (i.e., connecting embeddings before classification) produce high-dimensional, sparse representations, which increase computational and sample complexity and exacerbate overfitting. Conversely, average aggregation (i.e., element-wise averaging) assumes equal information content and reliability across channels; in practice, physiological flows vary in terms of SNR and complementary content, so uniform averaging washes out mode-specific cues and is vulnerable in the absence of sensors.
[0142] To address these limitations, this disclosure employs a gating mechanism that introduces learnable scalar weights assigned to each modality. This approach adaptively emphasizes modalities based on their information content, dynamically adjusting the contribution of each PSG channel. Consequently, it yields more expressive, compact, and task-oriented aggregated representations, enabling efficient downstream learning. This mechanism is well-tunable for efficient downstream learning.
[0143] experiment
[0144] Pre-training diagnostics: alignment and retrieval
[0145] Figure 5 A schematic diagram of t-SNE visualization of encoder embeddings comparing randomly initialized and pre-trained results according to embodiments of the present disclosure is shown. Left panel (Subject-Modality Alignment): Visualization of [CLS]-labeled embeddings shows that pre-training effectively clusters embeddings of different modalities into independent subject-specific groups, indicating that subject-level physiological states have been aligned. Right panel (Time-Modality Alignment): In the time-step embedding visualization, the size of the dots indicates the temporal order (larger → later). The pre-trained embeddings form structured trajectories, contrasting sharply with the scattered distribution observed before training.
[0146] To evaluate the effectiveness of multimodal alignment, Figure 5 (Left panel) The embedding maps of [CLS] labels before and after pre-training were visualized using t-SNE. In the initial state, the embedding maps were clustered by modality, reflecting the inherent modality-specific biases and heterogeneous signal characteristics. After pre-training, the embedding maps of different modalities corresponding to the same subject showed co-clustering, indicating improved modality alignment and preservation of subject-specific structures.
[0147] Further analysis of time-step embeddings of randomized subjects ( Figure 5 The right panel shows a structured trajectory after training, indicating effective modal alignment at a finer temporal resolution. The coordinated variation in the size of the midpoints of the concentric rings highlights the temporal consistency in the representation. This consistency is beneficial for downstream sequence tasks such as sleep staging, demonstrating the practical advantages of time-aligned sleep2vec embeddings.
[0148] Downstream fine-tuning results
[0149] Sleep stages
[0150] This disclosure first evaluates the quality of the learned representations in a sleep staging task. Experiments were conducted on the SHHS and WSC datasets, and the results are shown in Tables 1 and 5 below, respectively. The following trends can be observed from Table 1:
[0151] (i) Comprehensive studies on polysomnography (PSG) data remain extremely limited, with most existing methods focusing only on single-channel EEG or small subsets of physiological signals. Specialized methods typically achieve state-of-the-art performance with existing channel combinations, setting a highly challenging benchmark for basic models.
[0152] (ii) Baseline models typically perform weaker compared to dedicated sleep staging methods optimized for sleep data. This difference is particularly pronounced in scenarios using only EEG: dedicated models consistently outperform the baseline models SleepFM (86.6% accuracy) and sleep2vec (87.4% accuracy) on overall metrics. However, the difference is minimal, with sleep2vec nearly matching the dedicated models on specific metrics (Kappa value 0.82 vs 0.83).
[0153] (iii) Across all polysomnography channel subsets, sleep2vec consistently outperforms the baseline models. Its performance is particularly outstanding in configurations such as "IBI & RESP," achieving an accuracy of 83.0%, significantly surpassing baseline FMs (SleepFM 79.4%, SleepFounder 80.9%). sleep2vec's performance is comparable to, and in some cases even better than, dedicated models.
[0154] (iv) Increasing modal diversity has a positive effect; sleep2vec consistently demonstrates improved performance when more physiological signals are incorporated. This trend highlights the scalability and practicality of integrating multimodal data into the basic model framework, further confirming sleep2vec's ability to effectively integrate multimodal physiological signals.
[0155]
[0156] Table 1: Performance of five sleep stages (W / N1 / N2 / N3 / REM) on different PSG channel sets and models on the SHHS dataset.
[0157] The reporting metrics in Table 1 cover overall performance, including accuracy (Accumulation, %), Cohen Kappa coefficient (κ), macro F1 score (MF1, %), sensitivity (Sensitivity, %), and specificity (Spectrality, %). Individual classification F1 scores (%) are also listed. To ensure fair comparison, the baseline models reproduced in this disclosure are marked with †. It should be noted that these baseline (FM) models are pre-trained separately for each PSG channel subset, while sleep2vec is only pre-trained once across all modalities. Underlined numbers indicate best overall performance within each channel set; bold numbers indicate best performance within the baseline models; bold underlined numbers indicate cases where the baseline model outperforms the specialized models.
[0158] Leave-one-out analysis
[0159] To further explore the role of each modality, this disclosure employs a leave-one-out (LOO) method, excluding one modality during both the pre-training and fine-tuning phases. Using the same experimental setup as in the "Sleep Staging" section above, this disclosure evaluates the model on the SHHS dataset. Figure 6 As shown, excluding modalities such as EEG or IBI leads to a significant decrease in accuracy, while the impact of other modalities (such as SpO2 and EOG) is relatively small.
[0160] Figure 6 A schematic diagram of leave-one-out analysis for the SHHS sleep staging task according to an embodiment of this disclosure is shown. Each bar represents the model accuracy when any one of the nine modalities is excluded during pre-training and fine-tuning. The decrease in accuracy relative to the full-channel baseline (marked "none") reflects the contribution of each modality to the overall model performance and its relative importance.
[0161] Clinical Disease Prediction and Modal Expansion
[0162] Figure 7 A schematic diagram is shown of ROC-AUC scores for a disease prediction task on the SHHS dataset using different numbers of modalities (N) according to an embodiment of this disclosure. The result is the average of all possible combinations of N modalities.
[0163] In the clinical evaluation, this disclosure selected four common and clinically significant diseases from the SHHS dataset, including allergies / sinus problems, asthma, hypertension, and coronary artery disease (e.g., Figure 7(As shown). These diseases encompass two major physiological systems directly monitored by PSG: the respiratory and cardiovascular systems. By incorporating these diverse clinical outcomes, this disclosure explicitly examines whether cross-modal embeddings can achieve robust generalization across different organ systems and sensor subsets.
[0164] Specifically, for a given number of modes N, this disclosure enumerates all possible combinations of modes, constructs the corresponding ensemble model, and reports the average ROC-AUC value. Figure 7 The results show that: (i) there is a significant modality expansion effect—performance continuously improves with the increase in the number of modalities, indicating a robust expansion law across clinical prediction tasks; (ii) the proposed DASH-InfoNCE loss function consistently outperforms the standard InfoNCE baseline, demonstrating its ability to effectively uncover richer cross-modal physiological correlations. The performance advantage of DASH-InfoNCE becomes increasingly significant with the increase in the number of modalities, highlighting its superior efficacy in large-scale multimodal pre-training scenarios.
[0165] in conclusion
[0166] The sleep2vec basic model proposed in this disclosure achieves robust physiological feature representation learning by mapping multimodal sleep polygraph (PSG) signals to a unified embedding space. Based on over 42,000 overnight recordings and an innovative DASH-InfoNCE loss function (which controls for differences in demographics, age, monitoring locations, and medical history), this disclosure achieves significant performance improvements in sleep staging and clinical prediction tasks. Experiments validate the robustness of sleep2vec to incomplete sensor data and reveal a clear scalability law that increases with modal diversity and model size. The results of this disclosure establish sleep2vec as a scalable and versatile tool, enabling generalized physiological monitoring and clinical decision support in the field of sleep medicine.
[0167] The following appendix examples are provided to provide a further understanding of the content provided in this disclosure.
[0168] Appendix A
[0169] A.1 Dataset Overview
[0170] Table 3 summarizes five large, publicly available polysomnography cohorts, covering individuals aged 1 to 109 years, with a data collection period spanning nearly thirty years (1995 to present). This dataset contains 42,249 overnight recordings from 30,852 participants. The HSP dataset has the broadest age range and the largest proportion of participants, while SHHS, WSC, MrOS, and MESA provide well-defined adult cohorts. The diverse demographic characteristics and data collection periods enable robust pre-training and evaluation of the models across heterogeneous sensor and electrode configurations.
[0171] A.2 Pre-training Configuration
[0172] During pre-training, this disclosure ensures consistency by using a fixed batch size of 320 across all models. [CLS] tokens with a prefix are included, and the maximum sequence length is set to L = 121. In contrastive learning, the temperature parameter is set to τ = 0.2. The AdamW optimizer (learning rate...) is employed. β=(0.9,0.95), The non-normalized weights are decayed by 0.01. The first 3% of steps use linear warm-up, and then switch to cosine decay.
[0173] By adjusting the architecture hyperparameters, sleep2vec variants with varying numbers of parameters were generated, as detailed in Table 4. Unless otherwise specified, all experiments used sleep2vec. medium Variants are performed.
[0174]
[0175] Table 2: Polysomnography (PSG) channels and interval-based derived features used in this study.
[0176] High-sampling-rate electrophysiological channels include EEG, EMG, EOG, and ECG; low-sampling-rate cardiopulmonary and oxygenation monitoring channels cover nasal airflow, abdominal-chest belt (ABD / Thor belt), and oxygen saturation (SpO2). Respiratory effort (i.e., the aforementioned respiratory movement effort, RESP) and IBI are interval-derived features, obtained from the ABD / Thor belt and ECG channels, respectively, and can also be measured via wearable devices. The sampling rate range is summarized from the minimum digital sampling rate recommended by AASM Berry et al.; actual device settings may vary.
[0177]
[0178] Table 3: Overview of the PSG dataset used in this study
[0179] For the sample weighting mechanism introduced in the above "Sample Weighting Mechanism" section, this disclosure uses the Laplace kernel function.
[0180]
[0181] Where bandwidth σ age = 20.0. This setting is based on clinical observations: sleep physiology and sleep-disordered breathing change gradually rather than abruptly with age. The 20-year extension captures meaningful lifetime differences while avoiding over-penalizing small age intervals.
[0182] The gender coefficient is set to γ same = 1.0 and γ diff = 0.8, to acknowledge gender differences in sleep structure and the prevalence / severity of sleep-disordered breathing—differences that exist at the individual record level but are not dominant. The site coefficient is set as δ. same = 1.3 and δ diff = 0.8, to account for systematic inter-site differences (equipment, electrode configuration, scoring protocol) that are often greater than the sex effect in multi-center cohorts. To ensure numerical stability, [the following values are used]. .
[0183] Finally, when adjusting for spurious negative outcomes for the same subject on the same night, a fixed amplitude term γm = 0.1 was applied, which reflects the high correlation of repeated segments in the recording and avoids treating them as completely independent negative outcomes.
[0184]
[0185] Table 4: Configuration schemes of the sleep2vec model at different scales
[0186] Each pre-training run uses two high-memory GPUs, with a maximum training time of up to 48 hours.
[0187] A.3 Retrieval Accuracy Matrix
[0188] The alignment quality was qualitatively validated by calculating the mean recall@1 across modalities on a test set containing 10,109 samples. The sleep2vec model achieved a recall@1 of 36.8%, significantly outperforming the baseline model (35.1%) using the original InfoNCE loss function. This highlights the representational quality and discriminative power of the obtained embedding vectors.
[0189] Figure 8 A schematic diagram of the Recall@1 retrieval accuracy matrix of the learned representation according to an embodiment of the present disclosure is shown. Rows correspond to query modes, and columns represent retrieval modes.
[0190] Figure 8The visualized Recall@1 cross-modal retrieval accuracy matrix reveals a clear modality-specific alignment pattern. Notably, respiratory motion and the ABD / Thor modality exhibit extremely high mutual retrieval accuracy (0.82), which is highly consistent with known physiological coupling relationships. The EEG and EOG modalities also show a strong correlation (0.69). Conversely, modalities such as SpO2 and EMG typically have lower retrieval accuracy (≈0.2–0.4), reflecting their relatively weak physiological relevance.
[0191] A.4 Complementary downstream fine-tuning results
[0192] A.4.1 Sleep Stage Configuration
[0193] In the sleep staging task, this disclosure employs Low-Rank Adaptive (LoRA) to fine-tune the Transformer backbone for efficient adaptive parameter adjustment. Specifically, while keeping the original backbone parameters frozen, LoRA adapters are integrated into the query, key, and value projections of each Transformer layer. Unless otherwise specified, this disclosure sets the rank r=8, scaling factor α=16, dropout probability p=0.05, and does not add any additional bias terms. For multimodal fine-tuning scenarios, all modalities share the same set of LoRA adapters. At each time step, this disclosure does not use a special classification token; instead, the output embedding of the final transformer layer is projected through two MLP classifiers, with its hidden dimension matching the output dimension of the main model. Optimization uses the AdamW optimizer with a learning rate of [missing value]. The weight decay coefficient is .
[0194] A.4.2 Performance Evaluation of the WSC Dataset
[0195]
[0196] Table 5: Performance of five sleep stages (wakefulness / N1 / N2 / N3 / REM) on the WSC dataset across different PSG channel sets and models.
[0197] The reporting metrics in Table 5 cover overall performance: accuracy (Accumulation, %), Cohen Kappa coefficient (κ), macro F1 score (MF1, %), sensitivity (Sensations, %), and specificity (Specions, %). F1 scores (%) for each category are also listed. To ensure fair comparison, the baseline models reproduced in this disclosure are marked with †. It should be noted that these baseline models were pre-trained separately for each PSG channel subset, while sleep2vec was only pre-trained once across all modalities. Other naming conventions follow the standards adopted in Table 1. Underlined numbers indicate best overall performance within each channel set; bold numbers indicate best performance in the base model; bold underlined numbers indicate cases where the base model outperforms the specialized model.
[0198] The following trends can be observed from Table 5:
[0199] (i) Comprehensive studies on PSG data remain insufficient, with existing methods typically focusing only on single-channel EEG or a small subset of physiological signals. Specialized models usually provide mature benchmarks, especially in cardiopulmonary configurations, establishing important evaluation criteria for general-purpose basic models.
[0200] (ii) The baseline model consistently demonstrates strong performance, often outperforming dedicated sleep staging methods. This is particularly evident in EEG configurations: sleep2vec achieves the best overall performance among the baseline models (accuracy: 86.3%, κ value: 0.80), surpassing the benchmark baseline model SleepFM (accuracy: 84.3%, κ value: 0.76).
[0201] (iii) sleep2vec demonstrates superior performance across various PSG channel subsets. Particularly in the "IBI & RESP" channel configuration, it significantly outperforms both the specialist and base models (accuracy 81.6%, compared to 79.8% for SleepFounder and 77.2% for SleepFM). Similarly, in the "EEG & EOG & EMG" subset configuration, sleep2vec not only surpasses the baseline base model (accuracy 86.8% vs 84.5%) but also significantly outperforms specialist methods (accuracy 86.8% vs 77.6%). In the full-channel configuration, sleep2vec achieves top performance across multiple metrics (accuracy: 87.4%, κ coefficient: 0.81), highlighting the effectiveness of the integrated modality combination.
[0202] A.4.3 Interpretability of Channel-Level Contributions
[0203] Figure 9 A schematic diagram of a feature fusion technique based on a gating mechanism according to an embodiment of the present disclosure is shown, in which the modality-specific weights learned for the four SHHS sleep staging configurations in Table 1 are visualized.
[0204] In addition to sleep staging performance, the gated scalar fusion technique used in the fine-tuning process provides transparent modality-level attribution for downstream decision-making. Specifically, the learned scalars quantify the contribution of each modality to the task-specific representation.
[0205] To illustrate this point, this disclosure analyzes the SHHS sleep staging results under the four modal configurations in Table 1. Figure 9 The normalized fusion weights are shown (visualized with a sharpening factor T=0.4 to enhance display clarity while maintaining relative proportions). EEG received the highest weight across all tasks, consistent with the fact that sleep staging is primarily based on EEG annotations. EOG and EMG provided important supplementary signals, while cardiopulmonary pathways (such as airflow, abdominal / chest straps, and IBI) had lower weights, and SpO2 contributed the least.
[0206] These weights represent global, task-level attributions, rather than explanations every 30 seconds, and do not capture higher-order interactions between modalities. Nevertheless, these weights still provide an interpretable summary of channel contributions that aligns with domain expectations, guiding sensor selection in channel-constrained deployments.
[0207] A.4.4 Scaling Law for Basic Model Parameters
[0208] Figure 10 A schematic diagram illustrating the performance comparison of sleep stage recognition on (a) SHHS and (b) WSC datasets with different model sizes (63.5 million, 133.7 million, and 238.2 million parameters) according to embodiments of the present disclosure is shown. The results exhibit a clear scaling law: increasing the number of parameters continuously improves accuracy (Acc.), Cohen's Kappa coefficient (κ), macro F1 score (MF1), sensitivity (Sens.), and specificity (Spec.), which fully demonstrates the feasibility of effectively capturing complex sleep dynamics by extending the physiological basis model.
[0209] Figure 10This study demonstrates the performance of the sleep2vec model as the parameter scale changes. On two benchmark datasets, SHHS(a) and WSC(b), the model performance consistently improves as the number of parameters increases from 63.5 million to 238.2 million. This improvement is particularly significant in key metrics, including accuracy (Acc.), Cohen's Kappa (κ), macro F1 score (MF1), sensitivity (Sens.), and specificity (Spec.). Notably, the scaling effect exhibits diminishing marginal returns, indicating that while larger-scale models can capture more complex physiological patterns in sleep data, the gains from each increase in parameters gradually decrease. Overall, these results confirm the robustness and scalability of the sleep2vec architecture, demonstrating its suitability for capturing fine-grained multimodal physiological dynamics in sleep research.
[0210] A.4.5 Clinical Diagnostic Configuration
[0211] For clinical diagnostic tasks, the main model parameters were kept frozen, and no LoRA adapter was applied. Predictions were derived from the [CLS] token output of the final Transformer layer, then processed by a two-layer MLP classifier whose hidden layer dimensions matched the main model's output dimensions. Training was performed using the AdamW optimizer with a learning rate of [missing value]. The weight decay coefficient is This is consistent with the sleep staging experiment setup.
[0212] In addition to the methods described above, this disclosure also provides corresponding apparatus and equipment, which will be discussed in conjunction with the appendix below. Figures 11 to 13 This needs to be explained.
[0213] Figure 11 This is a block diagram illustrating an apparatus 1100 for training a neural network model according to an embodiment of the present disclosure. The description of method 100 above also applies to apparatus 1100 unless explicitly stated otherwise. The neural network model may include a first multilayer perceptron, a second multilayer perceptron, and a deep learning network based on a self-attention mechanism.
[0214] Reference Figure 11 The device 1100 shown may include: an acquisition module 1110, a selection module 1120, a first acquisition module 1130, a second acquisition module 1140, a third acquisition module 1150, a determination module 1160, and a parameter adjustment module 1170.
[0215] According to embodiments of this disclosure, the acquisition module 1110 can be configured to acquire various types of physiological signals.
[0216] According to an embodiment of this disclosure, the selection module 1120 can be configured to select a first type of physiological signal and a second type of physiological signal of a predetermined duration from the plurality of physiological signals of the plurality of types.
[0217] According to an embodiment of this disclosure, the first obtaining module 1130 can be configured to input the first type of physiological signals of the plurality of predetermined durations into a first multilayer perceptron to obtain a first feature vector of the predetermined dimension of the plurality of predetermined durations, and to input the second type of physiological signals of the plurality of predetermined durations into a second multilayer perceptron to obtain a second feature vector of the predetermined dimension of the plurality of predetermined durations.
[0218] According to an embodiment of this disclosure, the second obtaining module 1140 can be configured to input the first feature vectors of the predetermined dimensions of the plurality of predetermined durations into a deep learning network based on a self-attention mechanism to obtain a first fused feature vector, and input the second feature vectors of the predetermined dimensions of the plurality of predetermined durations into the deep learning network based on a self-attention mechanism to obtain a second fused feature vector.
[0219] According to an embodiment of this disclosure, the third obtaining module 1150 can be configured to map the first fused feature vector and the second fused feature vector to the same feature space to obtain a plurality of third feature vectors corresponding to the first fused feature vector and a plurality of fourth feature vectors corresponding to the second fused feature vector.
[0220] According to an embodiment of this disclosure, the determining module 1160 can be configured to determine a first similarity between each third feature vector and each fourth feature vector.
[0221] According to an embodiment of this disclosure, the parameter adjustment module 1170 can be configured to calculate the training loss of the neural network model based on the first similarity, and adjust the parameters of the first multilayer perceptron, the second multilayer perceptron, and the deep learning network based on the self-attention mechanism based on the training loss.
[0222] Figure 12 This is a block diagram illustrating an apparatus 1200 for predicting a first task related to a user in a sleep scenario, according to an embodiment of the present disclosure. The description of method 200 above also applies to apparatus 1200, unless otherwise explicitly stated.
[0223] Reference Figure 12 The device 1200 shown may include: an acquisition module 1210 and an acquisition module 1220.
[0224] According to embodiments of this disclosure, the acquisition module 1210 can be configured to acquire at least one physiological signal of a user in a sleep scenario.
[0225] According to an embodiment of this disclosure, the obtaining module 1220 can be configured to input at least one acquired physiological signal into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task.
[0226] Figure 13 This is a block diagram illustrating an electronic device 1300 according to an embodiment of the present disclosure. The descriptions of methods 100 and 200 above also apply to device 1300 unless otherwise explicitly stated.
[0227] See Figure 13 The device 1300 may include a processor 1301 and a memory 1302. Both the processor 1301 and the memory 1302 can be connected via a bus 1303.
[0228] Processor 1301 can perform various actions and processes according to the program stored in memory 1302. Specifically, processor 1301 can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on x86 architecture or ARM architecture.
[0229] Memory 1302 stores computer instructions that, when executed by processor 1301, implement the methods described above. Memory 1302 may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0230] According to yet another embodiment of this disclosure, a computer-readable recording medium is also provided. Figure 14 A schematic diagram of the recording medium 1400 according to this disclosure is shown.
[0231] like Figure 14 As shown, the recording medium 1400 stores computer-executable instructions 1410. When the computer-executable instructions 1410 are executed by a processor, they can cause the processor to perform the method described with reference to the above figures according to embodiments of the present disclosure. The computer-readable recording medium in the embodiments of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0232] This disclosure also provides a computer program product including computer-executable instructions. When executed by a processor, the computer-executable instructions cause the processor to perform any of the methods described above.
[0233] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0234] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0235] The exemplary embodiments of the present invention described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the invention, and such modifications should fall within the scope of the invention.
Claims
1. A method for training a neural network model, characterized in that, The neural network model includes a first multilayer perceptron, a second multilayer perceptron, and a deep learning network based on a self-attention mechanism; the method includes: Acquire multiple types of physiological signals; Select a first type of physiological signal and a second type of physiological signal of a predetermined duration from the multiple types of physiological signals; The first type of physiological signals of the plurality of predetermined durations are input into a first multilayer perceptron to obtain a first feature vector of the predetermined dimension of the plurality of predetermined durations, and the second type of physiological signals of the plurality of predetermined durations are input into a second multilayer perceptron to obtain a second feature vector of the predetermined dimension of the plurality of predetermined durations. The first feature vectors of the predetermined dimensions with the multiple predetermined durations are input into the deep learning network based on the self-attention mechanism to obtain a first fused feature vector, and the second feature vectors of the predetermined dimensions with the multiple predetermined durations are input into the deep learning network based on the self-attention mechanism to obtain a second fused feature vector. The first fused feature vector and the second fused feature vector are mapped to the same feature space to obtain multiple third feature vectors corresponding to the first fused feature vector and multiple fourth feature vectors corresponding to the second fused feature vector; Determine the first similarity between each third feature vector and each fourth feature vector; and Based on the first similarity, the training loss of the neural network model is calculated, and the parameters of the first multilayer perceptron, the second multilayer perceptron, and the deep learning network based on the self-attention mechanism are adjusted based on the training loss.
2. The method according to claim 1, characterized in that, The step of calculating the training loss of the neural network model based on the first similarity includes: Determine the second similarity between feature vector pairs consisting of third and fourth feature vectors that satisfy predetermined conditions, wherein the predetermined conditions include the third and fourth feature vectors in the feature vector pair being obtained based on different types of physiological signals of the same object for the same predetermined duration; The training loss of the neural network model is calculated based on the first similarity and the second similarity.
3. The method according to claim 2, characterized in that, The step of calculating the training loss of the neural network model based on the first similarity and the second similarity includes: Based on the first similarity, a first weight corresponding to the first similarity is determined, wherein the greater the first similarity, the greater the value of the first weight corresponding to the first similarity; and The training loss of the neural network model is calculated based on the first similarity, the first weight, and the second similarity.
4. The method according to claim 3, characterized in that, Based on the first similarity, the first weight, and the second similarity, the training loss of the neural network model is calculated, including: A first similarity adjustment value is determined based on the first similarity and the value of the piecewise constant function corresponding to the first similarity; and The training loss of the neural network model is calculated based on the product of the first similarity adjustment value and the first weight, and the second similarity. Wherein, when the third feature vector and the fourth feature vector corresponding to the first similarity are obtained based on different types of physiological signals of the same object, the value of the piecewise constant function corresponding to the first similarity is a predetermined constant; otherwise, the value of the piecewise constant function corresponding to the first similarity is zero.
5. The method according to any one of claims 1 to 4, characterized in that, The sampling rate of the first type of physiological signal is different from that of the second type of physiological signal.
6. The method according to any one of claims 1 to 4, characterized in that, The various types of physiological signals include two or more of the following: electroencephalogram (EEG) physiological signals, electromyogram (EMG) physiological signals, electrooculogram (EOG) physiological signals, electrocardiogram (ECG) physiological signals, intercardiac interval (IBI) physiological signals, nasal airflow physiological signals, abdominal breathing band / thoracic breathing band (ABD / Thor Belt) physiological signals, respiratory movement physiological signals, and blood oxygen saturation (SpO2) physiological signals.
7. A method for predicting a user-related first task in a sleep context, characterized in that, The method includes: Acquire at least one physiological signal from the user during sleep; and The acquired physiological signal is input into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task. The first prediction model is obtained by fine-tuning the neural network model obtained by the method according to any one of claims 1 to 6.
8. The method according to claim 7, characterized in that, The first prediction model is obtained by fine-tuning the acquired neural network model through the following operations: Obtain the actual value corresponding to the first task; A third multilayer perceptron corresponding to the first task is set after the deep learning network based on the self-attention mechanism in the obtained neural network model. Based on the at least one physiological signal, the first multilayer perceptron, the second multilayer perceptron, the deep learning network based on the self-attention mechanism, and the third multilayer perceptron in the obtained neural network model are used to obtain the predicted value corresponding to the first task. Based on the difference between the predicted value corresponding to the first task and the actual value of the first task, the parameters of the first multilayer perceptron, the second multilayer perceptron, and the third multilayer perceptron corresponding to the first task are adjusted to obtain the first prediction model.
9. The method according to claim 7 or 8, characterized in that, The first task is any one of the following tasks: Tasks related to blood oxygenation prediction, sleep staging, apnea-hypopnea index estimation, sleep apnea detection, demographics, heart failure, and gastroesophageal reflux.
10. The method according to claim 7 or 8, characterized in that, The acquisition of at least one physiological signal of the user in a sleep scenario includes: acquiring the at least one physiological signal through relevant sensor data, wherein the relevant sensor data includes data obtained based on at least one of the following sensors: piezoelectric sensor, piezoresistive sensor, fiber optic sensor, audio sensor, heart rate sensor, nasal pressure sensor, motion sensor, and accelerometer.
11. An apparatus for training a neural network model, characterized in that, The neural network model includes a first multilayer perceptron, a second multilayer perceptron, and a deep learning network based on a self-attention mechanism; the device includes: The acquisition module is configured to acquire multiple types of physiological signals; The selection module is configured to select a first type of physiological signal and a second type of physiological signal of a predetermined duration from the multiple types of physiological signals; The first obtaining module is configured to input the first type of physiological signals of the plurality of predetermined durations into a first multilayer perceptron to obtain a first feature vector of the predetermined dimension of the plurality of predetermined durations, and to input the second type of physiological signals of the plurality of predetermined durations into a second multilayer perceptron to obtain a second feature vector of the predetermined dimension of the plurality of predetermined durations. The second obtaining module is configured to input the first feature vectors of the predetermined dimensions of the plurality of predetermined durations into the deep learning network based on the self-attention mechanism to obtain a first fused feature vector, and to input the second feature vectors of the predetermined dimensions of the plurality of predetermined durations into the deep learning network based on the self-attention mechanism to obtain a second fused feature vector. The third obtaining module is configured to map the first fused feature vector and the second fused feature vector to the same feature space to obtain a plurality of third feature vectors corresponding to the first fused feature vector and a plurality of fourth feature vectors corresponding to the second fused feature vector. The determination module is configured to determine a first similarity between each third feature vector and each fourth feature vector; and The parameter adjustment module is configured to calculate the training loss of the neural network model based on the first similarity, and adjust the parameters of the first multilayer perceptron, the second multilayer perceptron, and the deep learning network based on the self-attention mechanism based on the training loss.
12. An apparatus for predicting a first task related to a user in a sleep context, characterized in that, The device includes: The acquisition module is configured to acquire at least one physiological signal of the user in a sleep scenario; and The acquisition module is configured to input at least one acquired physiological signal into a first prediction model based on a neural network model to obtain a prediction result corresponding to the first task; The first prediction model is obtained by fine-tuning the neural network model obtained by the method according to any one of claims 1 to 6.
13. An electronic device, comprising: processor, and A memory storing computer-executable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-10.
14. A computer-readable recording medium storing computer-executable instructions, wherein, The computer-executable instructions, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-10.
15. A computer program product comprising computer-executable instructions, wherein, The computer-executable instructions, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-10.