Physiological signal prediction model training method, electronic equipment and product
By extracting and fusing the normal inverse gamma distribution features of multimodal data and calculating the training loss, the problem of low accuracy and reliability of physiological signal prediction models in existing technologies is solved, and higher accuracy physiological signal prediction is achieved.
Patent Information
- Application Number
- CN202511117596.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-01-16
AI Technical Summary
Existing technologies, when training physiological signal prediction models with multimodal data, focus only on the simple combination of feature values during the fusion process. This results in an inability to dynamically adapt to changes in the scene, leading to a decrease in the overall prediction accuracy of the model when a certain modality is disturbed.
Multimodal data is acquired in a non-contact manner. The first normal inverse gamma distribution feature of each modality is extracted, including predicted physiological signals, confidence features and variance distribution features. The second normal inverse gamma distribution feature is generated by fusing these features, and the feature training loss is calculated to iteratively update the model parameters.
It improves the prediction accuracy of physiological signal prediction models, avoids interference from noisy modes, enhances the model's focus on high-quality modal information, optimizes the multimodal collaboration mechanism, and reduces information loss or interference.
Smart Images

Figure CN121350554A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a training method, electronic device and product for a physiological signal prediction model. Background Technology
[0002] With the rapid development of the Internet of Things, wearable devices and artificial intelligence technologies, remote physiological signal measurement technology based on deep learning has received increasing attention.
[0003] Specifically, the model training process typically involves: acquiring multimodal data and corresponding real physiological signals using a non-contact method to construct a training dataset. Then, during model training, the initial model extracts features from each modality individually, obtaining the basic features corresponding to each modality. Next, the initial model directly concatenates the basic features from different modalities, or fuses them using a simple weighted summation, resulting in a comprehensive feature vector. Finally, physiological signal prediction is performed based on the fused comprehensive feature vector, and the training loss is calculated using the error between the predicted values and the real physiological signals. The model parameters are then iteratively updated using backpropagation to complete the training.
[0004] However, when training models using multimodal data, the fusion process focuses only on a simple combination of feature values, which can easily lead to the fused features failing to dynamically adapt to changes in the scene. Consequently, when a particular modality is disturbed, the overall prediction accuracy of the model can easily decrease. In other words, the accuracy and reliability of physiological signal prediction models trained using existing technologies are relatively low. Summary of the Invention
[0005] This application provides a training method, electronic device, and product for a physiological signal prediction model, which can solve the problem of low accuracy and reliability of physiological signal prediction models trained in the prior art.
[0006] In a first aspect, embodiments of this application provide a method for training a physiological signal prediction model, the method comprising:
[0007] Obtain a training dataset containing multiple training data points; the training data is acquired in a non-contact manner, and each training data point includes multiple modal data and corresponding real physiological signals;
[0008] For any training data, multiple modal data are input into the initial physiological signal prediction model, and the first normal inverse gamma distribution feature of each modal data is extracted. The first normal inverse gamma distribution feature includes the first predicted physiological feature for predicting physiological signals, the first confidence feature characterizing the reliability of the predicted physiological signals, the first distribution feature characterizing the variance distribution of the predicted physiological signals, and the first scale feature characterizing the variance scale of the predicted physiological signals.
[0009] By fusing the features of each first normal inverse gamma distribution, the features of the second normal inverse gamma distribution are obtained;
[0010] The feature training loss of each target's normal inverse gamma distribution is calculated based on real physiological signals; targets are divided into first and second categories.
[0011] The model parameters of the initial physiological signal prediction model are iteratively updated based on the training loss of all features to obtain the physiological signal prediction model.
[0012] In one embodiment, the modal data includes radar data; multiple modal data are input into an initial physiological signal prediction model, and the first normal inverse gamma distribution feature of each modal data is extracted, including:
[0013] A Fourier transform is performed on the radar data to generate a range matrix; each matrix value in the range matrix is used to describe the signal strength of the radar signal reflected by an object in space;
[0014] Based on the distance matrix, target radar data that meets preset conditions are identified from the radar data; the preset conditions include conditions for determining the category of the object as human.
[0015] The target radar data is input into the initial physiological signal prediction model to obtain the first normal inverse gamma distribution feature.
[0016] In one embodiment, based on the range matrix, target radar data that meets preset conditions is determined from the radar data, including:
[0017] Determine the target position of the object corresponding to the largest matrix value from the distance matrix;
[0018] Radar data from a preset range of the target location is identified as target radar data that meets preset conditions.
[0019] In one embodiment, fusing each first normal inverse gamma distribution feature to obtain a second normal inverse gamma distribution feature includes:
[0020] Based on each first confidence feature, a weighted sum is performed on each first predicted physiological feature to obtain the second predicted physiological feature;
[0021] Each first confidence feature is summed to obtain the second confidence feature;
[0022] Sum the first distribution features to obtain the second distribution features;
[0023] Calculate the feature difference between each first predicted physiological feature and each second predicted physiological feature;
[0024] The feature difference is adjusted based on the first confidence feature corresponding to each feature difference to obtain the target feature difference; the target feature difference is used to describe the uncertainty of the second predicted physiological feature.
[0025] The difference between each first-scale feature and the target feature is summed to obtain the second-scale feature;
[0026] The second predictive physiological feature, the second confidence feature, the second distribution feature, and the second scale feature are determined to be the second normal inverse gamma distribution feature.
[0027] In one embodiment, the feature training loss for each target's normal inverse gamma distribution features is calculated based on real physiological signals, including:
[0028] For any target with a normal inverse gamma distribution, calculate the similarity loss between the predicted physiological features of the target and the actual physiological signals;
[0029] Calculate the signal-to-noise ratio loss in the frequency domain between the predicted physiological characteristics of the target and the actual physiological signal;
[0030] Based on the target's predicted physiological characteristics, the actual physiological signal, the target's confidence level, the target's distribution characteristics, and the target's scale characteristics, the uncertainty loss of the initial physiological signal prediction model is calculated.
[0031] Based on similarity loss, signal-to-noise ratio loss, and uncertainty loss, the feature training loss corresponding to the target normal inverse gamma distribution is calculated.
[0032] In one embodiment, calculating the signal-to-noise ratio loss in the frequency domain between the predicted physiological features of the target and the actual physiological signals includes:
[0033] Determine the effective frequency corresponding to the maximum energy value of the real physiological signal in the frequency domain;
[0034] The frequency range is determined based on the effective frequency and the preset window.
[0035] Fourier transform is performed on the target predicted physiological characteristics to obtain the frequency domain predicted physiological parameters;
[0036] Determine the first predicted physiological parameter that is within the frequency range and the second predicted physiological parameter that is not within the frequency range from the frequency domain predicted physiological parameters;
[0037] The signal-to-noise ratio loss is calculated based on the first and second predicted physiological parameters.
[0038] In one embodiment, based on the predicted physiological characteristics of the target, the actual physiological signal, the target confidence level, the target distribution characteristics, and the target scale characteristics, the uncertainty loss of the initial physiological signal prediction model is calculated, including:
[0039] Calculate the prediction difference between the target predicted physiological characteristics and the actual physiological signals;
[0040] Based on the prediction difference, target confidence, target distribution characteristics, and target scale characteristics, the negative likelihood loss is calculated; the negative likelihood loss is used to describe the degree of uncertainty in the prediction of the initial physiological signal prediction model.
[0041] Based on the prediction difference, target confidence, and target distribution characteristics, the error loss is calculated; the error loss is used to describe the degree of error between the predicted physiological characteristics of the target and the actual physiological signal.
[0042] Based on negative likelihood loss and error loss, uncertainty loss is obtained.
[0043] In one embodiment, the modal data includes radar data; the model parameters of the initial physiological signal prediction model are iteratively updated based on the training loss of all features to obtain the physiological signal prediction model, including:
[0044] The feature training loss is obtained by weighted summation of all feature training losses; the weight of the feature training loss corresponding to radar data is less than the weight of the feature training loss corresponding to other modal data.
[0045] The physiological signal prediction model is obtained by iteratively updating the model parameters based on feature training loss.
[0046] Secondly, embodiments of this application provide a training device for a physiological signal prediction model, the device comprising:
[0047] The acquisition module is used to acquire a training dataset containing multiple training data points. The training data is acquired in a non-contact manner, and each training data point includes multiple modal data and corresponding real physiological signals.
[0048] The input module is used to input multiple modal data into the initial physiological signal prediction model for any training data, and extract the first normal inverse gamma distribution feature of each modal data respectively; the first normal inverse gamma distribution feature includes the first predicted physiological feature for predicting physiological signals, the first confidence feature characterizing the reliability of the predicted physiological signals, the first distribution feature characterizing the variance distribution of the predicted physiological signals, and the first scale feature characterizing the variance scale of the predicted physiological signals.
[0049] The fusion module is used to fuse each first normal inverse gamma distribution feature to obtain the second normal inverse gamma distribution feature;
[0050] The computation module is used to calculate the feature training loss of the normal inverse gamma distribution features of each target based on real physiological signals; the targets are divided into first and second targets.
[0051] The iterative module is used to iteratively update the model parameters of the initial physiological signal prediction model based on the training loss of all features, so as to obtain the physiological signal prediction model.
[0052] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0053] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect above.
[0054] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the method described in the first aspect.
[0055] The beneficial effects of this application embodiment compared with the prior art are as follows: It acquires training data containing multiple modalities and corresponding real physiological signals in a non-contact manner, retaining the convenience of non-contact measurement while covering physiological signal characteristics in different scenarios through multi-modal data. This provides a more comprehensive learning foundation for the initial physiological signal prediction model and avoids the scenario limitations of single-modal data. Subsequently, for each modal data, a first normal inverse gamma distribution feature is extracted, including a first predicted physiological feature, a first confidence feature, a first distribution feature, and a first scale feature. This breaks the limitation of traditional techniques that only extract numerical signal features. The first confidence feature quantifies the reliability of each modal prediction, while the first distribution feature and the first scale feature describe the uncertainty of signal variance distribution and scale. Therefore, based on the above feature information, the initial physiological signal prediction model can accurately identify the quality of the modal data. That is, the electronic device can fuse the first normal inverse gamma distribution feature of each modal data based on the confidence feature to obtain a second normal inverse gamma distribution feature, and variance distribution information can be integrated during the fusion process to suppress interference from noisy modalities. Furthermore, based on the aforementioned feature fusion method, the problem of noisy modes dragging down overall performance in traditional splicing can be solved, making the fused second normal inverse gamma distribution feature more focused on the effective information of high-quality modes. Finally, the feature training loss of each target (first and second) normal inverse gamma distribution feature is calculated separately, and the model parameters of the initial physiological signal prediction model are iteratively updated to obtain the physiological signal prediction model. By calculating the feature training loss for the first and second normal inverse gamma distribution features separately, the prediction accuracy of single-modal data features can be guaranteed (through the feature training loss calculated by the first normal inverse gamma feature), while also strengthening the consistency between the fused features and the real signal (through the feature training loss calculated by the second normal inverse gamma feature). Furthermore, through the dual training loss calculation, the model can optimize the multimodal collaborative mechanism while learning single-modal characteristics, avoiding information loss or interference during the fusion process. In this way, the prediction accuracy of the trained physiological signal prediction model can be improved. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart illustrating the implementation of a training method for a physiological signal prediction model according to an embodiment of this application.
[0058] Figure 2This is a schematic diagram illustrating an implementation method for generating the first normal inverse gamma distribution feature in a training method for a physiological signal prediction model provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram illustrating an implementation method for generating a second normal inverse gamma distribution feature in a training method for a physiological signal prediction model provided in an embodiment of this application;
[0060] Figure 4 This is a schematic diagram illustrating one implementation method of generating a physiological signal prediction model in a training method for a physiological signal prediction model provided in an embodiment of this application;
[0061] Figure 5 This is a schematic diagram of the structure of a physiological signal prediction model provided in an embodiment of this application;
[0062] Figure 6 is a schematic diagram of the results of predicting physiological signals in different application scenarios using different methods, according to an embodiment of this application.
[0063] Figure 7 This is a schematic diagram of the structure of a training device for a physiological signal prediction model provided in one embodiment of this application;
[0064] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0065] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0066] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0067] It should be noted that the information collection process (such as patient information collection process, physiological information collection process, etc.) / feature extraction process involved in this application is carried out with the user's knowledge and permission. That is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0068] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0069] With the rapid development of the Internet of Things, wearable devices and artificial intelligence technologies, remote physiological signal measurement technology based on deep learning has received increasing attention.
[0070] Specifically, the model training process typically involves: acquiring multimodal data and corresponding real physiological signals using a non-contact method to construct a training dataset. Then, during model training, the initial model extracts features from each modality individually, obtaining the basic features corresponding to each modality. Next, the initial model directly concatenates the basic features from different modalities, or fuses them using a simple weighted summation, resulting in a comprehensive feature vector. Finally, physiological signal prediction is performed based on the fused comprehensive feature vector, and the training loss is calculated using the error between the predicted values and the real physiological signals. The model parameters are then iteratively updated using backpropagation to complete the training.
[0071] However, when training models using multimodal data, the fusion process focuses only on a simple combination of feature values, which can easily lead to the fused features failing to dynamically adapt to changes in the scene. Consequently, when a particular modality is disturbed, the overall prediction accuracy of the model can easily decrease. In other words, the accuracy and reliability of physiological signal prediction models trained using existing technologies are relatively low.
[0072] Based on this, in order to improve the accuracy of the trained physiological signal prediction model, one embodiment of this application provides a training method for the physiological signal prediction model. This method can be applied to electronic devices such as tablet computers, laptops, ultra-mobile personal computers (UMPCs), and netbooks. This embodiment of the application does not impose any restrictions on the specific type of electronic device.
[0073] Please see Figure 1 , Figure 1 The following is a flowchart illustrating the implementation of a training method for a physiological signal prediction model provided in an embodiment of this application. The method includes the following steps:
[0074] S101. Obtain the training dataset containing multiple training data.
[0075] The training data was acquired in a non-contact manner, and each training data set included multiple modal data and corresponding real physiological signals.
[0076] In one embodiment, the training dataset is a set of samples used to train the physiological signal prediction model, consisting of multiple independent training data sets. Through a large amount of training data, the initial physiological signal prediction model can learn the correlation between various modalities of data and real physiological signals, thereby acquiring the ability to predict new data.
[0077] The training data described above consists of a single sample in the training dataset, serving as the basic unit for model learning. Each training data set contains two parts: input and output labels. The input consists of multiple modalities, while the output label represents the corresponding real physiological signal, forming a one-to-one mapping relationship between the two.
[0078] The aforementioned non-contact methods refer to data acquisition without direct physical contact with the human body, unlike traditional contact measurements (such as electrocardiograms which require attached electrodes). Non-contact methods avoid restricting the human body and can be applied to long-term health monitoring (e.g., home-based elderly care, chronic disease management), monitoring of special populations (e.g., infants, burn patients, psychiatric patients), monitoring of exercise and stress states (e.g., athlete training, virtual reality / augmented reality interaction), and public health scenarios (e.g., large-scale population health screening), without limitation.
[0079] In the context of modal data, "modality" refers to the source or form of data representation. Multiple modal data refers to data related to physiological signals acquired from different channels or through different principles. In remote physiological signal measurement, common modal data include video data and radio frequency (RF) modal data, which are not limited to these.
[0080] Among them, video data can be considered as a sequence of human images captured by a visible light camera or a near-infrared camera, containing information such as subtle color changes on the face caused by the pulse.
[0081] Radio frequency modal data can be considered as reflected signals transmitted and received by radio frequency sensors (e.g., millimeter-wave radar), containing information about the body's micro-movements caused by heartbeat and breathing.
[0082] It should be noted that multimodal data can be divided into at least two of the following: human image sequences captured by visible light cameras, human image sequences captured by near-infrared cameras, and radar data acquired by millimeter-wave radar. Even data consisting of only two modalities—human image sequences captured by visible light cameras and human image sequences captured by near-infrared cameras—can be considered multimodal data.
[0083] The aforementioned real physiological signals refer to actual physiological indicators of the human body, which serve as the target benchmark for model prediction, such as heart rate (heart rate per minute), respiratory rate (breathing rate per minute), and blood oxygen saturation. These real physiological signals can be synchronously collected using high-precision contact devices (such as medical electrocardiographs and pulse oximeters) to ensure accuracy.
[0084] In one embodiment, the training data described above can be collected non-contactly. For example, a visible light camera and a millimeter-wave radar can be deployed, placing the subject within the monitoring range of the devices (without physical contact). The visible light camera can continuously capture images of the subject's face or chest area, generating video data; the millimeter-wave radar can emit radar signals and receive signals reflected from the human body, generating radar data. The two devices operate synchronously, collecting data within the same time period, ensuring temporal consistency of the modal data. Furthermore, while collecting modal data, the subject can wear a high-precision contact-type physiological monitoring device (such as an ECG chest patch or a finger-clip pulse oximeter) to record real physiological signals such as heart rate and respiratory rate within the same time period, serving as tags for the training data.
[0085] Then, the electronic device can use video data and radar data collected within the same time period, along with the corresponding real physiological signals, as training data. This process is repeated to collect samples from different populations (age, gender, health status, etc.) under different scenarios (resting, mild activity, different lighting conditions, etc.), combining them into a training dataset containing multiple training data sets.
[0086] S102. For any training data, input multiple modal data into the initial physiological signal prediction model, and extract the first normal inverse gamma distribution feature of each modal data.
[0087] In one embodiment, the initial physiological signal prediction model described above is an untrained neural network architecture, typically consisting of a feature extraction layer, a fusion layer, and an output layer. The model parameters (such as the weights, biases, and hyperparameters of the probability distribution of the neural network) of the feature extraction layer, fusion layer, and output layer are usually randomly initialized parameters, providing only a framework for predicting physiological signals but lacking effective predictive capabilities.
[0088] For multimodal data, the initial physiological signal prediction model can include independent branches for each modality (such as video branch, radar branch) to extract features of a specific modality.
[0089] The first normal inverse gamma distribution feature mentioned above is a set of statistical features obtained after probabilistic modeling of each modality data, based on the normal-inverse-gamma distribution. This distribution is a commonly used conjugate prior distribution in Bayesian statistics, suitable for describing data with unknown mean and variance.
[0090] Among them, the first normal inverse gamma distribution feature mentioned above includes the first predictive physiological feature for predicting physiological signals, the first confidence feature for characterizing the reliability of predicting physiological signals, the first distribution feature for characterizing the variance distribution of predicting physiological signals, and the first scale feature for characterizing the variance scale of predicting physiological signals.
[0091] In one embodiment, the first predicted physiological feature is the direct predicted value of the physiological signal (such as heart rate and respiratory rate) by the initial physiological signal prediction model, which corresponds to the mean parameter in the first normal inverse gamma distribution and reflects the estimation of the real physiological signal by the modality data.
[0092] The aforementioned first confidence feature quantifies the reliability of the first predicted physiological feature and is typically related to the precision parameter of the distribution (the inverse of the variance). A higher confidence level corresponding to the first confidence feature indicates that the initial physiological signal prediction model believes the predicted first physiological feature is more accurate; conversely, a lower confidence level indicates that the predicted first physiological feature may have a larger error.
[0093] The aforementioned first distribution feature describes the variance distribution of the first predicted physiological characteristic, corresponding to the shape parameter (α) in the first normal inverse gamma distribution. This parameter determines the probability distribution shape of the variance, reflecting the range of uncertainty in the prediction.
[0094] The aforementioned first-scale feature is used to directly characterize the scale of variance, i.e., the degree of fluctuation of the first predicted physiological characteristic. A larger scale indicates that the first predicted physiological characteristic may fluctuate over a wider range, with higher uncertainty.
[0095] In one embodiment, the initial physiological signal prediction model may include a feature extraction network. This feature extraction network includes feature extraction branches for multiple modality data to extract a first normal inverse gamma distribution feature corresponding to each modality data.
[0096] Each feature extraction branch can include four channels, each containing one or more convolutional layers and fully connected layers. The last layer is used to output the parameters corresponding to the first normal inverse gamma distribution. That is, the four channels include a channel that outputs the first predicted physiological feature (μ), a passband that outputs the first confidence feature (v), a channel that outputs the first distribution feature (α), and a channel that outputs the first scale feature (β).
[0097] Based on the above explanation, the characteristics of the first normal inverse gamma distribution can be represented as follows:
[0098] NIG(xi)=[ui, vi, αi, βi];
[0099] Where xi represents the i-th modal data, NIG(xi) represents the first normal inverse gamma distribution feature of the i-th modal data, ui represents the first predicted physiological feature, vi represents the first confidence feature, αi represents the first distribution feature of the variance distribution of the predicted physiological signal, and βi represents the first scale feature of the variance scale of the predicted physiological signal.
[0100] It should be noted that after obtaining the four original parameters corresponding to the output of each channel, the output original parameters can be processed through a nonlinear transformation (e.g., an exponential function to ensure the parameters are positive) to map them into effective parameters that conform to the constraints of a normal inverse gamma distribution. At this time, the four effective parameters are the first normal inverse gamma distribution characteristics mentioned above.
[0101] The electronic device can perform the aforementioned nonlinear transformation processing using a preset activation function. For example, the activation function includes, but is not limited to, Softplus, Sigmoid, and ReLU activation functions.
[0102] It should be noted that inputting the original parameters into the activation function to obtain the first normal inverse gamma distribution feature can enhance the nonlinear transformation capability of each feature, making the first normal inverse gamma distribution feature satisfy the physical meaning, probabilistic properties or model constraints of specific parameters, and ensuring the rationality and effectiveness of the model output.
[0103] For example, in a normal inverse gamma distribution, the parameters characterizing variance (such as scale features and distribution features) must be positive (variance cannot be negative); in a probability model, parameters such as weights and standard deviations must be non-negative to ensure the validity of the probability distribution; in counting tasks (such as predicting the number of people or frequency), the output results must be non-negative to be meaningful in practice.
[0104] In this embodiment, for the physiological signal prediction scenario, the constraints of each parameter can be set by the user or generated based on the training data in the training dataset.
[0105] For example, let the inputs of video data and radar data be X, respectively. V and X R The corresponding real physiological signal is y. Based on the nature of the regression task, the target y can be viewed as a data distribution that follows a Gaussian distribution, i.e.:
[0106] y~N(μ,σ 2 );
[0107] Where μ represents the mean of all real physiological signals in the entire training dataset, and σ 2 This represents the variance of each real physiological signal in the entire training dataset. Furthermore, the parameters (μ, σ) 2This can be modeled as Gaussian prior and inverse gamma prior:
[0108] μ~Ν(γ,σ 2 υ -1 ),σ 2 ~Γ -1 9α,β);
[0109] Here, Γ represents the gamma function, γ∈R, σ>0, α>1, β>0. In this case, γ∈R, σ>0, α>1, β>0 can be considered as a constraint condition for each parameter. Here, R represents the set of real numbers.
[0110] Based on the above explanation, it can be considered that the objective of the regression task is to estimate the mean μ and variance σ. 2 That is, it can be modeled as a posterior distribution q(μ,σ). 2 )=p(μ,σ 2 |y). To simplify the calculation, the posterior distribution is factored into two independent distributions q(μ) and q(σ). Based on the Gaussian prior and inverse gamma prior mentioned above, the posterior distribution q(μ,σ) is... 2 This can be represented as a normal inverse gamma distribution NIG(γ,υ,α,β):
[0111]
[0112] Based on the above explanation, it can be considered that probabilistic modeling based on the training dataset can more comprehensively and accurately describe the distribution characteristics of physiological signals. Reasonably setting prior information can constrain the first normal inverse gamma distribution characteristics of each feature extraction branch output, so that the output features meet the essential requirements of the regression task and simplify the subsequent model calculation process.
[0113] In another embodiment, when multiple modalities of data are input into an initial physiological signal prediction model to extract the first normal inverse gamma distribution features of each modality, the structure of each modality of data may not conform to the input of the model (e.g., the feature extractor). Therefore, the initial physiological signal prediction model also needs to perform preprocessing operations on each modality of data to make it meet the requirements.
[0114] For example, for video data, the electronic device can pre-apply a preset face recognition model to perform face detection and alignment on each frame of the face image in the video data, and extract regions of interest (e.g., forehead, cheeks) to enhance blood volume change signals through color space conversion (e.g., RGB to YCbCr). Furthermore, the electronic device can also perform temporal filtering on the processed video data to remove low-frequency noise and convert it into a face image of a specific size.
[0115] As a specific example, electronic devices can use MTCNN (Multi-Task Cascaded Convolutional Networks) to perform face recognition on each frame of video data, crop the image based on the identified region of interest, and finally transform it into a 128x128 sized face image.
[0116] For radar data, electronic devices can perform range-Doppler processing on the radar data (reflected signals), locate the micro-movement area of the human body, and then extract the frequency components related to physiological signals (e.g., heartbeat, breathing), suppress environmental clutter, and obtain the target radar data after noise removal.
[0117] As a specific example, electronic devices can be based on, for example... Figure 2 Steps S201-S203, as shown, yield the first normal inverse gamma distribution characteristic corresponding to the radar data. Details are as follows:
[0118] S201. Perform a Fourier transform on the radar data to generate a range matrix.
[0119] Each matrix value in the aforementioned distance matrix describes the signal strength of radar signals reflected by objects in space.
[0120] In one embodiment, the aforementioned radar data refers to the raw data collected by the radar equipment during operation. The radar equipment can detect targets by emitting electromagnetic waves and receiving the reflected electromagnetic waves (reflected signals). The raw data set formed after preliminary processing of the reflected electromagnetic wave signals is the radar data. Typically, radar data contains various information about the target, such as the target's distance, speed, and angle, which will not be described in detail here.
[0121] The Fourier transform described above is a mathematical transformation method that converts a signal from the time domain to the frequency domain. In signal processing, time-domain signals describe how a signal changes over time, while frequency-domain signals describe the distribution of different frequency components within the signal. Through the Fourier transform, the frequency characteristics of a signal can be analyzed, such as the frequency components contained in the signal and their corresponding intensities.
[0122] The distance matrix described above is a data structure generated after performing a Fourier transform on radar data. It is typically a two-dimensional matrix, with each element (matrix value) corresponding to a specific location in space. This is usually achieved by dividing the radar data into grid points based on the radar's detection range and resolution, with each grid point representing a location. Each matrix value represents the signal strength of the radar signal reflected by an object at that location. Based on this, the distance matrix allows for a direct visual determination of the distribution of objects at different locations within the radar's detection range and their ability to reflect radar signals. A higher signal strength generally indicates that the object at that location is more easily detected by the radar, or that the object has stronger reflective properties.
[0123] The objects within the aforementioned space refer to all physical objects existing within the radar's detection range. These objects can include people, animals, or static obstacles, and there is no limitation on this.
[0124] The aforementioned reflected radar signal refers to the electromagnetic wave signal reflected back by the surface of an object after it has encountered an electromagnetic wave emitted by radar. When the electromagnetic waves emitted by radar strike an object, the object's surface reflects a portion of the electromagnetic wave energy back to the radar receiving antenna. The intensity and characteristics of the reflected signal depend on factors such as the object's shape, size, material, surface roughness, and relative position to the radar. By analyzing the reflected radar signal, information such as the object's type, location, and motion state can be determined.
[0125] The aforementioned signal strength refers to the power or amplitude of the reflected radar signal, reflecting an object's ability to reflect electromagnetic waves. Generally, the larger the object, the rougher its surface, and the closer it is to the radar, the stronger the reflected signal. Signal strength can be one of the parameters used by radar to detect and identify targets; by analyzing signal strength, different types of objects can be distinguished.
[0126] S202. Based on the range matrix, determine the target radar data that meets the preset conditions in the radar data.
[0127] Among them, the preset conditions include conditions for determining whether an object belongs to the human category.
[0128] In one embodiment, each matrix value in the distance matrix described above is used to describe the signal strength of radar signals reflected by an object in space, and it is explained that the signal strength refers to the power or amplitude of the reflected radar signal, which can reflect the object's ability to reflect electromagnetic waves.
[0129] Based on this, satisfying the above preset conditions can be considered as meeting the preset conditions when the matrix values in the distance matrix are within the preset matrix range. The preset matrix range can be set in advance and is not limited thereto.
[0130] It is important to note that in physiological signal prediction scenarios, objects in the space typically consist only of the human body and various static medical devices. Because the human body has stronger dynamic reflection characteristics than static medical devices (such as changes in reflected signals caused by breathing and slight limb movements), its corresponding matrix value is often significantly larger than that of static devices. Based on this characteristic, a preset condition can be determined when the matrix value reaches its maximum value. That is, once the maximum matrix value in the distance matrix is determined, the object category corresponding to that value can be identified as the human body, and its corresponding radar data is the target radar data.
[0131] In this embodiment, the method for determining whether the preset conditions are met is not limited.
[0132] As a specific example, an electronic device can determine the target location of the object corresponding to the largest matrix value from the distance matrix, and determine the radar data from the target location within a preset range as target radar data that meets preset conditions.
[0133] In one embodiment, the maximum matrix value is the maximum value of all matrix values in the distance matrix, which typically corresponds to the object with the strongest reflectivity in space (e.g., the human body).
[0134] The target location mentioned above is the spatial coordinate (distance and angle) of the maximum matrix value in the distance matrix, representing the spatial location of the object with the strongest reflectivity.
[0135] The aforementioned preset range can be considered as a pre-defined spatial area centered on the target location, used to capture relevant radar data around the target.
[0136] The aforementioned preset range can be set according to actual conditions and is not limited thereto. For example, since radar data of the human body needs to be collected, the preset range can be set to a spatial range according to the application scenario and target characteristics (e.g., human body size). For example, the preset range may include at least one of the following: distance range: ±0.5 meters (covering the thickness of the human body); angle range: ±5° (covering the horizontal width of the human body). That is, the preset range can include not only a distance range but also an angle range.
[0137] The aforementioned target radar data consists of radar signals from a preset range of the target location, containing the complete reflection characteristics of the target object (such as a human body), which are used for subsequent physiological signal prediction.
[0138] For example, after determining the target location, the electronic device can filter out all signal points within a preset range (e.g., ±0.5 meters, ±5° angle) of the target location from the original radar data to form the aforementioned target radar data, in order to avoid radar data generated by additional interference and retain the main physiological information.
[0139] In this embodiment, by directly locating the target with the strongest reflectivity (largest matrix value), interference from static backgrounds (such as medical equipment) can be effectively eliminated, improving the quality of the target radar data. Furthermore, extracting target radar data only from the local area (preset range) where the human body is located can further reduce the mixing of irrelevant signals, enabling subsequent models to more accurately capture physiological characteristics such as breathing and heartbeat.
[0140] S203. Input the target radar data into the initial physiological signal prediction model to obtain the first normal inverse gamma distribution feature.
[0141] In one embodiment, the electronic device may input only the target radar data into an initial physiological signal prediction model (e.g., a feature extractor corresponding to the radar data) to obtain the aforementioned first normal inverse gamma distribution feature.
[0142] The target radar data can be the range unit corresponding to the target position in step S202 above. In this case, the feature extractor corresponding to the radar data can be a simple one-dimensional convolutional neural network structure.
[0143] In this embodiment, the intensity distribution of reflected signals from various objects can be captured using a distance matrix. Therefore, differences in signal intensity can be used to distinguish between human bodies and interference sources such as static medical equipment. Based on this, combined with preset conditions, target radar data corresponding to human bodies can be efficiently filtered out, reducing the mixing of irrelevant data. Furthermore, when the target radar data is input into the initial physiological signal prediction model, it can provide high-quality input for the model to extract the first normal inverse gamma distribution feature. That is, while ensuring the convenience of non-contact measurement, the quality of the input radar data is enhanced.
[0144] S103. Combine the features of each first normal inverse gamma distribution to obtain the features of the second normal inverse gamma distribution.
[0145] In one embodiment, after obtaining the first normal inverse gamma distribution features corresponding to each modality of data, each first normal inverse gamma distribution feature can be concatenated and fused. Specifically, each first predicted physiological feature can be concatenated to obtain a second predicted physiological feature; each first confidence feature can be concatenated to obtain a second confidence feature; each first distribution feature can be concatenated to obtain each second distribution feature; and each first scale feature can be concatenated to obtain a second scale feature. At this time, the second predicted physiological feature, the second confidence feature, the second distribution feature, and the second scale feature are the aforementioned second normal inverse gamma distribution features.
[0146] However, this fusion method is relatively simple, and the splicing is merely a mechanical stacking of parameters, without considering the statistical correlation between features corresponding to different modalities, thus disrupting the overall probability structure of the distribution features. Furthermore, it fails to differentiate the reliability of the first predicted physiological feature; if the data of a certain modality is severely affected by noise, the spliced first predicted physiological feature will contaminate the overall features, reducing the robustness of subsequent predictions.
[0147] Therefore, in order to improve the fusion effect of the second inverse gamma distribution features, the electronic device can, as shown in the example... Figure 3 Steps S301-S307, as shown, generate the second inverse gamma distribution feature. Details are as follows:
[0148] S301. Based on each first confidence feature, perform a weighted summation on each first predicted physiological feature to obtain the second predicted physiological feature.
[0149] The first predicted physiological features of different modalities may be affected differently by environmental interference (such as radar signal obstruction and equipment noise), and the first confidence feature can quantify the aforementioned reliability differences. Based on this, by using a weighted summation method, modal data with high first confidence (such as radar channels with clear signals and small fluctuations) can be given a higher weight in the summation, reducing the interference of modal data with low first confidence (such as data contaminated by noise) on the results and improving the accuracy of the second predicted physiological features.
[0150] As an example, electronic devices can directly perform a weighted summation of each first confidence feature with its corresponding first predicted physiological feature to obtain the second predicted physiological feature.
[0151] Alternatively, each first confidence feature can be normalized, and the normalized first confidence feature can be used as a weight to participate in the above weighted summation process.
[0152] In another embodiment, the electronic device may further sort each first confidence feature from largest to smallest, and then assign weights to each first predicted physiological feature sequentially based on the order, so as to perform a weighted sum of each weight and the corresponding first predicted physiological feature to obtain a second predicted physiological feature. The earlier the order (i.e., the larger the first confidence feature), the larger the corresponding weight.
[0153] It should be noted that the first predicted physiological feature corresponding to each modality of data is typically a time series, which includes physiological signals corresponding to multiple time points. Based on this, during weighted summation, for the physiological signal corresponding to each time point in each first predicted physiological feature, the physiological signals at that time point can be weighted and summed with their corresponding weights to obtain the fused physiological signal at that time point. The sequence composed of the fused physiological signals at each time point then constitutes the aforementioned second predicted physiological feature.
[0154] In another embodiment, the first confidence feature output by the initial physiological signal prediction model can also be a time series, which includes sub-confidence features corresponding to multiple time points. In this case, during weighted summation, for the physiological signal corresponding to each time point in each first predicted physiological feature, the weights of each physiological signal at that time point and the corresponding sub-confidence features at the same time point can be weighted and summed to obtain the fused physiological signal at that time point.
[0155] In this embodiment, the method of generating the second predicted physiological feature by weighted summation is not limited.
[0156] S302. Sum each first confidence feature to obtain the second confidence feature.
[0157] S303. Sum the first distribution features to obtain the second distribution features.
[0158] In one embodiment, summing can be considered as concatenating each first confidence feature to obtain a second confidence feature, and summing each first distribution feature to obtain a second distribution feature.
[0159] S304. Calculate the feature difference between each first predicted physiological feature and the second predicted physiological feature.
[0160] In one embodiment, the generation method of the second predicted physiological feature has been described above. Therefore, when obtaining the feature difference, for any first predicted physiological feature, the physiological signals at the same time point in the first predicted physiological feature and the second predicted physiological feature can be subtracted to obtain the difference physiological signal. At this time, the sequence of difference physiological signals at each time point is the aforementioned feature difference.
[0161] When there are multiple first predicted physiological features, there will also be multiple feature differences.
[0162] It should be noted that the purpose of calculating the feature difference between each first predicted physiological feature and the second predicted physiological feature is to quantify the degree of deviation between single-modal prediction and fusion prediction in order to assess the uncertainty of the second predicted physiological feature.
[0163] Understandably, the first predicted physiological feature is the physiological signal (such as heart rate and respiratory rate sequences) independently output by each modality (such as different radar channels and camera channels), while the second predicted physiological feature is the consensus result based on the first confidence level weighted fusion. The difference between the two essentially reflects the discrepancy between single-modal prediction and global consensus.
[0164] Therefore, if the difference between the first and second predicted physiological features of a certain modality is small (small feature difference), it can be considered that the predictions of this modality are consistent with those of other modalities, indicating a high degree of consensus among multimodal information. Otherwise, if the difference is large (large difference value), it indicates that there is a discrepancy between the predictions of this modality and those of other modalities, indicating strong conflict among multimodal information. That is, based on the above feature difference, the uncertainty of the initial physiological signal prediction model when making predictions based on the data of this modality can be described.
[0165] S305. Adjust the feature difference based on the first confidence feature corresponding to each feature difference to obtain the target feature difference.
[0166] The target feature difference is used to describe the uncertainty of the second predicted physiological feature.
[0167] In one embodiment, based on the above explanation of the feature differences, it can be understood that the feature differences can describe the uncertainty of the initial physiological signal prediction model when making predictions based on the modality data. On this basis, the target feature difference obtained by fusing multiple feature differences can comprehensively quantify the overall uncertainty of the second predicted physiological feature after multimodal fusion. That is, not only can the prediction biases of each modality be integrated, but the impact of high-reliability modality divergence on the uncertainty of the fusion result can also be highlighted by adjusting the first confidence feature, thereby accurately quantifying the credibility of the second predicted physiological feature.
[0168] In one embodiment, adjusting the feature difference based on the first confidence feature can be achieved by weighted summation of the feature differences based on the first confidence feature to obtain the target feature difference. The weighted summation method can be described in detail below, referring to the method in S301.
[0169] S306. Sum the differences between each first-scale feature and the target feature to obtain the second-scale feature.
[0170] In one embodiment, summing can be considered as summing each first-scale feature and the target difference feature to obtain a second-scale feature.
[0171] It should be noted that the first scale feature is a measure of the uncertainty of each modal sensor in independent prediction (corresponding to the scale parameter in the first normal inverse gamma distribution), reflecting the degree of dispersion caused by noise and interference (such as signal blockage and equipment error) within a single modal sensor. Generally, the larger the first scale feature, the more unstable the prediction result output based on the modal data collected by that modal sensor.
[0172] However, the target feature difference, which is the deviation between the single-modal prediction and the fusion result after adjustment by the first confidence feature (modal reliability), can reflect the additional uncertainty caused by the discrepancy between the single modality and the global consensus when fusing the first normal inverse gamma features of multiple modalities. In this case, the larger the target feature difference, the more significant the conflict between the first normal inverse gamma features of each modality is during the fusion process, and the greater the impact on the reliability of the fusion result.
[0173] Based on the above explanation, the fused second-scale features need to fully reflect the total uncertainty of physiological feature prediction, which typically consists of the following two parts:
[0174] The inherent uncertainty of single-modal data: Even without fusion, the prediction of each modal data itself contains noise, which is quantified by the first-scale feature and serves as the basic uncertainty.
[0175] Uncertainty in fusion discrepancies: When fusing the first normal inverse gamma features of multiple modalities, if the single-modal prediction deviates significantly from the fusion result (and the deviation is more significant for feature differences corresponding to higher first confidence levels), consensus conflicts will occur. In this case, this part represents an additional uncertainty in the fusion process, which needs to be quantified by the aforementioned target feature differences.
[0176] Understandably, if only the first-scale features are retained, the additional risks brought about by fusion discrepancies will be ignored; if only the target feature differences are retained, the inherent noise of the single-modal data itself will be lost, and neither can fully reflect the true uncertainty.
[0177] Based on this, by generating the second-scale feature through the above steps S304-S306, the second-scale feature can not only retain the basic error information of each modality data, but also accommodate the conflict risk in the fusion process, so as to achieve a comprehensive and accurate quantification of the uncertainty of the prediction of the second physiological feature after fusion.
[0178] S307. The second predictive physiological feature, the second confidence feature, the second distribution feature, and the second scale feature are determined as the second normal inverse gamma distribution feature.
[0179] In one embodiment, the second normal inverse gamma distribution feature is similar to the first normal inverse gamma distribution feature described above, both containing the four parameters mentioned above. Therefore, it can be considered that obtaining the second predicted physiological feature, the second confidence feature, the second distribution feature, and the second scale feature constitutes obtaining the second normal inverse gamma distribution feature.
[0180] As an example, the first normal inverse gamma distribution feature NIG(γ) corresponding to video data is used. V ,υ V ,α V ,β VThe first normal inverse gamma distribution characteristic NIG(γ) corresponding to radar data R ,υ R ,α R ,β R Taking (e.g.) as an example, the second inverse gamma distribution feature is generated. The fusion formula can be explained as follows:
[0181]
[0182] in, This is an addition operation of two NIG distributions, where F represents the fusion result, V represents video data, R represents radar data, γ represents predicted physiological features, υ represents confidence features, α represents distribution features, and β represents scale features. For example, γ... F Indicates the second predicted physiological characteristic after fusion, γ V , representing the first predicted physiological feature corresponding to the video data, γ R This represents the first predicted physiological characteristic corresponding to the radar data. The parameters can be explained as in the example above, and will not be described in detail here.
[0183] The calculation method for each internal parameter is as follows:
[0184] γ F =(υ V +υ R ) -1 (υ V γ V +υ R γ R );
[0185] υ F =υ V +υ R ;
[0186] α F =α V +α R ;
[0187]
[0188] Among them, (υ V +υ R ) -1 (υ V γ V +υ R γ R This can be considered as a weighted sum of the first predicted physiological feature after normalizing each first confidence level; γ V -γ F It can be considered as the feature difference between the first predicted physiological feature and the second predicted physiological feature corresponding to the video data, γ.R -γ F It can be considered as the feature difference between the first predicted physiological feature and the second predicted physiological feature corresponding to the radar data. It can be considered as the difference in target features corresponding to the video data. It can be considered as the target feature difference corresponding to the radar data.
[0189] In this embodiment, the weighted summation of the prediction results based on the first confidence feature not only highlights the contribution of high-reliability modal data but also enables dynamic adaptive fusion, making the second predicted physiological feature closer to the true value. Furthermore, by accumulating the first confidence and the first distribution feature, and adding the difference between the first scale feature and the target feature adjusted by the first confidence, the basic error information and fusion discrepancy risk of each modal data can be fully preserved, allowing the second normal inverse gamma distribution feature to accurately describe the uncertainty range of the prediction result. Therefore, based on the above fusion method, not only can the robustness of the initial physiological signal prediction model to noise and outliers be enhanced, but it can also provide a quantitative basis for subsequent uncertainty-based decision-making.
[0190] S104. Calculate the feature training loss of the normal inverse gamma distribution features of each target based on real physiological signals.
[0191] The objectives are divided into first and second. That is, calculating the feature training loss for each first normal inverse gamma distribution feature also includes calculating the feature training loss for the second normal inverse gamma distribution feature.
[0192] In one embodiment, the method for calculating the feature training loss includes, but is not limited to, using functions such as mean squared error loss, logarithmic loss, and cross-entropy loss.
[0193] Based on the above explanation of the first normal inverse gamma distribution features and the second normal inverse gamma distribution features, it can be seen that the target normal inverse gamma distribution features include four parameters. Therefore, the training loss corresponding to each parameter can be calculated based on the above loss function, and then the four training losses can be summed to obtain the above feature training loss.
[0194] In another embodiment, the electronic device can be based on, for example... Figure 4 Steps S401-S405, as shown, calculate the feature training loss described above. Details are as follows:
[0195] S401. For any target's normal inverse gamma distribution characteristics, calculate the similarity loss between the predicted physiological characteristics of the target and the actual physiological signals.
[0196] In one embodiment, similarity loss measures the degree of difference between the target predicted physiological features output by the initial physiological signal prediction model and the actual observed physiological signals. Quantifying this difference can guide the optimization direction of model parameters. That is, minimizing the similarity loss makes the model prediction results as close as possible to the true values, thereby improving prediction accuracy.
[0197] The electronic device can calculate the aforementioned similarity loss using methods such as mean squared error or cosine distance, and there are no limitations on this. For example, the electronic device can calculate the similarity loss using the following formula:
[0198]
[0199] in, Let y represent the predicted physiological features of the target, y represent the actual physiological signal, and T represent the sampling time point of the actual physiological signal. Each sampling time point corresponds to a physiological signal value, in order to calculate the above similarity loss.
[0200] S402. Calculate the signal-to-noise ratio loss in the frequency domain between the predicted physiological characteristics of the target and the actual physiological signal.
[0201] In one embodiment, the aforementioned frequency domain is the transformation of the time domain (a signal that varies over time) to the frequency dimension using a Fourier transform (FFT) to analyze the frequency components of the signal (e.g., the dominant frequency of a heart rate signal corresponds to the heart rate value, and the dominant frequency of a respiratory signal corresponds to the respiratory rate). The frequency domain can represent the periodic characteristics of a signal and is generally easier to separate noise from the valid signal than the time domain.
[0202] The aforementioned signal-to-noise ratio (SNR) loss is a quantitative indicator of the difference in SNR between the predicted physiological features and the actual physiological signals, used to measure the degree of quality degradation of the predicted physiological features. Generally, a larger SNR loss indicates that the predicted physiological features contain more noise and have weaker effective components compared to the actual physiological signals.
[0203] As an example, an electronic device can perform a Fourier transform on the predicted physiological features of the target to obtain frequency domain predicted physiological features. Then, it can separate the effective physiological features from the noisy physiological features within the frequency domain predicted physiological features. Finally, it calculates the ratio of the energy corresponding to the effective physiological features to the energy corresponding to the noisy physiological features, obtaining the predicted signal-to-noise ratio (SNR). The difference between the predicted SNR and the SNR corresponding to the actual physiological signal is then determined as the aforementioned SNR loss.
[0204] In another embodiment, the electronic device can further determine the effective frequency corresponding to the maximum energy of the real physiological signal in the frequency domain. Then, based on the effective frequency and a preset window, a frequency range is determined. Next, a Fourier transform is performed on the target predicted physiological feature to obtain frequency domain predicted physiological parameters. From these parameters, a first predicted physiological parameter within the frequency range and a second predicted physiological parameter outside the frequency range are determined. Finally, the signal-to-noise ratio loss is calculated based on the first and second predicted physiological parameters.
[0205] The effective frequency mentioned above refers to the frequency value corresponding to the highest energy frequency component in the frequency domain after Fourier transform of the real physiological signal (such as heart rate and respiratory rate). Typically, physiological signals are periodic signals (heart rate is about 0.5-2.5Hz, and respiratory rate is about 0.1-1Hz), with energy concentrated near the dominant frequency. The effective frequency is usually the dominant frequency of the physiological signal (for example, a heart rate of 60 beats / minute corresponds to a dominant frequency of 1Hz).
[0206] In one embodiment, the electronic device can perform a Fourier transform on the real physiological signal to obtain the frequency domain real physiological signal. Then, the energy (usually the square of the spectral amplitude) corresponding to each frequency in the frequency domain real physiological signal can be calculated to determine the frequency value corresponding to the frequency point with the highest energy as the aforementioned effective frequency.
[0207] The preset window mentioned above can be a manually set frequency range width used to select the frequency range of valid physiological signals. Typically, physiological signals exhibit frequency fluctuations (for example, heart rate changes with exercise, and the dominant frequency may drift within a small range). Therefore, a window can be used to cover a reasonable fluctuation range near the dominant frequency.
[0208] The aforementioned frequency range can be considered as an interval range obtained by adding or subtracting a preset window centered on the effective frequency, used to distinguish the frequency components of effective physiological signals from the frequency components of noise / interference.
[0209] In one embodiment, the above-described method of performing Fourier transform on the target predicted physiological characteristics to obtain frequency domain predicted physiological parameters is similar to the above-described method of performing Fourier transform on the real physiological signal to obtain the frequency domain real physiological signal, and will not be described in detail.
[0210] The first predicted physiological parameter is a parameter whose frequency falls within a preset window in the frequency domain predicted physiological parameters. Typically, this first predicted physiological parameter can be considered a frequency component of the valid physiological signal. The second predicted physiological parameter is a parameter whose frequency falls outside the preset window in the frequency domain predicted physiological parameters; typically, this second predicted physiological parameter can be considered a frequency component of noise or interference.
[0211] As an example, when calculating the signal-to-noise ratio loss based on the first and second predicted physiological parameters, the effective energy corresponding to the first predicted physiological parameter and the noise energy corresponding to the second predicted physiological parameter can be used as the aforementioned signal-to-noise ratio loss.
[0212] For example, the formula for calculating the signal-to-noise ratio loss mentioned above can be:
[0213]
[0214] f0 = argmax f Y(f);
[0215] in, Let Y(f) represent the signal-to-noise ratio loss, and Y(f) represent the frequency domain real physiological signal after Fourier transform of the real physiological signal. The frequency domain predicted physiological parameters are represented by the Fourier transform of the target predicted physiological features, f0 represents the effective frequency corresponding to the maximum energy of the real physiological signal in the frequency domain, and h represents the preset window.
[0216] It should be noted that the dominant frequency of a true physiological signal is a core feature (directly corresponding to physiological parameters such as heart rate and respiratory rate), but the dominant frequency will fluctuate slightly due to the body's state (exercise, emotions). Therefore, by locating the effective frequency corresponding to the maximum energy value to determine the dominant frequency of the true physiological signal, and by covering the fluctuation range with a preset window, the effective physiological frequency range can be accurately determined, avoiding the misjudgment of effective components as noise due to frequency drift.
[0217] Furthermore, the target predicted physiological features output by the initial physiological signal prediction model should be similar to the actual physiological signals. Therefore, the frequency range determined based on the effective frequency and the preset window can be applied to predict physiological parameters in the frequency domain. Subsequently, the electronic device can calculate the aforementioned signal-to-noise ratio loss based on the first predicted physiological parameter within the frequency range and the second predicted physiological parameter not within the frequency range.
[0218] In this embodiment, the effective frequency is first locked by the peak energy of the real physiological signal in the frequency domain, and an effective frequency range that conforms to the natural fluctuations of the real physiological signal can be determined by combining it with a preset window. Then, a Fourier transform is performed on the target predicted physiological feature to obtain the frequency domain predicted physiological parameters, and effective first and second predicted physiological parameters can be determined based on the frequency range. Finally, the signal-to-noise ratio (SNR) loss is calculated based on the first and second predicted physiological parameters, allowing the SNR loss to quantify the quality of the target predicted physiological feature in the key frequency bands of the physiological signal. Furthermore, the optimization process of the initial physiological signal prediction model can focus on preserving the frequency domain features of the real physiological signal and suppressing irrelevant noise interference.
[0219] S403. Based on the target predicted physiological characteristics, real physiological signals, target confidence, target distribution characteristics, and target scale characteristics, calculate the uncertainty loss of the initial physiological signal prediction model.
[0220] In one embodiment, the aforementioned uncertainty loss is used to quantify the degree of matching between the predicted distribution (target predicted physiological characteristics) and the true value (true physiological signal), while penalizing predictions with high uncertainty. Typically, this can be calculated by combining the likelihood of the normal inverse gamma with uncertainty regularization, which will not be described in detail here.
[0221] As an example, an electronic device can calculate the predicted difference between the predicted physiological characteristics of a target and the actual physiological signal. Then, based on the predicted difference, target confidence, target distribution characteristics, and target scale characteristics, a negative likelihood loss is calculated. Additionally, based on the predicted difference, target confidence, and target distribution characteristics, an error loss is calculated. Finally, based on the negative likelihood loss and the error loss, the uncertainty loss is obtained.
[0222] The negative likelihood loss describes the degree of uncertainty in the initial physiological signal prediction model. The error loss describes the degree of error between the predicted physiological features and the actual physiological signals.
[0223] In one embodiment, the above-mentioned prediction difference is the difference between the target predicted physiological feature and the corresponding element of the real physiological signal, which is used to intuitively reflect the degree of deviation between the prediction result and the actual situation.
[0224] The aforementioned negative likelihood loss is based on probability distribution theory and measures the degree of matching between the predicted distribution and the true value, reflecting the uncertainty of the model's prediction. For example, assuming that the predicted structure follows a normal inverse gamma distribution, and the distribution is parameterized by the target confidence, distribution characteristics, and scale characteristics, then the negative likelihood loss can be calculated based on the predicted probability distribution (target normal inverse gamma distribution characteristics).
[0225] A larger negative likelihood loss indicates a lower degree of match between the predicted distribution and the true value, and a higher level of model uncertainty. Conversely, a smaller negative likelihood loss indicates a higher degree of match between the predicted distribution and the true value, and a lower level of model uncertainty.
[0226] The aforementioned error loss is the direct error between the predicted distribution and the true value, used to constrain the model's prediction error and its fit to uncertainty, balancing precise fitting with reasonable uncertainty estimation. Specifically, the error loss for electronic devices can be calculated by combining target confidence level and target distribution characteristics.
[0227] For example, the electronic device can first determine the weights corresponding to the target confidence level and the target distribution, and then sum the weights to obtain the total weight. The product of the total weight and the feature difference can then be determined as the aforementioned error loss.
[0228] In one embodiment, after obtaining the negative likelihood loss and the error loss, the negative likelihood loss and the error loss can be weighted and summed to obtain the uncertainty loss. The weights corresponding to the negative likelihood loss and the error loss can be preset and are not limited thereto.
[0229] As an example, electronic devices can calculate the aforementioned uncertainty loss using the following formula:
[0230]
[0231] Ω = 2β(1+γ);
[0232]
[0233] in, Indicates loss due to uncertainty. Indicates negative likelihood loss. denoted by τ, representing error loss; τ represents hyperparameters used to balance uncertainty and model fit; v represents target confidence; α represents target distribution characteristics; β represents target scale characteristics; y represents the true physiological signal; γ represents the predicted physiological characteristics of the target; and Γ represents the gamma function.
[0234] S404. Based on similarity loss, signal-to-noise ratio loss, and uncertainty loss, calculate the feature training loss corresponding to the target normal inverse gamma distribution.
[0235] In one embodiment, after obtaining each loss, the various losses can be weighted and summed to obtain the feature training loss. The weights corresponding to each loss can be the same or different; this is not limited.
[0236] In this embodiment, the negative likelihood loss describes the uncertainty of the initial physiological signal prediction model from the perspective of the normal inverse gamma distribution, enabling the initial physiological signal prediction model to quantitatively understand the reliability of its own predictions. Furthermore, the error loss focuses on the direct deviation between the predicted and actual values, intuitively reflecting the accuracy of the prediction. Based on the negative likelihood loss and the error loss, the uncertainty loss is derived, ensuring that the initial physiological signal prediction model, while pursuing prediction accuracy, can reasonably and realistically assess its own uncertainty, avoiding overconfidence or conservatism. This allows the initial physiological signal prediction model to possess accurate fitting capabilities in physiological signal data prediction, improving its robustness and practicality in complex scenarios.
[0237] S105. Iteratively update the model parameters of the initial physiological signal prediction model based on the training loss of all features to obtain the physiological signal prediction model.
[0238] In one embodiment, the aforementioned model parameters are learnable variables in the initial physiological signal prediction model, determining the model's predictive behavior. The physiological signal prediction model is a model whose parameters have been iteratively optimized, capable of outputting accurate physiological signals with reasonable uncertainty on test data.
[0239] The iterative methods mentioned above include, but are not limited to, gradient descent and adaptive learning rate methods; no specific limitations are imposed on these methods.
[0240] In one embodiment, based on step S104 above, the feature training loss includes the feature training loss corresponding to a single modality of data, and also includes the feature training loss corresponding to the fused features. Therefore, the electronic device can perform a weighted summation of all feature training losses to obtain the total training loss, and then update the model parameters based on the total training loss. The weights corresponding to each feature training loss can be the same or different; this is not limited.
[0241] As an example, the electronic device performs a weighted sum of the training losses for all features to obtain the feature training loss, with the weight of the feature training loss corresponding to radar data being less than the weight of the feature training losses corresponding to the other modalities. Then, the model parameters are iteratively updated based on the feature training loss to obtain the physiological signal prediction model.
[0242] In one embodiment, in the actual use scenario of physiological signals, radar equipment is susceptible to interference in a hospital environment. Specifically, during clinical use, radar data collected by the equipment is easily affected by electromagnetic radiation from the metal frame of the hospital bed and surrounding medical instruments, resulting in a large amount of non-physiological signal noise mixed in with the radar data. In this case, if the weight corresponding to the radar data is too high, the initial physiological signal prediction model may overfit the noise and ignore other reliable modal data, leading to a decrease in the accuracy of the trained physiological signal prediction model. Therefore, the weight of the feature training loss corresponding to the radar data can be set to be less than the weight of the feature training loss corresponding to other modal data.
[0243] The feature training loss for the remaining modal data can include the feature training loss corresponding to the fused second normal inverse gamma distribution feature. As an example, the weight for video data can be 1, the weight for radar data can be 0.1, and the weight for the feature training loss corresponding to the second normal inverse gamma distribution feature can be 1.
[0244] In this embodiment, training data containing multiple modalities and corresponding real physiological signals is acquired non-contactly. This retains the convenience of non-contact measurement while covering physiological signal characteristics in different scenarios through multi-modal data, providing a more comprehensive learning foundation for the initial physiological signal prediction model and avoiding the scenario limitations of single-modal data. Subsequently, a first normal inverse gamma distribution feature, comprising a first predicted physiological feature, a first confidence feature, a first distribution feature, and a first scale feature, is extracted for each modality. This overcomes the limitation of traditional techniques that only extract numerical signal features. The first confidence feature quantifies the reliability of each modality prediction, while the first distribution feature and the first scale feature describe the uncertainty of signal variance distribution and scale. Therefore, based on the above feature information, the initial physiological signal prediction model can accurately identify the quality of the modal data. That is, the electronic device can fuse the first normal inverse gamma distribution feature of each modality based on the confidence feature to obtain a second normal inverse gamma distribution feature, and variance distribution information can be integrated during the fusion process to suppress interference from noisy modalities. Furthermore, based on the aforementioned feature fusion method, the problem of noisy modes dragging down overall performance in traditional splicing can be solved, making the fused second normal inverse gamma distribution feature more focused on the effective information of high-quality modes. Finally, the feature training loss of each target (first and second) normal inverse gamma distribution feature is calculated separately, and the model parameters of the initial physiological signal prediction model are iteratively updated to obtain the physiological signal prediction model. By calculating the feature training loss for the first and second normal inverse gamma distribution features separately, the prediction accuracy of single-modal data features can be guaranteed (through the feature training loss calculated by the first normal inverse gamma feature), while also strengthening the consistency between the fused features and the real signal (through the feature training loss calculated by the second normal inverse gamma feature). Furthermore, through the dual training loss calculation, the model can optimize the multimodal collaborative mechanism while learning single-modal characteristics, avoiding information loss or interference during the fusion process. In this way, the prediction accuracy of the trained physiological signal prediction model can be improved.
[0245] To more clearly illustrate the solutions in this application, specific embodiments are used below to explain the solutions. See details below. Figure 5 , Figure 5 This is a schematic diagram of the physiological signal prediction model provided in one embodiment of this application. Taking radar data and video data as examples of multimodal data, the physiological signal prediction model can be divided into a network structure including an input layer, an Uncertainty Regression Head, an Evidential Multi-Modal Fusion (EMMF) layer, and a loss calculation layer.
[0246] In the input layer, the input for video data is a series of image frames with dimensions H (height) × W (width) × T (time step), which are processed by a feature extractor to obtain a feature representation. The input for radar data is a range matrix, which is processed by a Fast Fourier Transform (FFT) and then processed by feature extraction to obtain a feature representation.
[0247] For the uncertainty regression layer, feature representations from video and radar data can be received respectively. Then, through neural network processing, the parameters of the first normal inverse gamma distribution (NIG) corresponding to the video and radar data are output. That is, γ (first predicted physiological feature), ν (first confidence feature), α (first distribution feature), and β (first scale feature) are output. Furthermore, the uncertainty regression layer can also calculate the feature training loss corresponding to each first normal inverse gamma distribution feature.
[0248] The multimodal fusion layer can receive and fuse the first normal inverse gamma distributions corresponding to video data and radar data, respectively. Furthermore, the multimodal fusion layer can also calculate the feature training loss corresponding to the features of the second normal inverse gamma distribution.
[0249] For the loss calculation layer, the training loss corresponding to the video data, radar data, and the fused second normal inverse gamma distribution feature can be connected to the total loss calculation module to calculate the total training loss.
[0250] During model training, the initial physiological signal prediction model can be iteratively updated based on the total training loss to obtain the physiological signal prediction model. Furthermore, during use, the physiological signal prediction model can directly output the second predicted physiological feature from the second normal inverse gamma distribution features as the prediction result.
[0251] Referring to Figure 6, which is a schematic diagram illustrating the results of predicting physiological signals using different methods in different application scenarios according to an embodiment of this application, Figure 6 visualizes the results of predicting physiological signals using video, radio frequency (RF), and fusion methods in three different application scenarios (Varying Light, Occlusion, and Motion). The prediction accuracy is measured using the mean absolute error (MAE). The results for each application scenario are explained in detail below:
[0252] 1. Varying Light Scene
[0253] Video: MAE is 12.00. As can be seen from the figure, in scenarios with changing lighting, there is a large deviation between the physiological signal predicted based solely on video data (blue curve) and the actual physiological signal (red curve), with significant fluctuations. This indicates that under changing lighting conditions, the single video modality is greatly affected by lighting, resulting in low prediction accuracy.
[0254] RF (Radar): MAE is 5.86. In scenarios with changing lighting, the physiological signal curve predicted solely based on radar data is closer to the actual signal than that predicted from video data, with smaller fluctuations. This indicates that the radar mode is less affected by lighting conditions and exhibits better stability.
[0255] Fusion: MAE is 6.14. In scenarios with changing lighting, the fused prediction combines the advantages of video and radar data. Although the MAE is slightly higher than radar, it is generally closer to the real signal than the prediction based solely on video data. This indicates that the fusion prediction of two modalities can, to some extent, compensate for the shortcomings of a single modality in scenarios with changing lighting.
[0256] 2. Occlusion Scene
[0257] Video: MAE is 0.50. In occluded scenarios, the physiological signal curve predicted solely from video data almost perfectly matches the actual signal, resulting in a very small MAE. This may be because, under these specific occlusion conditions, the video can still capture sufficient physiological signal features, or the occlusion has a relatively small impact on the video.
[0258] RF (Radar): MAE is 36.10. In obstructed scenarios, the physiological signal curve predicted solely based on radar data deviates significantly from the actual signal, almost appearing as a straight line. This indicates that the radar is severely interfered with in obstructed scenarios and cannot accurately capture physiological signals.
[0259] Fusion: MAE is 0.50. In occluded scenarios, the fused prediction result is consistent with the video prediction result, and the MAE is also 0.50. This indicates that in occluded scenarios, the fusion prediction of the two modalities mainly relies on the prediction result of the video modality, while ignoring the inaccurate prediction of the radar modality, thus ensuring high prediction accuracy.
[0260] 3. Motion Scene
[0261] Video: MAE is 5.85. In motion scenarios, the physiological signal curves predicted based solely on video data deviate somewhat from the actual signals, but the overall trend is relatively close, indicating that the video modality has a certain degree of adaptability in motion scenarios.
[0262] RF (Radar): MAE is 6.10. In motion scenarios, the deviation between the physiological signal curve predicted solely based on radar data and the actual signal is similar to that in video, indicating that radar can also provide a certain predictive capability in motion scenarios.
[0263] Fusion: MAE is 5.85. In motion scenarios, the fused prediction result is the same as the video prediction result, with an MAE of 5.85. This indicates that in motion scenarios, the fusion prediction of the two modalities combines the prediction results of video and radar, further improving the stability and accuracy of the prediction.
[0264] Based on the above explanation, it can be concluded that different modalities of data have their own advantages and disadvantages in different application scenarios: video data performs well in occluded scenarios, but its accuracy is lower in scenarios with changing lighting. Radar data performs well in scenarios with changing lighting, but is severely affected by interference in occluded scenarios. Fusion data can combine the advantages of video and radar, and can provide relatively accurate and stable prediction results in most application scenarios. In particular, in scenarios with changing lighting and motion, the performance of fusion of two modalities of data is better than that of a single modality.
[0265] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a training device for a physiological signal prediction model provided in this application embodiment. The training device for the physiological signal prediction model in this embodiment includes modules for executing... Figures 1 to 4 The steps in the corresponding embodiments. Please refer to the details. Figures 1 to 4 as well as Figures 1 to 4 The relevant descriptions in the corresponding embodiments are shown below. For ease of explanation, only the parts relevant to this embodiment are shown. See also... Figure 7 The training device 700 for the physiological signal prediction model may include: an acquisition module 710, an input module 720, a fusion module 730, a calculation module 740, and an iteration module 750, wherein:
[0266] The acquisition module 710 is used to acquire a training dataset containing multiple training data. The training data is acquired in a non-contact manner, and each training data includes multiple modal data and corresponding real physiological signals.
[0267] The input module 720 is used to input multiple modal data into the initial physiological signal prediction model for any training data, and extract the first normal inverse gamma distribution feature of each modal data respectively; the first normal inverse gamma distribution feature includes the first predicted physiological feature for predicting physiological signals, the first confidence feature characterizing the reliability of the predicted physiological signals, the first distribution feature characterizing the variance distribution of the predicted physiological signals, and the first scale feature characterizing the variance scale of the predicted physiological signals.
[0268] The fusion module 730 is used to fuse each first normal inverse gamma distribution feature to obtain the second normal inverse gamma distribution feature.
[0269] The calculation module 740 is used to calculate the feature training loss of the normal inverse gamma distribution features of each target based on real physiological signals; the targets are divided into first and second.
[0270] Iteration module 750 is used to iteratively update the model parameters of the initial physiological signal prediction model based on the training loss of all features, so as to obtain the physiological signal prediction model.
[0271] In one embodiment, the modal data includes radar data; the input module 720 is further configured to:
[0272] A Fourier transform is performed on the radar data to generate a range matrix; each matrix value in the range matrix is used to describe the signal strength of the radar signal reflected by an object in space; based on the range matrix, target radar data that meets preset conditions is determined; the preset conditions include conditions for determining the category of the object as human; the target radar data is input into the initial physiological signal prediction model to obtain the first normal inverse gamma distribution feature.
[0273] In one embodiment, the input module 720 is further configured to:
[0274] The target location of the object corresponding to the largest matrix value in the distance matrix is determined; radar data from the target location within a preset range is determined as target radar data that meets preset conditions.
[0275] In one embodiment, the fusion module 730 is further configured to:
[0276] Based on each first confidence feature, a weighted sum is performed on each first predicted physiological feature to obtain a second predicted physiological feature; each first confidence feature is summed to obtain a second confidence feature; each first distribution feature is summed to obtain a second distribution feature; the feature difference between each first predicted physiological feature and the second predicted physiological feature is calculated; the feature difference is adjusted based on the first confidence feature corresponding to each feature difference to obtain a target feature difference; the target feature difference is used to describe the uncertainty of the second predicted physiological feature; the difference between each first scale feature and the target feature is summed to obtain a second scale feature; the second predicted physiological feature, the second confidence feature, the second distribution feature, and the second scale feature are determined as the second normal inverse gamma distribution feature.
[0277] In one embodiment, the calculation module 740 is further configured to:
[0278] For any target's normal inverse gamma distribution characteristics, calculate the similarity loss between the predicted physiological features and the actual physiological signals; calculate the signal-to-noise ratio loss between the predicted physiological features and the actual physiological signals in the frequency domain; calculate the uncertainty loss of the initial physiological signal prediction model based on the predicted physiological features, actual physiological signals, target confidence, target distribution characteristics, and target scale characteristics; and calculate the feature training loss corresponding to the target's normal inverse gamma distribution based on the similarity loss, signal-to-noise ratio loss, and uncertainty loss.
[0279] In one embodiment, the calculation module 740 is further configured to:
[0280] Determine the effective frequency corresponding to the maximum energy of the real physiological signal in the frequency domain; determine the frequency range based on the effective frequency and a preset window; perform Fourier transform on the target predicted physiological features to obtain frequency domain predicted physiological parameters; determine the first predicted physiological parameter that is within the frequency range and the second predicted physiological parameter that is not within the frequency range from the frequency domain predicted physiological parameters; calculate the signal-to-noise ratio loss based on the first and second predicted physiological parameters.
[0281] In one embodiment, the calculation module 740 is further configured to:
[0282] Calculate the prediction difference between the predicted physiological features and the actual physiological signals; calculate the negative likelihood loss based on the prediction difference, target confidence, target distribution characteristics, and target scale characteristics; the negative likelihood loss is used to describe the degree of uncertainty in the initial physiological signal prediction model; calculate the error loss based on the prediction difference, target confidence, and target distribution characteristics; the error loss is used to describe the degree of error between the predicted physiological features and the actual physiological signals; and obtain the uncertainty loss based on the negative likelihood loss and the error loss.
[0283] In one embodiment, the modal data includes radar data; the iteration module 750 is further configured to:
[0284] The feature training loss is obtained by weighted summation of all feature training losses; the weight of the feature training loss corresponding to radar data is less than the weight of the feature training loss corresponding to other modal data; the model parameters are iteratively updated based on the feature training loss to obtain the physiological signal prediction model.
[0285] When it is understood that, Figure 7 In the schematic diagram of the training device for the physiological signal prediction model shown, each module is used to execute... Figures 1 to 4 The steps in the corresponding embodiments, and for Figures 1 to 4 The steps in the corresponding embodiments have been explained in detail in the above embodiments. Please refer to them for details. Figures 1 to 4 as well as Figures 1 to 4 The relevant descriptions in the corresponding embodiments will not be repeated here.
[0286] Figure 8 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Figure 8 As shown, the electronic device 800 of this embodiment includes: a processor 810, a memory 820, and a computer program 830 stored in the memory 820 and executable on the processor 810, such as a program for training a physiological signal prediction model. When the processor 810 executes the computer program 830, it implements the steps of each embodiment of the training method for the various physiological signal prediction models described above, for example... Figure 1 S101 to S104 are shown. Alternatively, the processor 810 implements the above when executing the computer program 830. Figure 7 The functions of each module in the corresponding embodiments, for example, Figure 7 For details on the functions of each module shown, please refer to [link / reference]. Figure 7 The relevant descriptions in the corresponding embodiments.
[0287] For example, the computer program 830 can be divided into one or more modules, one or more of which are stored in the memory 820 and executed by the processor 810 to implement the training method of the physiological signal prediction model provided in the embodiments of this application. One or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 830 in the electronic device 800. For example, the computer program 830 can implement the training method of the physiological signal prediction model provided in the embodiments of this application.
[0288] Electronic device 800 may include, but is not limited to, processor 810 and memory 820. Those skilled in the art will understand that... Figure 8 This is merely an example of electronic device 800 and does not constitute a limitation on electronic device 800. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device may also include input / output devices, network access devices, buses, etc.
[0289] The processor 810 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0290] The memory 820 can be an internal storage unit of the electronic device 800, such as a hard disk or RAM of the electronic device 800. The memory 820 can also be an external storage device of the electronic device 800, such as a plug-in hard disk, smart memory card, flash memory card, etc., equipped on the electronic device 800. Furthermore, the memory 820 can include both internal storage units and external storage devices of the electronic device 800.
[0291] This application provides a computer-readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the training method for the physiological signal prediction model as described in the above embodiments.
[0292] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the training method for the physiological signal prediction model in the above embodiments.
[0293] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method of training a physiological signal prediction model, characterized by, The method comprises: obtaining a training data set comprising a plurality of training data; the training data are obtained in a non-contact manner, and each of the training data comprises a plurality of modal data and a corresponding real physiological signal; for any of the training data, inputting the plurality of modal data into an initial physiological signal prediction model, and extracting a first normal inverse gamma distribution feature of each of the modal data; the first normal inverse gamma distribution feature comprises a first predicted physiological feature of a predicted physiological signal, a first confidence feature representing reliability of the predicted physiological signal, a first distribution feature representing variance distribution of the predicted physiological signal, and a first scale feature representing variance scale of the predicted physiological signal; fusing each of the first normal inverse gamma distribution features to obtain a second normal inverse gamma distribution feature; calculating a feature training loss of each target normal inverse gamma distribution feature based on the real physiological signal; the target is divided into first and second; updating model parameters of the initial physiological signal prediction model based on all the feature training losses to obtain a physiological signal prediction model.
2. The method of claim 1, wherein, The modal data comprise radar data; the inputting of the plurality of modal data into the initial physiological signal prediction model and the extracting of the first normal inverse gamma distribution feature of each of the modal data comprise: performing Fourier transform on the radar data to generate a distance matrix; each matrix value in the distance matrix is used to describe signal intensity of an object in space reflecting radar signals; determining target radar data in the radar data that meet a preset condition based on the distance matrix; the preset condition comprises a condition for judging that the object is of a human category; inputting the target radar data into the initial physiological signal prediction model to obtain the first normal inverse gamma distribution feature.
3. The method of claim 2, wherein, The determining of the target radar data in the radar data that meet the preset condition based on the distance matrix comprises: determining a target position of the object corresponding to a maximum matrix value from the distance matrix; determining radar data from a preset range of the target position as the target radar data that meet the preset condition.
4. The method of claim 1, wherein, The fusing of each of the first normal inverse gamma distribution features to obtain the second normal inverse gamma distribution feature comprises: weighting and summing each of the first predicted physiological features based on each of the first confidence features to obtain a second predicted physiological feature; summing each of the first confidence features to obtain a second confidence feature; summing each of the first distribution features to obtain a second distribution feature; calculating a feature difference value between each of the first predicted physiological features and the second predicted physiological feature; adjusting the feature difference value based on the first confidence feature corresponding to each of the feature difference values to obtain a target feature difference value; the target feature difference value is used to describe uncertainty of the second predicted physiological feature; summing each of the first scale features and the target feature difference value to obtain a second scale feature; The second predicted physiological feature, the second confidence feature, the second distribution feature, and the second scale feature are determined as the second inverse normal gamma distribution feature.
5. The method of claim 1, wherein, The feature training loss of each target inverse normal gamma distribution feature is calculated based on the real physiological signal, including: For any target inverse normal gamma distribution feature, a similarity loss between the target predicted physiological feature and the real physiological signal is calculated; A signal-to-noise ratio loss between the target predicted physiological feature and the real physiological signal in the frequency domain is calculated; An uncertainty loss of the initial physiological signal prediction model is calculated based on the target predicted physiological feature, the real physiological signal, a target confidence, a target distribution feature, and a target scale feature; The feature training loss corresponding to the target inverse normal gamma distribution is calculated based on the similarity loss, the signal-to-noise ratio loss, and the uncertainty loss.
6. The method of claim 5, wherein, The signal-to-noise ratio loss between the target predicted physiological feature and the real physiological signal in the frequency domain is calculated, including: An effective frequency corresponding to an energy maximum value of the real physiological signal in the frequency domain is determined; A frequency range is determined based on the effective frequency and a preset window; A frequency domain predicted physiological parameter is obtained by performing Fourier transform on the target predicted physiological feature; A first predicted physiological parameter located in the frequency range and a second predicted physiological parameter not located in the frequency range are determined from the frequency domain predicted physiological parameter; The signal-to-noise ratio loss is calculated based on the first predicted physiological parameter and the second predicted physiological parameter.
7. The method of claim 5, wherein, The uncertainty loss of the initial physiological signal prediction model is calculated based on the target predicted physiological feature, the real physiological signal, a target confidence, a target distribution feature, and a target scale feature, including: A prediction difference between the target predicted physiological feature and the real physiological signal is calculated; A negative likelihood loss is calculated based on the prediction difference, the target confidence, the target distribution feature, and the target scale feature; the negative likelihood loss is used to describe an uncertainty degree when the initial physiological signal prediction model is predicted; An error loss is calculated based on the prediction difference, the target confidence, and the target distribution feature; the error loss is used to describe an error degree between the target predicted physiological feature and the real physiological signal; The uncertainty loss is obtained based on the negative likelihood loss and the error loss.
8. The method according to any one of claims 1 to 7, characterized in that, The modal data includes radar data; the model parameters of the initial physiological signal prediction model are iteratively updated based on all the feature training losses to obtain a physiological signal prediction model, including: All the feature training losses are weighted and summed to obtain a feature training loss; the weight of the feature training loss corresponding to the radar data is smaller than the weights of the feature training losses corresponding to the remaining modal data; The model parameters are iteratively updated based on the feature training loss to obtain a physiological signal prediction model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method in any of claims 1 to 8. The processor executes the computer program to implement the method in any of claims 1 to 8.
10. A computer program product, characterised in that, When the computer program product is run on an electronic device, it causes the electronic device to perform a method as claimed in any of claims 1 to 8.