Cross-modal rppg signal perception method and device based on large language model and electronic equipment

A cross-modal rPPG signal sensing method combining a large language model with a dual-domain stationary algorithm and multi-scale feature fusion solves the accuracy and robustness issues of rPPG signal sensing in complex scenarios, improving the prediction accuracy and robustness of the signal.

CN120412016BActive Publication Date: 2026-01-27BEIZHI TECHNOLOGY (ANJI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510477148.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2026-01-27
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing deep learning technologies lack accuracy and robustness in cross-modal rPPG signal sensing, especially in complex scenarios where it is difficult to maintain high measurement accuracy and is affected by factors such as illumination changes, motion artifacts, and head movements.

Method used

A cross-modal rPPG signal perception method based on a large language model is adopted. By extracting low-precision rPPG signals and fusing multi-scale visual features, prompt information is generated and prediction is performed using a large language model. Combined with a dual-domain stationarity algorithm and multi-scale feature fusion, the signal quality and the model's understanding ability are improved.

Benefits of technology

It improves the prediction accuracy and robustness of rPPG signals in complex environments, enhances robustness to illumination changes and motion artifacts, and achieves higher measurement accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412016B_ABST
    Figure CN120412016B_ABST
Patent Text Reader

Abstract

The application provides a cross-modal rPPG signal perception method and device based on a large language model and an electronic device. The method comprises the following steps: obtaining a face video segment; extracting a low-precision rPPG signal and multi-scale fusion visual features from the face video segment; generating prompt information about the face video segment and the rPPG signal; and predicting, by a large language model, the rPPG signal according to the rPPG signal, the multi-scale fusion visual features and the prompt information to obtain a high-precision rPPG signal. The application uses a large language model to comprehensively predict the rPPG signal by using multiple information sources such as the rPPG signal, the multi-scale fusion visual features and prompt information related to the rPPG signal, so that the final prediction accuracy and robustness can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically, to a cross-modal rPPG signal sensing method, apparatus, and electronic device based on a large language model. Background Technology

[0002] In the field of computer vision, vision-based remote photoplethysmography (rPPG) is an effective, non-contact technique for detecting physiological signals. This technique estimates physiological parameters (such as heart rate and respiratory rate) by analyzing optical information in facial videos. rPPG signal extraction primarily relies on the analysis of subtle changes in skin color, which are mainly reflected in the green channel in the RGB color space. Therefore, early rPPG signal extraction methods often employed color subspace transformation methods based on chromaticity and orthogonal projection planes of skin color to enhance signal discriminability. However, these methods heavily depend on manually selecting regions of interest and predefined filtering processes, making them difficult to adapt to complex lighting environments and susceptible to head movements, shadow variations, and skin color differences, resulting in insufficient robustness in rPPG signal extraction.

[0003] In recent years, the rise of deep learning technology has provided new directions for the study of rPPG signals. In particular, the introduction of convolutional neural networks (CNNs) allows researchers to learn the rPPG signal extraction process end-to-end without manually designing features, reducing reliance on manual intervention. To improve robustness to non-physiological factors (such as motion artifacts and lighting changes), researchers have explored various deep learning architectures, including 2D-CNNs, 3D-CNNs, end-to-end spatiotemporal networks (Smart Transport Network, STN), and Long Short-Term Memory (LSTM). Although these techniques have improved the accuracy and robustness of rPPG signal extraction to some extent, they still face many challenges. For example, since rPPG signals depend on minute color changes, variations in ambient lighting can cause signal distortion. Simultaneously, non-rigid facial movements (such as facial expressions and head movements) can introduce noise, interfering with the extraction of effective signals. Furthermore, although existing methods have attempted to use temporal modeling techniques such as LSTM and 3D-CNN, their ability to capture long-range dependencies remains limited, resulting in insufficient modeling performance for long-term signals.

[0004] Therefore, how to provide a novel and efficient method to improve the accuracy and robustness of rPPG physiological signals extracted in complex scenarios is a technical problem that needs to be solved. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a cross-modal rPPG signal sensing method, device, and electronic device based on a large language model, so as to improve the accuracy and robustness of rPPG physiological signals extracted in complex scenes. To achieve the above objective, the technical solution adopted by this invention is as follows:

[0006] In a first aspect, the present invention provides a cross-modal rPPG signal perception method based on a large language model, the method comprising: obtaining a face video segment; extracting a low-precision rPPG signal and multi-scale fused visual features from the face video segment; generating prompt information about the face video segment and the rPPG signal; and obtaining a high-precision rPPG signal by the large language model based on the rPPG signal, the multi-scale fused visual features and the prompt information.

[0007] Secondly, the present invention provides a cross-modal rPPG signal sensing device based on a large language model, comprising: an acquisition module for acquiring a face video segment; an extraction module for extracting low-precision rPPG signals and multi-scale fused visual features from the face video segment; a generation module for generating prompt information about the face video segment and the rPPG signal; and a prediction module for using the large language model to predict, based on the rPPG signal, the multi-scale fused visual features, and the prompt information, to obtain a high-precision rPPG signal.

[0008] Thirdly, the present invention provides an electronic device including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the cross-modal rPPG signal sensing method based on a large language model as described in any of the foregoing embodiments.

[0009] The present invention provides a cross-modal rPPG signal perception method, apparatus, and electronic device based on a large language model. First, a face video segment is obtained. Then, low-precision rPPG signals and multi-scale fused visual features are extracted from the face video segment. This generates prompts about the face video segment and the rPPG signal. Finally, the large language model predicts the rPPG signal based on the rPPG signal, multi-scale fused visual features, and prompts to obtain a high-precision rPPG signal. Compared to traditional methods, which are often limited by spatiotemporal feature extraction and long-term dependency processing in complex scenes when modeling rPPG signals, resulting in low measurement accuracy in complex environments, the present invention constructs specific prompts for the rPPG task to enhance the large language model's understanding of the rPPG task and its generalization ability in complex environments. Finally, the large language model integrates multiple information sources, including the rPPG signal and multi-scale fused visual features, for prediction, thereby improving the final prediction accuracy and robustness.

[0010] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic flowchart illustrating the cross-modal rPPG signal sensing method based on a large language model provided in an embodiment of the present invention;

[0013] Figure 2 This is a functional module diagram of the multi-scale interaction module provided in an embodiment of the present invention;

[0014] Figure 3 This is a schematic diagram illustrating the generation process of three types of prompt descriptions provided in embodiments of the present invention;

[0015] Figure 4 This is an overall schematic diagram of the cross-modal rPPG signal sensing process provided in an embodiment of the present invention;

[0016] Figure 5 A functional block diagram of a cross-modal rPPG signal sensing device based on a large language model provided in an embodiment of the present invention;

[0017] Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0020] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] Considering the shortcomings of existing deep learning technologies in terms of accuracy and stability during cross-modal rPPG signal perception, this invention provides an rPPG signal perception technology framework that integrates deep learning technology and a large language model, which can effectively solve the defects caused by signal modeling, noise interference, temporal consistency and insufficient generalization ability.

[0022] Please see Figure 1 , Figure 1 This is a schematic flowchart of a cross-modal rPPG signal sensing method based on a large language model provided in an embodiment of the present invention. The execution subject of this method can be an electronic device, and includes steps S101 to S104, as described below:

[0023] S101: Obtain facial video footage;

[0024] S102: Extracting low-precision rPPG signals and multi-scale fused visual features from face video clips;

[0025] S103: Generate prompts about the face video clip and rPPG signal;

[0026] S104: A high-precision rPPG signal is obtained by the large language model based on the rPPG signal, multi-scale fused visual features and cue information.

[0027] In the scheme of steps S101 to S104 provided in this embodiment of the invention, a face video segment is first obtained. Then, low-precision rPPG signals and multi-scale fused visual features are extracted from the face video segment. This process generates prompt information about the face video segment and the rPPG signal. Finally, a large language model predicts based on the rPPG signal, multi-scale fused visual features, and prompt information to obtain a high-precision rPPG signal. Compared with traditional methods, which are often limited by spatiotemporal feature extraction and long-term dependency processing in complex scenes when modeling rPPG signals, resulting in the inability to maintain high measurement accuracy in complex environments, this embodiment of the invention constructs specific prompt information for the rPPG task to enhance the powerful language model's understanding of the rPPG task and its generalization ability in complex environments. Finally, the large language model integrates multiple information sources such as the rPPG signal and multi-scale fused visual features for prediction, which can improve the final prediction accuracy and robustness.

[0028] Next, the embodiments of the present invention will provide a detailed and clear description of the above-described rPPG signal sensing process in conjunction with the accompanying drawings.

[0029] In step S101, the device can first obtain a face video clip for extracting the rPPG signal. The face video clip can come from video data obtained by the device. This video data can be video captured by the electronic device's own image acquisition device, video stored locally, or video received in real time. This embodiment of the invention does not limit the specific data.

[0030] After acquiring video data, the device can perform preprocessing to obtain facial video clips. This embodiment of the invention provides a preprocessing procedure, including steps a1 to a4, as described below:

[0031] Step a1: Extract a fixed-size rectangular region from the original video; this region covers the face and part of the background, while reducing the impact of irrelevant background noise on subsequent analysis;

[0032] In this embodiment of the invention, a multi-task cascaded convolutional network (MTCNN) algorithm can be used, but is not limited to, to detect the face region in the first frame and output a rectangular box coordinate (i.e., the face rectangle) to represent the position range of the face in the image.

[0033] Step a2: Expand the rectangular area mentioned above to a fixed-size area so that the area not only includes the entire face but also retains some background information;

[0034] Step a2 provides more contextual information, which helps to better understand the impact of environmental changes (such as illumination, shading, etc.) on rPPG signals.

[0035] Step a3: Crop the expanded face rectangle in the remaining video frames, and scale the cropped image to a uniform size.

[0036] In this embodiment of the invention, the size can be flexibly set by relevant personnel, such as 128×128 pixels. In this way, the input data can be standardized so that all video clips have the same resolution, which is convenient for subsequent processing by deep learning models.

[0037] Step a4: Reassemble the processed image sequence into a new video clip, which will serve as the face video clip.

[0038] In this embodiment of the invention, the length of the face video segment can be flexibly set by relevant personnel according to actual needs. For example, the face video segment consists of 128 consecutive frames, with a single frame width and height of 128, to ensure sufficient timing information for rPPG signal extraction.

[0039] After obtaining the face video segment in step S101, the rPPG signal can be extracted from it for the first time, as shown in step S102. In step S102, this embodiment of the invention can not only extract the rPPG signal, but also extract multi-scale visual features. These preliminary extraction results will serve as the basis for subsequent signal optimization.

[0040] In one embodiment of the present invention, step S102 may include steps b1 to b3, as described below:

[0041] Step b1: Extract rPPG signals and multi-scale visual features from face video clips using the trained deep learning model;

[0042] Step b2: Perform time-domain and frequency-domain weighted smoothing on the rPPG signal;

[0043] Step b3: Fuse the multi-scale visual features to obtain multi-scale fused visual features.

[0044] In step b1, to ensure efficiency and quality, a deep learning network can be used to perform steps b1 to b3. Optionally, the deep learning network can be, but is not limited to, the PhysNet model. The PhysNet model consists of a video encoder and a decoder. The video encoder is responsible for extracting high-level spatiotemporal features from the input face video clips. These features capture information such as color changes and motion patterns in the video, which are closely related to physiological signals. The decoder generates a preliminary rPPG signal prediction result based on the features extracted by the video encoder. Since PhysNet is a pre-trained model, its output rPPG signal usually has a certain degree of noise and instability, and is therefore referred to as a "low-quality signal." The feature set of the intermediate M layers of the encoder consists of multi-scale visual features.

[0045] Based on this, the rPPG signal sequence obtained in the embodiments of the present invention can be represented as: Where B is the batch size of the data, and T is the length of the face video segment, such as 128. This sequence contains the predicted physiological signal value at each time point. Although the quality is low, it contains important temporal information. Multi-scale visual features can be represented as: Wherein, each feature f i This represents information at a specific scale. H×W refers to the height and width of a video frame. These feature sets contain rich spatiotemporal information, providing a deeper feature representation for subsequent signal prediction and helping to further improve signal quality.

[0046] Low-quality rPPG signals typically contain significant noise due to factors such as illumination variations and motion artifacts, resulting in poor signal quality. To address this, this invention proposes a novel dual-domain stabilization algorithm, employing weighted smoothing operations in both the time and frequency domains, which effectively removes these noise components and improves signal quality.

[0047] In the time domain: the sequence of the rPPG signal obtained in step b1 is represented as follows Time-domain stabilization can be performed on it. Specifically: first, the mean μ and variance σ of the rPPG sequence x can be calculated, and then x can be standardized to obtain x′, as shown in the following formula:

[0048]

[0049] For simplicity, the above standardization process will be denoted as S(·) in the following text. To ensure global stationarity, a time-domain smoothing operation can be performed as shown in Equation (2):

[0050]

[0051] Where α is the smoothing factor; x i ′ represents the standardized x in the rPPG sequence i , Indicates x i Time-domain stabilization results; z i-1 Indicates x′ i-1 The result after performing time-domain smoothing. In the following text, these smoothing operations will be collectively referred to as... Therefore, the operation of the above formula (2) can be simplified to: z time It is the time-domain stabilization result of the rPPG signal x.

[0052] By performing time-smoothing processing on the rPPG signal using the above-described methods, drastic fluctuations in a short period of time can be reduced.

[0053] In the frequency domain: This embodiment of the invention can use Discrete Wavelet Transform (DWT) to perform frequency domain decomposition to separate useful physiological signal components while suppressing noise. Discrete Wavelet Transform decomposes the input signal into approximation coefficients (ac) and detail coefficients (dc) at multiple scales. Specifically, DWT is as shown in Equation (3):

[0054] x ac ,[x dc,1 ,…,x dc,J ]=DWT(x) (3)

[0055] Where J is the decomposition level (e.g., J = 3); x ac With x dc,1..J Let represent the approximate coefficients and detail coefficients of each decomposition level of the discrete wavelet transform decomposition, respectively. These coefficients represent different frequency components of the signal. The decomposed coefficients are normalized to eliminate amplitude differences. Then, the processed coefficients are recombined using the inverse discrete wavelet transform (IDWT) to reconstruct a smooth frequency domain representation, as shown in Equation (4).

[0056]

[0057] Among them, z fre It is the result of frequency domain stabilization of the rPPG signal.

[0058] To combine the advantages of time-domain and frequency-domain processing, this embodiment of the invention also introduces an adaptive weighting mechanism. Therefore, the final smoothed output z of the rPPG signal sequence is calculated using the following formula:

[0059] z = (1-β)·ztime +β·z fre (5)

[0060] Here, β∈[0,1] is a learnable parameter, ranging from [0,1]. When β is large, more emphasis is placed on the frequency domain stabilization result; when β is small, more emphasis is placed on the time domain stabilization result; z is the result of weighted smoothing operations in both the time and frequency domains of the rPPG signal sequence x.

[0061] This invention employs time-domain stabilization to better capture these periodic features, ensuring signal temporal consistency. Frequency-domain stabilization helps capture signal characteristics across different frequency ranges while suppressing high-frequency noise. Finally, an adaptive weighted fusion mechanism combines the results of both approaches to generate a more stable and reliable rPPG signal, laying a solid foundation for further optimization.

[0062] Furthermore, regarding the multi-scale visual features extracted in step b1, Shallow features (such as f1) mainly capture local details and high-frequency changes, while deep features (such as f) M This focuses more on global structure and low-frequency changes. In order to effectively integrate multi-scale visual features from different modalities, this embodiment of the invention fuses the multi-scale visual features to obtain multi-scale fused visual features, i.e., step b3.

[0063] In step b3, this embodiment of the invention first fuses multi-scale shallow features to obtain finer-grained visual features; then, it applies a self-attention mechanism to the fused shallow features to enhance the representation of internal dependencies within the features; finally, this embodiment of the invention weightedly fuses deep and shallow features to obtain multi-scale fused visual features. Therefore, step b3 can be divided into the following processes:

[0064] The first step is to extract shallow and deep features from multi-scale visual features;

[0065] In this step, embodiments of the present invention can use a set of feature projection layers to... Features at different scales are mapped to a unified space, namely F. i =Projection(f i ,l target ), where l target This is the compressed length setting for subsequent processing.

[0066] The second step is to extract and fuse shallow features using deep features to obtain fused shallow features.

[0067] To achieve strong semantic alignment between deep features and the text space, this embodiment of the invention employs a cross-attention mechanism, using deep features as the query to dynamically extract missing details from shallow features. This mechanism effectively establishes associations between features at different scales, enhancing the richness of feature representations. The fused shallow features can be represented as:

[0068] F cross =CrossAttention(F M ,X,X) (6)

[0069] Where CrossAttention represents the cross-attention mechanism; X = Concat(f1, f2, ..., f M-1 Concat represents a concatenation operation. To enhance the representation by capturing the internal dependencies in the cross-participation features and further improve the feature representation capability, a self-attention mechanism is applied to the shallow features obtained in equation (5), thereby merging them to obtain the feature mapping of the shallow features as shown in equation (6):

[0070] F self =SelfAttention(F cross (7)

[0071] The third step is to combine the deep features and the fused shallow features to obtain multi-scale fused visual features.

[0072] Finally, by weighted fusion, the outputs of the cross-attention mechanism and the self-attention mechanism are weighted by two learnable parameters and then fused with the deep features to obtain the final multi-scale fused visual features, as shown in Equation (8):

[0073] F visual =F M +γ2·(F cross +γ1·F self (8)

[0074] Where γ1 and γ2 are two learnable parameters used to control the contribution ratio of the two mechanisms, F visual For multi-scale fusion of visual features; F M This represents a deep feature.

[0075] By fusing multi-scale visual features through the above implementation methods, the overall capability of feature representation can be enhanced, redundant information can be removed while key features are highlighted, and a high-quality feature foundation can be provided for further optimization.

[0076] In one embodiment of the present invention, a multi-scale interaction module is also designed to complete the above-mentioned multi-scale visual feature fusion process. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a functional module diagram of the multi-scale interaction module provided in this embodiment of the invention. It adopts a hierarchical attention structure to dynamically synthesize multi-scale feature representations. Specifically: First, a set of feature projection layers is used to project multi-scale visual features. Then, the projected multi-scale shallow features are concatenated, i.e., a concatenation operation. A cross-attention mechanism is used to fuse deep features and multi-scale shallow features after concatenation. At the same time, a self-attention mechanism is applied to the fused visual features to enhance the representation of internal dependencies of features. Finally, the shallow features obtained under the cross-attention mechanism and the self-attention mechanism and the fused deep features are linearly weighted and fused to obtain multi-scale fused visual features.

[0077] The above implementation method yields the rPPG signal z after dual-domain weighted smoothing and the multi-scale fused visual feature F. visual Subsequently, in order to improve the understanding of the large language model for the rPPG task and ensure the accuracy of the prediction results, this embodiment of the invention also generates prompt information about the face video clip and the rPPG signal through step S103.

[0078] In step S103, to better adapt the Large Language Model (LLM) to the rPPG task, this embodiment of the invention provides prompts that match its input format. These prompts include not only descriptions of the visual content but also task characteristics and signal statistics. In this way, the semantic understanding capability of the Large Language Model (LLM) for the task can be significantly improved. Therefore, for step S103, this embodiment of the invention provides an implementation method including steps d1 to d3:

[0079] Step d1: Obtain various types of prompt descriptions.

[0080] In this embodiment of the invention, the prompt description has three types: rPPG task description, statistical feature description of rPPG signal, and visual description corresponding to the intermediate frame of face video segment.

[0081] For visual description, key visual information can be extracted from the video, such as lighting conditions, head posture, and facial expression changes. Specifically, a center frame can be extracted from a facial video clip, and questions related to physiological signals can be obtained, such as "Is the current ambient lighting stable?" or "Is there a significant head rotation?" The center frame and prompts are then input into a frozen multimodal large model (LLaVA). The LLaVA automatically generates and outputs a description in natural language, such as "In the current video, there is a slight shadow on the face, and the head is rotated about 15 degrees to the right." This description is then tokenized to obtain the visual description.

[0082] C vision =Tokenizer(LLaVA(I t ,Q)) (9)

[0083] Where Q is a cue question related to physiological cues, I t It is the center frame, C vision It is a visual token; Tokenizer is a text processing function used to convert text or textual representations into tokens that the model can understand.

[0084] For rPPG task descriptions: the task can be formalized using prior knowledge of task characteristics consistently described in the rPPG literature. This prior knowledge is encoded into LLMs-compatible notation, i.e., the task description: C task =Tokenizer(P), where P refers to this prior knowledge. Specifically: Literature related to the rPPG task can be collected, the task characteristics consistently described in it can be extracted, these characteristics can be converted into a text format that LLMs can understand and input into the multimodal large model LLaVA. The model outputs a natural language text describing the task characteristics, such as "The main frequency range of the rPPG signal is 0.5 to 3 Hz, which may be affected by motion artifacts and changes in illumination".

[0085] For the statistical feature description of rPPG signals, the rPPG signals from the pre-trained model are statistically analyzed. The statistical information is obtained, such as calculating basic statistics like the minimum, maximum, and median of the signal, then analyzing the trend changes of the signal (calculated through first-order differencing), extracting the top K significant feature points of the signal (such as the Top-5 peak positions), and then converting the above statistical results into a natural language description. Specifically:

[0086] For rPPG signals For each batch number b∈B, the signal-related information S can be calculated, as shown in formula (10):

[0087]

[0088] Where median(·) refers to the median of all values ​​of signal x in batch b; sgn(sign function) is the sign function, which extracts the sign of a number; T is the signal length, which is also the length of the face video segment; t and t-1 are both signal points; Γ(·) calculates the signal trend through first-order difference, and the calculation formula is: x i It is the observation value at time point i; x i-1 It is the observation value at the previous time point (i-1).

[0089] Finally, the labeling process for statistical feature description is formulated as C static =Tokenizer(S), where.

[0090] The above implementation methods generate visually relevant cues, task-related cues, and signal sequence statistical information cues, providing comprehensive task description information for large language models. These cues not only enhance LLMs' understanding of rPPG tasks but also improve their adaptability and robustness in complex scenarios.

[0091] Step d2: Generate the token features and weights corresponding to various prompt descriptions;

[0092] In this embodiment of the invention, the prompt description obtained through step d1 can be represented as C = {C} statistics C vision C task To integrate these prompts and generate prompt information, this embodiment of the invention first studies existing multimodal fusion methods, which typically use static weighted coefficients for integration. During the research process, the inventors discovered that this fusion method is limited by two factors: (1) a fixed combination ratio cannot adapt to learnable LLMs; and (2) it cannot adaptively select key cues across different data. Therefore, this embodiment of the invention first generates C using the following method. statistics C vision C task The corresponding attention-compressed token features:

[0093] The first step is to build a compressor for each type of prompt description;

[0094] The second step involves extracting key features from the cue description using a compressor and generating attention-compressed token features.

[0095] In this embodiment of the invention, each prompt description is processed by a compressor to compress the high-dimensional features into a unified dimension, as shown in formula (11):

[0096] ε k=AttentiveCompressor k (C k ),k∈{task,vision,statistic} (11)

[0097] Here, task, vision, and statistic represent three different modalities or task types of identifiers, respectively. task represents task-related features, vision represents vision-related features, and statistic represents statistical features. k iterates through each identifier and generates attention-compressed token representations, denoted as ε. k To analyze the characteristics of each identifier, each identifier corresponds to a compressor called AttentiveCompressor. k The compressor obtains the token feature ε corresponding to each cue description through a temporal context attention mechanism. k As shown in formula (12):

[0098]

[0099] Among them, Q k′ K k′ V k′ From C by linear projection k The exported query, key, and value matrix, where T′ refers to the transpose and d′ represents the key vector K. k′ Dimensions.

[0100] To better understand the process of generating the above prompt description, please refer to [link / reference]. Figure 3 , Figure 3 The diagram illustrates the generation process of three types of cue descriptions provided in embodiments of the present invention. (a) represents extracting visual priors (e.g., illumination, facial expressions, occlusion) via LLaVA and encoding them into visual representations. (b) represents symbolizing the textual description of the rPPG task to derive task-specific priors. (c) represents analyzing the statistical features of rPPG signals from the backbone network to generate statistical prior tokens. These cues integrate visual, semantic, and statistical priors to enhance physiological signal analysis.

[0101] Considering that key cues may differ in different scenarios during cross-modal remote physiological signal perception tasks, for example: visual cues may be more important in scenarios with significant changes in illumination; task-related cues may be more helpful in helping the model understand task requirements in complex task environments; and signal sequence statistics may play a greater role when analyzing data with poor signal quality. Therefore, this embodiment of the invention designs a flexible mechanism that can dynamically adjust the dependence on different types of cues according to specific scenarios, rather than simply using a fixed weight combination. That is, the weights of different cues are dynamically adjusted through a learnable weight mechanism, as shown in formula (13):

[0102]

[0103] Where L is a predefined hyperparameter relating the length of the target cue sequence, and d is the label embedding dimension.

[0104] W task W vision W statistic These represent the weight parameter matrices corresponding to the three prompt features. The weight parameter matrices are learnable and can automatically adjust the weights during training to adapt to the needs of different scenarios.

[0105] Step d3: Based on the weights, fuse the token features compressed from each attention to obtain the prompt information.

[0106] Finally, the three prompt features are weighted and fused using formula (13) to obtain the prompt information required in this embodiment of the invention. The fused prompt information (also known as the prompt token T) cue The form is as shown in formula (14):

[0107]

[0108] Here, ⊙ refers to the product weighting.

[0109] The above implementation method fuses the three types of prompt features by weighting them with a weight matrix to generate the final prompt feature representation. This enables adaptive selection and combination of information from different modalities, allowing the model to flexibly select the most important prompt information according to the specific scenario, thereby improving the model's generalization ability and enabling it to maintain high performance in various environments.

[0110] Next, this embodiment of the invention will use a large language model to apply the multi-scale fused visual features F obtained in step S102 of the above implementation method. visual and the rPPG signal sequence z and the prompt information T obtained in step S103 cue To perform high-precision rPPG signal prediction.

[0111] In step S104, the large language model selected in this embodiment of the invention can be any existing large language model. During rPPG signal prediction using LLMs, considering that the data processing object of LLMs is text, in order to unify the features of the rPPG signal with the language feature space of LLMs, it is first necessary to convert the rPPG signal and visual features into the semantic space of LLMs. To this end, this embodiment of the invention provides a Text Prototype Guidance (TPG) strategy, as shown in the following implementation:

[0112] Step e1: Obtain the text prototype library related to the rPPG signaling task;

[0113] Step e2: Construct a text prototype bootstrapping module using multiple different transformer layers;

[0114] Step e3: The text prototype guidance module converts the rPPG signal and multi-scale fused visual features into semantic representations based on the text prototype library.

[0115] In step e1, this embodiment of the invention uses a text prototype library. The process guides the transformation of rPPG signals and visual features, where V is the vocabulary size in the text prototype set, and D is the word vector length corresponding to each word. Considering that directly using a large-vocabulary text prototype library E would increase computational complexity, this invention proposes a simplified scheme: maintaining a small-scale text prototype library through linear projection, represented as... Where V′ << V. Using a small text prototype library E′ as a bridge, visual features and temporal features can be quickly mapped to the semantic space familiar to LLMs.

[0116] To achieve semantic space transformation, in step c2, this embodiment of the invention utilizes multiple different transformer layers to construct a text prototype guidance module, which enhances the interaction between rPPG signals, visual features, and the language model. The working process of the text prototype guidance module can be understood as follows:

[0117] In this embodiment of the invention, the working process of the text prototype guidance module is as follows: given an input A, the text prototype guidance module can perform transformation operations on the input A sequentially as shown in formulas (15) to (19):

[0118] A self =SelfAttention(A) (15)

[0119] E′ fusion =E′+A self (16)

[0120] y2(E′fusion1 A) = CrossAttention(E′) fusion ,A,A), (17)

[0121] E′ fusion2 =E′ fusion1 +y2(E′ fusion1 ;A), (18)

[0122] T out =FFN(E′) fusion2 (19)

[0123] Among them, A self It refers to the features generated by self-attention from input A, where FFN(·) refers to the linear transformation performed through a feedforward neural network; E′ fusion E′ fusion2 y1 is an intermediate result of A, and y2 is a coefficient. For simplicity, the conversion process of the above text prototype guidance module can be represented as: T out =TPG(A).

[0124] Therefore, embodiments of the present invention can combine the stationary signal z obtained by the dual-domain stationarity algorithm and the fused multi-scale visual features F obtained by multi-scale interaction. visual Input the text prototype guide module respectively, thereby realizing the control of z and F. visual Semantic space transformation. Here, in this embodiment of the invention, the transformed z and F visual These are called "signal tokens" and "visual tokens," denoted as T. signal and T vision Therefore, the visual token output by the model guided by the text prototype can be represented as T. vision =TPG(F visual The signal token can be represented as T. signal =TPG(z).

[0125] By mapping rPPG signals and visual features to an LLM-interpretable semantic space through the above implementation methods, cross-modal alignment is achieved and semantic understanding is enhanced. This enables the system to integrate multiple information sources more efficiently, allowing visual, temporal, and textual information to be effectively aligned, thereby improving the efficiency of cross-modal learning.

[0126] Obtain the video token T using the above method. vision and signal token T signal Then, combine them with the prompt token T generated in step S103. cue When input together into a pre-selected large language model, the model can predict high-precision rPPG signals.

[0127] In one embodiment of the present invention, in order to ensure that the prediction results of the model are as close as possible to the true values, the mean squared error (MSE) shown in formula (20) can also be used as the total loss for training the large language model, so as to make the model's rPPG signal prediction results more and more accurate.

[0128]

[0129] Where T is the signal length, corresponding to the frame length of the face video segment, and y is the actual rPPG signal label.

[0130] To facilitate a comprehensive understanding of the above-described perception process of cross-modal rPPG signals based on a large language model, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram illustrating the overall process of cross-modal rPPG signal sensing provided in an embodiment of the present invention. Figure 4 In this process, face video clips are processed by a visual encoder and predictor to obtain low-quality rPPG signals and multi-scale visual features. The low-quality rPPG signals are then processed by a dual-domain stabilization algorithm to obtain stable rPPG signals; the multi-scale visual features are processed by a visual aggregation module to obtain multi-scale fused visual features. The stable rPPG signals and multi-scale fused visual features are then transformed by a TPG module. After feature interaction by the TPG module, the features are combined with a text prototype library to obtain semantic representations of the rPPG signals and multi-scale fused visual features in the semantic space of the large language model, namely signaling tokens and visual tokens. Furthermore, using the center frame image of the face video clips and the designed prompt questions, the large language model generates three types of prompt descriptions: task prompts, visual prompts, and statistical feature prompts. These prompt descriptions are then segmented by a word segmenter to obtain three types of prompt features. Finally, through an adaptive prompt learning process, the three types of prompt features are fused to obtain the final prompt information, i.e., the prompt token. Finally, the cue token, visual token, and signaling are input into the large semantic model for prediction, and then the projector projects the prediction results of the large language model into a signal waveform.

[0131] Combination Figure 4 As can be seen from the above, the embodiments of the present invention integrate Large Language Models (LLMs) and Convolutional Neural Networks (CNNs) in the field of remote physiological signal sensing, forming a cross-modal remote physiological signal sensing framework. Compared with traditional methods, which are often limited by the extraction of spatiotemporal features in complex scenes and the processing of long-term dependencies when modeling rPPG signals, resulting in the inability to maintain high measurement accuracy in complex environments, the embodiments of the present invention have the following technical advantages:

[0132] First, this invention introduces a dual-domain stabilization mechanism to enhance the temporal stability and anti-interference capability of the rPPG signal. This mechanism optimizes the rPPG signal simultaneously in the time and frequency domains through an exponential decay adaptive coefficient modulation strategy, ensuring the periodic consistency of time-series data while reducing noise interference. This method effectively improves the robustness of rPPG measurements under interference factors such as illumination changes and motion artifacts.

[0133] Secondly, this invention constructs specific prompts for the rPPG task, utilizing learnable word vectors to introduce prompts such as physiological statistics, environmental factors, and task descriptions to enhance LLMs' understanding of the rPPG task. This mechanism enables the framework to dynamically adapt to different scenarios, improve generalization ability in complex environments, and enhance the model's environmental adaptability.

[0134] Furthermore, this invention proposes a text prototype guidance strategy to achieve cross-modal alignment and enhance the semantic understanding capability of rPPG signals. The TPG strategy projects hemodynamic features onto the semantic space interpretable by LLMs, enabling effective alignment of visual, temporal, and textual information, thereby improving the efficiency of cross-modal learning. This strategy significantly reduces the modal differences between rPPG processing and LLMs processing, allowing the system to more efficiently integrate multiple information sources and improve the final prediction accuracy.

[0135] Experimental results from the embodiments of this invention demonstrate that the proposed framework achieves superior performance in cross-domain tasks, exhibiting strong adaptability and robustness, particularly in complex environments. Through large-scale experiments on various datasets and real-world scenarios, the embodiments of this invention verify the effectiveness of the aforementioned perception framework in cross-modal remote physiological signal sensing tasks. Even when faced with interference factors such as illumination variations and motion artifacts, the aforementioned perception framework maintains high-precision rPPG signal extraction. Compared to traditional methods, the aforementioned perception framework demonstrates stronger generalization ability across different datasets, particularly its stability and accuracy on cross-domain data, further highlighting its advantages in real-world complex scenarios.

[0136] In summary, the embodiments of the present invention significantly improve the accuracy and robustness of cross-modal remote physiological signal sensing through a combination of technical means such as dual-domain stability algorithms, generation of prompt descriptions, and text prototype guidance strategies, providing new ideas and solutions for the application of rPPG technology in complex environments.

[0137] To perform the corresponding steps in the above embodiments and various possible methods, an implementation of a cross-modal rPPG signal sensing device based on a large language model is given below. Please refer to [link to relevant documentation]. Figure 5 , Figure 5This is a functional block diagram of a cross-modal rPPG signal sensing device based on a large language model provided in this embodiment of the invention. It should be noted that the basic principle and technical effects of the cross-modal rPPG signal sensing device based on a large language model provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments. The cross-modal rPPG signal sensing device 50 based on a large language model includes: an acquisition module 501, an extraction module 502, a generation module 503, and a prediction module 504.

[0138] Module 501 is used to acquire face video clips;

[0139] Extraction module 502 is used to extract low-precision rPPG signals and multi-scale fused visual features from the face video clip;

[0140] The generation module 503 is used to generate prompt information about the face video clip and the rPPG signal;

[0141] The prediction module 504 is used to predict the rPPG signal with high accuracy by the large language model based on the rPPG signal, the multi-scale fused visual features and the prompt information.

[0142] It is understandable that the acquisition module 501, extraction module 502, generation module 503, and prediction module 504 can be executed collaboratively. Figure 1 The various steps in the process are used to achieve the corresponding technical effects. The acquisition module 501, extraction module 502, generation module 503 and prediction module 504 can also be used to perform other steps in the above embodiments, which will not be described in detail here.

[0143] It should be noted that the module division in the above embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical entities, or have two or more units integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing executable program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0145] This invention also provides an electronic device, please refer to [link to relevant documentation]. Figure 6 , Figure 6 The structural block diagram of the electronic device provided in the embodiment of the present invention includes: a memory 601, a processor 602, and a communication interface 603. The memory 601, the processor 602, and the communication interface 603 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0146] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0147] In this embodiment of the invention, the processor 602 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The software modules may be located in the memory 601, and the processor 602 reads the program instructions from the memory 601 and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0148] In this embodiment of the invention, the memory 601 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as RAM. The memory can also be any other medium capable of carrying or storing desired executable program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in this embodiment of the invention can also be a circuit or any other device capable of implementing a storage function for storing instructions and / or data.

[0149] The memory 601 can be used to store software programs and modules, such as the instructions / modules of the cross-modal rPPG signal sensing device 50 based on a large language model provided in this embodiment of the invention. These can be stored in the memory 601 in the form of software or firmware, or in the operating system (OS) of the embedded electronic device 60. The processor 602 executes various functional applications and data processing by executing the software programs and modules stored in the memory 601. The communication interface 603 can be used to communicate with other node devices for signaling or data.

[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0151] Understandable. Figure 6 The structure shown is for illustrative purposes only; the electronic device 60 may also include components that are more advanced than those shown. Figure 6 The more or fewer components shown, or having the same Figure 6 The different configurations shown. Figure 6 The components shown can be implemented using hardware, software, or a combination thereof.

[0152] Based on the above embodiments, this application also provides a storage medium in which a computer program is stored. When the computer program is executed by a computer, the computer executes the cross-modal rPPG signal sensing method based on a large language model provided in the above embodiments.

[0153] Based on the above embodiments, this invention also provides a computer program that, when run on a computer, causes the computer to execute the cross-modal rPPG signal sensing method based on a large language model provided in the above embodiments.

[0154] Based on the above embodiments, this invention also provides a chip for reading a computer program stored in a memory and executing the cross-modal rPPG signal sensing method based on a large language model provided in the above embodiments.

[0155] This invention also provides a computer program product, including instructions that, when run on a computer, cause the computer to execute the cross-modal rPPG signal sensing method based on a large language model provided in the above embodiments.

[0156] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by instructions. These instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0159] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A cross-modal rPPG signal sensing method based on a large language model, characterized in that, The method includes: Obtain facial video clips; Low-precision rPPG signals and multi-scale fused visual features are extracted from the facial video clips. Generating prompt information about the face video clip and the rPPG signal includes: obtaining multiple types of prompt descriptions; wherein the prompt descriptions include an rPPG task description, a statistical feature description of the rPPG signal, and a visual description corresponding to the intermediate frames of the face video clip; generating token features and weights corresponding to each of the prompt descriptions; and fusing the token features based on the weights to obtain the prompt information. The large language model makes predictions based on the rPPG signal, the multi-scale fused visual features, and the cue information to obtain a high-precision rPPG signal.

2. The cross-modal rPPG signal sensing method based on a large language model according to claim 1, characterized in that, The high-precision rPPG signal is obtained by the large language model based on the rPPG signal, the multi-scale fused visual features, and the cue information, including: The rPPG signal and the multi-scale fused visual features are mapped to the semantic space of the large language model to obtain a semantic representation; The semantic representation and the prompt information are input into the large language model for prediction to obtain the high-precision rPPG signal.

3. The cross-modal rPPG signal sensing method based on a large language model according to claim 2, characterized in that, Mapping the rPPG signal and the multi-scale fused visual features onto the semantic space of the large language model yields a semantic representation, including: Obtain a text prototype library related to the rPPG signaling mission; A text prototype bootstrapping module is constructed using multiple different transformer layers; The text prototype guidance module converts the rPPG signal and the multi-scale fused visual features into the semantic representation based on the text prototype library.

4. The cross-modal rPPG signal sensing method based on a large language model according to claim 1, characterized in that, Generate token features corresponding to various prompt descriptions, including: Build a compressor for each type of prompt description; The compressor extracts key features from the cue description and generates attention-compressed token features.

5. The cross-modal rPPG signal sensing method based on a large language model according to claim 1, characterized in that, Extracting low-precision rPPG signals and multi-scale fused visual features from the facial video clips includes: The rPPG signal and multi-scale visual features are extracted from the face video segment using a trained deep learning model; The rPPG signal is subjected to time-domain and frequency-domain weighted smoothing operations; The multi-scale visual features are fused to obtain the multi-scale fused visual features.

6. The cross-modal rPPG signal sensing method based on a large language model according to claim 5, characterized in that, The multi-scale visual features are fused to obtain the multi-scale fused visual features, including: Extract the shallow and deep features from the multi-scale visual features; The shallow features are extracted and fused using the deep features to obtain the fused shallow features; The deep features and the fused shallow features are combined to obtain the multi-scale fused visual features.

7. The cross-modal rPPG signal sensing method based on a large language model according to claim 6, characterized in that, After extracting and fusing the shallow features using the deep features to obtain the fused shallow features, the method further includes: A self-attention mechanism is applied to the fused shallow features.

8. A cross-modal rPPG signal sensing device based on a large language model, characterized in that, include: The acquisition module is used to obtain facial video clips; The extraction module is used to extract low-precision rPPG signals and multi-scale fused visual features from the face video clips. A generation module is used to generate prompt information about the face video clip and the rPPG signal, including: obtaining multiple types of prompt descriptions; wherein the prompt descriptions include an rPPG task description, a statistical feature description of the rPPG signal, and a visual description corresponding to the intermediate frames of the face video clip; generating token features and weights corresponding to various prompt descriptions; and fusing the token features based on the weights to obtain the prompt information. The prediction module is used to predict the rPPG signal with high accuracy by the large language model based on the rPPG signal, the multi-scale fused visual features and the prompt information.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the cross-modal rPPG signal sensing method based on a large language model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Non-contact user heart rate monitoring method and device, electronic equipment and storage medium

    CN119360422A

  • Remote physiological signal estimation method and system based on diffusion model

    CN119670022A