A micro-motion health detection system based on mobile phone camera video

CN122556989APending Publication Date: 2026-08-14BEIJING AEROSPACE SCIENCE & TECHNOLOGY INNOVATION HEALTH MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]现有的非接触健康检测技术主要是基于普通图像分析的面部识别健康评估方法,主要通过提取面部颜色、纹理等静态特征进行粗略估计,但缺乏对细微生理动态变化的捕捉能力,检测精度有限,且易受环境光照、肤色差异等因素干扰,导致检测结果的稳定性和准确性难以保证

Benefits of technology

[0032] The advantages of this invention compared to the prior art are: this invention uses a mobile terminal camera as the only hardware entry point, and users do not need to wear any devices. They can complete the detection by sitting still for 30 seconds, making the barrier to entry extremely low.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122556989A_ABST
    Figure CN122556989A_ABST
Patent Text Reader

Abstract

This invention discloses a micro-motion feature health detection system based on mobile phone camera video, including a video acquisition module, a muscle micro-motion tracking module, an rPPG signal extraction module, a multimodal signal fusion module, and an AI health indicator generation module. The video acquisition module extracts temporal micro-motion signal sequences from various muscle regions; the rPPG signal extraction module extracts remote photoplethysmography (rPPG) signals; the multimodal signal fusion module performs spatiotemporal alignment and cross-validation between the temporal micro-motion signal sequences and the rPPG signals to generate a fused physiological feature vector; and the AI ​​health indicator generation module inputs the fused physiological feature vector into a pre-trained deep learning model to output multidimensional health indicators. The advantages of this invention compared to existing technologies are: it provides a convenient and routine home health monitoring system based on mobile phone camera video, enabling micro-motion feature health detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence health assessment, specifically to a micro-motion health detection system based on mobile phone camera video. Background Technology

[0002] With the increasing demand for proactive health management, non-contact physiological parameter detection technology has received widespread attention due to its advantages such as not requiring the use of any devices and ease of use.

[0003] Existing non-contact health detection technologies are mainly facial recognition health assessment methods based on ordinary image analysis. They mainly make rough estimates by extracting static features such as facial color and texture, but lack the ability to capture subtle physiological dynamic changes, resulting in limited detection accuracy. Furthermore, they are easily affected by factors such as ambient lighting and skin color differences, making it difficult to guarantee the stability and accuracy of the detection results.

[0004] In addition, most existing high-precision non-contact detection solutions rely on professional imaging hardware and controlled acquisition environments, which have high deployment thresholds and weak universality. They cannot rely on ordinary consumer-grade mobile phones to achieve convenient and routine home health monitoring, and cannot meet the public's practical needs for daily self-health assessment and risk screening. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the above-mentioned technical defects and provide a micro-motion feature health detection system based on mobile phone camera video to achieve convenient and routine home health monitoring relying on ordinary consumer mobile phones.

[0006] To solve the above-mentioned technical problems, the technical solution provided by the present invention is: a micro-motion feature health detection system based on mobile phone camera video, including a video acquisition module, a muscle micro-motion tracking module, an rPPG signal extraction module, a multi-modal signal fusion module, and an AI health indicator generation module;

[0007] The video acquisition module continuously acquires video streams of the user's face and neck area at a preset frame rate using the mobile terminal's camera. The muscle micro-motion tracking module analyzes the acquired video streams of the face and neck area frame by frame to extract the temporal micro-motion signal sequence of each muscle area.

[0008] The rPPG signal extraction module extracts remote photoplethysmography (rPPG) signals based on the color channel changes of the facial skin region in the video stream of the face and neck region. The multimodal signal fusion module performs spatiotemporal alignment and cross-validation between the temporal micro-motion signal sequence and the rPPG signal to generate a fused physiological feature vector.

[0009] The AI ​​health indicator generation module will integrate physiological feature vectors into a pre-trained deep learning model to output multi-dimensional health indicators.

[0010] Preferably, the multidimensional health indicators include brain function indicators, emotional stress indicators, traditional Chinese medicine constitution indicators, and meridian and organ indicators.

[0011] Preferably, the muscle micro-motion tracking module tracks the micro-motion changes of a preset number of muscle regions on the face and a preset number of muscle regions on the neck;

[0012] in:

[0013] The facial preset contains 44 muscle areas.

[0014] The preset number of muscle areas in the neck is 38.

[0015] Preferably, the preset frame rate of the video acquisition module is 30 frames per second.

[0016] Preferably, the rPPG signal extraction module includes a region of interest (ROI) localization module, a color channel separation module, and a motion artifact suppression module;

[0017] The Region of Interest (ROI) localization module locates the facial skin area based on a facial key point detection algorithm, excluding areas affected by hair, glasses, and shadows.

[0018] The color channel separation module separates the RGB color channels of the facial skin area and extracts the green channel signal and / or red channel signal as the original rPPG signal source;

[0019] The motion artifact suppression module constructs a motion interference reference signal based on a time-domain micro-motion signal sequence, performs adaptive filtering or independent component analysis on the original rPPG signal source, and obtains a pure rPPG signal after motion artifact suppression.

[0020] Preferably, the spatiotemporal alignment and cross-validation of the multimodal signal fusion module includes temporal alignment, spatial mapping, and cross-validation;

[0021] The time-domain alignment is based on the acquisition timestamp of the video stream, and the time-domain micro-motion signal sequence is synchronized with the rPPG signal in terms of timestamp.

[0022] The spatial mapping establishes a spatial mapping relationship between facial muscle regions and facial skin ROIs, and identifies the correlation between micro-motion signals and blood flow signals at the same anatomical location;

[0023] The cross-validation includes:

[0024] When the time-domain micromotion signal sequence indicates significant movement in a specific muscle region, the weight of the rPPG signal at the corresponding spatial location is reduced, and vice versa, the weight is increased to generate a confidence-weighted fused physiological feature vector.

[0025] Preferably, the deep learning model includes a shared feature extraction layer and a task branching layer;

[0026] The task branch layer includes a brain function assessment branch, an emotional stress assessment branch, a traditional Chinese medicine constitution identification branch, and a meridian and organ analysis branch. Each branch outputs health indicators corresponding to its dimension.

[0027] Preferably, the AI ​​health indicator generation module further includes an anomaly marker, which generates an anomaly marker and performs a graded response when any multidimensional health indicator exceeds a preset normal threshold range;

[0028] The graded response includes a first-level response, a second-level response, and a third-level response;

[0029] The first-level response is triggered when a single health indicator deviates from the normal range among the multidimensional health indicators, and lifestyle adjustment suggestions are pushed out.

[0030] The secondary response is triggered when more than one of the multidimensional health indicators deviates from the normal range, thus triggering an early warning notification.

[0031] The Level 3 response is triggered when all multidimensional health indicators deviate from the normal range, prompting an emergency medical consultation suggestion.

[0032] The advantages of this invention compared to the prior art are: this invention uses a mobile terminal camera as the only hardware entry point, and users do not need to wear any devices. They can complete the detection by sitting still for 30 seconds, making the barrier to entry extremely low.

[0033] In this invention, by tracking the micro-motion changes of 44 facial muscle regions and 38 neck muscle regions, a high-resolution motion reference signal is constructed, providing a precise basis for motion artifact suppression of rPPG signals; the muscle micro-motion signals and rPPG signals are spatiotemporally aligned and cross-validated, significantly improving the signal-to-noise ratio and stability of non-contact detection.

[0034] This invention breaks through the limitations of existing non-contact detection methods that only output basic indicators such as heart rate and blood oxygen; the abnormality marking and graded response mechanism realizes the hierarchical management of health risks, providing users with full-link health protection from lifestyle advice to emergency medical treatment. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the framework structure of a micro-motion feature health detection system based on mobile phone camera video. Detailed Implementation

[0036] The present invention will now be described in further detail with reference to the accompanying drawings.

[0037] Combined with appendix Figure 1As shown, a micro-motion health detection system based on mobile phone camera video is deployed on a mobile terminal and completes all data collection and processing by calling the device's built-in camera.

[0038] The system includes a video acquisition module, a muscle micro-motion tracking module, an rPPG signal extraction module, a multimodal signal fusion module, and an AI health indicator generation module;

[0039] The video capture module continuously captures video streams of the user's face and neck area at a preset frame rate of 30 frames per second using the mobile terminal's camera.

[0040] Before data collection, guide the user into a preset collection posture, namely, sitting still with their face facing the camera, maintaining a natural expression, and ensuring that the ambient lighting conditions meet the preset brightness threshold.

[0041] During the data acquisition process, the amplitude of the user's head movement is monitored in real time. When the amplitude of the movement exceeds a preset threshold, a prompt to re-acquire the data is triggered to ensure that the video quality meets the requirements for subsequent analysis and to provide a data foundation for high-precision tracking of subtle facial muscle changes.

[0042] In use, the muscle micro-motion tracking module analyzes the acquired video streams of the face and neck regions frame by frame. The system first locates the face region using a face detection algorithm, and then divides the face into 44 muscle regions and the neck into 38 muscle regions based on a facial key point detection algorithm.

[0043] Dense optical flow or phase change analysis techniques are used to track minute displacement changes in each muscle region between consecutive frames, extracting temporal micro-motion signal sequences for each muscle region. These temporal micro-motion signal sequences reflect the displacement amplitude, vibration frequency, and phase information of each muscle region within the detection period, providing a precise motion reference for subsequent motion artifact suppression.

[0044] In one embodiment, the rPPG signal extraction module includes a region of interest (ROI) localization module, a color channel separation module, and a motion artifact suppression module. The ROI localization module locates the facial skin region based on a facial key point detection algorithm. Specifically, it selects areas rich in capillaries and with less interference from hair and shadows, such as the cheeks on both sides and the center of the forehead, as effective ROIs. It automatically excludes areas covered by eyebrows, eyes, nostrils, beard, eyeglass frames, and hair to ensure the purity of the rPPG signal source.

[0045] The color channel separation module separates the RGB color channels within the effective ROI. Since hemoglobin has characteristic absorption peaks for green and near-infrared light, the system extracts the green channel signal and / or the red channel signal as the original rPPG signal source. The green channel is sensitive to changes in heart rate, and the red channel is sensitive to changes in blood oxygenation. Using the dual channels together can improve signal richness.

[0046] The motion artifact suppression module constructs a motion interference reference signal based on the time-domain micro-motion signal sequence output by the muscle micro-motion tracking module. The original rPPG signal source and the motion interference reference signal are simultaneously input into an adaptive filter or independent component analysis algorithm to separate and remove motion-related artifact components, and output a clean rPPG signal with motion artifact suppression.

[0047] Since the motion reference signal is directly derived from facial muscle micro-motion tracking, it shares the same spatiotemporal characteristics as the rPPG signal, and its artifact suppression accuracy is significantly better than traditional methods based on global motion estimation.

[0048] In one embodiment: the multimodal signal fusion module performs spatiotemporal alignment and cross-validation of the temporal micro-motion signal sequence and the pure rPPG signal to generate a fused physiological feature vector. The temporal alignment is based on the acquisition timestamp of the video stream, and the temporal micro-motion signal sequence and rPPG signal of each muscle region are synchronized by frame to ensure that the two types of signals correspond within the same time window.

[0049] Spatial mapping establishes the spatial mapping relationship between facial muscle regions and facial skin ROIs. Based on the facial anatomical coordinate system, each muscle region is spatially associated with adjacent skin ROIs, identifying the correlation between muscle micromotion signals and blood flow signals at the same anatomical location. For example, the micromotion changes in the frontalis muscle region directly correspond spatially to the rPPG signal of the forehead skin ROI.

[0050] Cross-validation includes: calculating the instantaneous motion amplitude of the time-domain micro-motion signal sequence of each muscle region; when the motion amplitude of a specific muscle region exceeds a preset threshold, it is determined that there is significant motion in that region, and the fusion weight of the rPPG signal at the corresponding spatial location is reduced.

[0051] Conversely, when the motion amplitude is below the threshold, the region is determined to be in a relatively static state, and the fusion weight of the corresponding rPPG signal is increased.

[0052] This adaptive confidence-weighted mechanism generates a confidence-weighted fused physiological feature vector, effectively avoiding the adverse effects of motion interference regions on the final health indicators.

[0053] The AI ​​health indicator generation module will integrate physiological feature vectors into a pre-trained deep learning model and output multi-dimensional health indicators. The deep learning model includes a shared feature extraction layer and a task branching layer. The shared feature extraction layer performs high-dimensional feature abstraction on the integrated physiological feature vectors and extracts deep physiological representations shared across tasks.

[0054] The task branching layer includes four independent fully connected branches:

[0055] The brain function assessment branch outputs cognitive load index and fatigue index, which are used to assess the user's current cognitive state and level of mental fatigue.

[0056] The emotional stress assessment branch outputs an anxiety index and a depression tendency index, which jointly assess the user's emotional stress status based on facial micro-expression features and heart rate variability features derived from rPPG.

[0057] The TCM constitution identification branch outputs the probability distribution of nine constitution types, and identifies TCM constitutions based on facial color, texture dynamic changes and physiological rhythm characteristics.

[0058] The meridian and organ analysis branch outputs energy status scores corresponding to the twelve meridians, and assesses energy status based on the mapping relationship between various facial areas and meridians and organs.

[0059] In one embodiment:

[0060] The AI ​​health indicator generation module also includes an anomaly marking and graded response mechanism, which presets the normal threshold range of each health indicator and compares the multi-dimensional health indicators output by the model with the corresponding thresholds in real time.

[0061] When only a single indicator deviates slightly from the normal range, a Level 1 response is triggered, and the system pushes targeted lifestyle adjustment suggestions to the user, such as optimizing work and rest, dietary adjustments, and mood regulation plans.

[0062] When multiple indicators deviate moderately from the normal range, or a single indicator deviates significantly, a level two response is triggered. The system generates an early warning notification and pushes it to the user's terminal, and suggests that the user conduct a retest and verification through wearable devices to achieve cross-confirmation through dual-mode linkage.

[0063] When all multidimensional health indicators deviate significantly from the normal range, or when there is a life-threatening abnormal combination, a level three response is triggered. The system will push emergency medical advice and simultaneously notify the preset emergency contact person to ensure timely handling of high-risk situations.

[0064] In a specific implementation, the present invention includes continuously capturing video streams of the user's face and neck area at a preset frame rate of 30 frames per second using a mobile terminal camera;

[0065] The video stream is analyzed frame by frame to track the micro-motion changes of 44 facial muscle regions and 38 neck muscle regions. The temporal micro-motion signal sequence of each muscle region is extracted. Based on the color channel changes of the facial skin region in the video stream, the remote photoplethysmography (rPPG) signal is extracted.

[0066] The temporal micro-motion signal sequence is spatiotemporally aligned and cross-validated with the rPPG signal to generate a fused physiological feature vector. The fused physiological feature vector is input into a pre-trained deep learning model to output multidimensional health indicators, including brain function indicators, emotional stress indicators, traditional Chinese medicine constitution indicators, and meridian and organ indicators.

[0067] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

[0068] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0069] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A micro-motion feature health detection system based on mobile phone camera video, characterized in that: It includes a video acquisition module, a muscle micro-motion tracking module, an rPPG signal extraction module, a multimodal signal fusion module, and an AI health indicator generation module; The video acquisition module continuously acquires video streams of the user's face and neck area through the mobile terminal camera at a preset frame rate. The muscle micro-motion tracking module analyzes the acquired video streams of the face and neck area frame by frame and extracts the temporal micro-motion signal sequence of each muscle area. The rPPG signal extraction module extracts remote photoplethysmography (rPPG) signals based on the color channel changes of the facial skin region in the video stream of the face and neck region. The multimodal signal fusion module performs spatiotemporal alignment and cross-validation between the temporal micro-motion signal sequence and the rPPG signal to generate a fused physiological feature vector. The AI ​​health indicator generation module will integrate physiological feature vectors into a pre-trained deep learning model to output multi-dimensional health indicators.

2. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The multidimensional health indicators include brain function indicators, emotional stress indicators, traditional Chinese medicine constitution indicators, and meridian and organ indicators.

3. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The muscle micro-motion tracking module tracks the micro-motion changes of a preset number of muscle areas on the face and a preset number of muscle areas on the neck. in: The facial preset contains 44 muscle areas. The preset number of muscle areas in the neck is 38.

4. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The preset frame rate of the video capture module is 30 frames per second.

5. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The rPPG signal extraction module includes a region of interest (ROI) localization module, a color channel separation module, and a motion artifact suppression module. The Region of Interest (ROI) localization module locates the facial skin area based on a facial key point detection algorithm, excluding areas interfered by hair, glasses, and shadows. The color channel separation module separates the RGB color channels of the facial skin area and extracts the green channel signal and / or red channel signal as the original rPPG signal source; The motion artifact suppression module constructs a motion interference reference signal based on a time-domain micro-motion signal sequence, performs adaptive filtering or independent component analysis on the original rPPG signal source, and obtains a pure rPPG signal after motion artifact suppression.

6. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The spatiotemporal alignment and cross-validation of the multimodal signal fusion module includes temporal alignment, spatial mapping, and cross-validation. The time-domain alignment is based on the acquisition timestamp of the video stream, and the time-domain micro-motion signal sequence is synchronized with the rPPG signal in terms of timestamp. The spatial mapping establishes a spatial mapping relationship between facial muscle regions and facial skin ROIs, and identifies the correlation between micro-motion signals and blood flow signals at the same anatomical location; The cross-validation includes: When the time-domain micromotion signal sequence indicates significant movement in a specific muscle region, the weight of the rPPG signal at the corresponding spatial location is reduced, and vice versa, the weight is increased to generate a confidence-weighted fused physiological feature vector.

7. A micro-motion feature health detection system based on mobile phone camera video according to claim 2, characterized in that: The deep learning model includes a shared feature extraction layer and a task branching layer; The task branch layer includes a brain function assessment branch, an emotional stress assessment branch, a traditional Chinese medicine constitution identification branch, and a meridian and organ analysis branch. Each branch outputs health indicators corresponding to its dimension.

8. The micro-motion feature health detection system based on mobile phone camera video according to claim 1, characterized in that: The AI ​​health indicator generation module also includes an anomaly marker. When any multidimensional health indicator exceeds the preset normal threshold range, an anomaly marker is generated and a graded response is performed. The graded response includes a first-level response, a second-level response, and a third-level response; The first-level response is triggered when a single health indicator deviates from the normal range among the multidimensional health indicators, and lifestyle adjustment suggestions are pushed out. The secondary response is triggered when more than one of the multidimensional health indicators deviates from the normal range, thus triggering an early warning notification. The Level 3 response is triggered when all multidimensional health indicators deviate from the normal range, prompting an emergency medical consultation suggestion.