Multi-modal health risk early warning system for home-based care for the elderly
Patent Information
- Application Number
- CN202611015415.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
AI Technical Summary
[0002]随着人口老龄化加剧,居家环境中老年人跌倒事件频发,传统监测手段受限于隐私顾虑、环境干扰及单一模态感知的局限性,难以实现高可靠、连续的姿态行为监测与跌倒风险判别
[0014]本发明提供的一种居家养老多模态健康风险预警方法,包括以下步骤:获取模块,用于获取目标区域中目标对象的连续视频帧数据以及对应的雷达回波数据;重建模块,用于对所述视频帧数据进行图像预处理和边缘检测,得到视频边缘特征图,并对所述雷达回波数据进行杂波抑制和点云重建,得到三维空间点云图;融合模块,用于基于预设的相机标定参数与雷达标定参数,将所述三维空间点云图与所述视频边缘特征图映射至同一空间坐标系,并进行数据关联与融合,提取所述目标对象的三维人体骨架序列;追踪模块,用于对所述三维人体骨架序列进行时序帧间关联与运动追踪,生成所述目标对象的姿态运动轨迹;检测模块,用于根据预设的跌倒判定规则对所述姿态运动轨迹进行跌倒事件检测,判断是否发生跌倒行为;输出模块,用于当判定发生跌倒行为时,对所述跌倒行为进行严重程度分级,并输出对应的风险预警信号,解决了如何利用多模态感知融合,实现对目标对象姿态行为的可靠连续监测与跌倒风险判别的问题,实现了在复杂家居光照变化、部分遮挡及低照度环境下仍能高精度、低误报地完成跌倒事件实时识别与分级预警的技术效果。
Smart Images

Figure CN122805248A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health early warning technology, and in particular to a multimodal health risk early warning system for home-based elderly care. Background Technology
[0002] With the increasing aging of the population, falls among the elderly are frequent in the home environment. Traditional monitoring methods are limited by privacy concerns, environmental interference, and the limitations of single-modal perception, making it difficult to achieve highly reliable and continuous postural behavior monitoring and fall risk assessment. Summary of the Invention
[0003] The purpose of this invention is to at least partially solve one of the technical problems existing in the prior art.
[0004] To achieve the above objectives, the present invention provides a home-based elderly care multimodal health risk early warning system, comprising: The acquisition module is used to acquire continuous video frame data of the target object in the target area and the corresponding radar echo data; The reconstruction module is used to perform image preprocessing and edge detection on the video frame data to obtain a video edge feature map, and to perform clutter suppression and point cloud reconstruction on the radar echo data to obtain a three-dimensional spatial point cloud map. The fusion module is used to map the three-dimensional spatial point cloud map and the video edge feature map to the same spatial coordinate system based on preset camera calibration parameters and radar calibration parameters, and to perform data association and fusion to extract the three-dimensional human skeleton sequence of the target object. The tracking module is used to perform temporal inter-frame correlation and motion tracking on the three-dimensional human skeleton sequence to generate the posture motion trajectory of the target object; The detection module is used to detect fall events on the posture movement trajectory according to the preset fall determination rules, and to determine whether a fall has occurred. The output module is used to classify the severity of a fall when it is determined that a fall has occurred, and to output a corresponding risk warning signal.
[0005] Furthermore, image preprocessing and edge detection are performed on the video frame data to obtain a video edge feature map, and clutter suppression and point cloud reconstruction are performed on the radar echo data to obtain a three-dimensional spatial point cloud map, including: The filtering module is used to perform grayscale conversion and pixel filtering on the video frame data to obtain a grayscale smoothed image, and to extract the contour of the grayscale smoothed image based on pixel gradient calculation to obtain a video edge feature map. The analysis module is used to perform static background cancellation and range frequency domain analysis on the radar echo data to obtain the dynamic target spectrum; The mapping module is used to perform three-dimensional coordinate mapping on the spectrum of the dynamic target based on spatial angle calculation to obtain a three-dimensional spatial point cloud map.
[0006] Furthermore, based on preset camera calibration parameters and radar calibration parameters, the 3D spatial point cloud map and the video edge feature map are mapped to the same spatial coordinate system, and data association and fusion are performed to extract the 3D human skeleton sequence of the target object, including: A spatial transformation matrix is constructed based on preset camera calibration parameters and radar calibration parameters. The coordinate rotation and translation calculations are performed on the three-dimensional spatial point cloud map and the video edge feature map through the spatial transformation matrix to obtain a set of spatial registration points in the same three-dimensional coordinate system. Based on spatial Euclidean distance, the radar points and visual edge points in the spatial registration point cloud are searched and matched in a neighborhood. The depth values of the successfully matched radar points are assigned to the corresponding visual edge points, and the depth edge point cloud containing three-dimensional coordinate information is generated. Density clustering of the depth edge point cloud is performed based on the topology of the human skeleton, and the spatial centroid of each cluster region is calculated as the human joint point. Based on the topological connection of the human skeleton, the various human joints are connected and arranged according to the time dimension to generate a three-dimensional human skeleton sequence of the target object.
[0007] Furthermore, by performing coordinate rotation and translation calculations on the three-dimensional spatial point cloud map and the video edge feature map using the spatial transformation matrix, a spatial registration point set in the same three-dimensional coordinate system is obtained, including: Based on the spatial transformation matrix, a rotation and translation vector is extracted, and the three-dimensional spatial point cloud map is transformed into a three-dimensional coordinate system using the rotation and translation vector to obtain a visual point cloud. The depth gradient of the visual point cloud and the pixel gradient of the video edge feature map are extracted to perform cross-modal disparity analysis, and the disparity compensation matrix is obtained. Based on the disparity compensation matrix, the depth coordinates of the visual system point cloud are corrected to obtain a corrected 3D point cloud. Based on preset camera calibration parameters, the corrected 3D point cloud and the video edge feature map are physically scaled to obtain a spatial registration point cloud set in the same 3D coordinate system.
[0008] Furthermore, temporal inter-frame correlation and motion tracking are performed on the three-dimensional human skeleton sequence to generate the pose motion trajectory of the target object, including: Spatial coordinate difference calculation is performed on the three-dimensional human skeleton sequence based on the time difference between adjacent frames to obtain joint motion velocity values; Based on the joint motion velocity values, the state prediction and measurement update of the three-dimensional human skeleton sequence are performed to obtain a smooth skeleton sequence; Based on the spatial topology, the Euclidean distance between adjacent frame joints of the smooth skeleton sequence is calculated to obtain a node matching cost matrix. Then, based on the node matching cost matrix, the smooth skeleton sequence is assigned a minimum cost association to obtain a homologous joint sequence. The homologous joint sequences are arranged in timestamp order and spatial coordinate interpolation and curve fitting are performed to obtain a continuous set of spatial coordinates. Based on the spatial topology, the continuous spatial coordinate set is connected by multiple joints to generate the attitude motion trajectory of the target object.
[0009] Furthermore, based on the node matching cost matrix, the smooth skeleton sequence is subjected to minimum cost association allocation to obtain a sequence of homologous joints, including: The node matching cost matrix is subjected to extreme value threshold filtering and row and column minimum value deduction to obtain a local optimization matrix. The local optimization matrix is then subjected to independent zero element traversal and conflict node allocation to obtain the optimal allocation matrix. Based on the optimal allocation matrix, cross-frame index mapping is performed on the smooth skeleton sequence to obtain rearranged skeleton nodes; Based on the spatial topology, the identity identifiers of the rearranged skeleton nodes are inherited to obtain a sequence of homologous joints.
[0010] Furthermore, based on the spatial topology, the rearranged skeleton nodes are subjected to identity inheritance to obtain a sequence of homologous joints, including: Based on the spatial topology, the spatial vector difference calculation of adjacent nodes is performed on the rearranged skeleton nodes to obtain the skeleton connectivity vector; The rigid body length constraint and joint rotation angle are analyzed on the skeleton connectivity vector to obtain the topological weight matrix; Based on the topological weight matrix, the rearranged skeleton nodes are subjected to probability weighting of the preceding frame identifier and cross-frame transmission. Nodes with a transmission probability lower than a preset threshold are subjected to topological connectivity inference and identifier assignment to obtain a sequence of homologous joints.
[0011] Furthermore, based on preset fall determination rules, fall event detection is performed on the posture movement trajectory to determine whether a fall has occurred, including: Based on the preset gravity vector, the attitude motion trajectory is orthogonally projected and decoupled from the coordinates to obtain the vertical displacement sequence. Then, the vertical displacement sequence is subjected to time-domain first-order difference and squaring operations to obtain the falling impact sequence. Based on a preset impact threshold, extreme value search and interval truncation are performed on the falling impact sequence to obtain suspected fall segments; Calculate the variance of the trajectory coordinates at the end of the suspected fall segment. When the peak value of the fall impact sequence is greater than the preset impact threshold and the variance of the trajectory coordinates is less than the preset variance threshold, it is determined that a fall has occurred.
[0012] Furthermore, when a fall is determined to have occurred, the severity of the fall is classified, and a corresponding risk warning signal is output, including: When a fall is determined to have occurred, the fall impact sequence is used to perform time-domain integration within the time interval corresponding to the suspected fall segment to obtain the fall kinetic energy scalar. Based on the timestamp corresponding to the end of the suspected fall segment, the post-fall activity rate is obtained by time-domain truncation and displacement differentiation of the posture motion trajectory. The fall severity index is obtained by normalizing the scalar of fall kinetic energy and the post-fall activity rate based on preset weighting coefficients and then summing them by weight. The fall severity index is numerically compared and mapped to a range based on a preset grading threshold range to obtain a risk level identifier. Based on the risk level identifier, address matching is performed in a preset communication command set to extract the corresponding control message and generate the risk warning signal.
[0013] Furthermore, based on the timestamp corresponding to the end of the suspected fall segment, the post-fall activity rate is obtained by time-domain truncation and displacement differentiation of the posture motion trajectory, including: Based on the timestamp corresponding to the end of the suspected fall segment, the posture motion trajectory is truncated by forward sliding along the time axis to obtain the post-fall trajectory sequence. The spatial coordinate difference between adjacent frames is then performed on the post-fall trajectory sequence to obtain the relative displacement vector. The relative displacement vector is differentiated over time based on the time difference between adjacent frames to obtain the instantaneous velocity vector. The norm modulus and root mean square time of the instantaneous velocity vector are then calculated to obtain the post-fall activity rate.
[0014] This invention provides a multimodal health risk early warning method for home-based elderly care, comprising the following steps: an acquisition module, used to acquire continuous video frame data of a target object in a target area and corresponding radar echo data; a reconstruction module, used to perform image preprocessing and edge detection on the video frame data to obtain a video edge feature map, and to perform clutter suppression and point cloud reconstruction on the radar echo data to obtain a three-dimensional spatial point cloud map; a fusion module, used to map the three-dimensional spatial point cloud map and the video edge feature map to the same spatial coordinate system based on preset camera calibration parameters and radar calibration parameters, and to perform data association and fusion to extract the three-dimensional human skeleton sequence of the target object; and a tracking module, used to track... The three-dimensional human skeleton sequence is correlated with time-series frames and motion tracked to generate the posture and motion trajectory of the target object. The detection module is used to detect fall events on the posture and motion trajectory according to preset fall judgment rules to determine whether a fall has occurred. The output module is used to classify the severity of the fall behavior when a fall is determined to have occurred and output the corresponding risk warning signal. This solves the problem of how to use multimodal perception fusion to achieve reliable and continuous monitoring of the posture and behavior of the target object and fall risk judgment. It achieves the technical effect of high accuracy and low false alarm in real-time identification and graded warning of fall events even in complex home lighting changes, partial occlusion and low light environment. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of a home-based elderly care multimodal health risk early warning system according to an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0018] Please see Figure 1 An embodiment of the present invention provides a multimodal health risk early warning system for home-based elderly care, comprising: Acquisition module 1 is used to acquire continuous video frame data of the target object in the target area and the corresponding radar echo data.
[0019] Specifically, in this system, the acquisition module 1 is deployed at a fixed location in the target area, simultaneously activating a visible light camera and a millimeter-wave radar sensor to achieve continuous observation of the target object within the same spatiotemporal range. The camera acquires an RGB video stream at a rate of no less than 15 frames per second, while the millimeter-wave radar transmits a frequency-modulated continuous wave at the same timestamp and receives the corresponding echo signal, ensuring strict temporal alignment between the video frames and the radar sampling data. To ensure spatial consistency of multi-source data, both devices must undergo extrinsic parameter calibration before installation, and synchronization deviations between heterogeneous sensors are eliminated through hardware triggering or software timestamp matching mechanisms. The acquired video frame data contains the appearance and motion information of the target object under natural lighting, while the radar echo data records its micro-Doppler characteristics and distance-angle distribution in three-dimensional space. In this embodiment, the above scheme provides a reliable data foundation for subsequent multimodal fusion through a high-precision spatiotemporal synchronization mechanism, effectively avoiding attitude misjudgment caused by asynchronous sampling.
[0020] Reconstruction module 2 is used to perform image preprocessing and edge detection on the video frame data to obtain a video edge feature map, and to perform clutter suppression and point cloud reconstruction on the radar echo data to obtain a three-dimensional spatial point cloud map.
[0021] Specifically, reconstruction module 2 first performs denoising, grayscale conversion, and histogram equalization on the acquired continuous video frame data to improve image contrast and suppress illumination fluctuation interference. Then, it uses the Canny operator for edge detection to generate a video edge feature map highlighting the human body contour and limb structure. Simultaneously, for the synchronously acquired radar echo data, a static clutter cancellation algorithm is used to filter out reflections from walls, furniture, and other static environments. Effective moving target points are extracted based on range-Doppler joint processing, and spatial mapping is completed using angle-of-arrival estimation to reconstruct a three-dimensional spatial point cloud map representing the distribution of the target object's torso and limbs. In this embodiment, the above scheme optimizes visual edge information and radar point cloud quality respectively, providing a high signal-to-noise ratio and structurally clear intermediate feature representation for subsequent cross-modal registration and skeleton extraction.
[0022] The fusion module 3 is used to map the three-dimensional spatial point cloud map and the video edge feature map to the same spatial coordinate system based on preset camera calibration parameters and radar calibration parameters, and to perform data association and fusion to extract the three-dimensional human skeleton sequence of the target object.
[0023] Specifically, fusion module 3 first constructs a unified mapping relationship from the image pixel coordinate system to the radar's three-dimensional spatial coordinate system based on pre-completed camera intrinsic and extrinsic parameters and the radar's installation pose parameters relative to the world coordinate system. Then, salient contour points in the video edge feature map are transformed to this common coordinate system through projection transformation, and nearest neighbor matching and confidence-weighted association are performed with the corresponding human reflection points in the three-dimensional spatial point cloud map to eliminate outlier pairings caused by occlusion or noise. Based on this, multi-frame temporal consistency constraints and a priori human motion model are used to jointly optimize the associated cross-modal features, ultimately generating a continuous three-dimensional human skeleton sequence containing coordinates of key nodes such as the head, shoulders, elbows, hips, and knees. In this embodiment, the above scheme significantly improves the integrity and robustness of human posture reconstruction in complex home environments through precise cross-modal spatial alignment and structural complementary fusion.
[0024] Tracking module 4 is used to perform temporal inter-frame correlation and motion tracking on the three-dimensional human skeleton sequence to generate the posture motion trajectory of the target object.
[0025] Specifically, tracking module 4 receives the 3D human skeleton sequence output by the fusion module. First, based on the Euclidean distance and motion velocity constraints of joint coordinates between adjacent frames, it constructs an inter-frame joint matching map and uses the Hungarian algorithm to solve for the optimal associated path to eliminate identity jumps caused by brief occlusions or missing point clouds. Then, a Kalman filter is introduced to predict the state and update the observation of the trajectory of each joint, effectively suppressing jitter caused by sensor noise. On this basis, through posture continuity verification and biomechanical rationality judgment within a sliding time window, abnormal jumps or non-humanoid motion segments are eliminated, ultimately forming a smooth, coherent, and consistent target object posture motion trajectory. In this embodiment, the above scheme achieves high-precision, long-term stable tracking of human motion trajectories in complex home scenarios by fusing geometric constraints and dynamic models.
[0026] The detection module 5 is used to detect fall events on the posture movement trajectory according to the preset fall determination rules, and to determine whether a fall has occurred.
[0027] Specifically, the detection module 5 extracts key criteria such as the trunk axis tilt angle, vertical velocity of the center of mass, and hip joint height in real time based on the posture trajectory output by the tracking module 4. When the angle between the trunk axis and the vertical direction suddenly increases from less than 30 degrees to more than 60 degrees within 0.3 seconds, and the descent rate of the center of mass exceeds 1.2 m / s while the hip joint height is below a preset threshold (e.g., 0.4 meters from the ground), a fall detection rule is triggered. To avoid false alarms, the system further verifies whether the target maintains a low posture and does not recover significant displacement within the following 0.5 seconds. In this embodiment, the above scheme, by integrating multi-dimensional kinematic features and temporal dynamic constraints, effectively suppresses false alarms caused by daily movements such as bending over and sitting / lying down while ensuring high sensitivity, significantly improving the accuracy and reliability of fall detection for elderly people living at home.
[0028] Output module 6 is used to classify the severity of a fall when a fall is determined to have occurred, and to output a corresponding risk warning signal.
[0029] Specifically, after receiving the fall behavior assessment result, output module 6 classifies the severity based on three dimensions: the rate of change of trunk tilt angle in the posture trajectory, the peak value of the vertical acceleration of the center of mass, and the duration of stillness after the fall. If the vertical acceleration of the center of mass exceeds 2.5g and the stillness time after the fall exceeds 30 seconds, it is judged as high risk; if the acceleration is between 1.5g and 2.5g and the stillness time is between 10 and 30 seconds, it is judged as medium risk; other situations that meet the fall criteria but do not reach the above thresholds are classified as low risk. Subsequently, different levels of risk warning signals are triggered according to the classification results, including pushing high-priority alarms to the monitoring terminal, initiating voice inquiries, or simply recording event logs. In this embodiment, the above scheme achieves a refined distinction of risk levels by quantifying the dynamic characteristics of falls and subsequent behavioral responses, providing a reliable basis for timely intervention and medical resource allocation.
[0030] In a specific embodiment, image preprocessing and edge detection are performed on the video frame data to obtain a video edge feature map, and clutter suppression and point cloud reconstruction are performed on the radar echo data to obtain a three-dimensional spatial point cloud map, including: The video frame data is subjected to grayscale conversion and pixel filtering to obtain a grayscale smoothed image, and the contour is extracted from the grayscale smoothed image based on pixel gradient calculation to obtain a video edge feature map. Static background cancellation and range-frequency domain analysis are performed on the radar echo data to obtain the dynamic target spectrum; Based on spatial angle calculation, the spectrum of the dynamic target is mapped to three-dimensional coordinates to obtain a three-dimensional spatial point cloud map.
[0031] Specifically, the input video frame data is first converted to grayscale, mapping the original RGB three-channel image to a single-channel grayscale image using a weighted average method. The weights for the red, green, and blue components are set to 0.299, 0.587, and 0.114, respectively, to conform to the characteristics of human visual sensitivity. Subsequently, a 3×3 Gaussian kernel is used for spatial domain filtering, with a standard deviation set to 1.0 pixel. This effectively suppresses image sensor noise while preserving the integrity of edge structures, thereby obtaining a smooth grayscale image.
[0032] Based on a smoothed grayscale image, the Sobel operator is used to calculate the pixel gradient magnitudes in the horizontal and vertical directions, respectively. Continuous contour lines are then extracted using non-maximum suppression and double-threshold hysteresis connection strategies, ultimately generating a video edge feature map representing the boundaries of the human body and the transition features of limb joints. This edge feature map highlights the limb contours and key pose inflection points of the target object, providing a clear geometric prior for subsequent skeleton extraction.
[0033] Meanwhile, for radar echo data, static background cancellation processing is first implemented. This involves constructing a background template using the median echo amplitude of each range-angle unit within the most recent 64 frames, and then subtracting this template from the current frame to filter out fixed clutter such as walls and furniture. This operation significantly reduces strong reflection interference caused by static objects in the indoor environment, making dynamic human targets stand out.
[0034] Subsequently, a distance-dimensional Fast Fourier Transform was performed on the residual dynamic signal to identify the Doppler frequency shift components corresponding to subtle human movements in the frequency domain, forming the dynamic target spectrum. This spectrum comprehensively reflects the target's motion state in the radial direction, effectively separating weak dynamic features such as breathing and limb swaying from environmental noise.
[0035] Then, combining the phase difference information of the radar antenna array, the spatial pointing of the target in azimuth and elevation angles is calculated using MUSIC or FFT direction-of-arrival estimation algorithms, and three-dimensional coordinate mapping is completed by combining range gate information. The range resolution is determined by the transmitted signal bandwidth (for example, 4 GHz bandwidth corresponds to a theoretical resolution of about 3.75 cm), and the angular resolution is limited by the antenna aperture length. Finally, a three-dimensional spatial point cloud map containing spatial position and reflection intensity attributes is reconstructed.
[0036] In this embodiment, the above scheme effectively suppresses environmental interference while preserving key human structural information by coordinating and optimizing the visual edge representation and millimeter-wave radar point cloud reconstruction process, providing a high signal-to-noise ratio and geometrically consistent input basis for subsequent cross-modal fusion.
[0037] In a specific embodiment, based on preset camera calibration parameters and radar calibration parameters, the three-dimensional spatial point cloud map and the video edge feature map are mapped to the same spatial coordinate system, and data association and fusion are performed to extract the three-dimensional human skeleton sequence of the target object, including: A spatial transformation matrix is constructed based on preset camera calibration parameters and radar calibration parameters. The coordinate rotation and translation calculations are performed on the three-dimensional spatial point cloud map and the video edge feature map through the spatial transformation matrix to obtain a set of spatial registration points in the same three-dimensional coordinate system. Based on spatial Euclidean distance, the radar points and visual edge points in the spatial registration point cloud are searched and matched in a neighborhood. The depth values of the successfully matched radar points are assigned to the corresponding visual edge points, and the depth edge point cloud containing three-dimensional coordinate information is generated. Density clustering of the depth edge point cloud is performed based on the topology of the human skeleton, and the spatial centroid of each cluster region is calculated as the human joint point. Based on the topological connection of the human skeleton, the various human joints are connected and arranged according to the time dimension to generate a three-dimensional human skeleton sequence of the target object.
[0038] Specifically, after independently generating the video edge feature map and the 3D point cloud map, the system first constructs a rigid body space transformation matrix from the radar coordinate system to the camera coordinate system based on the pre-calibrated camera intrinsic parameter matrix, lens distortion coefficients, and radar extrinsic pose relative to the camera (including a 3×3 rotation matrix and a 3D translation vector). This transformation matrix is used to perform coordinate rotation and translation operations on the 3D point cloud obtained from radar echo reconstruction. At the same time, the 2D pixels in the video edge feature map, combined with their assigned depth information, are back-projected into the 3D space, thereby achieving spatial alignment of the two types of heterogeneous data in a unified world coordinate system, forming a spatial registration point set.
[0039] To achieve effective fusion of cross-modal data, the system uses each visual edge point as a reference and sets a spherical neighborhood search area around it with a radius not exceeding 15 centimeters—this value references the size of a local human limb while taking into account the typical ranging error of millimeter-wave radar (usually within 5 centimeters). Within this area, the Euclidean distance between all radar points and the visual point is calculated; if the minimum distance is less than a preset matching threshold (e.g., 8 centimeters), it is determined that the two originate from the same physical location, and the high-precision depth value carried by the radar point is assigned to the corresponding visual edge point, thereby generating a depth edge point cloud that integrates image contour semantics and radar ranging accuracy.
[0040] The system introduces a standard human skeletal topology model, which defines 17 key anatomical joints (such as the top of the head, both shoulders, both elbows, both hands, the center of the hip, both knees, and both feet) and their connections. Based on this prior structure, a density clustering algorithm (such as DBSCAN) is performed on the depth edge point clusters, with a neighborhood radius of 10 cm and a minimum number of cluster points of 5, to effectively separate the local point clusters corresponding to each limb part. For each successfully clustered point cluster, the arithmetic mean of the spatial coordinates of all its points is calculated as the 3D position estimate of the joint point of that anatomical part.
[0041] Based on the topological rules of the human skeleton, the joints are connected in anatomical order, and the skeleton structures generated from consecutive video frames are sequentially arranged along the time dimension to form a complete and temporally continuous three-dimensional human skeleton sequence. In this embodiment, the above scheme effectively overcomes the limitations of a single sensor in occlusion, illumination changes, or low reflectivity scenarios through high-precision spatial registration, cross-modal point matching, and anatomical prior-based clustering strategies, significantly improving the robustness and spatiotemporal continuity of three-dimensional skeleton reconstruction.
[0042] In a specific embodiment, the spatial transformation matrix is used to perform coordinate rotation and translation calculations on the three-dimensional spatial point cloud map and the video edge feature map to obtain a spatial registration point set in the same three-dimensional coordinate system, including: Based on the spatial transformation matrix, a rotation and translation vector is extracted, and the three-dimensional spatial point cloud map is transformed into a three-dimensional coordinate system using the rotation and translation vector to obtain a visual point cloud. The depth gradient of the visual point cloud and the pixel gradient of the video edge feature map are extracted to perform cross-modal disparity analysis, and the disparity compensation matrix is obtained. Based on the disparity compensation matrix, the depth coordinates of the visual system point cloud are corrected to obtain a corrected 3D point cloud. Based on preset camera calibration parameters, the corrected 3D point cloud and the video edge feature map are physically scaled to obtain a spatial registration point cloud set in the same 3D coordinate system.
[0043] Specifically, the system first separates a 3×3 rotation submatrix and a 3D translation vector from the constructed spatial transformation matrix. The rotation submatrix describes the attitude difference between the radar coordinate system and the camera coordinate system, while the translation vector represents the spatial offset between their origins. Using these two parameters, the system performs a rigid body transformation on each 3D point in the original 3D spatial point cloud: first, the point coordinates are left-multiplied by the rotation submatrix to achieve orientation alignment, and then the translation vector is superimposed to achieve position correction, thereby generating a preliminary visual point cloud mapped to the camera coordinate system.
[0044] To eliminate cross-modal geometric deviations caused by sensor installation errors, calibration residuals, or media propagation effects, the system introduces a joint analysis mechanism of depth gradient and pixel gradient. Specifically, it calculates the depth gradient of the visual point cloud in its local neighborhood (i.e., the rate of change of the Z-coordinate between adjacent points) and simultaneously extracts the pixel gradient magnitude of the video edge feature map in the corresponding image region. By spatially aligning and statistically analyzing the deviations of the two types of gradient responses within a common visible region, the system identifies the systematic disparity distribution and constructs a disparity compensation matrix indexed by image coordinates, where each element records the depth correction amount required for that image location.
[0045] The system performs depth coordinate fine-tuning on the visual point cloud based on the aforementioned disparity compensation matrix. For example, if the radar point cloud is detected to be 2.1 cm ahead of the visual edge in the shoulder edge region, the Z coordinate of all points in that region is uniformly subtracted from the deviation value. This type of correction is not a global scaling, but a dynamic adjustment based on local edge consistency, ensuring that key structures such as limb contours maintain geometric coherence after fusion, and ultimately outputting a corrected 3D point cloud.
[0046] Based on this, the system calls the preset camera intrinsic parameter matrix (including horizontal and vertical focal lengths, principal point coordinates, and possible distortion coefficients), backprojects the corrected 3D point cloud onto the image plane, and compares its physical scale with the original video edge feature map. If an overall scale inconsistency is found (e.g., due to radar distance units not being fully normalized to the metric system or image pixels not being associated with actual physical dimensions), the system uses the typical length of a human limb (e.g., shoulder width of approximately 40–50 cm) as a reference benchmark, calculates the scaling factor, and uniformly scales the point cloud coordinates to achieve millimeter-level physical scale alignment, ultimately obtaining a spatial registration point cloud set in the same 3D coordinate system.
[0047] In this embodiment, the above scheme introduces a gradient response-based cross-modal parallax resolution and depth correction mechanism, which achieves sub-pixel-level geometric consistency optimization on the basis of traditional rigid body registration, significantly improving the fusion accuracy and skeleton reconstruction reliability of multi-source sensing data in complex indoor scenes.
[0048] In a specific embodiment, temporal inter-frame correlation and motion tracking are performed on the three-dimensional human skeleton sequence to generate the pose motion trajectory of the target object, including: Spatial coordinate difference calculation is performed on the three-dimensional human skeleton sequence based on the time difference between adjacent frames to obtain joint motion velocity values; Based on the joint motion velocity values, the state prediction and measurement update of the three-dimensional human skeleton sequence are performed to obtain a smooth skeleton sequence; Based on the spatial topology, the Euclidean distance between adjacent frame joints of the smooth skeleton sequence is calculated to obtain a node matching cost matrix. Then, based on the node matching cost matrix, the smooth skeleton sequence is assigned a minimum cost association to obtain a homologous joint sequence. The homologous joint sequences are arranged in timestamp order and spatial coordinate interpolation and curve fitting are performed to obtain a continuous set of spatial coordinates. Based on the spatial topology, the continuous spatial coordinate set is connected by multiple joints to generate the attitude motion trajectory of the target object.
[0049] Specifically, the system first performs a difference operation on the three-dimensional coordinates of each joint point in two consecutive frames based on the time interval between adjacent frames in the three-dimensional human skeleton sequence (usually the reciprocal of the video frame rate, such as 30 fps corresponding to about 33 milliseconds), thereby calculating the displacement vector of the joint in a unit time, and further obtaining its motion velocity value; this velocity value is used to characterize the instantaneous dynamic characteristics of the joint and provide observation basis for subsequent state filtering.
[0050] The system employs a Kalman filter framework for state prediction and measurement updates of the skeleton sequence. In the prediction phase, the prior state of the current frame is deduced based on the joint positions and velocities of the previous moment, combined with a uniform or uniformly accelerated motion model. In the update phase, the observed 3D joint coordinates are used as measurements, and the prior estimates and observation data are fused through covariance weighting to output a smooth skeleton sequence with noise suppression, effectively reducing high-frequency jitter caused by point cloud jitter or edge mismatches.
[0051] To further ensure cross-frame joint identity consistency, the system calculates the Euclidean distance between all pairs of identical joints in adjacent frames of the smooth skeleton sequence based on the inherent spatial topology of the human body (such as the chain connection relationship of shoulder-elbow-wrist), and constructs a node matching cost matrix with frame order as rows and joint index as columns. On this basis, the Hungarian algorithm is introduced to solve the cost matrix globally to achieve joint association allocation with minimum total matching cost, thereby generating independent and continuous homologous joint sequences for each joint and avoiding identity jumps caused by occlusion or sudden pose changes.
[0052] The system performs time alignment on each co-originating joint sequence according to timestamp order, and uses cubic spline interpolation to complete the spatial coordinates in missing frames or low confidence intervals. At the same time, it performs local polynomial curve fitting (such as fifth-order sliding window fitting) on the complete time series to eliminate residual non-physical jitter, and finally forms a continuous spatial coordinate set with high continuity and low noise.
[0053] Based on the preset human skeleton connection rules (such as the spine being connected by three key points: neck, chest, and waist), the multi-joint coordinates corresponding to the continuous spatial coordinate set are connected sequentially according to the anatomical structure to construct a three-dimensional skeleton diagram that evolves over time, i.e., the posture and motion trajectory of the target object.
[0054] In this embodiment, the above-mentioned scheme achieves highly robust and highly continuous three-dimensional attitude tracking in complex dynamic scenarios by integrating kinematic difference, state filtering, optimal allocation and curve fitting techniques, which significantly improves the spatiotemporal smoothness and biomechanical rationality of the motion trajectory.
[0055] In a specific embodiment, the smooth skeleton sequence is assigned a minimum-cost association based on the node matching cost matrix to obtain a sequence of homologous joints, including: The node matching cost matrix is subjected to extreme value threshold filtering and row and column minimum value deduction to obtain a local optimization matrix. The local optimization matrix is then subjected to independent zero element traversal and conflict node allocation to obtain the optimal allocation matrix. Based on the optimal allocation matrix, cross-frame index mapping is performed on the smooth skeleton sequence to obtain rearranged skeleton nodes; Based on the spatial topology, the identity identifiers of the rearranged skeleton nodes are inherited to obtain a sequence of homologous joints.
[0056] Specifically, the system first applies extreme value threshold filtering to the node matching cost matrix composed of the Euclidean distance between joints in adjacent frames, eliminating distance terms that exceed the reasonable range of human joint movement (for example, setting an upper limit of 30 centimeters, corresponding to the maximum limb displacement in a single frame during normal walking), in order to exclude abnormal matching candidates; then, it performs row and column minimum value subtraction operations on the cost matrix, that is, subtracting the minimum value of each row from each row and the minimum value of each column from each column, so that at least one zero element appears in the matrix, thereby transforming it into a local optimization matrix, laying the foundation for subsequent independent zero element selection.
[0057] In this local optimization matrix, the system uses a depth-first traversal strategy to scan row by row, searching for independent sets of zero elements that are not in the same row or column. When multiple zero elements have allocation conflicts (such as two zeros in the same column), they are prioritized according to their original cost, retaining the matching items with lower costs, and temporarily storing the conflicting nodes in the processing queue. For uncovered rows or columns, matrix coverage is completed by introducing virtual rows / columns and assigning high penalty costs, ultimately generating the optimal allocation matrix that satisfies the one-to-one constraint.
[0058] Based on this optimal allocation matrix, the system remaps the joint indexes of the current frame in the smooth skeleton sequence to ensure strict alignment with the joint identities of the previous frame, outputting rearranged skeleton nodes with consistent structures. Subsequently, combining the inherent spatial topology of the human body (such as the left shoulder connecting only to the left elbow and the spine being distributed in a linear chain), the identity identifiers of each joint in the previous frame (such as "right knee" and "left wrist") are inherited along the skeletal connection path to the corresponding rearranged nodes in the current frame. This ensures that even under brief occlusion or drastic changes in posture, the identities of each joint remain continuous and without cross-contamination, ultimately forming a homologous joint sequence with independent joint trajectories and stable identities.
[0059] In this embodiment, the above scheme effectively solves the cross-frame identity ambiguity problem of multi-target key points in complex dynamic interaction scenarios by integrating threshold constraints, matrix preprocessing variants of the Hungarian algorithm and topology-guided identity inheritance mechanism, significantly improving the long-term consistency and individual distinguishability of 3D pose tracking.
[0060] In a specific embodiment, identity inheritance is performed on the rearranged skeleton nodes based on the spatial topology to obtain a sequence of homologous joints, including: Based on the spatial topology, the spatial vector difference calculation of adjacent nodes is performed on the rearranged skeleton nodes to obtain the skeleton connectivity vector; The rigid body length constraint and joint rotation angle are analyzed on the skeleton connectivity vector to obtain the topological weight matrix; Based on the topological weight matrix, the rearranged skeleton nodes are subjected to probability weighting of the preceding frame identifier and cross-frame transmission. Nodes with a transmission probability lower than a preset threshold are subjected to topological connectivity inference and identifier assignment to obtain a sequence of homologous joints.
[0061] Specifically, the system first performs three-dimensional coordinate difference calculations on each pair of adjacent joints in the rearranged skeletal nodes based on the inherent spatial topology of the human body (such as the chain connection relationship between shoulder-elbow-wrist), generating a skeletal connectivity vector that represents the direction and length of the bones; this vector not only reflects the relative pose between joints, but also provides a geometric basis for subsequent biomechanical constraints.
[0062] The system introduces a rigid body length constraint mechanism, which compares the magnitude of each bone connection vector with a preset individualized bone length (e.g., the length of an adult forearm is usually between 23 and 27 cm). If the deviation exceeds 5%, it is judged as an abnormal match and a penalty is imposed. At the same time, the system calculates the joint rotation angle (e.g., the elbow flexion and extension angle) by combining the angle between adjacent bone vectors, and verifies the consistency of this angle with the allowable range of human anatomy (e.g., 0° to 150°). The system generates a confidence score for each connection relationship by combining the length deviation and the reasonableness of the angle, and then constructs a topological weight matrix covering the entire skeleton, where high weights correspond to highly reliable topological connections.
[0063] The system uses the confirmed identity identifiers of each joint in the previous frame (such as "left hip" and "right ankle") as prior information, and performs probability-weighted transmission of the rearranged nodes in the current frame according to the topological weight matrix. Specifically, the probability of a node inheriting a certain identity is proportional to its spatial proximity to the corresponding node in the previous frame and the topological weight of its skeletal chain. When the transmission probability is lower than a preset threshold (such as 0.65, to deal with the observation loss caused by brief occlusion or violent movement), the system no longer relies on direct matching, but instead activates the topological connectivity inference mechanism. For example, if an unidentified node is connected to a node that has been identified as "left shoulder" through a high-weight edge, and its spatial position conforms to the clavicle extension direction, it is inferred to be "left elbow" based on human symmetry and chain structure, and is assigned the corresponding identity identifier, thereby completing the closed-loop assignment of the identity of the joints in the whole frame, and finally outputting a temporally continuous and identity-consistent homologous joint sequence.
[0064] In this embodiment, the above-mentioned scheme, by integrating rigid body constraints of skeletons, reasonable joint kinematics, and topology-guided identifier transfer mechanism, can still maintain high-precision identity continuity in scenarios with missing observations or ambiguous matching, significantly enhancing the robustness and long-term stability of 3D human posture tracking in complex interactive environments.
[0065] In a specific embodiment, fall event detection is performed on the posture motion trajectory according to a preset fall determination rule to determine whether a fall has occurred, including: Based on the preset gravity vector, the attitude motion trajectory is orthogonally projected and decoupled from the coordinates to obtain the vertical displacement sequence. Then, the vertical displacement sequence is subjected to time-domain first-order difference and squaring operations to obtain the falling impact sequence. Based on a preset impact threshold, extreme value search and interval truncation are performed on the falling impact sequence to obtain suspected fall segments; Calculate the variance of the trajectory coordinates at the end of the suspected fall segment. When the peak value of the fall impact sequence is greater than the preset impact threshold and the variance of the trajectory coordinates is less than the preset variance threshold, it is determined that a fall has occurred.
[0066] Specifically, the system first orthogonally projects the posture trajectory of the target object based on the preset gravity vector (usually defined as the unit vector pointing to the ground in the world coordinate system, such as [0, 0, -1]). It then decouples the key points (such as the center of the spine or the center of the pelvis) in the three-dimensional skeleton sequence along the gravity direction and extracts their coordinate components on the vertical axis to form a vertical displacement sequence that characterizes the change in human height. This sequence is indexed by timestamps and reflects the dynamic rise and fall of the overall center of gravity of the target object.
[0067] The vertical displacement sequence is subjected to a first-order time-domain difference operation to obtain approximate vertical velocities between adjacent frames. The difference results are squared one by one to eliminate directionality and highlight energy characteristics, thereby generating a fall impact sequence. The magnitude of this sequence is proportional to the square of the human body's fall acceleration, which can effectively amplify the signal response of sudden fall events while suppressing interference from non-falling actions such as slow sitting or lying down.
[0068] The system sets a preset impact threshold (e.g., the theoretical impact energy corresponding to a free fall of 0.5 meters, which is 0.8 m² / s² after actual measurement and calibration). It performs a sliding window extreme value search on the falling impact sequence to locate all time intervals where the local peak exceeds the threshold, and extracts several frames before and after (e.g., ±15 frames, corresponding to about 1 second) as suspected fall segments. Based on this, it calculates the spatial coordinate variance of the posture trajectory in the last few frames of the segment (e.g., the last 5 frames) to measure whether the human body has entered a static or low-activity state. If the variance is less than the preset variance threshold (e.g., 0.02 m², corresponding to the body being basically still on the ground), and the peak value of the falling impact sequence is indeed higher than the impact threshold, then it is jointly determined that a fall has occurred.
[0069] In this embodiment, the above-mentioned scheme integrates three criteria: vertical dynamic feature extraction, impact energy quantification, and static state verification. This avoids misjudging rapid squatting or bending over as a fall while ensuring high sensitivity in capturing real fall events, significantly improving the accuracy and reliability of fall detection in home monitoring and elderly care scenarios.
[0070] In a specific embodiment, when a fall is determined to have occurred, the severity of the fall is graded, and a corresponding risk warning signal is output, including: When a fall is determined to have occurred, the fall impact sequence is used to perform time-domain integration within the time interval corresponding to the suspected fall segment to obtain the fall kinetic energy scalar. Based on the timestamp corresponding to the end of the suspected fall segment, the post-fall activity rate is obtained by time-domain truncation and displacement differentiation of the posture motion trajectory. The fall severity index is obtained by normalizing the scalar of fall kinetic energy and the post-fall activity rate based on preset weighting coefficients and then summing them by weight. The fall severity index is numerically compared and mapped to a range based on a preset grading threshold range to obtain a risk level identifier. Based on the risk level identifier, address matching is performed in a preset communication command set to extract the corresponding control message and generate the risk warning signal.
[0071] Specifically, once the system determines that a fall has occurred, it immediately initiates a severity grading process. First, for the time interval corresponding to the identified suspected fall segment, a time-domain integral operation is performed on the fall impact sequence. This integral result, in a physical sense, approximately represents the kinetic energy accumulated by the body's center of gravity during the fall, and the output is a dimensionless but calibrated scalar of fall kinetic energy. The larger the value of this scalar, the more severe the impact process and the higher the potential risk of injury.
[0072] Using the timestamp corresponding to the end of the suspected fall segment as a benchmark, a fixed-duration (e.g., 3 seconds) postural movement trajectory is extracted, and the time derivative of the spatial coordinate sequence of key points (e.g., the pelvis or the center of the thoracic spine) is estimated. The modulus mean of the displacement change rate is calculated to form the post-fall activity rate. This rate reflects whether an individual has the ability to move independently after a fall—a value close to zero usually means that the individual is unable to get up or has lost consciousness, while a higher value may indicate rapid recovery after a minor fall.
[0073] The system introduces preset weighting coefficients (e.g., kinetic energy weight 0.7, activity rate weight 0.3, calibrated based on clinical statistical data). First, the scalar of fall kinetic energy and the activity rate after the fall are normalized to the [0,1] interval to eliminate the difference in dimensions. Then, the weights are linearly combined to generate a fall severity index between 0 and 1. This index integrates the impact intensity and subsequent behavioral response to avoid misjudgment by a single indicator.
[0074] Based on preset risk threshold ranges (e.g., 0.0–0.4 for low risk, 0.4–0.7 for medium risk, and 0.7–1.0 for high risk), the fall severity index is mapped to the corresponding risk level identifier. Finally, the identifier is used for address matching in a preset set of communication instructions. For example, high risk corresponds to the control message "SOS emergency call + family push + local sound and light alarm", medium risk triggers the "voice inquiry + delayed confirmation" instruction, and low risk only records the log, thereby generating a structured and executable risk warning signal.
[0075] In this embodiment, the above-mentioned scheme constructs a multi-level fall risk quantification model that conforms to the logic of medical observation by integrating dynamic energy assessment and behavioral response characteristics. While ensuring timely response, it significantly reduces the false alarm rate and false negative rate, and improves the clinical applicability and intervention accuracy of the intelligent monitoring system in elderly care scenarios.
[0076] In a specific embodiment, the post-fall activity rate is obtained by performing time-domain truncation and displacement differentiation on the posture motion trajectory based on the timestamp corresponding to the end of the suspected fall segment, including: Based on the timestamp corresponding to the end of the suspected fall segment, the posture motion trajectory is truncated by forward sliding along the time axis to obtain the post-fall trajectory sequence. The spatial coordinate difference between adjacent frames is then performed on the post-fall trajectory sequence to obtain the relative displacement vector. The relative displacement vector is differentiated over time based on the time difference between adjacent frames to obtain the instantaneous velocity vector. The norm modulus and root mean square time of the instantaneous velocity vector are then calculated to obtain the post-fall activity rate.
[0077] Specifically, after determining that a fall has occurred, the system uses the timestamp corresponding to the end of the suspected fall segment as the starting point and slides forward along the time axis to extract a pre-set duration (e.g., 3 seconds) of posture and motion trajectory, forming a post-fall trajectory sequence. This sequence contains the three-dimensional spatial coordinates of key human joints (such as the center of the pelvis or the midpoint of the spine) in consecutive frames, used to characterize the body's dynamic response after the fall. Subsequently, the system performs a spatial coordinate difference operation between adjacent frames on the post-fall trajectory sequence, that is, subtracting the coordinates of the previous frame from the coordinates of the current frame to generate a series of relative displacement vectors. Each vector accurately reflects the spatial movement direction and amplitude of the joints within a single frame interval.
[0078] The system acquires the actual time difference between adjacent frames (usually determined by the video frame rate, e.g., 30 frames / second corresponds to a time difference of approximately 33 milliseconds), and uses this time difference as the denominator to perform element-wise division on the aforementioned relative displacement vector, completing the time derivative process to obtain the instantaneous velocity vector corresponding to each frame. The three components of this vector represent the instantaneous motion velocity of the joint in the x, y, and z directions in the world coordinate system, respectively. To further quantify the overall activity level, the system calculates the Euclidean norm (i.e., modulus) of each instantaneous velocity vector to obtain a scalarized instantaneous velocity value. Then, it performs a root mean square (RMS) operation on all instantaneous velocity values throughout the entire post-fall time period—that is, first square, then average, and finally take the square root—to suppress instantaneous jitter interference and highlight continuous motion characteristics, ultimately outputting a stable and physically meaningful post-fall activity rate. The lower this rate value, the weaker the individual's physical activity ability after the fall, and the higher the potential risk.
[0079] In this embodiment, the above-mentioned scheme effectively filters out high-frequency noise while preserving dynamic details by deriving velocity based on real time intervals and performing root mean square smoothing. This makes the assessment of post-fall mobility both conform to kinematic principles and has good engineering robustness, providing a highly reliable behavioral basis for subsequent fall severity grading.
Claims
1. A multimodal health risk early warning system for home-based elderly care, characterized in that, include: The acquisition module is used to acquire continuous video frame data of the target object in the target area and the corresponding radar echo data; The reconstruction module is used to perform image preprocessing and edge detection on the video frame data to obtain a video edge feature map, and to perform clutter suppression and point cloud reconstruction on the radar echo data to obtain a three-dimensional spatial point cloud map. The fusion module is used to map the three-dimensional spatial point cloud map and the video edge feature map to the same spatial coordinate system based on preset camera calibration parameters and radar calibration parameters, and to perform data association and fusion to extract the three-dimensional human skeleton sequence of the target object. The tracking module is used to perform temporal inter-frame correlation and motion tracking on the three-dimensional human skeleton sequence to generate the posture motion trajectory of the target object; The detection module is used to detect fall events on the posture movement trajectory according to the preset fall determination rules, and to determine whether a fall has occurred. The output module is used to classify the severity of a fall when it is determined that a fall has occurred, and to output a corresponding risk warning signal.
2. The home-based elderly care multimodal health risk early warning system according to claim 1, characterized in that, The video frame data undergoes image preprocessing and edge detection to obtain a video edge feature map. Clutter suppression and point cloud reconstruction are performed on the radar echo data to obtain a three-dimensional spatial point cloud map, including: The filtering module is used to perform grayscale conversion and pixel filtering on the video frame data to obtain a grayscale smoothed image, and to extract the contour of the grayscale smoothed image based on pixel gradient calculation to obtain a video edge feature map. The analysis module is used to perform static background cancellation and range frequency domain analysis on the radar echo data to obtain the dynamic target spectrum; The mapping module is used to perform three-dimensional coordinate mapping on the spectrum of the dynamic target based on spatial angle calculation to obtain a three-dimensional spatial point cloud map.
3. The home-based elderly care multimodal health risk early warning system according to claim 1, characterized in that, Based on preset camera calibration parameters and radar calibration parameters, the 3D spatial point cloud map and the video edge feature map are mapped to the same spatial coordinate system, and data association and fusion are performed to extract the 3D human skeleton sequence of the target object, including: A spatial transformation matrix is constructed based on preset camera calibration parameters and radar calibration parameters. The coordinate rotation and translation calculations are performed on the three-dimensional spatial point cloud map and the video edge feature map through the spatial transformation matrix to obtain a set of spatial registration points in the same three-dimensional coordinate system. Based on spatial Euclidean distance, the radar points and visual edge points in the spatial registration point cloud are searched and matched in a neighborhood. The depth values of the successfully matched radar points are assigned to the corresponding visual edge points, and the depth edge point cloud containing three-dimensional coordinate information is generated. Density clustering of the depth edge point cloud is performed based on the topology of the human skeleton, and the spatial centroid of each cluster region is calculated as the human joint point. Based on the topological connection of the human skeleton, the various human joints are connected and arranged according to the time dimension to generate a three-dimensional human skeleton sequence of the target object.
4. The home-based elderly care multimodal health risk early warning system according to claim 3, characterized in that, The spatial transformation matrix is used to perform coordinate rotation and translation calculations on the 3D spatial point cloud map and the video edge feature map to obtain a spatial registration point set in the same 3D coordinate system, including: Based on the spatial transformation matrix, a rotation and translation vector is extracted, and the three-dimensional spatial point cloud map is transformed into a three-dimensional coordinate system using the rotation and translation vector to obtain a visual point cloud. The depth gradient of the visual point cloud and the pixel gradient of the video edge feature map are extracted to perform cross-modal disparity analysis, and the disparity compensation matrix is obtained. Based on the disparity compensation matrix, the depth coordinates of the visual system point cloud are corrected to obtain a corrected 3D point cloud. Based on preset camera calibration parameters, the corrected 3D point cloud and the video edge feature map are physically scaled to obtain a spatial registration point cloud set in the same 3D coordinate system.
5. The home-based elderly care multimodal health risk early warning system according to claim 1, characterized in that, Performing temporal inter-frame correlation and motion tracking on the three-dimensional human skeleton sequence to generate the pose motion trajectory of the target object includes: Spatial coordinate difference calculation is performed on the three-dimensional human skeleton sequence based on the time difference between adjacent frames to obtain joint motion velocity values; Based on the joint motion velocity values, the state prediction and measurement update of the three-dimensional human skeleton sequence are performed to obtain a smooth skeleton sequence; The smooth skeleton sequence is subjected to Euclidean distance calculation of adjacent frame joints based on spatial topology to obtain node matching cost matrix, and the smooth skeleton sequence is subjected to minimum cost association allocation based on the node matching cost matrix to obtain homologous joint sequence. The homologous joint sequences are arranged in timestamp order and spatial coordinate interpolation and curve fitting are performed to obtain a continuous set of spatial coordinates. Based on the spatial topology, the continuous spatial coordinate set is connected by multiple joints to generate the attitude motion trajectory of the target object.
6. The home-based elderly care multimodal health risk early warning system according to claim 5, characterized in that, Based on the node matching cost matrix, the smooth skeleton sequence is assigned a minimum cost association to obtain a sequence of homologous joints, including: The node matching cost matrix is subjected to extreme value threshold filtering and row and column minimum value deduction to obtain a local optimization matrix. The local optimization matrix is then subjected to independent zero element traversal and conflict node allocation to obtain the optimal allocation matrix. Based on the optimal allocation matrix, cross-frame index mapping is performed on the smooth skeleton sequence to obtain rearranged skeleton nodes; Based on the spatial topology, the identity identifiers of the rearranged skeleton nodes are inherited to obtain a sequence of homologous joints.
7. The home-based elderly care multimodal health risk early warning system according to claim 6, characterized in that, Based on the spatial topology, the rearranged skeleton nodes are subjected to identity inheritance to obtain a sequence of homologous joints, including: Based on the spatial topology, the spatial vector difference calculation of adjacent nodes is performed on the rearranged skeleton nodes to obtain the skeleton connectivity vector; The rigid body length constraint and joint rotation angle are analyzed on the skeleton connectivity vector to obtain the topological weight matrix; Based on the topological weight matrix, the rearranged skeleton nodes are subjected to probability weighting of the preceding frame identifier and cross-frame transmission. Nodes with a transmission probability lower than a preset threshold are subjected to topological connectivity inference and identifier assignment to obtain a sequence of homologous joints.
8. The home-based elderly care multimodal health risk early warning system according to claim 1, characterized in that, The posture motion trajectory is subjected to fall event detection according to preset fall determination rules to determine whether a fall has occurred, including: Based on the preset gravity vector, the attitude motion trajectory is orthogonally projected and decoupled from the coordinates to obtain the vertical displacement sequence. Then, the vertical displacement sequence is subjected to time-domain first-order difference and squaring operations to obtain the falling impact sequence. Based on a preset impact threshold, extreme value search and interval truncation are performed on the falling impact sequence to obtain suspected fall segments; Calculate the variance of the trajectory coordinates at the end of the suspected fall segment. When the peak value of the fall impact sequence is greater than the preset impact threshold and the variance of the trajectory coordinates is less than the preset variance threshold, it is determined that a fall has occurred.
9. The home-based elderly care multimodal health risk early warning system according to claim 8, characterized in that, When a fall is detected, the severity of the fall is classified, and a corresponding risk warning signal is output, including: When a fall is determined to have occurred, the fall impact sequence is used to perform time-domain integration within the time interval corresponding to the suspected fall segment to obtain the fall kinetic energy scalar. Based on the timestamp corresponding to the end of the suspected fall segment, the post-fall activity rate is obtained by time-domain truncation and displacement differentiation of the posture motion trajectory. The fall severity index is obtained by normalizing the scalar of fall kinetic energy and the post-fall activity rate based on preset weighting coefficients and then summing them by weight. The fall severity index is numerically compared and mapped to a range based on a preset grading threshold range to obtain a risk level identifier. Based on the risk level identifier, address matching is performed in a preset communication command set to extract the corresponding control message and generate the risk warning signal.
10. The home-based elderly care multimodal health risk early warning system according to claim 9, characterized in that, Based on the timestamp corresponding to the end of the suspected fall segment, the post-fall activity rate is obtained by time-domain truncation and displacement differentiation of the posture motion trajectory, including: Based on the timestamp corresponding to the end of the suspected fall segment, the posture motion trajectory is truncated by forward sliding along the time axis to obtain the post-fall trajectory sequence. The spatial coordinate difference between adjacent frames is then performed on the post-fall trajectory sequence to obtain the relative displacement vector. The relative displacement vector is differentiated over time based on the time difference between adjacent frames to obtain the instantaneous velocity vector. The norm modulus and root mean square time of the instantaneous velocity vector are then calculated to obtain the post-fall activity rate.