Old people safety monitoring method and device based on multi-modal sensor fusion

By combining multimodal sensor fusion technology and edge computing architecture with AI intelligent voice interaction, a full-process emergency response plan is generated, which solves the problems of decreased accuracy and high false alarm rate of fall detection systems for the elderly in special environments, and realizes intelligent, safe and reliable home safety monitoring for the elderly.

CN121545288APending Publication Date: 2026-02-17SHENZHEN NO 1 VOCATIONAL & TECH SCHOOL
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511735606.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing fall detection and safety monitoring systems for the elderly suffer from decreased detection accuracy and high false alarm rates in special environments such as insufficient lighting, obstructions, and bathrooms. They also lack intelligent voice interaction, have low emergency response efficiency, pose a significant risk of privacy leaks, and have poor system scalability, making it difficult to achieve intelligent, continuous, safe, and reliable home safety monitoring.

Method used

Employing multimodal sensor fusion technology, multimodal monitoring data is collected through millimeter-wave radar, infrared thermal imaging, microphone arrays, and cameras. Combined with edge data encryption and local caching, standardized multidimensional feature vectors are generated. Utilizing a dynamic weighted multimodal fusion engine and an AI intelligent agent voice interaction confirmation mechanism, a full-process emergency response execution plan is generated. Relying on edge computing architecture and cloud-based collaborative optimization, safety monitoring for the elderly is achieved.

Benefits of technology

It improves the accuracy and reliability of fall detection, reduces the false alarm rate, enables intelligent voice interaction, enhances emergency response efficiency and data security, improves the system's scalability and adaptability, and ensures the safety of the elderly at home.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545288A_ABST
    Figure CN121545288A_ABST
Patent Text Reader

Abstract

The invention discloses an old people safety monitoring method and equipment based on multi-modal sensor fusion, which are applied to the technical field of data processing, and the method comprises the steps: collecting multi-modal monitoring, environment state and equipment operation data, and carrying out synchronous correction, feature extraction and edge end encryption caching to generate a standardized multi-dimensional feature vector; by means of a dynamic weight multi-modal fusion engine, in combination with adaptive weight distribution and scene association analysis, suspected tumble events are recognized, and then the real tumble situation is confirmed through AI agent voice interaction. Based on an emergency linkage decision matrix, a priority notification and a home custom control rule are fused, an emergency linkage scheme is generated, and finally, based on an edge-cloud collaboration architecture and an adaptive adjustment strategy, the security dynamic protection and continuous optimization of the old people are realized, and the home security of the old people is comprehensively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an old person safety monitoring method and device based on multi-modal sensor fusion. BACKGROUND

[0002] Although the existing old person fall detection and safety monitoring system can realize home monitoring and alarm to some extent, it still has many defects. First, most systems rely only on a single type of sensor for detection, such as millimeter wave radar or camera, lack multi-modal data fusion capability, resulting in a significant decrease in detection accuracy in low light, obstruction, bathroom and other special environments, and easy false positives or false negatives. Second, the existing technology generally lacks intelligent voice interaction and confirmation mechanism. When the system detects a suspected fall event, it usually directly issues an alarm or pushes information, and cannot confirm the status through voice with the old person, so the false alarm rate is high and the actual use experience is poor.

[0003] In addition, the existing system stays in the simple alarm stage in the emergency disposal link and cannot form a complete closed loop from fall detection, voice confirmation, contact notification to medical emergency linkage, with low emergency response efficiency. On the other hand, many visual recognition-based solutions rely on cameras or cloud analysis and processing, which not only have privacy leakage risks, but also affect real-time and reliability when the network is interrupted or delayed. Furthermore, some systems use centralized cloud processing architecture, lack edge computing capability and local encryption mechanism, and have poor data security. Finally, the existing fall detection system has insufficient scalability and compatibility, different brands of devices are difficult to interconnect, the algorithm model cannot adapt to environmental changes, the system function is single and maintainability is poor, and it is difficult to meet the needs of long-term home safety monitoring of the elderly. In summary, the existing technology generally has the problems of single detection means, insufficient recognition accuracy, lack of interaction capability, imperfect emergency response, weak privacy protection and poor system scalability, and it is difficult to realize intelligent, continuous and safe and reliable home safety monitoring of the elderly. SUMMARY

[0004] To solve the above technical problems, the present application provides the following technical solutions: The application discloses an old-age safety monitoring method based on multi-modal sensor fusion, and comprises the following steps: acquiring multi-modal monitoring data, environment state data and device operation data; performing data preprocessing on the multi-modal monitoring data, introducing a synchronous correction algorithm and a multi-modal feature extraction algorithm, combining edge data encryption and local caching mechanism, and generating a standardized multi-dimensional feature vector; processing the environment state data and the device operation data based on a dynamic weight multi-modal fusion engine, combining fusion confidence calculation and abnormal data filtering through an adaptive weight distributor and scene correlation analysis, and generating suspected fall event identification information; processing the standardized multi-dimensional feature vector and the suspected fall event identification information, and generating a real fall event judgment result through an AI intelligent body voice interaction confirmation mechanism; processing the real fall event judgment result, preset emergency contact information and home device control authority in combination with an emergency linkage decision matrix, fusing priority notification strategies and home scene customized control rules, and generating a full-process emergency linkage execution scheme; processing the real fall event judgment result, the full-process emergency linkage execution scheme and abnormal feature identification information based on an edge computing architecture and a cloud collaborative optimization engine in combination with a multi-modal adaptive adjustment strategy, and generating an old-age safety monitoring dynamic protection and continuous optimization result.

[0005] The application discloses an old-age safety monitoring device based on multi-modal sensor fusion, and the device comprises a construction module and a processing module; the construction module is used for acquiring multi-modal monitoring data, environment state data and device operation data; the processing module is used for performing data preprocessing on the multi-modal monitoring data, introducing a synchronous correction algorithm and a multi-modal feature extraction algorithm, combining edge data encryption and local caching mechanism, and generating a standardized multi-dimensional feature vector; processing the environment state data and the device operation data based on a dynamic weight multi-modal fusion engine, combining fusion confidence calculation and abnormal data filtering through an adaptive weight distributor and scene correlation analysis, and generating suspected fall event identification information; processing the standardized multi-dimensional feature vector and the suspected fall event identification information, and generating a real fall event judgment result through an AI intelligent body voice interaction confirmation mechanism; processing the real fall event judgment result, preset emergency contact information and home device control authority in combination with an emergency linkage decision matrix, fusing priority notification strategies and home scene customized control rules, and generating a full-process emergency linkage execution scheme; processing the real fall event judgment result, the full-process emergency linkage execution scheme and abnormal feature identification information based on an edge computing architecture and a cloud collaborative optimization engine in combination with a multi-modal adaptive adjustment strategy, and generating an old-age safety monitoring dynamic protection and continuous optimization result.

[0006] The beneficial effect is that the application provides an old people safety monitoring method based on multi-modal sensor fusion, through collecting multi-modal monitoring, environment state and equipment operation data, generating standardized multi-dimensional feature vectors through synchronous correction, feature extraction, edge encryption cache. With the help of dynamic weight multi-modal fusion engine, combined with adaptive weight distribution and scene correlation analysis, identify suspected fall events, and then confirm the real fall situation through AI agent voice interaction. Based on the emergency linkage decision matrix, the priority notification and the home customized control rule are fused to generate an emergency linkage scheme. Finally, relying on the edge-cloud collaborative architecture and adaptive adjustment strategy, the dynamic protection and continuous optimization of the old people safety are realized, and the home safety of the old people is fully guaranteed. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 A flow chart of an old people safety monitoring method based on multi-modal sensor fusion provided by the embodiment of the application is provided. Figure 2 A module schematic diagram of an old people safety monitoring device based on multi-modal sensor fusion provided by the embodiment of the application is provided. DETAILED DESCRIPTION

[0008] The preferred embodiments of the application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to explain and illustrate the application, and are not used to limit the application. The following describes an old people safety monitoring method based on multi-modal sensor fusion according to the exemplary embodiments of the application in conjunction with the accompanying drawings. Figure 1 The old people safety monitoring method based on multi-modal sensor fusion according to the exemplary embodiments of the application is described below in conjunction with the accompanying drawings.

[0009] In the embodiments of the application, an old people safety monitoring method based on multi-modal sensor fusion is as shown in Figure 1 S101, multi-modal monitoring data, environment state data and equipment operation data are acquired.

[0010] ​In an embodiment, the multi-modal monitoring data is synchronously collected by four types of sensors: millimeter wave radar, infrared thermal imaging, microphone array, and camera. All sensors are uniformly connected to the system LAN, and data is collected at fixed time intervals (100 ms) and attached with a uniform UTC timestamp to ensure data time sequence consistency. A 24 GHz frequency-modulated continuous wave radar is selected and installed 1.5 meters from the side of the bed in the bedroom. The radar beam covers the human sleep area, and spatial point cloud data and motion trajectory information of the human body are collected. The radar captures the chest fluctuation and limb movement of the elderly during sleep in real time, and the output point cloud data can present the human body contour coordinates (such as the torso center point coordinates X=120 cm, Y=80 cm, Z=50 cm). The motion trajectory data records the displacement vector of the turning-over action (such as moving 15 cm along the X-axis, taking 2 seconds). An 8-14 pm waveband infrared thermal imager is used and installed on the same side above the radar to collect human thermal distribution characteristics. In the dark environment at night, the thermal imager generates a gray thermal image, and the human thermal center point coordinates (X=118 cm, Y=78 cm) are extracted. The temperature change rate data shows that the chest region temperature fluctuation is ≤0.3°C during calm breathing, and the local temperature drops 0.8°C due to limb contact with the ground during a fall.

[0011] A 4-channel microphone array is used and placed on the bedside cabinet in the bedroom to collect environmental sound, voice, and fall impact sound. Through MFCC feature extraction technology, the ground impact sound during a fall (amplitude ≥ 60 dB, lasting 0.5 seconds), the audio features of the old person's distress voice "help" (frequency range 300-3000 Hz, energy concentrated near 1000 Hz), and the background noise in the daily environment (amplitude ≤ 30 dB) are captured. A 1080P near-infrared camera is selected, which supports 850 nm light supplement, and is installed on the wall opposite the bed in the bedroom to collect environmental pictures and human posture changes. The camera generates a gray-scale image every 15 fps, clearly presenting the human skeleton key points, such as the head key point coordinates (X=125 cm, Y=160 cm) when standing, and the head key point coordinates changing to (X=123 cm, Y=65 cm) after falling, and the center of gravity height dropping from 110 cm to 55 cm.

[0012] The environmental parameters of the monitoring scene are collected in real time by the environmental perception module of the sensor and the external environmental sensor, and are bound to the same timestamp as the multi-modal monitoring data to ensure the relevance of the scene and the monitoring data. The light intensity is collected by the built-in light sensor of the camera, with a quantization range of 0-1000 lux, used to determine the light and dark state of the environment. The light intensity is 5 lux after turning off the light at night, meeting the low-light environment determination standard; the light intensity is 800 lux after opening the curtains during the day, belonging to the normal light scene. The radar signal penetration rate and the camera image occlusion recognition result are combined to determine the occlusion state value, with a quantization range of 0-1 (0 for no occlusion and 1 for complete occlusion). When the old man covers a thin quilt, the radar signal penetration rate is 90%, and the occlusion state value is 0.1; when covering a thick quilt, the radar signal penetration rate is 60%, and the occlusion state value is 0.4; when the seat blocks the camera's line of sight, making it impossible to identify the chest and abdomen area of the human body, the occlusion state value is 0.8. The temperature range is 0-50°C and the humidity range is 20%-90%RH, collected by the external temperature and humidity sensor. The temperature in the bedroom is 24°C and the humidity is 55%RH, which is in the comfortable environment interval; the temperature in the bathroom is 30°C and the humidity is 85%RH, which is recorded as a high-humidity scene. The microphone array collects and filters the data, with a quantization range of 0-100dB. The daily environmental noise is 35dB, meeting the quiet monitoring conditions; the noise transmitted from the kitchen when cooking is 65dB, marked as environmental interference noise.

[0013] The system automatically scans all connected sensor devices in the local area network, collects device running state parameters in real time, and stores them in the local device management table (DeviceID, Type, Status, Parameter), ensuring that device failures can be timely warned and switching to backup collection strategies. The power supply state, network connection state and collection state of the sensor are monitored in real time, and are identified as "normal / abnormal". The millimeter wave radar DeviceID is RAD001, and the working state is displayed as "normal" with a network connection delay of ≤10ms; the camera is disconnected due to loose wiring, DeviceID is CAM001, working state is displayed as "abnormal", and device failure alarm is triggered. The actual sampling frequency, resolution and other core parameters of each sensor are recorded. The infrared thermal imager has a sampling frequency of 30fps and a resolution of 640x480; the microphone array has a sampling rate of 20000Hz and a quantization bit of 16bit; the millimeter wave radar has a sampling time interval of 10ms and a point cloud density of 100 points / frame. The signal quality is evaluated by calculating the signal-to-noise ratio (SNR) and data integrity (percentage of non-missing frames) of the sensor output signal. The millimeter wave radar signal-to-noise ratio is 45dB, and the data integrity is 99.8%; the camera signal-to-noise ratio is 28dB in low-light environment, and the data integrity is 98.5%; the microphone data integrity decreases to 95% when the network fluctuates, marked as signal quality degradation.

[0014] S102, data preprocessing is performed on the multi-modal monitoring data, a synchronization correction algorithm and a multi-modal feature extraction algorithm are introduced, and an edge data encryption and local cache mechanism is combined to generate a standardized multi-dimensional feature vector.

[0015] In an embodiment, the multi-modal monitoring data is classified and preprocessed for analysis, and an adaptive association result of data type and processing algorithm is generated. The multi-modal monitoring data includes millimeter wave radar spatial point cloud and motion trajectory data, infrared thermal imaging human body heat distribution data, microphone environmental sound and voice data, and camera environmental picture and human body posture data. First, the multi-modal monitoring data is classified and sorted according to the source and type, the core attributes and processing requirements of each data are determined, and the corresponding preprocessing algorithm is matched to form the adaptive association result of data type and processing algorithm, so as to ensure the pertinence and effectiveness of the processing process. The millimeter wave radar spatial point cloud and motion trajectory data contain a set of human body three-dimensional spatial coordinate points and a time sequence sequence of motion trajectories, and there are environmental noise (such as furniture reflection interference) and isolated point redundancy. The adaptive associated preprocessing algorithm is a Gaussian filter algorithm (used to remove high-frequency noise) and a DBSCAN clustering algorithm (used to remove isolated points and aggregate human body point cloud clusters). The original point cloud data collected by the radar in a certain period of time contains 500 spatial coordinate points, of which 30 are isolated points of furniture reflection. After filtering noise by Gaussian filtering (standard deviation σ=0.8), the data is processed by DBSCAN clustering (neighborhood radius ε=0.3m, minimum point number min_samples=5), and finally a human body point cloud cluster containing 470 points and corresponding motion trajectory data (such as continuous displacement sequence along the Y-axis direction) is obtained.

[0016] The infrared thermal imaging human body heat distribution data is a gray thermal imaging image, which contains human body heat radiation distribution information, and has background thermal noise and image blur problems. The adaptive associated preprocessing algorithm is a median filter algorithm (used to suppress background thermal noise) and a histogram equalization algorithm (used to enhance the thermal contrast between the human body and the background). In the infrared thermal imaging image collected at night, the thermal contrast between the human body area and the background area is low (gray difference ≤20). After removing the background random thermal noise by 3x3 median filtering, the human body area gray value range is expanded from 80-120 to 50-180 after histogram equalization processing, and the heat distribution contour is clear and distinguishable, and the human body heat center position can be accurately extracted.

[0017] The microphone ambient sound and voice data are audio time domain signals, containing environmental noise, human voice and fall impact sound, etc., with noise interference and signal amplitude fluctuation. The adaptive associated preprocessing algorithm is Hanning windowing processing (used to reduce spectral leakage) and adaptive noise suppression algorithm (used to separate target sound and environmental noise). In the original audio signal collected by the microphone, there are 35 dB of environmental background noise and 60 dB of fall impact sound. After Hanning windowing processing (window length 20 ms) of the signal, the adaptive noise suppression algorithm (convergence factor μ = 0.02) is used to filter the background noise, and finally the pure impact sound signal and the voice signal segment (such as the old man's cry for help) are obtained. The camera environment picture and human body posture data are continuous near-infrared image frames, containing human body posture and environmental scene information, with image distortion, uneven illumination and background redundancy problems. The adaptive associated preprocessing algorithm is Zhang's calibration algorithm (used to correct image distortion) and adaptive threshold segmentation algorithm (used to separate human body and background). The original image collected by the camera causes the human body contour to be inclined due to lens distortion. After the Zhang's calibration algorithm (based on a chessboard calibration board, the correction parameters are focal length f = 1200px, principal point coordinates (u0, v0) = (640, 360)), the adaptive threshold segmentation (threshold range 120-200) is used to separate the human body and the background, and the binary image containing only the human body region is obtained, laying a foundation for subsequent posture key point extraction.

[0018] The adaptive associated results are processed in a targeted manner to generate a feature extraction variable list, including radar signal filtering and point cloud clustering variables, infrared thermal image human body heat center extraction variables, visual posture skeleton key point extraction variables, audio MFCC features and energy envelope analysis variables. Based on the adaptive associated results of the above data types and processing algorithms, the core feature extraction dimensions are targetedly selected for each data type, the definition, calculation method and physical meaning of the feature variables are clarified, the standardized feature extraction variable list is generated, and the uniformity and integrity of feature extraction are ensured. The radar signal filtering and point cloud clustering variables focus on human body movement and spatial features, including three core variables. Variable 1 is the point cloud cluster barycenter coordinates (Xc, Yc, Zc), the calculation method is the mean of all point coordinates in the human body point cloud cluster; variable 2 is the point cloud cluster volume (V), the calculation method is the length x width x height of the point cloud cluster bounding box; variable 3 is the motion trajectory speed (v), the calculation method is the ratio of the Euclidean distance of the barycenter coordinates of adjacent time instants to the time interval. The barycenter coordinates of the human body point cloud cluster at a certain time instant are (1.2 m, 0.8 m, 0.5 m), the point cloud cluster bounding box size is 0.6 m x 0.4 m x 1.5 m, the volume V = 0.36 m³, and the Euclidean distance of the barycenter of the adjacent two frames (time interval 100 ms) is 0.05 m, then the motion speed v = 0.5 m / s.

[0019] The variable focusing on the human thermal distribution characteristics contains two core variables. Variable 1 is the human thermal center coordinates (Hx, Hy), which is calculated by the weighted center of gravity of the human region in the thermal image. Variable 2 is the thermal center temperature change rate (ΔT / Δt), which is calculated by the ratio of the temperature difference and the time interval of the thermal center corresponding to adjacent time. In a certain frame of infrared thermal image, the human thermal center coordinates are (320px, 240px), the corresponding temperature is 36.5℃, the thermal center temperature of the next frame (time interval 300ms) is 36.3℃, and the temperature change rate ΔT / Δt is -0.67℃ / s. The variable focusing on the human posture characteristics contains 10 core variables (selecting the human key skeleton nodes), which are the head key point coordinates (Hx, Hy), the neck key point coordinates (Nx, Ny), the left shoulder key point coordinates (Sx1, Sy1), the right shoulder key point coordinates (Sx2, Sy2), the left hip key point coordinates (Hx1, Hy1), the right hip key point coordinates (Hx2, Hy2), and the center of gravity height (H) (calculated by the average of the Y coordinates of the hip key point and the head key point). In a certain frame of image, the extracted human posture key point coordinates are head (330px, 120px), neck (330px, 150px), left shoulder (300px, 160px), right shoulder (360px, 160px), left hip (310px, 280px), and right hip (350px, 280px). The center of gravity height H is (280+120) / 2=200px.

[0020] The variable focusing on the audio signal characteristics contains 8 core variables, which are 12-order MFCC coefficients (C1-C12), audio energy (E) (calculated by the sum of the square of the audio signal amplitude), energy envelope peak (Emax), and energy envelope rise time (Tr) (calculated by the time from the energy 10% peak to 90% peak). The MFCC feature extraction is performed on the audio segment (0.5s) of the fall impact sound, and C1=12.5, C2=8.3, …, C12=2.1, audio energy E=5.2×10 4 , energy envelope peak Emax=8.6×10 4 , and energy envelope rise time Tr=0.08s.

[0021] The adaptive association result, the feature extraction variable list and the original multi-modal monitoring data are synchronized, corrected and encrypted, and combined with the edge local cache mechanism to generate a standardized multi-dimensional feature vector. The adaptive association result, the feature extraction variable list and the original multi-modal monitoring data are time-synchronized, corrected and encrypted, and finally all feature variables are standardized into a unified format of multi-dimensional feature vector, ensuring the time sequence consistency, security and fusion of the data. Taking the system unified UTC timestamp as the basis (accurate to millisecond level), the collection time of the four types of modal data is aligned and calibrated. For the sampling frequency difference of radar data (sampling interval 10 ms), infrared thermal imaging data (sampling interval 33 ms), microphone data (sampling interval 5 ms) and camera data (sampling interval 67 ms), the linear interpolation method is used to supplement the feature variable values of the missing time, ensuring the complete correspondence of the feature variables of the four types of data at the same time node. At a certain time node t = 1699999999800 ms, the radar data and the microphone data have collection values, and the infrared thermal imaging data and the camera data have no collection values. The linear interpolation method is used to calculate the infrared thermal center temperature change rate and the camera gravity height at this time, realizing the time synchronization of the four types of data.

[0022] The synchronized feature data is encrypted at the edge using the AES-256 encryption algorithm, and the key is generated and stored locally by the system (not uploaded to the cloud), and sensitive data (such as human posture images, voice clips) are desensitized (retain feature variables, remove original image / audio data). After the human body gravity height, thermal center temperature and other feature variables are synchronized, the encrypted data is stored in the local cache, and the desensitized voice data only retains the MFCC coefficient and energy features, and the original voice waveform data is automatically deleted. The encrypted synchronized data is stored in the local cache of the edge device (capacity ≥ 100 GB), and the "first-in, first-out" strategy is used to manage the cache data (cache validity period 7 days), and the data collection time, device ID, encryption state and other metadata are recorded, facilitating subsequent tracing and calling. A single data record stored in the local cache includes a timestamp, a device ID list, an encrypted feature variable set, and a data integrity check code, ensuring data storage security and fast reading.

[0023] All feature extraction variables are arranged in the order of "radar features-infrared features-visual features-audio features", and each variable is normalized (mapped to the [0, 1] interval) to finally generate a standardized multi-dimensional feature vector with a fixed dimension (3+2+10+8=23 dimensions). The normalized feature vector at a certain time node is [0.62, 0.35, 0.48, 0.51, 0.32, 0.45, 0.38, 0.52, 0.49, 0.53, 0.41, 0.56, 0.43, 0.55, 0.47, 0.50, 0.39, 0.46, 0.54, 0.37, 0.42, 0.57, 0.40], where the first 3 are radar features, 4-5 are infrared features, 6-15 are visual features, and 16-23 are audio features. All values are in the [0, 1] interval and can be directly input into the subsequent fusion model for analysis.

[0024] In S103, the environment state data and the equipment operation data are processed based on a dynamic weight multi-modal fusion engine, through an adaptive weight distributor and scene correlation analysis, combined with fusion confidence calculation and abnormal data filtering, to generate suspected fall event identification information.

[0025] In an implementation, for edge devices with limited computing resources, the multi-modal body motion feature information is limited to near-infrared visual body motion feature information and radar radio frequency micro-motion feature information; wherein the visual body motion feature includes whole body posture mutation amplitude P, local limb abnormal motion frequency f, center of gravity height change rate H / t, the radar radio frequency feature includes body displacement amplitude S, motion acceleration change rate a / t, micro-Doppler frequency ; b. Based on the multi-modal cooperative anti-interference target, the near-infrared visual feature is associated with the radar radio frequency feature, and the partial information decomposition theory (PID) is used to quantify the redundant information of the two types of modalities , visual unique information , radar unique information and cooperative information to generate multi-modal body motion feature fusion basis information; c. According to the scene adaptation and individual adaptation targets, the low light / shading environment interference factor and the individual body shape / body motion frequency adaptation factor are respectively input into the feature weight dynamic adjustment model and the fusion strategy adaptive optimization model, and the formula =0.5×(1-L)×(1-O)+0.3× 、 =1- The visual and radar modal weights are calculated to generate dynamic weight distribution adaptation information; d. The single modal confidence is calculated by substituting into the visual fall determination model (Sigmoid function) and the radar signal displacement analysis model respectively, and the fusion confidence is calculated by using a weighted summation formula, without a special conflict processing mechanism, and the result of the high weight mode is adopted by default.

[0026] Specifically, the environmental state data and equipment operation data are classified, extracted and feature matched to generate near-infrared visual body motion feature information, radar radio frequency micro-motion feature information, low light / shading environment interference factor information, and individual body shape / body motion frequency adaptation factor information. The environmental state data (light intensity, shading degree, temperature and humidity) and equipment operation data (sensor signal-to-noise ratio, sampling frequency, signal integrity) are classified and disassembled, and the core features strongly related to body motion recognition, scene adaptation and individual adaptation are extracted. Through feature matching, data correlation is established to generate four types of key information, providing basic input for subsequent fusion analysis. The environmental picture and human posture data collected from the near-infrared camera are extracted, focusing on human motion state features, and the core features include the whole body posture mutation amplitude (normalized to the 0-1 interval, 1 represents the maximum posture mutation), local limb abnormal motion frequency f (unit: times / s), and center of gravity height change rate (unit: cm / s, negative value represents center of gravity descent). Feature extraction is achieved through human skeleton key point algorithm. During the process of the old people from standing to falling, the head key point Y coordinate of the continuous two frames of images decreases from 160px to 65px, and the trunk key point displacement deviation is 12px. The calculation result is =0.85, the upper limb abnormal motion frequency f=3.2 times / s, and the center of gravity height change rate =(65-160) / 0.5=-190cm / s, and the near-infrared visual body motion feature information =0.85, f=3.2 times / s, =-190cm / s.

[0027] The spatial point cloud and motion trajectory data collected from the millimeter wave radar are extracted, focusing on the penetrating human micro-motion features, and the core features include the body overall displacement amplitude S (unit: m), motion acceleration change rate (unit: m / s³), and micro-Doppler frequency (unit: Hz). Feature extraction is achieved through radar signal filtering (Gaussian filtering, standard deviation =0.6 and point cloud clustering (DBSCAN algorithm, neighborhood radius ​​​​= 0.4 m. The radar monitors the displacement sequence of the human body within 3 seconds as [0, 0.3, 0.9] m, and the overall body displacement amplitude S = 0.9 m is calculated; the acceleration sequence is [0.2, 0.5, 1.1] m / s², and the acceleration rate of change / = (1.1-0.2) / 3 = 0.3 m / s³; the micro-Doppler frequency = 1.5 Hz, and the radar radio frequency micro-motion feature information {S = 0.9 m, / = 0.3 m / s³}, = 1.5 Hz} is integrated.

[0028] The low light / shading environment interference factor information is extracted from the environment state data, quantifying the interference degree of the environment on the monitoring, and the core factors include the low light intensity factor L and the shading degree factor O (both normalized to the interval of 0-1, and the greater the value, the stronger the interference). The low light intensity factor is calculated based on the built-in light sensor data of the camera, and the formula is L = 1-min (light intensity / 500, 1) (500 lux is the normal light threshold); the shading degree factor is calculated in combination with the radar signal penetration rate T and the camera image shading recognition result, and the formula is = 0.6x (1-T) + 0.4x image shading rate. The night bedroom light intensity is 8 lux, and L = 1-8 / 500 = 0.984 is calculated; the old man covers a thick quilt, resulting in a radar signal penetration rate T = 55%, and the body chest and abdomen area shading rate in the camera image is 60%, and = 0.6x (1-0.55) + 0.4x 0.6 = 0.27 + 0.24 = 0.51 is calculated, and the low light / shading environment interference factor information {L = 0.984, O = 0.51} is integrated.

[0029] Based on the individual historical monitoring data and real-time body motion data extracted from the device operation data, individual differences are adapted, and the core factors include the body type adaptation factor B and the body motion frequency adaptation factor M (both normalized to the interval of 0-1). The body type adaptation factor is calculated according to the radar point cloud cluster volume V, and the formula is B = min (V / 0.5, 1) (0.5 m³ is the standard body volume threshold); the body motion frequency adaptation factor is calculated based on the real-time body motion frequency and the historical average body motion frequency comparison, and the formula is (taking 1.2 when it exceeds 1.2, indicating that the body motion frequency is significantly higher than the historical level). The human body volume V = 0.42 m³ is calculated by the radar point cloud cluster, and B = 0.42 / 0.5 = 0.84 is calculated; the historical average body motion frequency = 0.8 times / minute, and the real-time body motion frequency =2.5 times / min, calculated M=2.5 / 0.8=3.125 (take 1.2), integration of individual body shape / body motion frequency adaptation factor information {B=0.84, M=1.2}.

[0030] Based on the multi-modal collaborative anti-interference target, the whole body posture mutation feature in near-infrared vision, the local limb abnormal motion feature and the penetrating body displacement feature in radar radio frequency are associated, and the partial information decomposition theory is used to quantify the redundancy, unique and collaborative information of the two types of modalities in body motion recognition, to generate multi-modal body motion feature fusion basis information. Based on the multi-modal collaborative anti-interference target, the whole body posture mutation feature in near-infrared vision, the local limb abnormal motion feature, and the penetrating body displacement feature in radar radio frequency (time synchronization is realized through unified UTC timestamp, accuracy 1ms) are associated, and the partial information decomposition theory (PID) is used to quantify the redundancy information , unique information (visual unique , radar unique and collaborative information of two types of modalities in body motion recognition, to generate multi-modal body motion feature fusion basis information.

[0031] The core calculation formula is as follows, single modal information entropy: =H(Y)-H(Y| )( modal feature, Y is the body motion recognition target, H is the information entropy); joint information entropy: =H(Y)-H(Y| ); redundancy information: =I( ;Y| ) I( ;Y| ) (take the intersection of the conditional information entropy of the two types of modalities); unique information: =I( ;Y)- , =I( ;Y)- ; collaborative information: =I( , ;Y)-I( ;Y)-I( ;Y)+ . After calculation, the visual modal information entropy I( ;Y)=0.78, the radar modal information entropy I( ;Y)=0.86, the joint information entropy =1.35, the redundancy information =0.52, then visual unique information =0.78-0.52=0.26, unique radar information =0.86-0.52=0.34, Collaborative Information =1.35-0.78-0.86+0.52=0.23, generating multimodal volume motion feature fusion based on information. =0.52, =0.26, =0.34, =0.23}, which indicates that the unique information of the radar mode is more significant. There is a certain redundancy between the two modes (such as both can reflect body displacement) and synergy (visual posture + radar penetration displacement can improve the recognition accuracy of occluded scenes).

[0032] Based on the goals of scene adaptation and individual adaptation, the low-light / occlusion environment interference factor and the individual body size / body movement frequency adaptation factor are respectively assigned to the feature weight dynamic adjustment model and the fusion strategy adaptive optimization model to generate dynamic weight allocation adaptation information. Based on the goals of scene adaptation (addressing low-light / occlusion interference) and individual adaptation (matching differences in body size / body movement frequency), the low-light / occlusion environment interference factor and the individual body size / body movement frequency adaptation factor are respectively input into the feature weight dynamic adjustment model and the fusion strategy adaptive optimization model to calculate modal weights and fusion threshold parameters, generating dynamic weight allocation adaptation information.

[0033] The feature weight dynamic adjustment model (computational vision / radar modal weights) uses the low light intensity factor L and the occlusion degree factor. and multimodal unique information , As input, output visual modality weights Radar mode weights (satisfy + =1), the core formula is as follows: =0.5×(1-L)×(1-O)+0.3× , =1- Substituting L=0.984, =0.51、 =0.26、 =0.34, calculated as follows =0.5×(1-0.984)×(1-0.51)+0.3× =0.5×0.016×0.49+0.3×0.433≈0.0039+0.13=0.134, =1-0.134=0.866, the result shows that the radar modal weight needs to be significantly improved (utilizing its penetration detection advantage) under low light shielding environment.

[0034] Take body size adaptation factor B and body motion frequency adaptation factor M as input, output fusion threshold adjustment coefficient k (used for dynamically adjusting the fall determination threshold, adapting to individual differences), the core formula is as follows, k=1+0.2×B+0.3×min(M,1) (M exceeds 1, take 1, avoid excessive adjustment). Substituting B=0.84, M=1.2 (take 1), k=1+0.2×0.84+0.3×1=1+0.168+0.3=1.468 is calculated. The result shows that for individuals with this body size (obese) and high body motion frequency, the fusion determination threshold needs to be improved to reduce false positives caused by normal body motion. Integrate the above calculation results to generate dynamic weight distribution adaptation information =0.134, =0.866, k=1.468.

[0035] Based on the multi-modal body motion feature fusion information and dynamic weight distribution adaptation information, the core body motion monitoring data in the environment state data and device operation data are marked and calculated to generate the initial identification result of suspected fall events. Mark the core data in the near-infrared visual body motion feature 、 / ), substitute the visual fall determination model (Sigmoid function normalized confidence), the core formula is as follows, the visual fall confidence , is the Sigmoid function, 200 is the normalized coefficient of / , unit cm / s). Substitute =0.134、 =0.85、 / =-190cm / s, calculate ≈0.530, the visual dimension determination standard, if ≥k×0.5 (k=1.468), that is, the threshold value=0.734), it is determined as a suspected fall; here 0.530<0.734, the visual dimension is not determined as a suspected fall.

[0036] Mark the core radio frequency body motion data, substitute the radar signal displacement analysis model, calculate the body overall displacement amplitude and motion acceleration change rate, identify the radio frequency dimension penetrating suspected fall displacement feature, combine the dynamic weight distribution strategy to fuse two kinds of features, through fusion confidence calculation and abnormal data filtering, generate the initial identification result of suspected fall events. Mark the core data in the radar radio frequency micro-motion feature (S, ), the radar fall confidence is calculated by substituting the radar signal displacement analytical model, and the core formula is as follows: radar fall confidence (1.2 is the normalized coefficient of S, unit m; 0.6 is the normalized coefficient of S, unit m / s³).

[0037] Substituting =0.866, S=0.9m, =0.3m / s³, it is calculated that ≈0.632. The radar dimension judgment standard is that if ≥k×0.5=0.734, it is determined to be a suspected fall; here, 0.632<0.734, the radar dimension is not determined to be a suspected fall; if subsequent monitoring is S=1.1m, =0.5m / s³, it is calculated that ≈ (0.866×(0.458+0.417))≈ (0.75)≈0.682, which still does not reach the threshold value; when S=1.2m, =0.6m / s³, ≈ (0.866×(0.5+0.5))≈ (0.866)≈0.704, close to the threshold value.

[0038] The initial identification result of the suspected fall event is integrated and rule checked to generate the suspected fall event identification information. The weighted sum formula is used to fuse the vision and radar confidence, and the formula is C= × + × , substituting =0.530, =0.704, it is calculated that C=0.134×0.530+0.866×0.704≈0.071+0.610≈0.681. Abnormal data filtering, the abnormal data judgment rule is set to eliminate invalid data caused by sensor failure and environmental mutation, and the rules include: ① radar signal signal-to-noise ratio SNR<25dB; ② vision image blur rate>35%; ③ single mode confidence and fusion confidence deviation>0.3. At a certain moment, the radar signal-to-noise ratio SNR=32dB, the vision image blur rate is 18%, and the single mode and fusion confidence deviation is less than 0.2, which is determined to be valid.

[0039] ​In another embodiment, for full-featured devices with sufficient computing power. Multimodal body movement feature information covers radar speed and micro-Doppler features, visual posture and motion features, audio impact and distress features, infrared thermal distribution features; among them, the audio features include impact sound amplitude, distress speech MFCC coefficient, and infrared features include human thermal center coordinates, thermal center temperature change rate; b. Calculate the environmental quality index and confidence for each modality to obtain the credibility score (Radar based on signal-to-noise ratio, vision based on light clarity and occlusion rate, audio based on signal amplitude, infrared based on temperature resolution); adjust the environment adaptation coefficient dynamically in combination with the home scene (Radar in high-humidity bathroom scene =0.8, thermal image =0.9, visual in low-light bedroom scene =0.5, radar =0.9; c. Calculate the weight of each modality through a dynamic weight formula , where the adjustment parameter is calibrated based on 1000 groups of historical fall case data; d. Use an interpretable Bayesian or D-S evidence fusion mechanism to handle modality information conflicts (such as radar determining a fall and vision determining normal, suppressing the weight of visual low-confidence data to below 0.1), combined with time series judgment, analyze the fusion features of the last 10 frames, and set the fusion confidence ≥0.8 and the last 3 frames meet the conditions as the suspected fall determination standard.

[0040] Specifically, through an adaptive weight distribution mechanism, the weight of each modality in the fusion is automatically adjusted according to real-time environmental noise and sensor quality. The system first synchronizes and preprocesses the data of each modality, extracts the speed and micro-Doppler features of the radar, the posture and motion information of the vision, and the impact and distress features of the audio. Then, for each modality, calculate the environmental quality index and confidence (such as signal-to-noise ratio, light clarity, occlusion rate, etc.), and obtain the credibility score . In the fusion stage, the model dynamically calculates the weight of each modality according to these real-time indicators , the formula is ; where represents the environment adaptation coefficient, is the adjustment parameter. Finally, the system fuses the outputs of each modality in a weighted manner, combines with time series judgment, outputs the fall state and risk level, and automatically corrects the parameters according to the feedback, realizing environmental self-adaptation and long-term learning.

[0041] Compared to traditional fixed-weight, voting, or single-modal detection methods, this application's advantages lie in its "dynamic adaptability" and "multimodal robustness." In traditional schemes, weights are often preset constants, leading to a rapid decline in system performance when visual blur or noise increases. This application, however, introduces quality perception and uncertainty assessment, enabling weights to adjust in real-time according to the environment, automatically switching the dominant modality in nighttime, occluded, or noisy scenarios. Furthermore, employing interpretable Bayesian or D-S evidence fusion mechanisms can suppress unreliable data when different modal information conflicts, effectively reducing false positives and false negatives.

[0042] Time synchronization and preprocessing were performed on the data from each modality to extract radar velocity and micro-Doppler features, visual attitude and motion information, and audio impact and distress call features. Time synchronization was based on the system's unified UTC timestamp (accurate to 100ms), and linear interpolation was used to supplement missing data from different sampling frequency modes to ensure temporal consistency. Preprocessing included radar signal filtering, visual image denoising, and audio signal enhancement to improve the quality of the feature data.

[0043] For each modal computing environment quality index and confidence level, a credibility score is obtained. Radar modes are based on signal-to-noise ratio (SNR≥45dB). =0.9, 30dB≤SNR<45dB =0.6, SNR < 30dB =0.3); Visual modality is based on illumination sharpness (≥500 lux). =0.9, when 50 lux ≤ light intensity < 500 lux =0.6, <50 lux =0.3) and occlusion rate (occlusion rate ≤20%) No deduction is made when 20% < occlusion rate ≤ 50%. Deduct 0.2; deduct 0.5 if >50%; thermal imaging modality is based on temperature resolution (≤0.1℃). =0.8, >0.1℃ =0.5); Audio mode is based on signal amplitude (when the impact sound is ≥60dB). =0.8, 30dB≤amplitude<60dB =0.5, <30dB =0.2).

[0044] During the integration phase, credibility scoring is used as a basis. Environmental adaptability coefficient Calculate the modal weights Environmental adaptability coefficient The radar dynamically adjusts based on the home environment, especially in high-humidity bathroom scenarios (humidity > 80% RH). = 0.8, thermal image = 0.9, visual in low light bedroom scenario (light < 100 lux) = 0.5, radar = 0.9, visual in multi-occlusion living room scenario = 0.6, radar = 0.8, all modalities in normal scenario (light 200-500 lux, humidity 40-60% RH, no occlusion) = 1.0; the adjustment parameter is set to = 1.2, = 0.8, = 0.5, calibrated based on 1000 sets of historical fall case data.

[0045] An interpretable Bayesian or D-S evidence fusion mechanism is used to handle the conflict of modal information. For example, when radar = 0.8 (determines a fall), visual = 0.2 (determines normal), the mechanism suppresses the low-confidence data of the visual and reduces its weight to 0.1, giving priority to the radar result, effectively reducing the false positive and false negative rates. In combination with time series judgment, the fusion features of 10 consecutive frames (1 second in length) are analyzed, and the fusion confidence is calculated. When the fall confidence value exceeds the set threshold (for example, 0.8), and 3 consecutive frames meet the condition, suspected fall event recognition information is generated, accompanied by the weight of each modality, the confidence change curve and the key feature data.

[0046] S104, the standardized multi-dimensional feature vector and the suspected fall event recognition information are processed, and a real fall event determination result is generated through an AI agent voice interaction confirmation mechanism.

[0047] In one embodiment, the speech features, posture features in the standardized multi-dimensional feature vector and the confidence features in the suspected fall event identification information are subjected to keyword extraction, sentiment semantic analysis to generate speech response effectiveness features, posture abnormal correlation features, fall confidence matching features to form interactive confirmation basis information. The speech features, posture features in the standardized multi-dimensional feature vector and the confidence features in the suspected fall event identification information are subjected to directional processing, the core correlation features are quantified through keyword extraction and sentiment semantic analysis to generate three types of key feature information to form interactive confirmation basis information. Speech response effectiveness feature extraction: the speech response content is extracted from the speech features (such as MFCC coefficients, energy envelope, audio time sequence) in the standardized multi-dimensional feature vector, the keyword extraction model based on BERT (the pre-training corpus contains daily conversations of the elderly and fall help-seeking sentences) is used to extract keywords, and the effectiveness is determined by combining the sentiment semantic analysis model (the classification labels are "help-seeking", "negative", "meaningless", "no response"). The core features include keyword matching degree K (normalized to 0-1, 1 indicates complete matching of the target keyword) and emotional tendency value E (-1~1, 1 for strong help-seeking tendency, -1 for strong negative tendency). After the old person falls, the speech "come and help me, I have fallen" is emitted, the keywords "fall" and "help me" and other help-seeking keywords are extracted, K=0.92 is calculated, the sentiment semantic analysis determines the emergency help-seeking tendency, E=0.85, and the speech response effectiveness feature {K=0.92, E=0.85} is generated.

[0048] From the posture features (such as human body skeleton key point coordinates, center of gravity height change rate, limb motion angle) in the standardized multi-dimensional feature vector, posture correlation features related to the suspected fall event are extracted, the similarity between the real-time posture and the fall typical posture (such as body inclination angle > 45°, center of gravity height < 60 cm) is calculated to generate posture abnormal correlation degree R (normalized to 0-1, 1 indicates complete matching of the fall typical posture) and posture continuous abnormal duration T (unit: s). After the suspected fall event occurs, the real-time monitoring shows that the body inclination angle of the old person is 48° and the center of gravity height is 55 cm, the similarity with the fall typical posture is calculated to be R=0.88; the abnormal posture lasts for 15 s without recovery, T=15 s, and the posture abnormal correlation feature {R=0.88, T=15 s} is generated.

[0049] The fusion confidence (such as C=0.681 calculated in the previous process) in the suspected fall event identification information is matched with the determination results of the speech and posture features to generate a confidence matching coefficient M (normalized to 0-1, 1 indicates that the multiple features completely match the fall determination) and a confidence fluctuation value F (unit: 0-0.5, 0 indicates stable confidence). The calculation logic is M=0.4×K+0.3×R+0.3×C, F=| - Substituting K=0.92, R=0.88, and C=0.681, we calculate M=0.4×0.92+0.3×0.88+0.3×0.681≈0.368+0.264+0.204=0.836; Real-time confidence level =0.681, confidence level at the previous time step =0.675, F=0.006, generating fall confidence matching features {M=0.836, F=0.006}. Integrating the three types of features, basic information for interactive confirmation is formed {voice response validity: K=0.92, E=0.85; posture abnormality correlation: R=0.88, T=15s; fall confidence matching: M=0.836, F=0.006}.

[0050] The keyword matching degree and response duration features in the voice response validity characteristics are transformed into judgment rules and valid intervals are divided to generate negative statement judgment parameters, distress call statement recognition parameters, and no response judgment threshold parameters, forming voice interaction constraint information. (The time interval from the suspected fall incident to the voice response) is used for rule transformation and interval division to clarify the criteria for judging negative statements, distress calls, and no response, generating three types of constraint parameters to form voice interaction constraint information. Negative statement judgment parameter generation: By analyzing the keyword matching degree and sentiment value distribution of historical negative statements (such as "I didn't fall, I just sat down"), negative judgment intervals are divided, and negative keyword matching thresholds are generated. Threshold for negative emotional tendency When K < And E < When this condition is met, it is determined to be a negative statement. Based on statistics of 1000 historical negative statements, 95% of the negative statements satisfy K < 0.3 and E < -0.4, therefore, we set... =0.3、 =-0.4, generating parameters for negative statements. =0.3, =-0.4.

[0051] Similarly, by analyzing the feature distribution of historical distress calls (such as "Help!" or "I fell down"), distress call judgment intervals are defined, and distress call keyword matching thresholds are generated. Threshold for Emotional Tendency in Calling for Help Rule setting: When K≥ And E≥ When this occurs, it is determined to be a distress call. 95% of historical distress calls satisfy K≥0.6 and E≥0.5, therefore, we set... =0.6、 =0.5, generating distress call recognition parameters = 0.6, = 0.5.

[0052] Based on the average speech response speed of the elderly (e.g., responding within 3-10s after a suspected fall trigger is normal), the no response determination interval is divided, and the maximum response duration threshold is generated , no response retry interval When > and still no response after 2 retries, it is determined that there is no response. The average response duration of the elderly after a fall is 5s, so set = 10s (reserve sufficient reaction time), = 3s, generate no response determination threshold parameter = 10s, = 3s. Integrate the three types of parameters to form the voice interaction constraint information {negative determination: = 0.3, = -0.4; help identification: = 0.6, = 0.5; no response determination: = 10s, = 3s}.

[0053] The real fall event determination target is combined with the suspected fall event identification information to perform adaptation analysis and determination logic optimization, generate multi-feature collaborative determination target parameters and event authenticity verification standards, and form determination target optimization information. With the real fall event determination target ("accurately identify real falls and reduce false positives / negatives") as the core, combined with suspected fall event identification information (such as fusion confidence C, event occurrence scene), adaptation analysis is performed, and determination logic is optimized to generate two types of optimization information, forming determination target optimization information.

[0054] Analyze the relevance of voice, posture, and confidence features to real fall events, and optimize multi-feature collaborative determination logic through weight distribution to generate feature collaboration weight , (satisfying = 1), collaborative determination threshold Based on historical real fall cases, the recognition accuracy of each feature (voice accuracy 92%, posture accuracy 88%, confidence accuracy 85%) is calculated, and the entropy weight method is used to distribute the weight, with the formula being ( is the feature information entropy).

[0055] The voice information entropy = 0.35, the posture information entropy = 0.42, and the confidence information entropy = 0.48, then =(1-0.35) / (0.65+0.58+0.52)=0.65 / 1.75≈0.371, =0.58 / 1.75≈0.331, =0.52 / 1.75≈0.298; Based on the false positive rate control target (false positive rate < 5%), set... =0.7, generating multi-feature collaborative judgment target parameters {W=(0.371,0.331,0.298), =0.7)}.

[0056] To address the varying fall risks across different scenarios (such as bedrooms, bathrooms, and living rooms), we optimized the authenticity verification rules and generated scenario-adaptive verification coefficients. (0.8-1.2, 1.2 for high-risk scenarios), time series verification threshold (Unit: seconds, the shortest duration for continuously meeting the judgment conditions). The bathroom is a high-risk scenario (slippery floor, difficulty getting up after a fall), so the following settings are used... =1.2; To avoid instantaneous attitude interference, the coordination judgment condition must be met for 3 consecutive seconds, i.e. =3s, generate event authenticity verification criteria { =1.2, =3s}. Integrate the two types of information to form the target optimization information {cooperative decision parameters: W=(0.371,0.331,0.298)}. =0.7; Authenticity verification standard: =1.2, =3s.

[0057] Integrating basic interactive confirmation information, voice interaction constraint information, and judgment target optimization information, the system generates a record of the actual fall event judgment result and interactive confirmation process. The validity of the voice response is verified based on the voice interaction constraint information, and posture and confidence features are verified by combining the judgment target optimization information. Voice response K=0.92≥ =0.6, E=0.85≥ =0.5, deemed a valid distress call; Posture R=0.88, T=15s≥ =3s, determined to be a valid abnormal pose; confidence level C=0.681, combined with scene coefficients =1.2 then adjusted to =0.681×1.2≈0.817, which is considered a valid confidence level. Substitute this into the multi-feature collaborative decision formula to calculate the final collaborative confidence level. = ×K+ ×R+ × .,For example, = 0.371 * 0.92 + 0.331 * 0.88 + 0.298 * 0.817 = 0.341 + 0.291 + 0.243 = 0.875. The result is determined as follows: if ≥ = 0.7, it is determined as a "real fall event"; otherwise, it is determined as a "suspected fall exclusion". 0.875 > 0.7, it is determined as a real fall event.

[0058] Record the whole process information from suspected fall triggering to real determination, including triggering time (such as 2025-11-07 08:30:15), voice interaction content ("Come and help me, I fell down"), posture change data (inclination angle 48°→ 50°, center of gravity height 55 cm→ 53 cm), confidence change curve (0.675→ 0.681→ 0.817→ 0.875), constraint parameter application result (help determination effective, no response determination not triggered), optimization parameter adjustment record (scene coefficient 1.2, synergy weight application), form a complete interactive confirmation process record. The final output: real fall event determination result (event type: real fall; final synergy confidence: 0.875; risk level: high risk) and interactive confirmation process record (including the whole process data mentioned above).

[0059] S105, the real fall event determination result, the preset emergency contact information and the home device control authority are combined with the emergency linkage decision matrix, the priority notification strategy and the home scene self-defined control rule are fused, and the whole process emergency linkage execution scheme is generated.

[0060] In an embodiment, the event authenticity level, the emergency level in the real fall event determination result, and the priority ranking and contact method type in the preset emergency contact information are matched with the notification strategy, the response time limit is classified, the first contact person instant push feature, the second contact person delay linkage feature, and the medical emergency trigger condition feature are generated, and the contact person notification strategy information is formed. The notification method, response time limit, and medical emergency trigger condition of different contact persons are determined by matching and classifying the real fall event determination result and the preset emergency contact information, and the contact person notification strategy information is formed. The real fall event determination result includes the event authenticity level (such as “complete confirmation” and “highly suspected”) and the emergency level (such as “high risk”, “medium risk”, and “low risk”); the preset emergency contact information includes the priority ranking (first and second) and the contact method type (phone, short message, APP push, and WeChat notification). The matching logic is that the emergency level is one-to-one corresponding to the contact person priority, and the contact method type is selected according to the contact person preference and response efficiency. A real fall event is determined as “complete confirmation + high risk”, the first contact person (child) in the preset emergency contact person prefers APP push + phone notification, the second contact person (relative) prefers short message + WeChat notification, and the medical emergency contact method is the local 120 emergency phone, and the strategy matching is completed.

[0061] The first contact person (core relatives) adopts the “multi-channel instant push” strategy, and the core features include the notification channel combination, the response time limit requirement, and the information containing elements. The notification channel of the first contact person is “APP push + phone call + short message backup”, the response time limit requirement is “must respond within 5 minutes”, the push information contains “old person identity (Zhang XX), fall time (2025-11-07 09:15:30), fall location (bedroom), real-time state (suspected unable to get up), real-time picture link (encrypted)”, and the first contact person instant push feature {channel: APP + phone + short message, response time limit: ≤5 minutes, information elements: identity + time + location + state + picture link} is generated. The second contact person (secondary relatives / caregivers) adopts the “delay linkage” strategy, and the core features include the delay time, the notification channel, and the linkage trigger condition. The delay time is set to “3 minutes after the first contact person does not respond”, the notification channel is “short message + WeChat notification”, and the linkage trigger condition is “the first contact person does not click the APP to confirm the response or does not answer the phone within 5 minutes”, and the second contact person delay linkage feature {delay time: 3 minutes, channel: short message + WeChat, trigger condition: first contact person does not respond} is generated.

[0062] The "double trigger" strategy is adopted for medical emergency, and the core features include trigger time, trigger condition, and emergency information elements. The trigger time is set to "2 minutes after the secondary contact fails to respond", the trigger condition is "no effective response within 10 minutes of the primary and secondary contacts combined", and the emergency information includes "home address (XX City, XX District, XX Community, Building 3, Room 502), old person's health status (history of hypertension), fall status (high risk, suspected fracture), and contact phone number (primary contact mobile phone number)". The medical emergency trigger condition feature is generated {trigger time: 2 minutes after the secondary fails to respond, trigger condition: no effective response within 10 minutes, information elements: address + health status + status + contact phone number}. Integrating the three types of features, the contact notification strategy information is formed {primary contact: multi-channel immediate push, ≤5 minutes response; secondary contact: 3 minutes delay linkage, SMS + WeChat; medical emergency: 10 minutes no response trigger, including complete emergency information}.

[0063] The device linkage feasibility analysis, temporary permission allocation evaluation are performed on the door lock control permission, light adjustment permission, camera invocation permission in the home equipment control permission and the emergency mode configuration, rescue assistance setting in the home scene custom control rule, and the door lock automatic unlocking path feature, emergency light startup parameter, camera real-time sharing scheme feature are generated, and the home linkage basic information is formed. The adaptation analysis is performed on the home equipment control permission and the home scene custom control rule, the device linkage feasibility and the permission allocation rationality are evaluated, the three types of core features are generated, and the home linkage basic information is formed. Based on the door lock control permission and the emergency mode configuration, the unlocking feasibility (such as whether the door lock supports remote control and whether there is an emergency unlocking protocol) is analyzed, the unlocking path, permission validity period, and security verification mechanism are determined. The home door lock supports Wi-Fi remote control and emergency unlocking protocol, the unlocking path is "system temporarily obtains door lock control permission → pushes unlocking verification code to primary contact → contact inputs verification code to confirm unlocking (or system automatically unlocks, leaves trace record)", the permission validity period is "2 hours (from the unlocking time)", and the security verification mechanism is "automatically records operator, operation time, and operation log after unlocking", the door lock automatic unlocking path feature is generated {path: permission acquisition → verification code confirmation → unlocking, validity period: 2 hours, verification mechanism: operation trace + log record}.

[0064] Based on the light adjustment permission and rescue assistance settings, the type of starting light, brightness parameter, and starting time are determined. The emergency light type is "bedroom main light + corridor light + door indicator light", the brightness parameter is "bedroom main light 50% brightness (to avoid strong light stimulation), corridor light 100% brightness (to guide rescue), door indicator light flashing mode (frequency 2 times / sec)", and the starting time is "contact notification sent at the same time". The emergency light starting parameter is generated {type: bedroom + corridor + door indicator light, brightness: 50% / 100% / flashing, starting time: synchronized with contact notification}.

[0065] Based on the camera calling permission and rescue assistance settings, the sharing feasibility is analyzed (such as whether the camera supports encrypted sharing and whether there is a privacy protection mechanism), and the sharing object, sharing duration, and encryption method are determined. The camera supports encrypted link sharing, the sharing object is "primary contact + secondary contact + emergency personnel (authorized by primary contact)", the sharing duration is "4 hours (from the sharing time)", and the encryption method is "dynamic random password (sent to the sharing object through SMS) + link 15-minute automatic invalidation refresh". The camera real-time sharing scheme feature is generated {sharing object: authorized contact + emergency personnel, duration: 4 hours, encryption method: dynamic password + link refresh}. The three types of features are integrated to form the home linkage basic information {door lock: verification code unlocking, 2-hour validity period; light: multi-region hierarchical starting, synchronous notification sending; camera: encrypted authorization sharing, 4-hour duration}.

[0066] In combination with the emergency degree of the fall event and the rescue scene demand, the contact person notification strategy information and the home linkage basic information are executed in sequence, the permission conflict is resolved, the emergency linkage execution sequence characteristics and the multi-measure coordination coefficient are generated, and the linkage optimization information is formed. In combination with the emergency degree of the fall event (high risk) and the rescue scene demand (indoor rescue, which requires rapid guidance and safe entry), the contact person notification strategy information and the home linkage basic information are processed in the process of planning and conflict, and the linkage optimization information is generated. According to the principle of “rescue efficiency first, safety controllable”, the execution sequence is planned, and the starting time and dependent relationship of each link are determined. The execution sequence is “1. Synchronous start of emergency light (0 seconds) → 2. Send multi-channel notification to primary contact person (0 seconds) → 3. Generate camera encryption sharing link (5 seconds) → 4. Push door lock unlocking code after primary contact person responds (≤5 minutes) → 5. Send notification to secondary contact person after primary contact person does not respond for 3 minutes (8 minutes) → 6. Trigger medical emergency after secondary contact person does not respond for 2 minutes (10 minutes) → 7. Emergency personnel contact primary contact person to obtain camera sharing password and door lock unlocking permission (10-15 minutes)”, and the emergency linkage execution sequence characteristics {sequence: light → primary notification → camera link → primary unlocking → secondary notification → medical emergency → emergency personnel authorization, dependent relationship: unlocking permission depends on contact person response, emergency trigger depends on two-level non-response} are generated.

[0067] The coordination degree between each measure is quantified (0-1 interval, 1 for complete coordination), which includes “notification and light coordination coefficient”, “unlocking and camera coordination coefficient” and “emergency and home linkage coordination coefficient”. Notification and light coordination coefficient = 0.9 (synchronous start, no delay), unlocking and camera coordination coefficient = 0.85 (unlocking automatically synchronizes camera link, only needs 1 authorization), emergency and home linkage coordination coefficient = 0.8 (emergency personnel can obtain unlocking and camera permission through contact person authorization, without additional operation), and the multi-measure coordination coefficient {notification-light: 0.9, unlocking-camera: 0.85, emergency-home: 0.8} is generated. Integrate the two types of characteristics to form linkage optimization information {execution sequence: 0 seconds (light + primary notification) → 5 seconds (camera link) → ≤5 minutes (primary unlocking) → 8 minutes (secondary notification) → 10 minutes (medical emergency), coordination coefficient: 0.8-0.9}.

[0068] Integrate contact notification strategy information, home linkage basic information, and linkage optimization information to generate an emergency linkage execution scheme covering the whole process of notification, permission, and linkage. Integrate contact notification strategy information, home linkage basic information, and linkage optimization information to form an emergency linkage scheme covering the whole process of "notification, permission, and linkage" and directly executable, with clear operation details, responsible subjects, and safety mechanisms. Trigger condition: automatically start this scheme after the system determines a "real fall event (complete confirmation + high risk)". The execution steps (including time nodes) are as follows: 0 seconds: start emergency lights (bedroom 50% brightness, corridor 100% brightness, door indicator light flashing); send APP push + phone call + SMS to primary contact (children), with content including the old person's identity, fall time / location / status, and encrypted camera link. 5 seconds: generate camera encryption sharing link, synchronize to local database, and link is automatically refreshed every 15 minutes. ≤5 minutes: if the primary contact responds (clicks APP confirmation or answers the phone), the system pushes the door lock unlocking code to the contact, and the contact inputs the code to automatically unlock the door (unlocking log is synchronized and recorded); at the same time, the contact is authorized to view the real-time camera screen.

[0069] 8 minutes: if the primary contact does not respond within 5 minutes, the system sends an SMS + WeChat notification to the secondary contact (relative), with content including fall basic information and explanation of the primary contact's non-response. 10 minutes: if the secondary contact does not respond within 3 minutes, the system automatically dials 120 emergency call, broadcasts the family address, old person's health status, fall status, and primary contact's mobile number; at the same time, the system pushes a temporary authorization link to the emergency personnel (which needs to be confirmed by the primary contact or automatically authorized by the system, with a validity period of 2 hours). 10-15 minutes: before the emergency personnel arrives, the system continuously refreshes the camera link and keeps the emergency lights on; if the contact responds in advance, the APP can be used to cancel the emergency call or adjust the home device status.

[0070] All home device control permissions (door lock, camera) are temporary authorization, with a maximum validity period of 4 hours, and automatically expire after the time limit; operation logs are recorded throughout the process and stored in a local encrypted database. The camera screen is encrypted using AES-256, and only authorized objects can view it through a dynamic password, while unauthorized personnel cannot access it. The primary contact can use the APP to urgently stop the linkage process (e.g., false fall judgment), and the system will immediately turn off the lights, revoke the door lock permission, and terminate the emergency call. The system automatically executes the core process, the primary contact leads the response and authorization, the secondary contact assists the response, and the emergency personnel is responsible for on-site rescue. The execution results of each link (such as whether the notification is delivered, whether the door lock is unlocked, and whether the emergency call is connected) are fed back to the system in real time and pushed to the primary contact's APP, ensuring traceability throughout the process. The final output is as follows: a whole-process emergency linkage execution scheme (including trigger conditions, step-by-step execution details, safety mechanisms, and feedback mechanisms) that can be directly called and executed by the AI intelligent agent decision-making module.

[0071] S106, based on the edge computing architecture, the cloud-side cooperative optimization engine combines the multi-modal adaptive adjustment strategy to process the real fall event judgment result, the whole process emergency linkage execution scheme and the abnormal feature recognition information, and generates the old-age safety monitoring dynamic protection and continuous optimization result.

[0072] In an implementation, the event characteristics in the real fall event judgment result, the execution effect data in the whole process emergency linkage execution scheme, and the local processing performance parameters and privacy protection encryption indicators of the edge computing architecture are subjected to adaptability analysis to generate a protection strategy adaptation quantitative factor, wherein the protection strategy adaptation quantitative factor represents the degree of fit of the emergency scheme with the edge end running characteristics and privacy security requirements. By subjecting the real fall event characteristics and the emergency linkage execution effect data to multi-dimensional adaptability analysis with the local processing performance and privacy protection encryption indicators of the edge computing architecture, a protection strategy adaptation quantitative factor is generated, which directly reflects the degree of fit of the emergency scheme with the edge end running characteristics and privacy security requirements. The event characteristics of the real fall event judgment result include the total event processing time (the total time from suspected fall recognition to emergency scheme activation), and the total multi-modal data transmission amount (the total amount of feature data and control instruction data generated during event judgment and emergency execution); the execution effect data of the whole process emergency linkage execution scheme includes device control response delay (such as the time delay from edge end to device execution of instructions such as door unlocking, light starting, etc.), and emergency operation success rate (such as contact person notification delivery rate, home device control success rate); the local processing performance parameters of the edge computing architecture include the maximum computing power of the edge device, the local data storage capacity, and the upper limit of single instruction processing delay; the privacy protection encryption indicators include data encryption processing time, encrypted data redundancy ratio, and privacy data leakage risk level (1-5 levels, with level 1 being the lowest risk).

[0073] The adaptability analysis is carried out from five dimensions of "processing time limit adaptation", "storage capacity adaptation", "operation success rate adaptation", "encryption performance adaptation", and "privacy security adaptation". The total event processing time of a certain real fall event is 12 seconds, the upper limit of single instruction processing delay of the edge device is 2 seconds, the device control response delay is 1.5 seconds, which meets the time limit requirement; the total multi-modal data transmission amount is 80MB, the local storage capacity of the edge device is 512GB, and the storage pressure is extremely small; the emergency operation success rate is 98% (only 1 time of secondary contact person SMS notification delay); the data encryption time is 0.8 seconds, the encrypted data redundancy ratio is 5%, and the transmission and storage burden is not significantly increased; the privacy data leakage risk level is level 1, which meets the security requirement. After comprehensive analysis, the protection strategy adaptation quantitative factor (in the range of 0-1, the larger the value, the higher the fit degree) is 0.68, indicating that the fit degree of the emergency scheme with the edge end characteristics and privacy requirements is at a medium to high level, and only the total event processing time needs to be optimized.

[0074] The sensor signal abnormal pattern in the abnormal feature recognition information, the multi-modal fusion deviation data, and the historical optimization case library of the cloud collaborative optimization engine, the model parameter adjustment record are compared and analyzed, and the model iterative optimization quantitative factor is generated. The sensor signal abnormal pattern in the abnormal feature recognition information, the multi-modal fusion deviation data, and the historical optimization case library of the cloud collaborative optimization engine, the model parameter adjustment record are compared and analyzed, and the model iterative optimization quantitative factor is generated. The internal correlation between data anomalies and model parameters is mined, and the model iterative optimization quantitative factor is generated. The adjustment direction and optimization space of the model parameter are clear. The sensor signal abnormal pattern in the abnormal feature recognition information includes single sensor signal quality abnormality (such as sudden drop of millimeter wave radar signal-to-noise ratio, missing of posture feature caused by camera image occlusion), multi-sensor signal synchronization deviation (such as radar and camera data collection time is not synchronized); multi-modal fusion deviation data includes fusion judgment result and real event deviation rate (such as false positive rate, false negative rate), and recognition deviation caused by unreasonable allocation of different modal feature weights; the historical optimization case library of the cloud collaborative optimization engine contains a large number of complete cases of "sensor abnormality-model adjustment-effect improvement" (such as weight adjustment case in low signal-to-noise ratio radar signal scene, feature compensation case in image occlusion scene); the model parameter adjustment record includes the adjustment trajectory and corresponding effect change data of the modal weight, fusion confidence threshold, feature extraction parameter of the multi-modal fusion model in the past 6 months.

[0075] The trend comparison focuses on "abnormal pattern similarity", "deviation reason matching degree" and "parameter adjustment correlation". The current millimeter wave radar signal-to-noise ratio drops from normal 45dB to 28dB, which causes the fusion recognition false positive rate to rise by 8%. The similarity of this sensor abnormal pattern to the 32 "low signal-to-noise ratio radar signal" cases in the cloud historical case library is 90%; the multi-modal fusion deviation is mainly caused by the failure to adjust the radar feature weight in time, and the matching degree with the "unreasonable weight allocation leading to deviation" in the historical cases is 85%; according to the model parameter adjustment record, in a similar scene in the past, the radar feature weight was increased from 0.377 to 0.45, and the false positive rate was reduced by an average of 5%. After comprehensive analysis, the model iterative optimization quantitative factor (0-1 interval, the larger the value, the larger the optimization space) is 0.72, indicating that the current abnormal data is highly matched with the historical optimization cases, and the expected optimization effect is significant after adjusting the model parameters.

[0076] The protection strategy adaptation quantization factor, the model iterative optimization quantization factor, the dynamic weight distribution logic in the multi-modal adaptive adjustment strategy, and the scene adaptive switching rule are combined to fuse the protection response time, the identification accuracy, and the strategy adjustment effectiveness in different monitoring scenes to generate the protection-optimization-adaptation associated quantization features. The protection strategy adaptation quantization factor, the model iterative optimization quantization factor, the dynamic weight distribution logic in the multi-modal adaptive adjustment strategy, and the scene adaptive switching rule are combined to comprehensively fuse the protection response time, the identification accuracy, and the strategy adjustment effectiveness in different home scenes to generate the protection-optimization-adaptation associated quantization features, and the collaborative relationship among the three is established. Three common home scenes of the elderly (bedroom: low light + static scene, bathroom: high humidity + dynamic scene, living room: normal light + multi-obstruction scene) are selected, and the core data of each scene includes: protection response time (time consumption from event triggering to protection measure starting), identification accuracy (correct identification rate of fall events in the scene), and strategy adjustment effectiveness (improvement proportion of identification accuracy / response time after scene switching or parameter adjustment). The dynamic weight distribution logic of the multi-modal adaptive adjustment strategy is that “for every increase of 0.1 in the environmental interference factor, the corresponding modal weight decreases by 0.05”; and the scene adaptive switching rule is that “the radar + audio modal weight is increased by default in the bathroom scene, and the multi-modal collaborative filtering algorithm is enabled by default in the living room scene”.

[0077] In the fusion processing, the protection response time, the identification accuracy, and the strategy adjustment effectiveness of each scene are weighted and calculated according to the scene weight (bedroom 0.4, bathroom 0.3, living room 0.3), and then are collaboratively corrected in combination with the protection strategy adaptation quantization factor and the model iterative optimization quantization factor. The protection response time of the bedroom scene is 10 seconds, the identification accuracy is 92%, and the strategy adjustment effectiveness is 5%; the protection response time of the bathroom scene is 8 seconds, the identification accuracy is 88%, and the strategy adjustment effectiveness is 4%; the protection response time of the living room scene is 11 seconds, the identification accuracy is 85%, and the strategy adjustment effectiveness is 6%; the weighted calculation of the scene comprehensive score is 0.87; in combination with the protection strategy adaptation quantization factor 0.68 and the model iterative optimization quantization factor 0.72, the protection-optimization-adaptation associated quantization features (including the overall feature value and the scene sub-feature value) are finally generated, the overall feature value is 0.75, and the bathroom scene sub-feature value 0.71 is the short board, and the sensor signal stability and the strategy adaptability in the high humidity environment are focused on.

[0078] Based on an edge computing and cloud-based collaborative architecture, the quantitative characteristics of the protection-optimization-adaptation correlation are analyzed and processed to generate dynamic protection and continuous optimization results for elderly safety monitoring, including dynamic protection strategy update schemes, multimodal fusion model iteration parameters, and robustness assessments for adaptation to different home scenarios. Deep analysis of the quantitative characteristics of the protection-optimization-adaptation correlation is performed based on this edge computing and cloud-based collaborative architecture. Combining the advantages of local real-time processing at the edge and the big data optimization capabilities of the cloud, complete optimization results are generated, including dynamic protection strategy updates, multimodal fusion model iterations, and scenario adaptation robustness assessments, enabling continuous self-improvement of the system. The dynamic protection strategy is optimized to address the shortcomings reflected by the quantitative factors of protection strategy adaptation and the correlation characteristics of various scenarios. The event handling process has been optimized by executing multimodal feature preprocessing and emergency command generation in parallel, reducing the total event handling time (from 12 seconds to within 8 seconds). For high-humidity bathroom scenarios, a new "sensor signal enhancement algorithm" (such as radar signal denoising and camera image defogging) has been added to improve signal quality in high-interference environments. The privacy encryption process has been optimized by adopting a lightweight encryption algorithm, reducing data encryption time from 0.8 seconds to 0.3 seconds without compromising security levels. An "emergency strategy emergency termination function" has been added, allowing primary contacts to quickly terminate accidentally triggered emergency operations (such as misjudging a fall and calling for emergency medical assistance, or unlocking door locks) via the app.

[0079] Based on model iteration optimization of quantization factors and cloud-based historical case association results, the core parameters of the multimodal fusion model were updated. The modal weight baseline value was adjusted, increasing the radar feature base weight from 0.377 to 0.42 to enhance its influence in low-light and occluded scenarios. Fusion confidence thresholds were set for different scenarios: 0.62 for bedroom, 0.60 for bathroom, and 0.65 for living room, adapting to the recognition difficulty of different scenarios. Abnormal data filtering parameters were optimized, adjusting the feature filtering threshold for radar signal-to-noise ratios below 30dB from 0.2 to 0.15 to reduce interference from low-quality signals on the fusion results. A new "sensor signal anomaly compensation parameter" was added, automatically increasing the feature weights and recognition weights of other modalities when a single sensor signal is abnormal.

[0080] The three-dimensional evaluation model of "response time-identification accuracy-strategy adaptability" is used to comprehensively evaluate the adaptation robustness of three core scenes: bedroom, bathroom, and living room. The evaluation result of the bedroom scene is "good" (response time 10 seconds + identification accuracy 92% + strategy adaptability 0.78), and only the response time needs to be optimized. The evaluation result of the bathroom scene is "medium" (response time 8 seconds + identification accuracy 88% + strategy adaptability 0.71), and the identification accuracy and strategy adaptability need to be improved. The evaluation result of the living room scene is "good" (response time 11 seconds + identification accuracy 85% + strategy adaptability 0.76), and the identification accuracy in the occlusion scene needs to be optimized. The overall system robustness evaluation result is "medium to high", and after optimization, the target is to make the three types of scenes reach the "good" level, and the overall robustness is improved to more than 0.8. Finally, the dynamic protection and continuous optimization results of the elderly safety monitoring are generated, including the dynamic protection strategy update scheme (4 specific optimization measures), the multi-modal fusion model iteration parameters (4 groups of core parameter update values), and the different home scene adaptation robustness evaluation report (including the scene level, short board and optimization direction). Through the edge-cloud collaborative mechanism, the optimization results are synchronized to the edge device and cloud model library in real time, realizing the dynamic iteration and upgrading of the system.

[0081] As shown in Figure 2 A safety monitoring device for the elderly based on multi-modal sensor fusion includes a construction module 201 for obtaining multi-modal monitoring data, environmental state data, and device operation data; a processing module 202 for data preprocessing of the multi-modal monitoring data, introducing a synchronization correction algorithm and a multi-modal feature extraction algorithm, combining edge data encryption and local caching mechanism, and generating a standardized multi-dimensional feature vector; processing environmental state data and device operation data based on a dynamic weight multi-modal fusion engine, through an adaptive weight distributor and scene correlation analysis, combining fusion confidence calculation and abnormal data filtering, and generating suspected fall event identification information; processing the standardized multi-dimensional feature vector and the suspected fall event identification information, and generating a real fall event judgment result through an AI intelligent agent voice interaction confirmation mechanism; processing the real fall event judgment result, the pre-set emergency contact information, and the home device control authority in combination with an emergency linkage decision matrix, fusing priority notification strategies and home scene customized control rules, and generating a full-process emergency linkage execution scheme; processing the real fall event judgment result, the full-process emergency linkage execution scheme, and the abnormal feature identification information based on the edge computing architecture and the cloud collaborative optimization engine in combination with the multi-modal adaptive adjustment strategy, and generating the dynamic protection and continuous optimization results of the elderly safety monitoring.

[0082] A computing device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to perform any one of the multi-modal sensor fusion based safety monitoring methods for the elderly.

[0083] The methods and / or embodiments in this application can be implemented as a computer program product, i.e., a computer program tangibly embodied in a machine-readable storage medium. The machine-readable storage medium can include one or more types of computer-readable storage media. For example, the machine-readable storage medium can include a computer-readable storage medium that is tangible, non-transitory, or a combination thereof. The machine-readable storage medium can include a program that, when executed by a processor, causes the processor to carry out the methods described herein.

[0084] It should be noted that the computer-readable medium in this application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In this application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0085] It is apparent that a person skilled in the art would not limit the present application to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Thus, the embodiments should be considered in all respects as illustrative and not restrictive, and the scope of the present application should be defined by the appended claims rather than the above description, and it is intended to cover all changes falling within the meaning and range of equivalents of the claims.

Claims

1. A method for elderly safety monitoring based on multimodal sensor fusion, characterized in that, include: Acquire multimodal monitoring data, environmental status data, and equipment operation data; Data preprocessing is performed on multimodal monitoring data, and a synchronous correction algorithm and a multimodal feature extraction algorithm are introduced. Combined with edge data encryption and local caching mechanisms, standardized multidimensional feature vectors are generated. Based on the dynamic weighted multimodal fusion engine, environmental status data and equipment operation data are processed. According to the hardware computing power or application scenario, the corresponding multimodal fusion method is selected. Through the adaptive weight allocator and scene correlation analysis, combined with fusion confidence calculation and abnormal data filtering, suspected fall event identification information is generated. The standardized multidimensional feature vector and suspected fall event identification information are processed, and the actual fall event judgment result is generated through the AI ​​intelligent agent voice interaction confirmation mechanism. The system combines the results of real fall incidents, preset emergency contact information, and home device control permissions with an emergency response decision matrix, integrates priority notification strategies and custom control rules for home scenarios, and generates a full-process emergency response execution plan. Based on an edge computing architecture and a cloud-based collaborative optimization engine, combined with a multimodal adaptive adjustment strategy, the system processes the results of real fall event assessments, full-process emergency response plans, and abnormal feature identification information to generate dynamic protection and continuous optimization results for elderly safety monitoring.

2. The method for elderly safety monitoring based on multimodal sensor fusion according to claim 1, characterized in that, Data preprocessing is performed on multimodal monitoring data, introducing synchronization correction algorithms and multimodal feature extraction algorithms. Combined with edge data encryption and local caching mechanisms, standardized multidimensional feature vectors are generated, including: The multimodal monitoring data is classified and preprocessed for analysis to generate the matching results of data types and processing algorithms. The multimodal monitoring data includes millimeter-wave radar spatial point cloud and motion trajectory data, infrared thermal imaging human body thermal distribution data, microphone ambient sound and voice data, and camera ambient image and human posture data. The adaptation and association results are processed in a targeted manner to generate a list of feature extraction variables, including radar signal filtering and point cloud clustering variables, infrared thermal imaging human body thermal center extraction variables, visual posture skeleton key point extraction variables, and audio MFCC feature and energy envelope analysis variables. The adaptation and association results, the list of feature extraction variables, and the original multimodal monitoring data are synchronously corrected and encrypted. Combined with the local caching mechanism at the edge, a standardized multidimensional feature vector is generated.

3. The method for elderly safety monitoring based on multimodal sensor fusion according to claim 1, characterized in that, Based on a dynamic weighted multimodal fusion engine, environmental status data and equipment operation data are processed. Depending on hardware computing power or application scenario, a corresponding multimodal fusion method is selected. Through an adaptive weight allocator and scene correlation analysis, combined with fusion confidence calculation and anomaly data filtering, suspected fall event identification information is generated, including: Environmental status data and equipment operation data are classified, extracted, and feature matched to generate multimodal body motion feature information, low light / occlusion environment interference factor information, and individual body size / body motion frequency adaptation factor information. The specific dimensions of the multimodal body motion feature information are determined according to the selected fusion method. The environmental interference factor information includes low light intensity factor and occlusion degree factor, and the individual adaptation factor information includes body size adaptation factor and body motion frequency adaptation factor. Multimodal monitoring data is synchronized and preprocessed in time. Based on the unified UTC timestamp of the system, linear interpolation is used to supplement the missing data of different sampling frequency modes to ensure time sequence consistency. The preprocessing includes Gaussian filtering of radar signals, denoising of visual images, adaptive noise suppression of audio signals, and equalization of infrared thermal image histograms to improve the quality of feature data. Based on hardware computing power or application scenario, select the corresponding multimodal fusion method to generate the initial identification result of suspected fall event; The initial identification results of suspected fall incidents are filtered for abnormal data, and corresponding invalid data is removed. The initial identification results of the filtered suspected fall events are validated according to rules. The validation includes data integrity and equipment operating status. After the validation is passed, suspected fall event identification information is generated, along with modal weights, confidence curves and key feature data.

4. The method for elderly safety monitoring based on multimodal sensor fusion according to claim 1, characterized in that, The standardized multidimensional feature vector and suspected fall event identification information are processed, and the actual fall event determination result is generated through an AI intelligent agent voice interaction confirmation mechanism, including: Keyword extraction and sentiment semantic analysis are performed on the speech features, posture features and confidence features in the suspected fall event identification information in the standardized multidimensional feature vector to generate speech response validity features, posture abnormality correlation features and fall confidence matching features, forming the basic information for interactive confirmation; The keyword matching degree and response duration in the voice response validity features are transformed into judgment rules and the effective interval is divided to generate negative statement judgment parameters, emergency call statement recognition parameters, and no response judgment threshold parameters, thus forming voice interaction constraint information; For the determination of real fall events, the system combines the identification information of suspected fall events to perform adaptation analysis and optimize the determination logic, generating multi-feature collaborative determination target parameters and event authenticity verification standards, thus forming optimized determination target information; Integrate basic information for interactive confirmation, constraint information for voice interaction, and optimization information for judgment targets to generate real fall event judgment results and interactive confirmation process records.

5. The method for elderly safety monitoring based on multimodal sensor fusion according to claim 1, characterized in that, The system processes the actual fall incident assessment results, preset emergency contact information, and home appliance control permissions in conjunction with an emergency response decision matrix. It integrates priority notification strategies with custom control rules for home scenarios to generate a full-process emergency response execution plan, including: The system matches the event authenticity level and urgency level in the real fall event determination results with the priority ranking and contact method type in the preset emergency contact information, performs notification strategy and response time classification, generates the instant push feature of the first-level contact, the delayed linkage feature of the second-level contact, and the medical emergency triggering condition feature, and forms the contact notification strategy information. Feasibility analysis and temporary permission allocation assessment are conducted on the device linkage feasibility of door lock control permissions, light adjustment permissions, and camera access permissions in the home device control permissions and emergency mode configuration and rescue assistance settings in the home scene custom control rules. This generates characteristics of automatic door lock unlocking path, emergency light activation parameters, and camera real-time sharing scheme characteristics, forming basic information for home linkage. Based on the urgency of fall incidents and the needs of rescue scenarios, the execution sequence of contact notification strategy information and home linkage basic information is planned and permission conflict resolution is handled. Emergency linkage execution sequence characteristics and multi-measure coordination coefficients are generated to form linkage optimization information. Integrate contact notification strategy information, home linkage basic information, and linkage optimization information to generate an emergency linkage execution plan covering the entire process of notification, permissions, and linkage.

6. The method for elderly safety monitoring based on multimodal sensor fusion according to claim 5, characterized in that, Based on an edge computing architecture and a cloud-based collaborative optimization engine, combined with a multimodal adaptive adjustment strategy, the system processes real fall event assessment results, full-process emergency response plans, and abnormal feature identification information to generate dynamic protection and continuous optimization results for elderly safety monitoring, including: The event characteristics in the judgment results of real fall events and the execution effect data in the full-process emergency linkage execution plan are analyzed for their compatibility with the local processing performance parameters and privacy protection encryption indicators of the edge computing architecture. A protection strategy adaptation quantification factor is generated, which represents the degree of fit between the emergency plan and the edge end's operating characteristics and privacy and security requirements. The abnormal sensor signal patterns and multimodal fusion deviation data in the abnormal feature identification information are compared and correlated with the historical optimization case library and model parameter adjustment records of the cloud-based collaborative optimization engine to generate model iterative optimization quantification factors. Based on the protection strategy adaptation quantization factor and the model iteration optimization quantization factor, combined with the dynamic weight allocation logic and scene adaptive switching rules in the multimodal adaptive adjustment strategy, the protection response timeliness, identification accuracy and strategy adjustment effectiveness under different monitoring scenarios are fused to generate protection-optimization-adaptation related quantization features. Based on the edge computing and cloud-based collaborative architecture, the quantitative characteristics of the protection-optimization-adaptation relationship are analyzed and processed to generate dynamic protection and continuous optimization results for elderly safety monitoring, including dynamic protection strategy update schemes, multimodal fusion model iteration parameters, and robustness assessment of adaptation to different home scenarios.

7. A safety monitoring device for the elderly based on multimodal sensor fusion, characterized in that, The device includes: Modules are built to acquire multimodal monitoring data, environmental status data, and equipment operation data; The processing module is used to preprocess multimodal monitoring data, introducing synchronous correction algorithms and multimodal feature extraction algorithms, combined with edge data encryption and local caching mechanisms, to generate standardized multidimensional feature vectors. Based on a dynamic weighted multimodal fusion engine, it processes environmental status data and equipment operation data, generating suspected fall event identification information through an adaptive weight allocator and scene correlation analysis, combined with fusion confidence calculation and abnormal data filtering. It then processes the standardized multidimensional feature vectors and suspected fall event identification information, generating a real fall event determination result through an AI intelligent agent voice interaction confirmation mechanism. Finally, it processes the real fall event determination result, preset emergency contact information, and home device control permissions in conjunction with an emergency linkage decision matrix, integrating priority notification strategies and custom control rules for home scenes to generate a full-process emergency linkage execution plan. Finally, based on an edge computing architecture and a cloud-based collaborative optimization engine, combined with a multimodal adaptive adjustment strategy, it processes the real fall event determination result, the full-process emergency linkage execution plan, and abnormal feature identification information to generate dynamic protection and continuous optimization results for elderly safety monitoring.

8. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the elderly safety monitoring method based on multimodal sensor fusion as described in any one of claims 1 to 6 by executing the executable instructions.

9. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the elderly safety monitoring method based on multimodal sensor fusion as described in any one of claims 1 to 6.

Citation Information

Cited By

  • LED intelligent control system for scene recognition

    CN121842914A

  • Multi-modal monitoring method, system and equipment based on hardware collaboration and storage medium

    CN121983272A

  • False fall filtering method and system based on multi-modal spatial-temporal feature fusion

    CN122024329A

  • A false fall filtering method and system based on multi-modal spatio-temporal feature fusion

    CN122024329B

  • Non-contact elderly health monitoring system and method based on multi-source information

    CN122177405A