Human body tumble radar perception correction method and system in combination with multi-modal data

By fusing radar echo and infrared thermal imaging data, a multimodal perception system was constructed and a fall detection deviation model was built. This solved the problems of time judgment deviation and attitude recognition misjudgment in complex environments, and achieved more accurate fall detection.

CN120802245APending Publication Date: 2025-10-17JIANGMEN YINXING ROBOTICS LTD

Patent Information

Application Number
CN202511294826.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing radar sensing technology is susceptible to interference from multiple targets and obstruction in complex environments, leading to time-judgment deviations and posture recognition errors in the detection of human falls, which affects the reliability of monitoring and the timeliness of rescue response.

Method used

By combining radar echo data and infrared thermal imaging data, a standardized multimodal dataset is constructed through spatiotemporal alignment preprocessing. A fall perception bias model is built using machine learning, and the radar perception bias value is output to correct the original perception results.

Benefits of technology

It improves the accuracy and reliability of human fall detection, ensures accurate time determination and posture recognition in complex environments, and provides technical support for timely rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802245A_ABST
    Figure CN120802245A_ABST
Patent Text Reader

Abstract

The invention discloses a human body tumble radar perception correction method and system combined with multi-modal data, and relates to the technical field of medical health monitoring. The method comprises the following steps: acquiring radar echo data and infrared thermal imaging data of a target monitoring area; preprocessing a unified timestamp and a space coordinate system through space-time alignment to obtain a standardized multi-modal data set; constructing a tumble perception deviation model based on the historical tumble sample and the real label; and finally, inputting the standardized data into the model to obtain a radar sensing deviation value, correcting an original radar result, and obtaining a precise fall sensing result. The system comprises a data acquisition module, a data preprocessing module, a model construction module and a perception correction module. According to the invention, the problem of time deviation or attitude misjudgment of a tumble sensing result caused by the fact that single radar data is easily influenced by factors such as multi-target interference and shielding in a complex environment in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical health monitoring, in particular to a human fall radar perception correction method and system combined with multi-modal data. BACKGROUND

[0002] In the medical health and safety monitoring scene, accurate perception of human fall events is of great significance for timely rescue and reduction of fall injury risk. Its technical application covers various scenes such as elderly care institutions, home care scenes, hospital wards, and other scenes that need to monitor the real-time safety state of the human body. In the prior art, radar perception technology has become one of the mainstream technologies for human fall monitoring due to its unique advantages such as non-contact monitoring, not limited by light conditions, and the ability to penetrate clothing.

[0003] However, in a complex environment, single radar data is easily affected by multi-target interference, object occlusion, environmental clutter and other factors, resulting in time determination deviation and posture recognition misjudgment in fall perception results, which seriously affects the reliability of monitoring and the timeliness of rescue response. SUMMARY

[0004] The present application provides a human fall radar perception correction method and system combined with multi-modal data to solve the technical problems that single radar data in the prior art is easily affected by multi-target interference, occlusion and other factors in a complex environment, resulting in time deviation or posture misjudgment in fall perception results.

[0005] The technical solution of the present application to solve the above technical problems is as follows: In a first aspect, the present application provides a human fall radar perception correction method combined with multi-modal data, comprising: Obtaining radar echo data and infrared thermal imaging data of a target monitoring area, wherein the radar echo data includes target motion trajectory and distance information, and the infrared thermal imaging data includes human body contour and temperature distribution; Performing spatio-temporal alignment preprocessing on the radar echo data and infrared thermal imaging data to unify data timestamps and spatial coordinate systems, and obtaining a standardized multi-modal data set; Based on the standardized multi-modal sample data of historical fall events and the corresponding true fall state label, a fall perception deviation model is constructed; Inputting the standardized multi-modal data set into the fall perception deviation model to output a radar perception deviation value, and based on the radar perception deviation value, the original radar fall perception result is corrected and adjusted to obtain an accurate fall perception result.

[0006] In a second aspect, the present application provides a human fall radar perception correction system combined with multi-modal data, comprising: A data acquisition module is used to acquire radar echo data and infrared thermal imaging data of the target monitoring area, wherein the radar echo data includes target motion trajectory and distance information, and the infrared thermal imaging data includes human body contour and temperature distribution; A data preprocessing module is used to perform spatiotemporal alignment preprocessing on the radar echo data and infrared thermal imaging data, unify the data timestamp and spatial coordinate system, and obtain a standardized multimodal data set; A model building module is used to build a fall perception bias model based on standardized multimodal sample data of historical fall events and the corresponding real fall status labels; The perception correction module is used to input the standardized multimodal dataset into the fall perception bias model, output a radar perception bias value, and correct and adjust the original radar fall perception result based on the radar perception bias value to obtain an accurate fall perception result.

[0007] The beneficial effects of the present invention are: Compared with the existing technology, this application first constructs a multimodal perception system by fusing radar echo data with infrared thermal imaging data, and uses radar motion trajectory, distance information, infrared human body contour, and temperature distribution to achieve data complementarity, breaking through the perception limitations of a single radar in a complex environment; secondly, through spatiotemporal alignment preprocessing, the data timestamp and spatial coordinate system are unified to ensure the synchronous matching of data in spatiotemporal and temporal space, avoid feature misassociation, and improve the effectiveness of fusion; thirdly, based on historical fall samples and real fall status labels, a fall perception deviation model is constructed through machine learning, the deviation law is quantified, and the correction is data-driven and scientific; finally, the original radar results are targetedly corrected through time deviation and posture deviation identification, effectively solving the problems of time judgment deviation and posture misjudgment, significantly improving the accuracy and reliability of fall perception, and providing reliable technical guarantee for timely rescue in scenarios such as elderly care and family care. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic flow chart of the method for correcting human fall radar perception using multimodal data provided by the present invention; Figure 2 This is a structural diagram of the human fall radar perception and correction system combined with multimodal data provided by the present invention.

[0009] In the accompanying drawings, the components represented by the reference numerals are as follows: Data acquisition module 11, data preprocessing module 12, model building module 13, perception correction module 14. DETAILED DESCRIPTION

[0010] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below, obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0011] In the description of the present application, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0012] In the description of the present application, the term "for example" is used to indicate "as an example, illustration or explanation". Any embodiment described as "for example" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. The following description is given in order to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that those skilled in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the shown embodiments, but is consistent with the broadest scope in accordance with the principles and characteristics disclosed.

[0013] Embodiment one, as shown, the embodiment of the present application provides a human fall radar perception correction method combined with multi-modal data, comprising: Figure 1 S10: obtaining radar echo data and infrared thermal imaging data of a target monitoring area, wherein the radar echo data includes target motion trajectory and distance information, and the infrared thermal imaging data includes human body contour and temperature distribution; Wherein, the target monitoring area refers to a specific physical space range covered by radar sensors and infrared thermal imaging sensors for human fall monitoring, which is the data acquisition object area of the whole perception correction system. The target monitoring area is usually demarcated according to the application scene demand, such as the room of the old care agency, the living room or bedroom in home care, the bed surrounding area of hospital ward, etc.

[0014] ​Radar echo data and infrared thermal imaging data of the target monitoring area are acquired, wherein the radar echo data refers to data formed by receiving target reflected signals after the radar sensor transmits electromagnetic waves, including target motion trajectory and distance information, the target motion trajectory includes dynamic information such as the moving path, speed change and motion direction of the human body in the monitoring area analyzed through continuous echo signals, which can reflect the action process of standing, moving and falling of the human body; the distance information refers to the straight-line distance data between the radar and the target human body, with the dimension of meters (m), and the distance change can assist in judging the height change of the human body, such as the vertical distance from the distance sensor when falling, which will change significantly.

[0015] Further, the infrared thermal imaging data is image data formed by capturing the thermal radiation of an object through an infrared sensor, including a human body contour and a temperature distribution, wherein the core of the human body contour is the body shape boundary outlined based on the temperature difference between the human body and the environment, since the normal body temperature of the human body is 36-37℃, and the environmental temperature is usually lower, the human body and non-human objects can be accurately distinguished, solving the problem of type misjudgment of static targets by the radar; the core of the temperature distribution is the temperature gradient data in the monitoring area, with the dimension of Celsius (℃), since the temperature distribution of the human body has continuity and a specific range, whether the target is a living human body can be further verified, and non-human heat source interference can be excluded.

[0016] In summary, by collecting data in a targeted manner for the target monitoring area, the space range that needs to be monitored can be focused on, such as an old room or a sickroom area, the interference data of irrelevant targets outside the area is reduced, and the data processing complexity is reduced. At the same time, by capturing key information in the monitoring area from different dimensions through radar echo data and infrared thermal imaging data, the features are complementary, the limitation that single radar data is easily disturbed by the environment is broken through, and more comprehensive basic feature support is provided for fall perception.

[0017] S20: performing spatio-temporal alignment preprocessing on the radar echo data and the infrared thermal imaging data, unifying data timestamps and spatial coordinate systems, to obtain a standardized multi-modal data set; The radar echo data and the infrared thermal imaging data are subjected to spatio-temporal alignment preprocessing to obtain a standardized multi-modal data set, including: The sampling timestamp sequence and the spatial coordinate reference point of the radar echo data are extracted, wherein the spatial coordinate reference point is the geometric center coordinate of the radar monitoring area; The frame acquisition time sequence and the pixel coordinate origin of the infrared thermal imaging data are extracted, wherein the pixel coordinate origin is the top left corner vertex coordinate of the infrared imaging picture; Based on the sampling timestamp sequence and the frame acquisition time sequence, a time interpolation algorithm is used to synchronize and calibrate the time dimensions of the two, to obtain time-aligned bimodal data; Based on the spatial coordinate reference point and the pixel coordinate origin, the pixel coordinates of the infrared thermal imaging data are mapped to the radar spatial coordinate system based on a preset coordinate conversion matrix, to obtain spatially matched dual-mode data. The dual-mode data are fused in time alignment and spatial matching to construct a standardized multi-modal data set.

[0018] After obtaining the radar echo data and infrared thermal imaging data of the target monitoring area, the radar echo data and thermal infrared imaging data need to be preprocessed in time and space to eliminate the heterogeneity of the two types of data in the time and space dimensions, providing standardized input for subsequent bias model construction.

[0019] First, the sampling timestamp sequence and spatial coordinate reference point of the radar echo data are extracted. The sampling timestamp sequence refers to the time record sequence when the radar sensor collects echo data, usually in the format of year-month-day-hour-minute-second-millisecond, such as 2025-08-21-10:00:00.000, 2025-08-21-10:00:00.010, which can reflect the time distribution of radar data, i.e., collecting once every 10 milliseconds. The spatial coordinate reference point refers to the geometric center coordinate of the radar monitoring area, represented by three-dimensional spatial coordinates (x, y, z) with a dimension of meters (m), such as a 5m x 5m room with the ground as the z=0 plane. The timestamp provides a reference for subsequent time synchronization, and the spatial reference point establishes a unified reference origin for the radar coordinate system, ensuring the consistency of spatial coordinates.

[0020] Second, the frame acquisition time sequence and pixel coordinate origin of the infrared thermal imaging data are extracted. The frame acquisition time sequence refers to the time record sequence when the infrared sensor generates thermal imaging pictures, with the same format as the radar timestamp, such as 2025-08-21-10:00:00.000, 2025-08-21-10:00:00.030, reflecting the frame interval of infrared images, i.e., generating a frame every 30 milliseconds. The pixel coordinate origin is defined as the top-left corner coordinate of the infrared imaging picture, represented by two-dimensional pixel coordinates (u, v) with a dimension of pixels, usually set to (0, 0), and other point coordinates in the picture are defined based on this origin, such as the point coordinate (100, 50) 100 pixels to the right and 50 pixels below. Since the sampling frequencies of infrared and radar are different, the extraction of the time sequence can provide a basis for time calibration, and the pixel origin can establish a reference for the infrared coordinate system, facilitating subsequent coordinate conversion.

[0021] Further, based on the sampling timestamp sequence and the frame acquisition time sequence, a time interpolation algorithm is used to synchronize and calibrate the time dimension of both, to obtain time-aligned bimodal data. Because the sampling / frame generation frequencies of the radar and the infrared sensor are different, that is, the timestamps do not coincide, the missing values at non-coinciding time points of the two types of data need to be supplemented through mathematical calculation, so that the radar echo data and the infrared thermal imaging data form a one-to-one corresponding mapping relationship in the time dimension, and time synchronization is achieved. Therefore, a time interpolation algorithm is used for time dimension synchronization and calibration. Specifically, the time interpolation algorithm supplements reasonable intermediate values between the observation data at known time points, to realize time synchronization and calibration of the two types of data.

[0022] For example, based on the sampling timestamp sequence of the radar and the frame acquisition time sequence of the infrared, a linear interpolation algorithm is used to supplement the time gap data, to realize time dimension alignment. Radar timestamp: =0.010s, position( , ), =0.030s, position( , ), infrared frame timestamp: =0.020s, the data of the radar at needs to be supplemented.

[0023] Suppose that within the time interval from to , the radar data changes linearly with time, then the radar position at is ( + ) / 2, ( + ) / 2), so that the data of the radar and the infrared at are time-synchronized.

[0024] Further, while synchronizing and calibrating in the time dimension, based on the spatial coordinate reference point and the pixel coordinate origin, the pixel coordinates of the infrared thermal imaging data are mapped to the radar spatial coordinate system based on a preset coordinate conversion matrix, to obtain spatially matched bimodal data. Because the radar data is based on a three-dimensional spatial coordinate system, reflecting the physical space position, while the infrared data is based on a two-dimensional pixel coordinate system, reflecting the image internal position, the heterogeneous coordinate systems will cause the coordinates of the same point in the two types of data to be unable to correspond, so matching needs to be realized in the spatial dimension.

[0025] The infrared pixel coordinates (u, v) are mapped to the radar space coordinates (x, y, z) based on a preset coordinate conversion matrix, wherein the preset coordinate conversion matrix is a mathematical matrix determined through a calibration experiment before sensor installation, is used for converting a two-dimensional pixel coordinate system of infrared thermal imaging data into a three-dimensional space coordinate system of radar data, and realizes space dimension matching of the two types of data. The matrix is obtained through a calibration experiment: a calibration object with known radar coordinates, such as a reference point with a heat source, is placed in the monitoring area, pixel coordinates of the calibration object in the infrared image are recorded, conversion matrix parameters are calculated through a linear regression method or the like, and are preset into the system. In application, the formula is applied, wherein M is the preset coordinate conversion matrix.

[0026] For example, a 3x3 matrix M is obtained through sensor installation calibration, a point (300, 200) in the infrared is converted to correspond to radar coordinates (2.5, 2.5, 1.5) m, and space position matching is realized.

[0027] Finally, the dual-mode data of time alignment and space matching are fused, and a standardized multi-modal data set is finally formed. The data set simultaneously contains dynamic motion features of the radar and static human body features of the infrared, and the two types of features correspond to each other in time and space dimensions, provides a consistent and reliable input basis for subsequent training and reasoning of a fall perception bias model, and avoids feature mis-association problems caused by data heterogeneity.

[0028] S30: based on the standardized multi-modal sample data of the historical fall events and the corresponding true fall state labels, a fall perception bias model is constructed; Specifically, based on the standardized multi-modal sample data of the historical fall events and the corresponding true fall state labels, a fall perception bias model is constructed, including: The standardized multi-modal sample data of the historical fall events are collected, the corresponding true fall state labels of each historical fall event are synchronously acquired, and a sample data set is formed; According to the sample data set, a bias value of a historical radar original perception result and a true fall state label is calculated; Based on machine learning, the standardized multi-modal sample data are used as input features, and the corresponding bias value is used as a supervision target, and the fall perception bias model is constructed.

[0029] The standardized multi-modal sample data of the historical fall events are collected, the corresponding true fall state labels of each historical fall event are synchronously acquired, and a sample data set is formed, including: From the historical monitoring logs of the target monitoring area, records of the fall events are screened, and a time interval of the historical fall events is determined; Within the time interval, the corresponding historical radar echo data and historical infrared thermal imaging data are extracted and standardized to obtain standardized multi-modal sample data; The third-party verification records of each historical fall event are synchronously called, and the real fall state identifier is extracted therefrom, wherein the real fall state identifier includes a fall occurrence time, a fall posture, and a misjudgment identifier; The standardized multi-modal sample data of the same historical fall event is associated and bound with the corresponding real fall state label to form a sample data set.

[0030] Specifically, to construct the fall perception bias model, the standardized multi-modal sample data based on historical fall events and the corresponding real fall state label are used as supervised learning samples to ensure that the model can learn the rules of radar perception bias. First, the standardized multi-modal sample data of historical fall events is collected, and the corresponding real fall state label of each historical fall event is synchronously obtained to form a sample data set. The records of the historical fall events are first screened from the historical monitoring logs of the target monitoring area to determine the time interval of the historical fall events. The historical monitoring logs refer to the archived files of radar echo data, infrared thermal imaging data, and system original perception results recorded by sensors continuously, which contain a large amount of daily monitoring data. However, only the fall event related data is meaningful for model training, so it is necessary to screen out the records of the fall events to reduce the interference of invalid data and reduce the data processing cost. The records of the fall events are screened out, and the time range of each event is accurately determined.

[0031] Secondly, within the determined time interval, the corresponding historical radar echo data, including the motion trajectory, distance information, and historical infrared thermal imaging data, including the human body contour and temperature distribution, are extracted. Since the original historical data may have problems such as time misalignment and non-uniform spatial coordinates, directly using them will cause feature distortion, so it is necessary to unify the time stamp and spatial coordinate system according to the spatio-temporal alignment preprocessing method to obtain standardized multi-modal sample data.

[0032] Further, the third-party verification records of each historical fall event are called, which are independent of the sensor system, such as on-site judgment records of medical personnel and manual annotation records of monitoring videos, and the real fall state identifier is extracted therefrom. The real fall state includes the fall occurrence time, which is used to judge the radar time bias; the fall posture, such as straight falling, falling on the side, and falling on the stomach, which is used as a reference to judge the radar posture misjudgment; and the misjudgment identifier, a binary identifier, 0 indicating that the radar original perception is correct and 1 indicating that the radar original perception is incorrect, which is used to strengthen the model's learning of misjudgment scenarios. The third-party verification records are objective, and the multi-dimensional information of the real state identifier can fully describe the types of radar perception bias, enabling the model to learn the bias rules in different dimensions.

[0033] Finally, the standardized multi-modal sample data of the same historical fall event is associated with the corresponding real fall state label by event ID or time interval to form a structured sample data set, each sample containing multi-modal features, time labels, posture labels and misjudgment labels. The one-to-one correspondence between such input features and bias labels can be directly used for training of machine learning models, ensuring the normativity and repeatability of the model training process.

[0034] After obtaining the sample data set, further, the bias value of the historical radar original perception result and the real fall state label is calculated, including: extracting the time point at which the radar determines the occurrence of a fall and the fall posture type recognized by the radar in the historical radar original perception result, wherein the fall posture type includes upright collapse, side fall and prone fall; extracting the corresponding actual fall occurrence time point and actual fall posture type in the real fall state label; calculating the time bias coefficient, wherein the time bias coefficient is the absolute difference between the time point at which the radar determines the occurrence of a fall and the time point at which the actual fall occurs, divided by the maximum value of all time bias values in the historical sample; setting the posture bias coefficient, wherein when the fall posture type recognized by the radar is consistent with the actual fall posture type, the posture bias coefficient is 0, and when it is inconsistent, the posture bias coefficient is 1; calculating the bias value of the historical radar original perception result and the real fall state label based on the preset time weight coefficient and the posture weight coefficient.

[0035] First, two core features are extracted from the historical radar records in the sample data set, including the time point at which the radar determines the occurrence of a fall, i.e. the specific time at which the radar sensor original algorithm determines the occurrence of a fall event, reflecting the radar's original judgment of the fall time; and the fall posture type recognized by the radar, i.e. the human fall posture recognized by the radar based on echo data, limited to three standardized types, upright collapse (human body vertically falling from standing state), side fall (human body falling to the side), and prone fall (human body falling forward), reflecting the radar's original judgment of the fall shape.

[0036] Secondly, from the corresponding real fall state label, the matching comparison features are extracted, including the actual fall occurrence time point and the actual fall posture type, which are respectively taken as the basis for time bias calculation and posture bias calculation.

[0037] Further, a time deviation coefficient is calculated, which is a dimensionless index quantifying the degree of deviation of the radar in determining the time of fall occurrence, and the time deviation coefficient = |t_radar-t_actual| / t_max, t_radar is the time of fall determined by the radar, t_actual is the actual time of fall, and t_max is the maximum value of all time deviation values in the historical sample set.

[0038] Suppose in a certain historical sample, t_radar = 09:31:18, converted into a timestamp value T_r = 1000s, t_actual = 09:31:15, converted into a timestamp value T_actual = 997s, and the maximum time deviation t_max in the historical sample set is 10s, then the absolute difference t_radar-t_actual = |1000-997| = 3s, and the time deviation coefficient = 3 / 10 = 0.3.

[0039] At the same time, a posture deviation coefficient is calculated, which is a binary index quantifying the accuracy of the radar in recognizing the posture of fall, and a rule is set as follows: when the type of fall posture recognized by the radar is consistent with the actual type of fall posture, the posture deviation coefficient is 0; when the two are inconsistent, the posture deviation coefficient is 1.

[0040] For example, if the radar recognizes a side fall, and the actual posture is a side fall, then the posture deviation coefficient = 0; if the radar recognizes an upright fall, and the actual posture is a side fall, then the posture deviation coefficient = 1.

[0041] Finally, based on the time weight coefficient and the posture weight coefficient, a comprehensive deviation value is calculated, and the comprehensive deviation value = wt×time deviation coefficient + wp×posture deviation coefficient. The time weight coefficient wt is a preset parameter quantifying the contribution of the time deviation in the comprehensive deviation value, and is used to reflect the importance of the accuracy of determining the time of fall occurrence in the monitoring scene. Its value is a non-negative number, and together with the posture weight coefficient wp satisfies wt+wp=1, ensuring that the comprehensive deviation value is within the interval [0, 1].

[0042] In addition, the posture weight coefficient wp is a preset parameter quantifying the contribution of the posture deviation in the comprehensive deviation value, and is used to reflect the importance of the accuracy of recognizing the posture of fall in the monitoring scene. Its value is also a non-negative number, and the sum of the time weight coefficient and the posture weight coefficient is 1, reflecting the difference in the demand for posture recognition accuracy in different scenes through weight allocation.

[0043] Specifically, the preset of the time weight coefficient wt and the posture weight coefficient wp needs to combine the demand differences of the time accuracy and the posture recognition accuracy of the monitoring scene, and is realized through experience setting or scene calibration. The core principle is to ensure that the contribution degrees of the two types of deviations in the comprehensive deviation value match the actual application demand. If applied to scenes such as old family care and lonely old people monitoring, the time sensitivity of timely rescue after falling is higher, the time weight coefficient wt = 0.7, wp = 0.3 can be increased to strengthen the modification priority of the time determination deviation.

[0044] The comprehensive deviation value is calculated, which is more suitable for the actual application demand and can be used as a supervised label of the sample data set, directly used for training of the subsequent fall perception deviation model, so that the model can learn the mapping relationship between the multi-modal features and the radar deviation, and provide a data-driven decision basis for the final perception result correction.

[0045] Further, based on machine learning, a standardized multi-modal sample data is used as an input feature, and a corresponding deviation value is used as a supervised target to construct a fall perception deviation model. Based on historical fall event data, the fall perception deviation model is constructed and trained. Since the feature interaction between the radar echo data and the infrared thermal imaging data is significant and the data relationship is complex, an exemplary neural network model is selected. The neural network is a computational model that simulates the structure and function of the biological nervous system. A large number of artificial neurons form a network through connection, can learn rules through data and make predictions or classifications. The basic structure includes an input layer, a hidden layer and an output layer. The input layer is used to receive raw data; the hidden layer is between the input layer and the output layer, and processes information through the connection weight between neurons; the output layer is used to output the final result. The core principle of the neural network is to flow the input data from the input layer, perform weighted calculation through the hidden layer neurons, and finally get the prediction result from the output layer. At the same time, by comparing the prediction result with the actual result, the connection weight of each layer of neurons is adjusted from the output layer, and the error is gradually reduced.

[0046] During specific training, the collected historical radar echo data, historical infrared thermal imaging data and corresponding real fall state labels are preprocessed to eliminate abnormal values, unify the time stamp and spatial coordinate system through space-time alignment, and eliminate interference caused by data heterogeneity. Then, the preprocessed data is divided in the ratio of 7:3: 70% as the training set for network learning parameter and weight update; 30% as the validation set for monitoring the generalization ability in the training process to avoid overfitting of the training data.

[0047] During the training process, the input layer receives the normalized standardized multi-modal sample data, and after weighted calculation by the hidden layer neurons, the predicted radar perception bias value (time bias coefficient, attitude bias coefficient) is obtained from the output layer. By comparing the error between the predicted value and the actual bias value, the weight gradient is calculated layer by layer using the back propagation algorithm, and then the optimizer dynamically adjusts the connection weights of each layer, and iterates repeatedly until the validation set error is stable and the accuracy reaches 95%, and finally a model capable of accurately predicting radar perception bias is formed. The model can take the real-time monitored standardized multi-modal data set as input and output the corresponding radar perception bias value, providing an accurate basis for correcting the original radar fall perception result.

[0048] S40: inputting the standardized multi-modal data set into the fall perception bias model, outputting a radar perception bias value, and correcting and adjusting the original radar fall perception result based on the radar perception bias value to obtain an accurate fall perception result.

[0049] Wherein, inputting the standardized multi-modal data set into the fall perception bias model to output a radar perception bias value comprises: inputting the standardized multi-modal data set into the fall perception bias model; The fall perception bias model performs feature fusion and bias calculation on the input standardized multi-modal data, and outputs a radar perception bias value, wherein the radar perception bias value includes a time bias amount and an attitude bias identifier.

[0050] First, the standardized multi-modal data set preprocessed by spatio-temporal alignment is input into the constructed fall perception bias model. The model performs feature fusion on the radar echo data (target motion trajectory, distance information) and infrared thermal imaging data (human contour, temperature distribution), combines the bias rules learned from history to perform bias calculation, and finally outputs a radar perception bias value. This bias value contains two types of key information: a time bias amount, which is used to quantify the difference between the radar-determined fall occurrence time point and the actual time, and corrects the time determination error; an attitude bias identifier, which is used to indicate whether the fall attitude type recognized by the radar is accurate, and corrects the attitude recognition error.

[0051] Further, based on the radar perception bias value, the original radar fall perception result is corrected and adjusted to obtain an accurate fall perception result, which comprises: extracting the original radar fall perception result, wherein the original radar fall perception result includes the radar-determined fall occurrence time point, fall attitude type, and fall occurrence determination; correcting and calibrating the radar-determined fall occurrence time point based on the time bias amount to obtain a corrected fall occurrence time point; updating the radar-determined fall attitude type based on the attitude bias identifier to obtain a corrected fall attitude type; The corrected fall occurrence time point, fall posture type and fall occurrence determination are integrated to form a precise fall perception result.

[0052] Specifically, first, core perception information is extracted from the original output of the radar system, including the radar-determined fall occurrence time point, fall posture type and fall occurrence determination; second, the time deviation amount output by the fall perception deviation model is used to adjust the fall occurrence time point determined by the radar. By calculating the original time point and the time deviation amount, the deviation amount is offset, and a corrected time point consistent with the actual fall time is obtained, solving the problem of lag or advance in time determination by the radar; further, according to the posture deviation identifier output by the deviation model, the fall posture type recognized by the radar is updated. If the identifier shows that the posture is misjudged, the correct posture type is replaced based on the model prediction result, improving the accuracy of posture recognition; finally, the corrected fall occurrence time point, the corrected fall posture type, and the original fall occurrence determination result are integrated to form a complete and precise fall perception result.

[0053] The result integrates the correction information of time and posture, ensuring the timeliness in the time dimension and the accuracy in the posture dimension, and more reliably supporting the timeliness judgment and rescue mode decision of fall rescue.

[0054] In summary, the embodiments of the present application have at least the following technical effects: Compared with the prior art, the present application first constructs a multi-modal perception system by fusing radar echo data and infrared thermal imaging data, uses the motion trajectory and distance information of the radar and the human body contour and temperature distribution of the infrared to realize data complementation, breaks through the perception limitation of single radar data in complex environment, and provides more comprehensive feature support for precise fall monitoring; second, the time stamp and spatial coordinate system of the multi-modal data are unified through spatio-temporal alignment preprocessing, ensuring that the radar and infrared data are synchronized in time dimension and matched in spatial dimension, avoiding feature mis-association problems caused by data mispositioning, and improving the effectiveness of multi-modal data fusion; third, a fall perception deviation model is constructed based on historical fall event samples and real labels, the radar perception deviation law is quantified through machine learning, so that the deviation correction has scientificity and accuracy driven by data, rather than relying on empirical adjustment; finally, the original radar result is corrected by the time deviation amount and posture deviation identifier output by the deviation model, effectively solving the problems of time determination deviation and posture recognition misjudgment, significantly improving the precision and reliability of the fall perception result, and providing more reliable technical support for timely rescue in elderly care, home care and other scenarios.

[0055] Embodiment two, as Figure 2As shown, based on the same inventive concept of the human fall radar perception correction method combined with multi-modal data provided in Embodiment One, the present embodiment also provides a human fall radar perception correction system combined with multi-modal data, comprising: a data acquisition module 11 configured to acquire radar echo data and infrared thermal imaging data of a target monitoring area, wherein the radar echo data comprises target motion trajectory and distance information, and the infrared thermal imaging data comprises human contour and temperature distribution; a data preprocessing module 12 configured to perform spatio-temporal alignment preprocessing on the radar echo data and the infrared thermal imaging data, unify data timestamps and spatial coordinate systems, and obtain a standardized multi-modal data set; a model construction module 13 configured to construct a fall perception bias model based on standardized multi-modal sample data of historical fall events and corresponding true fall state labels; a perception correction module 14 configured to input the standardized multi-modal data set into the fall perception bias model, output a radar perception bias value, correct and adjust an original radar fall perception result based on the radar perception bias value, and obtain an accurate fall perception result.

[0056] The data acquisition module 11 is specifically configured to: acquire data, including radar echo data and infrared thermal imaging data, in a target monitoring area; The radar echo data comprises target motion trajectory and distance information, and the infrared thermal imaging data comprises human contour and temperature distribution.

[0057] The data preprocessing module 12 is specifically configured to: perform spatio-temporal alignment preprocessing on the radar echo data and the infrared thermal imaging data to obtain a standardized multi-modal data set, including: extracting a sampling timestamp sequence and a spatial coordinate reference point of the radar echo data, wherein the spatial coordinate reference point is the geometric center coordinate of the radar monitoring area; extracting a frame acquisition time sequence and a pixel coordinate origin of the infrared thermal imaging data, wherein the pixel coordinate origin is the top-left corner vertex coordinate of the infrared imaging picture; based on the sampling timestamp sequence and the frame acquisition time sequence, synchronously calibrating the time dimensions of the two sequences using a time interpolation algorithm to obtain time-aligned bimodal data; based on the spatial coordinate reference point and the pixel coordinate origin, mapping the pixel coordinates of the infrared thermal imaging data to the radar spatial coordinate system based on a preset coordinate conversion matrix to obtain spatially matched bimodal data; fusing the time-aligned and spatially matched bimodal data to construct a standardized multi-modal data set.

[0058] The model construction module 13 is specifically configured to: construct a fall perception bias model based on standardized multi-modal sample data of historical fall events and corresponding true fall state labels, including: collecting standardized multi-modal sample data of historical fall events, synchronously obtaining corresponding true fall state labels of each historical fall event, and forming a sample data set; calculating the bias value of the historical radar original perception result and the true fall state label according to the sample data set; based on machine learning, using the standardized multi-modal sample data as input features, and using the corresponding bias value as a supervision target, to construct the fall perception bias model.

[0059] The collecting of standardized multi-modal sample data of historical fall events, the synchronously obtaining of corresponding true fall state labels of each historical fall event, and the forming of a sample data set include: filtering records of historical fall events from historical monitoring logs of a target monitoring area to determine a time interval of historical fall events; extracting corresponding historical radar echo data and historical infrared thermal imaging data within the time interval and performing standardization processing to obtain standardized multi-modal sample data; synchronously calling third-party verification records of each historical fall event and extracting true fall state labels therefrom, wherein the true fall state labels include a fall occurrence time, a fall posture, and a misjudgment identifier; associating and binding the standardized multi-modal sample data of the same historical fall event with the corresponding true fall state label to form a sample data set.

[0060] Further, the calculation of the bias value of the historical radar original perception result and the true fall state label includes: extracting a time point at which the radar determines that a fall occurs and a fall posture type recognized by the radar in the historical radar original perception result, wherein the fall posture type includes upright dumping, sideways falling, and prone falling; extracting a corresponding actual fall occurrence time point and an actual fall posture type in the true fall state label; calculating a time bias coefficient, wherein the time bias coefficient is the absolute difference between the time point at which the radar determines that a fall occurs and the actual fall occurrence time point divided by the maximum value of all time bias values in the historical sample; setting a posture bias coefficient, wherein when the fall posture type recognized by the radar is consistent with the actual fall posture type, the posture bias coefficient is 0, and when they are inconsistent, the posture bias coefficient is 1; Based on the preset time weight coefficient and posture weight coefficient, the deviation value between the historical radar original perception result and the actual fall state label is calculated.

[0061] The perception correction module 14 is specifically configured to: Inputting the standardized multimodal dataset into the fall perception bias model and outputting a radar perception bias value comprises: inputting the standardized multimodal dataset into a fall perception bias model; The fall perception deviation model performs feature fusion and deviation calculation on the input standardized multimodal data, and outputs a radar perception deviation value, wherein the radar perception deviation value includes a time deviation amount and a posture deviation identifier.

[0062] Furthermore, the original radar fall perception result is corrected and adjusted based on the radar perception deviation value to obtain an accurate fall perception result, including: Extracting original radar fall perception results, wherein the original radar fall perception results include the radar-determined fall occurrence time point, fall posture type, and fall occurrence determination; Correcting and calibrating the fall occurrence time point determined by the radar based on the time deviation to obtain a corrected fall occurrence time point; The fall posture type determined by the radar is updated based on the posture deviation identifier to obtain a corrected fall posture type; The corrected fall time point, fall posture type and fall occurrence determination are integrated to form an accurate fall perception result.

[0063] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0064] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

[0065] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, to the extent such modifications and variations fall within the scope of the present application and its equivalents, the present application is intended to include such modifications and variations.

Claims

1. A human fall radar perception correction method combining multimodal data, characterized in that: The method comprises: Acquire radar echo data and infrared thermal imaging data of the target monitoring area, wherein the radar echo data includes target motion trajectory and distance information, and the infrared thermal imaging data includes human body contour and temperature distribution; Performing spatiotemporal alignment preprocessing on the radar echo data and infrared thermal imaging data, unifying the data timestamps and spatial coordinate systems, and obtaining a standardized multimodal data set; Based on standardized multimodal sample data of historical fall events and the corresponding real fall status labels, a fall perception bias model is constructed; The standardized multimodal data set is input into the fall perception bias model, and a radar perception bias value is output. The original radar fall perception result is corrected and adjusted based on the radar perception bias value to obtain an accurate fall perception result.

2. The method according to claim 1, characterized in that The radar echo data and infrared thermal imaging data are preprocessed with time and space alignment to obtain a standardized multimodal data set, including: Extracting a sampling timestamp sequence and a spatial coordinate reference point of the radar echo data, wherein the spatial coordinate reference point is the geometric center coordinate of the radar monitoring area; Extracting the frame acquisition time series and pixel coordinate origin of the infrared thermal imaging data, wherein the pixel coordinate origin is the coordinate of the upper left corner vertex of the infrared imaging screen; Based on the sampling timestamp sequence and the frame acquisition time sequence, a time interpolation algorithm is used to synchronously calibrate the time dimensions of the two to obtain time-aligned bimodal data; Based on the spatial coordinate reference point and the pixel coordinate origin, the pixel coordinates of the infrared thermal imaging data are mapped to the radar spatial coordinate system based on a preset coordinate conversion matrix to obtain spatially matched dual-modal data; Fuse temporally aligned and spatially matched bimodal data to construct a standardized multimodal dataset.

3. The method according to claim 1, characterized in that Based on standardized multimodal sample data of historical fall events and the corresponding real fall status labels, a fall perception bias model is constructed, including: Collect standardized multimodal sample data of historical fall events, and simultaneously obtain the real fall status labels corresponding to each historical fall event to form a sample data set; Calculate the deviation between the original historical radar perception result and the actual fall state label according to the sample data set; Based on machine learning, the standardized multimodal sample data is used as input features and the corresponding deviation value is used as a supervision target to construct the fall perception deviation model.

4. The method according to claim 3, characterized in that Collect standardized multimodal sample data of historical fall events, and simultaneously obtain the actual fall status labels corresponding to each historical fall event to form a sample dataset, including: Filter the records of falls from the historical monitoring logs of the target monitoring area and determine the time interval of the historical falls; During the time interval, corresponding historical radar echo data and historical infrared thermal imaging data are extracted and standardized to obtain standardized multimodal sample data; Synchronously retrieve third-party verification records of each historical fall event and extract the true fall status identifier from it, wherein the true fall status identifier includes the time of fall, fall posture and misjudgment identifier; The standardized multimodal sample data of the same historical fall event are associated and bound with the corresponding real fall status labels to form a sample dataset.

5. The method according to claim 3, characterized in that Calculate the deviation between the historical radar raw perception results and the actual fall status label, including: Extracting the time point at which the radar determined the fall occurred and the fall posture type identified by the radar from the historical radar raw perception results, wherein the fall posture type includes upright fall, sideways fall, and prone fall; Extracting the actual fall time point and the actual fall posture type corresponding to the real fall state label; Calculating a time deviation coefficient, wherein the time deviation coefficient is the absolute difference between the time point when the radar determines that the fall occurred and the time point when the fall actually occurred, divided by the maximum value of all time deviation values ​​in the historical samples; Set the posture deviation coefficient. When the fall posture type identified by the radar is consistent with the actual fall posture type, the posture deviation coefficient is 0; when they are inconsistent, the posture deviation coefficient is 1. Based on the preset time weight coefficient and posture weight coefficient, the deviation value between the historical radar original perception result and the actual fall state label is calculated.

6. The method according to claim 1, characterized in that Inputting the standardized multimodal dataset into the fall perception bias model and outputting a radar perception bias value comprises: inputting the standardized multimodal dataset into a fall perception bias model; The fall perception deviation model performs feature fusion and deviation calculation on the input standardized multimodal data, and outputs a radar perception deviation value, wherein the radar perception deviation value includes a time deviation amount and a posture deviation identifier.

7. The method according to claim 1, characterized in that The original radar fall perception result is corrected and adjusted based on the radar perception deviation value to obtain an accurate fall perception result, including: Extracting original radar fall perception results, wherein the original radar fall perception results include the radar-determined fall occurrence time point, fall posture type, and fall occurrence determination; Correcting and calibrating the fall occurrence time point determined by the radar based on the time deviation to obtain a corrected fall occurrence time point; The fall posture type determined by the radar is updated based on the posture deviation identifier to obtain a corrected fall posture type; The corrected fall time point, fall posture type and fall occurrence determination are integrated to form an accurate fall perception result.

8. A human fall radar perception and correction system combining multimodal data is characterized by: Used to perform the method according to any one of claims 1 to 7, comprising: A data acquisition module is used to acquire radar echo data and infrared thermal imaging data of the target monitoring area, wherein the radar echo data includes target motion trajectory and distance information, and the infrared thermal imaging data includes human body contour and temperature distribution; A data preprocessing module is used to perform spatiotemporal alignment preprocessing on the radar echo data and infrared thermal imaging data, unify the data timestamp and spatial coordinate system, and obtain a standardized multimodal data set; A model building module is used to build a fall perception bias model based on standardized multimodal sample data of historical fall events and the corresponding real fall status labels; The perception correction module is used to input the standardized multimodal dataset into the fall perception bias model, output a radar perception bias value, and correct and adjust the original radar fall perception result based on the radar perception bias value to obtain an accurate fall perception result.

Citation Information

Patent Citations

  • Real-time calibration method and device

    CN110132305A

  • Human body fall detection adaptive system and device based on UWB radar

    CN110703241A

  • Intelligent tumble posture classification and identification method based on mobile phone sensor

    CN114818952A

  • Fall detection method and device based on millimeter wave radar and millimeter wave radar equipment

    CN116027324A

  • Multi-mode tumble detection method, device and equipment

    CN120220230A

Cited By

  • Abnormal behavior intelligent alarm system and method based on ultra wide band spatial positioning and multi-mode sensing

    CN122266110A