Multi-mode tumble detection method, device and equipment
Through the multimodal fall detection method, radar and infrared sensors combined with time series analysis and deep learning, high-precision fall detection is achieved and user privacy is protected, solving the problem of difficulty in taking into account both accuracy and privacy in the existing technology.
Patent Information
- Application Number
- CN202510295573.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The existing fall detection methods are difficult to balance between accuracy and privacy protection. The sensor method has a high false alarm rate and violates privacy, while the visual method poses privacy risks.
Multimodal fall detection method is adopted to collect data through radar sensors and infrared image sensors, and combine time series analysis and deep learning behavior recognition to generate time series multimodal fusion data to achieve high-precision fall detection without infringing on user privacy.
It improves the accuracy of fall detection, reduces the false alarm rate, and effectively protects users' privacy, providing a fall detection solution that takes into account detection accuracy and privacy protection.
Smart Images

Figure CN120220230A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of intelligent care, and more particularly relates to a multi-modal fall detection method, device, and equipment. Background Art
[0002] Current fall detection methods are mainly divided into wearable sensor methods or vision methods.
[0003] Wearable fall detection methods rely on accelerometers, gyroscopes, heart rate sensors, etc., which are usually integrated in smart bracelets, smart watches, or devices worn around the waist. Fall detection is mainly based on features such as acceleration thresholds, posture changes, and speed impacts. In this method, the sensors are directly attached to the human body, and can accurately capture changes in the motion state. However, users must wear the device at all times, and may forget to wear or remove the device, resulting in detection failure, and this method has a high false alarm rate.
[0004] The camera-based fall detection method uses technologies such as object detection, pose estimation, and deep learning to analyze human postures, contour changes, and motion trajectories to detect fall behaviors. In this method, users do not need to wear a device, and the detection is more natural. However, the camera needs to continuously monitor the user, which involves privacy risks, especially in private places such as bedrooms and bathrooms.
[0005] Therefore, how to provide a fall detection solution that takes into account both detection accuracy and user privacy protection is a research topic worthy of study. Summary of the Invention
[0006] In view of the above analysis, embodiments of the present invention aim to provide a multi-modal fall detection method, device, and equipment, aiming to improve detection accuracy while protecting user privacy.
[0007] In the first aspect of this application, a multi-modal fall detection method is provided, including:
[0008] Obtain fall detection data, which is collected by a sensor module and includes radar point cloud data collected by a radar sensor and thermal infrared imaging data collected by an infrared image sensor;
[0009] Extract moving points from the radar point cloud data based on time series analysis, filter out moving points whose positions change within a continuous time period, and generate moving point cloud data;
[0010] Project the moving point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix;
[0011] Combine the radar point cloud projection matrix with the thermal infrared imaging data to calculate the three-dimensional coordinates of the target and generate target height change data;
[0012] When the target height change data meets the preset fall condition, input the constructed time-series multi-modal fusion data into the deep learning behavior recognition module to output the classification result of the fall behavior; wherein, the time-series multi-modal fusion data includes thermal infrared imaging data of consecutive frames and a radar point cloud projection matrix.
[0013] Optionally, the projecting the moving point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix includes:
[0014] Perform coordinate transformation on the filtered radar point cloud data and calculate the angular information of the radar point cloud data;
[0015] Perform angular mapping based on the field of view angle of the infrared image sensor and calculate the projected pixel coordinates of the point cloud in the thermal infrared imaging data;
[0016] Construct a radar point cloud projection matrix with the same size as the thermal infrared imaging data and mark the projected pixel coordinates in the radar point cloud projection matrix;
[0017] Wherein, the radar sensor and the infrared image sensor are coaxially installed.
[0018] Optionally, after the projecting the moving point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix, it further includes:
[0019] Based on the radar point cloud projection matrix, construct a binary mask of the radar point cloud to generate a radar point cloud mask;
[0020] Filter the high-temperature area based on the thermal infrared imaging data to generate an infrared target mask;
[0021] Calculate the intersection of the radar point cloud mask and the infrared target mask to generate multi-modal fusion data.
[0022] Optionally, the combining the radar point cloud projection matrix and the thermal infrared imaging data to calculate the three-dimensional coordinates of the target and generate target height change data includes:
[0023] Obtain the highest FOV angle measured by the radar sensor;
[0024] Obtain the highest FOV angle of the target human body contour in the thermal infrared imaging data;
[0025] Use the highest FOV angle of the target human body contour in the thermal infrared imaging data to correct the highest FOV angle measured by the radar sensor to obtain the highest point of the target and generate target height change data.
[0026] Optionally, the step of inputting the constructed time series multimodal fusion data into a deep learning behavior recognition module and outputting the classification result of the fall behavior includes:
[0027] The deep learning behavior recognition module performs time series modeling based on the time displacement module and classifies the target behavior, and the classification of the target behavior includes falling, squatting, sitting down, and object dropping.
[0028] Optionally, before inputting the constructed time series multimodal fusion data into the deep learning behavior recognition module, the method further includes:
[0029] Based on the target height change data, when it is detected that the time period from the start of the target height change to the end of the change is less than a preset threshold, radar point cloud data and thermal infrared image data of equal length are retrieved forward and backward for filling.
[0030] Optionally, the preset fall condition is: a target height change rate exceeds a preset rate; or a target height is lower than a preset height for a time longer than a preset time.
[0031] Optionally, after inputting the constructed time series multimodal fusion data into the deep learning behavior recognition module and outputting the classification result of the fall behavior, the method further includes:
[0032] If the deep learning behavior recognition module determines that the behavior is a fall, the pre-trained speech classification model is called to output the detection result of whether the audio data includes negative emotional speech;
[0033] If the audio detection result contains negative emotional speech, an alarm will be triggered directly;
[0034] If the audio detection result does not include negative emotional speech, the thermal infrared imaging data will be sent to the predetermined contact and an alarm will be played. If no cancellation command is received within the specified time, the alarm will be triggered.
[0035] According to a second aspect of the present application, a multimodal fall detection device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the multimodal fall detection method according to any one of the above is implemented.
[0036] In a third aspect of the present application, a multi-modal fall detection device is provided, comprising: a sensor device, an alarm device, and the multi-modal fall detection device according to the above;
[0037] The sensor module includes a radar sensor and an infrared image sensor; the radar sensor is configured to collect radar point cloud data; the infrared image sensor is configured to collect thermal infrared imaging data;
[0038] The multi-modal fall detection device is configured to detect a target fall based on the radar point cloud data and the thermal infrared imaging data;
[0039] The alarm module is configured to receive the detection result of the multi-modal fall detection device and execute an alarm reminder.
[0040] The multi-modal fall detection method provided by this application obtains fall detection data, including radar point cloud data collected by a radar sensor and thermal infrared imaging data collected by an infrared image sensor; performs motion point extraction on the radar point cloud data based on time series analysis, filters out motion points whose positions change within a continuous time, and generates motion point cloud data; projects the motion point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix; combines the radar point cloud projection matrix and the thermal infrared imaging data to calculate the three-dimensional coordinates of the target and generate target height change data; when the target height change data meets a preset fall condition, inputs the constructed time series multi-modal fusion data into a deep learning behavior recognition module to output a classification result of the fall behavior. By fusing radar sensor and thermal infrared imaging data and combining time series analysis and deep learning behavior recognition, this application achieves high-precision fall detection without infringing on user privacy.
[0041] In addition, this application also provides a multi-modal fall detection device and a readable storage medium having the above technical effects. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this specification, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0043] Figure 1 It is a flowchart of a specific implementation manner of the multi-modal fall detection method provided by this application;
[0044] Figure 2 It is a flowchart of generating a radar point cloud projection matrix in the embodiments of this application;
[0045] Figure 3 It is a flowchart of the implementation process of generating the radar point cloud data of the target in the embodiments of this application;
[0046] Figure 4 It is a structural block diagram of the multi-modal fall detection device provided by this application;
[0047] Figure 5Block diagram of the multi-modal fall detection device provided by this application. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. It should be noted that, without conflict, the implementation manners and features in the present disclosure can be combined with, separated from, interchanged with, and / or rearranged with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0049] The terms used here are for the purpose of describing specific embodiments and are not intended to be restrictive. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. In addition, when the terms "comprise" and / or "include" and their variants are used in this specification, it is indicated that there are the stated features, wholes, steps, operations, components, assemblies, and / or groups thereof, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, components, assemblies, and / or groups thereof. It should also be noted that, as used here, the terms "substantially", "about", and other similar terms are used as approximate terms and not as degree terms, so they are used to explain the inherent deviations of measured values, calculated values, and / or provided values that those of ordinary skill in the art will recognize.
[0050] The flowchart of a specific implementation manner of the multi-modal fall detection method provided by this application is as Figure 1 shown, and the method specifically includes:
[0051] S101: Obtain fall detection data, where the fall detection data is collected by a sensor module and includes radar point cloud data collected by a radar sensor and thermal infrared imaging data collected by an infrared image sensor.
[0052] The sensor module includes a radar sensor (for obtaining point cloud data) and an infrared image sensor (for obtaining thermal infrared imaging data). The collected data includes: radar point cloud data, which is the three-dimensional coordinate information of the target in the environment; and thermal infrared imaging data, which is the thermal imaging contour information of the human target. Among them, the radar sensor can be a millimeter-wave radar sensor.
[0053] As a specific implementation, the trigger condition for obtaining fall detection data can be: initially determining that there is a suspected fall behavior. Specifically, the radar sensor can continuously obtain the point cloud of the moving human body within the current time period, and based on the distance and angle information, divide the point cloud belonging to different individuals to obtain the height change of the target point cloud data. Combining time series analysis, it is judged whether there is a suspected fall situation. For example, if it is calculated that the height change of the highest point of the target drops by more than 1m within 1s, it is considered that the target has a suspected fall behavior, triggering a further fall judgment process.
[0054] When it is initially determined that there is a suspected fall behavior, the radar point cloud data collected by the radar sensor and the thermal infrared imaging data collected by the infrared image sensor during the period from the start of height change to the end of change are obtained as fall detection data.
[0055] In order to adapt to the input of the subsequent deep learning behavior recognition module, when the time period from the start of the detected target height change to the end of change is less than the preset threshold (for example, 2s), equal-length radar point cloud data and thermal infrared image data are retrieved forward and backward for filling.
[0056] S102: Extract moving points from the radar point cloud data based on time series analysis, screen out the moving points whose positions change within continuous time, and generate moving point cloud data.
[0057] Compare the continuous frame point cloud data, and extract the point cloud data of the moving target. For example, a time window (such as 1s) can be set to analyze the point cloud movement trajectory within this time period. Calculate the speed change of the point cloud, and eliminate the points with a speed close to zero (such as the ground, walls). Use the DBSCAN clustering algorithm to eliminate isolated points and retain the dense area (i.e., the target), and finally generate moving point cloud data.
[0058] S103: Project the moving point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix.
[0059] In this step, the three-dimensional point cloud data obtained by the radar sensor needs to be projected onto the coordinate system of the thermal infrared image for fusion calculation.
[0060] Such as Figure 2 As shown in the flowchart of generating the radar point cloud projection matrix in the embodiment of the present application, in this step, projecting the moving point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix specifically includes:
[0061] S201: Perform coordinate transformation on the screened radar point cloud data and calculate the angle information of the radar point cloud data.
[0062] Under normal circumstances, radar point cloud data is provided in polar coordinates or rectangular coordinates. The images of infrared image sensors are in a two-dimensional pixel coordinate system, and coordinate transformation is required to convert the radar point cloud from the radar coordinate system to the infrared imaging coordinate system. In this application, the radar sensor and the infrared image sensor are coaxially installed, that is, they share a rotation center, but the imaging methods are different, and coordinate mapping is required.
[0063] S202: Perform angle mapping based on the field of view angle of the infrared image sensor, and calculate the projected pixel coordinates of the point cloud in the thermal infrared imaging data.
[0064] According to the field of view angle (FOV) of the infrared image sensor, calculate the projected pixel coordinates of each point cloud in the thermal infrared imaging data.
[0065] S203: Construct a radar point cloud projection matrix with the same size as the thermal infrared imaging data, and mark the projected pixel coordinates in the radar point cloud projection matrix.
[0066] Create a projection matrix with the same size as the thermal infrared imaging data, with a size of H×W (the resolution of the thermal infrared image). Traverse all points in the radar point cloud, map the calculated projected pixel coordinates to the projection matrix, and fill in the values to finally form the radar point cloud projection matrix, where each pixel point corresponds to a radar point depth value.
[0067] As Figure 3 As shown in the flowchart of the implementation process of generating the radar point cloud data of the target in the embodiments of this application, after projecting the moving point cloud data to the coordinate system corresponding to the thermal infrared imaging data and generating the radar point cloud projection matrix, the following steps are further included:
[0068] S301: Based on the radar point cloud projection matrix, construct a binary mask of the radar point cloud to generate a radar point cloud mask.
[0069] Construct a binary mask to represent the coverage area of the radar point cloud on the infrared image, and generate a radar point cloud mask. To increase the coverage range of the point cloud, each point can be expanded with a 3×3 dilation operation.
[0070] S302: Screen the high-temperature areas based on the thermal infrared imaging data to generate an infrared target mask.
[0071] Since the human body temperature is usually higher than the environment, a temperature threshold can be set, such as 30°C (which can be adjusted according to the ambient temperature), for screening to generate an infrared target mask.
[0072] S303: Calculate the intersection of the radar point cloud mask and the infrared target mask to generate multimodal fusion data.
[0073] Match the regions of the radar point cloud mask and the infrared target mask, extract the radar point cloud data that truly belongs to the target, and use it for subsequent behavior recognition. By calculating the intersection of the radar point cloud mask and the infrared target mask, the radar point cloud and the high-temperature region in the infrared image are mutually supplemented, so as to eliminate the interference of low-temperature moving objects and stationary high-temperature objects and obtain the accurate position of the human body.
[0074] Among them, the radar sensor and the infrared image sensor are coaxially installed. In practice, the physical positions of the radar sensor and the infrared image sensor are a few centimeters apart, and it can be approximately considered that they are in the same position in the three-dimensional space.
[0075] Map the point cloud output by the radar to the corresponding pixels of the infrared image according to the angle information, the FOV of the infrared image sensor, and the pixel parameters. For example, if the angle of a moving point obtained by the radar is (45°, 30°), and the FOV of the infrared image sensor is 45° on the left and right and 30° on the top and bottom, this moving point can be projected onto the top-rightmost pixel of the infrared image. At this time, an empty image with the same size as the infrared image is generated. Each moving point that can be projected into the infrared field of view will generate a 3x3 mask on this empty image, and the intersection of the merged mask and the high-temperature part mask obtained by threshold segmentation (usually taking 28°C) of the infrared image is taken to exclude high-temperature stationary objects (such as heaters) and low-temperature moving objects (fans, dropped water cups, clothes, etc.).
[0076] S104: Combine the radar point cloud projection matrix and the thermal infrared imaging data to calculate the three-dimensional coordinates of the target and generate the target height change data.
[0077] By fusing the radar point cloud projection matrix and the infrared thermal imaging data, calculate the three-dimensional coordinates of the target, track its height change, and generate the target height change data.
[0078] Obtain the highest FOV angle measured by the radar sensor; obtain the highest FOV angle of the target human body contour in the thermal infrared imaging data; use the highest FOV angle of the target human body contour in the thermal infrared imaging data to correct the highest FOV angle measured by the radar sensor to obtain the highest point of the target and generate the target height change data.
[0079] Considering that the angle resolution of the radar on the vertical axis is poor, in this application, by focusing on the horizontal axis coordinates of the point cloud, first obtain the position of the human body contour in the infrared image through the horizontal axis coordinates and the high-temperature region, and then use the highest FOV angle corresponding to the human body contour in the infrared image to correct the highest FOV angle measured by the radar to obtain a more accurate relative height change curve of the human body.
[0080] S105: When the target height change data meets the preset fall condition, input the constructed time-series multi-modal fusion data into the deep learning behavior recognition module to output the classification result of the fall behavior.
[0081] Among them, the time-series multi-modal fusion data includes thermal infrared imaging data of consecutive frames and a radar point cloud projection matrix.
[0082] The preset fall condition is: the target height change rate exceeds the preset rate, for example, the target drops more than 1m within 1s; or the time when the target height is lower than the preset height exceeds the preset time, for example, lower than 0.5m for more than 2s.
[0083] The deep learning behavior recognition module is a video-based behavior recognition model. Typical video recognition models include C3D, SlowFast Network, etc. This application can perform temporal modeling based on the time displacement module to classify target behaviors. The classification of target behaviors includes falls, squats, sits, and object drops.
[0084] During the training process, the deep learning behavior recognition module uses the time displacement module to obtain temporal modeling capabilities, and uses 2000 segmented infrared image + point cloud sequences collected for training. Each segment is two seconds long and has 16 frames, and a total of four types of data including squats, falls, object drops, and sits are included. Use the method described above to map the 3D point cloud to the infrared image sequence for fusion enhancement, and input the data sequence of the corresponding time period into the trained deep learning behavior recognition module for behavior classification of behaviors such as falls, squats, and sits.
[0085] Before inputting the constructed time-series multi-modal fusion data into the deep learning behavior recognition module, it also includes: based on the target height change data, when the time period from the start of the target height change to the end of the change is less than the preset threshold, retrieve equal-length radar point cloud data and thermal infrared image data forward and backward for filling.
[0086] This application uses a radar sensor and an infrared image sensor, and accurately projects the 3D point cloud data of the radar sensor into the infrared image through a geometric mapping method for information enhancement. This mapping method utilizes the field of view (FOV) of the sensor and performs coordinate transformation through angle matching.
[0087] This application uses a 3×3 mask fusion method. After the radar point cloud is mapped, a mask matching the infrared image is generated for the point cloud, and an intersection operation is performed with the mask of the high-temperature region of the thermal imaging to eliminate high-temperature stationary objects (such as heaters) and low-temperature moving objects (such as fans, clothes, etc.). This method significantly reduces misjudgments and improves the accuracy of human target recognition.
[0088] Since the vertical axis resolution of the radar sensor is low (it is difficult to accurately measure the height change of the human body), this application first determines the horizontal position of the human body through the horizontal axis coordinates of the radar sensor point cloud data combined with the high temperature area of thermal imaging. The highest FOV angle of the radar sensor is corrected by the highest FOV angle of the human body outline in the infrared image, thereby improving the accuracy of the relative height curve measured by the radar.
[0089] In addition, this application adopts a behavior recognition model based on temporal displacement (PPTSM), which has stronger temporal modeling capabilities, can more effectively learn motion patterns, and adapt to data segments of different lengths. Combined with 2000 segments of infrared + point cloud data collected in real environments for training, it can better adapt to the multimodal fusion data characteristics of this system.
[0090] In some specific embodiments, the sensor module also includes a sound sensor for collecting audio data. In this way, the sound sensor can capture the possible sound when the user falls and the cry for help of the elderly after the fall. Specifically, after inputting the constructed time series multimodal fusion data into the deep learning behavior recognition module and outputting the classification result of the fall behavior, it also includes: if the deep learning behavior recognition module determines that it is a fall behavior, calling the pre-trained speech classification model, and outputting the detection result of whether the audio data includes negative emotional speech.
[0091] If the audio detection result includes negative emotional speech, the alarm will be triggered directly; if the audio detection result does not include negative emotional speech, the thermal infrared imaging data will be sent to the predetermined contact and an alarm will be played. If no cancellation command is received within the specified time, the alarm will be triggered.
[0092] The speech classification model may be a CAM++-based speech classification model, which may use a data set of public and self-collected audio data, with a total of about 13,000 audios ranging from 4 to 10 seconds, to train a binary classification model of normal speech and negative emotion speech.
[0093] If the deep learning behavior recognition module determines that it is a fall behavior, and the voice classification model outputs a detection result that includes negative emotional speech, the alarm will be directly triggered. Otherwise, while sending the relevant infrared video to the emergency contact through the network, a ringtone will be played to require the user to press the alarm cancellation button. If the alarm cancellation command is not received within 1 minute, the alarm will be directly triggered.
[0094] In the embodiments of the present application, by combining a thermal infrared image sensor, a radar sensor, and a sound sensor, more accurate postures and motion states of the human body are obtained, and sounds that may exist when a fall event occurs are captured. The CAM++ voice classification model is used to identify groans and pain sounds, which provides an additional confirmation mechanism for fall detection, reduces the false alarm rate, and can more accurately determine whether a fall behavior exists. In addition, the intelligent alarm strategy adopted is more intelligent than the traditional direct alarm method, avoiding false alarms while improving the response efficiency of the alarm.
[0095] In addition, the present application also provides a multi-modal fall detection device, as Figure 4 shown in the structural block diagram of the multi-modal fall detection device provided by the present application. The device specifically includes a memory 41 and a processor 42. The memory 41 stores a computer program, and when the computer program is executed by the processor 42, it implements the multi-modal fall detection method according to any one of the above.
[0096] In addition, the present application also provides a multi-modal fall detection device, as Figure 5 shown in the structural block diagram of the multi-modal fall detection device provided by the present application. The device specifically includes a sensor module 51, a data processing module 52, and an alarm module 53.
[0097] Among them, the sensor module 51 may include a radar sensor and an infrared image sensor. This module is used for environmental perception and obtains data including an environmental thermal map, environmental temperature, the distance and azimuth of a person, and body surface temperature. It may also include a microphone for collecting audio data existing in the environment.
[0098] The data processing module 52 is configured to detect a target fall based on the radar point cloud data and the thermal infrared imaging data.
[0099] Specifically, the data processing module 52 obtains fall detection data, including radar point cloud data collected by the radar sensor and thermal infrared imaging data collected by the infrared image sensor; performs motion point extraction on the radar point cloud data based on time series analysis, filters out motion points whose positions change within a continuous time, and generates motion point cloud data; projects the motion point cloud data onto the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix; combines the radar point cloud projection matrix with the thermal infrared imaging data, calculates the three-dimensional coordinates of the target, and generates target height change data; when the target height change data meets a preset fall condition, inputs the constructed time series multi-modal fusion data into the deep learning behavior recognition module, and outputs a classification result of the fall behavior.
[0100] The alarm module 53 is used to call the pre-trained speech classification model when the deep learning behavior recognition module determines that the behavior is a fall, and output the detection result of whether the audio data includes negative emotional speech; if the audio detection result includes negative emotional speech, the alarm is directly triggered; if the audio detection result does not include negative emotional speech, the thermal infrared imaging data is sent to the predetermined contact and the alarm sound is played, and the alarm is triggered if the release instruction is not received within the specified time.
[0101] In addition, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the multimodal fall detection method according to any one of the above-mentioned methods is implemented.
[0102] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0103] The professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0104] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art.
[0105] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A multimodal fall detection method, characterized in that: include: Acquire fall detection data, where the fall detection data is collected by a sensor module, including radar point cloud data collected by a radar sensor and thermal infrared imaging data collected by an infrared image sensor; Extracting motion points from the radar point cloud data based on time series analysis, screening out motion points whose positions change in continuous time, and generating motion point cloud data; Projecting the motion point cloud data to a coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix; Combining the radar point cloud projection matrix with the thermal infrared imaging data, calculating the three-dimensional coordinates of the target and generating target height change data; When the target height change data meets the preset fall condition, the constructed time series multimodal fusion data is input into the deep learning behavior recognition module to output the classification result of the fall behavior; wherein the time series multimodal fusion data includes continuous frames of thermal infrared imaging data and radar point cloud projection matrix.
2. The multimodal fall detection method according to claim 1, characterized in that: The projecting the motion point cloud data to the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix comprises: Perform coordinate transformation on the filtered radar point cloud data and calculate the angle information of the radar point cloud data; Performing angle mapping based on the field of view angle of the infrared image sensor to calculate the projected pixel coordinates of the point cloud in the thermal infrared imaging data; Constructing a radar point cloud projection matrix of the same size as the thermal infrared imaging data, and marking projection pixel coordinates in the radar point cloud projection matrix; Wherein, the radar sensor and the infrared image sensor are installed coaxially.
3. The multimodal fall detection method according to claim 2, characterized in that: After projecting the motion point cloud data to the coordinate system corresponding to the thermal infrared imaging data to generate a radar point cloud projection matrix, the method further includes: Based on the radar point cloud projection matrix, a binary mask of the radar point cloud is constructed to generate a radar point cloud mask; Screening high temperature areas based on the thermal infrared imaging data to generate an infrared target mask; The intersection of the radar point cloud mask and the infrared target mask is calculated to generate multimodal fusion data.
4. The multimodal fall detection method according to claim 1, characterized in that: The step of combining the radar point cloud projection matrix with the thermal infrared imaging data to calculate the three-dimensional coordinates of the target and generate target height change data includes: Get the highest FOV angle measured by the radar sensor; Get the highest FOV angle of the target human body outline in the thermal infrared imaging data; The highest FOV angle of the target human body contour in the thermal infrared imaging data is used to correct the highest FOV angle measured by the radar sensor, obtain the highest point of the target, and generate target height change data.
5. The multimodal fall detection method according to any one of claims 1 to 4, characterized in that: The constructed time series multimodal fusion data is input into the deep learning behavior recognition module, and the classification results of the fall behavior are outputted, including: The deep learning behavior recognition module performs time series modeling based on the time displacement module and classifies the target behavior, and the classification of the target behavior includes falling, squatting, sitting down, and object dropping.
6. The multimodal fall detection method according to claim 5, characterized in that: Before inputting the constructed time series multimodal fusion data into the deep learning behavior recognition module, it also includes: Based on the target height change data, when it is detected that the time period from the start of the target height change to the end of the change is less than a preset threshold, radar point cloud data and thermal infrared image data of equal length are retrieved forward and backward for filling.
7. The multimodal fall detection method according to claim 5, characterized in that: The preset falling condition is: the target height change rate exceeds the preset rate; or the target height is lower than the preset height for a time longer than the preset time.
8. The multimodal fall detection method according to claim 5, characterized in that: After inputting the constructed time series multimodal fusion data into the deep learning behavior recognition module and outputting the classification result of the fall behavior, the method further includes: If the deep learning behavior recognition module determines that the behavior is a fall, the pre-trained speech classification model is called to output the detection result of whether the audio data includes negative emotional speech; If the audio detection result contains negative emotional speech, an alarm will be triggered directly; If the audio detection result does not include negative emotional speech, the thermal infrared imaging data will be sent to the predetermined contact and an alarm will be played. If no cancellation command is received within the specified time, the alarm will be triggered.
9. A multimodal fall detection device, characterized in that: The device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the multimodal fall detection method according to any one of claims 1 to 8 is implemented.
10. A multimodal fall detection device, characterized in that: include: A sensor device, an alarm device, and a multimodal fall detection device according to claim 9; The sensor module includes a radar sensor and an infrared image sensor; the radar sensor is configured to collect radar point cloud data; the infrared image sensor is configured to collect thermal infrared imaging data; The multimodal fall detection device is configured to detect a target fall based on the radar point cloud data and the thermal infrared imaging data; The alarm module is configured to receive the detection result of the multi-modal fall detection device and execute an alarm reminder.
Citation Information
Cited By
Multi-modal human body situation awareness method based on data fusion, terminal equipment and storage medium
CN120472542A
Multi-modal human body posture perception method based on data fusion, terminal device and storage medium
CN120472542B
Human body tumble radar perception correction method and system in combination with multi-modal data
CN120802245A
Line inspection channel crowd behavior identification method and system based on multi-modal data fusion
CN120853224A
Edge multi-mode grading detection-based falling early warning method for old people
CN120932368A