Multi-modal fusion-based living perception method and system

By employing a multimodal fusion-based life perception method and utilizing multiple sensors to generate a dual verification mechanism, the problem of traditional single sensors being unable to accurately monitor the living conditions of the elderly is solved, enabling comprehensive perception and timely early warning of the elderly.

WO2026091847A1PCT designated stage Publication Date: 2026-05-07SHENZHEN TOPTECH MANUFACTORING CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHENZHEN TOPTECH MANUFACTORING CO LTD
Filing Date
2025-09-03
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Traditional single-sensor methods are insufficient to fully capture the living conditions and behavioral characteristics of the elderly, making it difficult to accurately identify sudden accidents such as falls, thus hindering timely assistance.

Method used

A multimodal fusion-based life perception method is adopted, which acquires multimodal sensor information through multiple different types of sensors, generates a first perception model and a second perception model, performs preliminary and secondary predictions, and generates early warning information.

Benefits of technology

It enables comprehensive perception and monitoring of the living conditions of the elderly, improves the reliability of the early warning system, ensures timely triggering of early warnings in emergency situations, and provides rapid rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025118593_07052026_PF_FP_ABST
    Figure CN2025118593_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a multi-modal fusion-based living perception method and system. The method comprises: first obtaining multi-modal sensor information by means of a plurality of different types of sensors, wherein the sensors are deployed on the basis of living habit information and living environment information of a user; then generating a first perception model and a second perception model on the basis of the deployment positions and types of the sensors; inputting the multi-modal sensor information into the first perception model to obtain a preliminary feature and a preliminary prediction result; then inputting the information and the preliminary feature into the second perception model to obtain a secondary prediction result; and generating early warning information on the basis of the degree of association between the preliminary prediction result and the secondary prediction result. In embodiments of the present application, by fusing multi-mode sensor information, the living state of a user can be comprehensively perceived and monitored, and by dual prediction, the behavior state of the user, particularly emergency situations such as falling down, can be more accurately determined, so that an early warning mechanism is triggered promptly.
Need to check novelty before this filing date? Find Prior Art

Description

A multimodal fusion method and system for life perception Technical Field

[0001] This application relates to, but is not limited to, the field of smart home-based elderly care technology, and particularly to a multimodal fusion life perception method and system. Background Technology

[0002] Monitoring the living conditions of elderly people living alone is a key concern and challenge for their children and the wider elderly community. Due to declining physical function, the likelihood of falls increases significantly, and the resulting injuries, such as fractures, soft tissue contusions, and psychological trauma, are serious concerns. If a fallen person does not receive timely assistance, their injuries may worsen. Traditional monitoring methods often rely on single sensors or devices, but these methods have significant limitations. For example, they may not comprehensively capture the user's living conditions and behavioral characteristics, or accurately identify sudden accidents such as falls. Therefore, a more advanced, accurate, and comprehensive method is urgently needed to perceive and monitor the living conditions of elderly people living alone. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0004] This application provides a multimodal fusion life perception method that can comprehensively perceive and monitor the user's life status and trigger an early warning mechanism in a timely manner.

[0005] In a first aspect, embodiments of this application provide a multimodal fusion-based life perception method, including:

[0006] Multimodal sensor information is acquired through multiple different types of sensors, which are deployed based on the user's lifestyle habits and living scenarios.

[0007] Based on the deployment location and type of the sensors, a first perception model and a second perception model are generated, wherein the computational complexity of the first perception model is lower than that of the second perception model.

[0008] The multimodal sensor information is input into the first perception model to obtain preliminary features and preliminary prediction results;

[0009] The multimodal sensor information and the preliminary features are input into the second perception model to obtain the secondary prediction result;

[0010] Early warning information is generated based on the correlation between the preliminary prediction results and the secondary prediction results.

[0011] In conjunction with the first aspect, in one embodiment of this application, the multimodal sensor information includes living behavior state data and continuous vital sign data, wherein the living behavior state data is provided by a specified image and the continuous vital sign data is provided by specified point cloud data.

[0012] In conjunction with the first aspect, in one embodiment of this application, the step of inputting the multimodal sensor information into the first perception model to obtain preliminary features and preliminary prediction results includes: performing denoising and normalization processing on the life behavior state data and the continuous vital sign data to obtain first multimodal behavior state data; inputting the first multimodal behavior state data into the first perception model; extracting multimodal state features from the first multimodal behavior state data; fusing the multimodal state features to obtain preliminary features; and obtaining preliminary prediction results based on the state feature representation information.

[0013] In conjunction with the first aspect, in one embodiment of this application, the acquisition of multimodal sensor information includes: when the preliminary prediction result indicates that the target user's perception is normal, acquiring first life behavior state data of the target user; when the preliminary prediction result indicates that the target user's perception is abnormal, acquiring second life behavior state data of the target user, and generating preliminary warning information based on the second life behavior state data; the first life behavior state data is provided by a first image, the second life behavior state data is provided by a second image, and the resolution of the second image is greater than the resolution of the first image.

[0014] In conjunction with the first aspect, in one embodiment of this application, the step of inputting the multimodal sensor information and the preliminary features into the second perception model to obtain a secondary prediction result includes: combining the denoised and normalized life behavior state data and the continuous vital signs data, as well as the preliminary features, to form second multimodal behavior state data; and inputting the second multimodal behavior state data into the second perception model for prediction to obtain a secondary prediction result.

[0015] In conjunction with the first aspect, in one embodiment of this application, generating early warning information based on the correlation between the preliminary prediction result and the secondary prediction result includes: when both the preliminary prediction result and the secondary prediction result show abnormalities, the early warning information is a first abnormality early warning information; when the secondary prediction result shows an abnormality and the preliminary prediction result shows normality, the early warning information is a second abnormality early warning information; when the secondary prediction result shows normality and the preliminary prediction result shows an abnormality, the early warning information is a third abnormality early warning information; the abnormality intensity of the first abnormality early warning information is greater than the abnormality intensity of the second abnormality early warning information, and the abnormality intensity of the second abnormality early warning information is greater than the abnormality intensity of the third abnormality early warning information.

[0016] In conjunction with the first aspect, in one embodiment of this application, generating a first perception model and a second perception model based on the deployment location and type of the sensor includes: constructing an initial first perception model and an initial second perception model based on the deployment location and type of the sensor; performing model training on the initial first perception model to obtain a first perception model; and performing model training on the initial second perception model to obtain a second perception model.

[0017] Secondly, embodiments of this application provide a multimodal fusion-based life sensing system, including a terminal, a server, and multiple sensors, wherein the multiple sensors, the terminal, and the server communicate with each other; the multiple sensors are deployed according to life habit information and life scene information; the terminal is used to acquire multimodal sensor information, construct an initial first sensing model according to the deployment location and type of the sensors, train the initial first sensing model to obtain a first sensing model, input the multimodal sensor information into the first sensing model to obtain preliminary features and preliminary prediction results; the server is used to acquire multimodal sensor information and the preliminary features, construct an initial second sensing model according to the deployment location and type of the sensors, train the initial second sensing model to obtain a second sensing model, input the multimodal sensor information and the preliminary features into the second sensing model to obtain secondary prediction results; and generate early warning information based on the correlation between the preliminary prediction results and the secondary prediction results.

[0018] In conjunction with the second aspect, in one embodiment of this application, the types of sensors include color cameras, depth cameras, millimeter-wave radars, and infrared cameras.

[0019] In conjunction with the second aspect, in one embodiment of this application, the system is provided with a communication module for sending information to a designated user based on the warning information.

[0020] This application provides a multimodal fusion-based life sensing method. First, multimodal sensor information is acquired through multiple sensors of different types, deployed according to the user's lifestyle habits and living scenarios. This customized configuration enables comprehensive and multi-angle collection of the user's life status and behavioral characteristics data, effectively avoiding monitoring blind spots caused by the limited perspective or function of a single sensor, ensuring a comprehensive perception of the user's life status. Next, a first perception model and a second perception model are generated based on the deployment location and type of the sensors. The multimodal sensor information is input into the first perception model to obtain preliminary features and preliminary prediction results; these are then input into the second perception model along with the preliminary features to obtain secondary prediction results. Subsequently, warning information is generated based on the correlation between the preliminary prediction results and the secondary prediction results. This dual verification mechanism significantly improves the reliability of the warning system. Due to the real-time nature of multimodal sensor information, warning information can be quickly transmitted to emergency contacts, ensuring that the user can receive rescue as soon as possible. In contrast, traditional single-sensor methods may lead to inaccurate identification of sudden unexpected behaviors such as falls due to insufficient data or misjudgment. The embodiments of this application, by fusing multimodal sensor information and combining preliminary features and secondary prediction results, can more accurately determine the user's behavioral state, especially in emergencies such as falls, thereby triggering an early warning mechanism in a timely manner. Attached Figure Description

[0021] Figure 1 is a flowchart of the multimodal fusion life perception method provided in an embodiment of this application;

[0022] Figure 2 is a schematic diagram of the basic framework of the multimodal fusion life perception method provided in the embodiments of this application;

[0023] Figure 3 is a schematic diagram of a sensor arrangement scheme for a home sensing space provided in a specific example of this application;

[0024] Figure 4 is a schematic flowchart of step 130 provided in an embodiment of this application;

[0025] Figure 5 is a schematic diagram of a high-resolution enabling strategy for a color camera provided in an embodiment of this application;

[0026] Figure 6 is a schematic diagram of the judgment and analysis mechanism for abnormal situations provided in the embodiments of this application;

[0027] Figure 7 is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0029] It should be noted that although the flowchart shows a logical order, in some cases, the steps shown or described may be performed in a different order than that shown in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the structures, proportions, sizes, etc., depicted in the drawings are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the implementation conditions of this application. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to size, without affecting the effects and purposes achieved by this application, should still fall within the scope of the technical content disclosed in this application. Similarly, the terms such as "upper," "lower," "left," "right," "middle," and "one" used in this specification are only for clarity of description and are not used to limit the scope of implementation of this application. Changes or adjustments in their relative relationships, without substantially altering the technical content, should also be considered within the scope of implementation of this application.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0031] The perception of the living conditions of elderly people living alone is a key focus and challenge for their children and the broader society. Elderly people spend most of their time at home, and in such environments, they often struggle to receive timely assistance in case of accidents, posing a significant threat to their health and well-being. A user's daily activities at home, such as walking, sitting, bathing, cooking, and using the toilet, involve not only individual behavioral and physiological characteristics but, more importantly, the ability to monitor sudden accidents, especially falls. Falls are a common and dangerous accident for the elderly. Due to declining physical function, the likelihood of falls increases significantly, and the resulting injuries, such as fractures, soft tissue contusions, and psychological trauma, are serious. If a fallen user does not receive timely assistance, their injuries may worsen. Traditional monitoring methods often rely on single sensors or devices, but these methods have significant limitations. For example, they may not comprehensively capture a user's living conditions and behavioral characteristics, or accurately identify sudden accidents such as falls. Therefore, a more advanced, accurate, and comprehensive method is urgently needed to perceive and monitor the living conditions of elderly people.

[0032] In view of this, embodiments of this application provide a multimodal fusion-based life sensing method and system. First, multimodal sensor information is acquired through multiple sensors of different types, deployed according to the user's lifestyle habits and living scenarios. This customized configuration enables comprehensive and multi-faceted collection of the user's living status and behavioral characteristics data, effectively avoiding blind spots caused by the limited perspective or function of a single sensor, ensuring a comprehensive perception of the user's living status. Next, a first sensing model and a second sensing model are generated based on the sensor deployment location and type. The multimodal sensor information is input into the first sensing model to obtain preliminary features and preliminary prediction results; these are then input together with the preliminary features into the second sensing model to obtain secondary prediction results. Subsequently, warning information is generated based on the correlation between the preliminary prediction results and the secondary prediction results. This dual verification mechanism significantly improves the reliability of the warning system. Due to the real-time nature of multimodal sensor information, warning information can be quickly transmitted to emergency contacts, ensuring that the user receives immediate assistance. In contrast, traditional single-sensor methods may lead to inaccurate identification of sudden unexpected behaviors such as falls due to insufficient data or misjudgment. The embodiments of this application, by fusing multimodal sensor information and combining preliminary features and secondary prediction results, can more accurately determine the user's behavioral state, especially in emergencies such as falls, thereby triggering an early warning mechanism in a timely manner.

[0033] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0034] Referring to Figure 1, which is a flowchart of a multimodal fusion life perception method provided in an embodiment of this application, the method includes, but is not limited to, steps 110 to 150.

[0035] Step 110: Acquire multimodal sensor information through multiple different types of sensors, wherein the sensors are deployed based on the user's living habits and living scenarios;

[0036] Step 120: Generate a first sensing model and a second sensing model based on the deployment location and type of the sensors;

[0037] Step 130: Input the multimodal sensor information into the first perception model to obtain preliminary features and preliminary prediction results;

[0038] Step 140: Input the multimodal sensor information and preliminary features into the second perception model to obtain the secondary prediction result;

[0039] Step 150: Generate early warning information based on the correlation between the preliminary prediction results and the secondary prediction results.

[0040] Steps 110 to 150 will be described in detail below.

[0041] In one feasible embodiment, user life status perception information typically encompasses two main aspects: behavioral characteristics and physiological characteristics. Behavioral characteristics mainly include activities within the residence, such as walking, sitting, lying down, personal hygiene habits (bathing), cooking, and toileting; while physiological characteristics mainly include core physiological indicators such as body temperature, respiratory rate, heart rate, sleep quality, and physical activity level.

[0042] In one feasible embodiment, to accurately and efficiently acquire user's perceived life status information, various sensors can be flexibly deployed based on the user's lifestyle habits and living scenarios. These sensors include, but are not limited to, color cameras, depth cameras, millimeter-wave radar devices, and infrared cameras. Specifically, for the user's dynamic gait characteristics (e.g., walking patterns within the home) and static posture characteristics (e.g., posture when sitting or lying down), color cameras can be deployed for periodic image capture to record and analyze these characteristics in detail. Simultaneously, the addition of depth cameras further enriches the information dimensions of color images, providing depth data to more accurately assess the gait parameters and static posture characteristics of the elderly. Millimeter-wave radar devices, as advanced detection equipment operating in the 30 to 300 GHz frequency band, can accurately measure key information such as the distance, speed, and direction of target objects by transmitting and receiving millimeter-wave signals. Their high-resolution characteristics enable them to identify minute targets. More importantly, since human life activities (such as movement, breathing, and heartbeat) cause changes in the reflected echo sequence, millimeter-wave radar can accurately capture and analyze these vital parameters. Combined with advanced point cloud human imaging algorithms, millimeter-wave radar can achieve comprehensive coverage of respiration, heartbeat, sleep status, body activity, and multi-target monitoring. Infrared cameras, on the other hand, can capture infrared radiation invisible to the human eye. It's understood that any object above thermodynamic zero radiates infrared radiation. When the ambient temperature is below normal human body temperature (approximately 37 degrees Celsius), human skin releases significant infrared radiation energy, the intensity and distribution of which directly reflect the object's temperature state. When abnormal pathological changes occur in a part of the body, they are often accompanied by subtle changes in the temperature field. By capturing these changes, infrared cameras can quickly and clearly reveal differences in body temperature caused by different pathogenic factors, providing important clues for health monitoring.

[0043] In one feasible embodiment, multimodal sensor information (i.e., life status perception information) is acquired comprehensively through various types of sensors deployed at multiple locations, encompassing user life behavior status data and continuous vital sign data. Life behavior status data records the user's dynamic gait characteristics and static posture characteristics, primarily provided by specific image sequences captured by color and depth cameras, which can intuitively demonstrate the user's activity patterns and postural changes in daily life. Simultaneously, continuous vital sign data records key physiological indicators such as the user's movement, respiration, and heart rate, which can be acquired through precise point cloud data from millimeter-wave radar devices.

[0044] In a feasible embodiment, both the first and second sensing models can be network models built based on convolutional neural networks. Specifically, the first sensing model is a lightweight model, which is more computationally efficient than the second sensing model (i.e., the large multimodal model), meaning it requires fewer computing resources and processes information more quickly. In practical applications, the first sensing model can be deployed on a local terminal, quickly responding to and processing information from multimodal sensors, making preliminary and rapid judgments, and providing concise decision analysis. The second sensing model, due to its larger computational load and stronger analytical capabilities, can be deployed on a server. It can further analyze and confirm the preliminary judgments of the first sensing model, or supplement and improve the sensed information. This deployment strategy not only optimizes the allocation of computing resources but also ensures the efficiency and accuracy of data processing, providing users with a more intelligent and precise home life experience.

[0045] In a feasible embodiment, a first sensing model and a second sensing model can be constructed based on the deployment location and type of the sensors. The specific process is as follows: First, based on the deployment location and category of the sensors, an initial first sensing model (i.e., a lightweight model) and a second sensing model (i.e., a large model) are constructed. Both models have the ability to quickly respond to and process information from multimodal sensors and can provide corresponding prediction results. It is worth mentioning that the initial second sensing model can also integrate preliminary features from the initial first sensing model, thereby providing more accurate predictions. Subsequently, the initial first sensing model is trained in depth using relevant training datasets, and the model parameters are optimized using algorithms to ensure that it can accurately extract valuable information from sensor data and provide corresponding sensing results. After this series of training, an optimized first sensing model can be obtained. Similarly, the initial second sensing model is also comprehensively trained to obtain a high-performance second sensing model. Through these two trained sensing models, accurate processing and analysis of data from different sensors can be achieved, providing users with more intelligent and accurate sensing services.

[0046] The following example illustrates the process of generating a first and second perception model based on the deployment location and type of sensors.

[0047] Imagine a smart home environment where a color camera and an infrared camera are installed in the living room. The color camera captures users' facial features and movement patterns, while the infrared camera monitors indoor temperature changes and human thermal radiation to detect abnormal body temperatures. A depth camera and millimeter-wave radar are deployed in the bedroom. The depth camera precisely measures the bedroom's spatial layout and furniture placement, while the millimeter-wave radar monitors vital signs such as breathing and heartbeat in the person in bed. For the color camera in the living room, a lightweight model (the first perception model) can be generated. This model can quickly process the low-resolution and high-resolution RGB images from the color camera, identify the facial features of users in the area, and make a preliminary judgment about their activity level (e.g., walking, stationary). For the infrared camera, the lightweight model can analyze temperature data to make a preliminary judgment about abnormal body temperatures. On a server, a larger model (the second perception model) is generated for the color and infrared cameras in the living room. This model receives the preliminary features from the lightweight model and combines them with its own high-resolution images and temperature data for more in-depth analysis. For example, it can more accurately identify the users in the area, analyze their behavioral patterns, and precisely detect the specific location and degree of abnormal body temperatures. For depth cameras and millimeter-wave radar in the bedroom, corresponding lightweight and large models can be generated respectively. The lightweight model can quickly process the depth images and vital sign data from the millimeter-wave radar for preliminary anomaly detection; while the large model can combine preliminary features, vital sign data, and depth images for more in-depth sleep monitoring and anomaly detection. In this way, perception models suitable for different scenarios and needs can be generated based on the deployment location and type of sensors, enabling the perception and monitoring of users' living conditions.

[0048] Referring to Figure 2, which is a schematic diagram of the basic framework of the multimodal fusion life perception method provided in this application embodiment, multiple sensors, including color cameras, depth cameras, millimeter-wave radar, and infrared cameras, are first deployed in a living space. These sensors capture and provide multimodal sensor information, specifically covering color images, depth images, point cloud data, and infrared images. Subsequently, convolutional neural networks (first perception model and second perception model) are used to process and analyze this multimodal information in depth. Each network makes independent decision outputs for its corresponding modality data. Based on this, the decision results from different modalities are further integrated, and through comprehensive judgment, information about the elderly person's physical condition is finally output. This method not only fully utilizes the complementary advantages of multimodal data but also achieves accurate extraction and output of life condition perception information through the powerful processing capabilities of convolutional neural networks, providing strong support for subsequent monitoring, early warning, and response.

[0049] In one feasible embodiment, the deployment location and type of sensors can reflect differentiated strategies and priorities based on users' lifestyle habits and specific living scenario characteristics. For the specific needs of different spatial areas in the home environment, the sensor configuration scheme can be flexibly adjusted to ensure that each area receives the most suitable monitoring methods, thereby providing users with more accurate and comprehensive home life status perception services. As shown in Figure 3, taking the home life scenario of the elderly as an example, in the living room area, color cameras, millimeter-wave radar, and depth cameras can be deployed to capture and record the elderly's activity images and movement details from all angles; in the bedroom, infrared cameras and millimeter-wave radar are configured to effectively monitor the elderly's sleep status and physiological characteristics even at night, ensuring the continuity and accuracy of monitoring; in the kitchen, the combined use of infrared cameras and color cameras can more accurately identify the elderly's cooking behavior and food processing, improving the detail of monitoring; as for the toilet space, considering privacy protection and practicality, millimeter-wave radar can be used for monitoring. The rich multimodal sensor information acquired in each space is further processed and analyzed using convolutional neural networks (first perception model and second perception model). When multiple sensors are deployed in a given area, decision-making results from different modalities can be integrated to accurately output the physical condition information of the elderly in that area through comprehensive judgment. Ultimately, the physical condition information of the elderly generated from all areas will be integrated to form a complete record of home-based physical condition indicators. Anomaly detection and notification will be performed based on these indicator records to ensure that the elderly receive timely and effective care.

[0050] In a feasible embodiment, the first sensing model, as a lightweight model, can perform a series of operations after receiving information from the multimodal sensor, including but not limited to data preprocessing, modal feature extraction, or preliminary identification analysis, thereby quickly providing simple conclusions such as decision analysis or anomaly alarms. Therefore, as shown in Figure 4, step 130 involves the process of inputting multimodal sensor information into the first sensing model to obtain preliminary features and preliminary prediction results, including but not limited to steps 410 to 430.

[0051] Step 410: Denoise and normalize the life behavior state data and continuous vital sign data to obtain the first multimodal behavior state data, and input the first multimodal behavior state data into the first perception model.

[0052] Step 420: Extract multimodal state features from the first multimodal behavioral state data, and fuse the multimodal state features to obtain preliminary features;

[0053] Step 430: Obtain preliminary prediction results based on preliminary feature representation information.

[0054] In a feasible embodiment, in step 410, the purpose of denoising is to remove outliers or noise signals from the data caused by factors such as sensor noise and environmental interference, thereby improving the accuracy and reliability of the data. Normalization aims to transform data of different dimensions and scales to the same scale range to facilitate model processing and comparison. Normalization helps to accelerate the training speed of the model and improve the predictive performance of the model. The first multimodal behavioral state data is obtained, which has been preprocessed (denoising and normalization, etc.) and is suitable as input to the first perception model.

[0055] In a feasible embodiment, in step 420, key information or features are extracted from the data of each modality using machine learning or deep learning algorithms. These features can reflect the user's life behavior status and vital signs. The extracted multimodal features are then fused to form a more comprehensive and richer preliminary feature set. The fusion strategy may include weighted averaging, maximum value selection, feature concatenation in deep learning, or attention mechanisms, etc.

[0056] In one feasible embodiment, in step 430, a preliminary prediction result can be derived based on the preliminary feature representation information. For example, if the preliminary features clearly indicate that the user is currently in a falling posture, then the preliminary prediction result will be determined as an abnormal state. This prediction result is based on in-depth analysis and accurate identification of user behavior state data, aiming to provide users with timely safety warnings or take corresponding assistance measures.

[0057] In a feasible embodiment, through steps 410 to 430, the first sensing model can receive multimodal sensor information and perform data preprocessing, feature extraction, and fusion to ultimately obtain preliminary features and preliminary prediction results. These results can provide strong support for subsequent decision analysis, anomaly alarms, or further processing. Meanwhile, since the first sensing model typically employs a lightweight design, it can reduce computational complexity and resource consumption while ensuring performance.

[0058] In a feasible embodiment, to comprehensively ensure user privacy and effective recording of emergency events, this application embodiment also provides a high-resolution activation strategy for a color camera. As shown in Figure 5, in normal mode, when the preliminary prediction result indicates a normal state, suggesting the target user is in a normal activity state, the color camera will use image downsampling technology to generate a low-resolution image for basic body feature recognition, including preliminary identification of dynamic gait and static posture. This approach maintains user privacy while ensuring appropriate information display, facilitating intuitive display on terminal devices or sending status notifications to associated child devices, allowing them to understand the user's situation at any time. However, once the preliminary prediction result indicates that the target user may be in an abnormal state, such as an emergency involving unauthorized intrusion by a stranger or abnormal displacement of an object when no one is present, the color camera will immediately switch to high-resolution image data stream mode. This switch aims to accurately capture and display the details of the abnormal event, providing crucial visual evidence for subsequent in-depth analysis and confirmation of the event's nature. Subsequently, an alarm mechanism is simultaneously activated on multiple terminals (such as smartphones and other smart devices) to notify of the anomaly, and the captured high-resolution image is securely stored on local devices or cloud servers, preparing for possible subsequent investigations and evidence preservation. In this context, the specific acquisition process of multimodal sensor information in step 110 covers at least the following aspects: First, based on the preliminary prediction of the first perception model, if the target user is in a normal state, a first dataset reflecting their living behavior is collected and processed. This dataset mainly comes from the first image after resolution downscaling. Conversely, if the preliminary prediction shows that the target user may have an anomaly, a more detailed second dataset is collected and processed to generate preliminary warning information. This data comes from a high-resolution second image, which has a significantly higher resolution than the first image, to provide more delicate and accurate visual details.

[0059] In a feasible embodiment, the specific steps for inputting multimodal sensor information and preliminary features into a second perception model to obtain more accurate secondary prediction results include: First, denoising and normalizing the life behavior state data and continuous vital sign data to improve data quality and consistency; then, integrating these processed data with the preliminary features to form second multimodal behavior state data. This dataset not only covers details of the user's life behavior but also incorporates preliminary feature analysis results. Next, the second multimodal behavior state data is input into the second perception model. This model, based on advanced algorithms and deep learning technology, can comprehensively analyze multiple data types, thereby more accurately predicting the user's state. It is worth noting that the secondary prediction results often have higher accuracy and reliability than the preliminary prediction results, providing strong support for real-time monitoring and early warning of user states.

[0060] In a feasible embodiment, during the process of generating early warning information based on the correlation between the preliminary prediction result and the secondary prediction result, when both the preliminary prediction result and the secondary prediction result show anomalies, the early warning information is a first anomaly early warning information; when the secondary prediction result shows anomalies and the preliminary prediction result shows normal, the early warning information is a second anomaly early warning information; when the secondary prediction result shows normal and the preliminary prediction result shows anomalies, the early warning information is a third anomaly early warning information. The anomaly intensity of the first anomaly early warning information is greater than that of the second anomaly early warning information, and the anomaly intensity of the second anomaly early warning information is greater than that of the third anomaly early warning information. Furthermore, processing strategies for strong anomalies, anomalies, and weak anomalies can be formulated based on the early warning information to ensure that target groups associated with the target user, such as children, relatives, and friends, can quickly and accurately learn about any potential anomalies in the target user's situation. Specifically, when the prediction results of both perception models show anomalies, it is determined to be a strong anomaly, and an emergency notification can be immediately sent to the target group online. If only the secondary prediction shows anomalies and the preliminary prediction is normal, it is determined to be a general anomaly, and an SMS notification can be sent initially; if confirmation is not received within a specified time, a second reminder will be sent. Conversely, if the second prediction is normal but the initial prediction is abnormal, it is judged as a weak anomaly, and only one SMS notification needs to be sent to reduce unnecessary disturbance.

[0061] In a feasible embodiment, the judgment and analysis mechanism for abnormal situations in the collaborative strategy of the first and second perception models is shown in Figure 6. The first perception model can efficiently process multimodal sensor information (such as images from a color camera, point clouds from millimeter-wave radar, etc.), perform data preprocessing, modal feature extraction, and preliminary identification and analysis tasks, and quickly provide preliminary prediction results (such as decision analysis or anomaly alarms). However, its preliminary prediction results still need to be combined with the in-depth analysis results of the second perception model for confirmation or supplementary perception analysis to ensure the accuracy and completeness of the prediction results. Specifically, when both the first and second perception models detect an anomaly, the system immediately determines it as a strong anomaly and immediately notifies the user's children, relatives, and other key contacts through the background online method. If the second perception model determines it as an anomaly, while the first perception model determines it as normal, the system treats it as a general anomaly, sends an SMS notification to the relevant contacts through the background, and requires confirmation within a specified time. If confirmation is not received within the specified time, the system will send another SMS reminder to ensure that the abnormal situation is addressed in a timely manner. Conversely, if the second perception model determines it to be normal, while the first perception model determines it to be abnormal, the system treats it as a weak anomaly and only sends a single SMS notification to the relevant contacts in the background. This tiered response strategy ensures timely reporting of abnormal situations while avoiding unnecessary interference to users caused by frequent false alarms. Furthermore, when both the first and second perception models determine it to be normal, it indicates that the user is in a healthy state, and the system does not need to perform any notification operations. Through this collaborative strategy and intelligent judgment and analysis mechanism, the system can provide users with more accurate and efficient life perception services, ensuring that users receive timely and effective attention and assistance in emergency situations.

[0062] The following example illustrates the multimodal fusion life perception method of this application.

[0063] In smart home applications for elderly health monitoring, various types of sensors are deployed to monitor the health status and lifestyle habits of elderly individuals. These sensors include: Color cameras: used to capture daily activities such as walking, sitting, lying down, and facial expressions; Depth cameras: generate 3D images by measuring the distance between objects and the camera, helping to more accurately identify the elderly person's movements and postures; Millimeter-wave radar devices: using human imaging algorithms based on millimeter-wave radar point clouds, they can monitor breathing, heart rate, sleep, movement, and multiple targets; Infrared cameras: utilizing infrared radiation imaging, they can capture the elderly person's body temperature distribution and activity in completely dark environments, helping to detect potential fall risks. These sensors are carefully deployed based on the elderly person's lifestyle habits (such as activity time and activity area) and living environment (such as bedroom, living room, and kitchen) to ensure comprehensive and accurate collection of their life information. Based on the sensor deployment location and type, machine learning techniques can be used to generate two perception models: a first perception model and a second perception model. The first perception model mainly focuses on extracting preliminary features and preliminary prediction results from information obtained from various types of sensors, such as the basic movement recognition analysis of the elderly person. The second sensing model combines information from multiple types of sensors with the preliminary features from the first sensing model for deeper fusion and prediction. Specifically, after inputting multimodal sensor information into the first sensing model, the model can automatically identify preliminary features such as the elderly person's dynamic gait and static posture, and make preliminary predictions based on these features, such as determining whether the elderly person is in a normal activity state. Next, the multimodal sensor information and preliminary features are input into the second sensing model for secondary prediction. The second sensing model can comprehensively consider more dimensions of information, such as the elderly person's movement trajectory and body temperature changes, thus obtaining more accurate and comprehensive prediction results, such as determining whether the elderly person has a risk of falling or poor sleep quality. Finally, based on the correlation between the preliminary and secondary prediction results, corresponding early warning information can be generated. For example, if both models predict a risk of falling, a strong anomaly warning will be immediately triggered, and an emergency notification will be sent to children or relatives via a mobile app. If only the second sensing model predicts an anomaly, while the first sensing model does not detect an anomaly, an anomaly warning will be sent, requiring confirmation within a specified time. If the predictions from the two models are inconsistent, a weak anomaly alert is sent for further observation and monitoring. This multimodal fusion-based approach to life perception allows for more comprehensive and accurate monitoring of the health status of elderly people at home, providing strong protection for their safety.

[0064] Furthermore, this application embodiment also provides a multimodal fusion-based life sensing system. This system integrates a terminal, a server, and multiple sensors, achieving seamless communication among the three. Based on user lifestyle and scene information, after deploying multiple types of sensors in various independent scenes within the home, a first sensing model and a second sensing model can be constructed according to the type and deployment location of the sensors. The first sensing model is deployed on the terminal, while the second sensing model is deployed on the server. It should be noted that both sensing models can be built using advanced neural network architectures. Specifically, the first sensing model adopts a lightweight design approach to maximize computational efficiency; while the second sensing model, as a large multimodal model, has more powerful analytical capabilities. After training and parameter tuning these two models separately, a model with superior performance can be obtained. In actual operation, the terminal first collects multimodal information from various sensors. Subsequently, the first sensing model can quickly and accurately analyze this multimodal information, extract key features, and provide preliminary prediction results. Simultaneously, the server also receives this multimodal information synchronously. The second sensing model not only receives preliminary feature information from the first sensing model but also integrates its own acquired multimodal information to perform more detailed and in-depth secondary predictive analysis, thereby further improving the accuracy of the prediction results. Through the system's built-in communication module, relevant notifications can be accurately pushed to designated users. During this process, the system can intelligently generate early warning information based on the correlation between the preliminary and secondary prediction results. This mechanism enables the system to issue timely alerts when users' life conditions become abnormal, ensuring their safety and health.

[0065] In one feasible embodiment, the types of sensors employed include color cameras, depth cameras, millimeter-wave radar, and infrared cameras to construct a comprehensive life-sensing system. The color camera is primarily responsible for capturing the user's gait characteristics in low-resolution mode (i.e., single-frame image resolution not exceeding 480*640, with RGB three-color channels). These characteristics encompass walking posture and behavioral patterns. Additionally, it periodically records the user's static characteristics, especially postures during falls. Once the system detects an intruder or an unusual situation, the color camera immediately switches to high-resolution mode (HD or UHD) to clearly capture and record the intruder's illegal behavior. The depth camera complements the color camera, focusing on acquiring depth images and accurately measuring the distance of real-world objects to the camera, thus providing accurate physical size information. This function not only enriches the depth information lacking in the color camera but also further enhances the detection accuracy of user gait and static posture characteristics. The millimeter-wave radar, through point cloud human imaging algorithms, achieves accurate capture of breathing, heartbeat, sleep status, movement, and multi-target monitoring. Infrared cameras can promptly detect and clearly display temperature changes in the human body caused by various factors. It's worth emphasizing that the modal information collected by these sensors can be automatically uploaded to the server and terminal at preset time intervals (e.g., every 10 seconds, 20 seconds) for real-time online sensing and analysis. This mechanism ensures that the system can monitor the user's living status around the clock without blind spots, and in the event of any sudden emergencies (such as falls, abnormal body temperature, abnormal sleep heart rate or breathing), it immediately notifies the user's children, relatives, and other key contacts via calls, text messages, or other communication methods, enabling them to take immediate action to ensure the user's safety and health.

[0066] It should be noted that since the multimodal fusion life perception system of this embodiment can realize the multimodal fusion life perception method of the previous embodiment, the multimodal fusion life perception system of this embodiment and the multimodal fusion life perception method of the previous embodiment have the same technical principle and the same beneficial effect. In order to avoid repetition, it will not be described again here.

[0067] Referring to Figure 7, this application also discloses an electronic device 700, which includes:

[0068] At least one processor 701;

[0069] At least one memory 702 is used to store at least one program;

[0070] When at least one program is executed by at least one processor 701, the life-aware method of multimodal fusion as described above is implemented.

[0071] This application also discloses a computer-readable storage medium storing a processor-executable computer program, which, when executed by a processor, is used to implement the aforementioned multimodal fusion life perception method.

[0072] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal fusion method for life perception, characterized in that, include: Multimodal sensor information is acquired through multiple different types of sensors, which are deployed based on the user's lifestyle habits and living scenarios. Based on the deployment location and type of the sensors, a first perception model and a second perception model are generated, wherein the computational complexity of the first perception model is lower than that of the second perception model. The multimodal sensor information is input into the first perception model to obtain preliminary features and preliminary prediction results; The multimodal sensor information and the preliminary features are input into the second perception model to obtain the secondary prediction result; Early warning information is generated based on the correlation between the preliminary prediction results and the secondary prediction results.

2. The multimodal fusion life perception method according to claim 1, characterized in that, The multimodal sensor information includes life behavior status data and continuous vital signs data. The life behavior status data is provided by a specified image, and the continuous vital signs data is provided by a specified point cloud data.

3. The multimodal fusion life perception method according to claim 2, characterized in that, The step of inputting the multimodal sensor information into the first perception model to obtain preliminary features and preliminary prediction results includes: performing denoising and normalization processing on the life behavior state data and the continuous vital sign data to obtain first multimodal behavior state data; inputting the first multimodal behavior state data into the first perception model; extracting multimodal state features from the first multimodal behavior state data; fusing the multimodal state features to obtain preliminary features; and obtaining preliminary prediction results based on the state feature representation information.

4. The multimodal fusion life perception method according to claim 3, characterized in that, The acquisition of multimodal sensor information includes: when the preliminary prediction result indicates that the target user's perception is normal, acquiring the target user's first life behavior status data; when the preliminary prediction result indicates that the target user's perception is abnormal, acquiring the target user's second life behavior status data, and generating preliminary warning information based on the second life behavior status data; the first life behavior status data is provided by a first image, the second life behavior status data is provided by a second image, and the resolution of the second image is greater than the resolution of the first image.

5. The multimodal fusion life perception method according to claim 2, characterized in that, The step of inputting the multimodal sensor information and the preliminary features into the second perception model to obtain a secondary prediction result includes: combining the denoised and normalized life behavior state data and the continuous vital signs data with the preliminary features to form second multimodal behavior state data; and inputting the second multimodal behavior state data into the second perception model for prediction to obtain a secondary prediction result.

6. The multimodal fusion life perception method according to claim 1, characterized in that, The step of generating early warning information based on the correlation between the preliminary prediction result and the secondary prediction result includes: when both the preliminary prediction result and the secondary prediction result show abnormalities, the early warning information is a first abnormality early warning information; when the secondary prediction result shows an abnormality and the preliminary prediction result shows normality, the early warning information is a second abnormality early warning information; when the secondary prediction result shows normality and the preliminary prediction result shows an abnormality, the early warning information is a third abnormality early warning information; the abnormality intensity of the first abnormality early warning information is greater than the abnormality intensity of the second abnormality early warning information, and the abnormality intensity of the second abnormality early warning information is greater than the abnormality intensity of the third abnormality early warning information.

7. The multimodal fusion life perception method according to claim 1, characterized in that, The step of generating a first perception model and a second perception model based on the deployment location and type of the sensors includes: constructing an initial first perception model and an initial second perception model based on the deployment location and type of the sensors; training the initial first perception model to obtain a first perception model; and training the initial second perception model to obtain a second perception model.

8. A multimodal fusion-based life perception system, characterized in that, The system includes a terminal, a server, and multiple sensors, which communicate with each other. The sensors are deployed based on lifestyle and scene information. The terminal acquires multimodal sensor information, constructs an initial first perception model based on the sensor deployment locations and types, trains the initial first perception model to obtain a second perception model, and inputs the multimodal sensor information into the first perception model to obtain preliminary features and preliminary prediction results. The server acquires the multimodal sensor information and the preliminary features, constructs an initial second perception model based on the sensor deployment locations and types, trains the initial second perception model to obtain a second perception model, and inputs the multimodal sensor information and the preliminary features into the second perception model to obtain a second prediction result. Warning information is generated based on the correlation between the preliminary prediction result and the second prediction result.

9. A multimodal fusion life sensing system according to claim 8, characterized in that, The types of sensors include color cameras, depth cameras, millimeter-wave radar, and infrared cameras.

10. A multimodal fusion life sensing system according to claim 9, characterized in that, The system is equipped with a communication module, which is used to send information to designated users based on early warning information.

Citation Information

Patent Citations

  • Old man falling detecting method and system based on multi-sensor fusion

    CN104586398A

  • Method and system for analyzing living environment of smart home suitable for aging

    CN116883905A

  • Fall risk early assessment method based on multi-modal information fusion

    CN117672529A

  • Data matching method and device and multi-modal model processing method and device

    CN118861697A

  • Multi-modal fusion life perception method and system

    CN119516602A