Old people hazard identification and risk assessment system based on multi-modal fusion
Through a hazard identification and risk assessment system for elderly people based on multimodal fusion, combined with infrared and RGB images for hazard detection and risk assessment, the problem that the existing technology cannot achieve early detection and prevention of hidden dangers is solved, and efficient risk identification and risk assessment is achieved, which significantly reduces the probability of accidents.
Patent Information
- Application Number
- CN202411977151.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology cannot achieve preliminary inspection and prevention of hidden dangers in the home environment of the elderly, and cannot identify and eliminate potential hazards before the danger occurs, resulting in a high probability of accidents.
The hazard identification and risk assessment system for the elderly based on multimodal fusion is adopted, infrared and RGB images are collected through the data collection module. The image fusion module fuses the two images into a fusion image with multimodal characteristics. The hazard detection module uses the fusion image for object detection. The risk assessment module evaluates the risk level and response measures based on the detection results.
It significantly improves the comprehensiveness and accuracy of environmental safety assessment, realizes pre-prevention, reduces the probability of accidents, provides targeted improvement suggestions, and improves the safety of the home environment of the elderly.
Smart Images

Figure CN119942587A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental risk assessment, and in particular to a hazard identification and risk assessment system for the elderly based on multi-modal fusion. Background Art
[0002] As the global population ages, the safety and well-being of the elderly has become a focus of attention in all fields. Changes in traditional family structures have significantly increased the number of elderly people living alone or in empty nests, and the safety hazards and potential risks they face in their daily lives have become more prominent. Due to the decline in mobility and cognitive function caused by aging, the elderly are more prone to accidents such as falls and burns in the home environment. Therefore, smart elderly care technology has emerged, aiming to build a smart living environment centered on the needs of the elderly through advanced Internet of Things, artificial intelligence and big data technologies. Among them, environmental safety protection is one of the core contents of smart elderly care technology.
[0003] However, current research has some shortcomings. First, there is a lack of comprehensive and systematic safety assessment of the elderly's home environment, and the identification of hazard types is too single, mainly limited to specific risks such as fire, falls, and gas leaks, and there is a lack of a way to identify multiple hazards at the same time.
[0004] Second, existing research focuses on emergency response after an accident occurs, for example, fall detection based on vision, radar or wearable devices, monitoring of emergency situations such as fire and gas leaks using environmental sensors, or monitoring of health status and abnormal behavior through wearable devices.
[0005] These technologies have improved the ability of the elderly to cope with emergencies to a certain extent and reduced the consequences of accidents. However, over-reliance on post-event response mechanisms cannot fundamentally solve the problem and fail to achieve early detection and prevention of hidden dangers. Therefore, it is necessary to emphasize the principle of prevention over remediation, discover dangerous hidden dangers in the environment in advance and solve them in time. The concealment and complexity of potential risks determine that we must pay attention to potential dangerous factors in the environment, such as whether the home design is reasonable, whether the ground is slippery, and whether there are exposed wires. These safety hazards are often the direct cause of accidents. If these hidden dangers can be identified and eliminated before the danger occurs, the probability of accidents will be significantly reduced, thereby better protecting the lives of the elderly. Summary of the invention
[0006] In view of the shortcomings of the existing technology, the present invention provides a hazard identification and risk assessment system for the elderly based on multimodal fusion, which solves the technical problem that the existing technology cannot achieve early screening and prevention of hidden dangers, and can identify and eliminate these hidden dangers before the danger occurs. The present invention can identify and eliminate these hidden dangers before the danger occurs, which will significantly reduce the probability of accidents and thus better protect the lives of the elderly.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a multi-modal fusion-based elderly hazard identification and risk assessment system, the hazard identification and risk assessment system comprising:
[0008] A data collection module, which is used to collect infrared images and RGB images in the living environment of the elderly, and obtain complementary image pair samples to form a data set of the elderly's home environment containing different dangerous scenes and various environmental conditions. The image pair sample is formed by each RGB image corresponding to one infrared image;
[0009] An image fusion module, the image fusion module is used to fuse the image pair samples in the elderly home environment data set into a fused image with at least two modality features;
[0010] A hazard detection module, which uses the fused image to perform target detection to identify potential risk factors and calculate their confidence levels;
[0011] A risk assessment module, which assesses the overall situation based on risk factors and their confidence levels, and ultimately generates an assessment result, including risk level, possible hazard types, and recommended countermeasures.
[0012] Furthermore, the image fusion module includes:
[0013] A multi-scale generation module, wherein the multi-scale generation module is used to upsample the image pair samples, connect the images element by element to obtain dual-channel features, and then use bilinear samples to divide the dual-channel features into features of three different scales;
[0014] A simple convolution module is used to perform preliminary feature extraction on features of three different scales and generate fusion features of three different scales. The simple convolution module is represented as follows:
[0015]
[0016] Among them, F out Represents the features after convolution filtering; represents a 5×5 convolution with input channel 2 and output channel C, and the padding size of the replication mode is 2; BN(·) and Re(·) are batch normalization and ReLU, respectively; {Iir ,I vi} represents the feature input after the two images are spliced;
[0017] A dual attention module, which is used to perform dimension enhancement on the fusion features of three different scales to maintain the consistency of the dimensions of the fusion features of different scales to obtain high-level semantic features;
[0018] Two transformer modules, the two transformer modules are used to perform remote context feature extraction on high-level semantic features to capture global context information, and the two transformer modules are represented as:
[0019] F SA =MSA(LN(F in ))+F in
[0020] F out =MLP(LN(F SA ))+F SA
[0021] Among them, F in Represents the input feature; F SA represents the output features through multi-head attention; F out is the output feature of the transformer module; LN(·) represents normalization; MLP(·) represents a multi-layer perceptron;
[0022] A multi-scale adaptive fusion module is used to obtain a fused image containing local important features and global important information from the global context information. The multi-scale adaptive fusion module is expressed as:
[0023]
[0024] Among them, I F represents the fused image; represents a 1×1 convolution of input channel C and output channel 1; Tanh(·) represents the tanh activation function.
[0025] Furthermore, the dual attention module includes a residual block, a channel attention block and a spatial attention block, wherein:
[0026] The residual block consists of two convolutional layers, expressed as:
[0027]
[0028] Among them, F in represents the input features, Represents a 3×3 convolution of input channel C and output channel C;
[0029] The channel attention block consists of five layers: the first layer is a global average pooling layer for generating channel descriptors, the second layer is a fully connected layer for dimensionality reduction, the third and fourth layers are ReLU and another FC layer for dimensionality increase, and the fifth layer is a sigmoid activation layer;
[0030] The spatial attention block consists of two convolutional layers, a ReLU and a sigmoid activation layer to highlight features in space.
[0031] Furthermore, the process of obtaining the fused image by the multi-scale adaptive fusion module includes:
[0032] First, a 1x1 convolutional layer is used for dimensionality enhancement to maintain the consistency of feature dimensions at different scales.
[0033] Then, each feature generates a weight corresponding to each scale through the CAB module, which is multiplied by the input feature, weighted and fused to obtain the fused feature;
[0034] Finally, a fused image is obtained through a 1x1 convolutional layer and a tanh activation function.
[0035] Furthermore, the danger detection module uses the YOLOv8 target detection model to perform target detection, and calculates the confidence of each detection frame by predicting whether there is an object P (Object) in the target frame and the conditional probability P (Class|Object) of the object class, that is:
[0036] Confindence=P(Object)×P(Class|Object)
[0037] The risk factor is selected by the maximum confidence, and the risk factor and its corresponding maximum confidence are output to the next module, wherein the maximum confidence is the confidence level.
[0038] Furthermore, the risk assessment module adopts a lightweight CNN neural network, including a convolution layer for performing two sets of convolution operations, a flattening layer and a fully connected layer;
[0039] The convolution layer includes a convolution layer with a convolution kernel of 3×3 for extracting local features of the input features, a ReLU activation function for performing a nonlinear transformation on the output of the convolution layer, and a pooling layer for downsampling feature maps, reducing data dimensions and retaining important information;
[0040] The flattening layer is used to flatten the multi-dimensional feature map obtained after convolution and pooling operations into a one-dimensional feature vector;
[0041] The fully connected layer includes two fully connected layers, which are used to gradually linearly combine the flattened one-dimensional feature vectors and comprehensively analyze the input features to capture higher-level risk assessment information.
[0042] Furthermore, the Bayesian optimization method is used to iteratively optimize the learning rate θ of the risk assessment module, the search range is set to [10-5, 10-1], and the expected improvement acquisition function is used to gradually find the optimal value α(θ) of the learning rate θ, which is expressed as:
[0043]
[0044] Among them, f(θ * ) represents the optimal performance of the current model; f(θ) represents the model performance of the current θ; E[f(θ)] is the performance expectation of the model given the hyperparameter θ.
[0045] The technical solution also provides an application of the above-mentioned hazard identification and risk assessment system, which is used to conduct comprehensive detection and assessment of potential risk factors in the living environment of the elderly, and provide targeted improvement suggestions.
[0046] By means of the above technical solution, the present invention provides a multimodal fusion-based elderly hazard identification and risk assessment system, which has at least the following beneficial effects:
[0047] 1. The hazard identification and risk assessment system proposed in the present invention can identify dangerous factors in the environment, assess the level of danger, predict possible dangerous situations, and provide targeted improvement suggestions by integrating the thermal characteristics of infrared images with the high spatial resolution information of visible light images. It not only significantly improves the comprehensiveness and accuracy of environmental safety assessment, but also shifts from traditional post-detection to pre-prevention, taking countermeasures before the danger actually turns into an accident, reducing the possibility of accidents.
[0048] 2. The present invention can achieve more efficient operation under different environmental conditions and significantly enhance the robustness and reliability of detection. By automatically assessing the danger level and providing improvement suggestions, the model not only provides higher security for the elderly, but also provides practical guidance for families and caregivers, helping to improve the safety of the elderly's home environment from the source.
[0049] 3. The present invention can combine infrared and visible light images in a home environment to obtain image information under various environmental conditions such as day and night, light intensity, etc. It is all-weather and suitable for home safety monitoring of the elderly in various environments. In addition, the fusion of these two images can improve the accuracy and robustness of recognition and evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0051] Figure 1 It is a network framework diagram of the hazard identification and risk assessment system in the present invention;
[0052] Figure 2 This is an example diagram of the elderly home environment data set in the present invention;
[0053] Figure 3 This is a result diagram comparing the existing image fusion algorithm in the present invention. DETAILED DESCRIPTION
[0054] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods, so that the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.
[0055] In order to fundamentally realize the early detection and prevention of hidden dangers, discover the dangerous hidden dangers in the environment in advance and solve them in time, so as to better protect the life safety of the elderly, please refer to Figure 1-Figure 3 This embodiment proposes a multimodal fusion-based elderly hazard identification and risk assessment system, which consists of a data collection module, an image fusion module, a hazard detection module, and a risk assessment module. It generates a fused image by fusing the information of infrared and visible light images, and classifies and identifies the risk factors in the environment in combination with the corresponding target detection network. Finally, a neural network is used to implement risk assessment. Figure 1 As shown in the figure, after collecting the infrared image and RGB image, the image fusion module fuses the two input modal data into a fused image with two characteristics; the danger detection module identifies the danger factors in the fused image, and the risk assessment module conducts a comprehensive assessment of these danger factors, calculates the safety level, and gives reasonable warnings or improvement suggestions.
[0056] The process of the entire hazard identification and risk assessment system is as follows: First, images are collected through infrared cameras and visible light cameras, and complementary image pairs are obtained to form a dataset of elderly home environments. These images are then input into the image fusion module and fused into a fused image with two modal features. The fused image is then used for target detection to identify potential risk factors and calculate their confidence levels. The detected risk factors and their confidence levels are input into the risk assessment module, which will evaluate the overall situation and ultimately generate an assessment result, including the hazard level, specific hazards that may occur, and recommended countermeasures.
[0057] Among various environmental modalities, the combination of infrared images and RGB images has advantages. Infrared images can capture the thermal characteristics of objects and the distribution of surface temperature in low light or bad weather conditions, and have advantages in detecting objects with large temperature differences, but the spatial resolution of infrared images is usually relatively low. RGB images can provide high spatial resolution and clear texture details of objects, and have the advantage of revealing clear long-distance correlations over infrared spectra, but RGB images require good lighting conditions. These two modalities can complement each other, and after fusion, rich environmental information can be obtained.
[0058] In terms of technology, multimodal technology has also been developing continuously in recent years. Compared with single-modal technology, multimodal technology can provide richer information and has a wider research space. By integrating information from different sensors, signals or data modes, more comprehensive and accurate features can be extracted to achieve more complex and precise task objectives. The core idea of this technology is to use the complementarity between different modalities to make up for the limitations of a single modality.
[0059] As a further implementation method, in this embodiment, the data collection module is used to collect infrared images and RGB images in the living environment of the elderly through infrared cameras and visible light cameras, and obtain complementary image pair samples to form an elderly home environment data set containing different dangerous scenes and various environmental conditions. The image pair sample is one sample for each RGB image corresponding to an infrared image. This embodiment uses an infrared binocular camera module with a picture size of 1920*1080, and each RGB image corresponds to an infrared image as a sample. Due to the uncertainty of time and space and equipment reasons when the binocular camera is shooting, there may be misalignment in each group of samples.
[0060] In the data collection of this embodiment, a dataset of elderly people’s home environment including infrared images and RGB images is collected in the homes of the elderly in the community or nursing homes. It is ensured that the dataset contains different dangerous scenes and various environmental conditions to ensure the generalization of the hazard identification and risk assessment system. Figure 2As shown, these images cover a range of scenes, such as living rooms, bedrooms, kitchens, bathrooms, stairs, corridors, etc. The data are all labeled with specific hazards, making them suitable for classification and object detection tasks. Hazard factor labels include: loose or uneven carpets, cats, chairs, steps / stairs, damp and moldy, peeling paint or wall skin. 300 infrared images were selected as experimental data, and data enhancement methods were used. Seven data enhancement methods were used, including mirror flipping, translation, rotation, brightness adjustment, adding noise, cropping, and flipping. After enhancement, 2100 pairs of infrared and visible light images can be obtained. The dataset is divided into training set, validation set, and test set in a ratio of 7:2:1, so as to train, verify, and test the entire hazard identification and risk assessment system to ensure the generalization of the hazard identification and risk assessment system.
[0061] In this embodiment, the image fusion module is used to fuse the image pair samples in the elderly home environment dataset into a fused image with at least two modal features. In the image fusion module, two methods, multi-scale fusion and dual attention transformer, are used. First, the input image is upsampled to obtain pictures of different scales. The three pictures of different scales are subjected to preliminary feature extraction through a simple convolution module (SFC). Then, a dual attention module (DARM) is used to obtain high-level semantic features, and then two transformer modules (TRM) are used to extract remote context features in order to fully extract global feature information.
[0062] To further supplement the details, this embodiment designs a multi-scale method. After connecting the image element by element to obtain dual-channel features, these features are divided into features of three different scales using bilinear samples. Then, three fused features of different scales are generated by the dual attention module. Then this embodiment designs a multi-scale adaptive fusion module (MSFM) to adaptively fuse features from three different scales to generate fusion results of features of different scales. The image fusion module includes a multi-scale generation module, a simple convolution module (SFC), a dual attention module (DARM), two transformer modules (TRM) and a multi-scale adaptive fusion module (MSFM), where:
[0063] The multi-scale generation module is used to upsample the image samples and connect the images element by element to obtain the dual-channel features. Then, the dual-channel features are divided into three features of different scales using bilinear samples.
[0064] The simple convolution module is used to perform preliminary feature extraction on features of three different scales and generate fusion features of three different scales. The simple convolution module is expressed as:
[0065]
[0066] Among them, F out Represents the features after convolution filtering; represents a 5×5 convolution with input channel 2 and output channel C, and the padding size of the replication mode is 2; BN(·) and Re(·) are batch normalization and ReLU, respectively; {I ir ,I vi} represents the feature input after the two images are spliced;
[0067] The dual attention module is used to enhance the dimension of the fused features of three different scales to maintain the consistency of the dimensions of the fused features of different scales and obtain high-level semantic features;
[0068] The dual attention module (DARM) combines residual learning with channel attention and spatial attention to form a feature-preserving module, which mainly includes residual block (RB), channel attention block (CAB) and spatial attention block (SAB).
[0069] The residual block (RB) consists of two convolutional layers, expressed as:
[0070]
[0071] Among them, F in represents the input features, Represents a 3×3 convolution with input channel C and output channel C.
[0072] The channel attention block (CAB) has five layers in total. The first layer is a global average pooling (GAP) layer to generate channel descriptors, the second is a fully connected (FC) layer for dimensionality reduction, followed by ReLU and another FC layer to increase the dimension, and the last layer is a sigmoid activation layer.
[0073] The Spatial Attention Block (SAB) consists of two convolutional layers, a ReLU and a sigmoid activation layer to highlight features in space.
[0074] Two transformer modules (TRMs) are used to perform long-range context feature extraction on high-level semantic features to capture global context information, including normalization, multi-head attention, and multi-layer perceptron, which can be expressed as:
[0075] F SA =MSA(LN(F in ))+F in
[0076] F out =MLP(LN(F SA ))+F SA
[0077] Among them, F inRepresents the input feature; F SA represents the output features through multi-head attention; F out is the output feature of the transformer module; LN(·) represents normalization; MLP(·) represents a multi-layer perceptron.
[0078] The multi-scale adaptive fusion module (MSFM) is used to obtain a fused image containing local important features and global important information from the global context information. The multi-scale adaptive fusion module uses a convolution for dimensionality reduction, which is expressed as:
[0079]
[0080] Among them, I F represents the fused image; represents a 1×1 convolution of input channel C and output channel 1; Tanh(·) represents the tanh activation function.
[0081] In this way, both local important features and global important information can be obtained at the same time, thus obtaining a fusion result with rich information, prominent targets and clear details. The process of obtaining the fused image by the Multi-Scale Adaptive Fusion Module (MSFM) includes:
[0082] First, a 1x1 convolutional layer is used for dimensionality enhancement to maintain the consistency of feature dimensions at different scales. Then, each feature generates a weight corresponding to each scale through the CAB module, which is multiplied by the input feature, weighted and fused to obtain the fused feature. Finally, a 1x1 convolutional layer and a tanh activation function are used to obtain the fused image.
[0083] This embodiment uses multi-scale fusion and Transformer fusion technology, which can achieve more efficient operation under different environmental conditions and significantly enhance the robustness and reliability of detection compared to existing technologies. By automatically assessing the danger level and providing improvement suggestions, it not only provides higher security for the elderly, but also provides practical guidance for families and caregivers, helping to improve the safety of the elderly's home environment from the source.
[0084] In this embodiment, the danger detection module uses the fused image for target detection to identify potential risk factors and calculate their confidence levels; the danger detection module uses the YOLOv8 target detection model for target detection, obtains the fused image after undergoing the image fusion module, and then inputs the YOLOv8 target detection model for detection. The confidence of each detection frame is calculated by predicting whether there is an object P (Object) in the target frame and the conditional probability P (Class|Object) of the object class, that is:
[0085] Confindence=P(Object)×P(Class|Object)
[0086] The risk factor is selected by the maximum confidence, and the risk factor and its corresponding maximum confidence are output to the next module, wherein the maximum confidence is the confidence level.
[0087] In this embodiment, the risk assessment module evaluates the overall situation based on the risk factors and their confidence levels, and finally generates an assessment result, including the risk level, the type of hazards that may occur, and the recommended countermeasures. This embodiment takes into account the mutual influence between hazards, so a neural network method from an uncertain non-statistical method is used for multimodal features in the risk assessment module, combined with a Bayesian method. The input of the risk assessment module is a feature composed of a risk factor (Riskfactor) and a confidence level (Confidence level). These features are usually generated by the previous YOLOv8 target detection module, including the detected risk factor category and its corresponding confidence level. A lightweight CNN is used to extract input features and reduce the computational burden of the model. The composition of this module includes a convolutional layer, a flattening layer, and a fully connected layer.
[0088] In the convolutional layer, the first set of convolution operations is (Conv 3×3+ReLU+Pooling 2×2), which includes a convolution layer with a convolution kernel of 3×3, which is used to extract local features of the input features. Then the ReLU activation function is used to nonlinearly transform the output of the convolution layer. The Pooling 2×2 pooling layer is used to downsample the feature map, reduce the data dimension and retain important information. The second set of convolution operations repeats the above structure (Conv3×3+ReLU+Pooling 2×2) to further extract deeper features.
[0089] The flatten layer flattens the multi-dimensional feature map obtained after convolution and pooling operations into a one-dimensional feature vector so that it can be input into the fully connected layer for further processing.
[0090] The fully connected layer includes two fully connected layers, which gradually perform linear combinations on the flattened feature vectors to comprehensively analyze the input features and capture higher-level risk assessment information.
[0091] In order to optimize the hyperparameters in the lightweight CNN neural network, the Bayesian optimization method is used. The core of Bayesian optimization is to model the hyperparameter θ (learning rate) as a probability distribution and maximize the performance of the model through iterative optimization, which is expressed as: E[f(θ)] is the performance expectation of the model given the hyperparameter θ. As the core variable to be optimized, its search range is set to [10-5, 10-1], and the expected improvement (EI) acquisition function is used to gradually find the optimal value of the learning rate θ. This function can be expressed as: Among them, f(θ * ) represents the optimal performance of the current model, and f(θ) represents the model performance of the current θ.
[0092] In this embodiment, the risk level is divided into three levels: low, medium, and high. The potential danger and environmental improvement suggestions are judged by predefined rules, which can be regarded as a multi-label classification problem. The predefined rules are shown in Table 1.
[0093] Table 1 Predefined rules
[0094]
[0095]
[0096] Each label represents a warning or improvement suggestion, and various warnings and suggestions are pre-set. Adam (Adaptive Moment Estimation) is used as the optimizer to minimize the loss function.
[0097] This embodiment proposes a new system framework that integrates hazard detection and risk assessment functions, and uses the fusion characteristics of infrared and visible light dual-modal images to comprehensively detect and evaluate potential risk factors in the living environment of the elderly. By fusing the thermal characteristics of infrared images with the high spatial resolution information of visible light images, it is possible to identify dangerous factors in the environment, assess the level of danger, predict possible dangerous situations, and provide targeted improvement suggestions. This model not only significantly improves the comprehensiveness and accuracy of environmental safety assessments, but also shifts from traditional post-detection to pre-prevention, taking countermeasures before the danger actually turns into an accident, reducing the possibility of accidents.
[0098] This example compares the proposed hazard identification and risk assessment system with three advanced infrared and visible light image fusion algorithms, and also includes two unprocessed data as a benchmark. The experimental results are shown in Figure 2. Figure 3 At the same time, an objective comparison of the three algorithms in the PHELE dataset is also conducted, and the results of the four evaluation indicators of precision, recall, mAP50 and mAP50-95 are shown in Table 2.
[0099] Table 2 Evaluation index results
[0100] Fusion Method Accuracy Recall mAP50 mAP50-95 Visible light images 0.6922 0.68308 0.61352 0.42687 Infrared pictures 0.60169 0.76416 0.59257 0.42152 DATFuse 0.73426 0.79108 0.70871 0.52985 IPLF 0.75043 0.77981 0.71317 0.53448 Res2Fusion 0.75252 0.79164 0.72704 0.50932 This embodiment 0.79628 0.84394 0.81609 0.63484
[0101] As can be seen from the above table, this embodiment achieves the highest performance in terms of precision, recall, mAP50, and mAP50-95 indicators, and demonstrates its significant advantages over other alternative solutions.
[0102] In summary, in a home environment, combining infrared and visible light images can obtain image information under various environmental conditions such as day and night, light intensity, etc. The system is all-weather and suitable for home safety monitoring of the elderly in various environments. And the fusion of these two images can improve the accuracy and robustness of recognition and evaluation.
[0103] By fusing the two images, we can use the information of visible light and infrared images at the same time to obtain richer and more diverse visual features. The self-attention mechanism in the Transformer module can help the model effectively capture the relationship between different channels and features, thereby achieving better feature fusion and information extraction. Through the self-attention mechanism, the importance of each channel and feature can be automatically learned, and fused and weighted accordingly.
[0104] The lightweight CNN of the risk assessment module has fewer parameters and computational complexity, and can achieve lower computational complexity and memory consumption while maintaining high performance. This makes the hazard identification and risk assessment system more efficient and practical when deployed on mobile devices or embedded systems.
[0105] This embodiment also proposes specific application fields based on the hazard identification and risk assessment system. The physical hazard identification and safety assessment method for the elderly at home based on infrared and visible light images has broad research prospects in the field of safety. It can be applied to the elderly home monitoring system to timely discover possible safety hazards and conduct early warnings and alarms. Deploying infrared and visible light image monitoring systems in smart elderly communities can realize intelligent monitoring and management of the elderly, improve the safety and management efficiency of the community, and provide a better living environment and services for the elderly. Applied in the field of hospital environment detection, it can help hospitals manage the safety of patients.
[0106] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, so the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0107] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the above embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0108] The above implementation methods have been described in detail. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A multimodal fusion-based elderly hazard identification and risk assessment system, characterized in that: The hazard identification and risk assessment system includes: A data collection module, which is used to collect infrared images and RGB images in the living environment of the elderly, and obtain complementary image pair samples to form a data set of the elderly's home environment containing different dangerous scenes and various environmental conditions. The image pair sample is formed by each RGB image corresponding to one infrared image; An image fusion module, the image fusion module is used to fuse the image pair samples in the elderly home environment data set into a fused image with at least two modality features; A hazard detection module, which uses the fused image to perform target detection to identify potential risk factors and calculate their confidence levels; A risk assessment module, which assesses the overall situation based on risk factors and their confidence levels, and ultimately generates an assessment result, including risk level, possible hazard types, and recommended countermeasures.
2. The hazard identification and risk assessment system according to claim 1, characterized in that: The image fusion module comprises: A multi-scale generation module, wherein the multi-scale generation module is used to upsample the image pair samples, connect the images element by element to obtain dual-channel features, and then use bilinear samples to divide the dual-channel features into features of three different scales; A simple convolution module is used to perform preliminary feature extraction on features of three different scales and generate fusion features of three different scales. The simple convolution module is represented as follows: Among them, F out Represents the features after convolution filtering; represents a 5×5 convolution with input channel 2 and output channel C; BN(·) and Re(·) are batch normalization and ReLU, respectively; {I ir ,I vi } represents the feature input after the two images are spliced; A dual attention module, which is used to perform dimension enhancement on the fusion features of three different scales to maintain the consistency of the dimensions of the fusion features of different scales to obtain high-level semantic features; Two transformer modules, the two transformer modules are used to perform remote context feature extraction on high-level semantic features to capture global context information, and the two transformer modules are represented as: F SA =MSA(LN(F in ))+F in F out =MLP(LN(F SA ))+F SA Among them, F in Represents the input feature; F SA represents the output features through multi-head attention; F out is the output feature of the transformer module; LN(·) represents normalization; MLP(·) represents a multi-layer perceptron; A multi-scale adaptive fusion module is used to obtain a fused image containing local important features and global important information from the global context information. The multi-scale adaptive fusion module is expressed as: Among them, I F Represents the fused image; Conv1 C,1 represents a 1×1 convolution of input channel C and output channel 1; Tanh(·) represents the tanh activation function.
3. The hazard identification and risk assessment system according to claim 2, characterized in that: The dual attention module includes a residual block, a channel attention block and a spatial attention block, wherein: The residual block consists of two convolutional layers, expressed as: Among them, F in represents the input features, Represents a 3×3 convolution of input channel C and output channel C; The channel attention block consists of five layers: the first layer is a global average pooling layer for generating channel descriptors, the second layer is a fully connected layer for dimensionality reduction, the third and fourth layers are ReLU and another FC layer for dimensionality increase, and the fifth layer is a sigmoid activation layer; The spatial attention block consists of two convolutional layers, a ReLU and a sigmoid activation layer to highlight features in space.
4. The hazard identification and risk assessment system according to claim 2, characterized in that: The process of obtaining the fused image by the multi-scale adaptive fusion module includes: First, a 1x1 convolutional layer is used for dimensionality enhancement to maintain the consistency of feature dimensions at different scales. Then, each feature generates a weight corresponding to each scale through the CAB module, which is multiplied by the input feature, weighted and fused to obtain the fused feature; Finally, a fused image is obtained through a 1x1 convolutional layer and a tanh activation function.
5. The hazard identification and risk assessment system according to claim 1, characterized in that: The danger detection module uses the YOLOv8 target detection model to perform target detection. It calculates the confidence of each detection box by predicting whether there is an object P (Object) in the target box and the conditional probability P (Class|Object) of the object category, that is: Confindence=P(Object)×P(Class|Object) The risk factor is selected by the maximum confidence, and the risk factor and its corresponding maximum confidence are output to the next module, wherein the maximum confidence is the confidence level.
6. The hazard identification and risk assessment system according to claim 1, characterized in that: The risk assessment module adopts a lightweight CNN neural network, including a convolution layer for performing two sets of convolution operations, a flattening layer, and a fully connected layer; The convolution layer includes a convolution layer with a convolution kernel of 3×3 for extracting local features of the input features, a ReLU activation function for performing a nonlinear transformation on the output of the convolution layer, and a pooling layer for downsampling feature maps, reducing data dimensions and retaining important information; The flattening layer is used to flatten the multi-dimensional feature map obtained after convolution and pooling operations into a one-dimensional feature vector; The fully connected layer includes two fully connected layers, which are used to gradually linearly combine the flattened one-dimensional feature vectors and comprehensively analyze the input features to capture higher-level risk assessment information.
7. The hazard identification and risk assessment system according to claim 6, characterized in that: The Bayesian optimization method is used to iteratively optimize the learning rate θ of the risk assessment module, the search range is set to [10-5, 10-1], and the expected improvement acquisition function is used to gradually find the optimal value α(θ) of the learning rate θ. The expression is: α(θ)=E[max(0,f(θ)-f(θ*))] Among them, f(θ * ) represents the optimal performance of the current model; f(θ) represents the model performance of the current θ; E[f(θ)] is the performance expectation of the model given the hyperparameter θ.
8. An application of the hazard identification and risk assessment system according to any one of claims 1 to 7, characterized in that: The hazard identification and risk assessment system is used to comprehensively detect and evaluate potential risk factors in the living environment of the elderly, and provide targeted improvement suggestions.