Anti-lost multi-mode risk assessment and intelligent early warning system for intermittent dementia old people
This anti-wandering early warning system, which collects multimodal data through a smart helmet and uses a tree-structured semantic map encoder for multi-dimensional evaluation, solves the digital divide problem when elderly people get lost. It enables real-time danger assessment and early warning for dementia patients, improving the success rate and safety of wandering warnings.
Patent Information
- Application Number
- CN202510980760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-04
AI Technical Summary
Existing smart devices require elderly people to actively operate or passively respond, failing to effectively consider the digital divide among the elderly, resulting in low efficiency in locating and alerting dementia patients when they get lost.
A multimodal hazard assessment and intelligent early warning system for preventing elderly people with intermittent dementia is designed. The system includes a smart helmet, a multimodal hazard identification network based on scene-location information, and early warning software. Multimodal data is collected through the smart helmet, and a tree-structured semantic map encoder is used to align one-dimensional GPS information with two-dimensional scene image information to perform multi-dimensional hazard assessment. Real-time early warning is then provided through the early warning software on the guardian's end.
It enables multi-dimensional and accurate assessment of the process of elderly people getting lost, improves the success rate and real-time performance of missing person warnings, reduces the probability of elderly people encountering danger outdoors, and solves the digital divide problem.
Smart Images

Figure CN120895231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of electronic information and computer technology, and particularly relates to a multi-modal risk assessment and intelligent early warning system for preventing the wandering of intermittent senile dementia patients. BACKGROUND
[0002] With the aging of the population, the number of dementia patients will increase. At present, about 55 million people worldwide suffer from dementia. There are about 10 million new cases diagnosed each year, and the total number of cases is expected to rise to 78 million by 2030. With the increase in the number of dementia patients, the incidence of missing events due to severe wandering is also rising. Because patients cannot effectively distinguish the danger of the surrounding environment and lack the ability to actively call for help, the risk of injury or even death is high once they are lost. At present, the search for such missing persons consumes a lot of resources and time, and the probability of successful recovery is very low. Therefore, doing a good job in preventing the wandering of senile dementia patients is an effective method to protect the life safety of such patients.
[0003] The problem of preventing the wandering of the elderly has always been a social hot issue. Many researchers have proposed solutions to this problem. According to the design principle, existing methods can be divided into two categories: the first category uses the information of the lost person to seek help from society to achieve the pursuit of the lost person. The second category uses a portable positioning device to achieve the pursuit of the lost person.
[0004] Although multi-angle application research has been carried out on the problem of the wandering of the elderly. However, the existing methods still have the following shortcomings: (1) The method of seeking help from society using the information of the elderly has a certain effect, but its effect is limited by the number of people who can provide help near the lost person. (2) The method of tracking the trajectory of the lost person can locate the lost person in time, but due to the lack of first-person scene information, it fails to consider the potential safety hazards that the environment brings to the elderly. In addition, existing smart devices such as mobile phones and bracelets require the active operation or passive response of the elderly, and do not consider the digital divide problem of the elderly. With the acceleration of social digitization and intelligentization, the gray digital divide problem is to some extent aggravated. The convenience of technology should not be exclusive to young people, and the elderly need to enjoy the convenience brought by intelligentization and digitization. SUMMARY
[0005] The purpose of the present application is to solve the problem in the prior art that smart devices such as mobile phones and bracelets require the active operation or passive response of the elderly, and do not consider the digital divide problem of the elderly.
[0006] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0007] The application discloses a multi-modal danger assessment and intelligent early warning system for an intermittent senile dementia patient, which comprises an intelligent helmet, a multi-modal danger state recognition network based on scene-position information and early warning software, wherein the intelligent helmet is used for collecting multi-modal data, transmitting the data to a server through a wireless network and storing the data in a Microsoft SQL Server database; the multi-modal data is input into the multi-modal danger state recognition network based on scene-position information, and the multi-modal data comprises first visual angle image information and position information of a user.
[0008] The multi-modal danger state recognition network based on scene-position information is used for evaluating the multi-modal data and displaying evaluation results in the form of intuitive texts on the early warning software of a guardian; a tree structure semantic map encoder is used in the multi-modal danger state recognition network based on scene-position information to align the dimensions of one-dimensional GPS information and two-dimensional scene image information.
[0009] Preferably, the multi-modal danger state recognition network based on scene-position information comprises a first visual angle scene feature extraction module, a causal heuristic position information feature extraction network and a scene-position based multi-modal fusion module, wherein the causal heuristic position information feature extraction network comprises a tree structure semantic map encoder and a pseudo map feature extraction module.
[0010] Preferably, the first visual angle scene feature extraction module is used for extracting scene features f s from a target first visual angle image; the tree structure semantic map encoder is used for mapping one-dimensional position information of the old person to a two-dimensional image space by using real-time position information and historical activity data of a time period of Δt, generating a pseudo map image and constructing a pseudo map image sequence of {x t ,x t-1}.
[0011] In the pseudo map feature extraction module, adjacent two pseudo maps x t , x t-1 are sequentially input into a SAN-18 network to extract static trajectory features f 1,t ,f 1,t-1 .
[0012] The features f 1,t ,f 1,t-1 are input into an LSTM network to mine dynamic trajectory change rules f2 in adjacent pseudo map images, and then the trajectory change rule features f2 and the static trajectory features f 1,t at the time t are spliced to constitute position features f loc .
[0013] Finally, scene features f sand location features f loc The fusion of features and the Softmax classifier are used to map fused features to dangerous states.
[0014] Preferably, the workflow of the tree-structured semantic map encoder is as follows:
[0015] S1: Generating the original pseudo-map
[0016] A blank pseudo-map is generated using Python's Folium library to plot the dynamically updated movement trajectory of the elderly during their travels, and a frame of the pseudo-map image is saved at intervals of Δt (Δt = 5 min in this application). After the pseudo-map is generated, the movement trajectory on the pseudo-map continues to accumulate until it is closed. When no new data is received in the database for a period of time, it is considered that the elderly have stopped using the device, and the pseudo-map is closed.
[0017] S2: Generation of Historical Activity Areas
[0018] Assuming the update frequency of the elderly's historical activity data is set to once a week (caregivers can adjust the update frequency of historical activity data through visual monitoring and intelligent early warning software), the location information of the previous week stored in the database is used as the elderly's historical activity data for the current week; the set of historical activity location points P of the elderly is extracted from the historical activity data. T,h ={p 0,h ,…,p n,h}, where the subscript T represents the time span of the historical activity data (set to 1 week in this application), and h represents the historical activity data. These locations are marked on the pseudo-map as heat points to form historical activity areas; then the extracted historical activity and location point set P is... T,h It is also passed to the data processing and feature extraction section;
[0019] S3: Data Processing and Feature Extraction
[0020] By analyzing the set {P} Δt,o ,T Δt V Δt} and P T,h The following data analysis process was performed to extract three key indicators: minimum safe distance (MSD), travel period, and average speed (Avspeed).
[0021] S4: Tree-structured reasoning model
[0022] Finally, the above three key indicators are input into the tree-structured reasoning model used to infer the additional semantic information of the current travel location danger status.
[0023] The color and width visualization rules of the elderly travel trajectory are formulated according to the inferred additional semantic information, as shown in Table III;
[0024] S5: Dynamic trajectory generation
[0025] Based on the visual representation rules shown in Table III, the dynamic trajectory of the elderly travel is generated on the pseudo map,
[0026] Table III Different risk levels and trajectory parameters
[0027]
[0028] Preferably, the specific steps in S1 are as follows:
[0029] First, the activity position point set P Δt,o ={p 0,o ,…,p n,o} (wherein the subscript o represents real-time data) is extracted from the observed elderly travel data, and the starting point coordinate p 0,o and the ending point coordinate p n,o of the elderly travel in the current time interval Δt are determined.
[0030] Next, the starting and ending positions in this frame of pseudo map image are determined on the blank map, and the starting point position p 0,o =(x 0,o ,y 0,o ) is marked with a pentagram icon, and the ending point position p n,o =(x n,o ,y n,o ) is marked with an "i" icon.
[0031] Finally, the travel time T Δt , the ground speed V Δt in the position information, and the activity position point set P Δt,o are extracted from the real-time data to form a set {P Δt,o ,T Δt ,V Δt}, which is transmitted to the data processing and feature extraction part.
[0032] Preferably, the specific steps in S5 are as follows:
[0033] The adjacent two frames of pseudo map images are arranged into a sequence image {x t ,x t-1}, which is input to the self-attention network SAN for conversion and aggregation of GPS state features, to obtain {f 1,t ,f 1,t-1}, which is then input to the neck LSTM to extract the GPS state change rule feature f2, and finally, f2 is combined with f 1,tsplicing constitutes GPS joint feature f loc ;
[0034] Among them, the self-attention network SAN adopts a structure similar to ResNet with residual connections and block connections; SAN-X represents a network composed of X SA blocks; the neck LSTM has 3 layers, and captures the relationship between the current input and the historical state by combining forget gate, input gate and candidate state to extract the pattern of GPS state change.
[0035] Preferably, scene features f are performed in the multimodal hazard assessment module. s and location features f loc The specific process of using the fusion and Softmax classifier to map fused features to dangerous states is as follows:
[0036] First, the scene features f output by the scene feature extraction module are... s Joint GPS features f output by the GPS extraction module loc Channel splicing is performed to obtain the scene-GPS joint feature f, as shown in equation (4):
[0037] f = Concat(f) s ,f loc (4)
[0038] Then, f is input into the SENet channel attention module to determine the importance (i.e., weight) of each channel of f, focusing on those channels containing key information and giving less attention to those channels with less information; then the output features are straightened, as shown in Equation (5):
[0039] f SE =Flatten(SE(f))(5)
[0040] f SE The input uses a fully connected layer consisting of two stacked linear layers to identify dangerous states, mapping the learned features to a sample label space, and outputting the probabilities p of four dangerous states. i , as in equation (6)
[0041] As shown:
[0042] p i =Softmax(Linear(Relu(Linear(f SE ))))i∈{safe, low risk, medium risk, high risk}(6).
[0043] Preferably, the application software is designed on the guardian's mobile phone and used to display the information output by the scene-position-based multi-modal fusion module, and the cross section of the application software is composed of four parts, i.e., danger state assessment, first perspective scene, dynamic map and real-time position of the user.
[0044] Preferably, the application software also has an active early warning mechanism, which automatically performs a warning action according to the danger state assessment level.
[0045] Preferably, the intelligent helmet comprises a safety helmet, an intelligent camera module and a solar cell, wherein the intelligent camera module is composed of a camera sensor, a GPS module, a 5G communication module and a power control module, the camera sensor is used to obtain the first perspective image information of the user, and the GPS module is used to obtain the position information of the user.
[0046] Compared with the prior art, the application has the following beneficial effects:
[0047] 1. In the application, a small-scale multi-modal data set simulating the process of the old people getting lost is constructed, including first perspective scene images and GPS position information. The data set completes the matching of scene images and GPS
[0048] position information, and labels of target danger state levels at this moment are made for different scenes and position information.
[0049] 2. A scene-GPS multi-modal danger state assessment network model is proposed, which is composed of a scene feature information extraction module, a causal heuristic position information feature extraction network and a multi-modal feature fusion module. The model fuses the target first perspective scene information and the old people's travel time, speed, direction and distance information features to accurately assess the target danger state in multiple dimensions. In addition, in order to realize the dimension alignment of one-dimensional GPS information and two-dimensional scene image information, a tree structure semantic map encoder is first proposed, which is used for semantic enhancement of trajectory features in the process of converting position information to trajectory images.
[0050] 3. A simulation system for anti-lost danger state assessment and early warning based on an intelligent helmet is designed and implemented to verify the effectiveness of the scene-position information-based multi-modal danger state recognition method. The system is composed of three parts: an intelligent helmet for easy-to-lose targets with real-time multi-modal data acquisition and wireless transmission function, a scene-position information multi-modal danger state assessment model deployed in the cloud, and an intelligent early warning application software based on the Android system for the guardian's mobile phone, which receives the first perspective images and position information of the old people transmitted by the intelligent helmet in real time and actively warns the guardian according to the cloud intelligent algorithm to recognize the danger state. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 The overall architecture of the intelligent helmet-based anti-lost danger assessment and intelligent early warning system in an embodiment of the present application;
[0052] Figure 2 The assembly effect of the intelligent helmet and the specific details of the intelligent camera module in an embodiment of the present application;
[0053] Figure 3 The multi-modal danger state recognition network based on scene-GPS information in an embodiment of the present application;
[0054] Figure 4 The tree structure semantic map encoder in an embodiment of the present application;
[0055] Figure 5 The tree structure map for reasoning danger state in an embodiment of the present application;
[0056] Figure 6 The map sample generated when HD=0 in an embodiment of the present application;
[0057] Figure 7 The map sample generated when HD=1 in an embodiment of the present application;
[0058] Figure 8 The monitoring software interface and early warning method in an embodiment of the present application;
[0059] Figure 9 The information collected by the intelligent helmet in an embodiment of the present application;
[0060] Figure 10 The matching process of the position information and the scene image in an embodiment of the present application;
[0061] Figure 11 The target danger state comprehensive evaluation rule in an embodiment of the present application;
[0062] Figure 12 The data label distribution in each stage in an embodiment of the present application;
[0063] Figure 13 The data pair instance graph in an embodiment of the present application;
[0064] Figure 14 The single-modal network structure graph in the ablation experiment in an embodiment of the present application;
[0065] Figure 15 The model structure graph of the network structure adjustment experiment in the multi-modal network in an embodiment of the present application;
[0066] Figure 16A model structure diagram in a multi-modal backbone network adjustment experiment in an embodiment of the present application;
[0067] Figure 17 A loss value change in a training process in an embodiment of the present application;
[0068] Figure 18 An example diagram of simulating an old person to perform a complete travel process of "departure-return" in an embodiment of the present application;
[0069] Figure 19 An actual scene test example diagram of each scene in an embodiment of the present application. DETAILED DESCRIPTION
[0070] The application will be further described below in conjunction with specific embodiments.
[0071] Please refer to Figure 1 A missing-prevention multi-modal danger assessment and intelligent early-warning system for intermittent senile dementia patients, comprising an intelligent helmet, a multi-modal danger state recognition network based on scene-position information, and early-warning software, wherein the intelligent helmet is used to collect multi-modal data and transmit the data to the multi-modal danger state recognition network based on scene-position information through a 5G network and a TP and TCP protocol transmission; the multi-modal danger state recognition network based on scene-position information is used to evaluate the multi-modal data and display the evaluation results in the form of intuitive text on the early-warning software of a guardian, and can trigger different danger early-warning modes according to the evaluation results, thereby providing real-time and remote cloud monitoring support for guardians and reducing the probability of outdoor missing and danger of the elderly.
[0072] In an embodiment, please refer to Figure 2 The intelligent helmet comprises a safety helmet, an intelligent camera module, and a solar cell, wherein the intelligent camera module is composed of a camera sensor, a GPS module, a 5G communication module, and a power control module, the camera sensor is used to obtain first-view image information of a user (reflecting potential safety hazards in the surrounding environment), and the GPS module is used to obtain position information of the user (including a travel location, speed, time, etc.).
[0073] The intelligent camera module takes a G8100 chip as a core controller, which is used to adjust a data collection time interval, specify an upload server, maintain a TCP long connection, etc. Moreover, the intelligent helmet transmits scene images and GPS data to the multi-modal danger state recognition network based on scene-position information in real time through the 5G communication module. The power control module can ensure that the output voltage of the battery or the solar cell panel is stable.
[0074] When installing, the smart camera module is fixed by copper columns, placed and fixed in the sensor loading box, the GPS module is embedded in the groove on the surface of the loading box, the solar cell is installed into the battery loading box, then the sensor loading is fixed at the front end of the safety helmet, the battery loading box is fixed at the rear end of the safety helmet, and finally the safety belt is used for reinforcement.
[0075] When the user wears the smart helmet and goes out, the image data of the first perspective of the user collected by the camera module is uploaded to the server for storage through the FTP protocol of the 5G communication module box, and the GPS data (including latitude, longitude, altitude, UTC time, positioning state, satellite usage quantity, ground speed, heading and first positioning time) is transmitted to the server through the TCP long connection and stored in the Microsoft SQL Server database.
[0076] Please refer to Figure 3 , the multi-modal dangerous state recognition network based on scene-position information can fully utilize the first perspective scene image and position information to realize accurate evaluation of the target dangerous state.
[0077] In an embodiment, the multi-modal dangerous state recognition network based on scene-position information comprises a first perspective scene feature extraction module, a causal heuristic position information feature extraction network and a scene-position based multi-modal fusion module, wherein the causal heuristic position information feature extraction network is composed of a tree structure semantic map encoder and a pseudo map feature extraction module.
[0078] When the multi-modal dangerous state recognition network based on scene-position information is applied, the scene feature f s is first extracted from the target first perspective image by the first perspective scene feature extraction module.
[0079] Then the tree structure semantic map encoder maps the one-dimensional position information of the old person's travel to the two-dimensional image space using the real-time position information and the historical activity data of the time period Δt, generates a pseudo map image and constructs a pseudo map image sequence of {x t ,x t-1}. In the pseudo map feature extraction module, the adjacent two frames of pseudo maps x t ,x t-1 extract static trajectory features f 1,t ,f 1,t-1 through the SAN-18 network in turn; then the {f 1,t ,f 1,t-1} are input into the LSTM network to mine the dynamic trajectory change law f2 in the adjacent pseudo map images, and then the trajectory change law feature f2 is spliced with the static trajectory feature f 1,t at t time to constitute the position feature floc .
[0080] Finally, the fusion of scene features f s and location features f loc and the Softmax classifier realize the mapping of the fusion features to the dangerous state.
[0081] Specifically, in an embodiment, the first perspective scene image feature extraction module utilizes a transfer learning strategy. Specifically, on the basis of retaining the model parameters of the originally trained ResNet50 network, the ResNet50 model is fine-tuned in combination with the 5 common scene image data (road, stairs, riverbank, forest, and street) actually collected in the present application to adapt to the brightness, definition, and other characteristics of the photos taken by the intelligent helmet of the present system. During the fine-tuning process, the parameters of the shallow convolutional layers are frozen, and the high-level network parameters are adjusted and updated, thereby enhancing the semantic feature extraction effect on the present data set on the basis of retaining the bottom feature extraction capability (such as edges and textures). After the first perspective scene image is subjected to convolution and pooling processing by the ResNet50, the feature dimension is reduced through the fully connected layer, and the scene feature vector f s is obtained, which facilitates the splicing and fusion of the location features generated by the pseudo-map feature extraction module, and constructs a joint feature space of scene and location.
[0082] In order to realize the dimensional alignment between the one-dimensional GPS data and the two-dimensional image data, a tree structure semantic map encoder is matched in the cloud server to facilitate the mapping of the one-dimensional GPS data to the two-dimensional image space. In the present application, the tree structure semantic map encoder is constructed to solve the problem of large dimensional difference between the one-dimensional location information and the two-dimensional scene image features. Specifically, in an embodiment, the one-dimensional location vector (location information) is mapped to the two-dimensional image space (pseudo-map image) to facilitate the subsequent location feature extraction by the two-dimensional convolution kernel and the multi-modal fusion with the first perspective scene features.
[0083] The tree structure semantic map encoder visually presents the activity trajectory of the elderly on the map while increasing the semantic information of the location dangerous state by changing the visual features of the action trajectory. The working process of the encoder is shown in Figure 4 , which can be specifically divided into 5 parts of original pseudo-map generation, historical activity area generation, data processing and feature extraction, tree structure reasoning model, and dynamic trajectory generation.
[0084] The working process of the tree structure semantic map encoder is shown as follows:
[0085] S1: Original pseudo-map generation
[0086] A blank pseudo map is generated using the Folium library in Python to plot the dynamic update of the old person's travel trajectory and save a frame of pseudo map picture at intervals of Δt (Δt = 5 min in this application). After the pseudo map is generated, the action trajectory on the pseudo map is always accumulated until it is closed. If there is no new data received in the database for a period of time, it is considered that the elderly has stopped using the device, and the pseudo map is closed.
[0087] First, the activity location point set P Δt,o 0,o ,…,p n,o} is extracted from the observed old person's travel data (where the subscript o represents real-time data), and the old person's travel starting point coordinate p 0,o and the ending point coordinate p n,o in the current time interval Δt are determined. Then the start and end positions in this frame of pseudo map image are determined on the blank map, and the starting point p 0,o =(x 0,o ,y 0,o ) is marked with a pentagram icon, and the ending point p n,o =(x n,o ,y n,o ) is marked with an "i" icon. Finally, the old person's travel time T Δt , the ground speed V Δt in the position information are extracted from the real-time data, and combined with the activity location point set P Δt,o to form the set {P Δt,o ,T Δt ,V Δt}, which is transmitted to the data processing and feature extraction part.
[0088] S2: Historical activity area generation
[0089] Assuming that the update frequency of the old person's historical activity data is set to once a week (the guardian can adjust the update frequency of the historical activity data through the visual monitoring and intelligent early warning software), the position information stored in the database last week is used as the old person's historical activity data this week. The old person's historical activity location point set P T,h ={p 0,h ,…,p n,h} is extracted from the historical activity data, where the subscript T represents the time span of the historical activity data (set to 1 week in this application), and h represents the historical activity data. These positions are marked on the pseudo map in the form of heat points to form the historical activity area. Then the extracted historical activity and position point set P T,h is also transmitted to the data processing and feature extraction part.
[0090] S3: Data processing and feature extraction
[0091] The minimum safe distance (MSD), the travel time period (Period), and the average speed (Avspeed) are extracted by performing a data analysis process shown below on the set {P Δt,o , T Δt , V Δt} and P T,h .
[0092] (1) The minimum safe distance (MSD) of the elderly from the historical activity area.
[0093] In the experiment, it is assumed that a circle region with the elderly's familiar place as the center and the MSD as the radius is the elderly's familiar region, and the risk coefficient in this range is low. If there is historical activity data (HD = 1) in the system, the distance d0, d1, …, d n,o of the end position p n,o = (x n,o , y T,h ) of each pseudo map image to all position points in the historical activity position set P m is calculated using formula (3.1), and the minimum value of these distances is taken as the minimum safe distance MSD = min{d0, d1, …, d m}.
[0094]
[0095] If there is no historical activity data (HD = 0) in the system, the distance d of the end position p n,o = (x n,o , y n,o ) to the start position p 0,o = (x 0,o , y 0,o ) is calculated using formula (3.2), and the distance MSD = d is taken as the minimum safe distance.
[0096]
[0097] According to whether the system has historical activity data HD = 1 (HD = 0), the MSD is divided into four levels of attributes of safe distance, low-risk distance, medium-risk distance, and high-risk distance according to different thresholds. The specific division rules are shown in Table I.
[0098] Table I Thresholds of Mindistance in Two Cases
[0099]
[0100] (2) The elderly's travel time period (Period). According to the time period T ΔtPeriod is divided into two categories of 0 / 1, if the current T Δt Between 08:00 and 20:00, Period = 1; if the current T Δt Between 20:00 and 08:00, Period = 0.
[0101] (3) The average speed of the elderly (Avspeed). First, the ground speed V Δt (unit: knots) collected by the GPS module is converted into the travel speed (v) of the elderly (unit: m / s) by formula (3.3).
[0102]
[0103] Wherein: k is the ground speed, 1852 is the conversion coefficient of nautical miles to meters, and 3600 is the conversion coefficient of hours to seconds.
[0104] Then the average value of all travel speed data in the interval Δt is taken as Avspeed. The estimated walking speed of the elderly over 65 years old is between 0.6-1.2 m / s. Therefore, the normal average speed of the elderly is set to 0.6 m / s<Avspeed<1.2 m / s; when Avspeed>1.2 m / s, the speed is too fast; when Avspeed<0.6 m / s, the speed is too slow, as shown in Table II. According to common sense, it is known whether the speed is too fast or too slow, and the probability of danger encountered by the elderly is higher than that of normal speed.
[0105] Table II Average walking speed of the elderly in different states
[0106]
[0107] S4: Tree structure reasoning model
[0108] Finally, the above three key indicators need to be input into the tree structure reasoning model for reasoning the position danger state as shown in Figure 5 The additional semantic information of the current travel position danger state is inferred.
[0109] According to the inferred additional semantic information, the color and width visualization rules of the elderly travel trajectory are formulated, as shown in Table III. When the inference result is safe state, the trajectory color is green, and the width is 8; when the inference result is low danger state, the trajectory color is red, and the width is 10; when the inference result is medium danger state, the trajectory color is purple, and the width is 12; when the inference result is high danger state, the trajectory color is black, and the width is 14.
[0110] Table III Different danger levels and trajectory parameters
[0111]
[0112] The method of converting the semantic representation of the dangerous state of the travel position of the elderly into color texture features in the image can help the computer understand the danger caused by the position distance and other factors in the travel process of the elderly, and can improve the accuracy of the risk state evaluation by fusing scene-position features.
[0113] S5: Dynamic trajectory generation
[0114] Based on the visual representation rules shown in Table III, the dynamic trajectory of the elderly travel is generated on the pseudo map. If HD=0, the generated pseudo map sample is as shown in Figure 6 ; if HD=1, the generated pseudo map sample is as shown in Figure 7 .
[0115] The adjacent two pseudo map images are arranged into a sequence image {x t ,x t-1}, which is input into a self-attention network SAN for conversion and aggregation of GPS state features to obtain {f 1,t ,f 1,t-1}, which is then input into a neck LSTM to extract GPS state change rule features f2, and finally, f2 and f 1,t are spliced to form GPS joint features f loc . The self-attention network SAN adopts a structure similar to the residual connection and block connection of ResNet. SAN-X represents a network connected by X SAblocks. The number of layers of the neck LSTM is 3, which captures the relationship between the current input and the historical state through the combination of the forgetting gate, the input gate and the candidate state, and extracts the GPS state change rule.
[0116] The scene-position based multi-modal fusion module is used to fuse the scene features f s output by the scene feature extraction module with the GPS joint features f loc output by the GPS extraction module, and to evaluate the risk level, and the specific process is as follows:
[0117] First, the scene features f s output by the scene feature extraction module are spliced with the GPS joint features f loc output by the GPS extraction module to obtain scene-GPS joint features f, as shown in equation (4):
[0118] f=Concat(f s ,f loc )(4)
[0119] Then, f is input into the SENet channel attention module to determine the importance (i.e., weight) of each channel of f, focusing on those channels containing key information and giving less attention to those channels with less information. The output features are then flattened, as shown in equation (5):
[0120] f SE = Flatten(SE(f)) (5)
[0121] f SE is input into a fully connected layer using a stack of two linear layers for hazard state recognition, mapping the learned features to the sample label space and outputting the probability p of the four hazard states i , as shown in equation (6).
[0122] p i = Softmax(Linear(Relu(Linear(f SE )))) i∈{safe, low, medium, high} (6)
[0123] The application software is designed on the guardian's mobile phone and is used to display the information output by the scene-location-based multi-modal fusion module. In an embodiment, please refer to Figure 8 The cross section of the application software is composed of hazard state evaluation, first perspective scene, dynamic map, and real-time location of the user, wherein the hazard state evaluation is located at the upper end of the cross section, as shown in the yellow box at the upper end in Figure 8 , which is used to display the level (including: safe, low, medium, and high) evaluated by the scene-location-based multi-modal fusion module; the first perspective scene is directly below the hazard state evaluation, which is used to directly display the current first right-angle scene image of the user, enabling the guardian to intuitively understand the current surrounding environment of the user; the dynamic map is below the first perspective scene map, which is used to track the specific location of the user in real time and display the activity track of the user within a certain time range through a heat map; and the real-time location of the user is located in the yellow box of the cross section section, which displays the positioning information of the user in the form of text.
[0124] In an embodiment, the application software also has an active early warning mechanism, which automatically performs a warning action according to the hazard state evaluation level:
[0125] When the elderly person is in a high-risk state, the software directly dials the guardian; when in a medium-risk state, it sends a pop-up window to remind the guardian; and when in a low-risk state, it notifies the guardian through a short message.
[0126] The above content is described in combination with specific experiments as follows:
[0127] I. Experimental data and environment
[0128] The system automatically extracts the file name and UTC time field in the location information of the scene image, and pairs the data with the same file name and time field as shown in Figure 9 . However, due to the different resolutions of image and location information collection, multiple location information is matched with the same first-view scene image based on the collection resolution of image data, as shown in . To reduce the data storage cost of the server, the stored data will be cleaned and archived regularly, and the historical data that is no longer needed will be cleaned in time.
[0129] Figure 11 The first-view scene image data collected is classified into four dangerous states using manual annotation, and the data is divided into four categories according to the content of the scene image, namely safe state (486 images), low-risk state (896 images), medium-risk state (773 images), and high-risk state (1075 images). According to the classification rules of the tree structure semantic map encoder, the data segments in the location information are classified into four dangerous states, namely safe state (1612 segments), low-risk state (1017 segments), medium-risk state (499 segments), and high-risk state (102 segments). Finally, the target dangerous state label of the final elderly travel data pair is obtained by comprehensive evaluation according to the dangerous state of the two kinds of data, as shown in
[0130] The distribution of the final target dangerous state label data is as follows: safe state (880 data), low-risk state (888 data), medium-risk state (369 data), and high-risk state (1093 data). The distribution of the data label at each stage is shown in Figure 12 , and the generated data pair example is shown in Figure 13 . Finally, the data is divided into training set and validation set in the ratio of 9:1.
[0131] This experiment was conducted on a server, and the experimental environment is shown in Table IV. The GPU card of the server is NVIDIA GeForce RTX 2080Ti, the CPU is Intel Core i9-10900X, the deep learning framework is PyTorch 2.1.2+CUDA 12.1, the operating system is Linux, the programming language is Python, the version is 3.10.9, and the database management system is Microsoft SQL Server Management Studio 18, which is used to store and manage the GPS and processed location information data collected by the wearable device, as well as system log information.
[0132] Table IV Experimental environment
[0133]
[0134] Before model training, the image data input into the model needs to be preprocessed as follows: the input image is scaled to a size of 227x227, then randomly horizontally flipped to enhance data diversity, and then normalized according to the predefined RGB channel mean [0.4725, 0.4652, 0.4438] and standard deviation [0.2471, 0.2447, 0.2542] to ensure balanced data distribution.
[0135] The model training uses the Adadelta optimizer with an initial learning rate of 0.01 and a weight decay rate of 0.5. The batchsize is set to 10, the training period is 150 epochs, and the learning rate is dynamically adjusted every 20 epochs to promote model convergence. The loss function is CrossEntropyLoss(), which is used to optimize the performance of the multi-classification task.
[0136] II. Evaluation indicators
[0137] In this application, considering the problem of data class imbalance, weighted precision (P weight ), weighted recall (R weight ), weighted F1 score (F1 weight ), average loss (Loss) and accuracy (Acc) are used to evaluate the model performance. The calculation formulas are shown in formulas 7, 8 and 9:
[0138]
[0139]
[0140] The weight value of each class of sample is calculated using formula 10.
[0141]
[0142] where N i represents the total number of samples of the i-th class in the entire data set, w i represents the proportion of each sample in the entire data set. L represents the length of the entire data set.
[0143] The weighted precision is calculated using formula 11, and R weight and F1 weight can be obtained in the same way.
[0144]
[0145] The accuracy (Acc) is calculated using formula 12.
[0146]
[0147] III. Ablation experiments
[0148] In the process of building the optimal model, various network structures and module placement positions are tried in 10 groups of ablation experiments, proving the effectiveness of the idea of using scene-location multimodal information for elderly dangerous state assessment. In the experiment, the position information is first converted from one-dimensional information to two-dimensional image space by the tree structure semantic map encoder, and then input into the network model for subsequent operation.
[0149] Single-modal network structure, using single data for elderly dangerous state assessment. In this application, 2 groups of single-modal experiments are conducted, namely map single-stream network (SSNM) and scene single-stream network (SSNS), and the network structure is as shown in Figure 14 SSNM: This model inputs the generated pseudo-map image into the backbone network ResNet18 for position information feature extraction, and then performs dangerous state recognition through the fully connected layer. SSNS: This model inputs single-modal scene image into the backbone network ResNet18 for scene feature extraction, and then performs dangerous state recognition through the fully connected layer.
[0150] Multi-modal network structure, using scene and location data for elderly dangerous state assessment. The subsequent 8 groups of experiments in this application belong to multi-modal experiments, and the 3rd-7th groups of experiments belong to network structure adjustment experiments in multi-modal network, which include the adjustment process from double-stream network structure to triple-stream network structure, and the network structure diagram in the experiment is as shown in Figure 15 .
[0151] The 8th-10th groups of experiments perform backbone network adjustment experiments in multi-modal network based on triple-stream network, and the network structure diagram is as shown in 16.
[0152] Specifically, the multi-modal fusion stage in the double-stream network only involves two feature fusions, as shown in Figure 15 , including three groups of experiments of map-scene double-stream network I (MSDSN I), map-scene double-stream network I (+SE) (MSDSN I (+SE)), and map-scene double-stream network II (MSDSN II). Among them, MSDSN I: input scene and pseudo-map image data into ResNet18 to extract scene and position information features, then perform channel dimension splicing fusion of the two features, and finally perform elderly dangerous state recognition through the fully connected layer.
[0153] MSDSN I (+SE): First, the multi-modal fused features in the MSDSN I network are input into the channel attention mechanism SENet to adjust the weight of the fused features, emphasizing important features and suppressing irrelevant features. Then the obtained features are input into the fully connected layer for hazard state recognition.
[0154] MSDSN II: In this application, ResNet18 is used instead of the head CNN network in the CNN-LSTM hybrid model framework, and then ResNet18-LSTM hybrid model is constructed to extract the dynamic trajectory features of the pseudo map image; ResNet18 is still used for scene feature extraction, then the dynamic trajectory features and scene features are multi-modal fused as described above, and finally input into the fully connected layer for hazard state recognition.
[0155] The multi-modal fusion stage in the three-stream network involves three feature fusions. The three-stream network includes map-scene three-stream network (MSTSN), map-scene three-stream network (+SE) (MSTSN (+SE)), map-scene three-stream network (MSTSN II (+SE)), map-scene three-stream network (MSTSN III (+SE)) and map-scene three-stream network (MSTSN IV (+SE)) 5 groups of experiments.
[0156] The first two groups adjust the model structure, as shown in Figure 15 (4) and (5), and the last three groups mainly adjust the backbone network of the model to improve the feature extraction ability of the model, as shown in Figure 16 .
[0157] MSTSN: Based on the double-stream network MSDSN II, the position feature extraction part is improved using the convolutional autoencoder (CAE) and CNN-LSTM hybrid network structure. ResNet18 is used instead of the convolutional autoencoder (CAE) in the literature to extract the static trajectory features of the pseudo map image, and ResNet18-LSTM hybrid network is still used to extract the dynamic trajectory features of the pseudo map image sequence, then form a three-stream network structure with the scene feature part. The obtained three features are spliced and fused in the channel dimension, and then input into the fully connected layer for hazard state recognition.
[0158] MSTSN (+SE): After the feature fusion of the MSTSN network is completed, the output features are input into the SENet module for feature channel weight redistribution, and finally input into the fully connected layer for hazard state recognition.
[0159] The network structure of the multi-modal dangerous state recognition network model based on scene-position information in the present application has been basically determined through the above-mentioned 7 groups of experiments. However, considering that the position information is mainly saved in the track changes of the pseudo map image, it is easy to be disturbed by other surrounding background information. This feature has a great difference from the scene information distribution, so the following experiments try to adjust the backbone network of the position information feature extraction part, and replace the original ResNet18 with a self-attention mechanism SAN network. In addition, in order to fully utilize the scene information, the backbone network of the scene feature extraction part is also adjusted in the subsequent experiments in the present application.
[0160] MSTSN II (+SE): replacing the ResNet18 part of the position information feature extraction part with a self-attention network SAN-14 based on MSTSN (+SE), and keeping other parts unchanged.
[0161] MSTSN III (+SE): replacing the ResNet network of the deepened scene feature extraction part with ResNet50 based on MSTSN II (+SE).
[0162] MSTSN IV (+SE): replacing the SAN-14 with SAN-18 based on MSTSN III (+SE) to improve the feature extraction capability of the backbone network of the position information feature extraction part.
[0163] The above is a detailed introduction to the network models in the ablation experiment, and the ablation experiment results are specifically shown in Table V and Figure 17 . Figure 17 The average loss value change process during the training process of each network model in the ablation experiment is shown, and from Figure 17 , it can be clearly seen that the loss value reduction process of the 10 groups of network models is divided into 3 categories.
[0164] The first category is the two groups of single-modal network experiments. It can be seen that although the training process of these two models converges rapidly, the loss value does not decrease significantly, and the final loss value stabilizes at a high level, reflecting that the difference between the model prediction value and the data set label in the model training process is large. According to the experimental results shown in Table V, it can be seen that the weighted precision (P weight ), weighted recall (R weight ), weighted F1 score (F1 weight ) and accuracy (Acc) of the two network models on the validation set are all lower than 60%, indicating that using single modal data cannot well recognize the dangerous state in the travel process of the elderly. However, overall, the experimental results of using single scene data for dangerous state recognition are better than those of using single position information for dangerous state recognition.
[0165] The second category mainly includes five groups of experiments on adjusting the network structure, Figure 17 The average loss value during the training process of these five groups of experiments is not much different, but the final average loss value of MSTSN(+SE) is slightly lower than that of the other four groups. According to the results of the five groups of experiments shown in Table V, there are some points worth discussing. Adding the SENet module to MSDSN I seems to have no effect on improving the experimental results of the model except for lowering the average loss value. MSDSN II is an adjustment to the position information feature extraction part, replacing the original ResNet18 network with a ResNet18-LSTM hybrid model, and the input data has also changed from the pseudo map image x t to the pseudo map image sequence {x t , x t-1}, hoping to use the time sequence features of the position information changes in the subsequent dangerous state recognition process, and removing the SENet module in the MSDSN II network, but the results of this network seem to be worse than the experimental results of MSDSN I(+SE), indicating that the position information features of the pseudo map at the current time have a greater impact on the accurate recognition of the dangerous state. Therefore, the MSTSN experiment is conducted, which adds the position information features of the pseudo map at the current time extracted by ResNet18 in the multi-modal fusion process of MSDSN II, and the experimental results have changed significantly, with an accuracy rate even better than MSDSN I. MSTSN(+SE) adds the SENet module to MSTSN, and the accuracy rate of this model is improved by 0.45% compared with MSTSN.
[0166] The third category includes three groups of experiments on backbone network improvement. According to Figure 17 It can be seen that after adjusting the backbone network of the position information feature extraction part from ResNet18 to SAN-14, the convergence speed of the entire model is significantly improved, and the loss value is more significantly reduced compared with the previous seven groups of experiments. According to Table V, compared with MSTSN(+SE), the P weight of MSTSN II(+SE) is improved by 5%, the R weight is improved by 4%, the F1 weight is improved by 4%, and the Acc is improved by 4.06%. This result shows that the self-attention mechanism can better capture the trajectory change rules in the pseudo map image than the global feature extraction network, and can effectively improve the model performance. On the basis of MSTSN II(+SE), the feature extraction capability of the scene feature extraction part is improved, and the P weight , R weight , F1 weight, Acc each index is improved by 1%~2% on the basis of MSTSN II (+SE), among which the average loss value decreases significantly, proving that the information in the scene image was not fully utilized in the previous experiment, and appropriately deepening the backbone network can help the model better understand the image information and improve the accuracy of the model in identifying dangerous states.
[0167] Finally, MSTSN IV (+SE) increases the number of blocks of the self-attention network on the basis of MSTSN III (+SE) from SAN-14 to SAN-18. Compared with the experimental results of MSTSN III (+SE), it is shown that improving the feature extraction capability of the position information feature extraction part can also effectively improve the ability of the model in identifying dangerous states, and the P weight of MSTSN IV (+SE) is improved by 4%, the R weight is improved by 5%, the F1 weight is improved by 5%, the Acc is improved by 5.1%, and the average loss value also decreases accordingly. Thus, the multi-modal dangerous state recognition network model based on scene-position information proposed in the present application is determined as MSTSN IV (+SE), and the superiority of the model will be proved by comparison with other models.
[0168] Table V Experimental results
[0169]
[0170] As Figure 10 shown, the map-scene three-flow network (+ attention channel) IV proposed in the present application is superior to other models in terms of convergence speed and convergence effect. Through the performance of each model on the test set, it can be observed that the accuracy of the multi-modal model in identifying the dangerous state of the elderly is significantly higher than that of the single-flow network model.
[0171] The model proposed in the present application exhibits more excellent performance than other models in each evaluation index. Among them, the optimal model map-scene three-flow network (+ attention channel) IV compared with the map-scene three-flow network (MSTSN) (+ attention channel), the precision is improved by 0.1, the recall is improved by 0.09, the F1 score is improved by 0.11, the loss value is reduced by 0.25, and the overall accuracy is increased by 9.16%. Compared with the map single-flow network (SSNM) (scene single-flow network (SSNS)) at the beginning, the improvement in each index is more significant. The precision is improved by 0.31 (0.21), the recall is improved by 0.25 (0.18), the F1 score is improved by 0.30 (0.20), the loss value is reduced by 0.41 (0.91), and the overall accuracy is increased by 25.09% (17.81%).
[0172] Four, comparative experiment
[0173] To verify the conversion of the user's text location information into a pseudo-map image in this application, the following verification method is adopted in this application. First, the text location information and the first-person view scene image of the old person are input into the MDFNet for dangerous state recognition. The text location information can contain all the information required for the map generator to generate a map. In addition, three other multi-modal fusion networks, FGT-Net, YOLOv5s-DMF, and FSA-UNet, are used for dangerous state discrimination comparative experiments. The specific results are shown in Table VI.
[0174] Among them, the MDFNet model attempts to perform dangerous state recognition by fusing text location information and scene image features. However, due to the large difference in feature dimensions between the two types of data, it is difficult to balance the weights of different features, which affects the overall performance of the model.
[0175] The FGT-Net model relies on typical sample images to guide the extraction of image features and combines corresponding text data for classification. This method is suitable for data sets with small group differences, but it does not match the properties of our data set, so it fails to achieve the expected results.
[0176] The YOLOv5s-DMF model uses YOLOv5s to extract the features of RGB images and infrared images for fusion. Although it can capture the global features of the scene, it is not enough to capture the small changes in the pseudo-map image, which leads to poor performance on our data set, and confirms the correctness of the selection of the self-attention module SAN in this application to extract small changes in the map image.
[0177] The FSA-UNet model extracts remote sensing and street view image features respectively and uses a self-attention mechanism for multi-modal fusion. Although it also focuses on the overall features of the image in the feature extraction stage, the feature fusion through the self-attention mechanism significantly improves the accuracy and precision, outperforming the YOLOv5s-DMF model.
[0178] The model proposed in this application uses a self-attention module to capture the small changes in the action trajectory in the map and uses ResNet to extract the overall features of the first-person view image of the old person. Then, the map and scene features are fused to accurately recognize the dangerous state of the elderly. Compared with the prior art, the model proposed in this application performs well in terms of precision, accuracy, and loss value, and is more suitable for our data set. This achievement not only demonstrates the superiority of the model in this application, but also provides a new technical path for the effective recognition of the dangerous state of the elderly.
[0179] Table VI Comparative Experiment Results
[0180]
[0181] V. Real scene test experiment
[0182] The system real scene test process mainly includes a "departure-return" complete travel activity by the experimental personnel wearing the wearable danger perception device Figure 18 ) and the danger state assessment of each scene with / without historical activity data Figure 19 . The system real scene test process sets the update frequency of the historical activity data in the system to the past one hour, and each subgraph in the real scene instance respectively shows the third person view of the old person's travel (left side) and the display result of the guardian end visualized guardianship intelligent early warning application software interface (right side).
[0183] As shown in Figure 18 , the simulation old person wears a smart helmet to perform a "departure-return" complete travel, wherein:
[0184] The third person view and the interface details of the visualized guardianship intelligent early warning software during the start, the middle, and the end of the departure process of the simulation old person from home to the end point of the riverside pavilion are shown in Figure 18 (a)-(c),
[0185] Figure 18 The third person view and the interface details of the visualized guardianship intelligent early warning software during the start, the middle, and the end of the return process of the simulation old person from the riverside pavilion to home are shown in
[0186] By comparison, it is not difficult to see that there is a significant difference in the system's assessment of the old person's current danger state between the departure and return processes under the same environmental conditions. In the departure process, since there is no historical travel data of the old person in the system, the target danger state will be higher than that with historical data. Therefore, the danger state is low at the beginning of the departure process at the door, the danger state becomes medium in the middle, and the danger state becomes high at the end of the riverside pavilion. In the return process, the danger state is medium at the beginning of the return process at the riverside pavilion, the danger state becomes low in the middle, and the danger state becomes safe at the end of the home. Through the real scene test of the "departure-return" complete travel, it is reflected that the system indeed combines the scene and location information to make a logical assessment of the old person's current danger state.
[0187] As shown in Figure 19 , the system function is tested in each actual scene. According to Figure 19 (a)-(b), when the old person turns from the normal walking road to the river beside, the first perspective scene perceived by the smart helmet changes from the road to the riverside. In the case where the location information basically does not change, the system can accurately identify that the probability of the old person's danger increases according to the switching of the first perspective scene, and the inferred danger state changes from safe to medium. Figure 19(c)-(d)、 Figure 19 (e)-(f) and Figure 19 (g)-(h) The test results of the three groups can prove that the system can sensitively combine the changes of the first perspective scene of the elderly to reasonably evaluate the current dangerous state of the elderly.
[0188] The test results of the above experiments and Figure 18 reflect that the system can reasonably evaluate the dangerous state of the elderly by combining scene and location information, that is, it can adjust the evaluation result of the dangerous state of the elderly in time according to the changes of the first perspective scene image or location information. According to Figure 19 (i)-(j), it can be found that the first perspective scene does not change, Figure 19 (i) there is no historical activity data of the elderly in the system under the scenario shown in (i), Figure 19 (j) there is historical activity data of the elderly in the system under the scenario shown in (j), but the final system evaluation result of the dangerous state of the elderly at that time is the same, which shows that the system can skillfully adjust the influence of the first perspective scene and location information on the evaluation result of the dangerous state, so as to make accurate evaluation.
[0189] Through real scene testing, the effectiveness of the system in identifying the dangerous state by combining scene and location information is verified from multiple angles. The occurrence of MCI elderly missing events has been affected by many factors such as environment, location, and the elderly's own situation, but the current solutions for this group of people missing often only consider one of these factors. The real scene test experiment of the system also proves the stability of the system in the running process, and the test results prove the effectiveness of combining multiple factors to reason the current dangerous state of the elderly, which is expected to provide a new research idea for the field of elderly care and anti-missing in the future.
[0190] In summary, the anti-missing multi-modal dangerous assessment and intelligent early warning system for intermittent senile dementia provided in the present application specifically solves the problem of the elderly with dementia encountering danger and missing outdoors. Through theoretical and real scene testing, it is shown that the early warning system in the present application can complement modalities by fusing the current GPS data of the user and the first perspective scene, and more objectively and comprehensively realize multi-dimensional dangerous state assessment of the elderly. In addition, the system of the present application has high detection accuracy and good real-time performance, and provides a new strategy for solving the problem of the elderly missing, which has certain social significance and wide application potential.
Claims
1. A multimodal risk assessment and intelligent early warning system for preventing elderly people with intermittent dementia from wandering, characterized in that: The invention includes a smart helmet, a scene-location information-based multimodal hazard identification network, and early warning software. The smart helmet is used to collect multimodal data and transmit the data to the scene-location information-based multimodal hazard identification network via a wireless network. The multimodal data includes first-view image information and the user's location information. The scene-location information-based multimodal hazard identification network is used to evaluate multimodal data and display the evaluation results in intuitive text form on the early warning software at the guardian's end. The scene-location information-based multimodal hazard identification network adopts a tree-structured semantic map encoder to align the dimensions of one-dimensional GPS information with those of two-dimensional scene image information.
2. The multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 1, characterized in that: The scene-location information-based multimodal hazard identification network comprises a first-view scene feature extraction module, a causal heuristic location information feature extraction network, and a scene-location-based multimodal fusion module. The causal heuristic location information feature extraction network consists of two parts: a tree-structured semantic map encoder and a pseudo-map feature extraction module.
3. The multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 2, characterized in that: The first-view scene feature extraction module is used to extract scene features f from the target first-view image. s Then, the tree-structured semantic map encoder uses real-time location information and historical activity data over a time interval of Δt to map the one-dimensional location information of the elderly person's travel to a two-dimensional image space, generating a pseudo-map image and constructing {x t ,x t-1 A sequence of pseudo-map images; In the pseudo-map feature extraction module, adjacent pseudo-map frames x t x t-1 Static trajectory features f are extracted sequentially through the SAN-18 network. 1,t ,f 1,t-1 ; Then {f 1,t ,f 1,t-1 The input is fed into an LSTM network to mine the dynamic trajectory change pattern f2 in adjacent pseudo-map images. Then, the trajectory change pattern feature f2 is compared with the static trajectory feature f at time t. 1,t splicing constitutes positional features f loc ; Finally, scene features f are performed in the multimodal hazard assessment module. s and location features f loc The fusion of features and the Softmax classifier are used to map fused features to dangerous states.
4. The multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 3, characterized in that: The workflow of the tree-structured semantic map encoder is as follows: S1: Generating the original pseudo-map A blank pseudo-map is generated using Python's Folium library to plot the dynamically updated movement trajectory of the elderly during their travels, and a frame of the pseudo-map image is saved at intervals of Δt (Δt = 5 min in this application). After the pseudo-map is generated, the movement trajectory on the pseudo-map continues to accumulate until it is closed. When no new data is received in the database for a period of time, it is considered that the elderly have stopped using the device, and the pseudo-map is closed. S2: Generation of Historical Activity Areas Assuming the update frequency of the elderly's historical activity data is set to once a week (caregivers can adjust the update frequency of historical activity data through visual monitoring and intelligent early warning software), the location information of the previous week stored in the database is used as the elderly's historical activity data for the current week; the set of historical activity location points P of the elderly is extracted from the historical activity data. T,h ={p 0,h ,…,p n,h }, where the subscript T represents the time span of the historical activity data (set to 1 week in this application), and h represents the historical activity data. These locations are marked on the pseudo-map as heat points to form historical activity areas; then the extracted historical activity and location point set P is... T,h It is also passed to the data processing and feature extraction section; S3: Data Processing and Feature Extraction By analyzing the set {P} Δt,o ,T Δt V Δt } and P T,h The following data analysis process was performed to extract three key indicators: minimum safe distance (MSD), travel period, and average speed (Avspeed). S4: Tree-structured reasoning model Finally, the above three key indicators are input into the tree-structured reasoning model used to infer the additional semantic information of the current travel location danger status. Based on the extra semantic information inferred, color and width visualization rules for the travel trajectories of the elderly were formulated, as shown in Table III; S5: Dynamic Trajectory Generation Based on the visual representation rules shown in Table III, dynamic travel trajectories of elderly people are generated on the pseudo-map. Table III. Different Hazard Levels and Trajectory Parameters 5. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 4, characterized in that: The specific steps in S1 are as follows: First, extract the activity location set P from the observed travel data of the elderly. Δt,o ={p 0,o ,…,p n,o (where the subscript o represents real-time data), and determine the coordinates o of the elderly person's starting point for travel within the current time interval Δt. 0,o and the endpoint coordinates o n,o ; Next, determine the start and end positions of this frame of the pseudo-map image on the blank map, and mark the starting position p with a pentagram icon. 0,o =(x 0,o ,y 0,o The endpoint p is marked with an "i" icon. n,o =(x n,o ,y n,o ); Finally, the travel time T of the elderly was extracted from the real-time data. Δt The ground velocity V in the location information Δt And combined with the activity location point set P Δt,o Constitute the set {P Δt,o ,T Δt V Δt } and then transmit it to the data processing and feature extraction section.
6. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 5, characterized in that: The specific steps in S5 are as follows: Arrange two adjacent pseudomap images into a sequence of images {x} t ,x t-1 }, which is then input into a self-attention network (SAN) to transform and aggregate GPS state features, resulting in {f 1,t ,f 1,t-1 Then, input the data into the neck LSTM to extract the GPS state change pattern feature f2. Finally, compare f2 with f... 1,t splicing constitutes GPS joint feature f loc ; Among them, the self-attention network SAN adopts a structure similar to ResNet with residual connections and block connections; SAN-X represents a network composed of X SA blocks; the neck LSTM has 3 layers, and captures the relationship between the current input and the historical state by combining forget gate, input gate and candidate state to extract the pattern of GPS state change.
7. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 6, characterized in that: Scene features f are performed in the multimodal hazard assessment module. s and location features f loc The specific process of using the fusion and Softmax classifier to map fused features to dangerous states is as follows: First, the scene features f output by the scene feature extraction module are... s Joint GPS features f output by the GPS extraction module loc Channel splicing is performed to obtain the scene-GPS joint feature f, as shown in equation (4): f=Concat(f s ,f loc ) (4) Then, f is input into the SENet channel attention module to determine the importance (i.e., weight) of each channel of f, focusing on those channels containing key information and giving less attention to those channels with less information; then the output features are straightened, as shown in Equation (5): f SE =Flatten(SE(f)) (5) f SE The input uses a fully connected layer consisting of two stacked linear layers to identify dangerous states, mapping the learned features to a sample label space, and outputting the probabilities p of four dangerous states. i , as in equation (6) As shown: p i =Softmax(Linear(Relu(Linear(f SE ))))i∈{safe, low risk, medium risk, high risk}(6).
8. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 7, characterized in that: The application software is designed on the guardian's mobile phone and is used to display the information output by the scene-location-based multimodal fusion module. The interface of the application software consists of four parts: danger state assessment, first-person perspective scene, dynamic map, and the user's real-time location.
9. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 8, characterized in that: The application software also has an active early warning mechanism, which automatically executes early warning actions based on the level of danger assessment.
10. A multimodal risk assessment and intelligent early warning system for preventing wandering in elderly people with intermittent dementia according to claim 9, characterized in that: The smart helmet includes a safety helmet, a smart camera module, and a solar cell. The smart camera module consists of four parts: a camera sensor, a GPS module, a 5G communication module, and a power control module. The camera sensor is used to acquire the user's first-person perspective image information, and the GPS module is used to acquire the user's location information.