Efficient multi-mode tumble detection method in smart medical environment

By adopting multimodal fall detection method in a smart medical environment, using multimodal feature extraction, time alignment and deep fusion technologies, the problems of poor single-modal robustness and high multimodal calculation complexity in the existing technology are solved, and efficient, real-time and reliable fall detection effects are achieved.

CN120145302APending Publication Date: 2025-06-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510225848.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing fall detection technology mainly relies on single-modal data, with poor robustness and high false alarm rate. The multimodal method has technical difficulties in time alignment and feature fusion, resulting in the detection performance not reaching the ideal level, high computational complexity, difficult to guarantee real-timeness, and insufficient detection accuracy and adaptability in complex scenarios.

Method used

An efficient multimodal fall detection method in a smart medical environment is adopted. By collecting acceleration, audio and video multimodal data, the trained multimodal fall detection model is input, and multimodal feature extraction, time alignment, deep fusion and classification modules are used for processing. It combines Shapley value modal contribution evaluation and dynamic time alignment methods to achieve efficient multimodal data alignment, and feature fusion is performed through bottleneck mechanisms and cross attention mechanisms.

Benefits of technology

It improves the accuracy and robustness of fall detection, reduces the computational complexity, meets the real-time and reliability of fall detection in smart medical care, and enhances the system's adaptability in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145302A_ABST
    Figure CN120145302A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of multi-modal data processing, and relates to an efficient multi-modal fall detection method in an intelligent medical environment, which comprises the following steps: collecting multi-modal data and inputting the multi-modal data into a trained multi-modal fall detection model to obtain a detection result; the training process of the tumble detection model comprises the steps of collecting multi-modal data and inputting the multi-modal data into the multi-modal feature extraction module to obtain multi-modal features; inputting the multi-modal features into a multi-modal time alignment module to obtain multi-modal time alignment features; inputting the time alignment feature into a depth fusion module to obtain a depth fusion feature; inputting the deep fusion features into a classification module to obtain a detection result; calculating a multi-task loss function value, and updating model parameters according to the multi-task loss function value until a trained fall detection model is obtained; alignment of the multi-modal data is achieved through modal contribution evaluation and dynamic time warping, multi-modal data fusion is achieved in combination with a key modal and a bottleneck mechanism, the detection precision and robustness can be improved, and the calculation complexity can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent healthcare and multimodal data processing, and relates to an efficient multimodal fall detection method in an intelligent healthcare environment. Background Art

[0002] With the rapid development of Internet of Things (IoT) and artificial intelligence (AI) technologies, intelligent healthcare, as an emerging medical service model, has played an important role in improving medical efficiency and health monitoring capabilities. Among them, fall detection technology, as a core application in intelligent healthcare, is widely used in scenarios such as elderly health monitoring, real-time monitoring of people with limited mobility, and intelligent elderly care. It can alarm in time when people fall, effectively reducing the risk of injuries caused by falls.

[0003] Existing fall detection technologies mainly rely on unimodal data, such as detection methods based on accelerometers, audio, or video. The acceleration modality method identifies fall events by monitoring the acceleration changes of human movements, but it is prone to interference from daily rapid movements (such as sitting down, standing up, etc.) in practical applications, resulting in a high false alarm rate. The audio modality method detects falls by capturing specific acoustic features in the environment, but its detection robustness is poor due to the influence of environmental background noise (such as conversations, TV sounds). The video modality method identifies falls by analyzing human postures or action trajectories. Although it can provide intuitive action information, it is easily restricted by factors such as light changes, camera perspectives, and occlusion problems. These unimodal methods are highly dependent on the environment and are difficult to provide stable and reliable detection results in complex and variable practical scenarios.

[0004] In recent years, multimodal data fusion methods have gradually become a research hotspot. By combining acceleration, audio, and video modality data, the deficiencies of unimodal methods can be effectively compensated. However, there are still many technical difficulties in time alignment and feature fusion in existing multimodal methods, resulting in their detection performance not reaching the ideal level. First, there are differences in the sampling frequencies and time steps of different modality data, and existing alignment methods lack an efficient mechanism, making it difficult to achieve precise time alignment of multimodal data in dynamic and complex scenarios. Second, most existing feature fusion methods use simple concatenation or decision-level fusion, failing to fully explore the deep associations between modalities, resulting in more redundant information and low feature utilization efficiency. In addition, existing technologies generally lack a dynamic evaluation mechanism for modality contributions and cannot adjust the weights of modalities to the final detection result according to their importance, which may lead to the neglect of key modality information or the interference of noise data on the overall performance.

[0005] In addition, when traditional multi-modal methods handle high-dimensional video data and long-time series data, due to the high feature dimension and complex temporal modeling, the computational overhead increases significantly, making it difficult to meet the real-time requirements of application scenarios such as fall detection that require quick responses. Especially in complex scenarios (such as multi-background noise, light changes, and non-standard fall actions), existing systems are difficult to maintain efficient and stable detection effects due to inaccurate alignment and fusion of multi-modal data, as well as the inability to flexibly adjust the importance of each modality, resulting in obvious impacts on accuracy and adaptability.

[0006] In summary, the current fall detection technologies mainly have the following problems: single-modal methods have poor robustness and high false alarm rates; multi-modal methods lack efficient time alignment and feature fusion strategies, resulting in the failure to fully utilize the complementarity between modalities; the computational complexity of the system is relatively high, and real-time performance is difficult to guarantee; the detection accuracy and adaptability in complex scenarios are insufficient. Summary of the Invention

[0007] To solve the above problems of the existing technologies, the present invention adopts an efficient multi-modal fall detection method in an intelligent medical environment, including: collecting multi-modal data, inputting the collected multi-modal data into a trained multi-modal fall detection model to obtain a fall detection result; the multi-modal fall detection model includes: a multi-modal feature extraction module, a multi-modal time alignment module, a deep fusion module, and a classification module;

[0008] The training process of the multi-modal fall detection model includes:

[0009] S1. Collect multi-modal data, input the multi-modal data into the multi-modal feature extraction module to obtain multi-modal features;

[0010] S2. Input the multi-modal features into the multi-modal time alignment module to obtain multi-modal time alignment features;

[0011] S3. Input the multi-modal time alignment features into the deep fusion module to obtain deep fusion features;

[0012] S4. Input the deep fusion features into the classification module to obtain a fall detection result;

[0013] S5. Calculate the multi-task loss function value according to the multi-modal time alignment features, deep fusion features, and fall detection result, update the model parameters according to the multi-task loss function value, and when the loss function value is the smallest, obtain the trained multi-modal fall detection model.

[0014] The multimodal data includes acceleration data, audio data, and video data; the multimodal feature extraction module includes: an acceleration modality feature extraction module, an audio modality feature extraction module, a video modality feature extraction module, and a mapping and normalization module; each multimodal feature extraction module extracts features from the input acceleration data, audio data, and video data respectively to obtain multimodal initial features, and inputs the multimodal initial features into the mapping and normalization module respectively to obtain multimodal features F m ”; where m represents the modality.

[0015] The multimodal time alignment module processes the multimodal features as follows: quantifies the weights of each modality using a modality contribution evaluation method according to the multimodal features, and performs time alignment on the multimodal features using the dynamic time warping method according to the weights of each modality to obtain multimodal time-aligned features.

[0016] Quantifying the weights of each modality using a modality contribution evaluation method includes: constructing a modality set M = {s, a, v}, randomly generating N modality subsets S that do not contain modality m according to the modality set M k , calculating the modality value function of each modality subset S k The modality value function of Calculating the weight of the corresponding modality m according to the modality value function of the modality subset S k where s, a, and v represent the acceleration, audio, and video modalities respectively, k is the index of the modality subset, -Error(S ) represents the error value of performing time alignment on the modality subset S k , T k represents the time step of the features of modality m, F m ”[t], F m ”[t] represent the features of modality m and n at time step t respectively, D n ” is the Euclidean distance, S error / {m} is the subset obtained by removing modality m from the modality subset S k k

[0017] Performing time alignment on the multimodal features includes: constructing a weighted local distance between modalities according to the weights of each modality, constructing a cumulative distance matrix according to the weighted local distance between modalities, optimizing the cumulative distance matrix to obtain an optimal time alignment path, and adjusting the multimodal features according to the optimal time alignment path to obtain multimodal time-aligned features.

[0018] The deep fusion module processes the features after alignment of each modality as follows: compresses the multimodal time-aligned features using a bottleneck mechanism to obtain bottleneck features of each modality; performs interaction and fusion on the bottleneck features of each modality through a cross-attention mechanism to obtain deep fusion features. ​​

[0019] Interact with and fuse the bottleneck features of each modality, including: based on the bottleneck feature B of the acceleration modality s Calculate the query matrix Q s , based on the bottleneck feature B of the audio modality a Calculate the key matrix and value matrix, based on the bottleneck feature B of the video modality v Calculate the key matrix and value matrix, based on the query matrix Q s and the bottleneck feature B of the audio modality a Calculate the interaction result Attention based on the key matrix and value matrix of the query matrix Q a→s and the bottleneck feature B of the video modality s Calculate the interaction result Attention based on the key matrix and value matrix of the query matrix Q v and the bottleneck feature B of the video modality v→s ; Concatenate the interaction results Attention a→s , Attention v→s to obtain the deep fusion feature F fusion .

[0020] The classification module processes the deep fusion feature, including: performing a residual connection on the deep fusion feature F fusion with the multi-modal time-aligned feature to obtain the classification input feature F residual ; Normalize the classification input feature F residual to obtain the normalized feature Use the classifier to process the normalized feature to obtain the fall detection result.

[0021] The loss function L total = L CE + λ 1 L consistency + λ 2 L center ; where L CE represents the classification loss function, L consistency represents the modality consistency loss function, L center represents the center loss function, λ 1 and λ 2 are the weight parameters of the modality consistency loss and the center loss respectively.

[0022] The modality consistency loss function is:[[]]

[0023]

[0024] where F i,s , F i,a , F i,vrespectively represent the time-aligned features of the acceleration, audio, and video modalities of the i-th multimodal data, w sa is the weight between the acceleration modality and the audio modality, w sv is the weight between the acceleration modality and the video modality, N tr is the number of multimodal data.

[0025] Beneficial effects:

[0026] 1. The present invention uses the Shapley value modal contribution evaluation method to evaluate the contribution of each modality in the time alignment process, and combines the contribution of each modality in the time alignment process with the dynamic time warping method to achieve efficient alignment of multimodal data, improving the detection accuracy and robustness; 2. The present invention uses a bottleneck mechanism to compress the multimodal time-aligned features, reducing redundant features and highlighting key information, and uses a cross-attention mechanism to achieve information interaction between the features of the compressed key modality and the basic modality, improving the fusion efficiency of the multimodal time-aligned features and reducing the computational complexity. Therefore, it can not only improve the detection accuracy and robustness, but also significantly reduce the computational complexity, meeting the requirements of real-time and reliability for fall detection in intelligent healthcare. Description of the drawings

[0027] Figure 1 is a flowchart of an efficient multimodal fall detection method in an intelligent healthcare environment provided by an embodiment of the present invention;

[0028] Figure 2 is a structural diagram of an efficient multimodal fall detection method in an intelligent healthcare environment provided by an embodiment of the present invention. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0030] As Figure 1 、 Figure 2 shown, the present invention adopts an efficient multimodal fall detection method in an intelligent healthcare environment, including: collecting multimodal data, inputting the collected multimodal data into a trained multimodal fall detection model to obtain a fall detection result; the multimodal fall detection model includes: a multimodal feature extraction module, a multimodal time alignment module, a deep fusion module, and a classification module;

[0031] The training process of the multimodal fall detection model includes:

[0032] S1. Collect multimodal data, input the multimodal data into the multimodal feature extraction module, and obtain multimodal features;

[0033] The multimodal data includes acceleration modal data, audio modal data, and video modal data; the acceleration modal data and audio modal data are collected by a smart bracelet worn by the elderly to capture the movement changes of the human body and the sounds generated during the falling process; the video modal data is collected by a surveillance camera installed indoors or in the activity area to capture the human body movements and position dynamics.

[0034] Specifically, the collection of audio modal data includes: the audio modal data is collected by a microphone built into the smart bracelet, and the sampling frequency is 16 kHz. The microphone records the sound signals in the environment in real time to capture the typical sound features in the falling event, such as impact sounds or calls for help. The audio signal undergoes basic noise reduction processing during the collection process to ensure that the recorded sound data can reflect the key sound changes in the environment. The collection of video modal data includes: the video modal data is collected by a high-definition indoor camera installed in the user's activity area, the camera resolution is 1920×1080, and the frame rate is 30 fps. The camera records the user's dynamic behaviors in real time, covering the main activity areas of the user and recording the user's dynamic behaviors in real time. The collected video data provides a reliable basis for subsequent analysis of the user's behavior trajectories and dynamic changes before and after falling. The collection of acceleration modal data includes: the acceleration modal data is collected by a three-axis acceleration sensor built into the smart bracelet, and the sampling frequency is 50 Hz. The acceleration sensor records the motion states of the user in the X, Y, and Z directions in real time, and can effectively reflect the user's motion characteristics, including the motion change process before and after falling.

[0035] During the collection process, the audio, video, and acceleration data are all recorded based on the same clock system, which ensures the synchronization of the three modal data at the hardware level. The three modal data are recorded in real-time synchronization, providing a high-quality data basis for subsequent feature extraction, time alignment at the feature level, and modal fusion.

[0036] The multimodal feature extraction module includes: an acceleration modal feature extraction module, an audio modal feature extraction module, a video modal feature extraction module, and a mapping and normalization module; the acceleration modal feature extraction module, the audio modal feature extraction module, and the video modal feature extraction module respectively generate multimodal initial features by extracting the features of the corresponding modalities, providing a basis for subsequent time alignment and feature fusion.

[0037] The extraction of acceleration modal initial features includes:

[0038] Process the collected acceleration modal data through a Temporal Convolutional Network (TCN) to extract short-term and long-term dependence features of human motion.

[0039] The collected acceleration modal data is expressed as:

[0040]

[0041] where x s is the acceleration modal data, representing the acceleration signal of the time series, and T s is the time step, representing the number of time segments included in the signal; d s is the dimension of the signal. For example, for a three-axis acceleration sensor, d s = 3, which are the acceleration values in the x, y, and z axis directions respectively.

[0042] Model the acceleration signal through a Temporal Convolutional Network (TCN) to generate a feature representation of the acceleration modality:

[0043]

[0044] where F s is the initial feature of the extracted acceleration modality, containing the time dependence features of the acceleration signal; d f represents the dimension of the extracted feature, which is determined by the network parameters.

[0045] The extraction of the initial audio modality features includes:

[0046] The collected audio modality data is expressed as:

[0047]

[0048] where x a represents the audio modality data, T a is the time step, representing the number of time segments of the audio modality data, and d a is the number of channels of the audio modality data. For monophonic audio, d a = 1.

[0049] First, use Mel Frequency Cepstral Coefficients (MFCC) to extract the frequency domain features of the audio modality data:

[0050]

[0051] where d m represents the dimension of the MFCC features, and the MFCC features can capture the most important frequency domain characteristics in the audio signal.

[0052] Then, the MFCC features are input into a Recurrent Neural Network (RNN) to capture the time series characteristics of the audio and generate the initial features of the audio modality:

[0053]

[0054] Among them, F a is the initial feature of the extracted audio modality, which contains the time context information of the audio signal, and d f is the feature dimension.

[0055] The extraction of the initial features of the video modality includes:

[0056] The collected video modality data is represented as:

[0057]

[0058] Among them, x v represents the video modality data, T v is the number of video frames, H and W are the height and width of the video frame respectively, and C is the number of channels of the video frame. For example, when it is an RGB video, C = 3.

[0059] The video modality data is processed by a three-dimensional convolutional neural network (3D-CNN) to extract the spatio-temporal dynamic features of the video modality and generate the initial features of the video modality:

[0060]

[0061] Among them, F v is the initial feature of the extracted video modality, which contains the spatio-temporal dynamic information of human actions in the video, and d f represents the dimension of the extracted feature.

[0062] The unified mapping and normalization of the modality features include:

[0063] After completing the feature extraction of acceleration, audio, and video modalities, it is necessary to uniformly map these features to a shared feature space to eliminate the scale differences between modalities and provide a basis for subsequent fusion.

[0064] First, the initial features of each modality are mapped to the shared feature space through a fully connected layer (FC):

[0065]

[0066] Among them, F m ' represents the modality feature after mapping of modality m, T m is the time step corresponding to modality m, and d crepresents the dimension of the shared feature space, and s, a, and v represent the acceleration, audio, and video modalities respectively.

[0067] Next, normalize the mapped modal features to adjust the feature distribution:

[0068]

[0069] where F m " is the feature after normalization for modality m, that is, the final multi-modal feature, and μ m represents the mean of the modal features, which represents the center of the feature in the shared space, and σ m is the standard deviation std of the feature, which is used to scale the range of the feature.

[0070] S2. Input the multi-modal features into the multi-modal time alignment module to obtain multi-modal time alignment features;

[0071] Time alignment is to perform refined matching on the time steps of the modal features. Although basic time synchronization has been achieved during data acquisition, due to the possible inconsistency of the time distribution and dynamic change characteristics of the modal features after extraction (such as sampling frequency differences, inter-modal time delays, etc.), time alignment still needs to be achieved at the feature level.

[0072] In step S1, the multi-modal feature F m " is extracted. Due to the different sampling frequencies of the multi-modal data, the features are inconsistent in the time dimension and cannot be directly fused. Therefore, the multi-modal time alignment module quantifies the weights of each modality using the modal contribution evaluation method based on the multi-modal features, and uses the dynamic time warping technique to perform time alignment on the multi-modal features according to the weights of each modality to obtain multi-modal time alignment features, providing a unified time basis for subsequent feature fusion.

[0073] Quantifying the weights of each modality using the Shapley value modal contribution evaluation method includes:

[0074] To evaluate the contribution of each modality in the time alignment process, it is first necessary to define the value function v(S) of the modal subset.

[0075] Let the modal set M = {s, a, v}, where s, a, and v represent the acceleration, audio, and video modalities respectively. The feature F m " of modality m = {F m "[1], F m "[2],..., F m "[T m}, where T m represents the time step of modality m, and F m "[t] represents the feature Fm "Features at time step t.

[0076] The value function v(S) of the modal subset S is defined as:

[0077]

[0078] where Error(S) represents the error value when performing time alignment using only the modal subset S. The introduction of the negative sign is to convert the minimization objective of the error into the maximization of the value function. The smaller the error, the larger the value of the value function, and D error is the Euclidean distance, and S / {m} is the subset obtained by removing modal m from the modal subset S, and F m "[t], F n "[t] represent the features of modes m and n at time step t respectively. For example, when the subset S = {s, a}, Error(S) represents the error using the acceleration and audio modes for alignment.

[0079] After defining the value function, the contribution weight of each mode is calculated through the Shapley value. The Shapley value is based on cooperative game theory and fairly distributes the importance of each mode in the time alignment process by quantifying the marginal contribution of the mode in all possible subsets. Its calculation formula is:

[0080]

[0081] In this formula, φ m represents the Shapley value of mode m, n represents the total number of modes, the value v(S) represents the value function of the modal subset S, that is, the contribution of this subset in the time alignment task, and the marginal contribution v(S∪{m}) - v(S) represents the degree of improvement in the time alignment performance after mode m is added to subset S. The weighting coefficient in the formula reflects the importance of the subset size and permutation order for the Shapley value distribution. By weighted summing the marginal contributions of all possible subsets, the Shapley value comprehensively quantifies the importance of mode m.

[0082] Directly calculating the Shapley value requires traversing all modal subsets, and its complexity is O(2 n ), and the computational cost will increase rapidly as the number of modes increases. To solve this problem, the present invention uses the Monte Carlo sampling method to approximately calculate the Shapley value. Specifically, N modal subsets S k that do not contain mode m are randomly generated, and the marginal contribution of each sampled subset is calculated. The formula is:

[0083]

[0084] where S kis the k-th sampling subset, and N is the total number of samplings.

[0085] Monte Carlo sampling avoids the high cost of traversing all possible subsets through selective calculation, and can reduce the computational complexity of the Shapley value from exponential O(2 n ) to linear level O(k), achieving a good balance between computational efficiency and result accuracy.

[0086] The calculated Shapley value φ m represents the contribution weight of modality m, which will be used for modality weighting in the subsequent time alignment process. For example: if the weight φ s of the acceleration modality is large, it indicates that the features of the acceleration modality contribute more to time alignment; the weight φ a of the audio modality may be more concentrated in the later stage of the event occurrence; the weight φ v of the video modality may describe dynamic behaviors more accurately.

[0087] DTW is a non-linear time alignment algorithm that finds the optimal matching path of the time steps of different modality features through dynamic programming, enabling the dynamic changes of multi-modal features to be consistent in the time dimension, and can align time series with different lengths or inconsistent sampling frequencies on the time axis; using the dynamic time warping technique for time alignment of multi-modal features includes:

[0088] Define the weighted local distance d t,t′ :

[0089] d t,t′ = φ s ||F s ”[t] - F a ”[t′]|| 2 + φ a ||F a ”[t] - F v ”[t′]|| 2 + φ v ||F v ”[t] - F s ”[t′]|| 2 (13)

[0090] where t′ is the time step, and φ s , φ a , φ v are the Shapley value contribution weights, reflecting the importance of the corresponding modality m in the alignment task. The local distance d t,t′ measures the difference of modality features at the time step, and is weighted by the Shapley value weights, enhancing the influence of high-contribution modalities in the modality alignment process, thereby improving the accuracy and robustness of the alignment path.

[0091] Construct the cumulative distance matrix D[t, t′] based on the weighted local distance, and its recurrence formula is as follows:

[0092] D[t, t′] = d t,t′ + min{D[t - 1, t′], D[t, t′ - 1], D[t - 1, t′ - 1]} (14)

[0093] Among them, the matrix D[t, t′] represents the minimum cumulative distance from the time step (0, 0) to the time step (t, t′).

[0094] To make the time alignment process smoother and support gradient optimization, the present invention introduces the Soft-DTW method to solve the optimization difficulty problem caused by non-differentiability and path jumps in traditional DTW, so as to achieve more accurate and efficient alignment of multimodal features, and replaces the minimum operation in DTW with a differentiable soft-min operation. The definition of the soft-min operation is as follows:

[0095] softmin γ (a, b, c) = -γ log(e -a / γ + e -b / γ + e -c / γ ) (15)

[0096] Among them, a, b, and c are control parameters for the smoothing operation, which are obtained through experimental optimization, and the specific values need to be adjusted according to the data and the model; γ > 0 is the smoothing parameter, which controls the smoothing degree of the soft-min. When γ → 0, the soft-min approaches the standard minimum operation. After introducing the soft-min, the recurrence formula of the cumulative distance matrix is adjusted to:

[0097] D[t, t′] = d t,t′ + softmin γ {D[t - 1, t′], D[t, t′ - 1], D[t - 1, t′ - 1]} (16)

[0098] By backtracking the cumulative distance matrix, obtain the optimal time alignment path P, which is used to adjust the time step length of the modal features, so that the dynamic changes of the multimodal features are consistent on the time axis;

[0099] Adjust the multimodal features according to the optimal time alignment path P to obtain the multimodal time alignment features:

[0100]

[0101] Among them, represents the time alignment feature of modality m, which has been unified in the time dimension and provides a basis for subsequent feature fusion. Align represents alignment.

[0102] Time alignment solves the following problems: Sampling frequency differences: Data from different modalities have different sampling frequencies (e.g., audio at 16 kHz, acceleration at 50 Hz, video at 30 fps), and direct fusion may lead to mismatches in the time dimension; Event response time differences: Modal features may have different response time delays before and after an event. For example, the acceleration modality may capture the falling trend first, while the audio modality captures the sound features after the fall occurs; Improving the effect of modal fusion: The consistency of the aligned modal features on the time axis improves the accuracy and robustness of modal fusion.

[0103] S3. Input the multi-modal time-aligned features into the deep fusion module to obtain deep fusion features;

[0104] Step S3 uses a bottleneck mechanism to compress the multi-modal time-aligned features to obtain bottleneck features for each modality, and uses a cross-attention mechanism to achieve information interaction between the key modality and the base modality, generating unified deep fusion features. This process effectively fuses the complementary information of multiple modalities while reducing the computational complexity, providing efficient and accurate input for the fall detection task. The specific steps are as follows:

[0105] Compressing the multi-modal time-aligned features using the bottleneck mechanism includes:

[0106] Based on the time-aligned multi-modal features, the high-dimensional features of each modality are compressed into a small number of bottleneck token representations through the bottleneck mechanism, reducing redundancy and highlighting key information. The specific formula is as follows:

[0107]

[0108] Among them, is the time-aligned feature of modality m, MLP is a multi-layer perceptron, and W bottleneck, is the bottleneck feature mapping matrix, which is used for dimensionality reduction and extracting key information.

[0109] The introduction of the number of bottleneck tokens B reduces the computational complexity from O(T 2 ) to O(B 2 ), where B << T.

[0110] The bottleneck mechanism compresses the modal features into a fixed number of tokens, enabling the efficient representation of internal modal information and providing a unified low-dimensional feature representation for interaction between modalities. The compressed bottleneck feature B m is used as the input for subsequent cross-attention calculations.

[0111] Among the bottleneck features B mBased on the proposed method, the cross-attention mechanism is used to realize the information interaction between modalities. The cross-attention mechanism is calculated in the bottleneck feature space, which further reduces the resource consumption of high-dimensional modal feature calculation.

[0112] Since the acceleration modality can usually capture the drastic changes in human motion first in a fall event, and the audio and video modalities focus more on supplementing the acoustic and visual features, the acceleration modality has an obvious temporal relationship and complementarity with the audio and video modalities. Therefore, the acceleration modality is taken as the key modality (active exploration modality), and the audio and video modalities are taken as the basic modalities (interacted modalities). The interaction features between the acceleration modality and the audio and video modalities are calculated through the cross-attention mechanism:

[0113]

[0114] Among them, Q s =B s W q Represents the query matrix of the key mode (i.e., acceleration mode), K m =B m W k V represents the key matrix of the basic mode (i.e., audio mode or video mode); m =B m W v is the value matrix of the basic mode; W q , W k , W a is a learnable linear transformation matrix; Attention m→s Represents the interaction result of the basic mode m to the key mode s.

[0115] After the cross attention calculation is completed, the splicing operation is performed to generate the final deep fusion feature F fusion :

[0116] F fusion =Concat(Attention a→s ,Attention v→s )(20)

[0117] Among them, Concat represents the feature concatenation operation, Attention a→s and Attention v→s They are the interactive information of audio and video modalities to the key modalities respectively.

[0118] S4, input the deep fusion features into the classification module to obtain the fall detection result;

[0119] In step S4, the deep fusion feature F generated in step S3 is used fusion, the classification of fall events is completed by designing a classifier. In this process, to further optimize the feature distribution and improve the adaptability and robustness of the classification model, a residual connection and a group normalization mechanism are combined to process the input features, and finally the classification results of fall events are output. The specific steps are as follows:

[0120] Apply a residual connection to the deep fusion feature F fusion and add it to the time-aligned feature to form the final classification input feature F residual ; The calculation formula of the residual connection is:

[0121]

[0122] After the residual connection, group normalization (GN) is adopted to optimize the feature distribution, reduce the deviation between features, and at the same time improve the adaptability of the model to the difference in feature distribution between modalities. The calculation formula of group normalization is as follows:

[0123]

[0124] where is the normalized feature representation, GN represents the group normalization operation, and group normalization divides the input features into several groups according to the feature dimension, and each group performs the normalization operation independently. This can avoid the normalization error during small-batch training and enhance the expressiveness after the feature fusion between modalities..

[0125] The residual connection retains the information of the original features, alleviates the problem of gradient disappearance that may occur in deep models, and enhances the adaptability and robustness of the deep fusion features to the classification task. By combining group normalization and residual connection, not only can the core information after the fusion between modalities be retained, but also good numerical stability is achieved.

[0126] The classification input features after the residual connection and normalization processing are input into the classifier for the classification task of fall events. The classifier consists of several layers of fully connected layers (FC) and Relu activation functions, and is used to perform feature mapping and classification prediction on the input features.

[0127] First, the classification input feature is mapped to an intermediate hidden feature representation through one or more fully connected layers. The calculation formula of the fully connected layer is as follows:

[0128]

[0129] where is the weight matrix of the fully connected layer, is the bias term, h is the output of the hidden layer, and d h is the feature dimension of the hidden layer. The role of the fully connected layer is to reduce the dimension and perform non-linear transformation on the classification input features, and extract the key features related to the classification task.

[0130] Subsequently, the output h of the hidden layer is input to the output layer of the classifier and mapped to the binary classification space. Through the softmax function, the result of the output layer is converted into the classification probabilities of falling and not falling:

[0131] P = softmax(W out ·h + b out )(24)

[0132] Among them, represents the probabilities that each sample belongs to the two categories of falling and not falling, is the weight matrix of the output layer, is the bias term of the output layer. Through the softmax function, the model can generate a binary classification probability distribution for each sample.

[0133] Through the classification result P output by the classifier, the falling event category corresponding to each input feature can be effectively predicted. In the whole classification task, the introduction of residual connection and group normalization improves the stability and adaptability of the input features. At the same time, the combination of the fully connected layer and the activation function further enhances the expression ability and discriminative ability of the classifier, thus providing guarantee for the efficiency and robustness of the classification model.

[0134] S5. Calculate the multi-task loss function value according to the multi-modal time-aligned features, deep fusion features and the falling detection result, update the model parameters according to the multi-task loss function value, and when the loss function value is the smallest, obtain the trained multi-modal falling detection model.

[0135] In the step S5, the present invention further improves the robustness and classification accuracy of the model through the loss function of multi-task joint optimization. By introducing classification loss, modality consistency loss and center loss, the alignment effect of multi-modal features, the compactness of intra-class features and the discriminative ability of inter-class features are optimized, effectively enhancing the adaptability of the system in complex scenarios.

[0136] The classification loss is used to measure the difference between the model prediction result and the true label. By minimizing the classification error, the accuracy of falling detection is improved. The specific formula is as follows:

[0137]

[0138] Among them, N tr represents the total number of samples; y iDenote the true label of the i-th sample (i.e., multimodal data); P fall,i Denote the predicted probability that the classifier predicts sample i belongs to the fall event.

[0139] To ensure the alignment effect of different modal features in the shared feature space, the present invention designs a modal consistency loss. By constraining the distance between different modal features, the complementarity between modalities is enhanced, and the interference of noisy modalities on the overall performance is reduced. The specific formula is as follows:

[0140]

[0141] Among them, F i,s 、F i,a 、F i,v respectively denote the time-aligned features of the acceleration, audio, and video modalities of the i-th sample, w sa is the weight of the acceleration modality and the audio modality, w sv is the weight of the acceleration modality and the video modality, used to reflect the contribution degree between the acceleration modality and the audio modality, and between the acceleration modality and the video modality, φ sa 、φ sv are respectively the Shapley value of the acceleration modality and the audio modality, and the Shapley value of the acceleration modality and the video modality. In formula (29), S k is the k-th sampling subset that does not include modalities s and a. In formula (30), S k is the k-th sampling subset that does not include modalities s and v. N is the total number of samplings.

[0142] To further enhance the compactness constraint of the classifier on intra-class samples and the discrimination ability between inter-class samples, the present invention introduces a center loss. The center loss effectively optimizes the feature distribution by forcing the feature vectors of samples in the same class to be close to their class centers. The formula is as follows:

[0143]

[0144] Among them, denote the deep fusion feature of the i-th sample, denote the feature center of the i-th sample corresponding to class y i The feature center of class y i is the mean of the deep fusion features of all samples belonging to class y i ; β is the weight of the center loss, used to adjust the influence of this loss on the optimization process.

[0145] The present invention realizes the global optimization of the model in the form of a multi-task loss function by jointly optimizing the classification loss, modal consistency loss, and center loss. The specific form of the total loss function is:

[0146] L total = L CE + λ 1 L consistency + λ 2 L center (32)

[0147] Among them, L CE represents the classification loss, L consistency represents the modality consistency loss, and L center represents the center loss; λ 1 and λ 2 are the weight parameters of the modality consistency loss and the center loss respectively. By adjusting the weight parameters, the system can achieve a dynamic balance of performance between the classification task and the modality alignment task.

[0148] The above-mentioned embodiments further illustrate the purpose, technical solutions and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An efficient multi-modal fall detection method in a smart medical environment, characterized in that: include: Collecting multimodal data, and inputting the collected multimodal data into a trained multimodal fall detection model to obtain a fall detection result; The multimodal fall detection model includes: a multimodal feature extraction module, a multimodal time alignment module, a deep fusion module, and a classification module; The training process of the multimodal fall detection model includes: S1, collecting multimodal data, inputting the multimodal data into a multimodal feature extraction module to obtain multimodal features; S2, inputting the multimodal features into a multimodal time alignment module to obtain multimodal time alignment features; S3, inputting the multimodal time alignment features into the deep fusion module to obtain deep fusion features; S4, input the deep fusion features into the classification module to obtain the fall detection result; S5. Calculate the multi-task loss function value according to the multi-modal time alignment features, deep fusion features and fall detection results, update the model parameters according to the multi-task loss function value, and obtain the trained multi-modal fall detection model when the loss function value is minimized.

2. According to claim 1, an efficient multi-modal fall detection method in a smart medical environment is characterized in that: Multimodal data includes acceleration data, audio data, and video data; The multimodal feature extraction module includes: an acceleration modal feature extraction module, an audio modal feature extraction module, a video modal feature extraction module, and a mapping and normalization module; the multimodal feature extraction module processes the multimodal data by: inputting the acceleration data, the audio data, and the video data into the acceleration modal feature extraction module, the audio modal feature extraction module, and the video modal feature extraction module, respectively, to obtain multimodal initial features, and inputting the multimodal initial features into the mapping and normalization module, respectively, to obtain multimodal features F m "; where m represents the mode.

3. According to claim 1, an efficient multi-modal fall detection method in a smart medical environment is characterized in that: The multimodal time alignment module processes the multimodal features by: quantifying the weight of each mode using the Shapley value modal contribution evaluation method according to the multimodal features, and aligning the multimodal features in time using the dynamic time warping method according to the weight of each mode to obtain the multimodal time alignment features.

4. According to claim 3, an efficient multi-modal fall detection method in a smart medical environment is characterized in that: The Shapley value modal contribution evaluation method is used to quantify the weight of each mode, including: constructing a modal set M = {s, a, v}, randomly generating N modal subsets S that do not contain mode m according to the modal set M k , calculate each modal subset S k The modal value function According to the modal subset S k The modal value function calculates the weight of the corresponding modality m Where s, a, and v represent acceleration, audio, and video modes, respectively, k is the index of the modality subset, and -Error(S k ) represents the modal subset S k The error value for time alignment, T m The time step that characterizes mode m, F m ”[t]、F n ”[t] represents the characteristics of mode m and mode n at time step t, respectively. error is the Euclidean distance, S k / {m} is the modal subset S k Subset obtained by removing mode m.

5. According to claim 3, the efficient multi-modal fall detection method in a smart medical environment is characterized in that: Temporal alignment of multimodal features includes: constructing a weighted local distance between modalities according to the weights of each modality, constructing a cumulative distance matrix according to the weighted local distance between modalities, optimizing the cumulative distance matrix to obtain an optimal time alignment path, adjusting the multimodal features according to the optimal time alignment path, and obtaining multimodal time alignment features.

6. According to claim 1, the efficient multi-modal fall detection method in a smart medical environment is characterized in that: The deep fusion module processes the aligned features of each modality by: compressing the multimodal time-aligned features using the bottleneck mechanism to obtain the bottleneck features of each modality; and interacting and fusing the bottleneck features of each modality through the cross-attention mechanism to obtain the deep fusion features.

7. The efficient multi-modal fall detection method in a smart medical environment according to claim 6, characterized in that: The bottleneck features of each mode are interacted and integrated, including: according to the bottleneck feature B of the acceleration mode s Calculate the query matrix Q s , according to the bottleneck feature B of the audio modality a Calculate the key matrix and value matrix according to the bottleneck feature B of the video modality v Calculate the key matrix and value matrix according to the query matrix Q s And the bottleneck feature B of the audio modality a The key matrix and value matrix of Attention calculate the interaction results a→s , according to the query matrix Q s And the bottleneck feature B of the video modality v The key matrix and value matrix of Attention calculate the interaction results v→s ; Attention the interaction results a→s 、Attention v→s Splice to get the deep fusion feature F fusion .

8. According to claim 1, the efficient multi-modal fall detection method in a smart medical environment is characterized in that: The classification module processes the deep fusion features by performing residual connection on the deep fusion features and the multimodal time alignment features to obtain the classification input features F residual ; For classification input feature F residual Normalize and get the normalized features Use the classifier to normalize the features Processing is performed to obtain the fall detection result.

9. The efficient multi-modal fall detection method in a smart medical environment according to claim 1, characterized in that: Loss function L total =L CE +λ1L consistency +λ2L center Among them, L CE represents the classification loss function, L consistency represents the modality consistency loss function, L center represents the center loss function, λ1 and λ2 are the weight parameters of modality consistency loss and center loss, respectively.

10. The efficient multi-modal fall detection method in a smart medical environment according to claim 9, characterized in that: The modality consistency loss function is: Among them, F i,s 、F i,a 、F i,v denote the time-aligned features of acceleration, audio, and video modalities of the i-th multimodal data, respectively, and w sa is the weight of acceleration mode and audio mode, w sv is the weight of acceleration mode and video mode, N tr is the number of multimodal data.

Citation Information

Cited By

  • Emotion recognition interaction method and device based on AI intelligent analysis and electronic equipment

    CN121919804A

  • Anti-falling prediction method and system for rehabilitation training scene

    CN122201618A

  • A method and system for fall prevention prediction in rehabilitation training scenarios

    CN122201618B