A cognitive impairment old person behavior early warning method and system based on cross-modal fusion
Patent Information
- Application Number
- CN202610908965.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-15
AI Technical Summary
[0004]为了弥补以上不足,本发明提供了一种基于跨模态融合的认知障碍老人行为预警方法及系统,旨在改善传统的监护系统大多采用单一传感器,由于单一模态在复杂环境中易受杂波干扰且无法交叉验证,从而造成隐匿危急事件漏报与高频误报的问题
1、本发明中,通过应用证据理论融合雷达与声场的跨模态概率赋值,进而实现对异常行为的高鲁棒性预警,从而改善了传统的监护系统大多采用单一传感器,由于单一模态在复杂环境中易受杂波干扰且无法交叉验证,从而造成隐匿危急事件漏报与高频误报的问题。
Smart Images

Figure CN122761576A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent health monitoring and signal processing technology, and in particular to a method and system for early warning of the behavior of elderly people with cognitive impairment based on cross-modal fusion. Background Technology
[0002] The number of elderly people suffering from early cognitive impairment at home or in care facilities is constantly expanding. This group often exhibits pathological behaviors such as aimless wandering, frozen gait, and confusion at specific times. Traditional home monitoring technologies typically use single-modal sensors for spatial motion capture or rely solely on audio pickup devices to identify violent impacts and cries for help.
[0003] Traditional home monitoring technologies suffer from severe modal fragmentation in terms of physical architecture and data processing. In complex home environments, they are prone to generating high-frequency false alarms when faced with multipath clutter and non-stationary environmental noise interference. Furthermore, single sensors cannot cross modal barriers to perform feature cross-validation, making it impossible for the system to capture highly concealed critical events such as "abnormal silence" in elderly people with cognitive impairment who are in a wandering state but have lost the ability to call for help due to the onset of their condition. At the same time, there is a lack of mathematical mechanisms to integrate isolated heterogeneous physical signs into long-term assessment indicators, resulting in extremely poor robustness of the monitoring system in providing instantaneous early warning of complex pathological behaviors and a lack of tracking of long-term cognitive decline trends. Summary of the Invention
[0004] To overcome the above shortcomings, this invention provides a method and system for early warning of behavior in elderly people with cognitive impairment based on cross-modal fusion. It aims to improve the problem that most traditional monitoring systems use a single sensor. Since single-modal systems are easily affected by clutter interference in complex environments and cannot be cross-verified, they cause problems such as missed reporting of hidden critical events and high-frequency false alarms.
[0005] In a first aspect, the present invention provides the following technical solution: a method for early warning of cognitive impairment in elderly people based on cross-modal fusion, comprising the following steps: S1, process the acquired millimeter-wave radar signal to generate three-dimensional point cloud data, and extract gait spatiotemporal features based on the three-dimensional point cloud data; S2, Based on the three-dimensional point cloud data, track and obtain trajectory features, input the trajectory features into the trajectory classification network, and output the wandering classification result; S3, process the acquired sound field signal to output sound source localization parameters, extract acoustic feature vectors, input the acoustic feature vectors into the acoustic classification model, and output the acoustic classification results; S4, obtain the radar probability assignment constructed based on the gait spatiotemporal features and the wandering classification result, obtain the acoustic probability assignment constructed based on the sound source localization parameters and the acoustic classification result, fuse the radar probability assignment and the acoustic probability assignment through evidence theory to obtain the fusion confidence level, and trigger an abnormal warning when the fusion confidence level exceeds the warning threshold; S5, merge the gait spatiotemporal features with the acoustic feature vector to construct a cross-modal time series index, perform trend testing on the cross-modal time series index to generate a decay composite index, and output a trend warning when the decay composite index meets the preset conditions.
[0006] By adopting the above technical solution, the cross-modal probability assignment of radar and sound field is fused with evidence theory, thereby achieving a highly robust early warning of abnormal behavior. This improves the problem that traditional monitoring systems mostly use a single sensor, which is susceptible to clutter interference in complex environments and cannot be cross-verified, resulting in missed reports of hidden critical events and high-frequency false alarms.
[0007] Secondly, this invention provides the following technical solution: a behavioral early warning system for elderly people with cognitive impairment based on cross-modal fusion, comprising the following modules: The radar perception and feature extraction module is used to process the acquired millimeter-wave radar signals to generate three-dimensional point cloud data, and extract gait spatiotemporal features based on the three-dimensional point cloud data. The trajectory tracking and wandering classification module is used to track and obtain trajectory features based on the three-dimensional point cloud data, input the trajectory features into the trajectory classification network, and output the wandering classification result. The sound field perception and acoustic classification module is used to process the acquired sound field signal, output sound source localization parameters, extract acoustic feature vectors, input the acoustic feature vectors into the acoustic classification model, and output acoustic classification results. The cross-modal fusion and anomaly warning module is used to obtain radar probability assignments constructed based on the gait spatiotemporal features and the wandering classification results, obtain acoustic probability assignments constructed based on the sound source localization parameters and the acoustic classification results, fuse the radar probability assignments and the acoustic probability assignments through evidence theory to obtain fusion confidence, and trigger an anomaly warning when the fusion confidence exceeds the warning threshold; The temporal merging and trend warning module is used to merge the gait spatiotemporal features and the acoustic feature vector to construct a cross-modal temporal index, perform trend testing on the cross-modal temporal index to generate a decay composite index, and output a trend warning when the decay composite index meets preset conditions.
[0008] The present invention has the following beneficial effects: 1. In this invention, by applying evidence theory to fuse the cross-modal probability assignment of radar and sound field, a highly robust early warning of abnormal behavior is achieved, thereby improving the problem that most traditional monitoring systems use a single sensor. Since a single mode is easily affected by clutter interference in complex environments and cannot be cross-verified, it causes the problem of missed reporting of hidden critical events and high-frequency false alarms.
[0009] 2. In this invention, a decay composite index is generated by performing trend testing on the temporal indicators composed of gait spatiotemporal features and acoustic feature vectors, and then a trend warning is output. This improves the problem that traditional monitoring methods mostly use instantaneous alarm mechanisms, which are prone to missing the slow decline trend in the early stage of cognition due to the lack of longitudinal correlation of heterogeneous signs.
[0010] 3. In this invention, by calculating the evidence conflict coefficient and switching the weighted average strategy when the threshold is exceeded, the confidence of cross-modal fusion is obtained, thereby improving the problem that traditional multimodal fusion mostly adopts fixed constraint combinations, which is prone to falling into the evidence conflict paradox under strong noise interference, thus causing the system fusion decision logic to fail.
[0011] 4. In this invention, by extracting spatial statistical features such as straightness and information entropy and combining them with coordinate sequence input into the classification network, the wandering classification result is output. This improves the problem that traditional behavior recognition mostly uses a single coordinate tracking model, which is difficult to quantify the spatial complexity of motion patterns, thus causing misjudgment of pathological wandering behavior. Attached Figure Description
[0012] Figure 1 The flowchart shows a method for early warning of cognitive impairment in elderly people based on cross-modal fusion proposed in this invention. Figure 2 This is a flowchart of radar signal processing and gait spatiotemporal feature extraction for a behavioral early warning method for elderly people with cognitive impairment based on cross-modal fusion proposed in this invention. Figure 3 This invention presents an evidence theory flowchart for a cross-modal fusion and anomaly warning method for cognitively impaired elderly behavior based on cross-modal fusion. Figure 4 This invention presents a flowchart of cross-modal time-series indicators and cognitive decline trend early warning methods for a behavioral early warning method for elderly people with cognitive impairment based on cross-modal fusion. Figure 5 The image shows a simulation of radar micro-Doppler spectrum and frozen gait recognition for a cross-modal fusion-based behavioral early warning method for elderly people with cognitive impairment proposed in this invention. Figure 6 This is a simulation diagram of a two-dimensional radar trajectory for a wandering pattern in a cognitive impairment elderly behavior early warning method based on cross-modal fusion proposed in this invention. Figure 7 This is a simulation diagram of acoustic-radio cross-modal DS evidence fusion and temporal confidence accumulation for a behavioral early warning method for elderly people with cognitive impairment based on cross-modal fusion proposed in this invention. Figure 8 This is a long-term cognitive trend and change point detection experimental curve of a behavioral early warning method for elderly people with cognitive impairment based on cross-modal fusion proposed in this invention; Figure 9 This is an architecture diagram of a behavioral early warning system for elderly people with cognitive impairment based on cross-modal fusion proposed in this invention. Detailed Implementation
[0013] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Example 1: In a first embodiment of the present invention, the present invention provides a method for early warning of the behavior of elderly people with cognitive impairment based on cross-modal fusion, such as... Figures 1-2 , Figures 5-6 As shown, it includes the following steps: S1, process the acquired millimeter-wave radar signal to generate three-dimensional point cloud data, and extract gait spatiotemporal features based on the three-dimensional point cloud data; Furthermore, in S1, the process of processing the acquired millimeter-wave radar signals to generate 3D point cloud data includes: The millimeter-wave radar signal is subjected to intermediate frequency beat frequency processing, and then the range-dimensional fast Fourier transform and the Doppler-dimensional fast Fourier transform are performed sequentially to obtain the range-Doppler image. Two-dimensional ordered statistical constant false alarm rate (CFAR) detection is performed on the distance Doppler map. The power ranking values of the training units around the unit to be detected are extracted as noise estimates. The detection threshold is calculated based on the CFAR coefficient to extract the target detection point. A multi-signal classification algorithm is performed on the target detection points using a multi-input multi-output virtual array. The noise subspace is extracted by eigenvalue decomposition of the covariance matrix and a pseudo-spectrum is constructed to obtain the azimuth angle estimate. Three-dimensional point cloud data is generated by density-based clustering algorithm.
[0015] In S1, the extraction of gait spatiotemporal features based on 3D point cloud data includes: Calculate the change in the position of the human centroid in three-dimensional point cloud data between consecutive time frames, and extract the instantaneous gait speed and average gait speed; Autocorrelation is performed on the time-series signal of instantaneous gait speed to extract the gait period, and the step length is calculated by combining the average gait speed. The coefficient of variation of multiple consecutive step lengths is then calculated as the step length coefficient of variation. The micro-Doppler spectrum is obtained by short-time Fourier transform. The oscillation period of the positive and negative Doppler frequency shifts in the micro-Doppler spectrum is extracted to calculate the gait asymmetry index. The ratio of the power spectral density of the high-frequency band to the motion band in the micro-Doppler frequency domain is extracted to calculate the frozen gait index. The instantaneous gait speed, step length, step length variation coefficient, gait asymmetry index and frozen gait index are combined to generate the spatiotemporal characteristics of gait.
[0016] Specifically, the system's underlying hardware acquires raw echo signals from indoor millimeter-wave radar as input data. The signals undergo multi-dimensional domain transformation, constant false alarm rate (CFAR) detection, and spatial clustering processing to output three-dimensional point cloud data of the target human body. Based on the acquired three-dimensional point cloud data, temporal and frequency domain analysis is performed to output instantaneous gait speed, stride length, stride length coefficient of variation, gait asymmetry index, and frozen gait index. These are combined to generate multi-dimensional gait spatiotemporal features, which are then output to the downstream processing network.
[0017] The acquired frequency-modulated continuous wave echo signal undergoes intermediate frequency beat frequency processing, followed by range-dimensional fast Fourier transform and Doppler-dimensional fast Fourier transform, converting the one-dimensional time series into a two-dimensional range-Doppler feature map. A two-dimensional ordered statistical constant false alarm rate (CFAR) detection algorithm is applied to the range-Doppler map, obtaining specific quantile values of the power ranking of training units surrounding the target unit as noise level estimates. A detection threshold is established based on the set CFAR coefficient to extract the target point. A multi-signal classification algorithm is performed on the target point using a multi-input multi-output virtual antenna array, and eigenvalue decomposition of the covariance matrix is used to extract the noise subspace and construct a pseudospectrum to obtain the azimuth angle. Finally, a density-based clustering algorithm is used to generate three-dimensional point cloud data.
[0018] This process maps the original electromagnetic wave reflection signal into the three-dimensional structure of the human body. By using ordered statistical sorting to estimate background noise, it eliminates mutual occlusion interference caused by multiple targets clustering in the home environment, ensuring the detection stability of human targets in edge areas.
[0019] The change in the human centroid position coordinates within consecutive time frames of 3D point cloud data is calculated, and divided by the inter-frame time interval to obtain the instantaneous gait velocity. Autocorrelation is performed on the temporal signal of the instantaneous gait velocity, and the time delay corresponding to the first significant peak with non-zero delay is extracted as the gait period. The step length is obtained by combining this with the average gait velocity over the observation time. The step length variation coefficient is obtained by aggregating multiple consecutive sample step lengths and calculating the ratio of the sample standard deviation to the average step length.
[0020] Basic gait calculations strip away the random swaying of the human torso, isolate the periodic lower limb stepping movements, and quantify the variability indicators that reflect the degeneration of neural control circuits.
[0021] The micro-Doppler spectrum is obtained by short-time Fourier transform. The oscillation period of the positive and negative Doppler frequency shifts in the spectrum is separated, and the gait asymmetry index is extracted. The calculation model is as follows: ; in It is the gait asymmetry index; The swing period of the left leg, extracted from the micro-Doppler spectrum, is approximately 1.28 seconds in practice. The swing period of the right leg is approximately 1.32 seconds in practice.
[0022] Extract the power spectral density of a specific frequency band in the micro-Doppler frequency domain and calculate the frozen gait index: ; in The frozen gait index at the current time; The power spectral density is the instantaneous velocity of the center of mass. The numerator is the differential of the frequency, indicating that the integration operation is performed in the frequency domain; the numerator integral is the high-frequency vibration band, with an integration range of 3 to 8 Hz; the denominator integral is the normal motion band, with an integration range of 0.5 to 3 Hz. The system is set to determine a frozen gait event when the frozen gait index is greater than the judgment threshold of 2.0 and the duration exceeds 1 second.
[0023] The aforementioned pathological feature model segmented and compared the power spectral density energy distribution in different frequency ranges, capturing local high-frequency rigid tremor features that are difficult to detect with conventional macroscopic motion tracking, thus solving the problem of monitoring blind spots in patients with cognitive impairment who suddenly freeze their gait.
[0024] S2, based on 3D point cloud data tracking, obtains trajectory features, inputs the trajectory features into the trajectory classification network, and outputs the wandering classification results; Furthermore, in S2, the trajectory features acquired based on 3D point cloud data tracking include: The 3D point cloud data is tracked by an extended Kalman filter, and an acceleration component is introduced into the state vector for prediction and updating to obtain the trajectory coordinate sequence. Spatial complexity estimation and straightness calculation are performed on the trajectory coordinate sequence, and trajectory statistical features including direction change frequency, path repetition rate, spatial coverage area, velocity variation coefficient, straightness exponent, fractal dimension estimated based on box counting method, and trajectory information entropy calculated based on spatial gridding are extracted. Merge trajectory coordinate sequences and trajectory statistical features to generate trajectory features.
[0025] In S2, the trajectory features are input into the trajectory classification network, and the output wandering classification results include: The trajectory coordinate sequence in the trajectory features is input into a one-dimensional convolutional neural network with a temporal attention layer. The output sequence of the convolutional network is self-attention weighted by the temporal attention layer to output the first hidden layer representation. The trajectory statistical features in the trajectory features are input into a multilayer perceptron with an activation function, and the output is the second hidden layer representation; The first hidden layer representation and the second hidden layer representation are concatenated. The probability distribution of four wandering modes—walking, pacing, circling, and random wandering—is directly output through a fully connected layer and a classification function to generate the wandering classification result.
[0026] Specifically, the data input / output process is as follows: The system receives the target human body's 3D point cloud data generated by the front end as the basic input parameter. Through extended Kalman filtering and statistical feature calculation modules, a comprehensive trajectory feature containing coordinate temporal and spatial dimensions is constructed. This feature is imported into a dual-branch classification network for deduction, and finally outputs the wandering classification result probability value representing the cognitive behavior pattern to the downstream multimodal fusion engine.
[0027] An extended Kalman filter is used to process continuously input 3D point cloud data. In constructing the state vector, in addition to the conventional position and velocity variables, an acceleration component is forcibly introduced. This state vector structure possesses the ability to model and predict highly dynamic movements such as sharp turns and sudden stops, overcoming the trajectory loss defect that conventional uniform velocity models easily produce when cognitively impaired targets undergo maneuvering or stagnation, and extracting a highly smooth trajectory coordinate sequence.
[0028] This module extracts statistical features from trajectory segments collected within a specific time window, with the default window length set to two minutes. It focuses on calculating core physical indicators reflecting spatial pathological patterns. For example, it calculates the straightness index. ; in Represents the straightness index. and These refer to the starting and ending coordinates of the trajectory segment, respectively. Indicates the Euclidean distance between two points. This represents the total length of the actual walking path. A value close to 1 indicates normal, direct walking, while a value close to 0 indicates pacing or circling in place.
[0029] The fractal dimension of the trajectory space complexity is estimated using box counting: ; in For fractal dimension, Mathematical operations representing the taking of limits. Represents logarithmic operations. The physical scale is set as a spatial grid. This represents the minimum number of squares required to completely cover the trajectory of the human target. Calculations show that the value for the random walk pattern is close to 2, the value for the pacing pattern is close to 1, and the value for the circling pattern is distributed between 1.3 and 1.6.
[0030] Perform spatial gridding to calculate trajectory information entropy: ; in For trajectory information entropy, Represents the total number of spatial grids. For the trajectory through the first The normalized frequency proportion of each grid. High information entropy values characterize spatially uniform, aimless random walks, while low values reflect stereotyped, repetitive pacing concentrated in a specific, finite region.
[0031] The aforementioned spatial features quantify the repeatability and disorder of behavior, transforming the original coordinate matrix into measurable clinical behavioral indicators, thus solving the problem of missed reports that is easily caused by traditional monitoring relying on manual visual observation. By merging coordinate sequences with statistical features such as the frequency of direction changes, path repetition rate, spatial coverage area, and velocity variation coefficient, a complete trajectory feature is generated.
[0032] The extracted trajectory features are fed into a trajectory classification network. The network's underlying layer employs an independent dual-input architecture, avoiding feature space pollution caused by directly concatenating heterogeneous one-dimensional temporal data with multi-dimensional statistical data. The first branch of the network processes temporal features, feeding the trajectory coordinate sequence into a one-dimensional convolutional layer and then connecting it to a temporal attention layer to calculate the first hidden layer representation. ; in This is the first hidden layer representation of the output; The input is a sequence of trajectory coordinates; and These represent one-dimensional convolutional cascade operations with kernel sizes of 3 and 5, respectively. This represents time-attention weighted computation.
[0033] The second branch of the network processes global features, inputting trajectory statistical features into the multilayer perceptron to calculate the second hidden layer representation: ; in This is the output of the second hidden layer representation; The input is an 8-dimensional trajectory statistical feature vector; and This is the weight matrix of each hidden layer in the multilayer perceptron; and This is the corresponding bias vector; It is a non-linear activation function.
[0034] After the two hidden layer features are concatenated, the comprehensive result is output through the classification mapping layer: ; in This represents the probability distribution results for four patterns: direct walking, pacing, circling, and random walk. For feature tensor splicing operations; Refers to the mapping of fully connected layers; It is a normalized exponential function.
[0035] The network structure integrates the motion distortion of the local temporal dimension with the complexity features of the global spatial dimension, eliminating the technical blind spot of conventional motion classification algorithms that easily misjudge normal housework activities of the elderly as pathological wandering.
[0036] S3 processes the acquired sound field signal, outputs sound source localization parameters, extracts acoustic feature vectors, inputs the acoustic feature vectors into the acoustic classification model, and outputs the acoustic classification results. Furthermore, in S3, the acquired sound field signal is processed to output sound source localization parameters, and acoustic feature vectors are extracted, including: Fourier transform is performed on the sound field signal acquired by the ring microphone array, and the cross-correlation function between microphone channel pairs is calculated using the generalized cross-correlation phase transform algorithm to extract the time delay estimate; The sound source angle of arrival is calculated based on time delay estimation and microphone channel pair spacing, and the confidence level of the angle of arrival is calculated by combining the standard deviation of the multi-channel angle of arrival. The sound source angle of arrival and the confidence level of the angle of arrival are combined to generate sound source localization parameters. A noise covariance matrix and a target direction steering vector are constructed. A minimum variance distortionless response algorithm is used to perform beamforming processing on the sound field signal. After framing, Mel frequency cepstral coefficients and chromaticity map features are extracted to form an acoustic feature vector.
[0037] In S3, the acoustic feature vector is input into the acoustic classification model, and the output acoustic classification results include: The acoustic feature vectors are aggregated by time period and then input into a bidirectional long short-term memory network model to extract the forward hidden state and the backward hidden state respectively. The forward and backward hidden states are concatenated, and the classification layer outputs the classification probabilities of normal speech, groaning, rapid breathing, helpless shouting, repetitive vocalization, crying, falling and impact sounds, coughing, environmental noise, and abnormal silence to generate acoustic classification results. The triggering condition for abnormal silence is that the millimeter-wave radar signal detects that the target human body is in an active state and the acoustic energy of the sound field signal is continuously lower than the environmental background for more than a preset time.
[0038] Specifically, the data input / output process is as follows: The system receives the raw indoor sound field signal collected by a circular microphone array as the basic input data. Through spatial localization algorithms, the sound source angle of arrival and its corresponding confidence level are output as spatial parameters. Enhanced and denoised acoustic feature vectors are extracted and imported into a sequence model for deduction. Finally, classification probabilities covering ten categories of non-verbal acoustic events are output, along with the aforementioned spatial parameters, to the downstream cross-modal fusion decision engine.
[0039] The sound field signal acquired by the six-channel ring microphone array is transformed to the frequency domain using a Fourier transform. A generalized cross-correlation phase transform algorithm is then used to calculate the cross-correlation function between each pair of microphone channels within the array, and the time delay estimate corresponding to the peak value is extracted. Based on this time delay value and the physical array channel spacing, the angle of arrival of the sound source is calculated.
[0040] In normal indoor reverberation and multi-source interference can cause discrepancies in the measurement results for different microphone channel pairs. This module introduces an angle-of-arrival confidence assessment model: ; in This represents the confidence level assessment value for the angle of arrival. The standard deviation of the calculated sound source angle of arrival for 15 effective microphone pairs in the array; Configured for an omnidirectional spatial angle range, with a value of 360 degrees.
[0041] This calculation eliminates the shortcomings of traditional single-confidence outputs. By transforming isolated angle values into spatial parameters carrying reliability distributions, it eliminates direction-finding inaccuracies caused by short barks from pets or transient noises from electrical appliances in a home environment, providing mathematical criteria for subsequent spatial consistency matching with radar trajectories.
[0042] Using the calculated sound source arrival angle as the prior guidance direction, a spatial noise covariance matrix and a target direction guidance vector are constructed. A minimum variance distortionless response algorithm is applied to perform beamforming processing on the multi-channel sound field signal, outputting a directional enhanced audio signal while suppressing background noise in the non-target direction. The enhanced audio signal is then processed frame-by-frame, extracting 40-dimensional Mel-frequency cepstral coefficients and 12-dimensional chromaticity features frame by frame. These features are then combined with the spectral centroid, zero-crossing rate, and short-time energy to construct a 55-dimensional acoustic feature vector. This process eliminates reverberation interference and extracts the core physical features reflecting vocal cord vibration and respiratory status in patients with cognitive impairment.
[0043] Acoustic feature vectors over consecutive time periods are aggregated into a time series, with a default segment length of two seconds, and imported into a bidirectional long short-term memory network model. The model extracts the forward hidden state along the forward time step and the backward hidden state along the backward time step. The forward and backward hidden states are concatenated to construct a composite representation vector containing the global speech context. This vector is then mapped through a fully connected classification layer to output the probability of occurrence for ten specific acoustic events. The classified events include normal speech, groans, rapid breathing, helpless shouts, repetitive vocalizations, crying, sounds of falling or impact, coughing, environmental noise, and unusual silence.
[0044] To address the highly concealed pathological event of abnormal silence, the system is designed with strict cross-modal triggering conditions: the millimeter-wave radar mode must detect that the instantaneous walking speed of the target human body is greater than 0.1 meters per second to confirm that it is active, and the energy of the extracted sound field signal must be continuously lower than the ambient background level for a preset duration of more than 30 seconds.
[0045] The abnormal silence judgment mechanism breaks through the passive response limitation of traditional acoustic monitoring, which relies solely on abnormal sounds for alarms. By introducing radar's human movement status as a reverse reference constraint, the system accurately captures the silent critical moment when a patient suddenly loses consciousness or freezes their gait while wandering indoors and is unable to vocalize for help, eliminating the blind spot in conventional technical systems that monitors concealed pre-fall warning behaviors.
[0046] like Figure 3 , Figure 7 As shown, in S4, the radar probability assignment based on gait spatiotemporal features and wandering classification results is obtained, and the acoustic probability assignment based on sound source localization parameters and acoustic classification results is obtained. The radar probability assignment and acoustic probability assignment are fused through evidence theory to obtain the fusion confidence. When the fusion confidence exceeds the warning threshold, an abnormal warning is triggered. Furthermore, in S4, the fused confidence level is obtained by fusing radar probability assignments and acoustic probability assignments through evidence theory, including: Traverse the set of anomalous event hypotheses, calculate the sum of the cross products of radar probability assignments and acoustic probability assignments, and obtain the evidence conflict coefficient; When the evidence conflict coefficient is not greater than the preset conflict threshold, the combination rule is applied to calculate the first fusion probability; when the evidence conflict coefficient is greater than the preset conflict threshold, the weighted average strategy based on the historical reliability score of the sensor is switched to calculate the second fusion probability. The Gaussian consensus function value of the deviation between the radar target azimuth angle and the sound source arrival angle is calculated. The time context weight and the Gaussian consensus function value are multiplied by the first fusion probability or the second fusion probability, respectively. The time-series evidence accumulation is performed through the sliding window exponential weighted average mechanism, and the fusion confidence is output.
[0047] Specifically, the data input / output process is as follows: the system receives gait spatiotemporal features and wandering classification results from the front end to construct radar probability values, and receives sound source localization parameters and acoustic classification results to construct acoustic probability values. The cross-modal fusion engine performs conflict determination, spatial consistency constraints, and temporal cumulative calculations on the above two heterogeneous probability values, and finally outputs a fusion confidence score representing the probability of an abnormal event occurring. When the fusion confidence score exceeds the warning threshold, the system outputs an abnormal warning control command to the downstream intervention module.
[0048] In complex home environments, single sensors are highly susceptible to environmental clutter and are prone to outputting erroneous features. This module establishes a recognition framework that includes falls, sunset syndrome, wandering, cognitive disorientation crisis, frozen gait, and normal states. To address the discrepancy between radar and acoustic modal evidence, the coefficient of evidence conflict between the two is calculated: ; in The coefficient of evidence conflict; and These represent independent hypothetical events whose intersection within the identification framework is an empty set; Represents the empty set; For radar feature pairs hypothesis The basic probability assignment; For acoustic features, the hypothesis The basic probability is assigned.
[0049] Conventional evidence theory is prone to the Zadeh paradox, where fusion fails, when dealing with highly conflicting evidence. The system establishes a dynamic degradation mechanism based on the conflict coefficient. A preset conflict threshold of 0.7 is set. When the evidence conflict coefficient is not greater than the preset conflict threshold, the standard Dempster combination rule is applied to calculate the first fusion probability: ; in That is, the first fusion probability. To identify the target hypothetical event within the framework.
[0050] When the evidence conflict coefficient exceeds a preset conflict threshold, the above combination rule is blocked, and the second fusion probability is calculated using a weighted average strategy based on the historical reliability scores of the sensors. ; in That is, the second fusion probability; The radar weight is calculated by dividing the radar historical reliability score by the sum of the radar and acoustic historical reliability scores. Acoustic weight; and These assign probabilities to the target hypothetical event for each of the two modalities.
[0051] This dynamic fusion mechanism avoids the systemic logic collapse caused by clutter hijacking of a single sensor in a noisy environment, and ensures the robustness of early warning under extreme conflict conditions.
[0052] To verify the physical matching degree of heterogeneous sensors, a softened Gaussian consistency function is constructed: ; in The value of the Gaussian consistency function; The target azimuth angle obtained by radar detection; The angle of arrival of the sound source obtained for acoustic detection; The angular tolerance parameter is set to ten degrees in practice. This function replaces the traditional hard threshold decision with a continuously and smoothly decaying weight, imposing a severe probabilistic penalty on false anomalies caused by deviations in spatial location.
[0053] By incorporating temporal context weights tailored to specific pathogenesis patterns of cognitive impairment, the instantaneous fusion confidence score is calculated: ; in For instantaneous fusion confidence; The first fusion probability, calculated using the combination rules, is replaced by the second fusion probability when the weighted average strategy is triggered. ; Assumptions are made for specific abnormal events; For time context weights.
[0054] Isolated data frames are highly susceptible to false alarms due to transient interference. The system introduces a sliding window exponentially weighted averaging mechanism to accumulate time-series evidence. ; in The final fusion confidence score output at the current moment; The cumulative window length is set to 10 frames. The forgetting factor has a value of 0.85. This represents the frame offset within the sliding time window.
[0055] The data accumulation mechanism filters out high-frequency burst noise, and the system has a high sensitivity to persistent abnormal behavior. When the final fusion confidence exceeds the set warning threshold, the system determines that an abnormal event has occurred and triggers corresponding linkage intervention commands. In specific implementation, this warning threshold is set to 0.80.
[0056] like Figure 4 , Figure 8As shown in S5, gait spatiotemporal features and acoustic feature vectors are merged to construct cross-modal time series indicators. Trend tests are performed on the cross-modal time series indicators to generate a decay composite index. When the decay composite index meets the preset conditions, a trend warning is output. Furthermore, in S5, trend testing is performed on cross-modal time series indicators to generate a decay composite index, including: Establish longitudinal cognitive health records, perform sliding window linear regression tests on cross-modal time series indicators, and perform nonparametric trend tests on the test statistic of the rank difference of the measured value sequence to confirm the long-term decline slope. The cumulative offset of cross-modal time series indicators from the target mean is calculated using a cumulative sum and control chart algorithm, and abrupt acceleration point detection is performed. Cross-modal time-series indices with different physical dimensions are transformed into dimensionless relative deviations from historical baselines. These deviations are then weighted and combined with gait deviation, stride length variability deviation, gait asymmetry deviation, pause frequency deviation, and vocabulary richness deviation to construct a decline composite index.
[0057] Specifically, the data input / output process is as follows: The system receives multi-dimensional gait spatiotemporal features and acoustic feature vectors extracted by the front-end perception module as basic input data. This heterogeneous data is integrated and stored in a longitudinal cognitive health archive according to time series, constructing cross-modal time-series indicators. Through a trend analysis engine, linear regression, nonparametric tests, change point detection, and multi-dimensional exponential weighting are performed, ultimately outputting a long-term cognitive decline trend warning signal to authorized family members or medical intervention providers.
[0058] Cross-modal time-series indicators with a 30-day time window were extracted. Sliding-window linear regression analysis was performed on features such as daily walking speed, step length variability, and pause frequency to obtain the long-term decay slope characterizing the absolute rate of change. Conventional linear regression is highly susceptible to occasional abnormal sensor readings; therefore, the system simultaneously incorporates a Mankendall nonparametric trend test. The test statistic for the rank difference of the measured value sequence was calculated. ; in To test the statistic; This is the total length of the sequence; and These are the measured values of the indicators at different time points; This is a sign function; it takes a positive 1 when the difference between two values is greater than 0, a negative 1 when the difference is equal to 0, and a negative 1 when the difference is less than 0.
[0059] Calculate the standardized test statistic: ; in This is a standardized test statistic; This represents the estimated variance of the series. A confidence threshold of 1.96 is set; when the absolute value of the standardized test statistic exceeds this threshold, it statistically confirms a long-term degradation trend in the time-series index. The Mankendall test makes no assumptions about the data distribution, thus eliminating the interference of isolated outliers on the overall trend assessment.
[0060] In addition to slow, linear deterioration, patients with cognitive impairment often experience a sudden, accelerated deterioration. The system monitors the cumulative effect of indicators deviating from the target mean using a cumulative sum and control chart algorithm. ; ; in This is the positive cumulative offset; This is the negative cumulative offset; This is a function to find the maximum value. This is the measured value at the current time point; The initial baseline target mean of the archival records; To allow offset parameters, the default value in the actual implementation is [value to be filled in]. ,in The standard deviation of the historical baseline data. The system sets the decision threshold. In practice, the default value is When the positive cumulative offset or negative cumulative offset Exceeding the decision threshold At that time, the algorithm detected a sudden acceleration point. This algorithm accurately identified the accelerated stage in the cognitive decline process, where quantitative change transforms into qualitative change.
[0061] By integrating gait parameters from the mechanical dynamics dimension with acoustic parameters from the language cognition dimension, a decay composite index is constructed: ; in This is a recessionary composite index; superscript Indicates the mean of measurements within the current time window; superscript Indicates the baseline mean of historical data; Average walking speed; The step size is the coefficient of variation. It is the gait asymmetry index; The frequency of pauses; For vocabulary richness. to The weighting coefficients representing the various indicator dimensions are, in practice, taken as 0.30, 0.20, 0.15, 0.20 and 0.15 respectively.
[0062] The system integrates the above calculation results and sets four independent early warning trigger conditions. The first is that the average gait rate decline slope is less than -0.02 meters per second per month, and the coefficient of determination and test statistic meet the criteria. The second is that the pausing frequency rise slope is greater than +0.5 times per minute per month. The third is that the decline compound index is greater than 0.30. The fourth is that a change point is detected. Meeting any one of these conditions triggers a trend warning.
[0063] By converting independent time-series indicators with different physical dimensions into relative deviations, a multidimensional cognitive decline composite index is constructed. Combined with the change point detection mechanism of cumulative sum and control charts, this improves the technical defects of the traditional health monitoring system, which isolates the evaluation of various vital signs data and lacks a longitudinal change point localization mechanism. This avoids the problem of the system missing the slow deterioration trend in the early stage of cognitive impairment.
[0064] Example 2: In home-based elderly care or specialized cognitive care facilities, elderly individuals with early cognitive impairment often exhibit aimless wandering at night, sudden freezing of gait, or a state of "abnormal silence" due to confusion during an attack, rendering them unable to call for help. In this complex real-world monitoring scenario, existing technologies face serious technical limitations: First, home spaces are filled with multipath noise from furniture and non-stationary background sounds, making single radar or acoustic sensors less resistant to interference. In particular, single-modal sensors are prone to missing the hidden danger of elderly individuals being physically active but unable to vocalize for help. Second, when multi-sensor joint monitoring is introduced, conventional fusion algorithms may become paralyzed due to extreme evidence discrepancies when faced with highly contradictory features between modalities, leading to warning failures. Finally, current systems only provide a superficial, passive response to momentary dangerous actions, severing the longitudinal correlation between localized pathological gait and speech characteristics over time. They lack a unified mechanism for measuring the dimensions of multi-source heterogeneous physiological indicators and the ability to detect long-term cumulative changes, failing to clinically identify the critical inflection point where a patient's cognitive function transitions from slow decline to accelerated deterioration. To address the aforementioned problems, this invention provides a behavioral early warning system for elderly people with cognitive impairment based on cross-modal fusion, the structure of which is as follows: Figure 9 As shown. The specific implementation process of this system is as follows: A closed-loop technology system covering instantaneous anomaly detection and long-term disease tracking has been established. The radar perception and feature extraction module processes millimeter-wave radar signals to generate 3D point cloud data and extracts gait spatiotemporal features. This data is then passed to the trajectory tracking and wandering classification module, which uses the 3D point cloud data to track trajectory features and outputs wandering classification results. A parallel sound field perception and acoustic classification module performs spatial direction finding on the sound field signals to output sound source localization parameters and inputs the extracted acoustic feature vectors into a classification network to output acoustic classification results. The cross-modal fusion and anomaly warning module constructs radar and acoustic probability assignments based on the underlying perception features. It uses evidence theory to perform conflict discrimination and consistency combination on the radar and acoustic probability assignments to obtain fusion confidence, triggering anomaly warnings when thresholds are exceeded. The temporal merging and trend warning module continuously aggregates gait spatiotemporal features and acoustic feature vectors to construct cross-modal temporal indicators, performs non-parametric trend testing to generate a decay composite index, and triggers long-term trend warnings. At the spatial decision-making level, each hardware support module uses acoustic-matter modal cross-validation to block the inaccuracy of early warnings caused by background clutter from a single sensor. At the temporal dimension, it connects the underlying control chain for detecting instantaneous hidden crises and locating the inflection point of long-term cognitive degradation.
[0065] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for early warning of behaviors of the elderly with cognitive impairment based on cross-modal fusion, characterized in that, Includes the following steps: S1, process the acquired millimeter-wave radar signal to generate three-dimensional point cloud data, and extract gait spatiotemporal features based on the three-dimensional point cloud data; S2, Based on the three-dimensional point cloud data, track and obtain trajectory features, input the trajectory features into the trajectory classification network, and output the wandering classification result; S3, process the acquired sound field signal to output sound source localization parameters, extract acoustic feature vectors, input the acoustic feature vectors into the acoustic classification model, and output acoustic classification results; S4, obtain the radar probability assignment constructed based on the gait spatiotemporal features and the wandering classification result, obtain the acoustic probability assignment constructed based on the sound source localization parameters and the acoustic classification result, fuse the radar probability assignment and the acoustic probability assignment through evidence theory to obtain the fusion confidence level, and trigger an abnormal warning when the fusion confidence level exceeds the warning threshold; S5, merge the gait spatiotemporal features with the acoustic feature vector to construct a cross-modal time series index, perform trend testing on the cross-modal time series index to generate a decay composite index, and output a trend warning when the decay composite index meets the preset conditions. 2.The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S1, the process of generating three-dimensional point cloud data from the acquired millimeter-wave radar signal includes: The millimeter-wave radar signal is subjected to intermediate frequency beat frequency processing, and then the range-dimensional fast Fourier transform and the Doppler-dimensional fast Fourier transform are performed sequentially to obtain the range-Doppler image. Two-dimensional ordered statistical constant false alarm rate (CFAR) detection is performed on the distance Doppler map. The power ranking values of the training units around the unit to be detected are extracted as noise estimates. The detection threshold is calculated based on the CFAR coefficient to extract the target detection point. A multi-signal classification algorithm is performed on the target detection points using a multi-input multi-output virtual array. The noise subspace is extracted by eigenvalue decomposition of the covariance matrix and a pseudo-spectrum is constructed to obtain the azimuth angle estimate. The three-dimensional point cloud data is generated by a density-based clustering algorithm. 3.The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S1, the extraction of gait spatiotemporal features based on the three-dimensional point cloud data includes: Calculate the change in the position of the human centroid in the three-dimensional point cloud data between consecutive time frames, and extract the instantaneous walking speed and average walking speed; Autocorrelation is performed on the time-series signal of the instantaneous gait speed to extract the gait period, and the step length is calculated in combination with the average gait speed. The coefficient of variation of multiple consecutive step lengths is then calculated as the step length coefficient of variation. The micro-Doppler spectrum is obtained by short-time Fourier transform. The oscillation period of the positive and negative Doppler frequency shifts in the micro-Doppler spectrum is extracted to calculate the gait asymmetry index. The ratio of the power spectral density of the high-frequency band to the motion band in the micro-Doppler frequency domain is extracted to calculate the frozen gait index. The instantaneous gait speed, the step length, the step length variation coefficient, the gait asymmetry index, and the frozen gait index are combined to generate the gait spatiotemporal features.
4. The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S2, the trajectory feature acquisition based on the 3D point cloud data includes: The three-dimensional point cloud data is tracked by an extended Kalman filter, and an acceleration component is introduced into the state vector for prediction and updating to obtain a trajectory coordinate sequence. Spatial complexity estimation and straightness calculation are performed on the trajectory coordinate sequence to extract trajectory statistical features including direction change frequency, path repetition rate, spatial coverage area, velocity variation coefficient, straightness exponent, fractal dimension estimated based on box counting method, and trajectory information entropy calculated based on spatial gridding. The trajectory coordinate sequence is merged with the trajectory statistical features to generate the trajectory features.
5. The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S2, the step of inputting the trajectory features into the trajectory classification network and outputting the wandering classification result includes: The trajectory coordinate sequence in the trajectory features is input into a one-dimensional convolutional neural network with a temporal attention layer. The output sequence of the convolutional network is self-attentionally weighted by the temporal attention layer to output the first hidden layer representation. The trajectory statistical features in the trajectory features are input into a multilayer perceptron with an activation function, and the second hidden layer representation is output. By concatenating the first hidden layer representation and the second hidden layer representation, and mapping the output to the classification function through a fully connected layer, the probability distributions of four wandering modes—direct walking, pacing, circling, and random wandering—are generated to produce the wandering classification result.
6. The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S3, the sound field signal obtained through processing outputs sound source localization parameters, and the acoustic feature vector is extracted, including: The sound field signal acquired by the ring microphone array is subjected to Fourier transform, and the cross-correlation function between microphone channel pairs is calculated by the generalized cross-correlation phase transform algorithm to extract the time delay estimate; The sound source angle of arrival is calculated based on the time delay estimation and the microphone channel pair spacing, and the angle of arrival confidence is calculated by combining the standard deviation of the multi-channel angle of arrival. The sound source angle of arrival and the angle of arrival confidence are combined to generate the sound source localization parameters. A noise covariance matrix and a target direction steering vector are constructed. The minimum variance distortionless response algorithm is used to perform beamforming processing on the sound field signal. After framing, Mel frequency cepstral coefficients and chromaticity map features are extracted to form the acoustic feature vector.
7. The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S3, inputting the acoustic feature vector into the acoustic classification model and outputting the acoustic classification result includes: The acoustic feature vectors are aggregated by time period and then input into a bidirectional long short-term memory network model to extract the forward hidden state and the backward hidden state respectively. The forward hidden state and the backward hidden state are concatenated, and the classification probability of normal speech, groaning, rapid breathing, helpless shouting, repetitive vocalization, crying, falling and impact sound, coughing, environmental noise and abnormal silence is output through the classification layer to generate the acoustic classification result; The triggering condition for the abnormal silence is that the millimeter-wave radar signal detects that the target human body is in an active state and the acoustic energy of the sound field signal is continuously lower than the environmental background for more than a preset time. 8.The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S4, obtaining the fusion confidence level by fusing the radar probability assignment and the acoustic probability assignment through evidence theory includes: Traverse the set of anomalous event hypotheses, calculate the sum of the cross products of the radar probability assignment and the acoustic probability assignment respectively, and obtain the evidence conflict coefficient; When the evidence conflict coefficient is not greater than a preset conflict threshold, the combination rule is applied to calculate the first fusion probability; when the evidence conflict coefficient is greater than the preset conflict threshold, the weighted average strategy based on the historical reliability score of the sensor is switched to calculate the second fusion probability. The Gaussian consistency function value of the deviation between the radar target azimuth angle and the sound source arrival angle is calculated. The time context weight and the Gaussian consistency function value are multiplied by the first fusion probability or the second fusion probability, respectively. The time-series evidence accumulation is performed through a sliding window exponential weighted average mechanism, and the fusion confidence is output. 9.The cognitive impairment old person behavior early warning method based on cross-modal fusion of claim 1, characterized in that, In S5, performing trend testing on the cross-modal time series index to generate a decay composite index includes: Establish a longitudinal cognitive health record, perform a sliding window linear regression test on the cross-modal time series indicators, and perform a nonparametric trend test on the test statistic of the rank difference of the measured value sequence to confirm the long-term decline slope. The cumulative offset of the cross-modal time series index from the target mean is calculated using a cumulative sum and control chart algorithm, and abrupt acceleration point detection is performed. The cross-modal time series indices with different physical dimensions are transformed into dimensionless relative deviations from the historical baseline. The gait deviation, stride length variability deviation, gait asymmetry deviation, pause frequency deviation, and vocabulary richness deviation are weighted and combined to construct the decline composite index.
10. A behavioral early warning system for elderly people with cognitive impairment based on cross-modal fusion, characterized in that, A method for early warning of cognitive impairment in elderly people based on cross-modal fusion as described in any one of claims 1-9, comprising the following modules: The radar perception and feature extraction module is used to process the acquired millimeter-wave radar signals to generate three-dimensional point cloud data, and extract gait spatiotemporal features based on the three-dimensional point cloud data. The trajectory tracking and wandering classification module is used to track and obtain trajectory features based on the three-dimensional point cloud data, input the trajectory features into the trajectory classification network, and output the wandering classification result. The sound field perception and acoustic classification module is used to process the acquired sound field signal, output sound source localization parameters, extract acoustic feature vectors, input the acoustic feature vectors into the acoustic classification model, and output acoustic classification results. The cross-modal fusion and anomaly warning module is used to obtain radar probability assignments constructed based on the gait spatiotemporal features and the wandering classification results, obtain acoustic probability assignments constructed based on the sound source localization parameters and the acoustic classification results, fuse the radar probability assignments and the acoustic probability assignments through evidence theory to obtain fusion confidence, and trigger an anomaly warning when the fusion confidence exceeds the warning threshold; The temporal merging and trend warning module is used to merge the gait spatiotemporal features and the acoustic feature vector to construct a cross-modal temporal index, perform trend testing on the cross-modal temporal index to generate a decay composite index, and output a trend warning when the decay composite index meets preset conditions.