A method and system for detecting a fall of a human body based on an acoustic signal
By collecting acoustic signals through speakers and microphones and combining them with a dual-stream long short-term memory network classification model, this technology solves the problems of battery life, privacy, and noise interference in existing fall detection technologies, achieving fall detection with a low false alarm rate, and is suitable for home audio devices.
Patent Information
- Application Number
- CN202211361322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-11-02
AI Technical Summary
Existing fall detection technologies suffer from drawbacks such as wearable device battery life issues, privacy leaks, noise interference, performance degradation in noisy environments, and high costs. In particular, sound-based solutions have a high false alarm rate and are easily affected by environmental noise.
By emitting ultrasonic waves through a loudspeaker and collecting acoustic signals through a microphone, and using a dual-stream long short-term memory network classification model combined with low-frequency and high-frequency signal feature extraction, fall detection is achieved, including signal splitting, noise reduction, endpoint segmentation, and frequency leakage elimination.
It achieves fall detection with low false alarm rate and no privacy risk, can be widely deployed in home audio devices, and is unaffected by noise and obstacles, providing timely fall intervention.
Smart Images

Figure CN115841824B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and more particularly, to a human fall detection method and system based on acoustic signals. BACKGROUND
[0002] Since the development of medical technology, the life expectancy of humans has been greatly extended, and many countries have experienced an aging population. The elderly are prone to fractures and soft tissue contusions due to falls, and if a fall is not detected in time and measures are not taken, it can lead to death. Therefore, there is an urgent need for a convenient, effective and automatic detection technology for the elderly to fall, so as to detect the fall event in time and provide medical assistance in time, thereby reducing the severity of the harm caused by the fall.
[0003] Existing fall detection technologies are mainly based on wearable devices, vision, wireless signals or sound-based solutions. The wearable device-based solution requires users to wear wearable devices at all times, and there are problems with battery life. When users forget to carry the device at all times, forget to charge it, or take it off to charge it, the detection capability is lost. The vision-based detection method can effectively detect fall events, but there are significant privacy issues and a high dependence on good lighting environments, and it cannot detect human activity in blind spots. In the wireless signal-based solution, the WiFi-based solution affects its original communication function, and the radar-based solution is often very expensive. In the sound-based solution, the system that only collects fall sound has a high false positive rate and is easily disturbed by environmental noise, and the system that only uses ultrasonic waves has a sharp performance decline in non-line-of-sight environments due to the blocking of ultrasonic waves. SUMMARY
[0004] The purpose of the present application is to overcome the defects of the above-mentioned prior art, and to provide a human fall detection method and system based on acoustic signals, which can be more accurate, have a low false positive rate, have no privacy risk, be free of wearing and be unaffected by noise and obstacles, and achieve human fall detection by making full use of acoustic signals.
[0005] According to a first aspect of the present application, a human fall detection method based on acoustic signals is provided. The method comprises the following steps:
[0006] controlling a loudspeaker to generate ultrasonic waves at a set frequency, and using a microphone to collect acoustic signals generated when a human falls, the acoustic signals including sound signals generated by the human collision and reflection signals of the ultrasonic waves;
[0007] splitting and denoising the acoustic signals to obtain low-frequency signals generated by the human collision and high-frequency signals of the ultrasonic waves reflected by the human;
[0008] detect the start point and the end point of the human activity in the low-frequency signal and detect the start point and the end point of the human activity in the high-frequency signal, and then determine the start point and the end point of the acoustic signal;
[0009] extract features from the low-frequency signal and the high-frequency signal respectively to obtain low-frequency signal features and high-frequency signal features;
[0010] input the low-frequency signal features and the high-frequency signal features into a pre-trained double-flow long short-term memory network classification model to obtain a fall detection result.
[0011] According to the second aspect of the present application, a human fall detection system based on fully utilizing acoustic signals is provided. The system comprises:
[0012] a signal acquisition module: used for controlling a loudspeaker to generate ultrasonic waves at a set frequency, and collecting acoustic signals generated when a human falls by using a microphone, wherein the acoustic signals include sound signals generated by human collision and reflection signals of the ultrasonic waves;
[0013] a signal processing module: used for splitting and denoising the acoustic signals to obtain low-frequency signals generated by human collision and high-frequency signals of the ultrasonic waves reflected by the human; detecting the start point and the end point of the human activity in the low-frequency signal and detecting the start point and the end point of the human activity in the high-frequency signal, and then determining the start point and the end point of the acoustic signal;
[0014] a feature extraction module: used for extracting features from the low-frequency signal and the high-frequency signal respectively to obtain low-frequency signal features and high-frequency signal features;
[0015] a model classification module: used for inputting the low-frequency signal features and the high-frequency signal features into a pre-trained double-flow long short-term memory network classification model to obtain a fall detection result.
[0016] Compared with the prior art, the present application has the advantages that ordinary audio equipment (loudspeaker and microphone) is used, the acoustic signals (sound signals generated by human collision with the ground and reflection signals of the ultrasonic waves) generated when falling are fully utilized, real-time monitoring for falling is realized, the recognition of falling is accurate and the false positive rate is low, and the user who falls can be intervened and treated in time.
[0017] Other features and advantages of the present application will become apparent from the following detailed description of exemplary embodiments thereof, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments of the present application and, together with the description, serve to explain the principles of the application.
[0019] Figure 1 is a flow chart of a human fall detection method based on acoustic signals according to an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of a framework of a human fall detection system based on acoustic signals according to an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of the effect of a signal after endpoint segmentation but before frequency leakage elimination according to an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of the effect of a signal after endpoint segmentation and frequency leakage elimination according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangements, numerical expressions, and numerical values of components and steps set forth in these embodiments are not limiting to the scope of the present application unless otherwise specifically stated.
[0024] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way intended to limit the scope of the application its application or uses.
[0025] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein. However, where appropriate, such techniques, methods, and devices can be viewed as part of the specification and can be claimed as such.
[0026] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.
[0027] It should be noted that like references and characters herein relate to like items throughout the figures, and once an item is defined in one figure, it need not be discussed further in subsequent figures.
[0028] Referring to Figure 1 As shown, the provided human fall detection method based on acoustic signals includes the following steps.
[0029] Step S1, use a common speaker to emit ultrasonic waves, and a microphone to collect acoustic signals generated when a human falls.
[0030] Specifically, the loudspeaker and microphone (or other devices capable of emitting and collecting sound) are arranged in the corner of the use environment, the loudspeaker is controlled to emit 20000Hz ultrasonic wave, the microphone collects the acoustic signal generated when the human body falls (including the sound signal generated by the collision of the human body with the ground and the reflection signal of the ultrasonic wave) and converts it into a digital electrical signal for subsequent processing. For example, the sampling frequency of the microphone is set to 48000Hz.
[0031] Step S2, the acoustic signal collected by the microphone is split and denoised.
[0032] In one embodiment, step S2 includes the following sub-steps:
[0033] Step S21, for the original acoustic signal received by the microphone, a Butterworth low-pass filter is applied to filter, obtaining the sound signal generated by the collision of the human body with the ground (low-frequency signal). For example, the filter order is set to two orders, and the cutoff frequency is set to 18000Hz.
[0034] Step S22, frame the original acoustic signal and use Fast Fourier transform (FFT) to obtain the real-time spectrum diagram, thereby obtaining the spectrum diagram of the part of the ultrasonic wave reflected by the human body (high-frequency signal part).
[0035] Specifically, under the sampling rate of 48000Hz, the original signal is decomposed into a series of small overlapping signals, each signal has a length of 0.3s, and the overlap rate of two consecutive signals is 90%. Then, multiply each frame signal by a Hamming window, and apply a 16384-point Fast Fourier Transform to each frame. Finally, the spectrum matrix X(f) = [X1(f) X2(f) … X k (f)],k∈R,f∈[19.4kHz,20.6kHz]. The high-frequency band [19400Hz, 20600Hz] contains most of the reflection signals of the ultrasonic wave by the human body movement, so the spectrum diagram of the reflection signal of the ultrasonic wave by the human body is obtained by intercepting the frequency range [19400Hz, 20600Hz] of f.
[0036] Step S23, use spectral subtraction to remove noise in the high-frequency signal spectrum diagram;
[0037] First, directly set the energy on the spectrum between 19980 to 20020Hz to zero to eliminate the effect of the direct transmission signal from the loudspeaker to the microphone through the direct path after being emitted.
[0038] Next, eliminate the noise caused by the hardware defects of the microphone. Specifically, estimate the spectral noise of the inactive frequency (higher than 20600Hz), and construct an amplitude histogram of the noise. The spectral segmentation threshold N th can be defined as Nth = μ + kσ, k e R, where μ is the mean value and σ is the standard deviation of the Gaussian approximation of the noise histogram. The value of k is chosen to adjust the segmentation threshold in different environments to preserve weak energy components around the center frequency and filter noise energy. Then, the threshold N th is applied to the spectral spectrogram to obtain where S(f, t) represents the power at frequency f and time t on the spectrogram.
[0039] Step S3, endpoint segmentation is performed on the denoised signal.
[0040] In one embodiment, step S3 includes the following sub-steps:
[0041] Step S31, activity in the low-frequency signal is detected using the root mean square frame energy threshold method. Assuming that the time-domain signal of the low-frequency signal separated from the original acoustic signal collected by the microphone is x(n), the nth frame signal is x n (m) = x((n - 1) * l + m), where l is the frame jump 0.27s, m e [0, N - 1], and N is the frame length 0.3s; then the root mean square frame energy of the nth frame low-frequency signal x n (m) is: When performing signal detection, when the signal lasts more than the threshold value and maintains for a period of time t (t is greater than 1s), it is considered that the signal is an active signal, and the starting point t s1 and the ending point t e1 are further extracted. Specifically, first, the starting point of the active signal is judged, and for example, the sampling point before the first frame exceeding the threshold value (i.e., the last sampling point of the previous frame) is selected as the starting point; then, the end of the signal is judged, for example, the first sampling point after the energy of the signal continuously for M frames is lower than the threshold value.
[0042] Step S32, activity in the high-frequency signal is detected using the power burst curve (PBC) threshold method. PBC can be calculated by the following formula: where PBC represents the sum of signal powers between frequencies f l and f u . For positive frequencies, f l = 20020 Hz, and f u = 20600 Hz. A threshold PBC th : where Nth is the noise threshold estimated in step S23. When the PBC of the high-frequency signal spectrogram exceeds the predefined threshold PBC th , it is considered to be an active signal. By identifying the two intersection points between PBC and PBC th noise threshold, the starting position t s2and end position t e2 .
[0043] Step S33, determine the final starting point t s and end point t e When the low-frequency signal has active signal and the high-frequency signal has no active signal, t s = t s1 , t e = t e1 ; on the contrary, when the low-frequency signal has no active signal and the high-frequency signal has active signal, t s = t s2 , t e = t e2 ; when both have active signal, t s = min(t s1 , t s2 ), t e = max(t e1 , t e2 ).
[0044] Step S4, eliminate the influence of the frequency leakage of the fall sound on the high-frequency signal spectrum.
[0045] When the high-frequency signal is weak, the sound produced by the fall or similar fall activity will leak frequency components to 20000Hz and mix with the ultrasonic reflection signal. In an embodiment, a frequency leakage elimination algorithm is proposed. First, the envelope line of the high-frequency signal spectrum is extracted, and then the derivative of the envelope line is calculated. Next, the highest point and the next point of the lowest point of the derivative are found, and then the two points on the envelope line corresponding to the two points are connected to obtain a new envelope line, and finally the spectrum outside the new envelope line is set to zero. The frequency leakage elimination algorithm described in this embodiment can well eliminate the influence of frequency leakage on the Doppler signal, and the effect diagram before and after the alignment processing is shown in Figure 3 and Figure 4 , so it can be seen that after the frequency leakage elimination processing of the high-frequency signal spectrum, the spectrum of the active signal is more in line with the real activity, which is conducive to more accurate judgment of the fall signal.
[0046] Step S5, respectively extract the features of the low-frequency signal and the high-frequency signal.
[0047] In an embodiment, step S5 includes the following sub-steps:
[0048] Step S51, calculate the Mel-frequency cepstral coefficient (MFCC) coefficient and linear predictive coding (LPC) of the low-frequency signal, extract the first four coefficients of the MFCC and the second to fifth coefficients of the LPC, and form an 8-dimensional time sequence as the feature of the low-frequency signal.
[0049] Step S52, three groups of features in the high-frequency signal spectrum interval [19400Hz, 20600Hz] are obtained, including: speed curve features, extreme value ratio curves and spectrum entropy, etc.
[0050] For example, for the speed curve feature, the frequency curve is extracted as follows:
[0051]
[0052] wherein, represents the energy accumulated between the frequency range f l and f, represents the total energy of the Doppler signal at time t. The threshold of T(f, t) is set to 30%, 75% and 95%, which respectively represents the frequency f corresponding to the trunk movement, leg movement and arm movement. A total of 6 speed curves, respectively corresponding to the arm movement, leg movement and trunk movement on the positive side and the negative side.
[0053] For the extreme value ratio curve, it is represented as:
[0054]
[0055] wherein, f +max (t) represents the maximum frequency shift above 20000Hz; f -min (t) represents the minimum frequency shift below 20000Hz, and the calculation method is similar to f +max (t).
[0056] For the spectrum entropy, the spectrum entropy H(t) at a certain time t is represented as follows:
[0057]
[0058] wherein p(f, t) is the normalized power spectral density at frequency f and time t, wherein P(f, t) is the power spectral density.
[0059] Further, the above three groups of features form an 8-dimensional time sequence as the features of the high-frequency signal spectrum.
[0060] Step S6, the extracted features are input into a multi-modal classification network for fall detection.
[0061] For example, a double-stream LSTM (Long Short-Term Memory) classification network is used, and the extracted features are input into a neural network model for fall detection.
[0062] In one embodiment, step S6 includes the following sub-steps:
[0063] Step S61, constructing a neural network model.
[0064] The double-flow LSTM classification network model includes two LSTM networks respectively designed for low-frequency signals and high-frequency signals, and the input dimensions of the two LSTM networks are both set to 8, corresponding to 8 feature sequences of the low-frequency signals and the high-frequency signals respectively. The output category of the LSTM is set to 2, corresponding to two events of falling down and non-falling down. The number of LSTM network layers is set to 2, and the hidden dimension is set to 4. The last two LSTM networks are fused by a Trusted Multi-View Classification algorithm (Han Z, Zhang C, Fu H, et al. Trusted multi-view classification [J]. arXiv preprint arXiv:2102.02051, 2021.) for decision fusion, and are trained together.
[0065] In step S62, the double-flow LSTM classification network is trained and tested.
[0066] In the training phase, the falling down activity and the non-falling down activity are respectively marked as 0 and 1 as data labels for model training; in the test phase, the test results of the trained double-flow LSTM classification network are output.
[0067] In step S62, based on the detection method of the double-flow LSTM classification network, the extracted low-frequency features and high-frequency features are input into the double-flow LSTM classification network model to obtain detection results.
[0068] In step S7, the detection results are fed back to the user, and if a falling down activity is detected, the associated smartphone APP is displayed in time.
[0069] Correspondingly, the application also provides a falling down detection system based on acoustic signals, which is used to realize one or more aspects of the above method. For example, the system includes: a signal acquisition module for controlling the loudspeaker to emit ultrasonic waves and the microphone to collect acoustic signals; a signal processing module for splitting, denoising, endpoint segmentation and frequency leakage elimination of the acoustic signals received by the microphone; a feature extraction module for extracting features from low-frequency signals and high-frequency signals respectively; a model classification module for inputting the extracted features into a neural network model for falling down detection; and a result feedback module for feeding back the detection results to the user and displaying them in the system.
[0070] Further, the signal acquisition module further includes:
[0071] A signal acquisition unit for controlling the loudspeaker to emit ultrasonic waves and the microphone to collect acoustic signals.
[0072] Further, the signal processing module further includes:
[0073] The signal splitting unit splits the received original acoustic signal to obtain a frequency spectrum of a sound signal (low-frequency signal) generated by the collision of the human body with the ground and a frequency spectrum of a part of the ultrasonic wave reflected by the human body (high-frequency signal part).
[0074] The spectrum denoising module removes noise in the frequency spectrum of the high-frequency signal using spectral subtraction.
[0075] The endpoint segmentation unit detects the start point and end point of the user activity signal using a frame energy root mean square threshold method and a power burst curve threshold method to segment a single activity signal.
[0076] The frequency leakage elimination unit eliminates the influence of the frequency leakage of the fall sound on the ultrasonic Doppler signal using the proposed frequency leakage elimination algorithm.
[0077] Further, the feature extraction module further comprises:
[0078] The low-frequency signal feature extraction unit extracts the first four coefficients of MFCC and the second to fifth coefficients of LPC from the low-frequency signal to form an 8-dimensional time sequence feature.
[0079] The high-frequency signal feature extraction unit extracts three groups of features from the frequency spectrum interval [19400Hz, 20600Hz] of the high-frequency signal, including a speed curve feature, an extreme value ratio curve, and a spectral entropy. The three groups of features form an 8-dimensional time sequence feature.
[0080] Further, the model classification module further comprises:
[0081] The neural network construction unit constructs a dual-flow LSTM classification network based on an LSTM network and a multi-view classification algorithm Trusted Multi-View Classification.
[0082] The training and testing unit has a training phase and a testing phase. In the training phase, the fall activity and the non-fall activity are recorded as 0 and 1 respectively as data labels for model training; in the testing phase, the testing result of the trained dual-flow LSTM classification network is output.
[0083] The classification unit inputs the extracted low-frequency features and high-frequency features into the dual-flow LSTM classification network model based on the detection method of the dual-flow LSTM classification network and obtains a detection result.
[0084] Further, the result feedback module comprises a visualization unit for visualizing and displaying the detection result in the associated intelligent device APP.
[0085] In summary, the present application solves the problems of the prior art, such as the inconvenience of wearing devices, the privacy leakage of cameras, the poor robustness of wireless signals, and the problems of sound being easily disturbed by noise and ultrasonic waves being easily blocked. In addition, the detection method of the present application is novel, convenient and reliable, can be widely deployed in current home audio devices, and is harmless to the human body.
[0086] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.
[0087] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a magneto-optical or other optical medium, and / or any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0088] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0089] Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0090] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0091] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or nonvolatile memory, or a suitable combination of the different types of computer readable storage media. The computer readable program instructions can also be downloaded to a computer, other programmable data processing apparatus, or other device from a computer readable storage medium or to an external computer or external storage device via a data signal that can be transmitted for example via a wired medium or a wireless medium such as the Internet or wireless media.
[0092] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0093] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0094] Embodiments of the present application have been described above, and the description is intended to be illustrative, and not restrictive, of the various embodiments of the present application. Many modifications and variations of the described embodiments of the present application are possible, given the benefit of the present disclosure, without departing from the scope and spirit of the described embodiments of the present application. The scope of the present application is defined by the appended claims.
Claims
1. A method for human fall detection based on acoustic signals, comprising the following steps: controlling a speaker to generate ultrasonic waves at a set frequency, and using a microphone to collect acoustic signals generated when a human falls, the acoustic signals including sound signals generated by the human body colliding and reflection signals of the ultrasonic waves; splitting and denoising the acoustic signals to obtain low-frequency signals generated by the human body colliding and high-frequency signals of the human body reflecting the ultrasonic waves; detecting the start point and end point of human activity in the low-frequency signals and detecting the start point and end point of human activity in the high-frequency signals, and then determining the start point and end point of the acoustic signals; extracting features from the low-frequency signals and the high-frequency signals respectively to obtain low-frequency signal features and high-frequency signal features; inputting the low-frequency signal features and the high-frequency signal features into a pre-trained dual-stream long short-term memory network classification model to obtain a fall detection result; wherein the low-frequency signal features are obtained according to the following steps: obtaining the inverse Mel-frequency cepstral coefficients (MFCC) and linear predictive coding (LPC) of the low-frequency signals, extracting the first four coefficients of the MFCC and the second to fifth coefficients of the LPC, and combining them into an 8-dimensional time sequence feature as the low-frequency signal features; the high-frequency signal features correspond to three groups of features in a set frequency spectrum interval [19400 Hz, 20600 Hz], which are a velocity curve feature, an extreme value ratio curve, and a spectral entropy, and the three groups of features form an 8-dimensional time sequence feature; wherein, for the velocity curve feature, the frequency curve is extracted as: wherein, represents the energy accumulated in the frequency range to , represents the total energy of the Doppler signal at time , the threshold of is set to 30%, 75% and 95%, respectively, representing the frequency corresponding to the trunk movement, leg movement and arm movement, a total of 6 speed curves, respectively, corresponding to the positive and negative side arm movement, leg movement and trunk movement; for the extreme value ratio curve, it is expressed as: wherein represents the maximum frequency deviation above 20000 Hz, represents the minimum frequency deviation below 20000 Hz; For the spectral entropy, at a certain moment The spectral entropy H(t) is expressed as: wherein is the frequency and the normalized power spectral density at time , is the power spectral density, and denotes the frequency value. 2. The method of claim 1, wherein, the sampling frequency of the microphone is set to 48000 Hz, and the speaker is controlled to emit ultrasonic waves at 20000 Hz.
3. The method of claim 1, wherein, Splitting and denoising the acoustic signals includes: for the acoustic signals, a Butterworth low-pass filter is used to filter to obtain the low-frequency signals generated by the human body colliding; for the acoustic signals, frame processing is performed and the fast Fourier transform is used to obtain a high-frequency signal spectrum graph of the human body reflecting the ultrasonic waves; spectral subtraction is used to remove noise in the high-frequency signal spectrum graph.
4. The method of claim 1, wherein, The start point and end point of the acoustic signals are determined according to the following steps: detecting the starting point and the ending point of the human activity in the low-frequency signal by using a root mean square threshold method based on frame energy and ending point of the human activity in the low-frequency signal detecting an onset point of the human activity in the high frequency signal using a power burst curve threshold method and an end point Determine the starting point of the acoustic signal. and end point When there is an active signal in the low-frequency signal but no active signal in the high-frequency signal, , When there is no activity in the low-frequency signal but there is activity in the high-frequency signal, , When both low-frequency and high-frequency signals have active signals, ) , .
5. The method of claim 1, wherein, before extracting the high-frequency signal features, the frequency leakage of the fall sound is eliminated from the high-frequency signals according to the following steps: extracting the high-frequency signal spectrum envelope and calculating the derivative of the envelope; finding the highest point of the derivative as the first point, finding the next point of the lowest point of the derivative as the second point, and connecting the two points on the envelope corresponding to the first point and the second point to obtain a new envelope; setting the spectrum outside the new envelope to zero.
6. The method of claim 1, wherein, Further comprising: feeding back the detection result to the user, and if a fall activity is detected, displaying it in the associated smart device APP. 7.A human fall detection system based on fully utilizing acoustic signals, comprising: a signal collection module: for controlling a speaker to generate ultrasonic waves at a set frequency, and using a microphone to collect acoustic signals generated when a human falls, the acoustic signals including sound signals generated by the human body colliding and reflection signals of the ultrasonic waves; A signal processing module is configured to split and denoise the acoustic signal, and obtain a low-frequency signal generated by human body collision and a high-frequency signal reflected by the human body; The start point and the end point of human activity in the low-frequency signal are detected, and the start point and the end point of human activity in the high-frequency signal are detected, so as to determine the start point and the end point of the acoustic signal; A feature extraction module is configured to extract features from the low-frequency signal and the high-frequency signal respectively, and obtain low-frequency signal features and high-frequency signal features; A model classification module is configured to input the low-frequency signal features and the high-frequency signal features into a pre-trained double-flow long short-term memory network classification model, and obtain a fall detection result; The low-frequency signal features are obtained according to the following steps: The inverse Mel spectrum coefficients MFCC and linear predictive coding LPC of the low-frequency signal are calculated, the first four coefficients of the MFCC and the second to fifth coefficients of the LPC are extracted, and an 8-dimensional time sequence feature is formed as the low-frequency signal features; The high-frequency signal features correspond to three groups of features in the frequency spectrum interval [19400Hz, 20600Hz], which are the speed curve feature, the extreme value ratio curve and the spectral entropy, and the three groups of features form an 8-dimensional time sequence feature; For the speed curve feature, the frequency curve is extracted and expressed as: wherein, represents the energy accumulated in the frequency range to , represents the total energy of the Doppler signal at time , the threshold values of are set to 30%, 75% and 95%, respectively, representing the frequency corresponding to the trunk movement, leg movement and arm movement, respectively; a total of 6 speed curves, corresponding to the positive and negative sides of the arm movement, leg movement and trunk movement, respectively. For the extreme value ratio curve, it is expressed as: wherein represents the maximum frequency deviation above 20000 Hz, represents the minimum frequency deviation below 20000 Hz; For the spectral entropy, at a certain moment The spectral entropy H(t) is expressed as: wherein is the frequency and the normalized power spectral density at time , , is the power spectral density, and denotes the frequency value.
8. A computer readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to realize the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Fall detection method and device, electronic equipment and storage medium
CN113450537A