Fall detection method, device, earphones and storage medium
By collecting ear canal audio signals in headphones and using a fall detection model, the problem of deploying multiple sensors on the user in existing technologies is solved, convenient fall detection is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202280004515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-06-15
AI Technical Summary
In the existing technology, falls are detected by deploying multiple sensors or wearing sensor clothing on the target population, which may cause discomfort and movement disorders to people with limited mobility.
The feedback microphone in the headset is used to collect audio signals in the ear canal. Through feature extraction and filtering technology, features to be identified are generated and input into the fall detection model to detect whether the user has fallen.
There is no need to add additional sensors to the user, which simplifies the detection process and improves the convenience of detection and user experience.
Smart Images

Figure CN117597065B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information processing technology but is not limited to the field of information processing technology, and in particular to a fall detection method, device, earphone, and storage medium. Background Art
[0002] With the development of technology, more and more electronic devices have appeared in various application scenarios, and different electronic devices can realize different functions in corresponding application scenarios.
[0003] With the widespread use of health monitoring devices, the health status of the subjects can be monitored. For example, fall detection can be performed by having the target person wear appropriate sensors or clothing equipped with corresponding sensors. The sensor signals can then be used to determine whether the target person has fallen. This target group can include the elderly and people with limited mobility. Summary of the Invention
[0004] Embodiments of the present disclosure provide a fall detection method, device, earphones, and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a fall detection method is provided, which is applied to a headset, wherein the headset includes a feedback microphone. The method includes:
[0006] The feedback microphone collects audio signals in the ear canal to obtain an ear canal audio signal; wherein the ear canal audio signal includes: an audio signal generated by vibration generated by the user's body colliding with the ground when the headset is worn by the user and is transmitted to the ear canal via bone conduction when the user falls;
[0007] Extracting features of the ear canal audio signal to obtain audio signal feature parameters;
[0008] Filtering the ear canal audio signal according to a preset frequency range to obtain a periodic signal within the preset frequency range; wherein the periodic signal includes a peak value or a valley value of a waveform;
[0009] generating a feature to be identified according to the audio signal characteristic parameter and the number of the peak values or the valley values;
[0010] The feature to be identified is input into a fall detection model to obtain a detection result; wherein the detection result is used to indicate that the user has fallen.
[0011] According to a second aspect of an embodiment of the present disclosure, a fall detection device is provided, which is applied to a headset, wherein the headset includes a feedback microphone, and the device includes:
[0012] an ear canal audio signal detection module configured to collect audio signals in the ear canal through the feedback microphone to obtain an ear canal audio signal; wherein the ear canal audio signal includes: an audio signal generated by vibration generated by the user's body colliding with the ground when the headset is worn by the user and transmitted to the ear canal via bone conduction when the user falls;
[0013] an audio signal characteristic parameter acquisition module, configured to extract features of the ear canal audio signal to obtain audio signal characteristic parameters;
[0014] a period information determination module configured to filter the ear canal audio signal according to a preset frequency range to obtain a periodic signal within the preset frequency range; wherein the periodic signal includes a peak value or a valley value of a waveform;
[0015] a feature generation module to be identified, configured to generate a feature to be identified based on the audio signal characteristic parameter and the number of the peak values or the valley values;
[0016] The detection module is configured to input the features to be identified into a fall detection model to obtain a detection result; wherein the detection result is used to indicate that the user has fallen.
[0017] According to a third aspect of an embodiment of the present disclosure, there is provided an earphone, comprising a shell and a controller, a feedback microphone, a feedforward microphone and a speaker arranged on the shell; the feedforward microphone is connected to the controller for collecting audio data outside the ear canal and sending it to the controller; the feedback microphone is connected to the controller for collecting audio data inside the ear canal and sending it to the controller; the controller comprises a memory and a processor, the memory storing executable computer instructions, and the processor being capable of calling the computer instructions stored in the memory to execute the method described in any one of the above embodiments.
[0018] A fourth aspect of the embodiments of the present disclosure provides a computer storage medium storing an executable program. After the executable program is executed by a processor, the fall detection method provided in the first aspect can be implemented.
[0019] The fall detection method provided in the embodiments of the present disclosure can be applied to headphones. Whether a user has fallen can be determined through the headphones without the need for other detection sensors, thereby improving the convenience of detecting user falls and enhancing the user experience.
[0020] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory and are not restrictive of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the embodiments of the present invention.
[0022] Figure 1 is a schematic diagram showing a fall detection method according to an exemplary embodiment;
[0023] Figure 2 is a schematic diagram of an earphone according to an exemplary embodiment;
[0024] Figure 3 is a schematic diagram showing an earphone being worn by a user according to an exemplary embodiment;
[0025] Figure 4 is a schematic diagram showing another fall detection method according to an exemplary embodiment;
[0026] Figure 5 is a schematic diagram showing a periodic signal according to an exemplary embodiment;
[0027] Figure 6 is a schematic diagram of a fall detection device according to an exemplary embodiment;
[0028] Figure 7 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0029] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible implementations consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of the present invention.
[0030] The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present disclosure. The singular forms "a," "the," and "the" used in the present disclosure are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0031] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the embodiments of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0032] Typically, fall detection for the elderly or people with mobility impairments, such as those with leg problems, is performed using multiple sensors. For example, multiple accelerometers are deployed on these individuals' bodies, or they wear clothing equipped with accelerometers. Changes in the accelerometer signals are used to determine if the user has fallen.
[0033] Since these users have limited mobility, adding multiple sensors to these users through this method will have an impact on them, causing discomfort and further inconvenience in their mobility.
[0034] refer to Figure 1 , which shows a schematic diagram of a fall detection method provided by an embodiment of the present disclosure. The method can be applied to at least a headset, and the headset can at least include a feedback microphone.
[0035] like Figure 1 As shown, the method includes:
[0036] Step S100, collecting audio signals in the ear canal through a feedback microphone to obtain an ear canal audio signal; wherein the ear canal audio signal includes: when the earphone is worn by the user, the vibration generated by the user's body colliding with the ground when the user falls is transmitted to the ear canal through bone conduction.
[0037] Step S200: extract features of the ear canal audio signal to obtain audio signal feature parameters.
[0038] Step S300: Filter the ear canal audio signal according to a preset frequency range to obtain a periodic signal in the preset frequency range; wherein the periodic signal includes peak values or valley values of the waveform.
[0039] Step S400: Generate features to be identified based on audio signal characteristic parameters and the number of peaks or valleys.
[0040] Step S500: Input the features to be identified into a fall detection model to obtain a detection result; wherein the detection result is at least used to indicate that the user has fallen.
[0041] Headphones can be of different shapes, including in-ear, semi-in-ear, and head-mounted. Headphones can be wired or wireless. Wireless headphones can include Bluetooth headphones, such as True Wireless Stereo (TWS). Headphones can also include hearing aids and other devices with feedback microphones that can implement this solution.
[0042] The feedback microphone in the earphones can be located near the sound output channel of the earphones. When the earphones are worn, the feedback microphone is located in the ear canal and can collect audio signals in the ear canal. For example, in-ear earphones. In other earphone types, such as semi-in-ear earphones and headphones, the feedback microphone can only collect audio signals in the ear canal when the earphones are worn.
[0043] refer to Figure 2 , is a schematic diagram of an earphone, including a feedback microphone A. When the earphone is in a worn state, the feedback microphone A can be located in the ear canal. The earphone can also include a feedforward microphone B. The feedforward microphone B can be located on the earphone handle. When the earphone is in a worn state, the feedforward microphone is located outside the ear canal and can collect ambient audio signals from the external environment. The earphone can also include a call microphone C to collect audio signals emitted by the user in a call state. The audio signal collected by the feedback microphone A has a higher signal-to-noise ratio than that collected by the feedforward microphone B, so the feedback microphone A collects ear canal audio signals with less noise and higher quality.
[0044] When the earphones are head-mounted earphones, a certain degree of ear-blocking effect will also be formed. When the earphones are worn, the feedback microphone A can be located outside the ear canal or facing the ear canal, and can only collect the ear canal audio signal.
[0045] refer to Figure 3 , a schematic diagram of a headset being worn by a user. Earphone 1 blocks ear canal 2 to a certain extent, creating an ear-blocking effect. The ear canal audio signal collected by the feedback microphone includes: The audio signal generated by the vibrations generated by the user's breathing when the headset is worn is transmitted to the ear canal via bone conduction, i.e., audio signal 3.
[0046] For step S100, when the earphones are in the worn state, the ear canal is blocked to a certain extent, forming a certain degree of ear-blocking effect. The reason for this is that some sound is conducted to the inner ear through the bones. For example, when walking, the feet will vibrate in contact with the ground, and the vibration is transmitted to the audio signal of the ear canal through bone conduction. When the earphones are not in the worn state, part of the sound from bone conduction diffuses outward through the outer ear. However, when the earphones are in the worn state, the ear canal is blocked to a certain extent, reducing the amount of sound from bone conduction that diffuses outward through the ear canal, forming a certain degree of ear-blocking effect, also known as the occlusion effect. The sound characteristics produced by the occlusion effect are characterized by the enhancement of low-frequency signals and the weakening of high-frequency signals.
[0047] Because earphones block the ear canal to a certain extent, creating a varying degree of occlusion, the earphones prevent external audio signals from entering the ear canal, reducing their impact on the audio signals within the ear canal. A feedback microphone can collect audio signals within the ear canal to generate a breathing audio signal. This breathing audio signal includes the audio signal generated by the vibrations generated by the user's feet contacting the ground while walking while the earphones are worn, which are transmitted to the ear canal via bone conduction.
[0048] After a certain degree of ear-blocking effect is formed, the vibration generated by the user's feet contacting the ground when walking can be conducted to the ear canal through the bones, generating an audio signal. The ear-blocking effect can amplify the audio signal, thereby facilitating the feedback microphone to collect the audio signal in the ear canal and obtain the ear canal audio signal.
[0049] In step S200, after obtaining the ear canal audio signal, feature extraction can be performed on the ear canal audio signal to obtain audio signal feature parameters. Feature extraction methods may include various methods, such as extracting corresponding features through a feature extraction algorithm. The obtained audio signal feature parameters may be Mel-spectral coefficients and Mel-spectral cepstral coefficients (MFCCs), both of which are 40-dimensional feature parameters. Of course, other features of the ear canal audio signal may also be used.
[0050] When the ear canal audio signals are different, the corresponding audio signal characteristics will also be different, and different ear canal audio signals have their own audio signal characteristics. The audio signal characteristics of the audio signal generated by the vibration generated by the contact between the user's body and the ground when the user falls is transmitted to the ear canal through bone conduction are different from the audio signal characteristics of the audio signal generated by the vibration generated by the collision between the user's body and the ground when the user is not falling and transmitted to the ear canal through bone conduction. The audio signal characteristics are also different from the audio signal characteristics of other vibrations transmitted to the ear canal through bone conduction, for example, the audio signal characteristics of speech and external environmental sound signals.
[0051] In step S300, after obtaining the ear canal audio signal, the ear canal audio signal can be filtered according to a preset frequency range to obtain a periodic signal within the preset frequency range, wherein the periodic signal includes the peak value of the waveform. The preset frequency range can be determined according to actual usage requirements or can be preset. It can also be determined based on the walking frequency of a preset number of users. For example, the preset frequency range can be 1 Hz to 50 Hz, and the ear canal audio signal is low-pass filtered according to the preset frequency range.
[0052] According to the preset frequency range, the audio signals outside the preset frequency range in the ear canal audio signal are filtered out to obtain the audio signals with frequencies within the preset frequency range. The audio signals within the preset frequency range may be periodic signals, and the periodic signals may be signals in the time domain.
[0053] This periodic signal is represented as a waveform, including the relationship between time and amplitude. The peak values of the peaks in the waveform can be determined based on the periodic signal. The peak value of the peak in the waveform indicates that the user's foot made contact with the ground at the time corresponding to that peak value. Each peak value represents one contact point between the user's foot and the ground. The number of peak values can be used to determine the number of steps taken by the user.
[0054] The periodic signal can be used to determine the troughs of the waveform. Each trough in the waveform indicates that the user's foot was in contact with the ground at that trough. Each trough represents one contact with the ground. The number of troughs can be used to determine the number of steps the user has taken.
[0055] There is no necessary order between step S200 and step S300. Step S200 can be executed first, or step S300 can be executed first.
[0056] For step S400, after determining the number of peaks or valleys and the characteristic parameters of the audio signal, the feature to be identified can be generated based on the number of peaks and the characteristic parameters of the audio signal, or based on the number of valleys and the characteristic parameters of the audio signal. For example, the number of peaks and the characteristic parameters of the audio signal can be used together as a feature to be identified, and the feature to be identified includes information of two dimensions, the number of peaks and the characteristic parameters of the audio signal, for determining the detection result. The number of valleys and the characteristic parameters of the audio signal can also be used together as a feature to be identified, and the feature to be identified includes information of two dimensions, the number of valleys and the characteristic parameters of the audio signal, for determining the detection result.
[0057] When the number of peaks and / or audio signal characteristic parameters are different, the resulting features to be identified are different. When the number of valleys and / or audio signal characteristic parameters are different, the resulting features to be identified are different. When at least one of the number of peaks and the audio signal characteristic parameters changes, the features to be identified change. When at least one of the number of valleys and the audio signal characteristic parameters changes, the features to be identified change. This can reduce the situation where the resulting features to be identified remain unchanged when the number of peaks and the audio signal characteristic parameters change simultaneously, or when the number of valleys and the audio signal characteristic parameters change simultaneously, thereby improving the accuracy of the detection results.
[0058] In step S500, the features to be identified are input into a fall detection model to obtain a detection result. The detection result is used to at least indicate that the user has fallen. The fall detection model is a detection model that has been trained in advance.
[0059] Since when the earphones are worn by the user, the vibration generated by the collision of the user's body with the ground when the user falls is transmitted to the ear canal through bone conduction, the audio signal characteristics of the audio signal generated are different from the audio signal characteristics of other audio signals, and the number of peaks or valleys in the corresponding periodic signal is also different, so the features to be identified are input into the fall detection model to determine whether the user has fallen.
[0060] Of course, it can also be determined in other ways, such as a mapping table, which includes a mapping relationship between the features to be identified and the detection results. By looking up the mapping table, the detection result can be determined according to the features to be identified.
[0061] The disclosed example can obtain the user's ear canal audio signal through headphones, then process the ear canal audio signal and use a fall detection model to detect whether the user has fallen. By utilizing the existing built-in microphone, without the need for new hardware costs or the use of various other sensors or other monitoring equipment, the headset can be used to determine whether the user has fallen. This reduces the difficulty and inconvenience of detecting user falls, improves the convenience of detecting user falls, and reduces the discomfort and inconvenience caused to the user during the process of detecting a user's fall, thereby improving the user experience.
[0062] In another embodiment, when the ear canal audio signal does not include an audio signal generated when the earphone is worn by the user and the vibration generated by the user's body colliding with the ground when falling is transmitted to the ear canal through bone conduction, then through steps S200 to S500, a detection result indicating that the user has not fallen can be obtained, and the detection result can be output at this time.
[0063] In one embodiment, reference Figure 4 , Figure 4FIG. 1 is a schematic diagram of another fall detection method. The method further includes:
[0064] Step S10 , determining a target frame length of at least one frame period signal according to a duration corresponding to a preset number of steps.
[0065] Step S20 , determining the number of peak values or valley values included in each frame period signal within the target frame length; wherein each step corresponds to a peak value or valley value.
[0066] Before generating the features to be identified, it is also necessary to determine the number of peaks or valleys, and the target frame length of at least one frame periodic signal can be determined based on the duration corresponding to the preset number of steps. The preset number of steps can be determined based on actual application requirements, for example, it can be the duration of two steps, the duration of three steps, or the duration of four steps, etc. The duration corresponding to the preset number of steps can be in seconds, and the duration corresponding to the preset number of steps can be determined as the target frame length of one frame periodic signal, or the duration corresponding to the preset number of steps can be determined as the target frame length of N frame periodic signals. In this embodiment, taking N equal to 1, that is, the duration corresponding to the preset number of steps determines the target frame length of one frame periodic signal as an example, this can improve the accuracy and convenience of determining the number of peaks or valleys.
[0067] After determining the target frame length, the number of peaks or valleys included in each frame periodic signal can be determined within the target frame length. Since the periodic signal is represented in the form of a waveform, the peak or valley of the waveform can also be determined based on the periodic signal. The peak value of the peak can be determined based on the peak, and the valley value of the valley can be determined based on the valley interface. According to the target frame length, the number of peaks or valleys within the target frame length is determined. Each step corresponds to a peak or valley. Since vibration is generated when the user's foot contacts the ground, the intensity of the audio signal is greater than the intensity of the corresponding audio signal when the foot is not in contact with the ground. Therefore, the number of peaks or valleys within the target frame length can represent the number of steps taken by the user, that is, the number of times the user's foot contacts the ground.
[0068] The method can be used to determine the number of corresponding peak values or valley values in each frame period signal.
[0069] The number of peaks can be determined by detecting each frame period signal using a peak detection algorithm to determine the peak value in each frame period signal. The number of valleys can be determined by detecting each frame period signal using a valley detection algorithm to determine the valley value in each frame period signal. The peak detection algorithm can detect the peaks of the waveform in the periodic signal and then determine the corresponding peaks, while the valley detection algorithm can detect the valleys of the waveform in the periodic signal and then determine the corresponding valley values.
[0070] In another embodiment, step S400 may generate the feature to be identified based on the number of peaks in a one-frame periodic signal, or based on the number of peaks in a multi-frame periodic signal. For example, the feature to be identified may be generated based on the average value of the number of peaks in a multi-frame continuous periodic signal and the audio signal characteristic parameter.
[0071] The features to be identified are generated based on the number of peaks in the multi-frame periodic signal. This can reduce the impact of abnormal peaks in the periodic signal of a few frames on the number of peaks, thereby reducing the impact on the features to be identified, reducing the impact on the detection results, and improving the accuracy of the detection results.
[0072] refer to Figure 5 , Figure 5 A schematic diagram of a periodic signal. Figure 5 The periodic signal shown is a periodic signal obtained by filtering the audio signal corresponding to normal walking without falling according to a preset frequency range. Figure 5 Each target frame length corresponds to the duration of two steps, which is 1 second. In the positive direction of the amplitude, each periodic signal includes three peaks within the target frame length. Taking the periodic signal corresponding to the leftmost target frame length as an example, each peak indicates that the foot is in contact with the ground, indicating that the user is moving with their feet. The duration from the leftmost peak to the middle peak is the duration of the first step, and the duration from the middle peak to the rightmost peak is the duration of the next step.
[0073] Figure 5 The peak value of the peak or the valley value of the trough corresponding to each target frame length is represented by a dot, and the peak value or valley value represented by the dot can be determined by the corresponding detection algorithm.
[0074] In another embodiment, the ear canal audio signal includes an audio signal generated when the earphone is worn by the user and the vibration generated by the collision of the user's body with the ground when the user falls is transmitted to the ear canal through bone conduction. Since multiple parts of the body will contact and collide with the ground when falling, the number of peaks or valleys included in each cycle signal of the target frame length is greater than the number of corresponding peaks or valleys when the user is walking normally without falling.
[0075] according to Figure 5 It can be concluded that when walking normally without falling, the corresponding peaks or troughs are periodic and have a certain regularity. The difference between each peak value remains within a certain range, and the size of each peak value is similar, without any sudden changes in size. When falling, the peaks or troughs corresponding to the collision of multiple different parts of the body with the ground are irregular, and the corresponding peaks or troughs are also irregular. The impact force of different parts of the body with the ground is also different, resulting in different peaks of the corresponding peaks or troughs.
[0076] In another embodiment, the fall detection model is obtained by pre-training an initial neural network model using a machine learning approach based on a training sample set of fall information. The structure of the initial neural network model is not limited; after being trained using the training sample set using machine learning, it can output recognition results based on the features to be recognized.
[0077] In one embodiment, the training sample set includes a positive sample set, and the positive sample set includes a plurality of positive samples. Each positive sample includes: an audio feature of a target collision, a first quantity, and a first label.
[0078] The audio signature of a target collision is derived from the audio signal captured by the feedback microphone in the ear canal when the user's body collides with the ground during a fall. After the feedback microphone in the earphones captures the low-frequency audio signal in the ear canal, the earphones perform feature extraction on the captured audio signal to generate the audio signature of the target collision.
[0079] The first number is: the number of peaks or valleys in a waveform included in a periodic signal within a preset frequency range obtained after filtering the audio features of the target collision according to the preset frequency range.
[0080] The first label is used to represent the audio feature of the target collision and the first quantity corresponds to the output of the initial neural network model.
[0081] Each positive sample in the positive sample set is input into the initial neural network model, with the first label used as the output of the initial neural network. The initial neural network model is trained to obtain a fall detection model. The number of positive samples in the positive sample set can be determined based on actual needs. A larger number of positive samples results in higher detection accuracy for the trained fall detection model.
[0082] The audio features of the target collision may include Mel spectrum feature parameters and Mel cepstral feature parameters, and the Mel spectrum coefficients and Mel cepstral coefficients may both be 40-dimensional coefficients.
[0083] In another embodiment, the audio features of the target collision included in different positive samples correspond to different postures and / or collision times of the user's body colliding with the ground when falling.
[0084] Each positive sample may be an audio feature of an audio signal generated when different parts of the body of different users collide with the ground in different falling postures, and different positive samples include target collision audio features corresponding to different postures, parts of the body colliding with the ground, and / or number of collisions when the user falls.
[0085] Each positive sample may also be the audio features of the audio signal of the same user when different parts of the body collide with the ground in different falling postures, and different positive samples include target collision audio features corresponding to different postures, parts of the body that collide with the ground, and / or the number of collisions when the user falls.
[0086] For example, positive sample 1 includes the audio features of the body colliding with the ground when user 1 falls in posture 1, the number of times different parts of the body of user 1 collide with the ground when user 1 falls in posture 1 is number 1, and the parts that collide with the ground include the hands and knees. Positive sample 2 includes the audio features of the body colliding with the ground when user 2 falls in posture 2, the number of times different parts of the body of user 2 collide with the ground when user 2 falls in posture 2 is number 2, and the parts that collide with the ground include the hands and buttocks. Positive sample 3 includes the audio features of the body colliding with the ground when user 1 falls in posture 2, the number of times different parts of the body of user 1 collide with the ground when user 1 falls in posture 2 is number 3, and the parts that collide with the ground include the back and head.
[0087] In another embodiment, the first number is the number of peaks or valleys within the target frame length. The first number and the number of peaks or valleys in the positive sample that generate the feature to be identified are both determined within the same frame length, which can reduce variables and improve detection accuracy.
[0088] In one embodiment, the training sample set further includes a negative sample set, wherein the negative sample set includes a plurality of negative samples, each of which includes: an audio feature of a non-target collision, a second quantity, and a second label.
[0089] The audio characteristics of non-target collisions are audio characteristics obtained by collecting audio signals in the ear canal through a feedback microphone when the earphones are worn by the user, except when the user collides with the ground when falling.
[0090] The second number is: the number of peaks or valleys of a waveform included in a periodic signal within a preset frequency range obtained after filtering the audio features of the non-target collision according to the preset frequency range.
[0091] The second label is used to represent the audio features of the non-target collision and the output of the initial neural network model corresponding to the second quantity.
[0092] The audio features included in the negative samples are different from those included in the positive samples. The positive samples include the audio features of the audio signals of the body colliding with the ground in various fall states, which are collected by the feedback microphone. The negative samples include the audio features of various audio signals in the ear canal in non-fall states, which are collected by the feedback microphone. In the non-fall state, the audio signals in the ear canal can include audio features of ambient audio, audio of speech, and audio generated by other user interactions. These audio features are different from the audio features of the target collision included in the positive samples.
[0093] Each negative sample in the negative sample set is input into the initial neural network model, and the second label is used as the output of the initial neural network. The initial neural network model is trained to obtain a fall detection model. The number of negative samples in the negative sample set can be determined based on actual needs. A larger number of negative samples will improve the detection accuracy of the trained fall detection model.
[0094] The audio features of the non-target collision may include Mel spectrum feature parameters and Mel cepstral feature parameters, and the Mel spectrum coefficients and Mel cepstral coefficients may both be 40-dimensional coefficients.
[0095] In another embodiment, the second number is the number of peaks or valleys within the target frame length. The second number and the number of peaks or valleys in the negative sample that generate the feature to be identified are both determined within the same frame length, which can reduce variables and improve detection accuracy.
[0096] The initial neural network model is trained using positive and negative samples to obtain a fall detection model, which improves the detection capability of the fall detection model and makes the detection results more accurate.
[0097] In another embodiment, the fall detection method further includes:
[0098] When the detection result indicates that the user has fallen, a prompt message is sent to a preset device that has established a communication connection with the headset. The preset device can be a mobile phone, tablet computer, or a device held by a user who has a social relationship with the detected user. For example, if the user is an elderly person, the preset device can be the electronic device of the caregiver.
[0099] A prompt message is sent to a preset device that has established a communication connection with the headset to notify relevant personnel, thereby facilitating assistance to the user.
[0100] The prompt information can be a pop-up message, a sound prompt message or a short message.
[0101] In another embodiment, reference Figure 6 , is a schematic diagram of a fall detection device, wherein the device is applied to headphones, the headphones including a feedback microphone, and the device includes:
[0102] The ear canal audio signal detection module 1 is configured to collect audio signals in the ear canal through the feedback microphone to obtain an ear canal audio signal; wherein the ear canal audio signal includes: an audio signal generated by vibration generated by the user's body colliding with the ground when the headset is worn by the user and transmitted to the ear canal via bone conduction when the user falls;
[0103] An audio signal characteristic parameter acquisition module 2 is configured to extract characteristics of the ear canal audio signal to obtain audio signal characteristic parameters;
[0104] The periodic signal determination module 3 is configured to filter the ear canal audio signal according to a preset frequency range to obtain a periodic signal in the preset frequency range; wherein the periodic signal includes a peak value or a valley value of a waveform;
[0105] A feature generation module 4 for identifying is configured to generate a feature for identifying based on the audio signal characteristic parameter and the number of peak values or valley values;
[0106] The detection module 5 is configured to input the features to be identified into a fall detection model to obtain a detection result; wherein the detection result is at least used to indicate that the user has fallen.
[0107] In another embodiment, the apparatus further comprises:
[0108] a target frame length determination module, configured to determine a target frame length of at least one frame of the periodic signal according to a duration corresponding to a preset number of steps;
[0109] The quantity determination module is configured to determine the number of the peak values or the valley values included in the periodic signal of each frame within the target frame length; wherein each step corresponds to one peak value or one valley value.
[0110] In another embodiment, the fall detection model is obtained by pre-training an initial neural network model based on a fall information training sample set using a machine learning method.
[0111] In another embodiment, the training sample set includes a positive sample set, and the positive sample set includes a plurality of positive samples;
[0112] Each of the positive samples includes: an audio feature of the target collision, a first quantity, and a first label;
[0113] The audio feature of the target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the earphone is worn by the user and the user's body collides with the ground when falling;
[0114] The first number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the target collision according to the preset frequency range;
[0115] The first label is used to represent the audio feature of the target collision, and the first quantity corresponds to the output of the initial neural network model.
[0116] In another embodiment, the first number is the number of peaks or valleys within a target frame length.
[0117] In another embodiment, the audio features of the target collision included in different positive samples correspond to different postures, parts of the body that collide with the ground, and / or the number of collisions when the user falls.
[0118] In another embodiment, the training sample set further includes a negative sample set, and the negative sample set includes a plurality of negative samples;
[0119] Each of the negative samples includes: an audio feature of a non-target collision, a second quantity, and a second label;
[0120] The audio feature of the non-target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the headset is worn by the user, except for a collision between the user's body and the ground when falling;
[0121] The second number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the non-target collision according to the preset frequency range;
[0122] The second label is used to represent the audio feature of the non-target collision, and the second quantity corresponds to the output of the initial neural network model.
[0123] In another embodiment, the second number is the number of peaks or valleys within the target frame length.
[0124] In another embodiment, the apparatus further comprises:
[0125] The prompt information sending module is configured to send prompt information to a preset device; wherein a communication connection is established between the preset device and the headset.
[0126] In another embodiment, a headset is provided, comprising a housing and a controller, a feedback microphone, a feedforward microphone, and a speaker disposed on the housing;
[0127] The feedforward microphone is connected to the controller and is used to collect audio data outside the ear canal and send it to the controller;
[0128] The feedback microphone is connected to the controller and is used to collect audio data in the ear canal and send it to the controller;
[0129] The controller includes a memory and a processor. The memory stores executable computer instructions. The processor can call the computer instructions stored in the memory to execute the method described in any one of the above embodiments.
[0130] In another embodiment, a computer storage medium is provided, wherein the computer storage medium stores an executable program; after the executable program is executed by a processor, the method described in any one of the above embodiments can be implemented.
[0131] Figure 7 is a block diagram of an electronic device 800 according to an exemplary embodiment.
[0132] Reference Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0133] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0134] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0135] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0136] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0137] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0138] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0139] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0140] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0141] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0142] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0143] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow from the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0144] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A fall detection method, applied to a headset including a feedback microphone, comprising: The feedback microphone collects audio signals in the ear canal to obtain an ear canal audio signal; wherein the ear canal audio signal includes: an audio signal generated by vibration generated by the user's body colliding with the ground when the headset is worn by the user and is transmitted to the ear canal via bone conduction when the user falls; Extracting features of the ear canal audio signal to obtain audio signal feature parameters; Filtering the ear canal audio signal according to a preset frequency range to obtain a periodic signal within the preset frequency range; wherein the periodic signal includes a peak value or a valley value of a waveform; generating a feature to be identified according to the audio signal characteristic parameter and the number of the peak values or the valley values; The feature to be identified is input into a fall detection model to obtain a detection result; wherein the detection result is at least used to indicate that the user has fallen.
2. The method according to claim 1, wherein The method further comprises: Determining a target frame length of at least one frame of the periodic signal according to a duration corresponding to a preset number of steps; The number of the peak values or the valley values included in the periodic signal of each frame is determined within the target frame length; wherein each step corresponds to one peak value or one valley value.
3. The method according to claim 1 or 2, wherein: The fall detection model is obtained by pre-training an initial neural network model based on a fall information training sample set using a machine learning method.
4. The method according to claim 3, wherein: The training sample set includes a positive sample set, and the positive sample set includes a plurality of positive samples; Each of the positive samples includes: an audio feature of the target collision, a first quantity, and a first label; The audio feature of the target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the earphone is worn by the user and the user's body collides with the ground when falling; The first number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the target collision according to the preset frequency range; The first label is used to represent the audio feature of the target collision, and the first quantity corresponds to the output of the initial neural network model.
5. The method according to claim 4, wherein The first number is the number of peak values or valley values within the target frame length.
6. The method according to claim 4, wherein: The audio features of the target collision included in different positive samples correspond to different postures, parts of the body that collide with the ground, and / or the number of collisions when the user falls.
7. The method according to claim 3, wherein: The training sample set also includes a negative sample set, and the negative sample set includes multiple negative samples; Each of the negative samples includes: an audio feature of a non-target collision, a second quantity, and a second label; The audio feature of the non-target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the headset is worn by the user, except for a collision between the user's body and the ground when falling; The second number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the non-target collision according to the preset frequency range; The second label is used to represent the audio feature of the non-target collision, and the second quantity corresponds to the output of the initial neural network model.
8. The method according to claim 7, wherein: The second number is the number of peak values or valley values within the target frame length.
9. The method according to claim 1, wherein: The method further comprises: Sending a prompt message to a preset device; wherein a communication connection is established between the preset device and the headset.
10. A fall detection device, applied to a headset, the headset including a feedback microphone, the device comprising: an ear canal audio signal detection module configured to collect audio signals in the ear canal through the feedback microphone to obtain an ear canal audio signal; wherein the ear canal audio signal includes: an audio signal generated by vibration generated by the user's body colliding with the ground when the headset is worn by the user and transmitted to the ear canal via bone conduction when the user falls; an audio signal characteristic parameter acquisition module, configured to extract features of the ear canal audio signal to obtain audio signal characteristic parameters; a periodic signal determination module, configured to filter the ear canal audio signal according to a preset frequency range to obtain a periodic signal within the preset frequency range; wherein the periodic signal includes a peak value or a valley value of a waveform; a feature generation module to be identified, configured to generate a feature to be identified based on the audio signal characteristic parameter and the number of the peak values or the valley values; The detection module is configured to input the feature to be identified into a fall detection model to obtain a detection result; wherein the detection result is at least used to indicate that the user has fallen.
11. The device according to claim 10, wherein The device further comprises: a target frame length determination module, configured to determine a target frame length of at least one frame of the periodic signal according to a duration corresponding to a preset number of steps; The quantity determination module is configured to determine the number of the peak values or the valley values included in the periodic signal of each frame within the target frame length; wherein each step corresponds to one peak value or one valley value.
12. The device according to claim 10 or 11, wherein The fall detection model is obtained by pre-training an initial neural network model based on a fall information training sample set using a machine learning method.
13. The device according to claim 12, wherein The training sample set includes a positive sample set, and the positive sample set includes a plurality of positive samples; Each of the positive samples includes: an audio feature of the target collision, a first quantity, and a first label; The audio feature of the target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the earphone is worn by the user and the user's body collides with the ground when falling; The first number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the target collision according to the preset frequency range; The first label is used to represent the audio feature of the target collision, and the first quantity corresponds to the output of the initial neural network model.
14. The device according to claim 13, wherein The first number is the number of peak values or valley values within the target frame length.
15. The device according to claim 13, wherein The audio features of the target collision included in different positive samples correspond to different postures, parts of the body that collide with the ground, and / or the number of collisions when the user falls.
16. The device according to claim 12, wherein The training sample set also includes a negative sample set, and the negative sample set includes multiple negative samples; Each of the negative samples includes: an audio feature of a non-target collision, a second quantity, and a second label; The audio feature of the non-target collision is an audio feature obtained by collecting an audio signal in the ear canal through the feedback microphone when the headset is worn by the user, except for a collision between the user's body and the ground when falling; The second number is: the number of peaks or valleys in the waveform included in the periodic signal within the preset frequency range obtained after filtering the audio features of the non-target collision according to the preset frequency range; The second label is used to represent the audio feature of the non-target collision, and the second quantity corresponds to the output of the initial neural network model.
17. The device according to claim 16, wherein The second number is the number of peak values or valley values within the target frame length.
18. The device according to claim 10, wherein The device further comprises: The prompt information sending module is configured to send prompt information to a preset device; wherein a communication connection is established between the preset device and the headset.
19. An earphone comprising a housing and a controller, a feedback microphone, a feedforward microphone, and a speaker disposed on the housing; The feedforward microphone is connected to the controller and is used to collect audio data outside the ear canal and send it to the controller; The feedback microphone is connected to the controller and is used to collect audio data in the ear canal and send it to the controller; The controller includes a memory and a processor, wherein the memory stores executable computer instructions, and the processor can call the computer instructions stored in the memory to execute the method according to any one of claims 1 to 9.
20. A computer storage medium storing an executable program; after being executed by a processor, the executable program can implement the method provided in any one of claims 1 to 9.
Citation Information
Patent Citations
Fall-down monitoring system for vulnerable group
CN104065776A
Wearable-equipment-based fall-down warning method, wearable equipment and storage medium
CN112438726A