Multi-sleep behavior activity identification method and system
By collecting acoustic signals in sleep earbuds and applying dynamic threshold adaptive detection and segmentation algorithms, cumulative spectrum power classification and multi-dimensional feature extraction, the problem of high energy consumption and high cost of sleep earbud recognition methods in the prior art is solved, and low-power, low-cost and high-precision sleep behavior activity recognition is achieved.
Patent Information
- Application Number
- CN202510689701.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the sleep behavior activity recognition method based on sleep earplugs has problems of high energy consumption and high cost, and it is difficult to support the need for long-term continuous monitoring and identification.
The built-in microphone of the sleep earbud collects the acoustic signals of the ear canal body, and uses preprocessing, dynamic threshold adaptive detection and segmentation algorithm, cumulative spectrum power classification and multi-dimensional feature extraction methods to achieve accurate detection and recognition of various sleep behavior activities.
It realizes low-power, low-cost and high-precision sleep behavior activity recognition, breaks through the energy consumption and cost limitations of traditional methods, and supports long-term continuous monitoring and identification.
Smart Images

Figure CN120199283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of acoustic sensing technology, and particularly to a method and system for identifying multiple sleep behavior activities. Background Art
[0002] Human sleep is often accompanied by a series of behavior activities, such as body turning, limb movement, snoring, etc. These behavior activities may interrupt the sleep cycle and cause partial cerebral arousal, thus affecting sleep quality. Therefore, the detection of sleep behavior activities plays an important role in understanding and evaluating sleep health. Most of the existing work is based on multiple sensors on mobile phones or wrist-worn devices to identify sleep behavior activities. However, the sleep events that can be identified by these works are relatively limited and are not sufficient to support the multi-faceted evaluation of sleep health.
[0003] In recent years, as an auxiliary tool for isolating external noise and ensuring sleep quality, sleep earplugs have gradually become an indispensable part of people's daily lives. Due to the special nature of their wearing position, more and more sensors are integrated into sleep earplugs, providing new possibilities for the identification of sleep behavior activities. Most of the existing work deploys multiple modalities of sensors on sleep earplugs to detect human electroencephalogram signals and then identify the behavior activities that occur during sleep. However, such methods are often accompanied by high energy consumption and complex integration problems, which are not only costly but also difficult to support the actual needs of long-term continuous monitoring and identification. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and system for identifying multiple sleep behavior activities. By collecting the ear canal body-conducted acoustic signals through the built-in microphone of the sleep earplug, the accurate detection of the user's multiple sleep behavior activities is realized on the basis of low power consumption, breaking through the limitations of high energy consumption and high cost of the existing identification methods based on sleep earplugs, and providing a solution that combines low cost, low power consumption and high precision.
[0005] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a method for identifying multiple sleep behavior activities, including:
[0007] Preprocessing the collected dual-channel acoustic signals in the user's ear canal to obtain the processed dual-channel acoustic signals;
[0008] Based on the processed dual-channel acoustic signals, performing sleep behavior event detection and segmentation by an adaptive detection and segmentation algorithm based on a dynamic threshold to obtain a set of sleep events;
[0009] Based on the set of sleep events, classify sleep behavior events based on cumulative spectral power to obtain a set of sleep body activity events and a set of sleep acoustic activity events;
[0010] Based on the set of sleep body activity events, perform fine-grained identification of body activity events based on acoustic attenuation and cumulative spectral power;
[0011] Based on the set of sleep acoustic activity events, extract multi-dimensional features and use the extracted multi-dimensional features to perform fine-grained identification of acoustic activity events.
[0012] Optionally, the preprocessing includes:
[0013] Normalize the collected two-channel acoustic signals in the user's ear canal using the min-max normalization method, and subtract the average of all normalized signals from each normalized signal to obtain a preliminary preprocessed signal;
[0014] Replace the outliers in the preliminary preprocessed signal with local medians through a Hampel filter, and eliminate the heartbeat noise interference through a high-pass filter to obtain the processed two-channel acoustic signal.
[0015] Optionally, the sleep behavior event detection and segmentation based on the adaptive detection and segmentation algorithm with dynamic thresholds includes:
[0016] Perform frame segmentation on the processed two-channel acoustic signal through a sliding Hamming window to obtain a signal frame sequence;
[0017] Calculate the short-time energy STE of each frame in the signal frame sequence, and calculate the envelope of the short-time energy STE based on the root mean square function RMS;
[0018] Scale the envelope of the short-time energy STE to obtain a reconstructed envelope;
[0019] Add all the frames in the time period of the signal frame sequence to the non-event frame set , and initialize the event threshold and the set of sleep events ;
[0020] Traverse all the frames in the time period of the signal frame sequence and make the following judgments:
[0021] If and , then consider the th frame as a non-event, add it to the non-event frame set , and update the event threshold ;
[0022] Otherwise, consider the frame as an event and add it to the sleep event set ;
[0023] For each frame in the non-event frame set , if the time interval between consecutive frames is less than the preset value, merge the consecutive frames into one event. If the event is greater than or equal to the minimum duration, add the event to the sleep event set ;
[0024] Among them, represents the preset time, represents the end time of the signal frame sequence, represents the envelope value of the th frame, represents the feature set of the th frame, represents the mean value of the feature set th frame, represents the envelope reconstruction value of the
[0025] Optionally, scaling the envelope of the short-time energy STE is achieved through the following formula: ; Among them, represents the envelope value of the th frame, represents the scaling threshold, represents the adjustment factor, represents as the base to the power of the logarithmic function.
[0026] Optionally, the update formula for the event threshold is: ; Among them, represents the envelope value of the th frame, represents the standard deviation of the feature set ;
[0027] Optionally, the classification of sleep behavior events includes: traversing each event in the sleep event set, calculating the cumulative spectral power of the frequency band above its set frequency . If the cumulative spectral power of the event is less than the preset cumulative power threshold , then determine that the event is an acoustic activity event, otherwise determine that the event is a body activity event.
[0028] Optionally, the fine-grained recognition of the body activity events includes: Dividing the set of sleep body activity events into a set of micro-movement events and a set of macro-movement events based on the duration difference between micro-movements and macro-movements; Traverse each micro-movement event in the set of micro-movement events: divide the micro-movement event into multiple frames, measure the zero-crossing rate of the micro-movement event to obtain the acoustic attenuation coefficient; based on the acoustic attenuation coefficient, calculate the delay between the left and right channels of the micro-movement event using the cross-correlation function; input the mean and standard deviation of the zero-crossing rate, acoustic attenuation coefficient, and inter-channel delay into the classifier to obtain the output result of head rotation or body tremor;
[0029] Traverse each macro-movement event in the set of macro-movement events: calculate the cumulative spectral power of the macro-movement event at multiple preset quantiles, and input it into the classifier to obtain the output result of limb movement or body roll.
[0030] Optionally, the formula for calculating the inter-channel delay is:
[0031]
[0032] Where represents the inter-channel delay of the th frame, represents the maximum value function, represents the cross-correlation function, represents the right-channel signal of the th frame, represents the left-channel signal of the th frame.
[0033] Optionally, the multi-dimensional features include periodic features, harmonic structure features, energy distribution features, spectral statistical features, and formant features;
[0034] The fine-grained recognition of the acoustic activity events includes:
[0035] Traverse each acoustic activity event in the set of sleep acoustic activity events and perform the following steps: Calculate the RMS envelope of the signal in the time period of the acoustic activity event, and calculate the autocorrelation sequence of the RMS envelope using the autocorrelation function, and select the first peak position of the autocorrelation sequence as the periodic feature; where , and represent preset parameters, RMS represents the root mean square function, and the RMS envelope represents the envelope calculated based on the root mean square function; Calculate the cumulative spectral power within the preset frequency range in this acoustic activity event, and select the first five peaks in the cumulative spectral power as the harmonic feature frequency set. , and use the standard deviation and mean of the harmonic feature frequency set as the harmonic structure features; among them, represents the first peak in the cumulative spectral power, represents the second peak in the cumulative spectral power, represents the third peak in the cumulative spectral power, represents the fourth peak in the cumulative spectral power, represents the fifth peak in the cumulative spectral power;
[0036] Divide the frequency band within the set frequency range in this acoustic activity event into multiple non-overlapping sub-bands, and use the ratio of the energy of each sub-band to the total energy of the multiple non-overlapping sub-bands as the energy distribution feature, and use the skewness and kurtosis of the frequency components of each sub-band as the spectral statistical features;
[0037] Extract the Mel-frequency cepstral coefficients of this acoustic activity event as the formant features;
[0038] Input the periodic features, harmonic structure features, energy distribution features, spectral statistical features and formant features into the classifier, and the output result is one of grinding teeth, snoring, coughing, swallowing and talking in sleep.
[0039] In a second aspect, the present invention provides a multi-sleep behavior activity recognition system, including:
[0040] A signal acquisition and processing module, configured to: preprocess the acquired dual-channel acoustic signal in the user's ear canal to obtain the processed dual-channel acoustic signal;
[0041] A sleep behavior event detection and segmentation module, configured to: based on the processed dual-channel acoustic signal, perform sleep behavior event detection and segmentation based on an adaptive detection and segmentation algorithm with a dynamic threshold to obtain a set of sleep events;
[0042] A sleep behavior event classification module, configured to: based on the set of sleep events, classify sleep behavior events based on cumulative spectral power to obtain a set of sleep body activity events and a set of sleep acoustic activity events;
[0043] A body activity event fine-grained recognition module, configured to: based on the set of sleep body activity events, perform fine-grained recognition of body activity events based on acoustic attenuation and cumulative spectral power;
[0044] An acoustic activity event fine-grained recognition module, configured to: extract multi-dimensional features based on the set of sleep acoustic activity events, and perform fine-grained recognition of acoustic activity events by using the extracted multi-dimensional features.
[0045] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: By performing preprocessing on the captured signals, detecting and segmenting sleep behavior events based on an adaptive detection and segmentation algorithm with a dynamic threshold, classifying sleep behavior events based on the cumulative spectral power, and finally realizing fine-grained recognition of various sleep behavior activities based on acoustic attenuation, cumulative spectral power, and multi-dimensional signal features; compared with other sleep activity recognition methods based on sleep earplugs, it can recognize various sleep behavior activity events (8 types in total, including head rotation, body tremor, limb movement, rolling over, teeth grinding, snoring, coughing, swallowing, and talking in sleep) based on a single-modal sensor with high precision, providing sufficient basis for sleep health assessment; solving the problems of high energy consumption and high cost in traditional solutions, and being able to continuously monitor and recognize sleep behavior activities for a long time; realizing sleep behavior activity recognition only by relying on a microphone, with simple integration and low hardware cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of a multi-sleep behavior activity recognition method provided according to an embodiment of the present invention;
[0047] Figure 2 It is a flowchart of preprocessing provided according to an embodiment of the present invention;
[0048] Figure 3 It is a flowchart of an adaptive detection and segmentation algorithm with a dynamic threshold provided according to an embodiment of the present invention;
[0049] Figure 4 It is a flowchart of fine-grained recognition of body activity events provided according to an embodiment of the present invention;
[0050] Figure 5 It is a flowchart of fine-grained recognition of acoustic activity events provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of this patent are detailed descriptions of the technical solution of this patent, rather than limitations on the technical solution of this patent. Without conflict, the technical features in the embodiments of this patent and the embodiments can be combined with each other.
[0052] It should be noted that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.
[0053] Embodiment 1:
[0054] The embodiment of the present invention discloses a method for identifying multiple sleep behavior activities. Referring to Figure 1 as shown, the specific steps are as follows:
[0055] S1. Preprocess the collected dual-channel acoustic signals in the user's ear canal to obtain the processed dual-channel acoustic signals;
[0056] S2. Based on the processed dual-channel acoustic signals, perform sleep behavior event detection and segmentation using an adaptive detection and segmentation algorithm based on a dynamic threshold to obtain a set of sleep events;
[0057] S3. Based on the set of sleep events, classify sleep behavior events based on the cumulative spectral power to obtain a set of sleep body activity events and a set of sleep acoustic activity events;
[0058] S4. Based on the set of sleep body activity events, perform fine-grained recognition of body activity events based on acoustic attenuation and cumulative spectral power;
[0059] S5. Based on the set of sleep acoustic activity events, extract multi-dimensional features, and use the extracted multi-dimensional features to perform fine-grained recognition of acoustic activity events.
[0060] Specifically, in step S1, referring to Figure 2 as shown, since the user wears the sleep earplugs, collect the dual-channel acoustic signals in the user's ear canal through the built-in microphone of the sleep earplugs; set a 60s buffer on the audio signal stream, and after it is full, take out the signal, perform normalization processing through the min-max normalization method, and subtract the average value of all normalized signals from the normalized signal to eliminate individual differences and the DC offset introduced by the microphone hardware; replace the outliers caused by the microphone circuit with the local median through the Hampel filter, and eliminate the heartbeat noise interference through a high-pass filter with a cut-off frequency of 100Hz.
[0061] In step S2, due to the high diversity of sleep behavior activities, a fixed threshold cannot effectively detect and segment activity implementations, so an adaptive method with a dynamic threshold is used to achieve event segmentation; referring to Figure 3 as shown, perform sleep behavior event detection and segmentation using an adaptive detection and segmentation algorithm based on a dynamic threshold, including:
[0062] The processed two-channel acoustic signal is framed by using a sliding Hamming window with a length of 0.2 seconds and a step size of 0.1 seconds to obtain a signal frame sequence;
[0063] Calculate the short-time energy STE of each frame in the signal frame sequence, and calculate the envelope of the short-time energy STE based on the root mean square function RMS;
[0064] The envelope of the short-time energy STE is scaled by the following formula to obtain a reconstructed envelope;
[0065]
[0066] where, represents the scaling threshold, represents the th frame's envelope reconstruction value, represents the th frame's envelope value, represents the adjustment factor, which is set to 3 in this embodiment, represents as the base the logarithmic function of the power;
[0067] All frames in the time period of the signal frame sequence are added to the non-event frame set , and the event threshold and the sleep event set are initialized; in this embodiment, the initial event threshold , represents the preset time, which is set in this embodiment, with the unit of minutes, represents the end time of the signal frame sequence;
[0068] Traverse all frames in the time period of the signal frame sequence and make the following judgments:
[0069] If and (using the value of the previous frame to avoid instantaneous noise interference), then the th frame is considered a non-event and is added to the non-event frame set , and the event threshold is updated by the following formula:
[0070]
[0071] where, represents the th frame's envelope value, represents the th frame's feature set, Represents the standard deviation of, represents the envelope value of the frame; represents the
[0072] mean value of; Otherwise, consider the frame as an event and add it to the sleep event set
[0073] ; Considering the possibility of incorrect segmentation, for each frame in the non-event frame set if the time interval between consecutive frames is less than the preset value, merge the consecutive frames into one event. If the event is greater than or equal to the minimum duration, add the event to the sleep event set
[0074] In step S3, the classification of the sleep behavior events includes:
[0075] Traverse each event in the sleep event set and calculate the cumulative spectral power of the frequency band above 200 Hz of it. If the cumulative spectral power of the event is less than the preset cumulative power threshold , then determine that the event is an acoustic activity event; otherwise, determine that the event is a body activity event.
[0076] In step S4, the body activities during sleep are divided into micro-movements (head turning, body trembling) and macro-movements (body rolling, limb movement). The two can be distinguished by the duration of the event, and the specific events that occur need to be distinguished using a specific algorithm; refer to Figure 4 shown, the data processing flow for the fine-grained recognition of the body activity events includes:
[0077] Based on the difference in the duration of micro-movements and macro-movements, divide the sleep body activity event set into a micro-movement event set and a macro-movement event set; in this embodiment, the events with a duration of 5 - 10 s are divided into macro-movements, and the events with a duration of 1 - 3 s are divided into micro-movements;
[0078] Traverse each micro-movement event in the micro-movement event set: divide the micro-movement event into multiple frames, measure the zero-crossing rate of the micro-movement event to obtain the acoustic attenuation coefficient; based on the acoustic attenuation coefficient, calculate the delay between the left and right channels of the micro-movement event using the cross-correlation function; input the mean and standard deviation of the zero-crossing rate, acoustic attenuation coefficient, and delay between the left and right channels into the classifier to obtain the output result of head turning or body trembling;
[0079] Traverse each macro motion event in the set of macro motion events: Calculate the cumulative spectral power of the macro motion event at the 0%, 25%, 50%, 75%, and 100% quantiles, and input it into the classifier to obtain the output result of limb movement or body roll.
[0080] The formula for calculating the delay between the left and right channels is:
[0081]
[0082] Where represents the delay between the left and right channels of the th frame, represents the maximum value function, represents the cross-correlation function, represents the th frame of the right channel signal, represents the th frame of the left channel signal.
[0083] In step S5, since snoring and teeth grinding have more obvious periodic patterns compared to coughing, swallowing, and talking in sleep, they can serve as identification bases; due to the differences in bandwidth and harmonic patterns among teeth grinding, snoring, coughing, swallowing, and talking in sleep, event discrimination can be performed by extracting harmonic features; the multi-dimensional features include periodic features, harmonic structure features, energy distribution features, spectral statistical features, and formant features;
[0084] Refer to Figure 5 shown, the fine-grained identification of the acoustic activity event includes:
[0085] Traverse each acoustic activity event in the set of sleep acoustic activity events and perform the following steps:
[0086] Calculate the RMS envelope of the signal in the time period of this acoustic activity event, and calculate the autocorrelation sequence of the RMS envelope using the autocorrelation function, and select the first peak position of the autocorrelation sequence as the periodic feature; where , and represent preset parameters, RMS represents the root mean square function, and the RMS envelope represents the envelope calculated based on the root mean square function;
[0087] Calculate the cumulative spectral power within Hz of this acoustic activity event, and select the first five peaks in the cumulative spectral power as the harmonic feature frequency set , and use the standard deviation and mean value of as the harmonic structure feature; where represents the first peak in the cumulative spectral power, represents the second peak in the cumulative spectral power, represents the third peak in the cumulative spectral power, represents the fourth peak in the cumulative spectral power, represents the fifth peak in the cumulative spectral power;
[0088] In the acoustic activity event, the frequency band of [[Hz]] is divided into 4 non - overlapping sub - bands, and the ratio of the energy of each sub - band to the total energy of the 4 non - overlapping sub - bands is used as the energy distribution feature, and the skewness and kurtosis of the frequency components of each sub - band are used as the spectral statistical features;
[0089] Extract the Mel - Frequency Cepstral Coefficients of the acoustic activity event as the formant features;
[0090] Input the periodicity feature, harmonic structure feature, energy distribution feature, spectral statistical feature and formant feature into the classifier, and the output result is one of grinding teeth, snoring, coughing, swallowing and talking in sleep.
[0091] In summary, the multi - sleep behavior activity recognition method proposed in this embodiment uses the built - in microphone of the sleep earplug to capture the acoustic signals caused by the sleep behavior in the ear, pre - processes the captured signals, and then uses the method based on dynamic threshold to segment and classify the sleep events, and finally realizes the recognition of multiple sleep behavior activities based on the multi - dimensional features of the signals; in this process, the following steps are included: (1) Detect and segment sleep behavior events based on the adaptive detection and segmentation algorithm of dynamic threshold; (2) Sleep behavior event classification algorithm based on cumulative spectral power; (3) Fine - grained recognition of body activity events based on acoustic attenuation and cumulative spectral power; (4) Fine - grained recognition of acoustic activity events based on multi - dimensional features; compared with other sleep activity recognition methods based on sleep earplugs, it can recognize multiple sleep behavior activities using only a single modality, solves the problems of high energy consumption and high cost, and shows higher recognition accuracy.
[0092] Embodiment 2:
[0093] Based on the same inventive concept as Embodiment 1, the present invention embodiment discloses a multi - sleep behavior activity recognition system. The multi - sleep behavior activity recognition system is applicable to the multi - sleep behavior activity recognition method described in Embodiment 1, and includes:
[0094] A signal acquisition and processing module, which is used for: pre - processing the acquired dual - channel acoustic signals in the user's ear canal to obtain the processed dual - channel acoustic signals;
[0095] A sleep behavior event detection and segmentation module, configured to: perform sleep behavior event detection and segmentation on the processed two-channel acoustic signal based on an adaptive detection and segmentation algorithm with a dynamic threshold, to obtain a set of sleep events;
[0096] A sleep behavior event classification module, configured to: classify sleep behavior events based on the cumulative spectral power according to the set of sleep events, to obtain a set of sleep body activity events and a set of sleep acoustic activity events;
[0097] A body activity event fine-grained recognition module, configured to: perform fine-grained recognition of body activity events based on acoustic attenuation and cumulative spectral power according to the set of sleep body activity events;
[0098] An acoustic activity event fine-grained recognition module, configured to: perform multi-dimensional feature extraction according to the set of sleep acoustic activity events, and perform fine-grained recognition of acoustic activity events by using the extracted multi-dimensional features.
[0099] The specific functional implementation of each of the above modules refers to the relevant content in the method of Embodiment 1, and will not be elaborated here.
[0100] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows Figure 1 or a combination of one or more flows and / or blocks
[0102] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the flowFigure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or boxes Figure 1 one box or multiple boxes.
[0104] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these fall within the protection scope of the present invention.
Claims
1. A method for identifying multiple sleep behavior activities, characterized in that Including: Preprocess the collected binaural acoustic signals in the user's ear canal to obtain processed binaural acoustic signals; Based on the processed binaural acoustic signals, perform sleep behavior event detection and segmentation using an adaptive detection and segmentation algorithm based on a dynamic threshold to obtain a set of sleep events; Based on the set of sleep events, classify sleep behavior events based on cumulative spectral power to obtain a set of sleep body activity events and a set of sleep acoustic activity events; Based on the set of sleep body activity events, perform fine-grained recognition of body activity events based on acoustic attenuation and cumulative spectral power; Based on the set of sleep acoustic activity events, extract multi-dimensional features and use the extracted multi-dimensional features to perform fine-grained recognition of acoustic activity events.
2. The multi-sleep behavior activity recognition method according to claim 1, wherein The preprocessing includes: Normalize the collected binaural acoustic signals in the user's ear canal using the min-max normalization method, and subtract the average value of all normalized signals from each normalized signal to obtain a preliminary preprocessed signal; Replace the outliers in the preliminary preprocessed signal with local medians through a Hampel filter, and eliminate heartbeat noise interference through a high-pass filter to obtain processed binaural acoustic signals.
3. The multi-sleep behavior activity recognition method according to claim 1, characterized in that The sleep behavior event detection and segmentation using the adaptive detection and segmentation algorithm based on a dynamic threshold includes: Perform frame segmentation on the processed binaural acoustic signals through a sliding Hamming window to obtain a sequence of signal frames; Calculate the short-time energy STE of each frame in the signal frame sequence, and calculate the envelope of the short-time energy STE based on the root mean square function RMS; Scale the envelope of the short-time energy STE to obtain a reconstructed envelope; Add all frames in the time period of the signal frame sequence to the non-event frame set , and initialize the event threshold and the sleep event set ; Traverse all frames in the time period of the signal frame sequence and make the following judgments: If and , then the th frame is considered a non-event and added to the non-event frame set , and the event threshold is updated; Otherwise, consider the frame as an event and add it to the sleep event set ; For each frame in the non-event frame set if the time interval between consecutive frames is less than a preset value, the consecutive frames are merged into one event, and if the event is greater than or equal to the minimum duration, the event is added to the sleep event set ; Among them, represents a preset time, represents the end time of the signal frame sequence, represents the envelope value of the th frame, represents the feature set of the th frame, represents the envelope reconstruction value of the 4. The multi-sleep behavior activity recognition method according to claim 3, wherein, The scaling of the envelope of the short-time energy STE is achieved through the following formula: ; Among them, represents the envelope value of the nth frame, represents the scaling threshold, represents the adjustment factor, is the base and is the logarithmic function with as the exponent.
5. The multi-sleep behavior activity recognition method according to claim 3, wherein The event threshold The update formula is as follows: ; Among them, represents the envelope value of the nth frame, represents the standard deviation of the feature set .
6. The multi-sleep behavior activity recognition method according to claim 1, characterized in that, The sleep behavior event classification includes: Traverse each event in the set of sleep events and calculate the cumulative spectral power in the frequency band above its set frequency , if the cumulative spectral power of this event is less than the preset cumulative power threshold , then determine that this event is an acoustic activity event, otherwise determine that this event is a body activity event.
7. The multi-sleep behavior activity recognition method according to claim 1, characterized in that The fine-grained recognition of the body activity events includes: Based on the duration difference between micro-movements and macro-movements, divide the set of sleep body activity events into a set of micro-movement events and a set of macro-movement events; Traverse each micro-movement event in the set of micro-movement events: divide the micro-movement event into multiple frames, measure the zero-crossing rate of the micro-movement event to obtain an acoustic attenuation coefficient; based on the acoustic attenuation coefficient, calculate the delay between the left and right channels of the micro-movement event using the cross-correlation function; input the mean and standard deviation of the zero-crossing rate, acoustic attenuation coefficient, and inter-channel delay into a classifier to obtain an output result of head rotation or body tremor; Traverse each macro-movement event in the set of macro-movement events: calculate the cumulative spectral power of the macro-movement event at multiple preset quantiles and input it into a classifier to obtain an output result of limb movement or body roll.
8. The multi-sleep behavior activity recognition method according to claim 7, characterized in that, The formula for calculating the inter-channel delay is: ; Among them, represents the delay between the left and right channels of the th frame, represents the maximum value function, represents the cross-correlation function, represents the right channel signal of the th frame, represents the left channel signal of the 9. The multi-sleep behavior activity recognition method according to claim 1, characterized in that, The multi-dimensional features include periodic features, harmonic structure features, energy distribution features, spectral statistical features, and formant features; The fine-grained recognition of the acoustic activity events includes: Traverse each acoustic activity event in the set of sleep acoustic activity events and perform the following steps: Calculate the RMS envelope of the signal in the time period of the acoustic activity event, calculate the autocorrelation sequence of the RMS envelope using the autocorrelation function, and select the position of the first peak of the autocorrelation sequence as the periodic feature; where , , and represent preset parameters, RMS represents the root mean square function, and the RMS envelope represents the envelope calculated based on the root mean square function; Calculate the cumulative spectral power within a preset frequency range in this acoustic activity event, and select the first five peaks in the cumulative spectral power as the harmonic feature frequency set , and use the standard deviation and the mean value of the harmonic feature frequency set as the harmonic structure features; where represents the first peak in the cumulative spectral power, represents the second peak in the cumulative spectral power, represents the third peak in the cumulative spectral power, represents the fourth peak in the cumulative spectral power, represents the fifth peak in the cumulative spectral power; Divide the frequency bands within the set frequency range in the acoustic activity event into multiple non-overlapping sub-bands, and use the ratio of the energy of each sub-band to the total energy of the multiple non-overlapping sub-bands as the energy distribution feature, and use the skewness and kurtosis of the frequency components of each sub-band as the spectral statistical features; Extract the Mel-frequency cepstral coefficients of the acoustic activity event as the formant features; Input the periodicity feature, harmonic structure feature, energy distribution feature, spectral statistical feature and formant feature into a classifier, and obtain an output result of one of grinding teeth, snoring, coughing, swallowing and talking in sleep.
10. A multi-sleep behavior activity recognition system, characterized in that, Comprising: A signal acquisition and processing module, configured to: preprocess the acquired dual-channel acoustic signal in the user's ear canal to obtain a processed dual-channel acoustic signal; A sleep behavior event detection and segmentation module, configured to: perform sleep behavior event detection and segmentation on the basis of the processed dual-channel acoustic signal by using an adaptive detection and segmentation algorithm based on a dynamic threshold, so as to obtain a sleep event set; A sleep behavior event classification module, configured to: classify sleep behavior events on the basis of the cumulative spectral power according to the sleep event set, so as to obtain a sleep body activity event set and a sleep acoustic activity event set; A body activity event fine-grained recognition module, configured to: perform fine-grained recognition of body activity events on the basis of acoustic attenuation and cumulative spectral power according to the sleep body activity event set; An acoustic activity event fine-grained recognition module, configured to: perform multi-dimensional feature extraction according to the sleep acoustic activity event set, and perform fine-grained recognition of acoustic activity events by using the extracted multi-dimensional features.
Citation Information
Patent Citations
Sleep sound event identification method and device, storage medium and electronic equipment
CN115995237A
Sleep event detection method, system and device for user subject recognition and medium
CN118709038A
User identity authentication method and system based on breath sound
CN120030518A
Method, computing device and computer program for analyzing sleep state of user through sound information
JP2024048399A
Apparatus, system, and method for detecting physiological movement from audio and multimodal signals
US20210275056A1
Cited By
Multi-modal sensing fusion swallowing rehabilitation evaluation system and method
CN120899193A