Snore extraction method, storage medium and earphone
By acquiring ambient audio, posture, and heart rate signals through headphones, and extracting and fusing features for snoring recognition, the problem of inaccurate snoring extraction under multiple snoring sources is solved, and the accuracy of snoring extraction and sleep quality analysis is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGXI RUISHENG ELECTRONIC CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-05
AI Technical Summary
Existing headphones struggle to accurately extract target snoring audio in scenarios with multiple snoring sources, leading to a decrease in the accuracy of sleep quality analysis.
By acquiring environmental audio signals, posture signals, and heart rate signals, noise features, correlation features, and heart rate features are extracted, and after fusion, a snoring recognition model is used to perform frame-level probability analysis to accurately extract the target snoring audio signal.
This improves the accuracy of target snoring audio extraction, thereby improving the accuracy of sleep quality analysis.
Smart Images

Figure CN121983085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone technology, and in particular to a method for extracting snoring sounds, a storage medium, and headphones. Background Technology
[0002] With the continuous development of electronic technology, users have increasingly higher demands for headphone functionality. Existing headphones have incorporated health monitoring functions, such as heart rate monitoring and sleep quality analysis. Headphones equipped with sleep quality analysis typically monitor snoring during sleep. However, in scenarios involving multiple snoring sources, current headphones cannot accurately extract the target snoring audio from the collected ambient audio signal, leading to a decrease in the accuracy of subsequent sleep quality analysis.
[0003] Therefore, it is necessary to provide a method for snoring extraction, a storage medium, and headphones to solve the above-mentioned problems. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the present invention provides a method, storage medium and earphone for snoring extraction, which can effectively solve the problem that the earphone has low accuracy in extracting target snoring audio in multi-snoring sound source scenarios, thus affecting the accuracy of sleep quality analysis.
[0005] To achieve the above objectives, a first aspect of the present invention provides a method for snoring extraction, the steps of which include:
[0006] Acquire stored environmental audio signals, attitude signals, and heart rate signals;
[0007] Extract noise features, correlation features, and heart rate features from environmental audio signals, posture signals, and heart rate signals;
[0008] By fusing noise features, correlation features, and heart rate features, a fused feature sequence is obtained.
[0009] Based on the fused feature sequence and the preset snoring recognition model, several corresponding frame-level probabilities are obtained;
[0010] Based on frame-level probability and stored environmental audio signals, the audio signal of the target snoring is extracted.
[0011] In one implementation, the steps of extracting noise features, correlation features, and heart rate features from environmental audio signals, posture signals, and heart rate signals include:
[0012] Vibration signals are obtained from attitude signals;
[0013] Based on timestamps, environmental audio signals, vibration signals, and heart rate signals are synchronized, aligned, and framed.
[0014] Noise features and heart rate features were extracted from environmental audio signals and heart rate signals, respectively.
[0015] Correlation characteristics are obtained based on environmental audio signals and vibration signals.
[0016] In one embodiment, the step of obtaining correlation characteristics based on ambient audio signals and vibration signals includes:
[0017] Based on the environmental audio signal and vibration signal of each frame, the corresponding audio self-power spectrum, vibration self-power spectrum and mutual power spectrum are obtained.
[0018] The average correlation value is obtained based on the audio self-power spectrum, vibration self-power spectrum and mutual power spectrum of each frame, and is used as a frame-level sub-feature.
[0019] Based on temporal sequence, frame-level sub-features are arranged to obtain correlation features.
[0020] In one implementation, the step of fusing noise features, correlation features, and heart rate features to obtain a fused feature sequence includes:
[0021] The target features are obtained by fusing noise features, correlation features, and heart rate features from the same frame.
[0022] Based on the time sequence, the target features are arranged to obtain the fused feature sequence.
[0023] In one implementation, the step of obtaining several corresponding frame-level probabilities based on the fused feature sequence and a preset snoring recognition model includes:
[0024] The fused feature sequence is input into a preset snoring recognition model;
[0025] Based on the preset recognition logic, the fused feature sequence is analyzed and the confidence value corresponding to each frame is output;
[0026] Map the confidence values and output the corresponding frame-level probabilities.
[0027] In one implementation, the step of extracting the audio signal of the target snoring sound based on frame-level probability and stored ambient audio signals includes:
[0028] Based on a preset threshold, determine whether the frame-level probability triggers audio extraction.
[0029] If so, extract the corresponding segment from the ambient audio signal to obtain the marked segment;
[0030] Based on a preset duration threshold, the marked segments are filtered to obtain the target segments;
[0031] Based on the time sequence, the target segments are spliced together to obtain the audio signal of the target snoring sound.
[0032] In one implementation, the determination threshold includes a first threshold and a second threshold; the step of determining whether the frame-level probability triggers audio extraction based on the preset determination threshold includes:
[0033] Determine whether the frame-level probability is less than the first threshold;
[0034] If not, then start audio extraction from the corresponding frame in the ambient audio signal;
[0035] Determine whether the frame-level probability is greater than the second threshold;
[0036] If not, then stop audio extraction from the corresponding frame in the ambient audio signal.
[0037] In one implementation, the step of filtering marked segments based on a preset duration threshold to obtain the target segment includes:
[0038] Get the duration of the tagged segment;
[0039] Determine whether the duration is less than a preset duration threshold;
[0040] If not, then the marked fragment is defined as the target fragment.
[0041] A second aspect of the present invention provides a computer-readable storage medium comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the snoring extraction method described above.
[0042] A third aspect of the present invention provides an earphone including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the snoring extraction method described above.
[0043] The beneficial effects of this invention are as follows: by utilizing multimodal integration and frame-level analysis of noise features, correlation features and heart rate features, it is possible to identify and extract target snoring sounds in complex snoring scenarios, thereby improving the accuracy of target snoring sound audio extraction and thus improving the accuracy of subsequent sleep quality analysis. Attached Figure Description
[0044] Figure 1 This is a schematic flowchart of the snoring extraction method disclosed in an embodiment of the present invention.
[0045] Figure 2 This is a schematic diagram of the module structure of the earphone disclosed in an embodiment of the present invention. Detailed Implementation
[0046] In this invention, the terms "set up," "equipped with," and "connected" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, elements, or components. Those skilled in the art can understand the specific meaning of these terms in this invention according to the specific circumstances.
[0047] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0048] Furthermore, in addition to indicating direction or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in certain situations to indicate a dependency or connection. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0050] The following is the content of the first aspect of the present invention:
[0051] Please refer to Figure 1 In this embodiment, the steps of the snoring extraction method include:
[0052] S1. Acquire the stored environmental audio signals, attitude signals, and heart rate signals;
[0053] S2. Extract noise features, correlation features, and heart rate features from environmental audio signals, posture signals, and heart rate signals.
[0054] The earphones contain several components, including but not limited to a microphone array, a posture sensor, and a heart rate sensor. The microphone array is used to acquire ambient audio signals of the current environment in which the earphones are located. The posture sensor is used to acquire posture signals corresponding to the user's current posture. The heart rate sensor is used to acquire heart rate signals of the user in the current state.
[0055] Noise features are the characteristic information of noise in the current environment extracted from the current ambient audio signal. Correlation features are the degree of synchronization between snoring and the user's vibration in terms of frequency and timing. Heart rate features are the characteristic information of the user's heart rate in the current state.
[0056] During sleep tracking using headphones, when snoring detection is triggered, the microphone array, posture sensor, and heart rate sensor inside the headphones collect and store the ambient audio signal, the user's current posture signal, and heart rate signal, respectively. This stored information is then retrieved before sleep quality analysis.
[0057] Before conducting sleep quality analysis, it is necessary to extract the target snoring audio from the environmental audio signal. Specifically, this involves acquiring and storing the environmental audio signal, posture signal, and heart rate signal while the user is snoring. After obtaining the environmental audio signal, posture signal, and heart rate signal, these signals are preprocessed and transformed. Then, noise features and heart rate features are extracted from the preprocessed environmental audio signal and heart rate signal, respectively. Finally, correlation features are obtained using the environmental audio signal and posture signal.
[0058] S3. By fusing noise features, correlation features, and heart rate features, a fused feature sequence is obtained.
[0059] S4. Based on the fused feature sequence and the preset snoring recognition model, obtain several corresponding frame-level probabilities;
[0060] S5. Extract the audio signal of the target snoring sound based on the frame-level probability and the stored environmental audio signal.
[0061] The fused feature sequence is the sequence information obtained by fusing multimodal feature data, which can be used for target snoring identification. The snoring identification model is a pre-set model used to analyze whether the snoring is the target snoring; this model can be a Transformer model or a Bi-LSTM model, which is not limited here. The frame-level probability is the likelihood that the corresponding frame in the fused feature sequence is the target snoring.
[0062] After obtaining the noise features, correlation features, and heart rate features, frame-level truncation can be performed on the noise features, correlation features, and heart rate features of the same frame. Then, the noise features, correlation features, and heart rate features of the same frame can be fused to obtain the target features. Then, based on the time order, the target features are arranged to obtain the required fused feature sequence.
[0063] After obtaining the fused feature sequence, it is input into a pre-defined snoring recognition model. This model uses noise features as the primary feature, correlation features as the distinguishing criterion, and heart rate features as an auxiliary verification mechanism to perform frame-by-frame recognition and analysis of the fused feature sequence, outputting the frame-level probability for each frame. Specifically, when the snoring recognition model performs frame-by-frame recognition and analysis of the fused feature sequence, once it determines that the feature information in a certain frame contains snoring features, it uses the correlation features contained in the feature information of the same frame to analyze the degree of correlation between the snoring features in that frame and the vibration features of the user's head. After distinguishing the degree of correlation between the snoring features and the vibration features of the user's head, it uses heart rate feature information from the same frame for auxiliary verification, and then outputs a comprehensive result.
[0064] In essence, after determining that the feature information of the same frame contains snoring features, the correlation features in the same frame are used to analyze the relationship between the snoring feature and the user's head vibration, thereby outputting the frame-level probability corresponding to each frame. After obtaining the frame-level probabilities, each frame-level probability is paired with the previously stored ambient audio signal for the same frame. Then, based on the specific situation of each frame-level probability, the target snoring is identified and judged for each frame of the ambient audio signal, thereby extracting the relatively pure audio signal of the target snoring.
[0065] Understandably, by utilizing multimodal integration and frame-level analysis of noise features, correlation features, and heart rate features, it is possible to identify and extract target snoring sounds in complex snoring scenarios, thereby improving the accuracy of target snoring sound audio extraction and consequently improving the accuracy of subsequent sleep quality analysis.
[0066] Furthermore, in one embodiment, step S2, which extracts noise features, correlation features, and heart rate features from the environmental audio signal, posture signal, and heart rate signal, includes:
[0067] S21. Obtain vibration signals from attitude signals;
[0068] S22. Based on timestamps, synchronously align and frame environmental audio signals, vibration signals, and heart rate signals;
[0069] S23. Extract noise features and heart rate features from environmental audio signals and heart rate signals, respectively;
[0070] S24. Based on the environmental audio signal and vibration signal, obtain the correlation characteristics.
[0071] The vibration signal is a high-frequency, slight vibration of the user's head recorded by the posture detection sensor inside the earphone when the user snores. The timestamp is the start and end time of the event recorded by the timing hardware inside the earphone.
[0072] When a target user snores, it is usually accompanied by slight high-frequency vibrations in the target user's head, throat, and nose. These vibration signals are recorded by the posture detection sensor inside the headphones, and there is a high correlation between these vibration signals and the target user's snoring audio signal. That is, in a scenario with multiple snoring sources, when the target user snores, the microphone array inside the headphones will collect the snoring signal, and at the same time, the posture detection sensor inside the headphones will simultaneously record the corresponding vibration signal; while when other people snore, only the snoring signal is collected. Therefore, by analyzing the correlation between the snoring signal and the vibration signal, the target snoring sound of the target user can be determined even in the case of multiple snoring sources.
[0073] After acquiring the stored ambient audio signal, posture signal, and heart rate signal, preprocessing is performed on each signal, and the required vibration signal is extracted from the posture signal. After obtaining the vibration signal, the ambient audio signal, vibration signal, and heart rate signal are synchronized and aligned to a unified timeline according to the timestamp recorded by the timing hardware in the headset. Then, the ambient audio signal, vibration signal, and heart rate signal are segmented at a preset frame level to obtain several ambient audio signal segments, vibration signal segments, and heart rate signal segments.
[0074] The framed ambient audio signal is converted into an ambient audio spectrum, and the required noise features are extracted from the ambient audio spectrum. Specifically, the information that can be extracted from the ambient audio spectrum and used to constitute noise features includes, but is not limited to, spectral centroid, short-time energy, and spectral roll-off point. Then, the required heart rate features are extracted from the framed heart rate signal. Specifically, the information that can be extracted from the heart rate signal and used to constitute heart rate features includes, but is not limited to, instantaneous heart rate and heart rate variability.
[0075] Correlation characteristics refer to the degree of synchronization between snoring and the user's vibration in terms of frequency and timing. Therefore, it is necessary to utilize ambient audio signals and vibration signals to obtain correlation characteristics. In a preferred embodiment, step S24, which obtains correlation characteristics based on ambient audio signals and vibration signals, includes:
[0076] S241. Based on the environmental audio signal and vibration signal of each frame, obtain the corresponding audio self-power spectrum, vibration self-power spectrum and mutual power spectrum.
[0077] S242. Based on the audio self-power spectrum, vibration self-power spectrum and mutual power spectrum of each frame, obtain the average correlation value as a frame-level sub-feature.
[0078] S243. Based on the temporal sequence, arrange the frame-level sub-features to obtain the correlation features.
[0079] Among them, the audio self-power spectrum is the power distribution of the environmental audio signal itself at different frequencies. The vibration self-power spectrum is the power distribution of the vibration signal itself at different frequencies. The mutual power spectrum is the relationship between the environmental audio signal and the vibration signal at different frequencies.
[0080] Specifically, after obtaining the environmental audio and vibration signals for each frame, windowing and Fast Fourier Transform are applied to each frame's environmental audio and vibration signals to obtain the corresponding complex spectra. Then, based on the power spectrum calculation formula and the corresponding complex spectra, the corresponding audio self-power spectrum, vibration self-power spectrum, and mutual power spectrum can be calculated.
[0081] After obtaining the audio self-power spectrum, vibration self-power spectrum, and mutual power spectrum corresponding to each frame, the correlation average value can be obtained by windowing and segmenting the audio self-power spectrum, vibration self-power spectrum, and mutual power spectrum according to the correlation calculation formula. This correlation average value is then used as a frame-level sub-feature.
[0082] After obtaining the frame-level sub-features corresponding to each frame level, the obtained frame-level sub-features are sorted based on the time sequence of a unified time axis to obtain the correlation features.
[0083] It can be understood that by acquiring noise features, correlation features, and heart rate features, multimodal analysis of the subsequent snoring recognition model can be achieved, so as to form a judgment logic mechanism for snoring recognition, differentiation, and auxiliary confirmation, thereby improving the accuracy of subsequent target snoring extraction.
[0084] Furthermore, in one embodiment, step S3, which fuses noise features, correlation features, and heart rate features to obtain a fused feature sequence, includes:
[0085] S31. Fuse the noise features, correlation features, and heart rate features of the same frame to obtain the target features;
[0086] S32. Based on the time sequence, arrange the target features to obtain the fused feature sequence.
[0087] Before acquiring noise features, correlation features, and heart rate features, the environmental audio signal, vibration signal, and heart rate signal, which are aligned based on time synchronization, are divided into several frames of environmental audio signal, vibration signal, and heart rate signal according to the preset frame level. Then, the required noise features, correlation features, and heart rate features are acquired from the several frames of environmental audio signal, vibration signal, and heart rate signal.
[0088] After acquiring noise features, correlation features, and heart rate features, these features are fused within the same frame to obtain the target features corresponding to that frame. Specifically, the fusion method can be direct fusion or weighted fusion, which can be selected according to the actual design requirements. Since snoring is a continuous behavior, after obtaining the target features corresponding to several frames, the obtained target features are then arranged based on time sequence to obtain the fused feature sequence.
[0089] Furthermore, in one embodiment, step S4, which obtains several corresponding frame-level probabilities based on the fused feature sequence and a preset snoring recognition model, includes:
[0090] S41. Input the fused feature sequence into the preset snoring recognition model;
[0091] S42. Based on the preset recognition logic, analyze and fuse the feature sequence, and output the confidence value corresponding to each frame;
[0092] S43. Map the confidence value and output the corresponding frame-level probability.
[0093] The snoring recognition model can be either a Transformer model or a Bi-LSTM model, which can be selected according to the actual design requirements.
[0094] The Transformer model will be used as an example for explanation. Specifically, after the fused feature sequence is input into the snoring recognition model, the fused feature sequence is linearly projected and positionally encoded. The feature dimension of the fused feature sequence is projected to a higher-dimensional and more expressive model space to transform the fused feature sequence into a more easily processed representation. The corresponding positional code is embedded in the fused feature sequence so that the subsequent model can clearly know the temporal position of each frame in the fused feature sequence.
[0095] After completing the linear projection and positional encoding of the fused feature sequence, a corresponding query vector, key vector, and value vector are generated based on each embedding vector in the fused feature sequence. When processing the Nth frame of the fused feature sequence, the query vector corresponding to the target frame is multiplied by the key vectors of all frames in the fused feature sequence. The score obtained from the dot product is then scaled and applied to the Softmax function to generate the corresponding attention weight distribution. The resulting attention weights are then used as coefficients to perform a weighted summation of the value vectors of all frames to generate a specific vector rich in global information for the Nth frame. It is important to note that the aforementioned process of generating the specific vector for the target frame can be processed in parallel across multiple processes.
[0096] After obtaining several specific vectors, these vectors are concatenated and output to a linear layer for integration, resulting in an enhanced feature sequence. This enhanced feature sequence is then output to a feedforward neural network for nonlinear feature transformation and sequence deepening. After a predetermined number of feature enhancements and transformations, the target feature sequence is obtained.
[0097] The obtained target feature sequence is input into the output projection layer to obtain the confidence value corresponding to each frame; then the obtained confidence value is mapped through the Sigmod function to obtain the corresponding frame-level probability.
[0098] It is understandable that by utilizing fused feature sequences containing multimodal feature information and snoring recognition models, it is possible to output relatively more accurate probability values at the corresponding frame level based on the specific circumstances of different modalities, thereby improving the accuracy of identifying target snoring and thus improving the accuracy of extracting target snoring signals from environmental audio signals.
[0099] Furthermore, in one embodiment, step S5, which extracts the audio signal of the target snoring based on frame-level probabilities and stored ambient audio signals, includes:
[0100] S51. Based on a preset judgment threshold, determine whether the frame-level probability triggers audio extraction;
[0101] S52. If so, extract the corresponding segment from the ambient audio signal to obtain the marked segment;
[0102] S53. Based on a preset duration threshold, filter the marked segments to obtain the target segment;
[0103] S54. Based on the time sequence, splice the target segments to obtain the audio signal of the target snoring sound.
[0104] The determination threshold is a pre-set value used for frame-level probability comparison to determine whether to trigger audio extraction. The marked segment is an audio signal segment of the target snoring that initially meets the requirements and is extracted from the ambient audio signal. The duration threshold is a pre-set duration value used for screening marked segments to filter out marked segments whose duration does not meet the requirements.
[0105] After obtaining the frame-level probabilities, a judgment is made based on a preset threshold to determine whether to extract the corresponding audio segment from the ambient audio signal. When the frame-level probabilities meet the preset requirements, audio extraction is triggered, and the corresponding audio segment is extracted from the stored ambient audio signal as a marked segment; if the frame-level probabilities do not meet the preset requirements, the next frame-level probability judgment is performed. Specifically, the judgment threshold can be a single threshold or a dual threshold, which can be selected according to actual needs.
[0106] In a preferred embodiment, the determination threshold includes a first threshold and a second threshold; the step S51 of determining whether the frame-level probability triggers audio extraction based on the preset determination threshold includes:
[0107] S511. Determine whether the frame-level probability is less than the first threshold;
[0108] S512. If not, then start audio extraction from the corresponding frame in the ambient audio signal;
[0109] S513. Determine whether the frame-level probability is greater than the second threshold;
[0110] S513. If not, stop audio extraction from the corresponding frame in the ambient audio signal.
[0111] The first threshold is used to determine whether to start audio extraction, and the second threshold is used to determine whether to stop audio extraction. The first threshold is greater than the second threshold, and the first and second thresholds can be adjusted according to the actual design.
[0112] After obtaining several frame-level probabilities, a continuous sliding judgment is performed on the frame-level probabilities based on the temporal sequence, the first threshold, and the second threshold. Specifically, it is determined whether the current frame-level probability is less than the first threshold; if the current frame-level probability is less than the first threshold, it is further determined whether the next frame-level probability is less than the first threshold; if the current frame-level probability is greater than or equal to the first threshold, audio extraction is started from the corresponding frame in the ambient audio signal, and it is determined whether the next frame-level probability is greater than the second threshold to determine the stop of this audio extraction.
[0113] After initiating audio extraction, if the current frame-level probability is determined to be less than or equal to the second threshold, audio extraction stops from the corresponding frame in the ambient audio signal. If the current frame-level probability is determined to be greater than the second threshold, audio extraction continues, and then it is determined whether the probability of the next frame-level signal is greater than the second threshold.
[0114] Understandably, using a dual-threshold approach can achieve decision lag, effectively prevent frequent switching near the threshold, significantly improve the continuity and integrity of target snoring extraction, and maintain relatively low computational complexity.
[0115] After determining the frame-level probability, several marked segments were extracted from the environmental audio signal. These marked segments were then filtered and optimized to remove short segments caused by random errors, thereby obtaining the desired target segments.
[0116] In a preferred embodiment, step S53, which filters the marked segments based on a preset duration threshold to obtain the target segment, includes:
[0117] S531. Obtain the duration of the marked segment;
[0118] S532. Determine whether the duration is less than a preset duration threshold;
[0119] S533. If not, then the marked segment is defined as the target segment.
[0120] After obtaining several labeled segments, the labeled segments need to be filtered to reduce the impact of short segments caused by random errors on the overall extraction of target snoring. This can be achieved by filtering the duration of the labeled segments to obtain the desired target segments.
[0121] Specifically, the duration of a marked segment can be determined by its start and stop frames. After obtaining the duration, the duration of each marked segment is compared with a preset duration threshold, which can be adjusted according to the actual design. If the duration is less than the threshold, the marked segment is considered a short segment and is filtered out. If the duration is greater than or equal to the threshold, the marked segment meets the requirements and is defined as the target segment.
[0122] Understandably, by setting a duration threshold and filtering the marked segments by duration, short segments caused by random errors can be eliminated, thereby improving the accuracy of subsequent target snoring extraction and thus improving the accuracy of subsequent sleep quality analysis.
[0123] In summary, this application, by utilizing multimodal integration and frame-level analysis of noise features, correlation features, and heart rate features, enables the identification and extraction of target snoring sounds in complex snoring scenarios, improving the accuracy of target snoring audio extraction and thus enhancing the accuracy of subsequent sleep quality analysis.
[0124] The following is the content of the second aspect of the present invention:
[0125] The present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described snoring extraction method.
[0126] The following is the content of the third aspect of the present invention:
[0127] A third aspect of the present invention provides an earphone, such as Figure 2 As shown, the earphone includes a memory 10, a processor 20, and a snoring extraction method program instruction 30 stored in the memory 10 and executable on the processor 20. When the snoring extraction method program instruction 30 is executed by the processor 20, the aforementioned snoring extraction method is implemented.
[0128] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of the headphones. In this embodiment, the processor is used to run program code stored in a readable storage medium or to process data.
[0129] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0130] The above are merely specific embodiments of this application. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for snoring extraction, characterized in that, include: Acquire stored environmental audio signals, attitude signals, and heart rate signals; Noise features, correlation features, and heart rate features are extracted from the environmental audio signals, posture signals, and heart rate signals. By fusing the noise features, correlation features, and heart rate features, a fused feature sequence is obtained; Based on the fused feature sequence and the preset snoring recognition model, several corresponding frame-level probabilities are obtained; Based on the frame-level probability and the stored environmental audio signal, the audio signal of the target snoring is extracted.
2. The method for snoring extraction according to claim 1, characterized in that, The steps for extracting noise features, correlation features, and heart rate features from the environmental audio signal, posture signal, and heart rate signal include: The vibration signal is obtained from the attitude signal; Based on timestamps, the environmental audio signals, vibration signals, and heart rate signals are synchronized, aligned, and framed. Noise features and heart rate features are extracted from the environmental audio signal and heart rate signal, respectively. Correlation characteristics are obtained based on the environmental audio signal and the vibration signal.
3. The method for snoring extraction according to claim 2, characterized in that, The step of obtaining the correlation characteristics based on the environmental audio signal and the vibration signal includes: Based on the environmental audio signal and the vibration signal in each frame, the corresponding audio self-power spectrum, vibration self-power spectrum and mutual power spectrum are obtained. Based on the audio self-power spectrum, vibration self-power spectrum, and mutual power spectrum of each frame, the average correlation value is obtained as a frame-level sub-feature. Based on the temporal sequence, the frame-level sub-features are arranged to obtain the correlation features.
4. The method for snoring extraction according to claim 1, characterized in that, The step of fusing the noise features, correlation features, and heart rate features to obtain the fused feature sequence includes: The target features are obtained by fusing noise features, correlation features, and heart rate features from the same frame. Based on the time sequence, the target features are arranged to obtain a fused feature sequence.
5. The method for snoring extraction according to claim 1, characterized in that, The step of obtaining several corresponding frame-level probabilities based on the fused feature sequence and the preset snoring recognition model includes: The fused feature sequence is input into a preset snoring recognition model; Based on the preset recognition logic, the fused feature sequence is analyzed, and the confidence value corresponding to each frame is output; Map the confidence values and output the corresponding frame-level probabilities.
6. The method for snoring extraction according to claim 1, characterized in that, The step of extracting the audio signal of the target snoring based on the frame-level probability and the stored environmental audio signal includes: Based on a preset threshold, it is determined whether the frame-level probability triggers audio extraction. If so, extract the corresponding segment from the environmental audio signal to obtain the marked segment; Based on a preset duration threshold, the marked segments are filtered to obtain the target segments; Based on the time sequence, the target segments are spliced together to obtain the audio signal of the target snoring sound.
7. The method for snoring extraction according to claim 6, characterized in that, The determination threshold includes a first threshold and a second threshold; the step of determining whether the frame-level probability triggers audio extraction based on the preset determination threshold includes: Determine whether the frame-level probability is less than the first threshold; If not, then audio extraction is initiated from the corresponding frame in the ambient audio signal; Determine whether the frame-level probability is greater than the second threshold; If not, then audio extraction stops from the corresponding frame in the ambient audio signal.
8. The method for snoring extraction according to claim 6, characterized in that, The step of filtering the marked segments based on a preset duration threshold to obtain the target segment includes: Obtain the duration of the marked segment; Determine whether the duration is less than a preset duration threshold; If not, then the marked fragment is defined as the target fragment.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the snoring extraction method as described in any one of claims 1 to 8.
10. A pair of headphones, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the snoring extraction method as described in any one of claims 1 to 8.