Respiration monitoring method and system coexisting with interference, medium, product and terminal
By using frequency shift algorithm and short-time Fourier transform processing on smart devices to extract breath data features and construct a breath data recovery model, the problem of music interference when the device performs other audio functions is solved, and a highly accurate and robust breath recognition effect is achieved.
Patent Information
- Application Number
- CN202411893096.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
AI Technical Summary
How to effectively eliminate music interference to the acoustic perception system when the device performs other audio functions simultaneously, especially on smart devices with single speakers and single microphones.
The frequency shift algorithm is used to mix the real-time breathing data, and combined with the short-time Fourier transform processing, acoustic perception features and spectral features are extracted and feature fusion is performed. Then, based on the pre-constructed respiratory data recovery model, the interference is eliminated and the target respiratory characteristic data is restored.
It realizes effective removal of music interference and improves the accuracy and robustness of breath recognition without introducing additional hardware, providing a low-cost, robust acoustic sensing framework suitable for smart devices with single speakers and single microphones.
Smart Images

Figure CN119939232A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of respiratory monitoring technology, and in particular to a respiratory monitoring method, system, medium, product and terminal that coexists with interference. Background Art
[0002] With the popularity of smart devices, ultrasonic-based acoustic sensing technology has received widespread attention in various application scenarios, especially in the field of health monitoring. Ultrasonic-based acoustic sensing technology mainly relies on the built-in speakers and microphones of the device to effectively monitor and identify the target object by transmitting and receiving ultrasonic signals.
[0003] At present, frequency modulated continuous wave (FMCW) signals have become the mainstream of ultrasonic sensing technology. FMCW technology emits inaudible ultrasonic signals. When these signals collide with objects and reflect back, the device's microphone can receive these echo signals. By precisely analyzing the changes in the echo signals, key information such as the target's motion state and position can be accurately inferred. Compared with traditional visual or radio frequency sensing technologies, acoustic sensing technology not only has the significant advantage of not exposing personal privacy, but also excels in cost control and deployment flexibility. Therefore, in many fields such as smart homes, car cockpits, and health monitoring, acoustic sensing technology has shown broad application prospects.
[0004] However, although ultrasonic-based acoustic sensing technology has achieved remarkable results in theory and experiments, it still faces many challenges in practical application. In particular, when the device itself has other audio functions, such as playing music, the reception and processing of acoustic signals are easily interfered by music. These interference signals will seriously affect the performance of the acoustic sensing system, resulting in a significant reduction in the actual application effect of the system.
[0005] Most of the current research on acoustic perception technology has failed to fully consider the coexistence of perception and music. The only study tried to use multiple channels of a microphone array to eliminate the interference caused by music. However, this solution encountered considerable difficulties in practical application. Since most commercial smart devices are not equipped with multiple microphones and speakers, and many devices even have only one pair of microphones and speakers, the applicability of this solution on commercial devices is severely limited.
[0006] To sum up, how to effectively eliminate the interference of music on the acoustic perception system while ensuring the normal operation of other audio functions of the device has become a key issue that needs to be solved urgently. Summary of the invention
[0007] In view of the shortcomings of the prior art mentioned above, the purpose of the present application is to provide a respiratory monitoring method, system, medium, product and terminal that coexist with interference, so as to solve technical problems such as how to effectively eliminate the interference of music on the acoustic perception system when the device simultaneously performs other audio functions (such as playing music).
[0008] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a respiratory monitoring method coexisting with interference, including: collecting real-time respiratory data; performing feature extraction on the real-time respiratory data to obtain real-time respiratory acoustic perception feature data and real-time respiratory spectrum feature data; performing feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data; and obtaining target respiratory feature data after eliminating interference based on the real-time respiratory fusion feature data and a pre-constructed respiratory data recovery model.
[0009] In some embodiments of the first aspect of the present application, the method of extracting features from the real-time breathing data includes: performing mixing processing on the real-time breathing data based on a frequency shift algorithm; and performing fast Fourier transform processing on the real-time breathing data after the mixing processing to obtain real-time breathing acoustic perception feature data.
[0010] In some embodiments of the first aspect of the present application, the method of extracting features from the real-time respiratory data includes: performing short-time Fourier transform processing on the real-time respiratory data to obtain real-time respiratory spectrum feature data.
[0011] In some embodiments of the first aspect of the present application, the method of constructing the breathing data recovery model includes: collecting a historical breathing data set; the historical breathing data set includes undisturbed breathing data, disturbed breathing data, undisturbed stationary data, and disturbed stationary data; performing feature extraction and feature fusion on the undisturbed breathing data, disturbed breathing data, undisturbed stationary data, and disturbed stationary data, respectively, to obtain undisturbed breathing fusion feature data, disturbed breathing fusion feature data, undisturbed stationary fusion feature data, and disturbed stationary fusion feature data; superimposing the disturbed stationary fusion feature data with the undisturbed breathing fusion feature data and the undisturbed stationary fusion feature data, respectively, and constructing a training data set based on the disturbed breathing fusion feature data; using the training data set to train a neural network model to construct a breathing data recovery model.
[0012] In some embodiments of the first aspect of the present application, the training data set includes an interfered breathing training data set and an interfered breathing test data set; using the training data set to train a neural network model to construct a breathing data recovery model includes: using the interfered breathing training data set to train a neural network model to obtain a primary breathing data recovery model; using the interfered breathing test data set to test and evaluate the primary breathing data recovery model, and based on a defined contrast loss function and an optimization algorithm, adjusting and optimizing the parameters of the primary breathing data recovery model to obtain a final converged breathing data recovery model.
[0013] In some embodiments of the first aspect of the present application, the neural network model includes an encoder block, a residual block and a decoder block; the encoder block performs feature extraction on the input data; the residual block includes a plurality of residual units, and the output data of the encoder block passes through each residual unit in sequence; the decoder block upsamples the output data of the residual block to obtain the restored target features.
[0014] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a respiratory monitoring system that coexists with interference, including: a data acquisition module for collecting real-time respiratory data; a feature extraction module for performing feature extraction on the real-time respiratory data to obtain real-time respiratory acoustic perception feature data and real-time respiratory spectrum feature data; a feature fusion module for performing feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data; an interference elimination module for obtaining target respiratory feature data after eliminating interference based on the real-time respiratory fusion feature data and a pre-built respiratory data recovery model.
[0015] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the respiratory monitoring method coexisting with interference as described above.
[0016] To achieve the above-mentioned purpose and other related purposes, the fourth aspect of the present application provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer implements the respiratory monitoring method coexisting with interference as described above.
[0017] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides an electronic terminal, including a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement the respiratory monitoring method coexisting with interference as described above.
[0018] As described above, the respiratory monitoring method, system, medium, product and terminal that coexist with interference in the present application have the following beneficial effects: the real-time respiratory data is mixed through a frequency shift algorithm, and the real-time respiratory data is short-time Fourier transform processed, combined with acoustic perception features and spectral features, these two features can comprehensively reflect the characteristics of the respiratory signal, improve the accuracy and robustness of respiratory recognition, and by constructing a respiratory data recovery model, provide a low-cost, robust acoustic perception framework that can be deployed in real time on a smart device with a single speaker and a single microphone, and achieve effective removal of interference signals without introducing additional hardware, processing respiratory data interfered by music, so as to restore the interfered respiratory signal in a complex acoustic environment and obtain the respiratory characteristics of the target user. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Shown is a flow chart of a respiratory monitoring method coexisting with interference in one embodiment of the present application.
[0020] Figure 2 The diagram shows the frequency-time relationship before and after the signal frequency shift in one embodiment of the present application.
[0021] Figure 3 Shown is a schematic diagram of a process for constructing a respiratory data recovery model in one embodiment of the present application.
[0022] Figure 4 Shown is a distribution diagram of interference signals before frequency shift in one embodiment of the present application.
[0023] Figure 5 Shown is a distribution diagram of interference signals after frequency shift in one embodiment of the present application.
[0024] Figure 6 Shown is a structural schematic diagram of a neural network model in one embodiment of the present application.
[0025] Figure 7 Shown is a scenario diagram of a respiratory monitoring method coexisting with interference in one embodiment of the present application.
[0026] Figure 8 Shown is a schematic block diagram of a respiratory monitoring system that coexists with interference in an embodiment of the present application.
[0027] Fig. 9 Shown is a schematic diagram of the structure of an electronic terminal in one embodiment of the present application. DETAILED DESCRIPTION
[0028] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0029] Before further describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following interpretations:
[0030] <1> FMCW (Frequency Modulated Continuous Wave): Frequency modulated continuous wave radar, which is a radar system, transmits a continuous wave signal whose frequency changes linearly with time, and receives the echo signal reflected by the target. By mixing the transmitted signal and the echo signal, a difference frequency signal is obtained. According to the difference frequency signal, the distance, speed and other information of the target can be calculated.
[0031] <2> FFT (Fast Fourier Transform): The full name is Fast Fourier Transform, which is an algorithm for efficiently calculating discrete Fourier transform and its inverse transform.
[0032] <3> STFT (Short-Time Fourier Transform): is a technique for decomposing non-stationary signals (such as speech, music, etc.) into frequency components that vary over time. It combines the ideas of Fourier transform and window function, making it possible to observe the characteristics of the signal in both the time domain and the frequency domain.
[0033] <4> RNN (Recurrent Neural Network): Recurrent neural network is a special neural network structure, which is mainly used to process and predict sequence data. Unlike traditional neural networks, RNN can circulate in the direction of sequence evolution and maintain a hidden state that can capture the information of all time steps in the sequence so far.
[0034] <5> LSTM (Long Short-Term Memory): Long short-term memory network is a special recurrent neural network structure that aims to solve the gradient vanishing or gradient exploding problems faced by traditional RNN when processing long sequence data. LSTM introduces a clever design of self-loop, which enables information to flow over a long time scale, thereby capturing long-term dependencies in the sequence.
[0035] <6> GRU (Gated Recurrent Unit): Gated recurrent unit is a simplified LSTM (Long Short-Term Memory) network structure that aims to maintain the LSTM effect while reducing its complexity.
[0036] <7> Bin: In data processing, it usually refers to a data interval or grouping, which is used to divide continuous data values into discrete categories or ranges. Divide the data value range into several intervals of equal width, and each interval is a bin.
[0037] <8> ReLU (Rectified Linear Unit): Rectified linear unit is an activation function widely used in neural networks.
[0038] <9> IQ domain: It is a way of representing and processing signals in signal processing, especially suitable for application scenarios that require frequency conversion and phase demodulation. In the IQ domain, the signal is decomposed into a superposition of sine and cosine waves, which have different phases, amplitudes and frequencies.
[0039] In the application of ultrasonic-based acoustic sensing technology, especially frequency modulated continuous wave (FMCW) signals, although it has shown broad application prospects in smart homes, car cockpits, health monitoring and other fields, it still faces a key technical challenge in actual application: how to effectively eliminate the interference of music on the acoustic sensing system (especially the system that relies on speakers and microphones to transmit and receive ultrasonic signals) when the device is performing other audio functions (such as playing music) at the same time. This challenge limits the widespread application of acoustic sensing technology in commercial smart devices, because most commercial devices are only equipped with limited audio input and output devices (such as a pair of microphones and speakers), making some existing interference elimination solutions (such as using multiple channels of microphone arrays) difficult to implement.
[0040] In order to solve the problems in the above-mentioned background technology, the present application provides a respiratory monitoring method, system, medium, product and terminal that coexist with interference, aiming to solve technical problems such as how to effectively eliminate the interference of music to the acoustic perception system when the device simultaneously performs other audio functions (such as playing music).
[0041] To facilitate understanding of the embodiments of the present application, first Figure 1 Detailed description. Figure 1 The flowchart of the respiratory monitoring method coexisting with interference in an embodiment of the present invention is shown. The respiratory monitoring method coexisting with interference in this embodiment mainly includes the following steps:
[0042] S101: Collect real-time breathing data.
[0043] In this embodiment, a commercial device with a single speaker and a single microphone is used to collect real-time breathing data. The single speaker is used to transmit FMCW signals, that is, transmit signals. When these signals collide with the human body and reflect back, the microphone of the device can receive the reflected FMCW signals to extract real-time breathing data.
[0044] S102: Extracting features from the real-time breathing data to obtain real-time breathing acoustic perception feature data and real-time breathing spectrum feature data.
[0045] In this embodiment, the method of extracting features from the real-time breathing data includes: performing mixing processing on the real-time breathing data based on a frequency shift algorithm; and performing fast Fourier transform processing on the real-time breathing data after the mixing processing to obtain real-time breathing acoustic perception feature data.
[0046] In this embodiment, if Figure 2 As shown, it is a frequency-time relationship diagram before and after the signal frequency shift in an embodiment of the present invention. Wherein, the curve represents the interference signal in the collected breathing data. The collected breathing data includes a direct path and a target path. The frequency shift algorithm is to perform frequency shift on the reference signal (i.e., the transmitted signal) to obtain more comprehensive interference information, and the reference signal is subjected to a frequency shift (Δf) in a standard acoustic signal processing process, so that the captured interference features include not only a part of the frequency components in the traditional signal, but also more interference information. The frequency shift makes the interference signal components present a symmetrical distribution in the frequency spectrum, so that interference can be effectively extracted and eliminated.
[0047] In this embodiment, the existing signal processing method uses the transmission signal The mixed operation is performed with the received signal S(t), capturing only half of the symmetrical spectrum. Where T(t) represents the transmitted signal, S(t) represents the received signal, f0 is the starting frequency of the FMCW signal, B is the bandwidth, T is the cycle time of the FMCW signal, and t is the time sampling point.
[0048] It is worth noting that in order to extract more feature points, the present invention combines the real-time respiratory data S'(t) with the transmitted signal after frequency shift. Sine and cosine are used for mixing processing respectively, and then the real-time breathing data after mixing processing is fast Fourier transformed, and the obtained data is the ultrasonic perception feature point in the IQ domain. Among them, S'(t) represents real-time breathing data, T'(t) represents the transmitted signal after frequency shift, f0 is the starting frequency of the FMCW signal, B is the bandwidth, T is the cycle time of the FMCW signal, and Δf is the frequency size of the frequency offset. Based on the frequency shift algorithm, the real-time breathing data is mixed, so that the captured real-time breathing data not only includes part of the frequency components in the traditional signal, but also contains more interference information. The frequency shift algorithm obtains more comprehensive interference information by changing the frequency components of the signal, so that the interference signal that was originally not easy to detect is symmetrically distributed in the spectrum, so that the interference can be effectively extracted and eliminated.
[0049] In this embodiment, the real-time respiratory data after the mixing process is subjected to a fast Fourier transform (FFT) process to obtain real-time respiratory acoustic perception feature data, which includes amplitude and phase information at different frequencies.
[0050] In this embodiment, the method of extracting features from the real-time respiratory data includes: performing short-time Fourier transform processing on the real-time respiratory data to obtain real-time respiratory spectrum feature data. The real-time respiratory data is processed using short-time Fourier transform (STFT) to obtain time-frequency domain features. The Hanning window is used in the STFT process, which can be expressed as:
[0051]
[0052] Wherein, w(n) represents the window function, that is, the value of the Hanning window function at sampling point n, n is the sampling point of the signal, and N is the total number of sampling points in the window.
[0053] It is worth noting that the Hanning window used in the short-time Fourier transform processing of the present invention helps to reduce spectrum leakage, make the spectrum clearer, and improve the accuracy of frequency analysis. The Hanning window length used in the present invention is 1.2 times the FMCW signal cycle length, which is 120ms, and the sliding overlap is 20ms, so as to ensure that the sampling rate of the real-time respiratory spectrum feature data obtained by the short-time Fourier transform processing is consistent with the sampling rate of the real-time respiratory acoustic perception feature data.
[0054] S103: performing feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data.
[0055] In the present embodiment, the present invention uses short-time Fourier transform to decompose the real-time respiratory data in the time and frequency domains, thereby revealing the frequency component changes of the signal to obtain real-time respiratory spectrum feature data. These real-time respiratory spectrum feature data contain the key information of the target signal and also include interference components from internal sounds. Then, the real-time respiratory spectrum feature data is connected along the time axis with the real-time respiratory acoustic perception feature data obtained by the frequency shift algorithm, and the two types of features are fused to combine the acoustic signal and the spectrum feature, which is conducive to identifying the characteristics of the interference signal and removing the interference, and further improving the effect of interference elimination.
[0056] S104: According to the real-time breathing fusion feature data, and based on a pre-built breathing data recovery model, to obtain target breathing feature data after eliminating interference.
[0057] In this embodiment, if Figure 3 As shown, a schematic diagram of a process of constructing a breathing data recovery model in an embodiment of the present invention is shown. The method of constructing the breathing data recovery model includes:
[0058] S301: Collecting a historical breathing data set; the historical breathing data set includes non-interference breathing data, interference breathing data, non-interference stationary data, and interference stationary data.
[0059] In this embodiment, a commercial device with a single speaker and a single microphone is used to collect historical breathing data sets. The present invention uses an FMCW signal with a period of 100ms and a frequency of 18kHz-22kHz. Non-interference breathing data refers to breathing data collected on the device in a quiet situation. Interference breathing data refers to collecting disturbed breathing data when playing music using the same device. Non-interference static data refers to collecting static data without target motion (breathing) using the same device in a quiet situation. Interference static data refers to collecting static data without target motion when playing music using the same device.
[0060] S302: performing feature extraction and feature fusion on the non-interference breathing data, the interfered breathing data, the non-interference still data, and the interfered still data respectively, so as to obtain non-interference breathing fusion feature data, interfered breathing fusion feature data, non-interference still fusion feature data, and interfered still fusion feature data.
[0061] In this embodiment, the manner of extracting features from the non-interference breathing data, the interference breathing data, the non-interference stationary data, and the interference stationary data respectively includes:
[0062] (1) Based on the frequency shift algorithm, the non-interference breathing data, the interference breathing data, the non-interference stationary data, and the interference stationary data are mixed and processed respectively.
[0063] In this embodiment, the transmission signals corresponding to the collected undisturbed breathing data, disturbed breathing data, undisturbed stationary data, and disturbed stationary data are frequency shifted (Δf), and the undisturbed breathing data, disturbed breathing data, undisturbed stationary data, and disturbed stationary data are mixed with the corresponding frequency-shifted transmission signals using sine and cosine, and then fast Fourier transform is performed, and the obtained data are ultrasonic sensing feature points in the IQ domain. Based on the frequency shift algorithm, the data in the historical breathing data set are mixed, so that the captured historical breathing data set includes not only a part of the frequency components in the traditional signal, but also more interference information.
[0064] In this embodiment, if Figure 4 As shown in FIG. 1 , a distribution diagram of interference signals before frequency shifting in an embodiment of the present invention is shown. Figure 5 As shown in FIG. 1 , a distribution diagram of interference signals after frequency shifting in an embodiment of the present invention is shown. Figure 4 , 5 As shown, the frequency shift algorithm obtains more comprehensive interference information by changing the frequency components of the signal, so that the originally imperceptible interference signal is symmetrically distributed in the spectrum, so that the interference signal can be more comprehensively identified and removed during the signal processing process. Its principle and calculation formula have been described in detail in step S102, and for the sake of brevity, they will not be repeated here.
[0065] (2) Performing fast Fourier transform processing on the undisturbed breathing data, the disturbed breathing data, the undisturbed stationary data, and the disturbed stationary data after the mixing processing, respectively, to obtain undisturbed breathing acoustic perception feature data, disturbed breathing acoustic perception feature data, undisturbed stationary acoustic perception feature data, and disturbed stationary acoustic perception feature data.
[0066] In this embodiment, the historical breathing data set after the mixing process is subjected to a fast Fourier transform process to obtain corresponding acoustic perception feature data.
[0067] In the present embodiment, the mode of feature extraction of the undisturbed breathing data, the disturbed breathing data, the undisturbed stationary data, and the disturbed stationary data respectively also includes: the undisturbed breathing data, the disturbed breathing data, the undisturbed stationary data, and the disturbed stationary data are respectively subjected to short-time Fourier transform processing, so as to obtain undisturbed breathing spectrum characteristic data, disturbed breathing spectrum characteristic data, undisturbed stationary spectrum characteristic data, and disturbed stationary spectrum characteristic data. In the process of short-time Fourier transform processing, Hanning window is used, as shown in formula (one), it is helpful to reduce spectrum leakage, make spectrum clearer, and improve the accuracy of frequency analysis. The Hanning window length used in the present invention is 1.2 times of the FMCW signal cycle length, which is 120ms, and the sliding overlap is 20ms, so as to ensure that the sampling rate of the spectrum characteristic data obtained by short-time Fourier transform processing is consistent with the sampling rate of acoustic perception feature data.
[0068] In this embodiment, the manner of performing feature fusion on the non-interference breathing data, the interference breathing data, the non-interference stationary data, and the interference stationary data respectively includes:
[0069] (1) Performing feature fusion on the non-interference respiratory acoustic perception feature data and the non-interference respiratory spectrum feature data to obtain non-interference respiratory fusion feature data.
[0070] (2) performing feature fusion on the disturbed respiratory acoustic perception feature data and the disturbed respiratory spectrum feature data to obtain disturbed respiratory fusion feature data.
[0071] (3) Performing feature fusion on the non-interference stationary acoustic perception feature data and the non-interference stationary spectrum feature data to obtain non-interference stationary fused feature data.
[0072] (4) Fusing the interfered stationary acoustic perception feature data with the interfered stationary spectrum feature data to obtain interfered stationary fused feature data.
[0073] S303: Superimposing the interfering static fusion feature data with the non-interfering breathing fusion feature data and the non-interfering static fusion feature data respectively, and constructing a training data set based on the interfering breathing fusion feature data.
[0074] In this embodiment, the interfering static fusion feature data is superimposed with the non-interfering breathing fusion feature data, and the interfering static fusion feature data is superimposed with the non-interfering static fusion feature data to construct an interfering breathing training data set. The interfering breathing fusion feature data is used to construct an interfering breathing test data set.
[0075] In this embodiment, before superposition, the disturbed static fusion feature data and the non-interfered breathing fusion feature data are respectively multiplied by a random floating point number of 0-2, and the disturbed static fusion feature data and the non-interfered static fusion feature data are respectively multiplied by a random floating point number of 0-2 to simulate the different volumes of interference and breathing signals in actual scenes. Exemplarily, if a pieces of disturbed static data, b pieces of non-interfered breathing data, and c pieces of non-interfered static data are collected, a*b pieces of disturbed breathing data and a*c pieces of disturbed static data can be synthesized by the method of data superposition to jointly construct a disturbed breathing training data set for model training.
[0076] It is worth noting that the present invention combines the uninterrupted respiratory signal, the static signal with interference but not containing respiratory activity, and the static signal without interference and not containing respiratory activity by superimposing the interfering static fusion feature data with the uninterfered respiratory fusion feature data and the uninterfered static fusion feature data, thereby ensuring that the model can be effectively trained without real interference data, so that the model can learn how to separate and recover the respiratory signal from the complex signal, which has the following advantages:
[0077] (1) Enhanced robustness: By simulating multiple signal combinations, the model can learn more about how to identify and process interference signals, thereby improving its performance in complex environments.
[0078] (2) Improved generalization ability: Since the training data covers a variety of possible situations, the model can show better adaptability when faced with new, unseen data.
[0079] (3) Reduce dependence on real interference data: By training the model with synthetic data, the dependence on actually collecting a large amount of real interference data is reduced, thus reducing the cost and difficulty of data collection.
[0080] S304: Using the training data set to train a neural network model to construct a respiratory data recovery model.
[0081] In this embodiment, the training data set includes a disturbed breathing training data set and a disturbed breathing test data set; using the training data set to train a neural network model to construct a breathing data recovery model includes:
[0082] (1) Using the disturbed breathing training data set to train a neural network model to obtain a primary breathing data recovery model.
[0083] (2) Using the interfered breathing test data set to test and evaluate the primary breathing data recovery model, and based on the defined contrast loss function and optimization algorithm, adjusting and optimizing the parameters of the primary breathing data recovery model to obtain a final converged breathing data recovery model.
[0084] In this embodiment, if Figure 6 As shown, it is a schematic diagram of the structure of the neural network model in an embodiment of the present invention. The neural network model includes an encoder block, a residual block and a decoder block; the encoder block extracts features from the input data; the residual block includes a plurality of residual units, and the output data of the encoder block passes through each residual unit in turn; the decoder block upsamples the output data of the residual block to obtain the restored target features.
[0085] In this embodiment, a neural network model is used to construct an autoencoder network, and the neural network model is trained using a disturbed breathing training data set to model the mapping relationship from disturbed data to clean data. The network can not only process high-dimensional acoustic signal features, but also has low computational complexity, which is suitable for deployment on resource-constrained devices (such as smart watches, smart phones, etc.). In the encoder block, since the input feature is a two-dimensional tensor of time × frequency, the present invention uses two-dimensional convolution to extract features. In addition, the present invention uses dilated convolution to expand the receptive field, so that the network can capture the dependency between time and frequency without relying on RNN models (such as LSTM and GRU) with high computational overhead, thereby reducing the number of network parameters.
[0086] In this embodiment, the residual block is composed of multiple residual units, which maintain key information while enhancing features through jump connections. The residual design ensures that the network can learn deeper representations without losing the original feature information, thereby maintaining the stability and effectiveness of the feature recovery process. The residual structure creates a short connection between the previous layer and the current layer, which can be expressed as y=F(x)+x, where y represents the output of the current residual unit (or current layer), x represents the input of the residual block, and F(x) represents the result of a series of transformations (such as convolution, activation, etc.) performed by the current residual unit on the input x. X is directly added to F(x) through a jump connection to form the final output y. In the residual block, the output y is obtained by adding the transformed output F(x) of the current residual unit to the input x element by element, that is, y=F(x)+x. This addition operation retains the original information of the input x and combines the new features learned by the current layer, thereby enhancing the representation ability of the network. Residual connections help simplify the optimization process of the network while enhancing the representation ability of the network. In each residual layer, we employ one-dimensional convolution.
[0087] In this embodiment, in the decoder block, the features are upsampled using a 2D convolutional layer and decoded back to the target space. The output of the decoder block is then combined with the original target features to generate the restored target features. In each module, batch normalization and ReLU activation functions are performed after all convolutional layers.
[0088] In this embodiment, the neural network model is trained using the interfering breathing training data set, so that the model learns how to recover the original breathing signal from the interfering signal. During the training process, the model will continuously adjust its parameters to minimize the difference between the predicted value and the true value. The primary breathing data recovery model is tested and evaluated using the interfering breathing test data set. By defining a contrast loss function (such as mean square error, cross entropy loss, etc.), the difference between the model prediction value and the true value can be quantified. Based on the results of the contrast loss function and the optimization algorithm (such as gradient descent), the parameters of the primary breathing data recovery model are adjusted and optimized. This optimization process will continue until the model reaches the final convergence state, that is, the loss function value no longer decreases significantly, and the performance of the model remains stable on the test data set to obtain a finally converged breathing data recovery model.
[0089] In this embodiment, the neural network model is trained using the TensorFlow framework, which is an open source machine learning library designed to simplify the construction, training and deployment of machine learning models. It provides a comprehensive and flexible ecosystem suitable for research and production machine learning applications of all sizes. The core of TensorFlow is a data flow graph, in which nodes represent mathematical operations and edges represent multidimensional arrays (tensors) passed between these nodes. This graph structure enables TensorFlow to perform complex computing tasks efficiently, especially in distributed computing environments. TensorFlow can automatically calculate the gradients of nodes in the graph, which is crucial for training neural networks. Through the backpropagation algorithm, TensorFlow can automatically update the weights in the network to minimize the loss function. The trained model can be easily deployed to various platforms, including servers, mobile devices, web browsers, and edge devices.
[0090] In this embodiment, if Figure 7 As shown, a schematic diagram of a scenario of a respiratory monitoring method coexisting with interference in an embodiment of the present invention is shown. The trained respiratory data recovery model is deployed to commercial equipment to eliminate interference from the collected real-time respiratory data (audio data), including the following three steps:
[0091] The first step is to extract features from the real-time respiratory data, including the following: (1) After mixing using a frequency shift algorithm, fast Fourier transform processing is performed to obtain ultrasonic sensing feature points in the IQ domain, i.e., real-time respiratory acoustic sensing feature data. (2) Short-time Fourier transform (STFT) processing is performed on the real-time respiratory data, and a Hanning window is selected as the window function to intercept different time periods of the signal and obtain the spectrum information of the signal; then, real-time respiratory spectrum feature data can be extracted based on the spectrum information. (3) Feature fusion is performed on the real-time respiratory acoustic sensing feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data.
[0092] Step 2: Input the real-time respiratory fusion feature data into the pre-built respiratory data recovery model (i.e. Figure 7 The recovery network in the image is used to eliminate interference signals to obtain clean target breathing feature data.
[0093] Step 3: Perception application, using clean target breathing feature data for breathing monitoring, taking the upper half of the bin containing features in the target breathing feature data after the above interference is eliminated for breathing monitoring. The present invention uses a conventional breathing monitoring method, first using the energy intensity of each bin to make a judgment, and selecting the bin with the strongest energy as the bin corresponding to the breathing distance. Then the principal component analysis (PCA) algorithm is used to calculate the breathing rate.
[0094] In this embodiment, principal component analysis is a powerful data analysis tool and a technique for exploring high-dimensional data structures. Its main purpose is to reduce the dimensionality of the data while retaining as much information as possible in the data set. Through principal component analysis, high-dimensional respiratory feature data can be converted into low-dimensional principal component scores, thereby simplifying the data and removing noise and redundant information. After obtaining the principal component scores, these scores can be used to calculate the respiratory rate. Since the principal component scores have removed the redundant information in the original data, the calculated respiratory rate is more accurate and reliable.
[0095] It is worth noting that the breathing monitoring method coexisting with interference in the present invention performs mixing processing on the real-time breathing data through a frequency shift algorithm, and performs short-time Fourier transform processing on the real-time breathing data, and combines acoustic perception features and spectral features. These two features can fully reflect the characteristics of the breathing signal, improve the accuracy and robustness of breathing recognition, and by constructing a breathing data recovery model, provide a low-cost, robust acoustic perception framework that can be deployed in real time on a smart device with a single speaker and a single microphone, and achieves effective removal of interference signals without introducing additional hardware, processing breathing data interfered by music, so as to restore the interfered breathing signal in a complex acoustic environment and obtain the breathing characteristics of the target user.
[0096] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0097] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0098] Figure 8 is a schematic block diagram of a respiratory monitoring system coexisting with interference provided by an embodiment of the present application. Figure 8 As shown, the respiratory monitoring system 800 for coexistence with interference includes:
[0099] The data acquisition module 801 is used to acquire real-time respiratory data.
[0100] The feature extraction module 802 is used to extract features from the real-time breathing data to obtain real-time breathing acoustic perception feature data and real-time breathing spectrum feature data.
[0101] The feature fusion module 803 is used to perform feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data.
[0102] The interference elimination module 804 is used to obtain target breathing characteristic data after eliminating interference according to the real-time breathing fusion characteristic data and based on a pre-built breathing data recovery model.
[0103] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0104] It should also be understood that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0105] Fig. 9 is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. The electronic terminal includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the respiratory monitoring method coexisting with interference as described above. Fig. 9 As shown, the electronic terminal 900 includes: at least one processor 901, a memory 902, at least one network interface 903 and a user interface 905. The various components in the device are coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Fig. 9 In the specification, various buses are labeled as bus systems.
[0106] The user interface 905 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0107] It is understood that the memory 902 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable categories of memory.
[0108] The memory 902 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 900. Examples of these data include: any executable program for operating on the electronic terminal 900, such as an operating system 9021 and an application 9022; the operating system 9021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 9022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The respiratory monitoring method for coexistence with interference provided in the embodiment of the present invention may be included in the application 9022.
[0109] The method disclosed in the above embodiment of the present invention can be applied to the processor 901, or implemented by the processor 901. The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 901 or the instruction in the form of software. The above processor 901 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 901 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general processor 901 may be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0110] In an exemplary embodiment, the electronic terminal 900 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute the aforementioned method.
[0111] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: computer program code, when the computer program code is executed on a computer, the computer executes the respiratory monitoring method coexisting with interference as described above.
[0112] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the respiratory monitoring method coexisting with interference as described above.
[0113] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program and / or a computer. By way of illustration, both applications running on a computing device and a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).
[0114] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0115] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0117] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0119] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When loading and executing computer program instructions (programs) on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media integrations. Available media may be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid state disks (SSDs), etc.).
[0120] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0121] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0122] In summary, the present application provides a method, system, medium, product and terminal for monitoring breathing coexisting with interference, performs mixing processing on the real-time breathing data through a frequency shift algorithm, and performs short-time Fourier transform processing on the real-time breathing data, combines acoustic perception features and spectrum features, these two features can fully reflect the characteristics of breathing signals, improve the accuracy and robustness of breathing recognition, and by constructing a breathing data recovery model, provides a low-cost, robust acoustic perception framework, which can be deployed in real time on a single-speaker, single-microphone smart device, and achieves effective removal of interference signals without introducing additional hardware, processing breathing data interfered by music, so as to recover the interfered breathing signal in a complex acoustic environment, and obtain the breathing characteristics of the target user. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.
[0123] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A respiratory monitoring method coexisting with interference, characterized in that: include: Collect real-time respiratory data; Performing feature extraction on the real-time breathing data to obtain real-time breathing acoustic perception feature data and real-time breathing spectrum feature data; Performing feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data; According to the real-time breathing fusion feature data, and based on the pre-built breathing data recovery model, the target breathing feature data after eliminating interference is obtained.
2. The respiratory monitoring method coexisting with interference according to claim 1, characterized in that: The method of extracting features from the real-time breathing data includes: performing frequency mixing processing on the real-time breathing data based on a frequency shift algorithm; and performing fast Fourier transform processing on the real-time breathing data after the frequency mixing processing to obtain real-time breathing acoustic perception feature data.
3. The respiratory monitoring method coexisting with interference according to claim 1, characterized in that: The method of extracting features from the real-time respiratory data includes: performing short-time Fourier transform processing on the real-time respiratory data to obtain real-time respiratory spectrum feature data.
4. The respiratory monitoring method coexisting with interference according to claim 1, characterized in that: The method of constructing the breathing data recovery model includes: Collecting a historical respiratory data set; the historical respiratory data set includes non-interference respiratory data, interference respiratory data, non-interference static data, and interference static data; Performing feature extraction and feature fusion on the non-interference breathing data, the interfered breathing data, the non-interference static data, and the interfered static data, respectively, to obtain non-interference breathing fusion feature data, interfered breathing fusion feature data, non-interference static fusion feature data, and interfered static fusion feature data; Superimposing the interfering static fusion feature data with the non-interfering breathing fusion feature data and the non-interfering static fusion feature data, respectively, and constructing a training data set based on the interfering breathing fusion feature data; The training data set is used to train a neural network model to construct a respiratory data recovery model.
5. The respiratory monitoring method coexisting with interference according to claim 4, characterized in that: The training data set includes a disturbed breathing training data set and a disturbed breathing test data set; Using the training data set to train a neural network model to construct a respiratory data recovery model includes: Using the disturbed breathing training data set to train a neural network model to obtain a primary breathing data recovery model; The primary breathing data recovery model is tested and evaluated using the disturbed breathing test data set, and based on the defined contrast loss function and optimization algorithm, the parameters of the primary breathing data recovery model are adjusted and optimized to obtain a final converged breathing data recovery model.
6. The respiratory monitoring method coexisting with interference according to claim 1, characterized in that: The neural network model includes an encoder block, a residual block and a decoder block; the encoder block extracts features from input data; the residual block consists of a plurality of residual units, and the output data of the encoder block passes through each residual unit in sequence; the decoder block upsamples the output data of the residual block to obtain restored target features.
7. A respiratory monitoring system coexisting with interference, characterized in that: include: Data acquisition module, used to collect real-time respiratory data; A feature extraction module, used to extract features from the real-time respiratory data to obtain real-time respiratory acoustic perception feature data and real-time respiratory spectrum feature data; A feature fusion module, used for performing feature fusion on the real-time respiratory acoustic perception feature data and the real-time respiratory spectrum feature data to obtain real-time respiratory fusion feature data; The interference elimination module is used to obtain target breathing characteristic data after eliminating interference according to the real-time breathing fusion characteristic data and based on a pre-built breathing data recovery model.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the breathing monitoring method coexisting with interference according to any one of claims 1 to 6 is implemented.
9. A computer program product, characterized in that The computer program product includes computer program codes, and when the computer program codes are executed on a computer, the computer is enabled to implement the respiratory monitoring method coexisting with interference according to any one of claims 1 to 6.
10. An electronic terminal comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the breathing monitoring method coexisting with interference according to any one of claims 1 to 6.