Baby cry recognition method for white noise device

By acquiring and processing multiple audio signals, and combining adaptive echo cancellation and acoustic masking parameters, the anti-interference and low power consumption problems of infant crying sound recognition in white noise devices are solved, achieving high-precision crying sound pickup and reliable monitoring.

CN121922136BActive Publication Date: 2026-05-29深圳市迈远科技有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
深圳市迈远科技有限公司
Filing Date
2026-03-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for recognizing baby cries are easily affected by spontaneous noise interference and complex background noise in white noise devices, resulting in a decrease in recognition rate and high computational complexity, making it difficult to run in real time in low-power embedded devices.

Method used

The system acquires a clean reference signal, a first-channel spatial audio signal, and a second-channel environmental reference signal. Through adaptive echo cancellation algorithm and acoustic masking parameter calculation, combined with lightweight feature detection, voiceprint comparison, and multimodal information fusion, it achieves recognition with high anti-interference and low power consumption.

Benefits of technology

It improves recognition accuracy under strong spontaneous interference, realizes high-precision cry pickup in embedded devices, provides reliable monitoring for identity verification and multi-dimensional verification, and reduces false alarm rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922136B_ABST
    Figure CN121922136B_ABST
Patent Text Reader

Abstract

The application discloses a baby crying sound recognition method for white noise equipment, relates to the field of audio processing and intelligent acoustic recognition, and comprises the following steps: collecting pure reference signals, first signals and second signals for synchronization and preprocessing, and obtaining mixed audio signals based on the first signals; calculating residual signals through a self-adaptive echo cancellation algorithm; calculating acoustic masking parameters based on the residual signals, the pure reference signals and environmental noise estimation; constructing a three-level recognition processing path comprising light feature detection, registered voiceprint comparison and multi-modal information fusion; selecting a recognition processing path based on the numerical range of the acoustic masking parameters, confirming the crying event of the target baby from the residual signals; and triggering corresponding graded alarms based on the crying event confirmation result. The three-level recognition path is scheduled through the acoustic masking parameters, multi-modal information is fused, and accurate and low-power baby crying monitoring under strong interference is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing and intelligent acoustic recognition, specifically a method for recognizing the sound of a baby crying in a white noise device. Background Technology

[0002] With the development of smart home and infant care technologies, baby monitoring devices with integrated audio monitoring functions are becoming increasingly popular. These devices typically analyze the characteristics of an infant's crying in ambient sounds and alert caregivers to improve the timeliness of care.

[0003] However, when such devices also have white noise playback functionality enabled, the strong interference acoustic signals they generate can severely mask the target's crying sound, posing a significant challenge to audio monitoring. Furthermore, the complex background noises in a home environment, such as conversations and appliance noise, further increase the difficulty of accurate identification.

[0004] Existing recognition schemes perform reasonably well in quiet environments, but their recognition rate drops significantly in environments with strong interference (especially spontaneous noise interference). Furthermore, some highly robust algorithms often have high computational complexity, making them difficult to run in real-time in embedded baby monitoring devices that require low power consumption and long standby times. Therefore, how to achieve an accurate cry recognition method that combines high anti-interference capabilities with low power consumption under the typical operating conditions where the device itself continuously generates white noise interference has become a pressing technical problem in this field. Summary of the Invention

[0005] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a method for recognizing baby cries using white noise devices, in order to solve the aforementioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for recognizing infant cries using a white noise device, comprising:

[0007] S1: Acquire a clean reference signal, a first spatial audio signal, and a second environmental reference signal, and perform synchronization and preprocessing to obtain a mixed audio signal based on the first spatial audio signal;

[0008] S2: The residual signal is calculated based on the unified mixed audio signal using an adaptive echo cancellation algorithm;

[0009] S3: Calculate acoustic masking parameters based on residual signal, clean reference signal and ambient noise estimation;

[0010] S4: The recognition and processing path includes: obtaining the first path based on the residual signal through lightweight feature detection, obtaining the second path through registered voiceprint comparison, and obtaining the third path through multimodal information fusion;

[0011] S5: Based on the numerical range of acoustic masking parameters, select the recognition processing path and confirm the crying event of the target infant from the residual signal;

[0012] S6: Trigger corresponding tiered alerts based on the confirmation results of the crying event.

[0013] The present invention is further configured such that S1 specifically includes:

[0014] The digital audio stream of the device’s self-generated acoustic signal is directly obtained through the digital output of the audio codec as a clean reference signal.

[0015] The first spatial audio signal is acquired by a microphone array facing the target area, and the second environmental reference signal is acquired by an independent environmental microphone facing away from the target area.

[0016] DC bias removal and gain calibration are performed on the first spatial audio signal and the second environmental reference signal, respectively.

[0017] The calibrated first spatial audio signal, second environmental reference signal and clean reference signal are time-synchronized and aligned.

[0018] The first spatial audio signal after synchronization and alignment is subjected to directional beamforming to obtain an enhanced mixed audio signal, which mainly includes: sound sources from the target area, residual equipment self-echoing acoustic signals, and some environmental noise.

[0019] The present invention is further configured such that S2 specifically includes:

[0020] The mixed audio signal and the clean reference signal are input into a pre-trained adaptive filter;

[0021] The echo component of the device's spontaneous acoustic signal in the mixed audio signal is estimated by an adaptive filter, and the echo component is subtracted from the mixed audio signal to obtain a preliminary noise reduction signal;

[0022] The initial noise-reduced signal is spatially filtered to further suppress noise interference from non-target directions, resulting in a residual signal.

[0023] The present invention is further configured such that S3 specifically includes: a core frequency band energy and noise component analysis step, a potential cry energy estimation step, and an acoustic masking parameter calculation step.

[0024] The present invention is further configured such that the core frequency band energy and noise component analysis step specifically includes:

[0025] By analyzing the residual signal through bandpass filtering and short-time energy calculation, the energy of the residual signal in the preset core frequency band of infant crying is obtained;

[0026] The energy of the clean reference signal in the core frequency band is obtained by using the same bandpass filtering and short-time energy calculation method.

[0027] Based on the second environmental reference signal, the time-frequency information is calculated by the windowed short-time Fourier transform method, and the environmental noise energy estimate is obtained by the recursive least squares estimation algorithm.

[0028] The present invention is further configured such that the potential cry energy estimation step specifically includes:

[0029] Based on the energy of the residual signal in the core frequency band, the energy of the clean reference signal in the core frequency band, and the estimated energy of the environmental noise, the energy of the potential baby crying signal is estimated by weighted calculation of the noise contribution.

[0030] The present invention is further configured such that the acoustic masking parameter calculation step specifically includes:

[0031] Based on the psychoacoustic model, the masking effect coefficients corresponding to the pure reference signal type are obtained through a pre-set lookup table;

[0032] Based on the energy of the pure reference signal in the core frequency band and the corresponding masking effect coefficient, the masking contribution of the device's spontaneous acoustic signal is determined by multiplying the energy of the pure reference signal in the core frequency band with the masking effect coefficient.

[0033] Based on the environmental noise energy estimate, the environmental noise masking contribution is determined by directly using the environmental noise energy estimate as the masking contribution.

[0034] Based on the ratio of the masking contribution of the device's spontaneous acoustic signal, the masking contribution of environmental noise, and the energy of the potential baby crying signal, acoustic masking parameters are calculated and then smoothed.

[0035] The present invention is further configured such that S4 specifically includes:

[0036] Lightweight time-frequency characteristics are obtained by calculating the short-time zero-crossing rate, spectral centroid, and sub-band energy ratio based on the residual signal.

[0037] Based on lightweight time-frequency features, a fast decision is made through rule-based decision with a preset threshold to obtain the recognition result of the first path;

[0038] Voiceprint feature vectors are extracted from the residual signal using a Mel frequency cepstral coefficient extraction network. The voiceprint feature vectors are then compared with the pre-registered voiceprint model using cosine similarity to obtain the recognition result of the second path.

[0039] The residual signal is processed by an audio event detection model based on a deep neural network to obtain the audio recognition result;

[0040] By analyzing the triaxial acceleration data collected by the motion sensor, the time-domain variance and frequency-domain rhythmic characteristics are extracted to obtain motion sensing information;

[0041] Based on the audio recognition results and motion sensing information, the recognition results of the third path are obtained through weighted summation or decision-level fusion based on confidence.

[0042] The present invention is further configured such that S5 specifically includes:

[0043] The values ​​of the acoustic masking parameters are compared with a preset threshold range. Based on the comparison results, one of the first, second, or third paths is activated as the current recognition processing path.

[0044] The residual signal is processed by the activated recognition processing path to obtain the corresponding path recognition result;

[0045] Based on the path recognition results, contextual judgment is made in combination with the historical recognition status to confirm the crying event of the target infant. The crying event includes crying and non-crying.

[0046] The present invention is further configured such that S6 specifically includes:

[0047] Based on the confirmed crying events and their ongoing status, trigger the corresponding level of alarm.

[0048] The tiered alarms include at least: a local notification alarm triggered by the first crying event, a device-initiated response alarm triggered by a continuous crying event, and a remote notification alarm triggered by continuous crying exceeding a preset duration.

[0049] This invention provides a method for recognizing infant cries using a white noise device. The method comprises the following steps: S1: Acquiring a clean reference signal, a first spatial audio signal, and a second environmental reference signal, and performing synchronization and preprocessing to obtain a mixed audio signal based on the first spatial audio signal; S2: Calculating a residual signal based on the unified mixed audio signal using an adaptive echo cancellation algorithm; S3: Calculating acoustic masking parameters based on the residual signal, the clean reference signal, and environmental noise estimation; S4: The recognition processing path includes: obtaining a first path based on the residual signal through lightweight feature detection, obtaining a second path through registered voiceprint comparison, and obtaining a third path through multimodal information fusion; S5: Selecting the recognition processing path based on the numerical range of the acoustic masking parameters to confirm the target infant's crying event from the residual signal; S6: Triggering a corresponding graded alarm based on the crying event confirmation result. The beneficial effects include:

[0050] Improve the recognition capability of the device under self-noise interference: Quantify the recognition environment by defining and calculating "acoustic masking parameters", and adopt a preprocessing architecture that combines pure reference signal and independent environmental reference signal to effectively remove the white noise played by the device itself and complex environmental noise, so as to achieve high-precision crying sound pickup under strong spontaneous interference.

[0051] Achieving an adaptive optimal balance between recognition performance and system power consumption: By establishing a three-level recognition path (lightweight detection, voiceprint matching, and multimodal fusion) dynamically scheduled by acoustic masking parameters, the system can intelligently switch the algorithm complexity according to the degree of environmental interference, so that the high-precision model is activated only when necessary, thereby ensuring reliable recognition around the clock under the resource constraints of embedded devices.

[0052] Provides reliable monitoring with identity verification and multi-dimensional validation: Infant identity is confirmed through integrated voiceprint recognition, and fusion decision-making is achieved by combining multimodal sensor information, thereby reducing false alarm rates. A tiered alarm mechanism based on recognition results and historical status provides appropriate responses, from local alerts to remote notifications, achieving accurate and user-friendly intelligent monitoring.

[0053] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0055] Figure 1 The flowchart illustrates a method for recognizing a baby's cry using a white noise device, as an exemplary embodiment of the present invention. Detailed Implementation

[0056] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0057] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0058] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0059] Example:

[0060] Methods for recognizing baby cries using white noise devices, such as Figure 1 As shown, it includes:

[0061] S1: Acquire a clean reference signal, a first spatial audio signal, and a second environmental reference signal, and perform synchronization and preprocessing to obtain a mixed audio signal based on the first spatial audio signal;

[0062] S2: The residual signal is calculated based on the unified mixed audio signal using an adaptive echo cancellation algorithm;

[0063] S3: Calculate acoustic masking parameters based on residual signal, clean reference signal and ambient noise estimation;

[0064] S4: The recognition and processing path includes: obtaining the first path based on the residual signal through lightweight feature detection, obtaining the second path through registered voiceprint comparison, and obtaining the third path through multimodal information fusion;

[0065] S5: Based on the numerical range of acoustic masking parameters, select the recognition processing path and confirm the crying event of the target infant from the residual signal;

[0066] S6: Trigger corresponding tiered alerts based on the confirmation results of the crying event.

[0067] The present invention is further configured such that S1 specifically includes:

[0068] The digital audio stream of the device’s self-generated acoustic signal is directly obtained through the digital output of the audio codec as a clean reference signal.

[0069] The first spatial audio signal is acquired by a microphone array facing the target area, and the second environmental reference signal is acquired by an independent environmental microphone facing away from the target area.

[0070] DC bias removal and gain calibration are performed on the first spatial audio signal and the second environmental reference signal, respectively.

[0071] The calibrated first spatial audio signal, second environmental reference signal and clean reference signal are time-synchronized and aligned.

[0072] The first spatial audio signal, after synchronization and alignment, undergoes directional beamforming to obtain an enhanced mixed audio signal. This signal mainly includes: sound sources from the target area, residual echoes of the device's self-emitted acoustic signals, and some environmental noise. Specifically, the system acquires three signals in parallel: the first signal, through an embedded audio system, directly captures the white noise digital audio stream that the device is about to play from the digital audio output interface of the audio codec; this white noise digital audio stream is referred to as the clean reference signal. The second signal, called the first spatial audio signal, is acquired through a linear array of two omnidirectional microphones positioned facing the crib area and spaced at a fixed interval. The third signal, called the second environmental reference signal, is acquired through an independent omnidirectional microphone positioned away from the crib. All analog signals from the microphones are converted into digital signals at a sampling rate of 16,000 times per second. After acquisition, the signals from the microphones are preprocessed. First, the two channels of the first spatial audio signal and the second environmental reference signal are subjected to DC bias removal processing. Specifically, a first-order high-pass digital filter with a cutoff frequency of 10 Hz is applied. Next, a preset digital gain coefficient is applied to each channel for calibration. This digital gain coefficient is calculated and stored during the equipment manufacturing process by playing a standard test tone of 94 dB and based on the difference between the actual output of each microphone and the standard response. Subsequently, the three digital signals are strictly time-synchronized and aligned. The system adds a high-precision timestamp generated by a hardware timer to each audio data block and adjusts the starting point of the data block through an interpolation algorithm to ensure that the clean reference signal, the two channels of the first spatial audio signal, and the second environmental reference signal represent the exact same physical time starting point. Finally, the synchronized first spatial audio signal undergoes directional beamforming processing. The preset target sound source direction is directly in front of the device, i.e., at a 0-degree azimuth angle. Based on the known physical distance between the two microphones and the speed of sound propagation, the system calculates the theoretical time difference for sound from the 0-degree direction to reach the two microphones. Using this time difference, a fractional delay digital filter is applied to the audio data of one channel to achieve phase alignment of the sound from the target direction across the two channels. Then, the audio data from the two channels are added point-by-point to merge into a single-channel signal, which is the enhanced mixed audio signal. After this processing, sound from the preset 0-degree ±30-degree cone region is enhanced, while sound from other directions is suppressed. The final mixed audio signal mainly contains sound sources from the target area, residual white noise echoes from the device after air propagation, and some environmental noise that was not completely suppressed.

[0073] The present invention is further configured such that S2 specifically includes:

[0074] The mixed audio signal and the clean reference signal are input into a pre-trained adaptive filter;

[0075] The echo component of the device's spontaneous acoustic signal in the mixed audio signal is estimated by an adaptive filter, and the echo component is subtracted from the mixed audio signal to obtain a preliminary noise reduction signal;

[0076] The initial noise-reduced signal undergoes spatial filtering to further suppress noise interference from non-target directions, yielding a residual signal. Specifically, the system first receives the synchronized mixed audio signal and clean reference signal from the previous step. Processing is performed frame by frame, with a preset frame length of 32 milliseconds and a 16-millisecond overlap between frames. First, a Hanning window function is applied to each frame of the mixed audio signal and clean reference signal. Then, a Fast Fourier Transform (FFT) is used to convert the two time-domain signals into complex frequency-domain spectra. Before real-time processing begins, the adaptive filter requires an initialization training phase. During training, the device plays a continuous pink noise signal lasting at least 5 seconds as the training signal, and simultaneously acquires the corresponding mixed audio signal. The system uses a normalized least mean square algorithm, taking the frequency domain representation of the training reference signal as input and the frequency domain representation of the mixed audio signal as the desired output, iteratively updating the frequency domain coefficients of the adaptive filter until the prediction error converges stably. The filter coefficients obtained after training represent the acoustic transmission path response from the device's speaker to the main microphone and are stored and loaded as the initial coefficients for real-time processing. In the real-time echo cancellation stage, for each new frame of signal, the system performs a complex multiplication operation between the complex spectrum of the clean reference signal and the current frequency domain coefficients of the adaptive filter to obtain the complex spectrum of the predicted echo signal. Then, the predicted echo spectrum is directly subtracted from the complex spectrum of the mixed audio signal to obtain the complex spectrum of the preliminary denoised signal. Subsequently, according to the rules of the normalized least mean square algorithm, the system uses the power spectrum of the preliminary denoised signal and the power spectrum of the clean reference signal to fine-tune the frequency domain coefficients of the adaptive filter with a preset minimal update step size (e.g., 0.005), enabling it to slowly track long-term changes in the acoustic path. After echo cancellation, the system performs spatial post-filtering on the initial denoised signal to suppress residual noise. First, the current noise power spectrum needs to be estimated. The system maintains a historical noise power spectrum estimate. For each frame, the square of the amplitude spectrum of the initial denoised signal (i.e., the current frame power spectrum) is recursively smoothed and mixed with the historical noise power spectrum according to preset weights: the historical value has a weight of 0.95, and the current frame value has a weight of 0.05, thus obtaining an updated noise power spectrum estimate. Next, the spectral gain coefficient for each frequency point is calculated based on the Wiener filtering principle. For each frequency point, the power of the initial denoised signal at that frequency point is subtracted from the estimated noise power, and the difference is divided by the power of the initial denoised signal. The calculation result is limited to the range of 0 to 1, where 1 represents complete preservation and 0 represents complete suppression. Finally, each frequency point of the complex spectrum of the initial denoised signal is multiplied by the calculated corresponding spectral gain coefficient to achieve weighted suppression of the spectrum. The complex spectrum after the above frequency domain gain adjustment is converted back to the time domain through an inverse fast Fourier transform to obtain the time-domain frame signal.The same Hanning window function as before is applied to this time-domain frame, and it is then spliced ​​with the previous frame signal using an overlapping addition method to finally reconstruct a continuous and complete single-channel time-domain signal, which is the residual signal. At this point, the main linear echo component of the device white noise in the mixed audio signal has been effectively eliminated, and the remaining nonlinear components and environmental noise have been further suppressed. The residual signal mainly retains the target sound source and residual noise.

[0077] The present invention is further configured such that S3 specifically includes: a core frequency band energy and noise component analysis step, a potential cry energy estimation step, and an acoustic masking parameter calculation step. Specifically, the core frequency band energy and noise component analysis step extracts the energy and noise estimation of key frequency bands from multiple signals to provide basic data for subsequent calculations; the potential cry energy estimation step separates and estimates the possible infant cry energy from the mixed energy to quantify the intensity of the target signal; the acoustic masking parameter calculation step calculates the masking ratio based on the ratio of noise masking contribution to potential cry energy to objectively assess the degree of interference of the current acoustic environment on cry recognition.

[0078] The present invention is further configured such that the core frequency band energy and noise component analysis step specifically includes:

[0079] By analyzing the residual signal through bandpass filtering and short-time energy calculation, the energy of the residual signal in the preset core frequency band of infant crying is obtained;

[0080] The energy of the clean reference signal in the core frequency band is obtained by using the same bandpass filtering and short-time energy calculation method.

[0081] Based on the second environmental reference signal, the time-frequency spectrum information is calculated using a windowed short-time Fourier transform method. Then, an estimated environmental noise energy value is obtained using a recursive least squares estimation algorithm. Specifically, the system simultaneously reads in one frame of residual signal, a clean reference signal within the same time period, and the second environmental reference signal. First, for the residual signal, the system passes it through a pre-designed digital bandpass filter. This filter only allows frequency components between 400 Hz and 1200 Hz to pass. After filtering, the system extracts the filtered signal sequence, multiplies each data point in the sequence by itself (i.e., calculates the square), and then sums all the squared values ​​within the frame to obtain a total. This total is recorded and called the "core frequency band energy of the residual signal." Simultaneously, the system performs the same processing on the clean reference signal of the same frame: filtering it through the same 400-1200 Hz bandpass filter, then calculating the square of each point in the output signal sequence and summing the results. The result is recorded as the "core frequency band energy of the clean reference signal." On the other hand, the system processes the second environmental reference signal. After applying a Hanning window function to this second environmental reference signal frame, a Fast Fourier Transform is performed to obtain the spectrum of the signal containing amplitude and phase information. Then, ignoring the phase information, the amplitude value corresponding to each frequency point is squared to obtain the power distribution of the signal at different frequencies, called the current frame power spectrum. Subsequently, the system performs noise spectrum estimation. The "historical noise power spectrum" updated after the processing of the previous frame is stored in memory. For each frequency point in the current frame power spectrum, the system performs the following operations: taking 98% of the historical noise power spectrum value of that frequency point, adding 2% of the power spectrum value of that frequency point in the current frame, and summing these two values. The result is used as the updated new noise power spectrum estimate for that frequency point. This new estimate is saved for calculation in the next frame. This process is performed for all frequency points one by one. Finally, the system sums all the updated noise power spectrum estimates for all frequency points to obtain a value representing the total power of the entire environmental noise, which is recorded as the "environmental noise energy estimate".

[0082] The present invention is further configured such that the potential cry energy estimation step specifically includes:

[0083] Based on the energy of the residual signal in the core frequency band, the energy of the clean reference signal in the core frequency band, and the estimated energy of ambient noise, the potential energy of the infant cry signal is estimated through a weighted calculation that suppresses noise contributions. Specifically, the core frequency band is preset to a frequency range of 400 Hz to 1200 Hz. The estimation process uses two preset fixed weighting coefficients. The first coefficient is a white noise suppression coefficient, with a default value of 1.0 and a reasonable value range between 0.8 and 1.2, used to convert the energy of the clean reference signal into its equivalent masking contribution to the cry at the main microphone. The second coefficient is an ambient noise suppression coefficient, with a default value of 0.8 and a reasonable value range between 0.6 and 1.0, used to convert the estimated ambient noise energy into its equivalent masking contribution in the core frequency band. These coefficients can be determined and fixed before the product leaves the factory by playing noise of known intensity in a standard acoustic environment and collecting data, and then calibrating them through linear regression analysis. The specific calculation process is as follows: First, a weighted transformation of noise contribution is performed: the energy of the clean reference signal in the core frequency band is multiplied by the white noise suppression coefficient to obtain the equivalent masking contribution value of white noise; the estimated environmental noise energy is multiplied by the environmental noise suppression coefficient to obtain the equivalent masking contribution value of environmental noise. Then, the potential crying energy is estimated: from the energy of the residual signal in the core frequency band, the equivalent masking contribution values ​​of white noise and environmental noise are successively subtracted to obtain an intermediate difference. This intermediate difference is compared with the value 0: if the intermediate difference is greater than 0, it is used as the potential baby crying signal energy output for this frame; if the intermediate difference is less than or equal to 0, the potential baby crying signal energy output for this frame is 0. This process ensures that the output energy value is non-negative. Finally, the system outputs a non-negative potential baby crying signal energy value.

[0084] The present invention is further configured such that the acoustic masking parameter calculation step specifically includes:

[0085] Based on the psychoacoustic model, the masking effect coefficients corresponding to the pure reference signal type are obtained through a pre-set lookup table;

[0086] Based on the energy of the pure reference signal in the core frequency band and the corresponding masking effect coefficient, the masking contribution of the device's spontaneous acoustic signal is determined by multiplying the energy of the pure reference signal in the core frequency band with the masking effect coefficient.

[0087] Based on the environmental noise energy estimate, the environmental noise masking contribution is determined by directly using the environmental noise energy estimate as the masking contribution.

[0088] Based on the ratio of the masking contributions of the device's spontaneous acoustic signal, the masking contribution of environmental noise, and the energy of the potential baby crying signal, acoustic masking parameters are calculated and smoothed. Specifically, firstly, masking effect coefficients are obtained based on a psychoacoustic model. An internal lookup table is pre-defined, indexed by audio signal type and energy level. For the white noise type played by the device, the lookup table maps the energy value of the pure reference signal in the core frequency band to a pre-defined discrete energy range and outputs the corresponding masking effect coefficient. For example, for a medium energy range, the pre-defined masking effect coefficient is 1.8, which characterizes the enhanced masking ability of white noise per unit energy relative to broadband noise. Subsequently, the masking contribution of the device's spontaneous acoustic signal is calculated by directly multiplying the obtained masking effect coefficient by the energy value of the pure reference signal in the core frequency band; the product is the masking contribution value of the device's spontaneous acoustic signal. The masking contribution of environmental noise is directly calculated using the estimated environmental noise energy without additional weighting. Next, the instantaneous acoustic masking parameters are calculated. First, the masking contribution of the device's spontaneous acoustic signal is added to the masking contribution of the ambient noise to obtain the total noise masking contribution. Then, a protective bias is applied to the potential energy value of the infant cry signal, i.e., a very small constant (preset to 1×10⁻⁶) is added. - ¹ 0 To prevent division by zero errors, the total noise masking contribution is divided by the biased cry signal energy value, and the quotient is the instantaneous acoustic masking parameter. Finally, the instantaneous acoustic masking parameter is smoothed in the time domain to suppress jitter. The smoothing is achieved using a first-order infinite impulse response filter, and the smoothing coefficient of its recursive calculation formula is preset to 0.7. Specifically, the smoothed acoustic masking parameter value output from the previous frame is multiplied by 0.7, and then the instantaneous acoustic masking parameter value of the current frame is multiplied by 0.3. The sum of the two is the final smoothed acoustic masking parameter output for the current frame.

[0089] The present invention is further configured such that S4 specifically includes:

[0090] Lightweight time-frequency characteristics are obtained by calculating the short-time zero-crossing rate, spectral centroid, and sub-band energy ratio based on the residual signal.

[0091] Based on lightweight time-frequency features, a fast decision is made through rule-based decision with a preset threshold to obtain the recognition result of the first path;

[0092] Voiceprint feature vectors are extracted from the residual signal using a Mel frequency cepstral coefficient extraction network. The voiceprint feature vectors are then compared with the pre-registered voiceprint model using cosine similarity to obtain the recognition result of the second path.

[0093] The residual signal is processed by an audio event detection model based on a deep neural network to obtain the audio recognition result;

[0094] By analyzing the triaxial acceleration data collected by the motion sensor, the time-domain variance and frequency-domain rhythmic characteristics are extracted to obtain motion sensing information;

[0095] Based on the audio recognition results and motion sensing information, the recognition result of the third path is obtained through weighted summation or decision-level fusion based on confidence. Specifically, the system receives the residual signal from the previous step and maintains three independently operating recognition processing paths. The first path performs lightweight feature detection: this path processes the residual signal in frames, with a preset frame length of 20 milliseconds. For each frame, the system calculates three features: short-time zero-crossing rate, which is the number of times the signal crosses zero within the frame; spectral centroid, which is the weighted average frequency of the amplitude spectrum after performing a fast Fourier transform on the frame signal; and sub-band energy ratio, which is the proportion of the energy of each sub-band to the total energy of the four sub-bands when the spectrum from 0 to 4000 Hz is divided into four equal-width sub-bands. The system combines the calculated short-time zero-crossing rate, spectral centroid, and energy ratios of the four sub-bands into a six-dimensional feature vector. Then, it applies a pre-defined rule-based decision logic. For example, if the spectral centroid is greater than 800 Hz, the energy ratio of the second sub-band (1000-2000 Hz) is greater than 0.3, and the short-time zero-crossing rate is greater than 30, then the frame is marked as "suspicious." The system counts the number of frames marked as "suspicious" in the most recent ten consecutive frames. If this number is greater than or equal to six, the first path outputs a "confirmed cry" judgment, using the number divided by ten as the confidence level; otherwise, it outputs "not a cry" and uses the number divided by ten as the confidence level. The second path performs voiceprint comparison for registration: This path analyzes the residual signal with finer granularity, with a preset frame length of 25 milliseconds and a frame shift of 10 milliseconds. For each frame, 40-dimensional Mel-frequency cepstral coefficients are calculated. After accumulating 20 consecutive frames of Mel-frequency cepstral coefficients, the system inputs this 20x40 matrix into a pre-trained Mel-frequency cepstral coefficient extraction network. This network is a lightweight temporal convolutional network structure, ultimately outputting a 128-dimensional voiceprint feature vector. The system then calculates the cosine similarity between this vector and the pre-stored target infant voiceprint template vector. The preset similarity threshold is 0.75. If the calculated similarity is greater than this threshold, the second path outputs a "confirmed target infant cry" judgment, using this similarity value as the confidence level; otherwise, it outputs "non-target cry" and a low confidence level. The third path performs multimodal information fusion: This path synchronously processes data blocks with a duration of 1 second; Audio channel processing: The 1-second residual signal is converted into a Mel spectrogram and input into a pre-trained audio event detection model based on a convolutional recurrent neural network. This model outputs an audio confidence score between 0 and 1; Motion sensing channel processing: Triaxial acceleration data within 1 second is read synchronously. First, the variances of the X, Y, and Z axis acceleration sequences are calculated and summed to obtain the total motion intensity in the time domain; Second, the acceleration amplitude sequence is calculated, a fast Fourier transform is performed, and the total energy in the 1 to 4 Hz frequency band is extracted as the frequency domain rhythmic feature; The motion intensity and rhythmic energy are input into a pre-trained logistic regression classifier to obtain the motion confidence score.The decision-level fusion adopts a weighted summation strategy, with a preset audio weight of 0.7 and a motion weight of 0.3. The fusion score is calculated as (0.7 * audio confidence + 0.3 * motion confidence). The preset fusion score decision threshold is 0.5. If the fusion score is greater than 0.5, the third path outputs a "confirmed crying sound" judgment and the fusion score as confidence; otherwise, it outputs "not crying sound".

[0096] The present invention is further configured such that S5 specifically includes:

[0097] The values ​​of the acoustic masking parameters are compared with a preset threshold range. Based on the comparison results, one of the first, second, or third paths is activated as the current recognition processing path.

[0098] The residual signal is processed by the activated recognition processing path to obtain the corresponding path recognition result;

[0099] Based on the path recognition results and combined with historical recognition states, contextual judgment is performed to confirm the crying event of the target infant. The crying event includes both crying and non-crying. Specifically, the system receives acoustic masking parameters from the previous step. These parameters are dimensionless values ​​used to quantify the degree of interference from the current acoustic environment on the recognition of the target infant's crying sound; a larger value indicates stronger noise masking and a higher recognition challenge. The system internally presets two decision thresholds: for example, a default value of 2.0 for the lower threshold and a default value of 8.0 for the higher threshold. These two thresholds divide the numerical range of the acoustic masking parameter into three intervals, corresponding to the activation conditions of three recognition processing paths. First, path selection based on acoustic masking parameters is performed. The system compares the acoustic masking parameter value of the current frame with a low threshold and a high threshold. If the acoustic masking parameter value is less than the low threshold of 2.0, the first path is activated; if the acoustic masking parameter value is greater than or equal to the low threshold of 2.0 and less than the high threshold of 8.0, the second path is activated; if the acoustic masking parameter value is greater than or equal to the high threshold of 8.0, the third path is activated. The first path is a lightweight feature detection path, the second path is a registered voiceprint comparison path, and the third path is a multimodal information fusion path. Next, the activation path is invoked and identified. Based on the selection result, the system sends a processing request to the corresponding path processing unit and inputs the current residual signal (for the third path, motion sensor data must also be input simultaneously). The activated path processing unit works according to its predetermined algorithm and outputs a path identification result, which includes a judgment label (e.g., "crying" or "non-crying") and a confidence score between 0 and 1. Finally, a context-based historical state decision is executed to confirm the crying event: the system maintains a finite state machine to represent the historical identification states. The finite state machine contains three states: quiet state, suspicious state, and continuous crying state. The system updates the state according to the judgment label and confidence score of the current path identification result, combined with the current state of the finite state machine, based on preset state transition rules, and outputs the final crying event based on the state. The state transition rules are as follows: if the current state is quiet, and the current path identification result is a crying sound with a confidence score higher than 0.7, the state transitions to the suspicious state; otherwise, it remains quiet. If the current state is suspicious, and the path identification results for three consecutive frames are all crying sounds with an average confidence score higher than 0.6, the state transitions to the continuous crying state. If the path identification results for five consecutive frames in the suspicious state are all non-crying sounds, the state transitions back to the quiet state. If the current state is continuous crying, and the path identification results for ten consecutive frames are all non-crying sounds, the state transitions back to the quiet state; otherwise, it remains continuous crying.The output rule for crying events is as follows: when the finite state machine is in a continuous crying state, the crying event ultimately confirmed by the system is "crying"; when the finite state machine is in a quiet state or a suspicious state, the crying event ultimately confirmed by the system is "not crying". The entire calculation process is executed cyclically in units of frames. For example, the system reads the acoustic masking parameter value of the current frame as 5.2. Since 5.2 is greater than or equal to the low threshold of 2.0 and less than the high threshold of 8.0, the system activates the second path. The system inputs the residual signal of the current frame into the second path processing unit, which outputs a path recognition result that is determined to be "crying" with a confidence level of 0.82. Assuming that the end state of the previous frame of the finite state machine was a quiet state, since the current path recognition result is crying and the confidence level of 0.82 is higher than 0.7, according to the rule, there is... When a finite state machine transitions from a quiet state to a suspicious state, the system, based on the output rules, confirms the crying event in this frame as "non-crying" since the current state is suspicious. When the system processes the next frame, the historical state will continue to evolve from the suspicious state. If the finite state machine remains in the suspicious state in the next three consecutive frames due to obtaining a high-confidence crying result, the state will transition from the suspicious state to the continuous crying state. After that, the crying event output by the system will become "crying" until the condition for transitioning back to the quiet state from the continuous crying state is met.

[0100] The present invention is further configured such that S6 specifically includes:

[0101] Based on the confirmed crying events and their ongoing status, trigger the corresponding level of alarm.

[0102] The tiered alarms include at least: a local alert alarm triggered by the first crying event, a device-initiated response alarm triggered by a continuous crying event, and a remote notification alarm triggered when continuous crying exceeds a preset duration. Specifically, the system receives the crying event confirmed in the previous step and the continuous state from the finite state machine. The crying event value is either "crying" or "not crying," and the continuous state value is either quiet, suspicious, or continuous crying. The system maintains a timer variable to accumulate the duration of the continuous crying state; it also maintains three Boolean flags to indicate whether the local alert alarm, device-initiated response alarm, and remote notification alarm have been triggered within the current crying cycle. The system runs checks at a fixed period, with a default period of 100 milliseconds. The system continuously monitors crying events and continuous states. When the continuous state transitions from the suspicious state to the continuous crying state, the timer variable starts incrementing from 0. When the continuous state leaves the continuous crying state, the timer variable immediately resets to 0, and the flags for the local alert alarm, device-initiated response alarm, and remote notification alarm are all reset to false. Alarm triggering is based on conditional judgment of the continuous state, timer value, and alarm flag. The trigger conditions for a local notification alarm are: the continuous state has just entered a continuous crying state, and the local notification alarm trigger flag is false. When triggered, the LED indicator on the system control device enters a slow breathing flashing mode with a preset flashing period of 2 seconds; simultaneously, the local notification alarm trigger flag is set to true. The trigger conditions for a device-active response alarm are: the continuous state remains in a continuous crying state, the accumulated value of the timer variable exceeds the preset active response duration threshold, and the device-active response alarm trigger flag is false. The default value for the active response duration threshold is 5 seconds, corresponding to a timer variable value greater than 50. When triggered, the system starts the audio control program, linearly reducing the volume of the white noise played by the device from its current value to 50% within 3 seconds; simultaneously, the LED indicator mode switches to rapid flashing with a preset flashing period of 0.5 seconds; finally, the device-active response alarm trigger flag is set to true. The triggering conditions for the remote notification alarm are: the baby remains in a continuous crying state, the accumulated value of the timer variable exceeds the preset remote notification duration threshold, and the remote notification alarm triggered flag is false. The default value of the remote notification duration threshold is 30 seconds, corresponding to a timer variable value greater than 300. When triggered, the system sends a notification message to the preset server address through the network module. The message content must include at least the event type "baby crying" and a duration stamp. At the same time, the remote notification alarm triggered flag is set to true.The complete example is as follows: Initially, the continuous state is a quiet state, the timer is 0, and all alarm flags are false; when the continuous state changes to a continuous crying state, the timer starts counting. At the moment the state enters the continuous crying state, the local prompt alarm condition is met, triggering the LED to flash slowly and setting the corresponding flag; when the timer accumulates to 51 (i.e., 5.1 seconds), the device actively responds to the alarm condition, the system executes a white noise volume reduction program and switches the LED to fast flashing, setting the corresponding flag; when the timer accumulates to 301 (i.e., 30.1 seconds), the remote notification alarm condition is met, the system sends a network notification and sets the corresponding flag; if the continuous state subsequently leaves the continuous crying state, the system resets the timer to 0 and resets all alarm flags to false, and the entire alarm system returns to its initial standby state.

[0103] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for recognizing infant cries using a white noise device, characterized in that, include: S1: Acquire a clean reference signal, a first spatial audio signal, and a second environmental reference signal, and perform synchronization and preprocessing. Obtain a mixed audio signal based on the first spatial audio signal. The digital audio stream of the device's self-generated acoustic signal is directly acquired through the digital output of the audio codec as the clean reference signal. The first spatial audio signal is acquired through a microphone array facing the target area, and the second environmental reference signal is acquired through an independent environmental microphone facing away from the target area. S2: The residual signal is calculated based on the mixed audio signal using an adaptive echo cancellation algorithm; S3: Based on the residual signal, clean reference signal, and environmental noise estimation, calculate the acoustic masking parameters. S3 specifically includes: a core frequency band energy and noise component analysis step, a potential cry energy estimation step, and an acoustic masking parameter calculation step. The core frequency band energy and noise component analysis step specifically includes: analyzing the residual signal through bandpass filtering and short-time energy calculation to obtain the energy of the residual signal in the preset infant cry core frequency band; obtaining the energy of the clean reference signal in the core frequency band using the same bandpass filtering and short-time energy calculation method; calculating the time-spectrum information based on the second environmental reference signal using a windowed short-time Fourier transform method, and obtaining the environmental noise energy estimate using a recursive least squares estimation algorithm. The potential cry energy estimation step specifically includes: based on the energy of the residual signal in the core frequency band and the energy of the clean reference signal in the core frequency band... The energy and environmental noise energy estimates are used to estimate the potential energy of the infant crying signal by weighted calculation of the noise suppression contribution. The acoustic masking parameter calculation steps specifically include: obtaining the masking effect coefficient corresponding to the pure reference signal type through a pre-set lookup table based on the psychoacoustic model; determining the masking contribution of the device's spontaneous acoustic signal by multiplying the energy of the pure reference signal in the core frequency band with the masking effect coefficient based on the energy of the pure reference signal in the core frequency band and the corresponding masking effect coefficient; determining the masking contribution of the environmental noise by directly using the environmental noise energy estimate as the masking contribution based on the environmental noise energy estimate; calculating the acoustic masking parameter based on the ratio of the masking contribution of the device's spontaneous acoustic signal, the masking contribution of the environmental noise, and the potential energy of the infant crying signal, and smoothing the acoustic masking parameter. S4: The recognition and processing path includes: obtaining the first path based on the residual signal through lightweight feature detection, obtaining the second path through registered voiceprint comparison, and obtaining the third path through multimodal information fusion; S5: Based on the numerical range of the acoustic masking parameters, select an identification processing path and confirm the crying event of the target infant from the residual signal. Specifically, S5 includes: comparing the value of the acoustic masking parameters with a preset threshold range; activating one of the first, second, or third paths as the current identification processing path based on the comparison result; processing the residual signal through the activated identification processing path to obtain the corresponding path identification result; and making a contextual decision based on the path identification result and historical identification status to confirm the crying event of the target infant, wherein the crying event includes both crying and non-crying. S6: Trigger corresponding tiered alerts based on the confirmation results of the crying event.

2. The method for recognizing infant cries in a white noise device according to claim 1, characterized in that, S1 specifically includes: DC bias removal and gain calibration are performed on the first spatial audio signal and the second environmental reference signal, respectively. The calibrated first spatial audio signal, second environmental reference signal and clean reference signal are time-synchronized and aligned. The first spatial audio signal after synchronization and alignment is subjected to directional beamforming to obtain an enhanced mixed audio signal, which includes: sound sources from the target area, residual equipment self-echoing acoustic signal echoes, and some environmental noise.

3. The method for recognizing infant cries in a white noise device according to claim 1, characterized in that, S2 specifically includes: The mixed audio signal and the clean reference signal are input into a pre-trained adaptive filter; The echo component of the device's spontaneous acoustic signal in the mixed audio signal is estimated by an adaptive filter, and the echo component is subtracted from the mixed audio signal to obtain a preliminary noise reduction signal; The initial noise-reduced signal is spatially filtered to further suppress noise interference from non-target directions, resulting in a residual signal.

4. The method for recognizing infant cries in a white noise device according to claim 1, characterized in that, S4 specifically includes: Lightweight time-frequency characteristics are obtained by calculating the short-time zero-crossing rate, spectral centroid, and sub-band energy ratio based on the residual signal. Based on lightweight time-frequency features, a fast decision is made through rule-based decision with a preset threshold to obtain the recognition result of the first path; Voiceprint feature vectors are extracted from the residual signal using a Mel frequency cepstral coefficient extraction network. The voiceprint feature vectors are then compared with the pre-registered voiceprint model using cosine similarity to obtain the recognition result of the second path. The residual signal is processed by an audio event detection model based on a deep neural network to obtain the audio recognition result; By analyzing the triaxial acceleration data collected by the motion sensor, the time-domain variance and frequency-domain rhythmic characteristics are extracted to obtain motion sensing information; Based on the audio recognition results and motion sensing information, the recognition results of the third path are obtained through weighted summation or decision-level fusion based on confidence.

5. The method for recognizing infant cries in a white noise device according to claim 1, characterized in that, S6 specifically includes: Based on the confirmed crying events and their ongoing status, trigger the corresponding level of tiered alert; The tiered alarms include at least: a local notification alarm triggered by the first crying event, a device-initiated response alarm triggered by a continuous crying event, and a remote notification alarm triggered by continuous crying exceeding a preset duration.