Method and system for human voice prosody driven complex neuroacoustic signal generation

CN122676802APending Publication Date: 2026-09-01FERD MANSON MULTIMEDIA TECH SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611130987.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

一方面,固定频率的神经声学信号与人声内容自身的语音时间结构缺乏关联,容易产生节律错位,使听感显得生硬和突兀;另一方面,持续的神经声学信号层若不受控制地与人声叠加,会遮蔽人声的可懂度,尤其在需要准确听取语义信息的场景中,这一问题更为突出

Benefits of technology

第一,本发明实现了神经声学信号生成从固定频率预设向语音节律驱动的转变。现有技术中,双耳节拍、单耳节拍或调制噪声等神经声学信号通常以恒定目标频率或预设频段生成,其时间结构与人声内容无关。本发明通过从输入人声数字音频采样序列x[n]中提取语音幅度包络E[n]、包络调制谱M[l]、包络主导调制频率fp、音节节律频率fs、短语停顿节律频率fq及包络峰值周期频率fb等多种语音节律特征,并经加权融合计算得到语音节律频率fspeech(t),再根据目标使用状态和fspeech(t)的倍频或分频映射结果确定目标神经声学节律频率ftarget(t),使得所生成的复合神经声学信号层sneu[n]在时间结构上与人声内容的固有节律相匹配,避免了节律错位和听感突兀的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676802A_ABST
    Figure CN122676802A_ABST
Patent Text Reader

Abstract

The present application relates to a human voice rhythm driven complex neuroacoustic signal generation method and system. It includes: obtaining a digital audio sample sequence containing human voice; extracting the speech amplitude envelope, envelope modulation spectrum, envelope dominant modulation frequency, syllable rhythm frequency, phrase pause rhythm frequency and envelope peak period frequency; calculating the speech rhythm frequency; determining the candidate frequency according to the target use state, and performing frequency multiplication or frequency division mapping to obtain the mapping frequency, and calculating the target neuroacoustic rhythm frequency; generating the complex neuroacoustic signal layer; dynamically adjusting the gain, modulation depth or background layer coefficient according to the clear semantic stage, pause or envelope valley stage and burst sound event; performing digital sampling level mixing to output the mixed digital audio sequence. The present application forms a time structure coordination relationship between the complex neuroacoustic signal and the human voice rhythm, reduces the masking and jarring feeling caused by the continuous superposition of the auxiliary signal while maintaining the intelligibility of the human voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital audio signal processing and synthesis, specifically to a method and system for generating complex neuroacoustic signals driven by human voice rhythm. Background Technology

[0002] In recent years, neuroacoustic signals have received widespread attention in rhythmic auditory stimulation, attention aids, relaxation aids, sleep aids, and related research scenarios. Neuroacoustic signals in the form of binaural beats, monoaural beats, isochronous sequences, and amplitude-modulated noise or natural sounds have been used to generate auditory aid signals with specific temporal structures. However, most existing neuroacoustic signal generation methods employ fixed target frequencies or open-loop generation based on preset frequency bands. While these methods can achieve basic rhythmic stimulation when played alone, they reveal significant technical limitations when played simultaneously with content containing human voices (such as podcasts, lectures, conference speeches, news broadcasts, readings, online courses, and teacher lectures). On the one hand, the lack of correlation between fixed-frequency neuroacoustic signals and the inherent temporal structure of the human voice content easily leads to rhythmic misalignment, making the listening experience harsh and abrupt. On the other hand, uncontrolled superposition of continuous neuroacoustic signal layers with human voices can obscure the intelligibility of the voice, a problem particularly pronounced in scenarios requiring accurate semantic information. Meanwhile, human voice content itself has rich speech rhythm features, including fluctuations in speech amplitude envelope, distribution of the main peak of envelope modulation spectrum, syllable density variation, pause structure between phrases and sentences, and periodicity of envelope peaks. These natural rhythm information reflect the temporal organization rules of speech. However, existing public schemes still rarely utilize human voice speech rhythm features as direct input for neuroacoustic signal generation and dynamic control.

[0003] Furthermore, existing systems typically lack the ability to dynamically adjust the gain, modulation depth, and background layer coefficients of the neuroacoustic signal layer based on the semantic stage of speech, pauses, envelope valley values, or sudden sound events. This results in excessive interference from auxiliary signals when clear semantic transmission is required, while failing to effectively utilize rhythmic gaps to enhance the auxiliary effect during pauses or weak semantic stages.

[0004] Therefore, there is an urgent need for a method and system that can use human speech rhythm as the driving source and dynamically adjust the generation and mixing output of complex neuroacoustic signals according to the temporal structure of human speech content. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for generating composite neuroacoustic signals driven by human voice rhythm. This method extracts speech rhythm features from digital audio containing human voice, generates target neuroacoustic rhythm frequencies accordingly, and dynamically adjusts the mixing intensity of the composite neuroacoustic signal layer and the human voice content. This allows the auxiliary signal to be coordinated with the inherent rhythm of the human voice, improving the coordination between the composite neuroacoustic signal and the temporal structure of the human voice while maintaining the intelligibility of the human voice, and reducing the masking and abruptness caused by the continuous superposition of auxiliary signals.

[0006] To achieve the above objectives, a method for generating composite neuroacoustic signals driven by human voice rhythm is designed, comprising the following steps: acquiring a digital audio sampling sequence x[n] containing human voice; extracting speech rhythm features from the digital audio sampling sequence x[n], wherein the speech rhythm features include: speech amplitude envelope E[n], envelope modulation spectrum M[l], envelope dominant modulation frequency fp, syllable rhythm frequency fs, phrase pause rhythm frequency fq, and envelope peak period frequency fb; calculating the speech rhythm frequency fspeech(t) based on the speech rhythm features; determining the candidate frequency fstate based on the target usage state, and performing frequency doubling or frequency division mapping based on the speech rhythm frequency fspeech(t) to obtain the mapped frequency fspeechmapp. ed(t), then calculate the target neuroacoustic rhythm frequency ftarget(t) based on the candidate frequency fstate and the mapped frequency fspeechmapped(t); generate a composite neuroacoustic signal layer sneu[n] based on the target neuroacoustic rhythm frequency ftarget(t), wherein the composite neuroacoustic signal layer sneu[n] is a digital audio sampling sequence; dynamically adjust at least one of the gain, modulation depth, or background layer coefficient of the composite neuroacoustic signal layer sneu[n] based on the speech stage information or sudden event information of the human voice content; perform digital sampling level mixing of the composite neuroacoustic signal layer sneu[n] and the digital audio sampling sequence x[n] to output the mixed digital audio sequence y[n].

[0007] Preferably, the present invention further includes: the step of extracting the speech amplitude envelope E[n] from the digital audio sampling sequence x[n] specifically includes: performing a Hilbert transform operation on the digital audio sampling sequence x[n] to obtain a Hilbert component xH[n]; constructing an analytic signal z[n] = x[n] + j·xH[n], where j represents the imaginary unit; calculating the magnitude e0[n] = |z[n]|; performing a digital low-pass filter operation with a cutoff frequency of 10Hz to 20Hz on the magnitude e0[n] to obtain the speech amplitude envelope E[n] = LPFfLPF{e0[n]}, where LPFfLPF represents performing a digital low-pass filter operation with a cutoff frequency of fLPF on the digital sequence within the parentheses, and fLPF is different from the binaural beat carrier frequency fc.

[0008] Preferably, the present invention further includes: the envelope dominant modulation frequency fp is obtained by: extracting a segment of data Ew[n] of the speech amplitude envelope E[n] within a sliding analysis window; removing the average value of the data Ew[n] to exclude the DC component; performing a fast Fourier transform operation after applying a window function to obtain the envelope modulation spectrum M[l]; finding the energy peak value of the envelope modulation spectrum M[l] in the range of 0.5Hz to 12Hz, and determining the frequency corresponding to the peak value as the envelope dominant modulation frequency fp.

[0009] Preferably, the present invention further includes: the syllable rhythm frequency fs is obtained by estimating it through at least one of the following three methods: detecting local peaks exceeding the peak detection threshold θp in the speech amplitude envelope E[n] and counting the number of peaks Np, and calculating fspeak=Np / Tw based on the number of peaks Np, where Tw refers to the time length of the sliding analysis window in seconds; detecting energy rise points or short-term energy peaks in the effective speech segment after speech activity detection and counting the number Nv, and calculating fsvad=Nv / Tvoice based on the number Nv, where Tvoice refers to the total duration of the effective speech segment in seconds; estimating the number of syllables Ns based on the recognized text when speech recognition is enabled, and calculating fsasr=Ns / Tvoice based on the number of syllables Ns; the short The pause rhythm frequency fq is obtained as follows: detect continuous intervals where the speech amplitude envelope E[n] is lower than the low-energy pause threshold θq and lasts for more than 250ms, record the center time tq[i] of each pause interval, where i represents the ordinal index of the effective pause interval detected in chronological order, calculate the adjacent pause interval ΔTq[i]=tq[i]-tq[i-1], and take the median or mean as the phrase pause period Tq, fq=1 / Tq; the envelope peak period frequency fb is obtained as follows: detect the significant peak time tp[i] of the speech amplitude envelope E[n] within a sliding window, calculate the adjacent peak interval ΔTb[i]=tp[i]-tp[i-1], and take the median or robust mean as the peak period Tb, fb=1 / Tb.

[0010] Preferably, the present invention further includes: the speech rhythm frequency fspeech(t) is calculated by the following weighted fusion method: fspeech(t) = p1·fp + p2·fs + p3·fq + p4·fb, where p1, p2, p3, and p4 are weight coefficients and satisfy p1 + p2 + p3 + p4 = 1; the weight coefficients are selected from one of the following preset weight combinations according to the type of human voice content: under steady-state speech or course lecture type, p1 = 0.40, p2 = 0.20, p3 = 0.10, p4 = 0.30; under relaxed or low-awake speech type, p1 = 0.45, p2 = 0.15, p3 = 0.30, p4 = 0.10; under weak semantic or non-semantic prosodic type, p1 = 0.35, p2 = 0.15, p3 = 0.25, p4 = 0.25.

[0011] Preferably, the present invention further includes: obtaining the mapped frequency fspeechmapped(t) by performing frequency doubling or frequency division mapping based on the speech rhythm frequency fspeech(t), specifically including: multiplying the speech rhythm frequency fspeech(t) by a positive integer h, or dividing the speech rhythm frequency fspeech(t) by a positive integer h, such that the product or quotient falls within a preset range [Fmin, Fmax] of the target neuroacoustic rhythm frequency; when there are multiple positive integers h such that the product or quotient all fall within the preset range [Fmin, Fmax], selecting the product or quotient corresponding to the h that minimizes the absolute value of the difference between the product or quotient and the candidate frequency fstate(t). The quotient is the mapping frequency fspeechmapped(t); the target neuroacoustic rhythm frequency ftarget(t) is calculated as follows: ftarget(t) = Clip[α·fstate(t) + β·fspeechmapped(t), Fmin, Fmax], where Clip represents the range constraint operation, α and β are weighting coefficients, and Fmin and Fmax are the lower and upper limits of the target frequency, respectively; the target usage state includes relaxed or low arousal state, pre-sleep assistance state, steady-state speaking companionship state and focus assistance state, and different target usage states correspond to different candidate frequencies fstate and different frequency ranges [Fmin, Fmax].

[0012] Preferably, the present invention further includes: the composite neuroacoustic signal layer sneu[n] includes at least one of binaural beat signals and amplitude modulation noise or natural sound signals; the binaural beat signals are generated by setting the binaural beat carrier frequency fc, generating the left channel phase sequence ΦL[n]=ΦL[n-1]+2π·fc / Fs and the right channel phase sequence ΦR[n]=ΦR[n-1]+2π·(fc+ftarget[n]) / Fs respectively, and generating the left channel basic binaural beat signal sneu,L[n]=sin(ΦL[n]) and the right channel basic binaural beat signal sneu,R[n]=sin(ΦR[n]); the basic binaural beat signals are processed by the digital gain G of the composite neuroacoustic signal layer during the digital mixing output stage. n[n] is used for gain control, where Fs is the sampling rate; the amplitude modulation noise or natural sound signal is generated in the following way: amplitude modulation noise or natural sound basic signal q[n]=X[n]·[1+m[n]·sin(Φtarget[n])], where X[n] is the pink noise or natural sound digital sampling sequence, m[n] is the modulation depth, and Φtarget[n] is the target phase sequence obtained by accumulating according to the target neuroacoustic rhythm frequency ftarget; the composite neuroacoustic signal layer can be represented as sneu,L,total[n]=sneu,L[n]+k[n]·q[n] and sneu,R,total[n]=sneu,R[n]+k[n]·q[n], where k[n] is the background layer coefficient.

[0013] Preferably, the present invention further includes: dynamically adjusting at least one of the gain, modulation depth, or background layer coefficients of the composite neuroacoustic signal layer sneu[n] based on the speech stage information or sudden event information of the human voice content, specifically including: in the clear semantic stage, setting the digital gain Gn[n] of the composite neuroacoustic signal layer to a first gain value and setting the modulation depth m[n] to a first modulation depth value; in the pause or envelope valley stage, increasing the digital gain Gn[n] of the composite neuroacoustic signal layer from the first gain value to a second gain value and increasing the modulation depth m[n] from the first modulation depth value to a second modulation depth value; When a sudden event is detected, at least one of the following is reduced: digital gain Gn[n] of the composite neuroacoustic signal layer, modulation depth m[n], and background layer coefficient k[n]. The clear semantic stage is defined as a speech activity ratio greater than 70% and an average pause ratio less than 20%. The pause or envelope valley stage is defined as the speech amplitude envelope E[n] being below a threshold and lasting for more than 250ms. The sudden event is triggered by at least one of the following conditions: short-term loudness increases by more than 8dB within 200ms, output peak exceeds -3dBFS, high-frequency energy ratio above 4kHz exceeds 35%, or applause, laughter, or popping sounds are detected.

[0014] This invention also provides a composite neuroacoustic signal generation system driven by human voice rhythm, comprising: an audio input or decoding unit for acquiring a digital audio sampling sequence x[n] containing human voice; a speech rhythm feature extraction unit for extracting speech rhythm features from the digital audio sampling sequence x[n], wherein the speech rhythm features include: speech amplitude envelope E[n], envelope modulation spectrum M[l], envelope dominant modulation frequency fp, syllable rhythm frequency fs, phrase pause rhythm frequency fq, and envelope peak period frequency fb; a speech rhythm frequency calculation unit for calculating the speech rhythm frequency fspeech(t) based on the speech rhythm features; and a target rhythm parameter generation unit for determining a candidate frequency fstate based on a target usage state, and performing frequency doubling or frequency division mapping based on the speech rhythm frequency fspeech(t) to obtain the mapped frequency fspeechm. The system is configured to: 1) `apped(t)`, 2) `fstate`, 3) `fspeechmapped(t)`, 4) `ftarget(t)`, 5) `ftarget(t)`, 6) `ftarget(t)`, 7) `ftarget(t)`, 8) `ftarget(t)`, 9) `ftarget(t)`, 10) `ftarget(t)`, 11) `ftarget(t)`, 12) `ftarget(t)`, 13) `ftarget(t)`, 14) `ftarget(t)`, 15) `ftarget(t)`, 16) `ftarget(t)`, 17) `ftarget(t)`, 18) `ftarget(t)`, 19) `ftarget(t)`, 10) `ftarget(t)`, 18) `ftarget(t)`, 19 ...

[0015] Preferably, the present invention further includes: the dynamic control unit sets the digital gain Gn[n] of the composite neuroacoustic signal layer to a first gain value and the modulation depth m[n] to a first modulation depth value during the clear semantic stage; increases the digital gain Gn[n] of the composite neuroacoustic signal layer from the first gain value to a second gain value and increases the modulation depth m[n] from the first modulation depth value to a second modulation depth value during the pause or envelope valley stage; and reduces at least one of the digital gain Gn[n] of the composite neuroacoustic signal layer, the modulation depth m[n], and the background layer coefficient k[n] when a sudden event is detected; the clear semantic stage is defined as the proportion of speech activity being greater than 70% and the average pause proportion being less than 20%; the pause or envelope valley stage is defined as the speech amplitude envelope E[n] being lower than a threshold and lasting for more than 250ms; and the sudden event is triggered by at least one of the following conditions: the short-term loudness increases by more than 8dB within 200ms, the output peak exceeds -3dBFS, the proportion of high-frequency energy above 4kHz exceeds 35%, and applause, laughter, or popping sounds are detected.

[0016] Compared with the prior art, the advantages of this invention are: First, this invention realizes the transformation of neuroacoustic signal generation from fixed frequency preset to speech rhythm driven. In the prior art, neuroacoustic signals such as binaural beats, monoaural beats, or modulation noise are usually generated at a constant target frequency or preset frequency band, and their temporal structure is independent of the human voice content. This invention extracts various speech rhythm features from the input human voice digital audio sampling sequence x[n], such as speech amplitude envelope E[n], envelope modulation spectrum M[l], envelope dominant modulation frequency fp, syllable rhythm frequency fs, phrase pause rhythm frequency fq, and envelope peak period frequency fb, and obtains the speech rhythm frequency fspeech(t) through weighted fusion calculation. Then, the target neuroacoustic rhythm frequency ftarget(t) is determined according to the target usage state and the frequency overlay or frequency division mapping result of fspeech(t), so that the generated composite neuroacoustic signal layer sneu[n] matches the inherent rhythm of the human voice content in terms of temporal structure, avoiding the problems of rhythm misalignment and abrupt listening experience.

[0017] Secondly, this invention ensures the accuracy and controllability of signal fusion through digital sampling-level mixing output. The composite neuroacoustic signal layer sneu[n] is itself a digital audio sampling sequence. It is weighted and summed with the original human voice digital audio sampling sequence x[n] in the digital domain and then output as y[n], which is finally transduced into an air pressure wave p(t) by headphones or speakers. This processing method allows for control of the relative relationship between the human voice channel gain Gv[n] and the neuroacoustic signal layer gain Gn[n] at each sampling point. The composite neuroacoustic signal layer and the human voice signal are weighted and mixed at the same digital audio processing clock and audio buffer, which allows for unified control of the gain relationship and time correspondence between the two types of signals.

[0018] Third, this invention effectively balances the clarity of semantic transmission with the continuity of rhythmic assistance by dynamically adjusting the gain, modulation depth, or background layer coefficient of the composite neuroacoustic signal layer based on the speech phase information and sudden event information of the human voice content. In the clear semantic phase, the system sets the digital gain Gn[n] of the neuroacoustic signal layer to a lower first gain value and the modulation depth m[n] to a lower first modulation depth value, thereby reducing the masking of speech intelligibility; in the pause or envelope valley phase, the system increases Gn[n] from the first gain value to the second gain value and m[n] from the first modulation depth value to the second modulation depth value, using the rhythmic gaps of the human voice content to enhance the intensity of the auxiliary signal delivery; when a sudden increase in loudness, excessively high output peak, abnormal high-frequency energy ratio, or sudden events such as applause, laughter, or popping sounds are detected, the system promptly reduces at least one of Gn[n], m[n], and the background layer coefficient k[n], reducing abruptness and auditory discomfort.

[0019] Fourth, this invention can be flexibly deployed on various audio terminal devices. Since the entire processing flow takes digital audio sampling sequences as the operating object, the calculation steps can be executed by general-purpose processors, digital signal processors, audio processing chips, mobile terminal applications, or embedded audio processing units. Therefore, the system can be integrated into devices such as mobile phone apps, tablets, smart headphones, open-back headphones, DSP chips, smart speakers, learning terminals, or sleep aid terminals, and has good platform adaptability and wide application range. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall data flow of the present invention; Figure 2 This is a schematic diagram of speech rhythm feature extraction and speech rhythm frequency calculation according to the present invention; Figure 3 This is a schematic diagram illustrating the generation of the target neuroacoustic rhythm frequency in this invention. Figure 4This is a schematic diagram of the composite neuroacoustic signal generation and digital mixing output of the present invention; Figure 5 This is a schematic diagram illustrating the workflow of an online course and teacher voice-assisted listening embodiment of the present invention. Detailed Implementation

[0021] To make the purpose, principle and structure of the present invention clearer, the following description is provided in conjunction with the accompanying drawings and specific embodiments.

[0022] This invention provides a method and system for generating complex neuroacoustic signals driven by human voice rhythm. The invention will be further described in detail below with reference to the accompanying drawings. Specific embodiments of this invention can be performed according to the following steps, but the scope of protection of this invention is not limited thereto.

[0023] First, the system acquires a digital audio sampling sequence x[n] containing human voice. This sequence can originate from local audio files, network audio streams, microphone captures, or pulse-code modulation data obtained via application or Bluetooth decoding. The system decodes the content containing human voice into a unified digital audio sampling sequence x[n]. The sampling rate can be set to 44.1kHz or 48kHz, preferably 48kHz; the data format can be 24-bit PCM or 32-bit floating-point PCM. The system processes audio in 20ms frames. If the sampling rate Fs = 48000Hz, each frame contains 960 sampling points. The system uses a sliding analysis window of 10 to 30 seconds, preferably a 15-second window, and updates the control parameters every 1 to 5 seconds, preferably every 3 seconds.

[0024] Next, the system extracts speech rhythm features from the digital audio sampling sequence x[n]. Extraction of the speech amplitude envelope E[n] is the foundation for all subsequent rhythm analysis. The system first performs a Hilbert transform on x[n] to obtain the Hilbert component xH[n], which is orthogonal to x[n]. The Hilbert transform is a commonly used orthogonal transform method in digital signal processing. For a real-valued sequence x[n], its Hilbert transform xH[n] provides a component that is 90 degrees out of phase with the original sequence. The system constructs an analytic signal z[n] = x[n] + j·xH[n], where j represents the imaginary unit, and then calculates the magnitude of this analytic signal e0[n] = |z[n]| = sqrt(x[n]). 2 +xH[n] 2The system obtains the original amplitude envelope estimate. Since the original envelope estimate still contains high-frequency fluctuation components, the system performs a digital low-pass filtering operation on e0[n] to obtain a smooth speech amplitude envelope E[n]=LPFfLPF{e0[n]}, where LPFfLPF represents the digital low-pass filtering operation with a cutoff frequency of fLPF performed on the digital sequence within the parentheses. fLPF can be 10Hz to 20Hz, preferably 15Hz; fLPF is different from the binaural beat carrier frequency fc described later. The purpose of this low-pass filtering operation is to filter out fast fluctuation components above the cutoff frequency and retain only the low-frequency components related to the slowly changing envelope of speech amplitude, thereby obtaining an envelope signal that reflects the slow change of speech energy over time. This operation can be implemented by existing digital filtering algorithms such as finite-length unit impulse response filters, infinite-length unit impulse response filters, or moving averages. It does not mean that it must be set as an independent hardware component, but rather that it is executed as a digital signal processing step in a general-purpose processor, digital signal processor, or audio processing chip.

[0025] The system further extracts multiple rhythmic features from the speech amplitude envelope E[n]. The dominant modulation frequency fp of the envelope is obtained as follows: a segment of data Ew[n] is extracted from E[n] within a sliding analysis window. The average value of this segment is removed to eliminate the DC component. Then, a window function is applied, and a Fast Fourier Transform (FFT) operation is performed to obtain the envelope modulation spectrum M[l] = FFT{w[n]·(Ew[n]-mean(Ew[n]))}. The purpose of removing the DC component is to eliminate the interference of the envelope average energy or overall loudness bias on the spectrum analysis, because these bias components do not represent the rhythmic information of the speech. The system searches for the energy peak of M[l] in the range of 0.5Hz to 12Hz, and determines the frequency corresponding to this peak as the dominant modulation frequency fp of the envelope. This frequency reflects the most important periodic fluctuation speed in the speech amplitude envelope and is an important representation of the overall rhythm of the speech.

[0026] The syllable rhythm frequency fs is obtained by estimation using at least one of the following three methods.

[0027] The first method involves detecting local peaks exceeding the peak detection threshold θp within the speech amplitude envelope E[n] and counting the number of peaks Np. Based on the number of peaks Np, fspeak = Np / Tw is calculated, where Tw represents the duration of the sliding analysis window in seconds. The peak detection threshold θp is typically obtained by multiplying the maximum or average amplitude of the speech amplitude envelope E[n] within the current sliding window by a proportional coefficient, for example, set to 30% to 50% of the maximum envelope value within the window. Only local peaks exceeding this proportion are considered valid "syllable peaks," thus avoiding misinterpreting minute background noise fluctuations as syllables.

[0028] The second method is to detect energy rise points or short-term energy peaks in the effective speech segments after speech activity detection and count the number Nv. Based on the number Nv, calculate fsvad=Nv / Tvoice, where Tvoice represents the total duration of the effective speech segments in seconds.

[0029] The third approach: When speech recognition is enabled, the number of syllables Ns is estimated based on the recognized text, and fsasr = Ns / Tvoice is calculated based on the number Ns. The system can choose one of the above approaches, or it can fuse them according to confidence level. That is, in the main implementation, the system preferably uses fs = fspeak = Np / Tw as the syllable rhythm frequency; speech activity peak estimation fsvad or speech recognition syllable number estimation fsasr can be used as alternative implementations. Syllable rhythm frequency reflects the rate at which syllables appear in speech and is the basic unit of speech time structure.

[0030] The phrase pause rhythm frequency fq is obtained as follows: The system detects continuous intervals where the speech amplitude envelope E[n] is lower than the low-energy pause threshold θq and lasts for more than 250ms, and records the center time tq[i] of each pause interval, where i represents the ordinal index of the effective pause interval detected in chronological order. The low-energy pause threshold θq is usually obtained by multiplying the estimated noise floor value or the minimum envelope value in the current window by a relaxation coefficient, for example, set to a level slightly higher than the noise floor of silence, such as 5% to 10% of the maximum envelope value, to accurately identify true speech pauses or envelope valleys, rather than normal short consonant gaps during speech. The system calculates the adjacent pause interval ΔTq[i] = tq[i] - tq[i-1], where when i takes the value of 2, 3, 4... up to the total number of pauses detected N, the time interval between each adjacent pause is calculated sequentially, and finally the median or mean of all ΔTq[i] is taken as the phrase pause period Tq, fq = 1 / Tq. The frequency of phrase pauses reflects the temporal structure of phrases in speech, that is, the regularity of pauses between sentences.

[0031] The envelope peak period frequency fb is obtained as follows: the system detects the significant peak time tp[i] of the speech amplitude envelope E[n] within a sliding window, calculates the adjacent peak interval ΔTb[i] = tp[i] - tp[i-1], and takes the median or robust mean as the peak period Tb, fb = 1 / Tb. This frequency reflects the periodicity of the envelope peak occurrence, is related to the syllable rate but focuses more on the periodic recurrence of energy peaks.

[0032] The system calculates the speech rhythm frequency fspeech(t) based on the aforementioned speech rhythm characteristics. Specifically, a weighted fusion method is used: fspeech(t) = p1·fp + p2·fs + p3·fq + p4·fb, where p1, p2, p3, and p4 are weight coefficients and satisfy p1 + p2 + p3 + p4 = 1. The weight coefficients are selected from preset weight combinations based on the type of speech content. In steady-state speech or lecture-style speech, p1 = 0.40, p2 = 0.20, p3 = 0.10, and p4 = 0.30; this combination considers both the overall envelope and the energy peak period. In relaxed or low-awake speech, p1 = 0.45, p2 = 0.15, p3 = 0.30, and p4 = 0.10; this combination increases the weight of phrase pauses and decreases the weight of syllable details. In weak speech activity or non-semantic prosodic types, p1=0.35, p2=0.15, p3=0.25, and p4=0.25. This combination reduces the weight of syllable details and increases the weight of pauses and peak periods. Through the above weighted fusion, the system integrates multiple speech prosodic features into a single frequency parameter fspeech(t) that can reflect the overall temporal structure of human voice.

[0033] The system determines the candidate frequency fstate based on the target usage state and performs frequency doubling or subdivision mapping based on the speech rhythm frequency fspeech(t) to obtain the mapped frequency fspeechmapped(t). Specifically, the system multiplies or divides fspeech(t) by a positive integer h, such that the product or quotient falls within a preset range [Fmin, Fmax] of the target neuroacoustic rhythm frequency. When multiple positive integers h exist such that the product or quotient all fall within the preset range [Fmin, Fmax], the system selects the product or quotient corresponding to the h that minimizes the absolute value of the difference between the product or quotient and the candidate frequency fstate(t) as the mapped frequency fspeechmapped(t). The purpose of frequency doubling or subdivision mapping is to adjust the natural rhythm frequency of the human voice to the frequency band of the target neuroacoustic rhythm frequency, so that the auxiliary signal maintains temporal structural coordination with the human voice rhythm and can function in the target frequency band.

[0034] The system calculates the target neuroacoustic rhythm frequency ftarget(t) based on the candidate frequency fstate and the mapped frequency fspeechmapped(t): ftarget(t) = Clip[α·fstate(t) + β·fspeechmapped(t), Fmin, Fmax], where Clip represents the range constraint operation, α and β are weighting coefficients, and Fmin and Fmax are the lower and upper limits of the target frequency, respectively. The system first calculates fraw(t) = α·fstate(t) + β·fspeechmapped(t), and then executes ftarget(t) = min(max(fraw(t), Fmin), Fmax). The target usage states include relaxed or low arousal state, pre-sleep assistance state, steady-state speech companionship state, and focused assistance state. Different target usage states correspond to different candidate frequencies fstate and different frequency ranges [Fmin, Fmax]. For example, in a relaxed or low-wake state, fstate=6Hz, Fmin=3Hz, Fmax=8Hz; in a sleep aid state, fstate=4Hz, Fmin=2Hz, Fmax=6Hz; in a steady-state speaking companion state, fstate=5Hz, Fmin=3Hz, Fmax=7Hz; and in a focus aid state, fstate=14Hz, Fmin=12Hz, Fmax=18Hz. Let r be the control parameter update cycle number. The system calculates fspeech[r] and ftarget[r] in the r-th sliding analysis window. Within this control cycle, the system keeps the target frequency unchanged, or smoothly changes ftarget[r-1] of the previous cycle to ftarget[r] of the current cycle through a smooth transition function, thus obtaining the target frequency ftarget[n] for each sampling point. From the source perspective, ftarget(t) is the instantaneous target frequency obtained by weighted fusion of the candidate frequency fstate(t) corresponding to the target usage state and the speech rhythm frequency fspeech(t) extracted from human speech after frequency doubling or frequency division mapping, and then by range constraint Clip operation.

[0035] The system generates a composite neuroacoustic signal layer sneu[n] based on the target neuroacoustic rhythm frequency ftarget(t). This signal layer is a digital audio sampling sequence. The composite neuroacoustic signal layer includes at least one of binaural beat signals and amplitude-modulated noise or natural sound signals. The binaural beat signals are generated as follows: a binaural beat carrier frequency fc is set, and the left channel phase sequence ΦL[n]=ΦL[n-1]+2π·fc / Fs and the right channel phase sequence ΦR[n]=ΦR[n-1]+2π·(fc+ftarget[n]) / Fs are generated respectively. The basic binaural beat signals sneu,L[n]=sin(ΦL[n]) and sneu,R[n]=sin(ΦR[n]) are generated in the left channel and respectively. The basic binaural beat signals are then gain-controlled by the digital gain Gn[n] of the composite neuroacoustic signal layer during the digital mixing output stage, where Fs is the sampling rate. ftarget[n] is the instantaneous value of ftarget(t) at the discrete-time sampling point n, i.e., ftarget[n] = ftarget(n / Fs). The carrier frequency fc can be selected in the range of 160Hz to 440Hz, preferably 180Hz to 300Hz, and 220Hz can be used in the embodiment. Binaural beats require a certain degree of channel isolation between the left and right ears. Therefore, when the output terminal is a headphone, open-back headphone, ambient sound pass-through headphone, or near-ear device, the system can generate a binaural beat signal. The amplitude modulation noise or natural sound signal is generated in the following way: q[n] = X[n]·[1+m[n]·sin(Φtarget[n])], where X[n] is the pink noise or natural sound digital sampling sequence, m[n] is the modulation depth, and Φtarget[n] is the target phase sequence obtained by accumulating according to the target neuroacoustic rhythm frequency ftarget. The system can combine binaural beat signals with modulated noise or natural sound signals into composite signals: sneu,L,total[n]=sneu,L[n]+k[n]·q[n] and sneu,R,total[n]=sneu,R[n]+k[n]·q[n], where k[n] is a background layer coefficient used to control the proportion of modulated noise or natural sound in the composite layer. During the clear semantic stage, k[n] can be taken as 0.05 to 0.10 to avoid noise masking the human voice; it can be appropriately increased during low speech activity or pauses.

[0036] The system dynamically adjusts at least one of the gain, modulation depth, or background layer coefficients of the composite neuroacoustic signal layer sneu[n] based on the speech phase information or sudden event information of the human voice content. In the clear semantic phase, the system sets the digital gain Gn[n] of the composite neuroacoustic signal layer to a first gain value and the modulation depth m[n] to a first modulation depth value. The clear semantic phase is defined as a speech activity ratio greater than 70% and an average pause ratio less than 20%. In the pause or envelope trough phase, the system increases Gn[n] from the first gain value to a second gain value and increases m[n] from the first modulation depth value to a second modulation depth value. The pause or envelope trough phase is defined as a speech amplitude envelope E[n] below a threshold for more than 250ms. When a sudden event is detected, the system reduces at least one of Gn[n], m[n], and the background layer coefficient k[n]. Sudden events are triggered by at least one of the following conditions: a short-term loudness increase of more than 8dB within 200ms, an output peak exceeding -3dBFS, a high-frequency energy ratio above 4kHz exceeding 35%, or the detection of applause, laughter, or popping sounds. "Dynamic" in dynamic adjustment refers to the control module recalculating and writing digital processing parameters according to a preset update cycle Tu during operation, rather than manual knob adjustment or mechanical structural changes. In software applications, dynamic adjustment manifests as updating variables in the audio processing thread or real-time audio callback function; in digital signal processor chips, it manifests as updating registers, coefficient tables, frequency control words, gain multiplier coefficients, or filter coefficients. The sources of Gn[n] include output terminal type, target usage state, pause / valley state, high wake-up event, optional user feedback, and user intensity settings. The sources of m[n] include basic modulation depth mbase, speech stage coefficients, pause enhancement coefficients, high wake-up suppression coefficients, and optional feedback coefficients. The sources of k[n] include terminal type, target state, speech stage, and user feedback mapping.

[0037] The system performs digital sampling-level mixing of the composite neuroacoustic signal layer sneu[n] with the digital audio sampling sequence x[n], and outputs the mixed digital audio sequence y[n].

[0038] For dual-channel headphone output, the mixing method is as follows: uL[n]=Gv[n]·xL[n]+sneu,L,total[n]; uR[n]=Gv[n]·xR[n]+sneu,R,total[n]; yL[n]=Limiter{uL[n]}; yR[n]=Limiter{uR[n].

[0039] Where Gv[n] is the digital gain of the human voice channel, the sources of which include the root mean square value or short-time loudness of the human voice in the current window, the target human voice loudness, the limiting protection status, and the user's volume setting. Limiter indicates that the digital sequence within the brackets is subjected to limiting processing to avoid digital clipping caused by the output sample value exceeding the preset amplitude threshold. The output digital audio sequence y[n] is sent to the audio buffer, and output as an air pressure wave p(t) through a digital-to-analog converter or Bluetooth audio module, power amplifier, and headphone or speaker transducer. Digital sampling level mixing allows control of the relative relationship between the human voice channel gain Gv[n] and the neuroacoustic signal layer gain Gn[n] at each sampling point with precision. The composite neuroacoustic signal layer and the human voice signal are sample-level weighted mixing completed within the same digital audio processing clock and audio buffer, which allows for unified control of the gain relationship and time correspondence between the two types of signals.

[0040] The method of the present invention can be executed by a general-purpose processor, a digital signal processor, an audio processing chip, a mobile terminal application, a real-time audio processing thread, or an embedded audio processing unit. The system can be integrated into devices such as mobile applications, tablet computers, smart headphones, open-back headphones, digital signal processor chips, smart speakers, learning terminals, or sleep aid terminals.

[0041] The core technical solution of this invention can be summarized as the following data flow: x[n]->E[n]->fp,fs,fq,fb->fspeech(t)->ftarget(t)->sneu[n]->y[n]->p(t). Wherein, x[n] is the input human voice digital audio sampling sequence; E[n] is the speech amplitude envelope; fp, fs, fq, and fb are the envelope dominant modulation frequency, syllable rhythm frequency, phrase pause rhythm frequency, and envelope peak period frequency, respectively; fspeech(t) is the speech rhythm frequency; ftarget(t) is the target neuroacoustic rhythm frequency; sneu[n] is the composite neuroacoustic signal layer; y[n] is the digital mixed output sequence; and p(t) is the airborne sound pressure wave.

[0042] In this invention, LPF{·}, FFT{·}, Clip[·], Gate(·), Limiter{·}, Hilbert(·), NCO, etc., in the formulas all represent digital signal processing operations or algorithm steps, and do not imply that they must be set as independent hardware components. The above operations can be executed by a general-purpose processor, digital signal processor, audio processing chip, mobile terminal application, real-time audio processing thread, or embedded audio processing unit.

[0043] In this invention, the main parameters are defined as follows: the input human voice digital audio sampling sequence x[n] originates from local files, network audio streams, microphone acquisition, or pulse code modulation data decoded by applications and Bluetooth, and belongs to digital sampling sequences; the speech amplitude envelope E[n] is obtained by performing envelope calculation and low-pass filtering operations on x[n], and belongs to digital control and analysis sequences; the envelope modulation spectrum M[l] is the frequency domain sequence obtained by performing fast Fourier transform operations on E[n] within the window; the envelope dominant modulation frequency fp is the frequency corresponding to the main peak of M[l] in the candidate frequency band, in Hz; the syllable rhythm frequency fs is determined by the envelope peak density, speech activity peaks, or number of syllables. The quantity is estimated and the unit is Hz; the phrase pause rhythm frequency fq is calculated from the low-energy pause interval lasting more than 250ms or 300ms and the unit is Hz; the envelope peak periodic frequency fb is obtained by taking the reciprocal of the significant envelope peak interval and the unit is Hz; the speech rhythm frequency fspeech(t) is obtained by weighted fusion of fp, fs, fq, and fb and the unit is Hz; the target neuroacoustic rhythm frequency ftarget(t) is calculated from the target state frequency and the speech rhythm mapping frequency and the unit is Hz; the composite neuroacoustic signal layer sneu[n] is a digital audio sampling sequence generated by digital sine sequence, modulation noise, or natural sound. The digital gain Gv[n] of the human voice channel is determined by the loudness of the human voice, the loudness of the target voice, and the amplitude limiting protection, and belongs to the digital multiplication coefficient; the digital gain Gn[n] of the neuroacoustic signal layer is determined by the output terminal, the target state, the pause or valley value, the sudden event, and the feedback state, and also belongs to the digital multiplication coefficient; the modulation depth m[n] is used to control the amplitude modulation intensity of natural sound or noise, and is a dimensionless digital parameter; the background layer coefficient k[n] is used to control the proportion of modulation noise or natural sound in the composite layer, and belongs to the digital gain coefficient; the mixed output digital audio sequence y[n] is obtained by weighted summation and amplitude limiting of x[n] and sneu[n] through sampling stage; the final air sound pressure signal p(t) is the final physical sound wave output by the transducer of the headphones or speaker.

[0044] The specific operation of speech rhythm feature extraction is as follows: For input data x[n], the system decodes the human voice content into a unified PCM digital audio sampling sequence x[n]. The sampling rate can be 44.1kHz or 48kHz, preferably 48kHz; the data format can be 24bitPCM or 32bit floating-point PCM. The system processes audio in 20ms frames. If Fs=48000Hz, each frame contains 960 sampling points; the system uses a sliding analysis window of 10 to 30 seconds, preferably a 15-second window, and updates the control parameters every 1 to 5 seconds, preferably every 3 seconds.

[0045] Hilbert envelope and low-pass filtering operations: The Hilbert envelope is a commonly used envelope extraction method in digital signal processing, but it is not a concept defined in this invention. For a real-valued digital audio sampling sequence x[n], the system first performs a Hilbert transform operation on x[n] to obtain the Hilbert component xH[n] orthogonal to x[n], and then constructs an analytic signal: z[n] = x[n] + j·xH[n]; the system calculates the magnitude of the analytic signal to obtain the original amplitude envelope estimate: e0[n] = |z[n]| = sqrt(x[n]) 2 +xH[n] 2Subsequently, the system performs a digital low-pass filtering operation on e0[n] to obtain the speech amplitude envelope E[n] = LPFfLPF{e0[n]}; where LPFfLPF{·} represents performing a digital low-pass filtering operation on the digital sequence within the parentheses with a cutoff frequency of fLPF, where fLPF can be taken from 10Hz to 20Hz, preferably 15Hz. This operation can be implemented by existing digital filtering algorithms such as FIR, IIR, or moving average, and does not represent an independent hardware component. Envelope modulation spectrum M[l] and fp: The system extracts a segment of data Ew[n] from the speech amplitude envelope E[n] within the sliding analysis window, first removes the average value to exclude the DC component, then applies a window function w[n] and performs a fast Fourier transform operation: E0[n]=Ew[n]-mean(Ew[n]); M[l]=FFT{w[n]·E0[n]}; where DC represents the zero-frequency component or the average value component, which corresponds to the envelope average energy or overall loudness bias, and does not represent the speech rhythm. The system searches for the energy peak of M[l] in the range of 0.5Hz to 8Hz or 0.5Hz to 12Hz, and determines the frequency corresponding to this peak as the envelope dominant modulation frequency fp. Calculation of fs, fq, and fb: fs represents the syllable rhythm frequency, in Hz. The frequency of syllable rhythm can be estimated in three ways: First, detect local peaks exceeding the threshold θp in E[n] and count the number of peaks Np, fspeak=Np / Tw; Second, detect energy rise points or short-term energy peaks in the effective speech segments after speech activity detection, fsvad=Nv / Tvoice; Third, when speech recognition is enabled, estimate the number of syllables Ns based on the recognized text, fsasr=Ns / Tvoice. The system can choose one of these methods or fuse them according to confidence. In the main implementation, the system preferably uses fs=fspeak=Np / Tw as the syllable rhythm frequency; speech activity peak estimation fsvad or speech recognition syllable number estimation fsasr can be used as alternative implementation methods. fq represents the phrase pause rhythm frequency, in Hz. The system detects continuous intervals where E[n] is below the low energy threshold θq and lasts for more than 250ms or 300ms. It records the center time tq[i] of each pause interval, calculates the adjacent pause interval ΔTq[i] = tq[i] - tq[i-1], and takes the median or mean as the phrase pause period Tq, where fq = 1 / Tq. fb represents the envelope peak period frequency in Hz. The system detects the significant peak time tp[i] of E[n] within a sliding window, calculates the adjacent peak interval ΔTb[i] = tp[i] - tp[i-1], and takes the median or robust mean as the peak period Tb, where fb = 1 / Tb.

[0046] The calculation of speech rhythm frequency and target frequency is as follows: Speech rhythm frequency fspeech(t); The system calculates the speech rhythm frequency through weighted fusion: fspeech(t) = p1·fp + p2·fs + p3·fq + p4·fb; where p1, p2, p3, and p4 are weight coefficients, satisfying p1 + p2 + p3 + p4 = 1. To avoid unclear weight selection, this invention adopts a preset weight relationship defined manually according to requirements as the main implementation method.

[0047] When calculating speech rhythm frequencies, the system uses a preset weight combination based on different types of human speech content. For steady-state speech or lecture types, the weights p1, p2, p3, and p4 of fp, fs, fq, and fb are set to 0.40, 0.20, 0.10, and 0.30, respectively. This combination takes into account both the overall envelope and the peak period of energy. For relaxed or low-arousal speech types, the weights are set to 0.45, 0.15, 0.30, and 0.10, respectively. This combination increases the weight of phrase pauses while decreasing the weight of syllable details. For focused auxiliary course speech, the weight settings are consistent with the online course implementation example, at 0.40, 0.20, 0.10, and 0.30. For weak speech activity or non-semantic prosodic types, the weights are set to 0.35, 0.15, 0.25, and 0.25, respectively. This combination decreases the weight of syllable details while increasing the weight of pauses and peak periods. Through the different weight allocations mentioned above, the system can adaptively adjust the calculation emphasis of speech rhythm frequency according to the rhythmic characteristics of human speech content.

[0048] Target frequency ftarget(t): The system first determines the candidate frequency fstate(t) based on the target usage state, then performs frequency doubling or subdivision mapping based on the speech rhythm frequency to obtain fspeechmapped(t), and finally calculates the target frequency. ftarget(t)=Clip[α·fstate(t)+β·fspeechmapped(t),Fmin,Fmax]; where Clip[·] represents the range constraint operation. The system first calculates fraw(t)=α·fstate(t)+β·fspeechmapped(t), and then executes ftarget(t)=min(max(fraw(t),Fmin),Fmax).

[0049] In calculating the target neuroacoustic rhythm frequency, the system determines the corresponding candidate frequency fstate and frequency range constraints based on different target usage states. Specifically, for a relaxed or low-arousal state, the candidate frequency fstate is set to 6Hz, with a frequency range constraint of 3Hz to 8Hz. In this state, it can be fused with speech rhythms or their sub-frequency or octave frequencies. For a sleep aid state, the candidate frequency fstate is set to 4Hz, with a frequency range constraint of 2Hz to 6Hz. This state emphasizes low-frequency, low-intensity output and reduces abruptness. For a steady-state speech accompaniment state, the candidate frequency fstate is set to 5Hz, with a frequency range constraint of 3Hz to 7Hz. This state is suitable for podcasts, news broadcasts, or reading aloud. For a focus-assisted state, the candidate frequency fstate is set to 14Hz, with a frequency range constraint of 12Hz to 18Hz. This state is not fixed at a single 12Hz but rather preferably uses a sensorimotor rhythm or a nearby range of the low β frequency band. Through the above selection of candidate frequencies and frequency ranges based on different usage states, the system can flexibly adapt to the rhythm assistance needs of various application scenarios.

[0050] In an optional implementation, if a feedback sensor is connected, the following extended formula can be used: ftarget(t)=Clip[α·fstate(t)+β·fspeechmapped(t)+γ·ffeedback(t)+δ·fresp(t),Fmin,Fmax]; when no feedback or breathing sensor is connected, let γ=0 and δ=0, and they will not participate in the target frequency calculation.

[0051] The specific computational method for generating the composite neuroacoustic signal layer is as follows: the composite neuroacoustic signal layer is not an abstract concept, but rather one or more digital audio sampling sequences sneu[n]. This invention uses binaural beat signals and amplitude-modulated noise / natural sound signals as the core implementation method; monoaural beats and isochronous phonological sequences can be used as optional implementation methods.

[0052] Binaural beat signal: When the output terminal is a headphone, open-back headphone, ambient sound pass-through headphone, or near-ear device, the system can generate a low-intensity binaural beat signal. Binaural beat requires a certain degree of channel isolation between the left and right ears; therefore, ordinary external speakers are not the preferred terminal.

[0053] Left channel phase sequence ΦL[n] = ΦL[n-1] + 2π·fc / Fs; Right channel phase sequence ΦR[n]=ΦR[n-1]+2π·(fc+ftarget[n]) / Fs; The basic binocular beat signal for the left channel is sneu,L[n] = sin(ΦL[n]); The basic binaural beat signal for the right channel is sneu,R[n]=sin(ΦR[n]); Here, fc is the binaural beat carrier frequency, not the target neuroacoustic frequency. The target frequency is determined by the frequency difference fR-fL between the left and right vocal tracts. fc can be selected in the range of 160Hz to 440Hz, preferably 180Hz to 300Hz; in this embodiment, 220Hz can be used.

[0054] Amplitude-modulated noise or natural sound: The system can generate low-intensity pink noise, natural sound, or ambient sound modulation layers to mask the abruptness of pure tone beats.

[0055] q[n]=X[n]·[1+m[n]·sin(Φtarget[n])]; where X[n] is a digital sampling sequence of pink noise, rain sound, wind sound, water sound or ambient sound; m[n] is the modulation depth, usually taken as 0 to 0.20; φtarget[n] is the target phase sequence obtained by accumulating according to the target frequency ftarget.

[0056] sneu,L,total[n]=sneu,L[n]+k[n]·q[n]; sneu,R,total[n]=sneu,R[n]+k[n]·q[n]; k[n] is the background layer coefficient, which is usually taken as 0 to 1. It can be taken as 0.05 to 0.10 in the clear semantic stage to avoid noise masking the human voice; it can be appropriately increased in the low speech activity or pause stage.

[0057] Monoacoustic beats and isochronous sequencing are available: When the output terminal is a regular speaker, smart speaker, pillow speaker, or a device with insufficient isolation between the left and right channels, monoacoustic beats or modulated noise / natural sound can be used, but binaural beats are not preferred. Monoacoustic beats can generate two similar frequency signals in the same channel, making their difference frequency correspond to the target frequency; isochronous sequencing can control the amplitude envelope of carrier sound, natural sound, or noise through gating sequences.

[0058] The specific calculation method for dynamically adjusting and mixing digital output is as follows.

[0059] The meaning of dynamic adjustment: In this invention, dynamic adjustment refers to the control module recalculating and writing digital processing parameters according to a preset update cycle Tu during operation, rather than manual knob adjustment or mechanical structure changes. In software applications, dynamic adjustment is manifested in updating variables in the audio processing thread or real-time audio callback function; in DSP chips, it is manifested in updating registers, coefficient tables, frequency control words, gain multiplier coefficients, or filter coefficients.

[0060] The sources of Gv[n], Gn[n], m[n], and k[n]: Gv[n] is the digital gain of the human voice channel, which is derived from the current window's RMS or short-time loudness of the human voice, the target human voice loudness, the limiting protection status, and the user's volume settings. The system can first calculate Gvbase=10. (Ltarget-Lvoice) / 20 The exponent in the upper right corner of 10 represents the exponential operation with base 10, which is used to convert the decibel difference into a linear amplitude gain, and then obtain Gv[n] through range constraints.

[0061] Gn[n] represents the digital gain of the complex neuroacoustic signal layer, derived from factors including output terminal type, target usage status, pause / valley status, high arousal events, optional user feedback, and user intensity settings. The recommended calculation method is as follows: Gnraw[n]=Gbase·Cstage[n]·Cgap[n]·Carousal[n]·Cfeedback[n]; Gn[n]=ClipGn{Gnraw[n],Gnmin,Gnmax}; m[n] represents the modulation depth, which is derived from the base modulation depth mbase, speech stage coefficients, pause enhancement coefficients, high wake-up suppression coefficients, and optional feedback coefficients.

[0062] The recommended calculation method is as follows: mraw[n]=mbase·Mstage[n]·Mgap[n]·Marousal[n]·Mfeedback[n]; m[n]=Clipm{mraw[n],mmin,mmax}.

[0063] k[n] is the background layer mixing coefficient, used to control the proportion of modulation noise or natural sound in the composite layer. k[n] can be obtained by mapping terminal type, target state, voice stage and user feedback, or it can be set directly in the embodiment.

[0064] Simplified Speech Stages and Sudden Event Control Rules: To avoid the pressure of public disclosure caused by complex semantic scoring models, this invention recommends using simple speech state judgment as the primary implementation method. A clear semantic stage can be defined as: speech activity accounting for more than 70%, average pause ratio less than 20%, short-term loudness stable, and no obvious sudden events. A pause / low energy stage can be defined as: E[n] below the low energy threshold and lasting for more than 250ms. A weak semantic / non-semantic stage can be defined as: speech activity accounting for less than 50%, or the system marking it as background / low-intelligibility content.

[0065] High wake-up events are recommended to be triggered using a simple threshold: a short-term loudness increase of more than 8dB within 200ms, or an output peak exceeding -3dBFS, or a high-frequency energy ratio above 4kHz exceeding 35%, or the detection of applause, laughter, or popping sounds; any condition meeting these conditions will trigger high wake-up event processing. The processing method involves reducing Gn[n], m[n], and k[n], and limiting or dynamic compression can be enabled.

[0066] Digital mixing output: The system mixes the human voice with a complex neuroacoustic signal layer at the digital sampling level. For dual-channel headphone output: uL[n]=Gv[n]·xL[n]+sneu,L,total[n]; uR[n]=Gv[n]·xR[n]+sneu,R,total[n]; yL[n]=Limiter{uL[n]},yR[n]=Limiter{uR[n]}.

[0067] Limiter{·} indicates that the digital sequence within the brackets is subjected to amplitude limiting to prevent the output sample value from exceeding the preset amplitude threshold and causing digital clipping. The output digital audio sequence y[n] is fed into the audio buffer, and output as an airborne sound pressure wave p(t) through the DAC or Bluetooth audio module, power amplifier, and headphone / speaker transducer.

[0068] The system is specifically composed as follows: The system of the present invention may include the following functional units: audio input or decoding unit, audio frame buffer unit, speech rhythm feature extraction unit, speech rhythm frequency calculation unit, target rhythm parameter generation unit, composite neuroacoustic signal generation unit, dynamic control unit, digital mixing output unit, optional user feedback input unit, DAC or Bluetooth audio output unit, and headphone or speaker transducer unit.

[0069] The aforementioned functional units can be deployed in mobile apps, tablets, smart headphones, open-back headphones, DSP chips, smart speakers, learning terminals, sleep aid terminals, or neuroacoustic research equipment.

[0070] Example 1: Steady-state speech low-wake-up assistance scenario.

[0071] This embodiment is applicable to podcasts, news broadcasts, readings, special lectures, conference speeches, or steady-state speeches. The system decodes the input audio into a digital sampling sequence x[n] of Fs=48000Hz, 24-bit PCM or 32-bit floating-point PCM, with a frame size of 20ms, a sliding window size of 15 seconds, and updates the control parameters every 3 seconds.

[0072] Within a 15-second window, the system extracts the following frequencies: fp=3.20Hz, fs=4.10Hz, fq=0.80Hz, fb=2.60Hz. For the steady-state speech clarity semantic stage, p1=0.45, p2=0.15, p3=0.30, p4=0.10. fspeech=0.45∗3.20+0.15∗4.10+0.30∗0.80+0.10∗2.60=2.555Hz. Assuming low-wake-up candidate frequencies fstate=6.00Hz, Fmin=3Hz, Fmax=8Hz, α=0.50, β=0.40, and γ=0, δ=0 when no feedback or breathing is connected, ftarget=Clip[0.50∗6.00+0.40∗2.555,3,8]=4.022Hz. A suitable ftarget frequency is 4.02Hz. If using a binaural beat with headphones, fc = 220Hz, then fL = 220Hz and fR = 224.02Hz. For the clear semantic phase, set Gv = 0.90 and Gn = 0.025; during pauses, Gn can be briefly increased to 0.063; for sudden events or user discomfort, reduce Gn and m.

[0073] Example 2: Generation of focus-assisted composite neuroacoustic signals in online courses and teacher-assisted audio listening scenarios.

[0074] This embodiment illustrates the application of the present invention in classroom listening assistance scenarios such as online courses for students, recorded courses, online lectures, language learning, audio streams of teacher lectures, and audio transmission from the teacher's microphone to the student's end. This embodiment does not involve disease diagnosis or treatment; its purpose is to generate a low-intensity composite neuroacoustic signal layer coordinated with the teacher's speech rhythm, while ensuring the teacher's speech is clear and intelligible, for use in focusing assistance or maintaining attention during the listening process.

[0075] This embodiment is preferably applicable to scenarios where teacher voice input has already been digitized. In a typical offline classroom, if students hear the teacher speak directly through airborne transmission without a student-end audio terminal, the system cannot perform sample-level mixing and output of natural airborne sound. Therefore, offline classroom scenarios should be limited to listening assistance scenarios where the teacher's voice is collected via microphone or classroom sound reinforcement system and transmitted to student-end headphones, open-back headphones, pass-through headphones, learning terminals, or near-ear speakers.

[0076] The input can be teacher's lecture audio, online course audio, recorded course audio, classroom amplification audio, special lecture audio, or course lecture audio. The system will decode the input audio into x[n], Fs=48000Hz. If the input is mono teacher's audio, it will be copied as xL[n]= x[n], xR[n]=x[n].

[0077] Within a 15-second analysis window, the system extracted the following frequencies: fp=3.6Hz, fs=4.8Hz, fq=0.75Hz, and fb=3.2Hz. To ensure the clarity and intelligibility of the lecture audio, the weighting of syllable detail fs was reduced, while the weighting of the envelope-dominant modulation frequency fp and the envelope peak periodic frequency fb was increased. The values ​​were set as follows: p1=0.40, p2=0.20, p3=0.10, and p4=0.30.

[0078] Fspeech=0.40·3.60+0.20·4.80+0.10·0.75+0.30·3.20=3.435Hz.

[0079] The target usage state is listening focus assist. The system does not fix the focus assist frequency to a single value, but instead sets the focus assist candidate frequency range to Fatt=[12Hz, 18Hz]. In the absence of individual feedback data, the system defaults to using the calm focus candidate frequency fstate=14Hz.

[0080] Since fspeech=3.435Hz is below the Fatt range, the system maps fspeech to integer multiples: 2 times is 6.870Hz, 3 times is 10.305Hz, 4 times is 13.740Hz, and 5 times is 17.175Hz. Among them, 4·fspeech=13.740Hz falls within the Fatt range and is close to fstate=14Hz, so the system selects fmapped=13.740Hz.

[0081] ftarget=Clip[0.50*14.00+0.50*13.740,12.00,18.00]=13.870Hz.

[0082] Therefore, the target frequency of the composite neuroacoustic signal layer in this embodiment is 13.870 Hz. This frequency is not a fixed preset 12 Hz, nor is it simply a fixed β frequency, but is determined jointly by the focus-aid candidate range and the teacher's speech rhythm octave mapping results.

[0083] In this embodiment, headphones, open-back headphones, ambient sound pass-through headphones, or near-ear learning terminals are preferred. If binaural beats are used, a certain degree of channel isolation is required between the left and right ears; ordinary external speakers have crosstalk between the left and right channels and are not preferred. The headphones are not limited to closed-back headphones; for offline listening assistance, open-back headphones or pass-through headphones are preferred to avoid isolating the real classroom environment.

[0084] fc represents the binaural beat carrier frequency, not the focused auxiliary target frequency. The selectable range of fc is 160Hz to 440Hz, preferably 180Hz to 300Hz. In this embodiment, for ease of calculation, fc=220Hz, then fL=220Hz, fR=233.87Hz, and the frequency difference between the left and right channels is 13.87Hz.

[0085] φL[n]=φL[n−1]+2π∗220 / 48000; φR[n]=φR[n−1]+2π∗233.87 / 48000; sneu,L[n]sin(φL[n]); sneu,R[n]=sin(φR[n]).

[0086] To avoid disrupting normal learning, the semantic clarity stage was set with Gv[n] = 0.95, Gn[n] = 0.010 to 0.018, m[n] = 0.02 to 0.04, and k = 0.05 to 0.10. When a phrase pause or envelope trough occurred and lasted for more than 250ms, Gn[n] was smoothly increased from 0.015 to 0.030, and m[n] was smoothly increased from 0.03 to 0.06; after the pause ended, a smooth recovery occurred within 50ms to 100ms.

[0087] When events such as a teacher suddenly increasing their volume, students applauding, laughing, sharp background noise, or microphone popping are detected, the system reduces Gn[n], m[n], and k, and can apply peak limiting or dynamic compression to the teacher's main voice channel. If feedback is received from headphones, wristbands, or learning terminals, and increased head movement, increased body movement, increased heart rate, or subjective discomfort is detected in students, β, Gn[n], m[n], and k are reduced in the next control cycle, and ftarget is recalculated.

[0088] The digital mixed output is as follows: yL[n]=Limiter{0.95*xL[n]+0.015*sneu,L,total[n]}; yR[n]=Limiter{0.95*xR[n]+0.015*sneu,R,total[n]}; During the pause between sentences, Gn can be increased to 0.030. Finally, yL[n] and yR[n] are written into the audio output buffer, and output as an air sound pressure wave p(t) through the DAC or Bluetooth audio module, headphone amplifier and headphone transducer.

[0089] The composite neuroacoustic signal is no longer generated solely by a fixed frequency preset, but rather driven by the rhythmic characteristics of human speech. The system extracts E[n], M[l], fp, fs, fq, and fb from x[n] to form a calculable speech rhythm frequency fspeech(t). The target frequency ftarget(t) can be determined by combining the target usage state and the frequency doubling / division mapping of the speech rhythm, so that the auxiliary signal is coordinated with the natural rhythm of human speech. The composite neuroacoustic signal layer exists in the form of sneu[n], which belongs to the digital audio sampling sequence and can be generated by software, DSP, or audio chip. The mixed output is completed at the digital sampling level, and the final output is y[n], which is transduced into an air pressure wave p(t) by headphones or speakers. Reducing the strength of the auxiliary signal during the clear semantic stage and briefly enhancing it at pauses / enveloping valleys can reduce interference with the intelligibility of human speech. Reducing the strength of the auxiliary signal during loudness changes, applause, laughter, pops, and other events can reduce abruptness. The system can be deployed on mobile apps, smart headphones, open-back headphones, learning terminals, DSP chips, and smart audio devices.

[0090] The above description is merely a specific embodiment of the invention, but the scope of protection of the invention is not limited thereto. Any equivalent substitutions or changes made by those skilled in the art within the technical scope disclosed in the invention, based on the technical solution and concept of the invention, should be covered within the scope of protection of the invention.

Claims

1. A method for generating complex neuroacoustic signals driven by human voice rhythm, characterized in that, Includes the following steps: Obtain the digital audio sample sequence x[n] containing human voice; Speech rhythm features are extracted from the digital audio sampling sequence x[n]. The speech rhythm features include: speech amplitude envelope E[n], envelope modulation spectrum M[l], envelope dominant modulation frequency fp, syllable rhythm frequency fs, phrase pause rhythm frequency fq, and envelope peak period frequency fb. The speech rhythm frequency fspeech(t) is calculated based on the speech rhythm characteristics. Candidate frequencies fstate are determined based on the target usage state, and frequency-multiplication mapping or frequency-division mapping is performed based on the speech rhythm frequency fspeech(t) to obtain the mapped frequency fspeech_mapped(t). Then, the target neuroacoustic rhythm frequency ftarget(t) is calculated based on the candidate frequency fstate and the mapped frequency fspeech_mapped(t). A composite neuroacoustic signal layer sneu[n] is generated based on the target neuroacoustic rhythm frequency ftarget(t), wherein the composite neuroacoustic signal layer sneu[n] is a digital audio sampling sequence; The digital gain Gn[n], modulation depth m[n], or background layer coefficient k[n] of the composite neuroacoustic signal layer sneu[n] are dynamically adjusted based on the speech stage information or sudden event information of the human voice content. The composite neuroacoustic signal layer sneu[n] is digitally sampled and mixed with the digital audio sampling sequence x[n] to output the mixed digital audio sequence y[n].

2. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The step of extracting the speech amplitude envelope E[n] from the digital audio sampling sequence x[n] specifically includes: Perform a Hilbert transform operation on the digital audio sampling sequence x[n] to obtain the Hilbert component xH[n]; Construct an analytic signal z[n] = x[n] + j·xH[n], where j represents the imaginary unit; Calculate the magnitude of the analytic signal e0[n] = |z[n]|; Perform a digital low-pass filter operation with a cutoff frequency of 10Hz to 20Hz on the modulus value e0[n] to obtain the speech amplitude envelope E[n]=LPFfLPF{e0[n]}, where LPFfLPF represents performing a digital low-pass filter operation with a cutoff frequency of fLPF on the digital sequence within the parentheses.

3. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The envelope-dominant modulation frequency fp is obtained in the following manner: Extract a segment of data Ew[n] from the speech amplitude envelope E[n] within the sliding analysis window; Remove the average value of the data Ew[n] to eliminate the DC component; After applying a window function, a fast Fourier transform operation is performed to obtain the envelope modulation spectrum M[l]; The energy peak value of the envelope modulation spectrum M[l] is found in the range of 0.5Hz to 12Hz, and the frequency corresponding to the peak value is determined as the envelope dominant modulation frequency fp.

4. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The syllable rhythm frequency fs is obtained by estimating at least one of the following three methods: detecting local peaks exceeding the peak detection threshold θp in the speech amplitude envelope E[n] and counting the number of peaks Np; calculating fs_peak=Np / Tw based on the number of peaks Np, where Tw refers to the time length of the sliding analysis window in seconds; In the effective speech segment after speech activity detection, energy rise points or short-time energy peaks are detected and the number Nv is counted. Based on the number Nv, fs_vad = Nv / Tvoice is calculated, where Tvoice refers to the total duration of the effective speech segment in seconds. When speech recognition is enabled, the number of syllables Ns is estimated based on the recognized text, and fs_asr = Ns / Tvoice is calculated based on the number of syllables Ns. The phrase pause rhythm frequency fq is obtained through the following method: Detect continuous intervals where the speech amplitude envelope E[n] is lower than the low-energy pause threshold θq and lasts for more than 250ms, record the center time tq[i] of each pause interval, where i represents the ordinal index of the effective pause interval detected in chronological order, calculate the adjacent pause interval ΔTq[i]=tq[i]-tq[i-1], take the median or mean as the phrase pause period Tq, fq=1 / Tq; The envelope peak period frequency fb is obtained by detecting the significant peak time tp[i] of the speech amplitude envelope E[n] within a sliding window, calculating the adjacent peak interval ΔTb[i]=tp[i]-tp[i-1], and taking the median or robust mean as the peak period Tb, fb=1 / Tb.

5. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The speech rhythm frequency fspeech(t) is calculated using the following weighted fusion method: fspeech(t)=p1·fp+p2·fs+p3·fq+p4·fb, Where p1, p2, p3, and p4 are weighting coefficients and satisfy p1+p2+p3+p4=1; The weighting coefficients are selected from one of the following preset weighting combinations based on the type of human voice content: For steady-state speech or lecture types, p1=0.40, p2=0.20, p3=0.10, p4=0.30; In relaxed or low-arousal speech patterns, p1=0.45, p2=0.15, p3=0.30, and p4=0.

10. Under weak semantic or non-semantic prosodic types, p1=0.35, p2=0.15, p3=0.25, and p4=0.

25.

6. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The step of obtaining the mapped frequency fspeech_mapped(t) by performing frequency harmonic mapping or frequency division mapping based on the speech rhythm frequency fspeech(t) specifically includes: Multiply the speech rhythm frequency fspeech(t) by a positive integer h, or divide the speech rhythm frequency fspeech(t) by a positive integer h, such that the product or quotient falls within a preset range [Fmin, Fmax] of the target neuroacoustic rhythm frequency; when there are multiple positive integers h such that the product or quotient all fall within the preset range [Fmin, Fmax], select the product or quotient corresponding to the h that minimizes the absolute value of the difference between the product or quotient and the candidate frequency fstate(t), and use it as the mapped frequency fspeech_mapped(t); The target neuroacoustic rhythm frequency ftarget(t) is calculated in the following way: ftarget(t)=Clip[α·fstate(t)+β·fspeech_mapped(t),Fmin,Fmax], Where Clip represents the range constraint operation, α and β are weight coefficients, and Fmin and Fmax are the lower and upper limits of the target frequency, respectively; The target usage states include relaxed or low arousal state, pre-sleep assistance state, steady-state speaking companionship state, and focus assistance state. Different target usage states correspond to different candidate frequencies fstate and different frequency ranges [Fmin, Fmax].

7. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The composite neuroacoustic signal layer sneu[n] includes at least one of binaural beat signals and amplitude-modulated noise or natural sound signals; The binaural beat signal is generated in the following manner: Set the binaural beat carrier frequency fc, and generate the left channel phase sequence ΦL[n]=ΦL[n-1]+2π·fc / Fs and the right channel phase sequence ΦR[n]=ΦR[n-1]+2π·(fc+ftarget[n]) / Fs, respectively. Then generate the left channel basic binaural beat signal sneu,L[n]=sin(ΦL[n]) and the right channel basic binaural beat signal sneu,R[n]=sin(ΦR[n]). The basic binaural beat signal is gain controlled by the composite neuroacoustic signal layer digital gain Gn[n] in the digital mixing output stage, where Fs is the sampling rate. The amplitude-modulated noise or natural sound signal is generated in the following manner: The amplitude-modulated noise or natural sound base signal q[n] = X[n]·[1+m[n]·sin(Φtarget[n])], where X[n] is the pink noise or natural sound digital sampling sequence, m[n] is the modulation depth, and Φtarget[n] is the target phase sequence accumulated according to the target neuroacoustic rhythm frequency ftarget; the composite neuroacoustic signal layer can be represented as sneu,L,total[n] = sneu,L[n] + k[n]·q[n] and sneu,R,total[n] = sneu,R[n] + k[n]·q[n], where k[n] is the background layer coefficient.

8. The method for generating complex neuroacoustic signals driven by human voice rhythm as described in claim 1, characterized in that, The step of dynamically adjusting at least one of the gain, modulation depth, or background layer coefficients of the composite neuroacoustic signal layer sneu[n] based on the speech phase information or sudden event information of the human voice content specifically includes: In the clear semantic stage, the digital gain Gn[n] of the composite neuroacoustic signal layer is set to the first gain value, and the modulation depth m[n] is set to the first modulation depth value; During the pause or envelope valley phase, the digital gain Gn[n] of the composite neuroacoustic signal layer is increased from the first gain value to the second gain value, and the modulation depth m[n] is increased from the first modulation depth value to the second modulation depth value; When a sudden event is detected, at least one of the following is reduced: digital gain Gn[n] of the composite neuroacoustic signal layer, modulation depth m[n], and background layer coefficient k[n]; The clear semantic stage is defined as having a speech activity ratio greater than 70% and an average pause ratio of less than 20%. The pause or envelope valley phase is defined as the speech amplitude envelope E[n] being below a threshold and lasting for more than 250ms; The sudden event is triggered by at least one of the following conditions: the short-term loudness increases by more than 8dB within 200ms, the output peak exceeds -3dBFS, the proportion of high-frequency energy above 4kHz exceeds 35%, or applause, laughter, or popping sounds are detected.

9. A complex neuroacoustic signal generation system driven by human voice rhythm, characterized in that, The method for generating composite neuroacoustic signals as described in any one of claims 1-8 includes: An audio input or decoding unit is used to acquire a digital audio sampling sequence x[n] containing human voice; The speech rhythm feature extraction unit is used to extract speech rhythm features from the digital audio sampling sequence x[n]. The speech rhythm features include: speech amplitude envelope E[n], envelope modulation spectrum M[l], envelope dominant modulation frequency fp, syllable rhythm frequency fs, phrase pause rhythm frequency fq, and envelope peak period frequency fb. A speech rhythm frequency calculation unit is used to calculate the speech rhythm frequency fspeech(t) based on the speech rhythm characteristics. The target rhythm parameter generation unit is used to determine the candidate frequency fstate according to the target usage state, and perform frequency doubling or frequency division mapping according to the speech rhythm frequency fspeech(t) to obtain the mapped frequency fspeech_mapped(t), and then calculate the target neuroacoustic rhythm frequency ftarget(t) according to the candidate frequency fstate and the mapped frequency fspeech_mapped(t). A composite neuroacoustic signal generation unit is used to generate a composite neuroacoustic signal layer sneu[n] based on the target neuroacoustic rhythm frequency ftarget(t), wherein the composite neuroacoustic signal layer sneu[n] is a digital audio sampling sequence; A dynamic control unit is used to dynamically adjust at least one of the digital gain Gn[n], modulation depth m[n], or background layer coefficient k[n] of the composite neuroacoustic signal layer sneu[n] based on the speech stage information or sudden event information of the human voice content; The digital mixing output unit is used to perform digital sampling level mixing of the composite neuroacoustic signal layer sneu[n] and the digital audio sampling sequence x[n], and output the mixed digital audio sequence y[n].

10. The composite neuroacoustic signal generation system driven by human voice rhythm as described in claim 9, characterized in that, The dynamic control unit sets the digital gain Gn[n] of the composite neuroacoustic signal layer to a first gain value and the modulation depth m[n] to a first modulation depth value during the clear semantic phase. During the pause or envelope valley phase, it increases the digital gain Gn[n] of the composite neuroacoustic signal layer from the first gain value to a second gain value and increases the modulation depth m[n] from the first modulation depth value to a second modulation depth value. When a sudden event is detected, it reduces at least one of the digital gain Gn[n] of the composite neuroacoustic signal layer, the modulation depth m[n], and the background layer coefficient k[n]. The clear semantic phase is defined as a speech activity ratio greater than 70% and an average pause ratio less than 20%. The pause or envelope valley phase is defined as a speech amplitude envelope E[n] lower than a threshold and lasting for more than 250ms. The sudden event is triggered by at least one of the following conditions: short-term loudness increases by more than 8dB within 200ms, output peak exceeds -3dBFS, high-frequency energy above 4kHz accounts for more than 35%, or applause, laughter, or popping sounds are detected.