A bone conduction-based folding recording noise reduction method

CN122551813APending Publication Date: 2026-08-11SHENZHEN XINGMAN SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

该噪声不仅会掩蔽同时间段的语音信息,更可能对后续基于统计假设的通用降噪算法造成干扰,引发噪声估计偏差,进而导致语音失真或残留可闻的音频瑕疵,最终影响录音与通话的整体清晰度

Benefits of technology

[0015]本发明提供一种基于骨传导的折叠式录音降噪方法,通过S1:同步采集骨传导振动信号、空气传导音频信号、三轴加速度信号及铰链状态信号,并对采集的四路信号进行预处理;S2:基于预处理后的骨传导振动信号、三轴加速度信号、铰链状态信号以及历史脉冲区间信息,计算多模态瞬态干扰置信度,并据此识别并输出候选机械振动脉冲区间及其对应的置信度分数;S3:基于候选机械振动脉冲区间及其对应的置信度分数,结合预处理后的骨传导振动信号与空气传导音频信号在候选机械振动脉冲区间内的相关性,对候选机械振动脉冲区间进行确认,输出真实机械振动脉冲区间及其对应的置信度分数;S4:根据真实机械振动脉冲区间及其对应的置信度分数,对预处理后的骨传导振动信号在该真实机械振动脉冲区间内进行自适应滤波抑制处理,得到中间处理信号;S5:基于真实机械振动脉冲区间及其对应的置信度分数,对中间处理信号在真实机械振动脉冲区间的边界进行平滑处理,得到降噪后的骨传导振动信号,产生的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551813A_ABST
    Figure CN122551813A_ABST
Patent Text Reader

Abstract

This invention discloses a bone conduction-based folding recording noise reduction method, relating to the field of audio signal processing. The method includes: simultaneously acquiring four signals—bone conduction vibration, air conduction audio, triaxial acceleration, and hinge state—and performing preprocessing; calculating the confidence level of multimodal transient interference based on the preprocessed bone conduction vibration signal, triaxial acceleration signal, hinge state signal, and historical pulse information; identifying and outputting candidate mechanical vibration pulse intervals and their confidence scores; confirming the candidate mechanical vibration pulse intervals and outputting the actual mechanical vibration pulse intervals and their confidence scores; adaptively filtering and suppressing the bone conduction vibration signal in the corresponding intervals according to the actual mechanical vibration pulse intervals and their confidence scores to obtain intermediate processed signals; and smoothing the intermediate processed signals at the interval boundaries based on the actual mechanical vibration pulse intervals and their confidence scores to obtain the noise-reduced bone conduction vibration signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing, specifically to a folded recording noise reduction method based on bone conduction. Background Technology

[0002] With the evolution of mobile communication technology and the form factor of portable electronic devices, foldable terminals, with their unique balance between portability and display area, have become an important direction for the development of terminal products. In audio application scenarios of such devices, such as recording and voice calls, users' demand for clear sound quality is becoming increasingly prominent. To address the challenge of speech clarity in complex acoustic environments, bone conduction sensors, with their ability to acquire speech signals by picking up vibrations of the human jawbone, can effectively suppress some of the environmental noise conducted through the air. Therefore, they are gradually being applied to terminal devices, forming a multimodal speech pickup system together with traditional air conduction microphones. However, in the unique usage process of foldable devices, especially at the moment when the user performs the folding or unfolding action, mechanical structures such as hinges are subjected to stress changes, generating instantaneous mechanical impact and vibration. This vibration is directly transmitted through the device shell and internal structure to the bone conduction sensor vibrator, which is tightly coupled to the structural components. As a result, while the sensor is acquiring the target speech vibration signal, it inevitably couples in high-intensity transient mechanical vibration noise. This type of noise presents as short-duration, high-amplitude pulses in the time domain, and may excite specific structural resonance bands in the frequency domain. Its generation mechanism, propagation path, and signal characteristics are significantly different from conventional air-conducted background noise. This noise not only masks speech information within the same time period, but may also interfere with subsequent general noise reduction algorithms based on statistical assumptions, causing noise estimation bias, which in turn leads to speech distortion or residual audible audio flaws, ultimately affecting the overall clarity of recordings and calls.

[0003] Therefore, without relying on high-cost hardware structure modifications, this study investigates a signal processing method that can accurately detect and effectively suppress the transient mechanical vibration noise with distinct characteristics generated during the form-changing process of foldable devices. This method has a clear application need for improving the audio performance of foldable terminal devices, especially in bone conduction-based recording and call noise reduction. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a bone conduction-based foldable recording noise reduction method to solve the above-mentioned technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a foldable recording noise reduction method based on bone conduction, comprising: S1: Synchronously acquire bone conduction vibration signals, air conduction audio signals, triaxial acceleration signals, and hinge status signals, and preprocess the acquired four signals; S2: Based on the preprocessed bone conduction vibration signal, triaxial acceleration signal, hinge state signal and historical pulse interval information, calculate the confidence of multimodal transient interference, and identify and output candidate mechanical vibration pulse intervals and their corresponding confidence scores accordingly. S3: Based on the candidate mechanical vibration pulse intervals and their corresponding confidence scores, and combined with the correlation between the preprocessed bone conduction vibration signal and the air conduction audio signal within the candidate mechanical vibration pulse intervals, the candidate mechanical vibration pulse intervals are confirmed, and the real mechanical vibration pulse intervals and their corresponding confidence scores are output. S4: Based on the real mechanical vibration pulse range and its corresponding confidence score, the preprocessed bone conduction vibration signal is subjected to adaptive filtering and suppression processing within the real mechanical vibration pulse range to obtain the intermediate processed signal. S5: Based on the real mechanical vibration pulse range and its corresponding confidence score, the intermediate processed signal is smoothed at the boundary of the real mechanical vibration pulse range to obtain the noise-reduced bone conduction vibration signal.

[0006] The present invention is further configured such that S1 includes: Simultaneously acquire bone conduction vibration signals output by bone conduction sensors, air conduction audio signals output by air microphones, triaxial acceleration signals output by triaxial accelerometers, and hinge status signals output by hinge status sensors; The four acquired signals are input to a multi-channel synchronous sample-and-hold analog-to-digital converter for synchronous sampling and analog-to-digital conversion to obtain time-aligned digital bone conduction vibration signals, digital air conduction audio signals, digital triaxial acceleration signals, and digital hinge status signals. High-pass filters with the same preset cutoff frequency are applied to the digital bone conduction vibration signal and the digital air conduction audio signal respectively to obtain the preprocessed bone conduction vibration signal and air conduction audio signal. The digital triaxial acceleration signal is multiplied by calibration coefficients and zero bias calibration is performed to obtain a triaxial acceleration value sequence as the preprocessed triaxial acceleration signal. The digital hinge state signal is subjected to binarization event detection processing to obtain a binary event sequence representing the hinge rotation state as the preprocessed hinge state signal.

[0007] The present invention is further configured such that S2 includes: The preprocessed bone conduction vibration signal is then subjected to frame segmentation. For each signal frame, calculate the confidence level of the audio feature sub-score based on the preprocessed bone conduction vibration signal corresponding to that signal frame; For each signal frame, the device state sub-confidence is calculated based on the preprocessed triaxial acceleration signal and hinge state signal corresponding to that signal frame; For each signal frame, the confidence level of the historical pulse correlation sub-score is calculated based on a historical pulse interval record queue that stores a preset number of real mechanical vibration pulse intervals that have been confirmed in previous processing cycles. For each signal frame, the audio feature sub-confidence, device status sub-confidence, and historical pulse correlation sub-confidence are weighted and summed to obtain the multimodal transient interference confidence of that signal frame. The confidence level of the multimodal transient interference in each signal frame is compared with a preset detection threshold, and signal frames with a multimodal transient interference confidence level greater than the detection threshold are marked as candidate pulse frames. Clustering of temporally continuous candidate pulse frames generates candidate mechanical vibration pulse intervals. The arithmetic mean of the multimodal transient interference confidence scores of all signal frames within each candidate mechanical vibration pulse interval is used as the confidence score corresponding to that candidate mechanical vibration pulse interval.

[0008] The present invention is further configured such that the calculation process of the audio feature sub-confidence includes: The first feature value is obtained by the ratio of the short-time energy of the preprocessed bone conduction vibration signal corresponding to the signal frame to the average energy of the background noise of a preset number of historical frames. The second feature value is obtained by the absolute value of the difference between the spectral centroid of the preprocessed bone conduction vibration signal corresponding to the signal frame and the mean spectral centroid of the historical frames that were identified as speech frames. The third characteristic value is obtained based on the ratio of the signal energy of the preprocessed bone conduction vibration signal corresponding to the signal frame within the preset mechanical resonance frequency band to the signal energy of the full frequency band. The first, second, and third feature values ​​are mapped using their respective preset monotonically increasing functions, and the results of each feature value mapping are weighted and summed using preset weights to obtain the audio feature sub-confidence.

[0009] The present invention is further configured such that the calculation process of the device state sub-confidence includes: Based on the triaxial components of the preprocessed triaxial acceleration signal corresponding to the signal frame, the Euclidean norm of the difference vector between the triaxial components and the preset static reference vector is calculated. The calculation result is low-pass filtered to obtain the motion intensity characterization value, and then mapped to the first intermediate value through the Sigmoid function. Based on the preprocessed hinge state signal corresponding to the signal frame, the hinge state signal value within the time period of the signal frame is logically ORed to obtain the second intermediate value of the folding event. The first intermediate value and the second intermediate value are weighted and fused with preset weights to obtain the device status sub-confidence.

[0010] The present invention is further configured such that the calculation process of the historical pulse correlation confidence score includes: Retrieve a queue of historical pulse interval records containing a preset number of real mechanical vibration pulse intervals that have been confirmed in previous processing cycles; Based on the current signal frame center time and the end time of the most recent real mechanical vibration pulse interval in the historical pulse interval recording queue, the timing decay component is calculated. The timing decay component is an exponential function with the natural constant e as the base. Its exponent is the quotient of the negative difference between the current signal frame time and the end time of the most recent real mechanical vibration pulse interval in the historical pulse interval recording queue and the preset decay time constant. Based on the number of occurrences of real mechanical vibration pulse intervals obtained from the historical pulse interval record queue within a preset time window and the preset maximum possible number of pulses, a frequency statistical component is calculated. The frequency statistical component is the minimum value between the ratio of the number of occurrences to the maximum possible number of pulses and the value 1. The temporal decay component and the frequency statistics component are weighted and fused with preset weights to obtain the confidence level of the historical pulse correlation sub-component.

[0011] The present invention is further configured such that S3 includes: For each candidate mechanical vibration pulse interval, the first signal segment and the second signal segment corresponding to the candidate mechanical vibration pulse interval in time are extracted from the preprocessed bone conduction vibration signal and air conduction audio signal, respectively, and the time domain cross-correlation coefficient between the first signal segment and the second signal segment is calculated. Based on a preset basic correlation threshold, a preset adjustment coefficient, and the confidence score corresponding to the candidate mechanical vibration pulse interval, a dynamic confirmation threshold is calculated. The dynamic confirmation threshold is calculated by adding the basic correlation threshold to an adjustment term, wherein the adjustment term is the product of the adjustment coefficient and the value 1 minus the confidence score. The time-domain cross-correlation coefficient is compared with the dynamic confirmation threshold. If the time-domain cross-correlation coefficient is less than the dynamic confirmation threshold, the candidate mechanical vibration pulse interval is confirmed as the real mechanical vibration pulse interval, and the real mechanical vibration pulse interval and its corresponding confidence score are output.

[0012] The present invention is further configured such that S4 includes: For each real mechanical vibration pulse interval, the preprocessed air conduction audio signal of its corresponding time period and a preset duration before the start time of that time period is used as the reference input signal, and the preprocessed bone conduction vibration signal of its corresponding time period is used as the main input signal. Based on the confidence scores corresponding to the actual mechanical vibration pulse range, a step size parameter is set for the adaptive filter. The step size parameter is the sum of a preset minimum step size value and an adjustment amount. The adjustment amount is the product of the difference between the preset maximum step size value and the minimum step size value and the confidence score. Based on the reference input signal and the main input signal, an adaptive filter with the step size parameter is used to perform filtering operations; Subtracting the output signal of the adaptive filter from the main input signal yields the suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse range; During periods outside the actual mechanical vibration pulse range, the pre-processed bone conduction vibration signal is used as the output signal for that period and is maintained. The suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse interval is spliced ​​with the bone conduction vibration signal maintained outside the actual mechanical vibration pulse interval in a time sequence to obtain the intermediate processed signal.

[0013] The present invention is further configured such that S5 includes: For each real mechanical vibration pulse interval, the width of the smooth transition band at its front and rear boundaries is determined based on its corresponding confidence score. The smooth transition band width is calculated based on a preset minimum transition band length, a preset maximum transition band length, and the confidence score. The smooth transition band width is the sum of the minimum transition band length and an expansion amount. The expansion amount is the rounded value of the product of the difference between the maximum transition band length and the minimum transition band length and the confidence score. Based on the start and end times of the actual mechanical vibration pulse interval, the duration of the smooth transition band width is extended forward and backward respectively to form a left smooth transition band and a right smooth transition band; Within the left smooth transition band, a first time-varying gain linearly increases from 0 to 1 is applied to the intermediate processed signal, and a second time-varying gain linearly decreases from 1 to 0 is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the left smooth transition band. Within the right smooth transition band, a third time-varying gain, linearly decreasing from 1 to 0, is applied to the intermediate processed signal, and a fourth time-varying gain, linearly increasing from 0 to 1, is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the right smooth transition band. Within the actual mechanical vibration pulse range, the intermediate processed signal is used as the output signal for that actual mechanical vibration pulse range; In other time periods outside the left smooth transition zone, the real mechanical vibration pulse interval, and the right smooth transition zone, the preprocessed bone conduction vibration signal is used as the output signal for that time period. The output signals corresponding to the real mechanical vibration pulse interval, the left smooth transition band, the right smooth transition band, and other time periods are spliced ​​together in time sequence to obtain the noise-reduced bone conduction vibration signal.

[0014] The present invention is further configured such that the adaptive filter is a normalized minimum mean square adaptive filter.

[0015] This invention provides a bone conduction-based folding audio recording noise reduction method. The method comprises: S1: Simultaneously acquiring bone conduction vibration signals, air conduction audio signals, triaxial acceleration signals, and hinge state signals, and preprocessing the four acquired signals; S2: Calculating the confidence level of multimodal transient interference based on the preprocessed bone conduction vibration signals, triaxial acceleration signals, hinge state signals, and historical pulse interval information, and identifying and outputting candidate mechanical vibration pulse intervals and their corresponding confidence scores; S3: Combining the preprocessed bone conduction vibration signals and air conduction audio signals with the candidate mechanical vibration pulse intervals and their corresponding confidence scores. The correlation of the frequency signal within the candidate mechanical vibration pulse interval is used to confirm the candidate mechanical vibration pulse interval, and output the real mechanical vibration pulse interval and its corresponding confidence score; S4: Based on the real mechanical vibration pulse interval and its corresponding confidence score, the preprocessed bone conduction vibration signal is subjected to adaptive filtering and suppression processing within the real mechanical vibration pulse interval to obtain the intermediate processed signal; S5: Based on the real mechanical vibration pulse interval and its corresponding confidence score, the boundary of the intermediate processed signal within the real mechanical vibration pulse interval is smoothed to obtain the noise-reduced bone conduction vibration signal. The beneficial effects include: By constructing and calculating a multimodal transient interference confidence level that integrates audio features, real-time device motion status, and historical pulse correlation patterns, a joint probability assessment of mechanical vibration pulse noise is achieved. This method overcomes the limitations of false detection and missed detection that exist when relying solely on audio signal analysis. By introducing a historical pulse queue to perceive the temporal correlation patterns of noise events, it achieves higher robustness in identifying pulse sequences caused by continuous or repetitive mechanical actions in complex application scenarios. By outputting candidate mechanical vibration pulse intervals with confidence scores, it provides a refined and quantifiable decision-making basis for subsequent processing stages. By utilizing the physical characteristic of extremely low correlation between real mechanical pulses in bone conduction and air conduction dual-channel signals for secondary confirmation, acoustic interference such as speech plosives that are similar to pulses only in single-channel features is effectively eliminated. A dynamic confirmation threshold decision mechanism based on candidate interval confidence scores is designed to adaptively adjust the strictness of the dual-channel correlation test. This dynamic mechanism relaxes the correlation decision conditions for high-confidence candidate intervals, avoiding false rejections caused by weak airborne sound radiation induced by strong mechanical vibrations, while tightening the decision conditions for low-confidence candidate intervals to prevent false positives. Thus, while ensuring the utilization of the strong discriminative feature of uncorrelated dual-channel signals, the results of the previous multimodal transient interference confidence assessment are effectively integrated, improving the accuracy and reliability of the final mechanical pulse confirmation result.

[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 The flowchart illustrates a folded recording noise reduction method based on bone conduction, as an exemplary embodiment of the present invention. Detailed Implementation

[0018] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0019] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0020] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0021] A bone conduction-based folded recording noise reduction method, such as Figure 1 As shown, it includes: S1: Synchronously acquire bone conduction vibration signals, air conduction audio signals, triaxial acceleration signals, and hinge status signals, and preprocess the acquired four signals; S2: Based on the preprocessed bone conduction vibration signal, triaxial acceleration signal, hinge state signal and historical pulse interval information, calculate the confidence of multimodal transient interference, and identify and output candidate mechanical vibration pulse intervals and their corresponding confidence scores accordingly. S3: Based on the candidate mechanical vibration pulse intervals and their corresponding confidence scores, and combined with the correlation between the preprocessed bone conduction vibration signal and the air conduction audio signal within the candidate mechanical vibration pulse intervals, the candidate mechanical vibration pulse intervals are confirmed, and the real mechanical vibration pulse intervals and their corresponding confidence scores are output. S4: Based on the real mechanical vibration pulse range and its corresponding confidence score, the preprocessed bone conduction vibration signal is subjected to adaptive filtering and suppression processing within the real mechanical vibration pulse range to obtain the intermediate processed signal. S5: Based on the real mechanical vibration pulse range and its corresponding confidence score, the intermediate processed signal is smoothed at the boundary of the real mechanical vibration pulse range to obtain the noise-reduced bone conduction vibration signal.

[0022] The present invention is further configured such that S1 includes: Simultaneously acquire bone conduction vibration signals output by bone conduction sensors, air conduction audio signals output by air microphones, triaxial acceleration signals output by triaxial accelerometers, and hinge status signals output by hinge status sensors; The four acquired signals are input to a multi-channel synchronous sample-and-hold analog-to-digital converter for synchronous sampling and analog-to-digital conversion to obtain time-aligned digital bone conduction vibration signals, digital air conduction audio signals, digital triaxial acceleration signals, and digital hinge status signals. High-pass filters with the same preset cutoff frequency are applied to the digital bone conduction vibration signal and the digital air conduction audio signal respectively to obtain the preprocessed bone conduction vibration signal and air conduction audio signal. The digital triaxial acceleration signal is multiplied by calibration coefficients and zero bias calibration is performed to obtain a triaxial acceleration value sequence as the preprocessed triaxial acceleration signal. The digital hinge state signal is binarized for event detection processing to obtain a binary event sequence representing the hinge rotation state as the preprocessed hinge state signal. Specifically, to achieve accurate identification and subsequent suppression of mechanical vibration noise generated by the hinge movement of the folding device, synchronous acquisition and preprocessing of the signal are first performed. Specifically, four sensors built into the folding device are used, including a bone conduction sensor, an air microphone, a triaxial accelerometer, and a hinge state sensor. Among them, the bone conduction sensor can be a piezoelectric accelerometer or an electromagnetic vibration sensor, which is tightly attached and mechanically fixed inside the device body by a rigid bracket, special adhesive, or welding process, and has a direct contact with the folding hinge structure. A physical connection is established to ensure efficient transmission of mechanical impact vibrations to the sensor. This sensor is used to pick up bone conduction vibration signals, including mechanical impact vibrations directly coupled to the user's jawbone target voice vibrations and hinge movements, and converts them into corresponding bone conduction vibration analog voltage signals. The air microphone is located behind the acoustic opening of the device housing. An acoustic seal or buffer material is provided between it and the internal structure of the device to isolate structural vibrations transmitted through the solid. This microphone is used to pick up air conduction audio signals, mainly including target human voices and environmental noise transmitted through the air, and converts them into corresponding air conduction audio analog voltage signals. The triaxial accelerometer is installed inside the main body of the device, and its installation orientation is calibrated to ensure... The three measuring axes are aligned with the preset reference direction of the equipment housing to synchronously measure the linear acceleration of the equipment in three orthogonal spatial directions. These accelerations together constitute a triaxial acceleration signal characterizing the overall motion state of the equipment, which is then converted into a corresponding triaxial acceleration analog voltage signal. The hinge state sensor is a Hall sensor installed near the hinge axis, fixed to the stationary part of the hinge. When the hinge rotates, it generates an analog voltage output that is a function of the hinge's angular velocity or angle change due to changes in the magnetic field. This produces a hinge state analog voltage signal that directly reflects the dynamic motion of the hinge. This hinge state analog voltage signal constitutes the hinge state signal picked up by the hinge state sensor. To ensure subsequent calculations... The method can establish accurate instantaneous event correlations by inputting four analog voltage signals in parallel into a multi-channel synchronous sample-and-hold analog-to-digital converter for synchronous sampling and analog-to-digital conversion. Each channel of the analog-to-digital converter contains a sample-and-hold circuit, driven by the same master clock, and can be synchronously triggered at the effective edge of the clock to instantaneously capture and hold the analog voltage signal values ​​of all input channels. Subsequently, the analog-to-digital conversion core sequentially digitizes the analog voltage signal values ​​held by each channel, thereby ensuring that four digital signals with strict sample point-to-point time alignment are finally obtained in the discrete time series, including digital bone conduction vibration signal, digital air conduction audio signal, digital triaxial acceleration signal, and digital hinge status signal.The analog-to-digital converter (ADC) has a preset quantization bit depth to meet the system's dynamic range requirements, such as 16 bits or 24 bits. Its sampling frequency is set according to the Nyquist sampling theorem, for example, 16kHz, and is provided by a low-jitter, stable clock source to ensure the accuracy of sampling timing. The multi-channel synchronous sample-and-hold ADC can be a single-chip integrated multi-channel ADC or a combination of multiple independent ADC chips with strictly synchronized clocks. Subsequently, the four digital signals are preprocessed in parallel, specifically including: applying high-pass filters with the same preset cutoff frequency to the digital bone conduction vibration signal and the digital air conduction audio signal to filter out ultra-low frequency interference and DC offset, obtaining preprocessed bone conduction vibration signal and preprocessed air conduction audio signal; and performing calibration coefficient multiplication and zero-biasing on the digital triaxial acceleration signal. The calibration process involves sequentially multiplying by calibration coefficients and subtracting from zero bias values ​​to obtain a sequence of triaxial acceleration values, which serves as the preprocessed triaxial acceleration signal. The digital hinge state signal undergoes binarization event detection processing, where a comparator logic compares it to a preset state threshold. If the digital hinge state signal is greater than the preset threshold, a logic "1" is output to indicate that the hinge is in a rotating state; otherwise, a logic "0" is output to indicate that the hinge is in a stationary state. This yields a binary event sequence representing the hinge's rotational state, which serves as the preprocessed hinge state signal. After step S1 is completed, four sets of time-aligned preprocessed signals are output: the preprocessed bone conduction vibration signal, the preprocessed air conduction audio signal, the preprocessed triaxial acceleration signal, and the preprocessed hinge state signal. These form the data foundation for subsequent algorithm processing.

[0023] The present invention is further configured such that S2 includes: The preprocessed bone conduction vibration signal is subjected to frame segmentation. Specifically, the preprocessed bone conduction vibration signal output in step S1 is subjected to frame segmentation. This frame segmentation process uses analysis frames of a preset duration and slides them on the time axis with fixed frames of a shorter duration, thereby dividing the continuous signal stream into a series of signal frames that partially overlap in time. This frame segmentation operation completes the conversion from continuous time domain signal to discrete frame sequence, so that subsequent frame-by-frame fixed-point analysis and feature calculation can be performed. For each signal frame, an audio feature sub-confidence is calculated based on the preprocessed bone conduction vibration signal corresponding to that signal frame. The invention further specifies that the calculation process of the audio feature sub-confidence includes: obtaining a first feature value based on the ratio of the short-time energy of the preprocessed bone conduction vibration signal corresponding to that signal frame to the average energy of the background noise of a preset number of historical frames; obtaining a second feature value based on the absolute value of the difference between the spectral centroid of the preprocessed bone conduction vibration signal corresponding to that signal frame and the mean spectral centroid of the historical frames identified as speech frames; and obtaining a third feature value based on the ratio of the signal energy of the preprocessed bone conduction vibration signal corresponding to that signal frame within a preset mechanical resonance frequency band to the signal energy across the entire frequency band. The first, second, and third feature values ​​are mapped using corresponding preset monotonically increasing functions, and the results of each feature value mapping are weighted and summed with preset weights to obtain the audio feature sub-confidence. Specifically, the audio feature sub-confidence is used to quantify and evaluate the degree of matching between the time-frequency domain features of the current signal frame and the hinge mechanical vibration pulse noise features. Its calculation is based on the preprocessed bone conduction vibration signal corresponding to the current signal frame and is achieved through the following steps: three time-frequency domain features are extracted sequentially, including the first, second, and third feature values; the calculation logic of the first feature value is as follows: by calculating all preprocessed bone conduction vibration signals within the current signal frame... The short-time energy of a signal frame is obtained by summing the squares of the sampled values ​​of the conducted vibration signal. A sliding window of a preset length is maintained simultaneously to dynamically estimate the background noise energy. This sliding window stores the short-time energy of a preset number of recent consecutive signal frames. Whenever the short-time energy of a new signal frame is calculated, it is added to the end of the sliding window. When the sliding window is full, the oldest short-time energy is removed to maintain a constant window capacity. Based on the short-time energy sequence currently stored in the sliding window, a statistically representative value that is insensitive to instantaneous energy fluctuations is selected as the background noise energy corresponding to the current signal frame by calculating either the median or a preset low quantile of the short-time energy sequence. The average energy of the background noise is calculated as follows: the short-time energy of the current signal frame is divided by the average energy of the background noise, and the quotient is then divided by the logarithm to the base 10 and multiplied by 10 to obtain the first characteristic value. This first characteristic value is used to quantify the prominence of the instantaneous energy intensity of the current signal frame relative to the estimated background noise level. The calculation logic of the second characteristic value is as follows: a fast Fourier transform is performed on the preprocessed bone conduction vibration signal of the current signal frame to obtain its discrete amplitude spectrum. The frequency value of each discrete frequency point is multiplied by the amplitude spectrum value corresponding to that frequency point. The product of all frequency points is summed, and the summation result is divided by the sum of the amplitude spectrum values ​​of all frequency points. The quotient is the spectral centroid of the current signal frame.Simultaneously, a historical speech frame spectral centroid buffer, implemented using a first-in-first-out (FIFO) queue structure and with a preset fixed capacity, is maintained. When a new speech frame is determined to be a clean speech frame by the speech activity detection algorithm, its spectral centroid is added to the tail of the queue. If the queue is full, the oldest data is removed from the head. The arithmetic mean of all spectral centroids stored in the buffer is calculated as the historical reference centroid. The absolute value of the difference between the spectral centroid of the current signal frame and the historical reference centroid is calculated. This absolute value is the second feature value of the current signal frame. This second feature value is used to quantify the difference between the signal spectral energy distribution within the current signal frame and the concentration trend of historical typical speech spectral energy. The calculation logic for the third feature value... The following steps are taken: The resonant frequency band of the main mechanical vibration noise excited by the device hinge movement is determined through pre-conducted offline experimental calibration. This calibration is performed in an acoustically isolated experimental environment. The device hinge is induced to complete a preset number of opening and closing cycles while bone conduction vibration signals are simultaneously acquired. Spectral analysis is performed on the acquired bone conduction vibration signals, and the frequency bands where the signal energy is greater than a preset energy threshold and occurs repeatedly are determined as the preset mechanical resonant frequency band. The signal energy of the preprocessed bone conduction vibration signal of the current signal frame within this preset mechanical resonant frequency band is calculated. This signal energy calculation is based on the discrete amplitude spectrum obtained by performing a Fast Fourier Transform on the preprocessed bone conduction vibration signal of the signal frame. The first, second, and third eigenvalues ​​are obtained by summing the squares of the amplitude spectrum values ​​corresponding to all discrete frequency points within the preset mechanical resonance frequency band. Simultaneously, the total signal energy of the preprocessed bone conduction vibration signal across the entire frequency band is calculated, which is obtained by summing the squares of the amplitude spectrum values ​​of all effective frequency points after the fast Fourier transform. The signal energy of the preset mechanical resonance frequency band is divided by the total signal energy of the entire frequency band to obtain the third eigenvalue, which reflects the relative proportion of mechanical vibration energy in the total signal energy. The extracted first, second, and third eigenvalues ​​are mapped to a numerical range of zero to one using a preset monotonically increasing function. The mapping of the first eigenvalue uses a... The mapping of the second eigenvalue is implemented using an exponential function with the natural constant e as the base. Its mapping output value is 1 minus an exponential decay term. The exponent of this exponential decay term is the product of a preset first negative slope coefficient and the portion of the first eigenvalue that exceeds a preset short-time energy threshold. If it does not exceed the preset short-time energy threshold, it is zero. The mapping of the second eigenvalue is implemented using an exponential function with the natural constant e as the base. Its mapping output value is 1 minus an exponential decay term. The exponent of this exponential decay term is the product of a preset second negative slope coefficient and the ratio obtained by dividing the second eigenvalue by the standard deviation of the centroid of the historical speech frame spectrum. The standard deviation of the centroid of the historical speech frame spectrum is calculated based on all the centroids of the spectrum stored in the centroid buffer of the historical speech frame spectrum.The mapping of the third eigenvalue is achieved through normalization calculation. Its mapping output value is obtained by dividing the third eigenvalue of the current signal frame by the sum of the third eigenvalue and a reference value that is determined through pre-statistical analysis and represents the average energy proportion of a typical speech signal in a preset mechanical resonance frequency band. The three results obtained by the above mapping are weighted and summed according to a set of preset weight coefficients that sum to one. The summation result is the audio feature sub-confidence of the current signal frame. The values ​​of each weight coefficient are determined based on the ability of each feature to distinguish mechanical vibration impulse noise. For each signal frame, the device state sub-confidence is calculated based on the preprocessed triaxial acceleration signal and hinge state signal corresponding to that signal frame. The invention further specifies that the calculation process of the device state sub-confidence includes: calculating the Euclidean norm of the difference vector between the triaxial components and a preset static reference vector based on the triaxial components of the preprocessed triaxial acceleration signal corresponding to that signal frame; performing low-pass filtering on the calculation result to obtain a motion intensity characterization value; and mapping it to a first intermediate value using the Sigmoid function; performing a logical OR operation on the hinge state signal values ​​within the signal frame period based on the preprocessed hinge state signal corresponding to that signal frame to obtain a second intermediate value indicating the occurrence of the folding event; and then converting the first intermediate value into a second intermediate value. The first intermediate value and the second intermediate value are weighted and fused with preset weights to obtain the device state sub-confidence. Specifically, the device state sub-confidence is used to quantify the probability that the physical motion of the device and the hinge rotation event in the current signal frame jointly indicate the occurrence of a mechanical vibration pulse. Its calculation is based on the preprocessed triaxial acceleration signal and the preprocessed hinge state signal corresponding to the signal frame, and is achieved through the following steps: Calculate the first intermediate value, the purpose of which is to transform the triaxial acceleration signal into a scalar characterizing the overall motion intensity of the device. The specific process is as follows: Based on the component sequence of the preprocessed triaxial acceleration signal corresponding to the signal frame on the three spatial orthogonal axes, for each sampling time point in the sequence, calculate the triaxial acceleration component at that time point. The Euclidean norm of the difference vector between the instantaneous vector and a preset static reference vector representing the stationary state of the device is used to obtain the instantaneous motion intensity at that time point. To suppress instantaneous jitter and extract the low-frequency trend representing the overall displacement or rotation of the device, a first-order low-pass filter is applied to the instantaneous motion intensity sequence at all time points within the signal frame. The arithmetic mean or root mean square value of the filtered sequence is calculated to obtain the aggregated motion intensity representation value. This value is then mapped to a zero-to-one range using a preset Sigmoid function, and the result is the first intermediate value. A second intermediate value is calculated to perform binary decision based on the direct event signal provided by the hinge state sensor. The specific process is as follows: The preprocessed hinge state signal corresponding to the signal frame is a binary event sequence, where logic "1" indicates that a hinge rotation event was detected at the corresponding time point; based on the transient characteristics of the hinge's mechanical impact vibration, a logical OR operation is performed on the binary states at all time points within the signal frame. If the operation result is true, it is determined that a folding event has occurred within the signal frame, and the second intermediate value is set to logic "1". Otherwise, the second intermediate value is set to logic "0". This second intermediate value provides direct event evidence of whether the hinge has moved; the first intermediate value and the second intermediate value are weighted and summed according to a preset set of weight coefficients. The sum of these weight coefficients is one, and the summation result is the device state sub-confidence of the current signal frame; For each signal frame, a historical pulse correlation confidence score is calculated based on a historical pulse interval record queue containing a preset number of real mechanical vibration pulse intervals confirmed in previous processing cycles. The invention further specifies that the calculation process of the historical pulse correlation confidence score includes: acquiring a historical pulse interval record queue containing a preset number of real mechanical vibration pulse intervals confirmed in previous processing cycles; calculating a timing decay component based on the current signal frame center time and the end time of the most recent real mechanical vibration pulse interval in the historical pulse interval record queue, wherein the timing decay component is an exponential function with the natural constant e as its base, and its exponent is the sum of the current signal frame time and the historical pulse interval. The negative difference between the end time of the most recent real mechanical vibration pulse interval in the record queue and the preset decay time constant is used to calculate the frequency statistical component. Based on the number of occurrences of real mechanical vibration pulse intervals obtained from the historical pulse interval record queue within a preset time window and the preset maximum possible number of pulses, a frequency statistical component is calculated. This frequency statistical component is the minimum value between the ratio of the number of occurrences to the maximum possible number of pulses and the value 1. The time-series decay component and the frequency statistical component are then weighted and fused with preset weights to obtain the historical pulse correlation sub-confidence. Specifically, the calculation of the historical pulse correlation sub-confidence relies on a dynamically updated historical pulse interval record queue. Initially empty, its maintenance mechanism is as follows: A first-in-first-out (FIFO) queue data structure is maintained in memory. This queue stores a preset number of real mechanical vibration pulse intervals that have been confirmed in previous processing cycles. Each stored real mechanical vibration pulse interval record includes its start and end times. Whenever a new real mechanical vibration pulse interval is confirmed, it is added to the tail of the historical pulse interval record queue. If the length of the historical pulse interval record queue exceeds a preset limit, the oldest interval record is removed from the head. This allows the historical pulse interval record queue to dynamically maintain information on several recently occurred real mechanical vibration pulse intervals whose number of records does not exceed the preset limit. Next, the timing is calculated. The calculation logic for the attenuation component is as follows: Obtain the center time of the currently processed signal frame and query the historical pulse interval record queue to obtain the end time of the latest actual mechanical vibration pulse interval. Calculate the time difference between the center time of the current signal frame and the end time of the most recent actual mechanical vibration pulse interval in the historical pulse interval record queue. Negate this time difference and divide it by a preset attenuation time constant to obtain the intermediate quotient. Calculate the value of an exponential function using the natural constant e as the base and the intermediate quotient as the exponent. The result is the time-series attenuation component, which is used to quantify the attenuation effect caused by the elapsed time since the end of the last pulse event.Then, the frequency statistical component is calculated, and the calculation logic is as follows: Define a time window of a preset length, and trace back the length of this time window based on the center time of the current signal frame to delineate a historical statistical interval; count the total number of real mechanical vibration pulse intervals in the historical pulse interval record queue whose end time is within this historical statistical interval as the occurrence count, divide the occurrence count by the preset maximum possible number of pulses to obtain a preliminary ratio; take the smaller of the preliminary ratio and the value 1 as the frequency statistical component, which is used to quantify the density of pulse events in the recent history adjacent to the current frame; finally, the obtained time-series decay component and frequency statistical component are weighted and summed according to two preset weight coefficients that sum to one, and the sum is the confidence of the historical pulse correlator of the current signal frame; For each signal frame, the audio feature sub-confidence, device status sub-confidence, and historical pulse correlation sub-confidence are weighted and summed to obtain the multimodal transient interference confidence of the signal frame. The multimodal transient interference confidence of each signal frame is compared with a preset detection threshold; signal frames with a multimodal transient interference confidence greater than the detection threshold are marked as candidate pulse frames. Temporally continuous candidate pulse frames are clustered to generate candidate mechanical vibration pulse intervals, and the arithmetic mean of the multimodal transient interference confidence of all signal frames within each candidate mechanical vibration pulse interval is used as the confidence score corresponding to that candidate mechanical vibration pulse interval. Specifically, after independently calculating the audio feature sub-confidence, device status sub-confidence, and historical pulse correlation sub-confidence of the current signal frame, these three sub-confidences are weighted and summed according to preset weight coefficients. The summation result is defined as the multimodal transient interference confidence of the signal frame, which is used to quantify the current signal frame. The probability of the noise being a hinge mechanical vibration pulse is calculated. After calculating the confidence level of the multimodal transient interference for all signal frames, the pulse detection stage begins. In this stage, the confidence level of the multimodal transient interference calculated for each signal frame is compared with a preset detection threshold, and all signal frames with a multimodal transient interference confidence level greater than the detection threshold are marked as candidate pulse frames. Subsequently, clustering and merging processing is performed on the candidate pulse frames that are continuously distributed on the time axis. Specifically, by traversing the entire time axis, all temporally adjacent candidate pulse frames are merged and connected to form a continuous candidate mechanical vibration pulse interval. For each generated candidate mechanical vibration pulse interval, its corresponding confidence score is calculated. The calculation logic for this confidence score is as follows: extract the multimodal transient interference confidence levels of all signal frames covered by the candidate mechanical vibration pulse interval, calculate the arithmetic mean of the extracted multimodal transient interference confidence levels, and use this arithmetic mean as the confidence score corresponding to the candidate mechanical vibration pulse interval.

[0024] The present invention is further configured such that S3 includes: For each candidate mechanical vibration pulse interval, the first signal segment and the second signal segment corresponding to the candidate mechanical vibration pulse interval in time are extracted from the preprocessed bone conduction vibration signal and air conduction audio signal, respectively, and the time domain cross-correlation coefficient between the first signal segment and the second signal segment is calculated. Based on a preset basic correlation threshold, a preset adjustment coefficient, and the confidence score corresponding to the candidate mechanical vibration pulse interval, a dynamic confirmation threshold is calculated. The dynamic confirmation threshold is calculated by adding the basic correlation threshold to an adjustment term, wherein the adjustment term is the product of the adjustment coefficient and the value 1 minus the confidence score. The time-domain cross-correlation coefficient is compared with the dynamic confirmation threshold. If the time-domain cross-correlation coefficient is less than the dynamic confirmation threshold, the candidate mechanical vibration pulse interval is confirmed as the real mechanical vibration pulse interval, and the real mechanical vibration pulse interval and its corresponding confidence score are output. Specifically, based on the start time, end time, and corresponding confidence score recorded for each candidate mechanical vibration pulse interval output in step S2, the start time and end time are converted into precise sampling point indices according to the signal sampling rate. This conversion is achieved by multiplying the time value by the signal sampling rate. Based on the sampling point index, the bone conduction vibration signal and air conduction vibration signal, which have been preprocessed in step S1 and maintained strict time synchronization, are... In the audio signal, two signal segments that are perfectly aligned in time are extracted. The segment extracted from the bone conduction vibration signal is defined as the first signal segment, and the segment extracted from the air conduction audio signal is defined as the second signal segment. The first and second signal segments completely overlap in time, jointly representing the synchronization signals collected through different physical sensing paths within the same time period. The time-domain cross-correlation coefficient between the first and second signal segments is calculated. This time-domain cross-correlation coefficient is a statistical measure used to quantify the degree of linear correlation and waveform similarity between two discrete-time signal sequences. The calculation process is as follows: First, the covariance between the two signal segments is calculated, that is, the covariance of each segment is calculated separately. The covariance is calculated by taking the difference between all sampled values ​​in a signal segment and their respective arithmetic means. Then, corresponding points from the two sets of difference sequences are multiplied, and the sum of all products is divided by the total number of sampled points. This result represents the covariance, which characterizes the consistency in direction and amplitude of the deviations of two signals from their respective means. Next, the standard deviations of the first and second signal segments are calculated. The standard deviation is calculated by taking the average of the squares of the differences between all sampled values ​​and their arithmetic means in each signal segment, and then taking the square root of this average. This standard deviation characterizes the dispersion or fluctuation intensity of the signal amplitude around its mean. Finally, the calculated covariance value is divided by the first signal segment... The product of the standard deviation and the standard deviation of the second signal segment yields the time-domain cross-correlation coefficient. The calculation of the dynamic confirmation threshold adopts an adaptive mechanism. Its core logic lies in dynamically adjusting the judgment strictness of the correlation test of the dual-channel signal based on the confidence score corresponding to the candidate mechanical vibration pulse interval. Specifically, the dynamic confirmation threshold is obtained by adding the preset basic correlation threshold to an adjustment term. The preset basic correlation threshold represents the benchmark threshold for distinguishing between low-correlation mechanical noise and high-correlation speech events in typical scenarios. The adjustment term is calculated by multiplying a preset adjustment coefficient by a value of 1 and subtracting the difference in the confidence score corresponding to the candidate mechanical vibration pulse interval.According to this calculation rule, when the confidence score of a candidate mechanical vibration pulse interval increases, the value of the adjustment term decreases, leading to a lower dynamic confirmation threshold. A lower dynamic confirmation threshold means a more relaxed correlation judgment condition for the candidate mechanical vibration pulse interval, allowing it to be classified as a mechanical vibration pulse even when the dual-channel signals within that interval exhibit relatively high correlation. Conversely, when the confidence score decreases, the value of the adjustment term increases, leading to a higher dynamic confirmation threshold. An higher dynamic confirmation threshold means a tighter correlation judgment condition for the candidate mechanical vibration pulse interval, requiring the dual-channel signals within that interval to exhibit lower correlation to be classified as a mechanical vibration pulse. This will affect the determination of the correlation between the candidate mechanical vibration pulse interval and the dynamic confirmation threshold. The time-domain cross-correlation coefficient calculated for the vibration pulse interval is compared with its corresponding dynamic confirmation threshold. If the time-domain cross-correlation coefficient is less than the dynamic confirmation threshold, it indicates that the bone conduction vibration signal and the air conduction audio signal show low correlation within the candidate mechanical vibration pulse interval. This characteristic is consistent with the features exhibited by real mechanical vibration pulse noise in dual-channel signals. Based on this, the candidate mechanical vibration pulse interval is confirmed as a real mechanical vibration pulse interval, and the real mechanical vibration pulse interval and its corresponding confidence score are output. If the time-domain cross-correlation coefficient is greater than or equal to the dynamic confirmation threshold, it indicates that the bone conduction vibration signal and the air conduction audio signal show high correlation within the candidate mechanical vibration pulse interval, and the candidate mechanical vibration pulse interval is discarded.

[0025] The present invention is further configured such that S4 includes: For each real mechanical vibration pulse interval, the preprocessed air conduction audio signal of its corresponding time period and a preset duration before the start time of that time period is used as the reference input signal, and the preprocessed bone conduction vibration signal of its corresponding time period is used as the main input signal. Based on the confidence scores corresponding to the actual mechanical vibration pulse intervals, a step size parameter is set for the adaptive filter. The step size parameter is the sum of a preset minimum step size value and an adjustment amount. The adjustment amount is the product of the difference between the preset maximum step size value and the minimum step size value and the confidence score. The present invention is further configured such that the adaptive filter is a normalized minimum mean square adaptive filter. Based on the reference input signal and the main input signal, an adaptive filter with the step size parameter is used to perform filtering operations; Subtracting the output signal of the adaptive filter from the main input signal yields the suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse range; During periods outside the actual mechanical vibration pulse range, the pre-processed bone conduction vibration signal is used as the output signal for that period and is maintained. The suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse interval is concatenated with the bone conduction vibration signal maintained outside the actual mechanical vibration pulse interval in a time sequence to obtain an intermediate processed signal. Specifically, for each actual mechanical vibration pulse interval output in step S3, a signal segment corresponding to the time period is extracted from the bone conduction vibration signal preprocessed in step S1 according to its recorded start and end times, and this signal segment is defined as the main input signal. This signal contains the mechanical pulse noise to be suppressed and the target speech components that may coexist. Simultaneously, a segment corresponding to the main input is extracted from the preprocessed air conduction audio signal that is strictly time-synchronized with the bone conduction vibration signal. The signal segments with completely overlapping time periods are used, and an additional historical signal segment starting a predetermined duration before the start of the current time period is extracted. These two temporally consecutive air-conducted audio signal segments are collectively defined as the reference input signal. The purpose of introducing this additional historical signal segment is to provide a finite-duration initial convergence period for the subsequently activated adaptive filter algorithm. This allows the filter to use this historical signal to pre-learn and construct an initial estimation model of the linear transfer relationship between the reference input and the main input. This ensures that when the signal processing process enters the initial boundary of the real mechanical vibration pulse range, the filter coefficients have reached a basically stable state and possess preliminary effective noise estimation capabilities. This avoids the problem of insufficient noise suppression in the pulse initiation stage caused by filter convergence delay. This embodiment uses a normalized least mean square adaptive filter algorithm for noise suppression, where the step size parameter is the core control variable. The value of this step size parameter determines the rate and step size of filter coefficient updates, thus affecting the algorithm's convergence speed and the estimation error accuracy after reaching steady state. To achieve differentiated processing based on pulse confidence level, the step size parameter is configured to adaptively adjust according to the confidence score corresponding to the currently processed real mechanical vibration pulse interval. Its specific calculation and setting logic is as follows: For each real mechanical vibration pulse interval, first calculate the preset maximum step size value and the preset... The arithmetic difference between the minimum step size values ​​is then multiplied by the confidence score corresponding to the current real mechanical vibration pulse interval to obtain an adjustment amount. Finally, the preset minimum step size value is summed with this adjustment amount, and the sum is set as the final step size parameter applied to process the current real mechanical vibration pulse interval. According to this calculation rule, when the confidence score corresponding to the real mechanical vibration pulse interval is high, the calculated adjustment amount is increased, so that the final step size parameter approaches the preset maximum step size value. This setting will drive the adaptive filter to converge at a faster coefficient update rate and apply stronger suppression to the estimated noise components to eliminate high-confidence mechanical pulse noise.Conversely, when the confidence score corresponding to the real mechanical vibration pulse interval is low, the calculated adjustment amount decreases, causing the final step size parameter to approach the preset minimum step size value. This setting will cause the filter to update its coefficients at a more conservative and gradual pace, thereby gently suppressing the noise present in the real mechanical vibration pulse interval while minimizing damage to the useful speech signal components coexisting in that interval, achieving an optimal balance between noise suppression performance and speech signal fidelity. Based on the reference input signal and main input signal prepared in the aforementioned steps, the step size parameter is driven to complete the preset normalized minimum mean square adaptive filter based on the confidence score corresponding to the current real mechanical vibration pulse interval. The core operating mechanism of this filter is to continuously analyze the reference input signal and dynamically generate an optimal estimate of the component in the main input signal that has a linear correlation with the current reference input signal using an internal set of filter coefficients that are iteratively updated according to the minimum mean square error criterion. Within the actual mechanical vibration pulse range, since the hinge vibration energy is mainly conducted directly to the bone conduction sensor through the solid structure of the equipment, and the proportion of energy propagating through the air acoustic path to the air microphone is extremely low, these two signals exhibit low correlation in characterizing the mechanical pulse noise component. Under these physical conditions, the linear correlation component estimated by the adaptive filter mainly corresponds to the common signal from both sensors. The system incorporates ambient background noise and a portion of the speech signal slightly coupled to the bone conduction sensor via airborne sound leakage or secondary radiation from structural vibrations. This estimated signal, output in real-time by the adaptive filter, is then subtracted point-by-point from the original main input signal. The resulting difference is the noise-suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse range. In this suppressed bone conduction vibration signal, components linearly correlated with the reference input signal are effectively suppressed, thus achieving targeted reduction of mechanical pulse noise. Simultaneously, components uncorrelated or nonlinearly related to the reference input signal are well preserved, allowing the actual mechanical pulse noise itself to be highlighted. For all time periods outside the actual mechanical vibration pulse intervals on the time axis, the corresponding preprocessed bone conduction vibration signals are not processed in any way and are directly used as the output signals for that time period. After completing the adaptive noise suppression processing for all actual mechanical vibration pulse intervals, the suppressed bone conduction vibration signals corresponding to each actual mechanical vibration pulse interval and the maintained bone conduction vibration signals corresponding to all non-actual mechanical vibration pulse intervals are seamlessly spliced ​​and connected according to their temporal relationship in the original preprocessed bone conduction vibration signals. Through this integration process, a continuous and uninterrupted signal sequence in the time domain is generated. This signal sequence is the intermediate processing signal output in step S4.This intermediate processed signal effectively suppresses the identified mechanical transient interference noise while preserving, to the maximum extent possible, the speech information and other non-mechanical impulse signal components contained in the original preprocessed bone conduction vibration signal.

[0026] The present invention is further configured such that S5 includes: For each real mechanical vibration pulse interval, the width of the smooth transition band at its front and rear boundaries is determined based on its corresponding confidence score. The smooth transition band width is calculated based on a preset minimum transition band length, a preset maximum transition band length, and the confidence score. The smooth transition band width is the sum of the minimum transition band length and an expansion amount. The expansion amount is the rounded value of the product of the difference between the maximum transition band length and the minimum transition band length and the confidence score. Based on the start and end times of the actual mechanical vibration pulse interval, the duration of the smooth transition band width is extended forward and backward respectively to form a left smooth transition band and a right smooth transition band; Within the left smooth transition band, a first time-varying gain linearly increases from 0 to 1 is applied to the intermediate processed signal, and a second time-varying gain linearly decreases from 1 to 0 is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the left smooth transition band. Within the right smooth transition band, a third time-varying gain, linearly decreasing from 1 to 0, is applied to the intermediate processed signal, and a fourth time-varying gain, linearly increasing from 0 to 1, is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the right smooth transition band. Within the actual mechanical vibration pulse range, the intermediate processed signal is used as the output signal for that actual mechanical vibration pulse range; In other time periods outside the left smooth transition zone, the real mechanical vibration pulse interval, and the right smooth transition zone, the preprocessed bone conduction vibration signal is used as the output signal for that time period. The output signals corresponding to the actual mechanical vibration pulse interval, the left smooth transition band, the right smooth transition band, and other time periods are spliced ​​sequentially to obtain the noise-reduced bone conduction vibration signal. Specifically, for each actual mechanical vibration pulse interval, its corresponding smooth transition band width is calculated. This smooth transition band width defines the time domain range for signal mixing before and after the actual mechanical vibration pulse interval. The calculation logic for the smooth transition band width of the currently processed actual mechanical vibration pulse interval is as follows: Calculate the arithmetic difference between the preset maximum transition band length and the preset minimum transition band length, and then compare this difference with the set value corresponding to the current actual mechanical vibration pulse interval. The confidence score is multiplied to obtain a preliminary expansion value. This preliminary expansion value is then rounded to obtain an integer expansion value. The preset minimum transition band length is summed with this rounded expansion value. The summation result is the final smooth transition band width determined and applied to the current real mechanical vibration pulse range. This calculation rule establishes a positive correlation between the confidence score and the smooth transition band width. Its physical meaning is that for real mechanical vibration pulse ranges with higher confidence scores, they are more likely to be strongly suppressed by the adaptive filter in step S4, thus resulting in suppressed bone conduction at the edges of the real mechanical vibration pulse range. There is a more significant amplitude or spectral jump between the vibration signal and the original preprocessed bone conduction vibration signal. Therefore, a wider smooth transition band is needed to perform slow, gradual signal fusion, thereby achieving a smooth and natural auditory transition. Based on the smooth transition band width calculated independently for each real mechanical vibration pulse interval, a corresponding smooth transition region is constructed on its time boundary. Specifically, taking the start time of the real mechanical vibration pulse interval as a reference, a time range equal to the width of its smooth transition band is traced back along the time axis. The time range covered by this traversal is defined as the left smooth transition band of the real mechanical vibration pulse interval. Based on the end time of the vibration pulse interval, the time range is extended backward along the time axis by a duration equal to the width of its smooth transition band. The time range covered by this extension is defined as the right smooth transition band of the real mechanical vibration pulse interval. Thus, the processing range of each real mechanical vibration pulse interval on the time axis is extended to consist of three sequentially connected parts: the left smooth transition band, the real mechanical vibration pulse interval itself, and the right smooth transition band. Time-varying gain smoothing is performed within the left smooth transition band. This processing is used to achieve a seamless transition from fully using the preprocessed bone conduction vibration signal output from step S1 to fully using the intermediate processed signal output from step S4.To this end, two linear time-varying gain sequences with the same length as the duration of the left smooth transition band are generated. The first time-varying gain sequence is applied to modulate the intermediate processed signal. Its gain coefficient is set to 0 at the beginning of the left smooth transition band and then linearly and monotonically increases from 0 throughout the duration of the left smooth transition band until it reaches a maximum value of 1 at the end of the left smooth transition band, which is also the beginning of the actual mechanical vibration pulse interval. The second time-varying gain sequence is applied to modulate the preprocessed bone conduction vibration signal. Its gain coefficient is set to 1 at the beginning of the left smooth transition band and then linearly and monotonically decreases from 1 throughout the duration of the left smooth transition band until it drops to 0 at the end of the left smooth transition band. For each discrete sampling time point within the left smooth transition band, the following operations are performed synchronously: The sampled value of the intermediate processed signal at that time point is multiplied by the gain coefficient of the first time-varying gain sequence corresponding to that time point; simultaneously, the sampled value of the preprocessed bone conduction vibration signal at that time point is multiplied by the gain coefficient of the second time-varying gain sequence corresponding to that time point. These two products are then algebraically summed, and the sum is the output signal value after smoothing and mixing at that sampling time point. Time-varying gain smoothing is performed within the right smooth transition band. This process is logically symmetrical to the processing in the left smooth transition band, but the gain change trend is opposite. Therefore, two new linear time-varying gain sequences need to be generated, where the first... Three time-varying gain sequences are applied to the intermediate processed signal output from modulation step S4. Their gain coefficients are set to 1 at the start time of the right smooth transition band (i.e., the end time of the actual mechanical vibration pulse interval), and linearly and monotonically decrease from 1 at the start time throughout the entire duration of the right smooth transition band until the gain coefficient drops to 0 at the end time of the right smooth transition band. A fourth time-varying gain sequence is applied to the preprocessed bone conduction vibration signal output from modulation step S1. Its gain coefficient is set to 0 at the start time of the right smooth transition band, and linearly and monotonically increases from 0 at the start time throughout the entire duration of the right smooth transition band until the gain coefficient reaches its maximum value of 1 at the end time of the right smooth transition band. For each discrete sampling time point within the right smooth transition band, the following operations are performed synchronously: the sampled value of the intermediate processed signal at that time point is obtained and multiplied by the gain coefficient of the third time-varying gain sequence at that time point; simultaneously, the sampled value of the preprocessed bone conduction vibration signal at that time point is obtained and multiplied by the gain coefficient of the fourth time-varying gain sequence at that time point; then, the two product results are algebraically added, and the sum is the output signal value of that sampling time point after smoothing and mixing; within the time period covered by the real mechanical vibration pulse interval, no signal mixing processing is performed, and the intermediate processed signal output in step S4 is directly used as the output signal for that time period;For all other time periods on the time axis that do not belong to any real mechanical vibration pulse interval and its corresponding left and right smooth transition bands, the preprocessed bone conduction vibration signal output in step S1 is directly used as the output signal for that time period and is maintained. The output signals of all left smooth transition bands, all real mechanical vibration pulse intervals, all right smooth transition bands, and all other time periods are seamlessly connected and merged according to the temporal relationship of each output signal segment on the original global time axis, thereby constructing a continuous and uninterrupted signal sequence in the time domain. This signal sequence is the denoised bone conduction vibration signal obtained after the complete denoising and smoothing processing of this invention.

[0027] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A bone conduction-based folding recording noise reduction method, characterized in that, include: S1: Synchronously acquire bone conduction vibration signals, air conduction audio signals, triaxial acceleration signals, and hinge status signals, and preprocess the acquired four signals; S2: Based on the preprocessed bone conduction vibration signal, triaxial acceleration signal, hinge state signal and historical pulse interval information, calculate the confidence of multimodal transient interference, and identify and output candidate mechanical vibration pulse intervals and their corresponding confidence scores accordingly. S3: Based on the candidate mechanical vibration pulse intervals and their corresponding confidence scores, and combined with the correlation between the preprocessed bone conduction vibration signal and the air conduction audio signal within the candidate mechanical vibration pulse intervals, the candidate mechanical vibration pulse intervals are confirmed, and the real mechanical vibration pulse intervals and their corresponding confidence scores are output. S4: Based on the real mechanical vibration pulse range and its corresponding confidence score, the preprocessed bone conduction vibration signal is subjected to adaptive filtering and suppression processing within the real mechanical vibration pulse range to obtain the intermediate processed signal. S5: Based on the real mechanical vibration pulse range and its corresponding confidence score, the intermediate processed signal is smoothed at the boundary of the real mechanical vibration pulse range to obtain the noise-reduced bone conduction vibration signal.

2. The bone conduction-based folding sound recording noise reduction method according to claim 1, wherein, S1 includes: Simultaneously acquire bone conduction vibration signals output by bone conduction sensors, air conduction audio signals output by air microphones, triaxial acceleration signals output by triaxial accelerometers, and hinge status signals output by hinge status sensors; The four acquired signals are input to a multi-channel synchronous sample-and-hold analog-to-digital converter for synchronous sampling and analog-to-digital conversion to obtain time-aligned digital bone conduction vibration signals, digital air conduction audio signals, digital triaxial acceleration signals, and digital hinge status signals. High-pass filters with the same preset cutoff frequency are applied to the digital bone conduction vibration signal and the digital air conduction audio signal respectively to obtain the preprocessed bone conduction vibration signal and air conduction audio signal. The digital triaxial acceleration signal is multiplied by calibration coefficients and zero bias calibration is performed to obtain a triaxial acceleration value sequence as the preprocessed triaxial acceleration signal. The digital hinge state signal is subjected to binarization event detection processing to obtain a binary event sequence representing the hinge rotation state as the preprocessed hinge state signal.

3. The foldable recording noise reduction method based on bone conduction according to claim 1, characterized in that, S2 includes: The preprocessed bone conduction vibration signal is then subjected to frame segmentation. For each signal frame, calculate the confidence level of the audio feature sub-score based on the preprocessed bone conduction vibration signal corresponding to that signal frame; For each signal frame, the device state sub-confidence is calculated based on the preprocessed triaxial acceleration signal and hinge state signal corresponding to that signal frame; For each signal frame, the confidence level of the historical pulse correlation sub-score is calculated based on a historical pulse interval record queue that stores a preset number of real mechanical vibration pulse intervals that have been confirmed in previous processing cycles. For each signal frame, the audio feature sub-confidence, device status sub-confidence, and historical pulse correlation sub-confidence are weighted and summed to obtain the multimodal transient interference confidence of that signal frame. The confidence level of the multimodal transient interference in each signal frame is compared with a preset detection threshold, and signal frames with a multimodal transient interference confidence level greater than the detection threshold are marked as candidate pulse frames. Clustering of temporally continuous candidate pulse frames generates candidate mechanical vibration pulse intervals. The arithmetic mean of the multimodal transient interference confidence scores of all signal frames within each candidate mechanical vibration pulse interval is used as the confidence score corresponding to that candidate mechanical vibration pulse interval.

4. The bone conduction based folding recording noise reduction method according to claim 3, characterized in that, The calculation process for the audio feature sub-confidence includes: The first feature value is obtained by the ratio of the short-time energy of the preprocessed bone conduction vibration signal corresponding to the signal frame to the average energy of the background noise of a preset number of historical frames. The second feature value is obtained by the absolute value of the difference between the spectral centroid of the preprocessed bone conduction vibration signal corresponding to the signal frame and the mean spectral centroid of the historical frames that were identified as speech frames. The third characteristic value is obtained based on the ratio of the signal energy of the preprocessed bone conduction vibration signal corresponding to the signal frame within the preset mechanical resonance frequency band to the signal energy of the full frequency band. The first, second, and third feature values ​​are mapped using their respective preset monotonically increasing functions, and the results of each feature value mapping are weighted and summed using preset weights to obtain the audio feature sub-confidence.

5. The bone conduction based folding recording noise reduction method according to claim 3, wherein, The calculation process for the device state sub-confidence includes: Based on the triaxial components of the preprocessed triaxial acceleration signal corresponding to the signal frame, the Euclidean norm of the difference vector between the triaxial components and the preset static reference vector is calculated. The calculation result is low-pass filtered to obtain the motion intensity characterization value, and then mapped to the first intermediate value through the Sigmoid function. Based on the preprocessed hinge state signal corresponding to the signal frame, the hinge state signal value within the time period of the signal frame is logically ORed to obtain the second intermediate value of the folding event. The first intermediate value and the second intermediate value are weighted and fused with preset weights to obtain the device status sub-confidence.

6. The folded recording noise reduction method based on bone conduction according to claim 3, characterized in that, The calculation process for the confidence level of the historical pulse correlation sub-element includes: Retrieve a queue of historical pulse interval records containing a preset number of real mechanical vibration pulse intervals that have been confirmed in previous processing cycles; Based on the current signal frame center time and the end time of the most recent real mechanical vibration pulse interval in the historical pulse interval recording queue, the timing decay component is calculated. The timing decay component is an exponential function with the natural constant e as the base. Its exponent is the quotient of the negative difference between the current signal frame time and the end time of the most recent real mechanical vibration pulse interval in the historical pulse interval recording queue and the preset decay time constant. Based on the number of occurrences of real mechanical vibration pulse intervals obtained from the historical pulse interval record queue within a preset time window and the preset maximum possible number of pulses, a frequency statistical component is calculated. The frequency statistical component is the minimum value between the ratio of the number of occurrences to the maximum possible number of pulses and the value 1. The temporal decay component and the frequency statistics component are weighted and fused with preset weights to obtain the confidence level of the historical pulse correlation sub-component.

7. The bone conduction based folding recording noise reduction method according to claim 1, wherein, S3 includes: For each candidate mechanical vibration pulse interval, the first signal segment and the second signal segment corresponding to the candidate mechanical vibration pulse interval in time are extracted from the preprocessed bone conduction vibration signal and air conduction audio signal, respectively, and the time domain cross-correlation coefficient between the first signal segment and the second signal segment is calculated. Based on a preset basic correlation threshold, a preset adjustment coefficient, and the confidence score corresponding to the candidate mechanical vibration pulse interval, a dynamic confirmation threshold is calculated. The dynamic confirmation threshold is calculated by adding the basic correlation threshold to an adjustment term, wherein the adjustment term is the product of the adjustment coefficient and the value 1 minus the confidence score. The time-domain cross-correlation coefficient is compared with the dynamic confirmation threshold. If the time-domain cross-correlation coefficient is less than the dynamic confirmation threshold, the candidate mechanical vibration pulse interval is confirmed as the real mechanical vibration pulse interval, and the real mechanical vibration pulse interval and its corresponding confidence score are output.

8. The bone conduction based folding recording noise reduction method according to claim 1, wherein, S4 includes: For each real mechanical vibration pulse interval, the preprocessed air conduction audio signal of its corresponding time period and a preset duration before the start time of that time period is used as the reference input signal, and the preprocessed bone conduction vibration signal of its corresponding time period is used as the main input signal. Based on the confidence scores corresponding to the actual mechanical vibration pulse range, a step size parameter is set for the adaptive filter. The step size parameter is the sum of a preset minimum step size value and an adjustment amount. The adjustment amount is the product of the difference between the preset maximum step size value and the minimum step size value and the confidence score. Based on the reference input signal and the main input signal, an adaptive filter with the step size parameter is used to perform filtering operations; Subtracting the output signal of the adaptive filter from the main input signal yields the suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse range; During periods outside the actual mechanical vibration pulse range, the pre-processed bone conduction vibration signal is used as the output signal for that period and is maintained. The suppressed bone conduction vibration signal corresponding to the actual mechanical vibration pulse interval is spliced ​​with the bone conduction vibration signal maintained outside the actual mechanical vibration pulse interval in a time sequence to obtain the intermediate processed signal.

9. The bone conduction based folding recording noise reduction method according to claim 1, wherein, S5 includes: For each real mechanical vibration pulse interval, the width of the smooth transition band at its front and rear boundaries is determined based on its corresponding confidence score. The smooth transition band width is calculated based on a preset minimum transition band length, a preset maximum transition band length, and the confidence score. The smooth transition band width is the sum of the minimum transition band length and an expansion amount. The expansion amount is the rounded value of the product of the difference between the maximum transition band length and the minimum transition band length and the confidence score. Based on the start and end times of the actual mechanical vibration pulse interval, the duration of the smooth transition band width is extended forward and backward respectively to form a left smooth transition band and a right smooth transition band; Within the left smooth transition band, a first time-varying gain linearly increases from 0 to 1 is applied to the intermediate processed signal, and a second time-varying gain linearly decreases from 1 to 0 is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the left smooth transition band. Within the right smooth transition band, a third time-varying gain, linearly decreasing from 1 to 0, is applied to the intermediate processed signal, and a fourth time-varying gain, linearly increasing from 0 to 1, is applied to the preprocessed bone conduction vibration signal. The weighted sum of the two is then used as the output signal of the right smooth transition band. Within the actual mechanical vibration pulse range, the intermediate processed signal is used as the output signal for that actual mechanical vibration pulse range; In other time periods outside the left smooth transition zone, the real mechanical vibration pulse interval, and the right smooth transition zone, the preprocessed bone conduction vibration signal is used as the output signal for that time period. The output signals corresponding to the real mechanical vibration pulse interval, the left smooth transition band, the right smooth transition band, and other time periods are spliced ​​together in time sequence to obtain the noise-reduced bone conduction vibration signal.

10. The bone conduction based folding recording noise reduction method according to claim 8, wherein, The adaptive filter is a normalized minimum mean square adaptive filter.