A multi-mode voice processing system and method for an earbud earphone

By using a multi-mode voice processing system, combined with an air conduction microphone and a bone conduction sensor, the voice processing mode is dynamically selected, solving the problem of insufficient voice clarity in ear clip-on headphones under different noise environments, and achieving high-quality calls in various noise environments.

CN122340404APending Publication Date: 2026-07-03ZERO WORLD AUDIO (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZERO WORLD AUDIO (SUZHOU) TECHNOLOGY CO LTD
Filing Date
2026-05-20
Publication Date
2026-07-03

Smart Images

  • Figure CN122340404A_ABST
    Figure CN122340404A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of headphone voice processing technology, and specifically relates to a multi-mode voice processing system and method for clip-on headphones. The multi-mode voice processing system for clip-on headphones includes a first voice processing module, a second voice processing module, and a mode selection module. The first voice processing module performs beamforming processing on signals collected by an air conduction microphone. The second voice processing module fuses signals collected by the air conduction microphone and signals collected by a bone conduction sensor. The mode selection module selects the appropriate voice processing module based on real-time detected environmental characteristics and outputs the voice signal generated by the corresponding module as the call audio. This system, by employing a mode selection mechanism to choose the optimal processing method based on real-time environmental characteristics, effectively improves the call audio quality of clip-on headphones in all scenarios, taking into account the characteristics of different noise environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of headphone voice processing technology, specifically relating to a multi-mode voice processing system and method for clip-on headphones. Background Technology

[0002] Clip-on earphones are open-back headphones. Due to their non-sealed design, they are more susceptible to ambient noise during calls, thus reducing voice clarity. Therefore, voice noise reduction processing is necessary for clip-on earphones. Current technologies primarily employ the following two methods for voice noise reduction: First, beamforming technology based on multiple air conduction microphones enhances the target speech through spatial filtering, achieving good speech naturalness in low-to-medium noise environments. Second, a bone conduction sensor (VPU) is introduced, leveraging its insensitivity to environmental noise to improve speech extraction capabilities in high-noise environments.

[0003] However, during the actual research and development process, the applicant discovered that: in medium or low noise environments, the participation of bone conduction signals in fusion may lead to speech distortion; dual-microphone beamforming has limited effectiveness in strong noise environments; different speech processing methods show significant performance differences in different noise environments; and existing systems typically adopt a single processing mode and lack adaptive mechanisms.

[0004] Therefore, it is necessary to provide a system that can dynamically select the voice processing method according to environmental changes in order to improve the call performance of clip-on headphones. Summary of the Invention

[0005] The present invention aims to provide a multi-mode voice processing system and method for ear clip-on headphones, so as to achieve adaptive selection of appropriate voice processing modes according to different environmental characteristics, thereby adapting to different noise environments, improving call voice quality, and obtaining clear and natural call voice in different noise scenarios.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A multi-mode voice processing system for clip-on headphones is provided, wherein the clip-on headphones are equipped with at least two air conduction microphones and one bone conduction sensor, and the multi-mode voice processing system includes: A first voice processing module is used to perform a first signal processing procedure, which performs beamforming processing on the signal collected by the air conduction microphone to form a first output voice signal. The second voice processing module is used to perform a second signal processing process, which fuses the signal collected by the air conduction microphone and the signal collected by the bone conduction sensor to form a second output voice signal. The mode selection module is used to select the output voice signal, and selects between the first voice processing module and the second voice processing module according to environmental characteristics, and adopts the corresponding output voice signal.

[0007] Preferably, the first signal processing procedure specifically includes: The first speech processing module acquires speech signals through at least two air conduction microphones to obtain multi-channel input signals, performs time-frequency analysis processing on the multi-channel input signals, and converts the time-domain signals into time-frequency domain signals for representation. Based on the differences between the multi-channel input signals after time-frequency analysis processing, spatial feature information is extracted, including phase difference, time delay difference, and amplitude difference. According to the spatial feature information, the speech signal in the target direction is enhanced, while the signal in other directions is suppressed. Then, the enhanced and suppressed time-frequency domain signal is converted into the first output speech signal.

[0008] Preferably, the second signal processing procedure specifically includes: The second voice processing unit acquires ambient voice signals through at least one of the air conduction microphones and acquires the wearer's vibration signals through the bone conduction sensor. It preprocesses the ambient voice signals and vibration signals respectively, including gain adjustment, bandwidth limitation, noise reduction, or filtering. The preprocessed environmental speech signal and vibration signal are subjected to time synchronization processing, and the time synchronization processing methods include delay compensation and phase alignment. Speech-related features are extracted from the time-synchronized environmental speech signal and vibration signal. These speech-related features include energy features, spectral features, and temporal envelope. Based on the aforementioned speech-related features, a first fusion operation is performed on the time-synchronized environmental speech signal and vibration signal. The first fusion operation includes weighted fusion, selective enhancement, and / or feature-based fusion processing. The signal after the first fusion operation is reconstructed to obtain the second output speech signal.

[0009] Preferably, the mode selection module can also perform fusion processing on the first output voice signal and the second output voice signal, and perform a second fusion operation on the first output voice signal and the second output voice signal according to weights to obtain the third output voice signal; the mode selection module selects between the first output voice signal, the second output voice signal and the third output voice signal according to environmental characteristics, and adopts the corresponding output voice signal.

[0010] Preferably, it also includes a mode smoothing switching module, which is used to introduce a transition phase during the switching process between the first output voice signal, the second output voice signal and the third output voice signal. During the transition phase, the two output voice signals are fused. During the switching process between one output voice signal and another output voice signal, the two output voice signals are combined according to a certain weight to achieve a smooth transition from one voice signal to another.

[0011] Preferably, the selection of the output speech signal specifically includes: the mode selection module estimates the ambient noise intensity or signal-to-noise ratio information in real time based on the audio signal collected by the air conduction microphone, and selects the corresponding output speech signal according to the ambient noise intensity or signal-to-noise ratio information.

[0012] The present invention also provides a multi-mode speech processing method for clip-on headphones, wherein the clip-on headphones are provided with at least two air conduction microphones and one bone conduction sensor, and the multi-mode speech processing method includes: The mode selection module selects the output voice signal, choosing between a first, second, and third output voice signal based on environmental characteristics, and adopts the corresponding output voice signal. The formation methods of the first, second, and third output voice signals are as follows: The first voice processing module performs a first signal processing procedure, and performs beamforming processing on the signal collected by the air conduction microphone to form a first output voice signal. The second speech processing module is used to perform a second signal processing process, which fuses the signal collected by the air conduction microphone and the signal collected by the bone conduction sensor to form a second output speech signal. The mode selection module is used to fuse the first and second output speech signals to form a third output speech signal.

[0013] Preferably, the first signal processing procedure specifically includes: the first speech processing module acquiring speech signals through at least two air conduction microphones to obtain multi-channel input signals; performing time-frequency analysis processing on the multi-channel input signals to convert the time-domain signals into time-frequency domain signals for representation; extracting spatial feature information based on the differences between the multi-channel input signals, the spatial feature information including phase difference, time delay difference, and amplitude difference; enhancing the speech signal in the target direction and suppressing the signals in other directions according to the spatial feature information; and then converting the enhanced and suppressed time-frequency domain signals into the first output speech signal. The second signal processing procedure specifically includes: the second speech processing unit acquiring ambient speech signals through at least one air conduction microphone and acquiring the wearer's vibration signals through the bone conduction sensor; preprocessing the ambient speech signals and vibration signals respectively, including gain adjustment, bandwidth limitation, noise reduction, or filtering; performing time synchronization processing on the preprocessed ambient speech signals and vibration signals, including delay compensation and phase alignment; extracting speech-related features from the time-synchronized ambient speech signals and vibration signals, including energy features, spectral features, and temporal envelope; performing a first fusion operation on the time-synchronized ambient speech signals and vibration signals based on the speech-related features, including weighted fusion, selective enhancement, and / or feature-based fusion processing; and reconstructing the signal after the first fusion operation to obtain the second output speech signal. The processing of the third output speech signal specifically includes: performing a second fusion operation on the first output speech signal and the second output speech signal according to weights to obtain the third output speech signal.

[0014] Preferred options also include: During the switching process between the first, second, and third output voice signals using the mode smoothing switching module, a transition phase is introduced. In this transition phase, the two output voice signals are fused. That is, during the switching from one output voice signal to another, the two output voice signals are combined according to a certain weight. The weight is a fixed value or is dynamically adjusted with time and environmental changes to achieve a smooth transition from one voice signal to another.

[0015] The present invention also provides an ear clip-on earphone, including a front earphone shell and a rear earphone shell, a C-shaped connecting bridge between the front earphone shell and the rear earphone shell, and electronic devices in the front earphone shell and the rear earphone shell. The electronic devices include at least two air conduction microphones, a bone conduction sensor, a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the method described in any of the preceding claims.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This multi-mode voice processing system for clip-on earphones operates based on at least two air conduction microphones and one bone conduction sensor. Specifically, it includes a first voice processing module, a second voice processing module, and a mode selection module. The first voice processing module performs a first signal processing step, performing beamforming processing on the signals collected by the air conduction microphones to generate a first output voice signal. The second voice processing module performs a second signal processing step, fusing the signals collected by the air conduction microphones and the bone conduction sensor to generate a second output voice signal. The mode selection module selects between the first and second voice processing modules based on real-time detected environmental characteristics, directly outputting the corresponding module's voice signal as the final call voice. This system integrates different processing schemes, such as beamforming processing and air / bone conduction signal fusion, to address the characteristics of different noise environments. Through a mode selection mechanism, it selects the optimal processing method based on real-time environmental characteristics. In low-noise environments, beamforming processing ensures natural voice quality; in high-noise environments, fusion processing improves voice clarity; and in moderate-noise environments, it fuses the two processing results to balance clarity and naturalness, effectively improving the call voice quality of the clip-on earphones in all scenarios. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of an ear-clip earphone, which is part of an embodiment of the multi-mode voice processing system for ear-clip earphones of the present invention. Figure 1 -a) is a schematic diagram of a VPU (bone conduction sensor) placed inside the ear canal. Figure 1 -b) is a schematic diagram of the VPU placed on the motherboard behind the ear.

[0018] Figure 2 This is a flowchart illustrating the voice signal processing of the first voice processing module and the second voice processing module in one embodiment of the multi-mode voice processing system for the ear clip-on earphone of the present invention.

[0019] Figure 3 This is a schematic diagram of mode switching in the mode selection module of one embodiment of the multi-mode voice processing system for the ear clip-on earphone of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] In one embodiment, a multi-mode voice processing system for clip-on headphones is provided, the multi-mode voice processing system for clip-on headphones being based on, for example, Figure 1 The ear-clip earphone 20 shown is implemented, and the ear-clip earphone 20 is worn on the ear 10, as shown. Figure 1 As shown in -a), the ear clip-on earphone 20 includes a front earphone shell 21 and a rear earphone shell 22. A C-shaped connecting bridge 23 is provided between the front earphone shell 21 and the rear earphone shell 22. The front earphone shell 21 is used to be placed inside the ear 10, and the rear earphone shell 22 is used to be placed behind the ear 10. The C-shaped connecting bridge 23 is made of elastic or memory metal. Electronic devices are provided in the front earphone shell 21 and the rear earphone shell 22. The electronic devices include at least two air conduction microphones 24 and a bone conduction sensor 25 (VPU). The air conduction microphones 24 collect sound signals through the air, and the bone conduction sensor 25 collects sound signals through bone vibration. The electronic devices also include a speaker, a wireless module, a processor, and a memory, etc., to support the wireless transmission, playback, and acquisition of sound signals.

[0022] The multi-mode voice processing system of the ear clip-on earphone in this embodiment is a processor-executable program module. The multi-mode voice processing system includes a first voice processing module, a second voice processing module, and a mode selection module.

[0023] Combination Figure 2 As shown, the first voice processing module is used to perform a first signal processing procedure, which performs beamforming processing on the signal collected by the air conduction microphone to form a first output voice signal.

[0024] Beamforming processing involves processing the collected signals to form an output beam (first output voice signal). The first output voice signal can be used during a call as a carrier of what the earbud wearer says, and then sent to the other party in the call, so that the other party can decode it and obtain what the earbud wearer says.

[0025] The first signal processing procedure specifically includes: The first speech processing module acquires speech signals through at least two air conduction microphones 24 to obtain multi-channel input signals, performs time-frequency analysis processing on the multi-channel input signals, and converts the time-domain signals into time-frequency domain signals for representation.

[0026] The speech signals acquired by the two air conduction microphones 24 are processed by time-frequency analysis to obtain the corresponding time-frequency domain signals. The time-frequency domain signals are obtained by using Fourier transform to analyze the time-domain signals that change with time, thereby determining when and at what frequency the signal appeared.

[0027] Because the two air conduction microphones 24 are positioned differently on the ear clip-on headphones, the collected sound information will differ, and this difference can be analyzed through time-frequency domain signal analysis.

[0028] Based on the differences between the multi-channel input signals after time-frequency analysis, spatial feature information is extracted, including phase difference, time delay difference, and amplitude difference. According to the spatial feature information, the speech signal in the target direction is enhanced, while the signal in other directions is suppressed. Then, the enhanced and suppressed time-frequency domain signal is converted into the first output speech signal.

[0029] The phase difference, time delay difference, and amplitude difference here can respectively reflect the differences in the spatial propagation direction of the target speech and interference noise in the sound source of the two air conduction microphones 24. The phase difference reflects the degree of phase shift of the same frequency speech signals received by the two air conduction microphones. The time delay difference corresponds to the time difference of sound propagation to different microphones. The amplitude difference reflects the energy attenuation difference when the sound signals from different directions reach the microphones. The combination of the three can locate the spatial position of the target speaker's sound source, so that targeted speech enhancement and noise suppression can be performed. The speech spoken by the ear clip-on headphone wearer is taken as the target speech, and the speech is enhanced while other environmental noises are attenuated. In low to medium noise environments, the speech naturalness can be preserved. Moreover, this processing method has low complexity and can operate with low power consumption. Therefore, in the default mode, the mode selection module uses the first output speech signal generated by the first speech processing module.

[0030] The second voice processing module is used to perform the second signal processing process, which fuses the signals collected by the air conduction microphone and the signals collected by the bone conduction sensor to form the second output voice signal.

[0031] Similarly, the second output voice signal can also be used during a call as a carrier of what the earbud wearer says, sending it to the other party in the call so that the other party can decode it and obtain what the earbud wearer says.

[0032] The second signal processing procedure specifically includes: The second voice processing unit collects ambient voice signals through at least one air conduction microphone and collects the wearer's vibration signals through a bone conduction sensor. It preprocesses the ambient voice signals and vibration signals separately, including gain adjustment, bandwidth limitation, noise reduction, or filtering.

[0033] Air conduction microphones collect speech signals transmitted through the air, while bone conduction sensors collect signals transmitted through bone vibrations. In preprocessing, gain adjustment is used to unify the amplitudes of both types of signals to a suitable dynamic range, preventing an excessively low signal-to-noise ratio; bandwidth limiting is used to filter out unwanted frequency bands outside the effective range of the speech signal; denoising and filtering are used to pre-filter out noise such as impulse interference in the signal, improving the quality of the input signal.

[0034] Because air conduction signals (environmental speech signals collected by air conduction microphones) and bone conduction signals (vibration signals collected by bone conduction sensors) have different propagation paths, after preprocessing, the preprocessed environmental speech signals and vibration signals are time-synchronized. The time synchronization methods include delay compensation and phase alignment.

[0035] Because air conduction sound signals and bone conduction vibration signals propagate at different speeds, there is a fixed time difference between the two types of signals from the time of sound emission to the time of acquisition. Delay compensation and phase alignment can align the same frame of speech signals in the time dimension, avoiding speech distortion caused by phase cancellation in subsequent fusion processing.

[0036] Next, speech-related features are extracted from the time-synchronized environmental speech and vibration signals. These features include energy features, spectral features, and temporal envelope. Based on these features, effective speech and environmental noise can be distinguished, and the quality of which part of the bone conduction signal and air conduction signal can be determined.

[0037] Among them, energy features can intuitively reflect the overall energy level of the signal. The energy of effective speech is usually stable within a specific range, while the energy fluctuation of environmental noise is often greater. By judging the energy threshold, pure noise segments without speech can be initially screened out. Spectral features can reflect the difference in frequency distribution between the two types of signals. Bone conduction signals have a more obvious attenuation of high-frequency components, while air conduction signals retain more high-frequency components. By comparing the spectral differences, it can be distinguished which signal has a higher signal-to-noise ratio in the current environment. The temporal envelope reflects the amplitude change trend of the speech signal in the time domain. The envelope of effective speech has a clear change pattern of voiced and unvoiced sounds, while the envelope of noise is more gradual and irregular. Combining these three types of features can provide an accurate basis for subsequent fusion processing.

[0038] Based on speech-related features, a first fusion operation is performed on the time-synchronized environmental speech signal and vibration signal. The first fusion operation includes weighted fusion, selective enhancement, and / or feature-based fusion processing.

[0039] Weighted fusion can be performed based on fixed weight values ​​or dynamic weight values. In dynamic weighted fusion, the fusion weights are calculated based on the signal-to-noise ratios (SNRs) of the two types of signals. Signals with higher SNRs are assigned greater weights, and vice versa, resulting in a fused signal with weighted combinations.

[0040] Selective enhancement targets the quality of signals in different frequency bands, retaining the higher-quality frequency bands in a particular signal while suppressing the lower-quality frequency bands. For example, it preserves the clean speech components at the low-frequency end of bone conduction signals and the rich speech details at the high-frequency end of air conduction signals, and then combines the preferred components of different frequency bands into a complete signal.

[0041] Feature-based fusion processing determines whether the current signal is valid speech based on the extracted features, fuses only the parts determined to be valid speech, and directly suppresses pure noise frames to improve the signal-to-noise ratio of the output signal.

[0042] After the first fusion operation is completed, the fused signal is inversely transformed and reconstructed in the time domain to obtain the final second output speech signal. This processing method can effectively preserve the wearer's speech information and avoid environmental noise masking the target speech in high-noise environments by taking advantage of the anti-environmental interference characteristics of bone conduction sensors.

[0043] The mode selection module is used to select the output speech signal. Based on environmental characteristics, it selects between the first speech processing module and the second speech processing module and adopts the corresponding output speech signal. The specific process of selecting the output speech signal includes: the mode selection module estimates the environmental noise intensity or signal-to-noise ratio information in real time based on the audio signal collected by the air conduction microphone, and selects the corresponding speech processing module based on the environmental noise intensity or signal-to-noise ratio information.

[0044] Combination Figure 3 As shown, the specific selection logic of the mode selection module is as follows: First, the noise energy level of the current environment is detected in real time to obtain the environmental noise parameters. When the environmental noise parameters are detected to be lower than the switching threshold, it is determined that the current environment is a low-noise environment, and the first output speech signal is selected as the output. Only air-guided beamforming processing is used to ensure the naturalness of the speech and reduce the system's operating power consumption. When the environmental noise parameters are detected to be higher than the switching threshold, the second output speech signal is selected as the output to give full play to the bone conduction sensor's ability to resist environmental interference and ensure that the target speech can be effectively extracted.

[0045] Furthermore, the mode selection module can also perform fusion processing on the first output speech signal and the second output speech signal, and perform a second fusion operation on the first output speech signal and the second output speech signal according to weights to obtain the third output speech signal; the mode selection module selects between the first output speech signal, the second output speech signal and the third output speech signal according to environmental characteristics, and adopts the corresponding output speech signal.

[0046] Similarly, the third output voice signal can also be used during a call as a carrier of what the earbud wearer says, sending it to the other party so that the other party can decode it and obtain what the earbud wearer says.

[0047] The processing of the third output speech signal specifically includes: performing a second fusion operation on the first output speech signal and the second output speech signal according to weights to obtain the third output speech signal.

[0048] In the second fusion operation based on weights, the initial weights corresponding to the first and second output speech signals are first calculated based on the noise energy of the current environment. The higher the noise energy, the higher the weight ratio of the second output speech signal; the lower the noise energy, the higher the weight ratio of the first output speech signal. Then, the initial weights are corrected based on the cross-correlation between the two output speech signals. The lower the cross-correlation, the more severe the noise interference in one of the signals, and the weight of that signal is reduced accordingly. Finally, the corrected fusion weights are obtained, and the weighted fusion is completed to obtain the third output speech signal.

[0049] The third output speech signal processing is designed for medium-to-high noise environments, where the target speech and environmental noise come from different directions in complex scenarios. In such cases, the first signal processing alone cannot completely suppress strong environmental noise, while the second signal processing alone is prone to losing high-frequency details of the speech, leading to distortion. Therefore, the output results of the two processing processes are further fused according to weights, combining their advantages. This approach utilizes the anti-interference capability of bone conduction signals to lock onto the target speech while preserving the high-frequency details of the speech retained after air conduction beamforming processing, thus balancing speech clarity and naturalness.

[0050] The mode selection module selects from the first, second, and third output voice signals based on environmental characteristics, and adopts the corresponding output voice signal.

[0051] At this point, two switching thresholds are set. When determining the environmental state, the current scene is classified or graded according to environmental characteristics, into low-noise, medium-noise, and high-noise environments. When the environmental noise level is between the two switching thresholds, it is determined that the current scene is a complex medium-to-high noise environment, and the third output speech signal is selected as the final output.

[0052] Furthermore, the multi-mode voice processing system of the ear clip-on headset also includes a mode smoothing switching module. The mode smoothing switching module is used to introduce a transition phase during the switching process between the first output voice signal, the second output voice signal and the third output voice signal. During the transition phase, the two output voice signals are fused to achieve a smooth transition from one voice signal to another. The fusion operation specifically includes: combining the two output voice signals according to a certain weight during the switching process from one output voice signal to another.

[0053] During the combination process, the weighted fusion weights can be dynamically adjusted over time. First, the output signals to be switched out and the output signals to be switched in are time-aligned to ensure that the two signals are synchronized in the time dimension. Then, the corresponding weights are set according to the current switching stage. In the initial stage of switching, the signal to be switched out has a higher weight and the signal to be switched in has a lower weight. As the switching process progresses, the weight of the signal to be switched out is gradually reduced, and the weight of the signal to be switched in is increased accordingly. Until the switching is completed, all weights are allocated to the output signal to be switched in. Finally, the voice signal with the weights combined is output to avoid voice jumps, truncation, or obvious auditory differences during the switching process, ensuring a smooth and natural auditory experience during the call.

[0054] When the mode selection module collects environmental features, in addition to environmental noise intensity or signal-to-noise ratio information, it can also include information such as background noise type, multi-channel signal stability, and speech activity status. The weights in the weighted fusion can also be dynamically adjusted based on factors such as environmental noise intensity, signal-to-noise ratio, speech activity status, and signal quality to achieve a smooth transition between the two speech processing results. When the environmental noise intensity changes slowly near the switching threshold, the weights will be gradually fine-tuned to follow the changes in noise intensity, avoiding frequent switching of processing modes due to small fluctuations in noise. When it is detected that the current voice activity segment is in progress, the rate of weight adjustment is slowed down to avoid obvious changes in auditory perception during speech playback. When it is detected that the current voice gap segment is in progress, the rate of weight adjustment is accelerated to quickly complete the mode switching. At the same time, the weights can also be dynamically corrected according to the real-time quality of the two output signals to further ensure the stability of the output speech during the transition.

[0055] In one embodiment, a multi-mode speech processing method for an earcup-type headset is provided. The earcup-type headset on which the method is based is equipped with at least two air conduction microphones and one bone conduction sensor. The multi-mode speech processing method includes: The mode selection module selects the output voice signal, choosing from the first, second, and third output voice signals based on environmental characteristics, and adopts the corresponding output voice signal; wherein: The first speech processing module performs a first signal processing procedure, which involves beamforming based on the signals collected by the air conduction microphones to generate a first output speech signal. The first signal processing procedure specifically includes: the first speech processing module collects speech signals through at least two air conduction microphones to obtain multi-channel input signals; performs time-frequency analysis processing on the multi-channel input signals to convert the time-domain signals into time-frequency domain signals for representation; extracts spatial feature information based on the differences between the multi-channel input signals, including phase difference, time delay difference, and amplitude difference; enhances the speech signal in the target direction based on the spatial feature information while suppressing signals in other directions; and then converts the enhanced and suppressed time-frequency domain signals into the first output speech signal.

[0056] The second speech processing module is used to perform the second signal processing process, which fuses the signal collected by the air conduction microphone and the signal collected by the bone conduction sensor to form the second output speech signal. The second signal processing procedure specifically includes: the second speech processing unit acquires ambient speech signals through at least one air conduction microphone and acquires the wearer's vibration signals through a bone conduction sensor; preprocesses the ambient speech signals and vibration signals respectively, including gain adjustment, bandwidth limitation, noise reduction, or filtering; performs time synchronization processing on the preprocessed ambient speech signals and vibration signals, including delay compensation and phase alignment; extracts speech-related features from the time-synchronized ambient speech signals and vibration signals, including energy features, spectral features, and temporal envelope; performs a first fusion operation on the time-synchronized ambient speech signals and vibration signals based on the speech-related features, including weighted fusion, selective enhancement, and / or feature-based fusion processing; and reconstructs the signal after the first fusion operation to obtain a second output speech signal.

[0057] The third output speech signal processing procedure is executed by the mode selection module, which fuses the first output speech signal and the second output speech signal to form the third output speech signal. The third output speech signal processing procedure specifically includes: performing a second fusion operation on the first output speech signal and the second output speech signal according to weights to obtain the third output speech signal.

[0058] Furthermore, the multi-mode voice processing method for this clip-on earphone also includes: During the switching process between the first, second, and third output voice signals by the mode smoothing switching module, a transition phase is introduced. During the transition phase, the two output voice signals are fused to achieve a smooth transition from one voice signal to another. The fusion operation specifically includes: during the switching process from one output voice signal to another, the two output voice signals are combined according to a certain weight. The weight is a fixed value or dynamically adjusted with time and environmental changes.

[0059] The multi-mode voice processing method of clip-on earphones can automatically select the most suitable voice processing mode according to the actual situation of the current environmental noise. In low-noise environments, a low-power air conduction single processing mode is activated, which reduces the power consumption of the earphone while maintaining voice clarity and extending the battery life of the clip-on earphones. In high-noise environments, a bone conduction and air conduction fusion processing mode is activated, which utilizes the anti-environmental interference characteristics of bone conduction signals to effectively preserve target voice information and significantly improve call clarity in complex noise environments.

[0060] Through the mode selection module, in complex scenarios with medium to high noise, it combines the advantages of air conduction processing and bone conduction fusion processing. It can rely on bone conduction signals to lock onto the target speech and suppress environmental noise, while also preserving the rich high-frequency details brought by air conduction signals, balancing the clarity and naturalness of the speech and improving the listening experience of the other party in the call.

[0061] Through a smooth mode switching mechanism, voice jumps, cutoffs, or significant auditory drops are avoided during the switching between different processing modes. Furthermore, the switching rhythm can be dynamically adjusted according to changes in ambient noise and voice activity status to ensure a smooth and natural switching process. The switching can be completed quickly at the appropriate time, improving the stability and comfort of the overall call experience.

[0062] In one embodiment, an ear clip-on headphone is provided, combined with Figure 1 As shown, the ear-clip earphone 20 includes a front earphone shell 21 and a rear earphone shell 22, with a C-shaped connecting bridge 23 between them. The front earphone shell 21 is positioned inside the ear 10, and the rear earphone shell 22 is positioned behind the ear 10. The C-shaped connecting bridge 23 is made of elastic or shape memory metal. Electronic devices are housed in the front and rear earphone shells 21 and 22, including at least two air conduction microphones 24 and a bone conduction sensor 25 (VPU). The air conduction microphones 24 collect sound signals through the air, and the bone conduction sensor 25 collects sound signals through bone vibration. The electronic devices also include a speaker, a wireless module, a processor, and a memory, which support wireless transmission, playback, and acquisition of sound signals. A computer program is stored in the memory, and when the processor executes the computer program, it implements the multi-mode voice processing method for the ear-clip earphone in the previous embodiment.

[0063] Combination Figure 1 As shown, the bone conduction sensor (VPU) can be placed in the front earpiece shell. Figure 1 -a) can also be set in the rear earpiece shell ( Figure 1 -b)). The air conduction microphones are all located in the rear earpiece shell, with two air conduction microphones spaced apart to collect sound signals from different locations.

[0064] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0065] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-mode voice processing system for clip-on headphones, characterized in that, The clip-on earphone is equipped with at least two air conduction microphones and one bone conduction sensor. The multi-mode voice processing system includes: A first voice processing module is used to perform a first signal processing procedure, which performs beamforming processing on the signal collected by the air conduction microphone to form a first output voice signal. The second voice processing module is used to perform a second signal processing process, which fuses the signal collected by the air conduction microphone and the signal collected by the bone conduction sensor to form a second output voice signal. The mode selection module is used to select the output voice signal, and selects between the first voice processing module and the second voice processing module according to environmental characteristics, and adopts the corresponding output voice signal.

2. The multi-mode voice processing system for clip-on headphones according to claim 1, characterized in that, The first signal processing procedure specifically includes: The first speech processing module acquires speech signals through at least two air conduction microphones to obtain multi-channel input signals, performs time-frequency analysis processing on the multi-channel input signals, and converts the time-domain signals into time-frequency domain signals for representation. Based on the differences between the multi-channel input signals after time-frequency analysis processing, spatial feature information is extracted, including phase difference, time delay difference, and amplitude difference. According to the spatial feature information, the speech signal in the target direction is enhanced, while the signal in other directions is suppressed. Then, the enhanced and suppressed time-frequency domain signal is converted into the first output speech signal.

3. The multi-mode voice processing system for ear-clip headphones according to claim 1, characterized in that, The second signal processing procedure specifically includes: The second voice processing unit acquires ambient voice signals through at least one of the air conduction microphones and acquires the wearer's vibration signals through the bone conduction sensor. It preprocesses the ambient voice signals and vibration signals respectively, including gain adjustment, bandwidth limitation, noise reduction, or filtering. The preprocessed environmental speech signal and vibration signal are subjected to time synchronization processing, and the time synchronization processing methods include delay compensation and phase alignment. Speech-related features are extracted from the time-synchronized environmental speech signal and vibration signal. These speech-related features include energy features, spectral features, and temporal envelope. Based on the aforementioned speech-related features, a first fusion operation is performed on the time-synchronized environmental speech signal and vibration signal. The first fusion operation includes weighted fusion, selective enhancement, and / or feature-based fusion processing. The signal after the first fusion operation is reconstructed to obtain the second output speech signal.

4. The multi-mode voice processing system for clip-on headphones according to claim 1, characterized in that, The mode selection module can also perform fusion processing on the first output voice signal and the second output voice signal, and perform a second fusion operation on the first output voice signal and the second output voice signal according to weights to obtain the third output voice signal; the mode selection module selects between the first output voice signal, the second output voice signal and the third output voice signal according to environmental characteristics, and adopts the corresponding output voice signal.

5. The multi-mode voice processing system for clip-on headphones according to claim 4, characterized in that, It also includes a mode smoothing switching module, which is used to introduce a transition phase during the switching process between the first output voice signal, the second output voice signal and the third output voice signal. During the transition phase, the two output voice signals are fused. During the switching process between one output voice signal and another output voice signal, the two output voice signals are combined according to a certain weight to achieve a smooth transition from one voice signal to another.

6. The multi-mode voice processing system for clip-on headphones according to claim 1, characterized in that, The selection of the output voice signal specifically includes: the mode selection module estimates the ambient noise intensity or signal-to-noise ratio information in real time based on the audio signal collected by the air conduction microphone, and selects the corresponding output voice signal according to the ambient noise intensity or signal-to-noise ratio information.

7. A multi-mode speech processing method for clip-on headphones, characterized in that, The clip-on earphone is equipped with at least two air conduction microphones and one bone conduction sensor. The multi-mode speech processing method includes: The mode selection module selects the output voice signal, choosing between a first, second, and third output voice signal based on environmental characteristics, and adopts the corresponding output voice signal. The formation methods of the first, second, and third output voice signals are as follows: The first voice processing module performs a first signal processing procedure, and performs beamforming processing on the signal collected by the air conduction microphone to form a first output voice signal. The second speech processing module is used to perform a second signal processing process, which fuses the signal collected by the air conduction microphone and the signal collected by the bone conduction sensor to form a second output speech signal. The mode selection module is used to fuse the first and second output speech signals to form a third output speech signal.

8. The multi-mode voice processing method for clip-on headphones according to claim 7, characterized in that: The first signal processing procedure specifically includes: the first speech processing module acquiring speech signals through at least two air conduction microphones to obtain multi-channel input signals; performing time-frequency analysis processing on the multi-channel input signals to convert the time-domain signals into time-frequency domain signals for representation; extracting spatial feature information based on the differences between the multi-channel input signals, the spatial feature information including phase difference, time delay difference, and amplitude difference; enhancing the speech signal in the target direction and suppressing the signals in other directions based on the spatial feature information; and then converting the enhanced and suppressed time-frequency domain signals into the first output speech signal. The second signal processing procedure specifically includes: the second speech processing unit acquiring ambient speech signals through at least one air conduction microphone and acquiring the wearer's vibration signals through the bone conduction sensor; preprocessing the ambient speech signals and vibration signals respectively, including gain adjustment, bandwidth limitation, noise reduction, or filtering; performing time synchronization processing on the preprocessed ambient speech signals and vibration signals, including delay compensation and phase alignment; extracting speech-related features from the time-synchronized ambient speech signals and vibration signals, including energy features, spectral features, and temporal envelope; performing a first fusion operation on the time-synchronized ambient speech signals and vibration signals based on the speech-related features, including weighted fusion, selective enhancement, and / or feature-based fusion processing; and reconstructing the signal after the first fusion operation to obtain the second output speech signal. The processing of the third output speech signal specifically includes: performing a second fusion operation on the first output speech signal and the second output speech signal according to weights to obtain the third output speech signal.

9. The multi-mode voice processing method for clip-on headphones according to claim 7, characterized in that, Also includes: During the switching process between the first, second, and third output voice signals using the mode smoothing switching module, a transition phase is introduced. In this transition phase, the two output voice signals are fused. That is, during the switching from one output voice signal to another, the two output voice signals are combined according to a certain weight. The weight is a fixed value or is dynamically adjusted with time and environmental changes to achieve a smooth transition from one voice signal to another.

10. An ear clip-on earphone, comprising a front earphone shell and a rear earphone shell, wherein a C-shaped connecting bridge is provided between the front earphone shell and the rear earphone shell, and electronic components are provided in the front earphone shell and the rear earphone shell, characterized in that: The electronic device includes at least two air conduction microphones, a bone conduction sensor, a processor, and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 7-9.