A near-field private voice communication method based on whisper physiological principle
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]要解决的技术问题:针对现有技术中所有私密语音方案均围绕「发声后隔音 / 反向消音」的后置思路设计,无基于人类悄悄话生理原理的完整端到端私密通讯方法,无法同时解决公共场合说话扰民与正常流畅私密交流的核心矛盾,同时现有方案存在佩戴笨重、语音失真、隐私性不足、场景适配性差的缺陷,本发明旨在提供一种全新的近场私密语音通讯方法,从发声源头阻断人声外泄,同时保障语音交流的清晰度与流畅性,填补现有技术的全球空白
第一,从发声根源解决人声外泄问题,基于人类悄悄话的生理发声原理,通过近场气导采集,仅捕捉口腔呼出气流携带的悄悄话语音信号,超出近场范围后信号快速衰减至环境本底噪声水平,天然阻断人声向外部环境的辐射传播,无需复杂的隔音或主动消音结构,彻底解决了公共场合说话扰民的核心痛点;
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of voice signal processing, near-field private communication, and wearable device technology. Background Technology
[0002] With the increasing demand for remote work, online communication, and private interaction in public settings, users are becoming more and more eager for voice communication solutions that do not disturb others in public places while ensuring clarity and privacy.
[0003] All existing private voice solutions on the market are designed around the post-processing concept of "sound isolation / reverse noise cancellation after speech," and can be broadly divided into two categories: one is physical sound isolation solutions, such as noise-canceling masks and face shields, which use thick sound-absorbing materials to wrap around the mouth and block the transmission of voice after it is emitted. These solutions have drawbacks such as large size, being bulky and stuffy to wear, causing severe distortion of speech, and being unable to achieve smooth and natural communication, making them unsuitable for long-term daily wear; the other is sound pickup optimization solutions, such as bone conduction microphones and AI noise-canceling microphones, which can only optimize the sound pickup effect in noisy environments and reduce the interference of environmental noise on the sound pickup, but cannot completely block the radiation of human voice to the external environment. When the wearer speaks, people around can still hear clearly, and they cannot solve the core pain point of disturbing others when speaking in public places.
[0004] To date, there is no complete end-to-end private voice communication method worldwide based on "the physiological principle of human whispering + near-field air conduction acquisition + point-to-point private transmission". Existing technologies have always been unable to simultaneously resolve the core contradiction of "disturbing others by talking in public places" and "normal, smooth, and private communication", leaving a significant technological gap and failing to meet users' actual needs. Summary of the Invention
[0005] The technical problem to be solved: All existing private voice solutions are designed around the post-processing approach of "sound isolation / reverse noise cancellation after speech," lacking a complete end-to-end private communication method based on the physiological principles of human whispering. This fails to simultaneously resolve the core contradiction between disturbing others in public and maintaining normal, smooth, and private communication. Furthermore, existing solutions suffer from drawbacks such as bulky design, voice distortion, insufficient privacy, and poor scene adaptability. This invention aims to provide a novel near-field private voice communication method that blocks voice leakage at the source of speech while ensuring clarity and fluency in voice communication, filling a global gap in existing technology.
[0006] To achieve the above objectives, the technical solution adopted by this invention is a near-field private voice communication method based on the physiological principle of whispering, comprising the following complete steps executed in chronological order: Step 1: By using a pickup module fixed near the front of the wearer's lips, only the oral air conduction speech signal generated when the wearer speaks in a whispering manner is collected, naturally blocking the radiation and propagation of human voice to the external environment from the source of sound. Step 2: Preprocess the collected air conduction speech signal to filter out environmental noise and retain the specific effective frequency band speech signal corresponding to the physiological characteristics of whispering. Step 3: Using the trained human voice restoration model, the loudness and timbre of the preprocessed weak air-conduction speech signal are restored to generate a highly intelligible human voice signal; Step 4: Establish a point-to-point encrypted transmission channel without third-party intermediary devices through a low-power wireless transmission module, and transmit the repaired human voice signal directly to the pre-paired external audio receiving device, so that only the wearer of the paired device can listen clearly, and there is no perceptible human voice leakage in the surrounding environment.
[0007] Furthermore, in step 1, the pickup end of the pickup module is preferably fixed within a range of 0.3cm-1.5cm from the front of the wearer's lips, and optimally within a range of 0.5cm-1cm, to ensure that only the air-conducted speech signal carried by the exhaled airflow from the oral cavity is collected.
[0008] Furthermore, in step 2, the preprocessing prioritizes the retention of the 1kHz-8kHz effective frequency band signal for whispering through bandpass filtering, while simultaneously filtering out stable environmental noise and non-stationary clutter outside the effective frequency band through an adaptive filtering algorithm.
[0009] Furthermore, in the preprocessing process of step 2, dynamic gain control is performed on the selected effective speech signals to stabilize the signal amplitude within the standard range without clipping distortion.
[0010] Furthermore, the voice reconstruction model in step 3 is an edge-side inference model optimized based on a lightweight neural network architecture. The model supports completely offline local operation, and the single-frame inference latency meets the lip-sync requirements of real-time voice calls.
[0011] Furthermore, in step 4, the low-power wireless transmission module preferentially adopts a low-power Bluetooth module and uses a symmetric encryption algorithm to establish a full-link encrypted transmission channel. Only pre-paired devices can decrypt and play the encrypted voice signal.
[0012] Furthermore, in step 4, the low-power wireless transmission module can simultaneously establish a connection with the smart terminal, and connect the repaired voice signal to the smart terminal's call and audio / video conferencing link, adapting to private call scenarios in public places.
[0013] Furthermore, in step 4, a one-to-many encrypted networking channel can be established through a low-power wireless transmission module to synchronously transmit the repaired human voice signal to multiple pre-paired external audio receiving devices, thereby realizing a scenario of private voice communication among multiple people.
[0014] Furthermore, the pickup module in step 1 is compatible with any wearable form, including ear-hook, glasses, headband, and neckband, as long as the pickup end is within the near-field range where oral air conduction whisper signals can be effectively collected.
[0015] Furthermore, in step 1, the front end of the pickup module is equipped with an oil-proof and smoke-proof protective structure, and in the preprocessing process of step 2, a special filtering process for airflow impact noise is added to adapt to stable pickup under smoking and outdoor windy scenarios.
[0016] Furthermore, all signal acquisition, preprocessing, voice restoration, and encrypted transmission processes in steps 1 to 4 are completed locally on the device, without relying on a cloud server.
[0017] The core beneficial effects of this invention address the pain points of existing technologies, and can be summarized in the following four points: First, it addresses the issue of voice leakage at its source. Based on the physiological principle of human whispering, it captures only the whispered voice signal carried by the airflow exhaled from the mouth through near-field air conduction. Once the signal exceeds the near-field range, it quickly attenuates to the level of ambient background noise, naturally blocking the radiation and propagation of human voice to the external environment. It eliminates the need for complex sound insulation or active noise cancellation structures, thus completely solving the core pain point of disturbing others by talking in public places. Secondly, it achieves highly intelligible and fluent voice communication. By matching the physiological characteristics of whispers with a dedicated frequency band selection and human voice restoration model, it specifically repairs the loudness, timbre and formant characteristics of weak whisper signals, solving the problems of unclear voice, low intelligibility and long-term listening fatigue in traditional whispers. It ensures that the listening effect at the receiving end is no different from normal speech, and achieves natural and fluent private communication. Third, end-to-end privacy protection across the entire chain. Through point-to-point encrypted transmission without the need for third-party intermediary devices, it avoids leakage of voice signals during the relay process without going through intermediate devices such as mobile phones, tablets or cloud servers. At the same time, only pre-paired devices can decrypt and listen, realizing private communication throughout the entire process from collection to playback, with no risk of information leakage. Its privacy and security are far superior to existing solutions. Fourth, it has strong scenario adaptability, is lightweight and easy to implement. This method does not require a complex hardware structure, can be adapted to various wearable device forms, and the entire process can be completed on a local low-power chip without cloud dependence. It can be widely used in various scenarios such as public offices, public transportation, libraries, multi-person private meetings, and nighttime home use, solving the shortcomings of existing solutions such as poor adaptability, high threshold for use, and inability to be used for long periods of time in daily life. Attached image description: Figure 1 is a flowchart of the overall method of the present invention, showing the four core steps of the method executed in chronological order: step S1 air-guided whisper voice signal acquisition, step S2 voice signal preprocessing, step S3 human voice restoration and repair, and step S4 point-to-point encrypted transmission and playback. Figure 2 is a flowchart of the speech signal processing of the present invention, showing the complete end-to-end link of the method from the acquisition of the original signal to the final playback, which consists of: a sound pickup module, a preprocessing unit, a human voice restoration unit, an encrypted transmission unit, and an external audio receiving device. Detailed implementation method: The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. These embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can completely replicate and implement the method of the present invention based on these embodiments without creative effort. The complete process of the near-field private voice communication method based on the physiological principle of whispering described in this embodiment is shown in Figure 1. The specific execution steps and achievable details are as follows: (1) Step S1: Acquisition of whispered voice signals via air conduction This step corresponds to the sound pickup module in Figure 2. During execution, the pickup end of the sound pickup module is positioned in the near field at the front of the wearer's lips through the fixed structure of the wearable device, facing the wearer's mouth. This step is based on the physiological principle of human whispering: when the wearer whispers, the vocal cords do not vibrate, and the voice signal is generated only by the friction of the airflow exhaled from the mouth. The signal energy is highly concentrated in the near field range at the front of the lips. Beyond this range, the signal strength rapidly decays to the ambient noise level, naturally blocking the radiation and propagation of human voice to the external environment from the source of the sound. There is no perceptible leakage of human voice in the surrounding environment. In this embodiment, the microphone module uses a MEMS microphone with a signal-to-noise ratio ≥60dB and sensitivity of -38dBFS. The sampling rate is set to 16kHz and the bit depth to ensure complete acquisition of whispered air conduction speech signals. The microphone is preferably fixed within a range of 0.3cm-1.5cm from the front of the lips, and optimally within a range of 0.5cm-1cm, balancing acquisition effect and wearing comfort. The microphone module can be adapted to any wearable form, including ear-hook, glasses, headband, and neckband, as long as the microphone is within the near-field range where the whispered air conduction signal can be effectively acquired. The front end of the microphone module can be equipped with a replaceable oil-proof and smoke-proof protective structure made of polypropylene to meet the protection requirements of special scenarios. (2) Step S2: Speech signal preprocessing This step corresponds to the preprocessing unit in Figure 2, and processes the raw air-conduction speech signal acquired in step S1 sequentially: First, anti-aliasing filtering removes noise interference from the original signal. Second, bandpass filtering preserves the specific effective frequency band of speech signal corresponding to the physiological characteristics of whispering, prioritizing an 8th-order Butterworth bandpass filter to retain the 1kHz-8kHz band (where over 90% of the speech energy of human whispering is concentrated). Then, a least mean square (LMS) adaptive filtering algorithm is used to filter out stationary environmental noise (such as traffic noise and air conditioning noise) and non-stationary noise outside the effective frequency band. Finally, dynamic gain control is applied to the filtered signal to stabilize the signal amplitude within the clipping-free distortion standard range of -18dBFS to -3dBFS, preventing signal distortion in subsequent processing. In this embodiment, all preprocessing processes are completed in the local hardware DSP unit of the low-power main control chip, with a processing delay of ≤5ms, ensuring real-time performance; special filtering processing can be added according to the usage scenario, such as adding 200Hz-800Hz band-stop filtering for smoking or outdoor windy scenarios to specifically filter out airflow impact noise and ensure stable sound pickup effect. (3) Step S3: Voice restoration and repair This step corresponds to the voice restoration unit in Figure 2, which repairs the effective speech signal after preprocessing in step S2: The trained voice reconstruction model is used to restore the loudness and timbre of the preprocessed weak air-conduction speech signal, generating a highly intelligible voice signal. In this embodiment, the voice reconstruction model is a lightweight convolutional neural network model optimized based on the MobileNet architecture. The model size is ≤5MB, which can be directly deployed on the local side of the low-power main control chip. It supports completely offline operation without cloud dependency. The single-frame inference latency is ≤10ms, and the end-to-end processing latency is ≤15ms, which fully meets the lip-sync requirements of real-time voice calls. The model was pre-trained under supervision using a dataset of whispers and normal human voices paired with different genders, ages, and accents. The training objective was to map the weak whisper air conduction signal into a clear and full normal human voice signal, and to specifically repair the loudness, timbre, and formant features of the signal. The intelligibility of the repaired speech signal was ≥95%, and there was no difference in listening effect from normal speech. (4) Step S4 Point-to-point encrypted transmission and playback This step corresponds to the encrypted transmission unit in Figure 2, which encrypts and transmits the high-intelligibility human voice signal repaired in step S3: Employing a low-power wireless transmission module, prioritizing Bluetooth 5.2 and above, and supporting the BLE Audio protocol, the device automatically enters pairing mode upon power-on, completing point-to-point pairing with external Bluetooth headphones and audio receiving devices. After pairing, an end-to-end encrypted transmission channel is established using the AES-128 symmetric encryption algorithm. Without the need for third-party intermediary devices such as mobile phones or tablets, the repaired voice signal is directly transmitted through the encrypted channel to the pre-paired external audio receiving device. Only the paired device can decrypt and play the voice signal using the corresponding key; other surrounding devices cannot listen or decrypt, ensuring that only the wearer of the paired device can hear clearly, with no perceptible voice leakage in the surrounding environment. In this embodiment, the encrypted transmission latency is ≤20ms, and the transmission distance covers a range of 10 meters, fully meeting the needs of daily private communication; it can also be expanded to adapt to the following scenarios: 1. Smart terminal call adaptation: The low-power Bluetooth module synchronously establishes an HFP call protocol connection with smart terminals such as mobile phones, and connects the repaired human voice signal to the call and audio-visual conferencing link of the smart terminal, adapting to private calls in public places and online meeting scenarios; 2. Multi-person private networking scenario: A one-to-many encrypted networking channel is established through the low-power Bluetooth module, which can support up to 8 pre-paired audio receiving devices to connect at the same time. The repaired human voice signal is synchronously encrypted and transmitted to all paired devices, realizing multi-person private meetings and team remote communication scenarios. 3. Fully Local Offline Operation: All processes from step S1 to step S4 are completed locally on the device, without relying on cloud servers, reducing device power consumption and further enhancing privacy and security.
Claims
1. A near-field private voice communication method based on the physiological principle of whispering, characterized in that, Includes the following complete steps performed in chronological order: Step 1: By using a pickup module fixed near the front of the wearer's lips, only the oral air conduction speech signal generated when the wearer speaks in a whispering manner is collected, naturally blocking the radiation and propagation of human voice to the external environment from the source of sound. Step 2: Preprocess the collected air conduction speech signal to filter out environmental noise and retain the specific effective frequency band speech signal corresponding to the physiological characteristics of whispering. Step 3: Using the trained human voice restoration model, the loudness and timbre of the preprocessed weak air-conduction speech signal are restored to generate a highly intelligible human voice signal; Step 4: Establish a point-to-point encrypted transmission channel without third-party intermediary devices through a low-power wireless transmission module, and transmit the repaired human voice signal directly to the pre-paired external audio receiving device, so that only the wearer of the paired device can listen clearly, and there is no perceptible human voice leakage in the surrounding environment.
2. The method according to claim 1, characterized in that, In step 1, the pickup end of the pickup module is fixed within a range of 0.3cm-1.5cm from the front of the wearer's lips, preferably within a range of 0.5cm-1cm, to ensure that only the air-conducted speech signal carried by the exhaled airflow from the mouth is collected.
3. The method according to claim 1, characterized in that, The preprocessing in step 2 retains the effective frequency band signal for whispering (1kHz-8kHz) through bandpass filtering, while filtering out stable environmental noise and non-stationary clutter outside the effective frequency band through an adaptive filtering algorithm.
4. The method according to claim 1, characterized in that, In the preprocessing process of step 2, dynamic gain control is performed on the selected effective speech signals to stabilize the signal amplitude within the standard range without clipping distortion.
5. The method according to claim 1, characterized in that, The voice reconstruction model in step 3 is an edge-side inference model optimized based on a lightweight neural network architecture. The model supports completely offline local operation, and the single-frame inference latency meets the lip-sync requirements of real-time voice calls.
6. The method according to claim 1, characterized in that, The low-power wireless transmission module in step 4 is a low-power Bluetooth module, which uses a symmetric encryption algorithm to establish a full-link encrypted transmission channel. Only pre-paired devices can decrypt and play the encrypted voice signal.
7. The method according to claim 1, characterized in that, In step 4, the low-power wireless transmission module can simultaneously establish a connection with the smart terminal and connect the repaired voice signal to the smart terminal's call and audio / video conferencing link, adapting to private call scenarios in public places.
8. The method according to claim 1, characterized in that, In step 4, a one-to-many encrypted network channel can be established through a low-power wireless transmission module to synchronously transmit the repaired human voice signal to multiple pre-paired external audio receiving devices, thereby realizing a scenario of private voice communication among multiple people.
9. The method according to claim 1, characterized in that, The pickup module in step 1 is compatible with any wearable form, including ear-hook, glasses, headband, and neckband, as long as the pickup end is within the near-field range where oral air conduction whisper signals can be effectively collected.
10. The method according to claim 1, characterized in that, In step 1, the front end of the pickup module is equipped with an oil-proof and smoke-proof protective structure. In the preprocessing process of step 2, a special filtering process for airflow impact noise is added to adapt to stable pickup under smoking and outdoor windy scenarios.
11. The method according to claim 1, characterized in that, All signal acquisition, preprocessing, voice restoration, and encrypted transmission processes from step 1 to step 4 are completed locally on the device, without relying on a cloud server.