This invention provides an end-to-end voice
encryption method, apparatus, and
Bluetooth headset suitable for multi-hop lossy channels. The method includes front-end
noise reduction, framing, and segmentation of the acquired voice at the transmitting end. Based on the session mode, a group public key or a point-to-point private key is automatically selected as the
master key from a pre-configured key set. Then, multiple time-domain segments within each frame are scrambled, subjected to segment-by-segment
Fast Fourier Transform, and frequency-domain
encryption based on subkeys, according to the
master key. The encrypted voice is then generated through inverse transformation and overlapping weighted
smoothing concatenation, and transmitted through a multi-hop
voice communication channel containing multi-level lossy encoding and decoding. At the receiving end, the reverse process is performed: framing and segmentation, transformation, inverse frequency-domain
encryption, and inverse time-domain scrambling. Combined with U-shaped neural network
noise reduction at both the front-end and back-end, the voice is restored to intelligible speech. This method enables flexible
key management and high-quality
secure voice transmission in scenarios where group calls and point-to-point private calls coexist and undergo multiple
lossy compression operations.