End-to-end voice encryption methods, devices, and Bluetooth headsets applicable to multi-hop lossy channels
By employing an alternating encryption structure in the time domain and transform domain, along with front-end and back-end noise reduction processing, the reliability problem of end-to-end voice encryption under multi-hop lossy channels is solved, achieving a balance between voice quality and security, reducing the implementation complexity on the terminal side, and supporting unified management of group and point-to-point keys.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VISION INTELLIGENCE CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-06-30
AI Technical Summary
Existing technologies struggle to achieve reliable decryption and recovery of end-to-end voice encryption under multi-hop lossy channels, and lack encryption schemes designed jointly in the time and frequency domains. They cannot balance voice quality and security, have insufficient coordination between group and point-to-point key management, and are highly complex to implement on the terminal side.
It adopts an encryption structure that alternates between the time domain and the transform domain. Through frame division and segment subdivision processing, combined with time domain scrambling and frequency domain encryption controlled by the master key, it introduces front-end and back-end noise reduction processing, and performs missing detection and loss compensation at the receiving end. It supports unified management of group public keys and point-to-point private keys.
Achieving stable synchronization and reliable decryption recovery of encrypted voice under multi-hop lossy channels improves voice quality and intelligibility, reduces terminal-side implementation complexity, enhances key management flexibility and security, and adapts to complex voice service scenarios.
Smart Images

Figure CN121923865B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice signal processing and encrypted communication technology, specifically to an end-to-end voice encryption method, apparatus, and Bluetooth headset suitable for multi-hop lossy channels. Background Technology
[0002] This field relates to voice signal processing and secure communication technologies, particularly real-time voice call scenarios where mobile terminals, Bluetooth headsets, and internet conferencing software work together. With the widespread adoption of mobile internet and cloud conferencing systems, users increasingly connect their phones via Bluetooth headsets, and then the phones conduct multi-party voice interactions through conferencing software, instant messaging software, or voice call platforms. To ensure smooth calls under complex network conditions, lossy voice encoders such as mSBC, Opus, SILK, and Speex are widely used for compression on both the terminal and network sides, often involving multi-hop relays and cross-system transcoding. Simultaneously, with increasingly stringent privacy and data security requirements, there are risks of eavesdropping and hijacking in areas such as operator networks, open Wi-Fi, and cloud servers. Therefore, the industry is gradually evolving from link encryption to end-to-end voice encryption, aiming to ensure that voice content can only be decoded and understood by the two parties involved in the call or authorized groups, while maintaining low latency and voice quality.
[0003] In existing technologies, common voice encryption schemes can be broadly categorized into two types. One type involves encrypting the bitstream or data packets at the channel or transport layer after voice encoding. Examples include SRTP over RTP and TLS at the connection layer. This type of scheme encrypts the network data carrying the voice, not the voice signal itself. The advantage of this type is that it is almost transparent to the voice encoding / decoding and media processing links, resulting in low modification costs. However, once transcoding, mixing, or recording occurs at the server side or a relay node, the voice content needs to be decrypted into plaintext at that node. This still presents a risk of eavesdropping around that node, making true end-to-end confidentiality difficult to achieve. The other type of scheme directly encrypts or scrambles the voice waveform. For example, it involves scrambling, modulating, or encrypting the sampled data in the time or frequency domain. This type of scheme provides security at the signal level, but it often assumes the transmission channel is lossless or only has bit errors, and it doesn't adequately consider lossy compression and multiple transcoding. Once the encrypted voice passes through multiple levels of lossy codecs, it is easily unreliable to recover at the receiving end.
[0004] Furthermore, many existing solutions are designed for single voice links and single codec scenarios. For example, they only consider direct transmission between terminals using a fixed encoding format, ignoring the complex link in real-world applications: "headset—phone—conference software—phone—headset." The encrypted voice signal must be adapted to Bluetooth mSBC codec, network-side Opus / SILK / Speex codec, and then mSBC codec again, undergoing multiple lossy compression and decompression processes. Traditional voice encryption algorithms, without considering the bandwidth characteristics, quantization noise, and spectral truncation of lossy codecs, are prone to decryption failure or severe degradation of intelligibility after multiple channel hops, making them difficult to implement in real-world products.
[0005] To improve the compatibility of voice encryption with existing voice links, some "vocoder penetration" or "scrambling transmission" ideas have been proposed in existing technologies. For example, CN105071896A, entitled "Vocoder Penetration Encryption and Decryption Method and System and Voice Encryption and Decryption Method," discloses a vocoder penetration encryption process, which includes performing an N-point time-frequency domain forward transform on digitized plaintext speech to obtain a frequency domain digital speech signal, performing forward renormalization on the frequency domain signal according to the transform domain renormalization sequence, performing an N-point time-frequency domain inverse transform to generate ciphertext speech data, and inserting synchronization information into the ciphertext speech data to form the output ciphertext speech. However, in multi-hop lossy scenarios where mobile terminals and cloud conferencing collaborate, voice may undergo multiple lossy encoders in series and cross-system transcoding. Relying solely on the approach of overall transformation and reorganization and inserting synchronization information is still difficult to systematically address the statistical characteristic drift and error accumulation caused by multiple lossy encoding and decoding. Furthermore, it does not address the unified selection and collaborative management of the terminal-side key system when group calls and point-to-point private calls coexist. Therefore, there is still room for improvement in adapting to the aforementioned complex business links.
[0006] In terms of key management, existing technologies mostly adopt a single session key or a simple group key management mode. For complex voice scenarios involving group broadcasts and point-to-point private calls, some systems directly assign a shared key to the entire group chat, with all participants using the same key for encryption and decryption. While this is simple to implement, if a terminal leaks its key, the entire group's historical and future call content may be decrypted. Simultaneously, point-to-point private voice calls in a group chat can only reuse the group key, failing to achieve fine-grained security control. Some systems have attempted to assign point-to-point keys to each pair of terminals, aggregating and distributing them through a server within the group chat. However, in actual operation, frequent key switching and session reconstruction increase signaling overhead and implementation complexity, especially given the limited power consumption and computing power of mobile terminals. It is difficult to implement a unified logical judgment mechanism on the device side that allows public and private keys to work together.
[0007] Regarding voice quality, existing voice communication products typically rely on front-end noise reduction, echo cancellation, and various gain control algorithms on the network side to mitigate environmental noise and link jitter. Most of these algorithms are designed and optimized on unencrypted voice signals, assuming the voice waveform possesses typical speech-like statistical characteristics. Once complex time-frequency domain encryption or perturbation is applied to the voice at the front end, the original noise reduction model and back-end enhancement algorithms may fail to correctly identify the voice structure, leading to deterioration in noise reduction or even the introduction of new distortions. Existing technologies also include solutions that add a back-end noise reduction module at the receiver, but these are mostly used as general voice enhancement processing after decoding, lacking coordinated design with end-to-end encryption algorithms, making it difficult to balance noise suppression and encryption robustness.
[0008] To address these issues, several improvement approaches have been proposed in existing technologies. For example, some literature suggests embedding encryption logic within a specific speech encoder to scramble the encoding parameters, thereby providing a certain level of security while ensuring speech compatibility. However, this "codec-decoder coupled encryption" is often only applicable to specific encoding formats. Once cross-coding transcoding occurs in a multi-hop channel, such as from mSBC to Opus and back to mSBC, the original parameter-level encryption information will be destroyed, and the receiver cannot recover the original speech. Other schemes segment and scramble the speech in the time domain, then apply a mask or modulation to the spectral coefficients in the frequency domain, aiming to balance security and codec compatibility. However, these designs are mostly based on single-pass lossy compression scenarios and lack systematic analysis of the statistical characteristics and error accumulation in the case of multi-level lossy codecs cascaded, still making it difficult to guarantee the reliability of decryption after crossing multiple hop channels.
[0009] In terms of secure voice communication in groups, some existing systems maintain group keys and peer-to-peer keys at the application layer, coordinate the distribution of different keys through a server, and use signaling during the call to identify whether the current voice is addressed to the entire group or a single individual. While this approach can differentiate between group broadcasts and private peer-to-peer calls to some extent, the key selection logic, update strategies, and mapping relationships with specific encryption algorithms are quite fragmented on the terminal side. Developers need to perform extensive adaptation work between the business and security layers, which is not conducive to implementing a unified voice processing chain on lightweight terminals such as low-power headsets. Furthermore, existing solutions often rely on upper-layer protocols for key rotation, replay attack protection, and key consistency handling when multiple terminals simultaneously join or leave the group, failing to provide complete logical support at the voice encryption algorithm design level.
[0010] In summary, existing technologies still have shortcomings in end-to-end voice encryption and multi-hop lossy channel adaptation, collaborative group and point-to-point key management, and overall coordination with front-end / back-end noise reduction processing. On the one hand, there is a lack of an encryption scheme jointly designed in the time and frequency domains that can ensure encrypted voice retains decryptability and basic intelligibility after multiple mSBC and network speech encoding processes, while also preserving as much of the speech-like statistical characteristics as possible to facilitate compatibility with lossy codecs along the route. On the other hand, there is also a lack of a unified logical mechanism on the terminal side for automatic selection of group public keys and point-to-point private keys, tightly integrated with the end-to-end voice encryption process, to adapt to complex real-time voice service scenarios. Summary of the Invention
[0011] To address the aforementioned technical issues, this invention provides an end-to-end voice encryption method suitable for multi-hop lossy channels. Its purpose is to achieve stable synchronization and reliable decryption recovery of encrypted voice in application scenarios where group calls and point-to-point private calls coexist and voice needs to undergo multiple lossy encoding / decoding and network transmissions, ensuring that the receiving end can listen normally. Furthermore, it enhances the flexibility of key management and reduces the complexity of implementation on the terminal side.
[0012] An end-to-end voice encryption method suitable for multi-hop lossy channels includes:
[0013] The acquired voice signal to be encrypted is subjected to front-end noise reduction processing to obtain the front-end noise-reduced voice signal;
[0014] The front-end denoised speech signal is divided into frames, and each speech frame is divided into N consecutive time-domain segments in the time domain, where N>1.
[0015] Based on the current call session mode information, the master key for the voice frame is selected from the pre-configured key set. The key set includes a public key for group calls and a private key for point-to-point calls. When the current session is detected as a group call, the public key is selected as the master key. When the current session is detected as a point-to-point call, the private key shared only between the terminals of the two parties in the call is selected as the master key.
[0016] Based on the temporal permutation rules determined by the master key, the temporal arrangement of the N temporal segments is scrambled to obtain a temporal segment sequence after macroscopic temporal scrambling;
[0017] For each scrambled time-domain segment, the following operation is performed: Based on the master key and combined with the frame identifier and / or segment identifier corresponding to the time-domain segment, a frequency-domain encryption subkey corresponding to the time-domain segment is generated through key derivation, so that the frequency-domain encryption subkeys corresponding to different time-domain segments within the same voice frame are different from each other, and the receiving end can synchronously generate the corresponding frequency-domain encryption subkey based on the same master key and the frame identifier and / or segment identifier;
[0018] The time domain segment is transformed to obtain the transform domain coefficients of the time domain segment. The transform domain coefficients are then encrypted in the frequency domain using the frequency domain encryption subkey corresponding to the time domain segment. Finally, the encrypted transform domain coefficients are inversely transformed to obtain the encrypted time domain data of the time domain segment. The transformation and the inverse transformation are performed separately at the time domain segment granularity.
[0019] The encrypted time-domain segments are concatenated according to the scrambled time sequence, and smoothing is performed at the concatenation boundaries of adjacent encrypted time-domain segments to obtain an encrypted voice signal.
[0020] Furthermore, the encryption method employs an alternating encryption structure of the time domain and the transform domain, including: performing macroscopic structural processing on each speech frame in the time domain, that is, dividing the speech frame into N consecutive time domain segments, and macroscopically scrambling the temporal arrangement of the N time domain segments according to the time domain permutation rule determined by the master key; and processing the microstructure of each time domain segment in the transform domain, transforming each time domain segment after time domain scrambling to obtain transform domain coefficients, performing frequency domain encryption processing on the transform domain coefficients using the frequency domain encryption subkey corresponding to the time domain segment in the transform domain, and performing inverse transformation on the encrypted transform domain coefficients to obtain the corresponding encrypted time domain segment, wherein the frequency domain encryption processing is performed segment by segment sequentially on the time domain segments.
[0021] Furthermore, the frame identifier includes at least one of frame index, frame sequence number, and timestamp; the segment identifier includes at least one of segment index and segment sequence number; the key derivation method is a key derivation function, the input of which includes at least the master key and the frame identifier and / or segment identifier, and may further include at least one of session identifier, session sequence number, and random number, to generate a frequency domain encryption subkey corresponding one-to-one with each time domain sub-segment, thereby enabling the sending end and receiving end to achieve subkey synchronization without additionally transmitting the frequency domain encryption subkey.
[0022] Furthermore, the transformation performed on each of the time-domain sub-segments is a Fast Fourier Transform (FFT), and the number of transformation points of the FFT is less than the number of sampling points of the corresponding speech frame; the inverse transformation performed on the encrypted transform domain coefficients is an Inverse Fast Fourier Transform (IFT) corresponding to the FFT; thus, when encrypting the entire frame of speech, by performing the FFT and IFT on each time-domain sub-segment respectively, the transformation and inverse transformation of the sub-segments can be replaced by a single overall transformation and inverse transformation of the entire frame of speech.
[0023] Furthermore, the smoothing process performed at the splicing boundary of adjacent encrypted time-domain segments includes: setting an overlap region of a predetermined length at the splicing point of adjacent encrypted time-domain segments; applying a set of complementary weighting coefficients to the tail sample of the preceding encrypted time-domain segment and the head sample of the following encrypted time-domain segment respectively, and superimposing them to achieve cross-fading and / or overlapping addition processing of adjacent encrypted time-domain segments, so as to reduce the amplitude and phase abrupt changes at the splicing boundary and ensure the listening quality of the output encrypted speech, wherein the length of the overlap region is less than the length of the time-domain segment.
[0024] Furthermore, among the transform domain coefficients obtained by transforming each of the aforementioned time domain segments, frequency domain encryption processing is performed on the transform domain coefficients only within the first frequency interval corresponding to the intersection of the effective speech bandwidths of each lossy speech codec in the multi-hop speech communication channel, according to the frequency domain encryption subkey. The frequency domain encryption processing is not performed in the second frequency interval outside the first frequency interval. This ensures that the encrypted speech signal maintains speech-like statistical characteristics similar to the original speech signal in terms of overall power spectrum distribution and short-time energy trajectory, so that the encrypted speech signal can be transmitted in a multi-hop speech communication channel including one or more lossy speech codecs and transmission units, and can be recovered into intelligible speech at the receiving end through a decryption process that is the reverse of the encryption process.
[0025] Furthermore, the front-end noise reduction processing includes: performing a short-time Fourier transform on the acquired speech signal to be encrypted to obtain complex time-frequency features; inputting the complex time-frequency features into a U-shaped neural network containing an encoding part, an enhancement part, and a decoding part and having multiple levels of skip connections; the output of the U-shaped neural network having a complex value mask with an amplitude limited to the range of 0 to 1; using the complex value mask to weight the real and imaginary parts of the complex time-frequency features respectively; and performing an inverse short-time Fourier transform on the weighted time-frequency features to obtain the front-end noise-reduced speech signal as the encryption input.
[0026] Furthermore, after the receiving end completes the decryption of the received encrypted voice signal, it also includes performing back-end noise reduction processing on the decrypted voice signal. The back-end noise reduction processing includes: performing a short-time Fourier transform on the decrypted voice signal to obtain complex spectral features; inputting the complex spectral features into a U-shaped neural network with the same or reduced parameters as the U-shaped neural network used in the front-end noise reduction processing to obtain an initial complex value mask; optimizing the initial complex value mask using causal time-weighted depth filtering that only utilizes the time-frequency information of the current frame and its preceding frames to obtain an optimized mask; calculating the frequency point gain according to the estimated prior signal-to-noise ratio and posterior signal-to-noise ratio using the minimum mean square error criterion to correct the optimized mask to obtain a final mask; multiplying the final mask by the decrypted noisy spectrum and performing an inverse short-time Fourier transform to output the back-end noise-reduced voice signal.
[0027] Furthermore, the pre-configured key set includes at least one group public key and multiple peer-to-peer private keys. The group public key is shared among multiple terminals within the same group, and each peer-to-peer private key is shared in pairs between different terminal pairs. Each peer-to-peer private key is stored only in the corresponding terminal pair.
[0028] Furthermore, the current call session mode information includes a group identifier and / or a session participant identifier. The sending end determines whether the current voice segment belongs to a group broadcast voice or a point-to-point private voice based on signaling messages and / or session control information from upper-layer applications, and selects the group public key or the corresponding terminal pair's point-to-point private key as the master key for encryption. This ensures that in a group call scenario, only the terminal holding the group public key can decrypt the group broadcast voice, and in a point-to-point private call scenario within the same group, only the two terminals holding the point-to-point private key can decrypt the corresponding private voice. Other terminals that do not hold the group public key or the point-to-point private key can only receive incomprehensible voice.
[0029] Furthermore, the master key is negotiated between relevant terminals through a key negotiation protocol at the time of each session establishment, and is updated during the session according to a predetermined number of frames or time intervals, combined with random numbers and / or session sequence numbers. The updated master key is used for time-domain scrambling and frequency-domain encryption subkey derivation of subsequent voice frames to improve the ability to resist replay attacks and brute-force attacks.
[0030] Furthermore, during the process of framing and segmenting, transforming, inverse frequency domain encryption, and inverse time domain scrambling of the encrypted voice signal transmitted through the multi-hop voice communication channel at the receiving end, missing detection and loss compensation processing is also included. The missing detection and loss compensation processing includes:
[0031] Based on the frame sequence number and / or timestamp and / or segment sequence number carried in the received encrypted voice data, missing voice frames and / or time-domain segments are detected to determine missing voice frames and / or missing time-domain segments. When missing voice frames and / or time-domain segments are determined, in order to maintain the frame index and / or segment index consistent with the sending end for subkey synchronization derivation, placeholder decryption time-domain data corresponding to the missing voice frames and / or time-domain segments is generated, and the placeholder decryption time-domain data is used in subsequent inverse time-domain scrambling and splicing output. The placeholder decryption time-domain data is obtained through at least one of the following strategies: performing time-domain interpolation on the recovered voice data before and after the missing position; masking the missing position and replacing it with a copy of the adjacent voice frame or adjacent time-domain segment and applying energy attenuation and / or window function smoothing; or predictively replacing the missing position based on the transform domain coefficients of the adjacent voice frames before performing the inverse transform.
[0032] Furthermore, before transmitting the encrypted voice signal to the multi-hop voice communication channel, a bandwidth control process is included, which includes at least one of the following: low-pass and / or band-pass filtering of the encrypted voice signal to limit the target frequency band; downsampling of the filtered signal to reduce the sampling rate; and / or adaptively adjusting the transmitter's bit rate and / or frame rate based on network bandwidth, packet loss rate, and / or latency jitter, wherein the adjustment includes adjusting the target bit rate of the voice codec and / or adjusting at least one of the number of time-domain segments N per frame, frame length, and overlap length, to reduce the amount of transmitted data and maintain decryptability and recovery at the receiver.
[0033] Furthermore, it includes a processor and a memory, the memory storing program instructions executable on the processor, which, when executed by the processor, cause the processor to perform the end-to-end voice encryption method for the voice signal according to any one of claims 1 to 13.
[0034] Furthermore, a Bluetooth headset includes a microphone, a speaker, a Bluetooth communication module, a processor, and a memory. The memory stores program instructions that can run on the processor. When the program instructions are executed by the processor, the processor is configured to: perform front-end noise reduction processing on the voice signal collected by the microphone, determine the call session mode, select and encrypt the master key, generate encrypted voice signals for each voice frame, and send them to an external terminal device through the Bluetooth communication module; and / or receive the encrypted voice signal sent by the external terminal device from the Bluetooth communication module and perform a decryption process opposite to the end-to-end voice encryption method to obtain a decrypted voice signal, which is then played through the speaker.
[0035] Furthermore, in the Bluetooth headset, the voice processing module for implementing the front-end noise reduction and end-to-end voice encryption methods is integrated with the Bluetooth communication module in the same audio processing chip.
[0036] The end-to-end voice encryption method for multi-hop lossy channels provided by this invention has the following advantages:
[0037] This invention, through the joint design of an encryption structure in the time and transform domains, ensures end-to-end security while also considering decryptability and intelligibility under multi-hop lossy encoding and decoding conditions. Specifically, this invention employs a frame-by-frame and segment-by-segment processing method, dividing the entire frame of speech into multiple consecutive time-domain segments, and macroscopically scrambling the order of these time-domain segments under the control of the master key. On the one hand, the segmented subdivision combined with macroscopic scrambling disrupts the continuity of speech in the time domain, thereby improving the encryption level and making the semantic content of the uncracked speech less likely to be discernible. On the other hand, the processing at the granular level of the time domain segments results in fewer FFT points during subsequent frequency domain transformations of the segments compared to the number of transformation points in the entire frame. This reduces the computational load of the transformation and inverse transformation, making it more suitable for real-time computation in resource-constrained terminals such as Bluetooth headsets. Furthermore, the encryption of each segment using a transformation domain with fewer points than the number of points in the entire frame, followed by inverse transformation reconstruction, ensures that the encrypted speech maintains time-frequency statistical characteristics and energy distribution features as close as possible to the original speech. This makes it easier for lossy speech encoders such as mSBC, Opus, and Speex to process the speech as normal speech signals. This helps reduce the risk of abnormal distortion and decoding failure in multi-hop combination scenarios such as Bluetooth links and web conferencing links, improving the stability and intelligibility of end-to-end speech recovery. A key derivation mechanism driven by a master key, frame index, and segment index allows different frequency domain encryption subkeys to correspond to different time-domain segments within the same frame. Simultaneously, the sending and receiving ends can maintain key synchronization without transmitting additional subkeys. This design increases the difficulty for eavesdroppers to perform brute-force or multiplexing attacks against a single key. Furthermore, because encryption and decryption are organized at the segment level, even if some segments suffer significant distortion during multi-hop transmission and multiple lossy encoding / decoding processes, the impact is more likely to be limited to a localized area. This helps maintain the recoverability and continuity of the entire frame's audio, achieving a balance between security and availability.
[0038] To address the common packet loss and erasure issues in real-world networks, this invention further introduces a missing detection and loss compensation mechanism into the reverse decryption process at the receiving end. The receiving end performs missing detection on the speech frame and its temporal sub-segments based on frame sequence number, timestamp, or segment sequence number to locate the missing position. When a missing speech frame or missing temporal sub-segment is detected, to maintain consistent frame and segment indices with the sending end for subkey synchronization and to ensure structural mismatch during inverse temporal scrambling and frame reconstruction, placeholder temporal data corresponding to the missing position is constructed and participates in subsequent inverse temporal scrambling and splicing output. The placeholder data can be generated by interpolation from the recovered speech before and after the missing position, or by copying adjacent sub-segments or frames with attenuation and smoothing processing. In cases of severe missing data, silence or comfortable noise can also be used as a substitute. Therefore, even in the presence of packet loss or erasure, the decryption process can maintain index consistency and processing link continuity, thereby improving the coherence of the recovered speech and reducing abrupt artifacts.
[0039] This invention introduces bandwidth control processing after encryption at the transmitting end to reduce the amount of transmitted data and adapt to link conditions. The transmitting end can limit the effective frequency band of the target by using low-pass or band-pass filtering, and downsample the filtered encrypted voice to reduce the sampling rate, thereby reducing the amount of data that needs to be transmitted. The transmitting end can also adaptively adjust the target bit rate or frame rate based on bandwidth estimation results, packet loss rate, or latency jitter, and adjust parameters such as the number of time-domain segments N per frame, frame length, or overlap area length in a linked manner, so that the transmission overhead dynamically matches the network conditions, while ensuring that the receiving end can complete decryption and recovery according to the corresponding parameters and maintain intelligibility.
[0040] Furthermore, this invention introduces session mode judgment and master key selection logic on the terminal side to achieve unified management of group public keys and peer-to-peer private keys. In group broadcast mode, the group public key is used to ensure that authorized terminals within the same group can decrypt and listen. In peer-to-peer private mode, a private key shared only between the two parties in the call is used to prevent other members from decrypting and recovering the data. This allows for flexible switching between different session modes without changing the underlying transmission protocol, improving the granularity of privacy protection and the efficiency of key management.
[0041] Regarding speech quality and auditory continuity, this invention introduces front-end noise reduction processing before encryption to improve the signal-to-noise ratio of the input speech, thereby helping to retain more effective speech information during subsequent lossy compression processes. After inverse transformation and inverse scrambling at the decryption end, back-end enhancement processing can be further introduced to suppress the effects of superimposed noise and quantization noise. Simultaneously, in segment reconstruction and boundary processing, by setting overlapping areas at the splicing points of adjacent sub-segments and employing smoothing methods such as cross-fading or overlapping addition, the popping and discontinuity caused by abrupt changes in amplitude and phase between segments can be alleviated, allowing the decrypted speech to maintain a relatively natural waveform transition even under complex noise and multi-hop lossy conditions.
[0042] Furthermore, this invention supports periodic updates of the master key during the session and incorporates information such as the session sequence number and random numbers into the key derivation process, thereby improving protection against replay attacks and long-term key leakage risks. This makes it difficult for attackers to decrypt subsequent voice content even if they obtain some historical keys. In summary, this invention achieves a synergistic effect through key technical features such as joint encryption in the time domain and transform domain, index-driven subkey derivation, multi-hop lossy encoding / decoding compatibility, packet loss and erasure compensation, adaptive bandwidth control, and group and peer-to-peer key management. It is applicable to end-to-end secure voice communication in various devices and systems, including Bluetooth headsets, mobile terminals, and cloud conferencing platforms. Attached Figure Description
[0043] Appendix Figure 1 This is a schematic diagram of the overall structure of an end-to-end voice encryption communication device.
[0044] Appendix Figure 2 This is a schematic diagram of the end-to-end voice encryption method.
[0045] Appendix Figure 3 This is a schematic diagram of the session mode determination and master key selection process.
[0046] Appendix Figure 4 This is a schematic diagram of the framing, segmentation, and temporal scrambling structure for a single frame of speech.
[0047] Appendix Figure 5 This is a schematic diagram of the transform domain encryption processing structure for the time domain sub-segment.
[0048] Appendix Figure 6 This is a schematic diagram of a multi-hop lossy voice communication channel structure.
[0049] Appendix Figure 7 This is a schematic diagram of the U-shaped neural network structure for front-end noise reduction.
[0050] Appendix Figure 8 This is a schematic diagram of the backend noise reduction processing flow and network structure.
[0051] Appendix Figure 9 This is a block diagram of an end-to-end voice encryption device.
[0052] Appendix Figure 10 This is a structural schematic diagram of an embodiment of a Bluetooth headset.
[0053] Appendix Figure 11 This is a schematic diagram of the missing detection and loss compensation process at the receiving end.
[0054] Appendix Figure 12 This diagram illustrates the bandwidth control, adaptive parameter selection, and synchronization at the transmitting end.
[0055] Figure Labels: Transmitter Voice Encryption Device - 10; Voice Acquisition Unit - 101; Front-end Noise Reduction Unit - 102; End-to-End Encryption Unit - 103; Key Management Unit - 104; Bluetooth / Network Transmission Unit - 105; Multi-hop Voice Communication Channel - 20; First Lossy Codec Unit - 201; Relay Node - 202; Second Lossy Codec Unit - 203; Receiver Voice Decryption Device - 30; Receiver Unit - 301; Decryption Unit - 302; Back-end Noise Reduction Unit - 303; Playback Unit - 304; Voice Acquisition and Front-end Noise Reduction Steps - 1; Frame Segmentation and Segmentation Steps - 2; Session Mode Judgment and Master Key Selection Steps - 3; Time Domain Scrambling Step Based on Master Key Steps - 4; Segment-by-Segment Transformation Domain Encryption Processing Steps - 5; Segment Reassembly and Boundary Smoothing Steps - 6; Encrypted Voice Transmission Steps - 7; Receiver Reverse Decryption Steps - 8; Receiver Back-end Noise Reduction Steps - 9; Session Mode Information Input Steps - 01; Group Public Key Selection Steps - 21 Point-to-point private key selection step - 22 Default preset key selection step - 23 Master key determination step - 31 Key derivation function step - 41 Subkey set generation step - 51 End-to-end voice encryption device - 900 Processor - 901 Memory - 902 Voice acquisition module - 903 Front-end noise reduction module - 904 Session mode judgment and master key selection module - 905 Frame segmentation and segmentation module - 906 Temporal scrambling module - 907 Transform domain encryption module - 908 Segment reconstruction and boundary smoothing module - 909 Communication interface module - 910 Bluetooth headset body - 1000 Microphone - 1001 Speaker - 1002 Bluetooth communication module - 1003 Voice processing chip - 1004 Front-end noise reduction and frame segmentation module - 1004a End-to-end encryption / decryption module - 1004b Session mode and key management module - 1004c Power supply / battery module - 1005 Detailed Implementation
[0056] The technical solution of this embodiment will now be clearly and completely described with reference to the accompanying drawings. In the description of this embodiment, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings and are only for the convenience of describing this embodiment and simplifying the description. They do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this embodiment. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying their relative importance.
[0057] In the following embodiments, the end-to-end voice encryption method and apparatus for multi-hop lossy channels proposed in this embodiment will be described in detail with reference to the accompanying drawings, enabling those skilled in the art to implement it without creative effort. It should be understood that the following description is merely an exemplary embodiment and does not constitute a limitation on the scope of protection of this embodiment.
[0058] In one embodiment, such as Figure 1 As shown, the end-to-end voice encryption communication system includes a transmitting voice encryption device 10, a multi-hop voice communication channel 20, and a receiving voice decryption device 30. The transmitting voice encryption device 10 internally comprises a voice acquisition unit 101, a front-end noise reduction unit 102, an end-to-end encryption unit 103, a key management unit 104, and a Bluetooth or network transmission unit 105. The voice acquisition unit 101 acquires analog voice signals from a microphone array or a single pickup channel and converts them into digital voice data. The front-end noise reduction unit 102 performs front-end noise reduction processing on the digital voice output by the voice acquisition unit 101 to reduce environmental noise and channel interference. The front-end noise-reduced voice signal is sent to the end-to-end encryption unit 103, where, under the control of the master key and sub-keys provided by the key management unit 104, it undergoes joint encryption processing in the time domain and transform domain. The resulting encrypted voice data is then encapsulated by the Bluetooth or network transmission unit 105 into a data stream conforming to the communication protocol and sent to the multi-hop voice communication channel 20. The multi-hop voice communication channel 20 can be composed of one or more lossy voice encoding / decoding and transmission units. Figure 1 The example includes a first lossy encoding / decoding unit 201, a relay node 202, and a second lossy encoding / decoding unit 203. These units encode, transmit, and decode encrypted voice between different network nodes or different devices. The receiving-end voice decryption device 30 includes a receiving unit 301, a decryption unit 302, a back-end noise reduction unit 303, and a playback unit 304. The receiving unit 301 obtains the encrypted voice stream from the multi-hop voice communication channel 20. The decryption unit 302 performs the decryption process in reverse to that of the sending-end end-to-end encryption unit 103. The back-end noise reduction unit 303 further enhances the decrypted voice with back-end noise reduction. Finally, the playback unit 304 outputs the understandable voice to the user, thus forming an end-to-end encrypted communication path from the sending-end voice encryption device 10 to the receiving-end voice decryption device 30.
[0059] In this communication system, the structure of the multi-hop voice communication channel 20 can be further understood as follows: Figure 6The diagram illustrates the structure of a multi-hop lossy voice communication channel. The first lossy codec unit 201 corresponds to the codec function of the Bluetooth or other short-range wireless link between the transmitting end and the first relay device. The relay node 202 corresponds to the network-side conferencing server or cloud voice processing node. The second lossy codec unit 203 corresponds to the lossy voice codec process between the receiving mobile terminal and the receiving headset. Multiple lossy compression and decompression operations can occur in the multi-hop voice communication channel 20. Each lossy codec operation may introduce distortions such as amplitude attenuation, quantization noise, and spectral truncation into the waveform of the encrypted voice. This embodiment, through the joint design of time-domain segments and transform-domain coefficients, maintains speech-like statistical characteristics similar to the original voice in terms of overall power spectrum distribution and short-time energy trajectory. This ensures that the first lossy codec unit 201 and the second lossy codec unit 203 can still achieve stable codec performance when processing the encrypted voice, guaranteeing that the encrypted voice transmitted through the multi-hop voice communication channel 20 can still be correctly decrypted into understandable voice at the receiving end.
[0060] In order to implement the above method flow at the device level, Figure 9 A block diagram of an end-to-end voice encryption device 900 is shown. The end-to-end voice encryption device 900 includes a processor 901 and a memory 902. The memory 902 stores program instructions that can run on the processor 901. When executed by the processor 901, these program instructions logically form multiple functional modules. The voice acquisition module 903 corresponds to... Figure 1 The voice acquisition unit 101 in the key management unit 104 is used to perform analog-to-digital conversion on the analog voice signal acquired by the external microphone and output digital voice data. The front-end noise reduction module 904 corresponds to the front-end noise reduction unit 102, which performs short-time Fourier transform and neural network front-end noise reduction on the digital voice output by the voice acquisition module 903 to obtain a noise-suppressed voice signal. The session mode determination and master key selection module 905 corresponds to the session mode determination and master key selection part in the key management unit 104, which is used to obtain mode information such as the group identifier and participant identifier of the current call session, and select a group public key or a point-to-point private key from a pre-configured key set as the master key for the current voice frame based on the mode information. The framing and segmentation module 906 corresponds to... Figure 2The framing and segmentation step 2 shown is used to divide the front-end noise-reduced speech signal into frames, and further divide each speech frame into multiple consecutive temporal sub-segments. The temporal scrambling module 907 corresponds to the master key-based temporal scrambling step 4, which macroscopically scrambles the order of adjacent temporal sub-segments according to the temporal permutation rules determined by the master key. The transform domain encryption module 908 corresponds to the segment-by-segment transform domain encryption processing step 5, which performs a Fast Fourier Transform (FFT) on each scrambled temporal sub-segment and generates transform domain coefficients, using the key derived function step 41 (see step 41) combining the master key, frame index, and segment index. Figure 3 The generated frequency domain encryption subkey is used to encrypt the effective frequency range within the sub-segment in the frequency domain, and then the encrypted time domain sub-segment is output through inverse fast Fourier transform (IFFT). The segment reassembly and boundary smoothing module 909 corresponds to step 6 of segment reassembly and boundary smoothing. It splices multiple encrypted time domain sub-segments according to the scrambled time sequence, sets an overlapping area at the splicing boundary and applies complementary weights to achieve cross-fading or overlapping addition processing, thereby smoothing the boundary of encrypted voice. The communication interface module 910 corresponds to the Bluetooth or network transmitting unit 105, and is used to send the encrypted voice signal to the multi-hop voice communication channel 20 according to a predetermined communication protocol, and receive encrypted voice feedback data from the multi-hop voice communication channel 20 when necessary.
[0061] In terms of specific terminal product form, this embodiment can be integrated into the Bluetooth headset body 1000. Figure 10A schematic diagram of a Bluetooth headset embodiment is shown. The Bluetooth headset body 1000 includes a microphone 1001, a speaker 1002, a Bluetooth communication module 1003, a voice processing chip 1004, and a power supply or battery module 1005. The microphone 1001 is used to collect the voice from the wearer's side and convert it into an electrical signal. The voice processing chip 1004 integrates a front-end noise reduction and framing module 1004a, an end-to-end encryption or decryption module 1004b, and a session mode and key management module 1004c. The front-end noise reduction and framing module 1004a performs similar functions to the front-end noise reduction module 904 and the framing and segmentation module 906 within the chip, performing short-time Fourier transform, complex-valued mask noise reduction, and frame and segment division on the voice from the microphone 1001. The end-to-end encryption or decryption module 1004b encrypts the voice signal according to the time-domain and transform-domain joint encryption structure of this embodiment in the transmission path, and performs the corresponding decryption in the reception path, ensuring that the encrypted voice signal received from the Bluetooth communication module 1003 can be restored to intelligible voice. The session mode and key management module 1004c maintains at least one group public key and multiple point-to-point private keys within the Bluetooth headset body 1000, and automatically selects the master key based on the group identifier and session participant identifier during the call, and uses it in conjunction with the end-to-end encryption or decryption module 1004b. The Bluetooth communication module 1003 is responsible for establishing a wireless connection with the external mobile terminal to realize the transmission and reception of encrypted voice signals. The power supply or battery module 1005 supplies power to the microphone 1001, voice processing chip 1004, Bluetooth communication module 1003, and speaker 1002, enabling the Bluetooth headset body 1000 to achieve end-to-end voice encryption and decryption under low power conditions. In one embodiment, the voice processing module for implementing front-end noise reduction and end-to-end voice encryption methods and the Bluetooth communication module 1003 can be integrated into the same audio processing chip to further reduce overall latency and power consumption.
[0062] In terms of methodology and process, such as Figure 2As shown, the end-to-end voice encryption method in this embodiment includes, in chronological order, a voice acquisition and front-end noise reduction step 1, a framing and segmentation step 2, a session mode determination and master key selection step 3, a master key-based temporal scrambling step 4, a segment-by-segment transform domain encryption processing step 5, a segment reconstruction and boundary smoothing step 6, an encrypted voice transmission step 7, a receiver-side reverse decryption step 8, and a receiver-side back-end noise reduction step 9. The voice acquisition and front-end noise reduction step 1 is jointly completed by the voice acquisition unit 101 or the voice acquisition module 903 and the front-end noise reduction unit 102 or the front-end noise reduction module 904. The framing and segmentation step 2 is completed by the framing and segmentation module 906. The session mode determination and master key selection step 3 is executed by the key management unit 104 and the session mode determination and master key selection module 905. The time-domain scrambling step 4 based on the master key is executed by the time-domain scrambling module 907; the segment-by-segment transform-domain encryption processing step 5 is implemented by the transform-domain encryption module 908; the segment reconstruction and boundary smoothing step 6 is handled by the segment reconstruction and boundary smoothing module 909; and the encrypted voice transmission step 7 sends encrypted voice data to the multi-hop voice communication channel 20 via Bluetooth or network transmission unit 105 and communication interface module 910. The reverse decryption step 8 at the receiving end is executed by the receiving unit 301 and decryption unit 302 inside the receiving end voice decryption device 30; the back-end noise reduction step 9 at the receiving end is executed by the back-end noise reduction unit 303; and finally, the enhanced decrypted voice is played by the playback unit 304.
[0063] In step 1 of voice acquisition and front-end noise reduction, in order to improve the signal-to-noise ratio of the voice before encryption, this embodiment adopts the following... Figure 7 The diagram illustrates a U-shaped neural network structure for front-end denoising. The front-end denoising unit 102 or the front-end denoising module 904 first performs a short-time Fourier transform on the digital speech output by the speech acquisition unit 101 or the speech acquisition module 903 to obtain complex time-frequency features. These complex time-frequency features are then input into a U-shaped neural network containing an encoding part, an enhancement part, and a decoding part, and equipped with multi-level skip connections. The encoding part progressively extracts high-level time-frequency features through multi-level downsampling, while the decoding part restores the original frequency resolution through multi-level upsampling. The skip connections preserve low-level detail information. The network output is a complex-valued mask with an amplitude limited to the range of zero to one. This mask weights the real and imaginary parts of the complex time-frequency features, respectively. After inverse short-time Fourier transform, the resulting speech signal exhibits suppressed noise components in both the spectral and temporal domains, thus forming the front-end denoised speech signal used as the input for subsequent encryption.
[0064] In step 2 of framing and segmentation, the noise-reduced speech signal is divided into consecutive speech frames on the time axis. Each speech frame is further divided into multiple consecutive time-domain sub-segments. Figure 4The diagram illustrates the framing, segmentation, and temporal scrambling structure of a single-frame speech. The framing and segmentation module 906 can divide the speech data into multiple temporal segments within each speech frame according to a fixed length or other preset rules. The number of sampling points in each temporal segment is less than the total number of sampling points in the corresponding speech frame. This two-level structure design at the frame and segment levels allows for macroscopic scrambling at the segment level in the temporal domain, while encryption is performed at the segment spectrum level in the transform domain. This approach helps to improve the fine-grainedness and flexibility of encryption while ensuring speech intelligibility.
[0065] The logic of session mode determination and master key selection step 3 can be combined. Figure 3 The following explains the session mode determination and master key selection process. In the session mode information input step 01, the session mode determination and master key selection module 905 obtains the mode information of the current call session. This mode information may include at least the group identifier and the session participant identifier. The key management unit 104 or the session mode and key management module 1004c maintains a set of pre-configured keys. This set includes at least one group public key and multiple point-to-point private keys. Each group public key is shared among multiple terminals within the same group, and each point-to-point private key is stored in pairs between different terminal pairs. In the group public key selection step 21, when the session mode determination and master key selection module 905 determines that the current voice segment belongs to group broadcast voice based on the mode information, it selects the group public key of the corresponding group from the pre-configured key set as the master key of the current voice frame. In the point-to-point private key selection step 22, when it is determined that the current voice segment belongs to point-to-point private voice in the same group, it selects the private key that is shared only between the two current terminals from the multiple point-to-point private keys as the master key. In the default key selection step 23, for scenarios where a specific group or terminal pair cannot be matched, a default key can be selected as the master key for compatibility or low-security call requirements. In the master key determination step 31, the final determined master key will be used for subsequent temporal scrambling and subkey derivation. The key derivation function step 41 generates a frame-level or segment-level subkey set based on the master key, the frame index of the voice frame, and the segment index of each temporal segment. The subkey set generation step 51 establishes a correspondence between each subkey and a specific temporal segment, so that different temporal segments within the same voice frame use different frequency domain encryption subkeys when encrypting in the transform domain. The master key can be negotiated between relevant terminals through a key negotiation protocol at the time of each session establishment, and is updated during the session according to a predetermined number of frames or time intervals combined with random numbers and session sequence numbers. The updated master key continues to be used for temporal scrambling and frequency domain encryption subkey derivation for subsequent new voice frames, thereby improving the ability to resist replay attacks and brute-force attacks.
[0066] In the master key-based temporal scrambling step 4, the temporal scrambling module 907 uses the temporal permutation rules obtained from the master key through the key derivation function step 41 to macroscopically scramble the temporal arrangement of multiple temporal sub-segments obtained from segmentation within each speech frame. For example... Figure 4 As shown, the original sequential arrangement of multiple segments is disrupted by scrambling, thus altering the original continuity of the speech content on the timeline. Since the local time structure within each time-domain segment is maintained subsequently, short-term speech characteristics are preserved when using lossy encoding and decoding in the multi-hop speech communication channel 20, which helps maintain the encodeability and decodeability of encrypted speech.
[0067] In step 5 of the segment-by-segment transform-domain encryption process, the transform-domain encryption module 908 sequentially performs transform-domain encryption operations on each time-domain sub-segment after time-domain scrambling. Figure 5 This diagram illustrates the transform-domain encryption process for time-domain segments. For any selected time-domain segment, the transform-domain encryption module 908 first uses a Fast Fourier Transform (FFT) to convert the segment to the frequency domain, obtaining a set of transform-domain coefficients. Then, based on the frequency-domain encryption subkey corresponding to the segment, frequency-domain encryption is performed on the transform-domain coefficients within a first frequency interval corresponding to the intersection of the effective speech bandwidths of each lossy speech codec in the multi-hop speech communication channel 20. Frequency-domain encryption is not performed in a second frequency interval outside the first frequency interval to maintain the overall energy distribution close to the original speech. Finally, an Inverse Fast Fourier Transform (IFFT) is performed on the frequency-domain encrypted transform-domain coefficients to restore them to the time-domain signal, generating the encrypted time-domain data for that segment. By sequentially performing the above FFT and IFFT processes on all time-domain segments, the encryption process of the entire frame of speech is composed of multiple small-point FFTs and IFFTs, thus achieving better implementation in terms of computational complexity and latency, while simultaneously providing fine-grained encryption of the speech content at the transform-domain level.
[0068] In step 6 of the segment reassembly and boundary smoothing process, the segment reassembly and boundary smoothing module 909 reassembles each encrypted time-domain segment into a complete encrypted speech frame according to the scrambled time-domain order. To mitigate potential amplitude abrupt changes and phase discontinuities at the boundaries of different segments, the segment reassembly and boundary smoothing module 909 sets a predetermined overlap region at the junction of adjacent encrypted time-domain segments. Within the overlap region, a set of complementary weighting coefficients are applied to the samples at the end of the preceding encrypted segment and the samples at the beginning of the following encrypted segment, and then superimposed. Through cross-fading or overlapping addition, the two segments transition naturally within the overlap region, thereby ensuring that the entire encrypted speech waveform sounds smoother and more coherent. The overlap region length is shorter than the time-domain segment length, which can reduce additional computation and latency while ensuring the smoothing effect.
[0069] Before the encrypted voice signal is sent to the multi-hop voice communication channel at the sending end, the following can also be performed: Figure 12 The bandwidth control process is illustrated. The transmitting end can obtain link status input information, including bandwidth estimation, packet loss rate, and / or jitter, and the adaptive controller / gear selector selects a gear identifier or parameter set identifier based on the link status. The parameter linkage module adjusts at least one of the following parameters accordingly: number of time-domain segments per frame, frame length, overlap length, target bit rate, and / or frame rate. Subsequently, the bandwidth control process may include at least one of the following: low-pass and / or band-pass filtering of the encrypted voice signal to limit the target frequency band; anti-aliasing low-pass processing and downsampling of the filtered signal to reduce the sampling rate; and / or adaptive adjustment of the target bit rate and / or frame rate of the voice codec based on the link status, thereby reducing the amount of transmitted data and maintaining decryptability and recovery at the receiving end. The encapsulation and transmission module may carry the gear identifier / parameter set identifier when generating data packets, so that the receiving end performs framing, segmentation, and decryption recovery processing according to the corresponding parameters. In the encrypted voice transmission step 7, the Bluetooth or network transmission unit 105 and the communication interface module 910 are responsible for encapsulating the encrypted voice frame output by the segment reconstruction and boundary smoothing module 909 into an audio data stream adapted to the specific transmission protocol, and sending it to the receiving end voice decryption device 30 through the multi-hop voice communication channel 20. Figure 6 The multi-hop lossy voice communication channel structure shown can include multiple lossy voice codec units and relay nodes of different types. Through the time domain and transform domain encryption structure and frequency range selection strategy of this embodiment, the encrypted voice still retains the characteristics of being decryptable and understandable after passing through multiple processing stages such as the first lossy codec unit 201, the relay node 202 and the second lossy codec unit 203.
[0070] Before framing and segmenting, the receiving end can read the file segment identifier / parameter set identifier from the received data packet and configure the framing and segmentation parameters (such as the number of segments, frame length, overlap length, sampling rate, etc.) accordingly to ensure consistency with the processing parameters after bandwidth control at the sending end. In the reverse decryption step 8 at the receiving end, the receiving unit 301 receives the encrypted voice signal processed by multiple lossy encoding and decoding units from the multi-hop voice communication channel 20. Under the control of the master key provided by the session mode and key management module 1004c or other key management device, the decryption unit 302 executes the decryption process in reverse to the encryption process at the sending end. The decryption unit 302 first frames and segments the received encrypted voice data in the same way as at the sending end, and then performs a Fast Fourier Transform (FFT) on each encrypted time-domain segment to obtain the encrypted frequency-domain coefficients. Using the same master key as the sender, along with the same frame and segment indices, the corresponding frequency domain decryption subkey is derived through key derivation function step 41. Then, the inverse frequency domain decryption operation is performed on the encrypted frequency domain coefficients within the same first frequency interval as during encryption. Finally, an inverse fast Fourier transform (IFFT) is performed to obtain the decrypted time domain segment. In the reverse decryption step 8 at the receiver, to address potential packet loss and erasure issues in multi-hop voice communication channels, further mechanisms such as... can be introduced. Figure 11 The missing data detection and loss compensation process is illustrated. After receiving encrypted voice data / data packets, the receiving unit 301 parses the metadata carried therein (including frame sequence number, timestamp, and / or segment sequence number), and the expected frame sequence number and expected segment sequence number are generated by the expected index generator. The missing data detection comparator compares the expected index with the received index and outputs a missing marker and a list of missing positions. When it is determined that there are missing voice frames and / or missing time domain segments, in order to maintain the frame index and / or segment index consistent with the sending end for subkey synchronization derivation, the index advancement / synchronization module advances the corresponding frame sequence number and segment sequence number even in the case of missing data, and the placeholder data generator generates placeholder decryption time domain data corresponding to the missing position. The placeholder decryption time domain data can be obtained by at least one of the following strategies: performing time domain interpolation on the recovered voice data before and after the missing position; masking the missing position and replacing it with a copy of an adjacent voice frame or an adjacent time domain segment and applying energy attenuation and / or window function smoothing; or predictively replacing the missing position based on the transform domain coefficients of adjacent voice frames and then performing the inverse transform. Subsequently, the placeholder segments are merged into the normal decryption segments and participate together in the subsequent inverse temporal scrambling and splicing output. This ensures that the frame reconstruction structure remains intact and improves the coherence of the decrypted speech even in the event of packet loss or erasure. Then, inverse temporal scrambling is performed according to the same temporal permutation rules as the sender, but in the opposite direction, rearranging each segment to its original time position to form the decrypted speech frame, completing step 8 of the receiver's reverse decryption process.
[0071] The receiver's back-end noise reduction step 9 is executed by the back-end noise reduction unit 303, and its processing flow can be combined with... Figure 8 The back-end denoising process and network structure are explained below. The back-end denoising unit 303 first performs a short-time Fourier transform on the decrypted noisy speech to obtain complex spectral features. These spectral features are then input into a U-shaped neural network identical to that used in the front-end denoising or with reduced parameters in terms of channel and layer count, to obtain an initial complex-valued mask. Since the decrypted speech may be affected by quantization noise and artifact interference introduced by multi-hop lossy channels, this embodiment employs a causal temporal weighted deep filtering method. It utilizes only the time-frequency information of the current frame and its preceding frames to weight and filter the initial mask in both time and frequency directions, obtaining an optimized mask that considers both current and past frame information. Next, based on the estimated prior and posterior signal-to-noise ratios, the gain at each frequency point is calculated according to the minimum mean square error criterion. This gain is used to correct the optimized mask, forming the final mask. The final mask is multiplied point by point with the decrypted noisy spectrum, and then the inverse short-time Fourier transform is used to recover the time-domain speech signal. This speech signal is then played by the playback unit 304 or the speaker 1002, so that the user can obtain high-quality speech that has undergone end-to-end encryption and decryption as well as front-end and back-end noise reduction processing.
[0072] Through the above combination Figures 1 to 10As can be seen from the detailed description of the reference numerals in the accompanying drawings, the end-to-end voice encryption method and apparatus for multi-hop lossy channels provided in this embodiment comprises: a transmitting end voice encryption device 10, a voice acquisition unit 101, a front-end noise reduction unit 102, an end-to-end encryption unit 103, a key management unit 104, a Bluetooth or network transmitting unit 105, a multi-hop voice communication channel 20, a first lossy encoding / decoding unit 201, a relay node 202, a second lossy encoding / decoding unit 203, a receiving end voice decryption device 30, a receiving unit 301, and a decryption unit 302. The method involves the coordinated operation of backend noise reduction unit 303 and playback unit 304, combined with the complete process flow of voice acquisition and frontend noise reduction steps 1, framing and segmentation steps 2, conversation mode judgment and master key selection steps 3, master key-based temporal scrambling steps 4, segment-by-segment transform domain encryption processing steps 5, segment reconstruction and boundary smoothing steps 6, encrypted voice transmission steps 7, receiver-side reverse decryption steps 8, and receiver-side backend noise reduction steps 9, as well as conversation mode information input steps 01, group public key selection steps 21, and point-to-point private key selection steps 22. The key management process consists of three steps: default preset key selection (step 23), master key determination (step 31), key derivation function (step 41), and subkey set generation (step 51). Simultaneously, at the device level, it is constructed using an end-to-end voice encryption device 900, a processor 901, a memory 902, a voice acquisition module 903, a front-end noise reduction module 904, a session mode judgment and master key selection module 905, a framing and segmentation module 906, a temporal scrambling module 907, a transform domain encryption module 908, a segment reconstruction and boundary smoothing module 909, and a communication interface module 910. A practically deployable encryption device structure is proposed, and in the terminal form, an integrated end-to-end voice encryption and decryption scheme is realized through a structure including a Bluetooth headset body 1000, a microphone 1001, a speaker 1002, a Bluetooth communication module 1003, a voice processing chip 1004, a front-end noise reduction and framing module 1004a, an end-to-end encryption or decryption module 1004b, a session mode and key management module 1004c, and a power supply or battery module 1005, thereby fully supporting the technical solutions defined in the claims and overcoming the shortcomings of the prior art.
[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An end-to-end voice encryption method suitable for multi-hop lossy channels, characterized in that, include: The acquired voice signal to be encrypted is subjected to front-end noise reduction processing to obtain the front-end noise-reduced voice signal; The front-end denoised speech signal is divided into frames, and each speech frame is divided into N consecutive time-domain segments in the time domain, where N>1. Based on the current call session mode information, the master key for the voice frame is selected from the pre-configured key set. The key set includes a public key for group calls and a private key for point-to-point calls. When the current session is detected as a group call, the public key is selected as the master key. When the current session is detected as a point-to-point call, the private key shared only between the terminals of the two parties in the call is selected as the master key. Based on the temporal permutation rules determined by the master key, the temporal arrangement of the N temporal segments is scrambled to obtain a temporal segment sequence after macroscopic temporal scrambling; For each scrambled time-domain segment, the following operation is performed: Based on the master key and combined with the frame identifier and / or segment identifier corresponding to the time-domain segment, a frequency-domain encryption subkey corresponding to the time-domain segment is generated through key derivation, so that the frequency-domain encryption subkeys corresponding to different time-domain segments within the same voice frame are different from each other, and the receiving end can synchronously generate the corresponding frequency-domain encryption subkey based on the same master key and the frame identifier and / or segment identifier; The time domain segment is transformed to obtain the transform domain coefficients of the time domain segment. The transform domain coefficients are then encrypted in the frequency domain using the frequency domain encryption subkey corresponding to the time domain segment. Finally, the encrypted transform domain coefficients are inversely transformed to obtain the encrypted time domain data of the time domain segment. The transformation and the inverse transformation are performed separately at the time domain segment granularity. Each encrypted time-domain segment is spliced together according to the scrambled time sequence, and smoothing is performed at the splicing boundary of adjacent encrypted time-domain segments to obtain an encrypted voice signal.
2. The method according to claim 1, characterized in that, The encryption method employs an alternating encryption structure in the time domain and transform domain, comprising: performing macroscopic structural processing on each speech frame in the time domain, i.e., dividing the speech frame into N consecutive time domain segments, and macroscopically scrambling the temporal arrangement of the N time domain segments according to the time domain permutation rule determined by the master key; and processing the microscopic structure of each time domain segment in the transform domain, transforming each time domain segment after time domain scrambling to obtain transform domain coefficients, performing frequency domain encryption processing on the transform domain coefficients using the frequency domain encryption subkey corresponding to the time domain segment in the transform domain, and performing inverse transformation on the encrypted transform domain coefficients to obtain the corresponding encrypted time domain segment, wherein the frequency domain encryption processing is performed segment by segment sequentially.
3. The method according to claim 1, characterized in that, The frame identifier includes at least one of frame index, frame sequence number, and timestamp; the segment identifier includes at least one of segment index and segment sequence number; the key derivation method is a key derivation function, the input of which includes at least the master key and the frame identifier and / or segment identifier, and may further include at least one of session identifier, session sequence number, and random number, to generate a frequency domain encryption subkey corresponding one-to-one with each time domain sub-segment, thereby enabling the sending end and receiving end to achieve subkey synchronization without additionally transmitting the frequency domain encryption subkey.
4. The method according to claim 1, characterized in that, The transformation performed on each of the aforementioned time-domain sub-segments is a Fast Fourier Transform (FFT), and the number of transformation points of the FFT is less than the number of sampling points of the corresponding speech frame; the inverse transformation performed on the encrypted transform domain coefficients is an Inverse Fast Fourier Transform (IFT) corresponding to the FFT; thus, when encrypting the entire frame of speech, by performing the FFT and IFT on each time-domain sub-segment respectively, the multiple transformations and inverse transformations of the sub-segments are replaced by a single overall transformation and inverse transformation of the entire frame of speech.
5. The method according to claim 1, characterized in that, The smoothing process performed at the splicing boundary of adjacent encrypted time-domain segments includes: setting an overlap region of a predetermined length at the splicing point of adjacent encrypted time-domain segments; applying a set of complementary weighting coefficients to the tail sample of the preceding encrypted time-domain segment and the head sample of the following encrypted time-domain segment and superimposing them, thereby achieving cross-fading and / or overlapping addition processing of adjacent encrypted time-domain segments to reduce amplitude and phase abrupt changes at the splicing boundary and ensure the listening quality of the output encrypted speech, wherein the length of the overlap region is less than the length of the time-domain segment.
6. The method according to claim 1, characterized in that, Among the transform domain coefficients obtained by transforming each of the aforementioned time domain segments, frequency domain encryption processing is performed only on the transform domain coefficients within a first frequency interval corresponding to the intersection of the effective speech bandwidths of each lossy speech codec in the multi-hop speech communication channel, according to the frequency domain encryption subkey. The frequency domain encryption processing is not performed in a second frequency interval outside the first frequency interval. This ensures that the encrypted speech signal maintains speech-like statistical characteristics similar to the original speech signal in terms of overall power spectrum distribution and short-time energy trajectory, so that the encrypted speech signal can be transmitted in a multi-hop speech communication channel including one or more lossy speech codecs and transmission units, and can be recovered into intelligible speech at the receiving end through a decryption process that is the reverse of the encryption process.
7. The method according to claim 1, characterized in that, The front-end noise reduction process includes: performing a short-time Fourier transform on the acquired speech signal to be encrypted to obtain complex time-frequency features; inputting the complex time-frequency features into a U-shaped neural network containing an encoding part, an enhancement part, and a decoding part and having multiple levels of skip connections; the output of the U-shaped neural network having a complex value mask with an amplitude limited to the range of 0 to 1; using the complex value mask to weight the real and imaginary parts of the complex time-frequency features respectively; and performing an inverse short-time Fourier transform on the weighted time-frequency features to obtain the front-end noise-reduced speech signal as the encryption input.
8. The method according to claim 1, characterized in that, After decrypting the received encrypted voice signal at the receiving end, the process also includes performing back-end noise reduction on the decrypted voice signal. The back-end noise reduction process includes: performing a short-time Fourier transform on the decrypted voice signal to obtain complex spectral features; inputting the complex spectral features into a U-shaped neural network with the same or reduced parameters as the U-shaped neural network used in the front-end noise reduction process to obtain an initial complex value mask; optimizing the initial complex value mask using causal time-weighted depth filtering that utilizes only the time-frequency information of the current frame and its preceding frames to obtain an optimized mask; calculating the frequency gain based on the estimated prior signal-to-noise ratio and posterior signal-to-noise ratio according to the minimum mean square error criterion to correct the optimized mask to obtain a final mask; multiplying the final mask by the decrypted noisy spectrum and performing an inverse short-time Fourier transform to output the back-end noise-reduced voice signal.
9. The method according to claim 1, characterized in that, The pre-configured key set includes at least one group public key and multiple peer-to-peer private keys. The group public key is shared among multiple terminals within the same group, and each peer-to-peer private key is shared in pairs between different terminal pairs. Each peer-to-peer private key is stored only in its corresponding terminal pair.
10. The method according to claim 1, characterized in that, The current call session mode information includes a group identifier and / or a session participant identifier. The sending end determines whether the current voice segment belongs to a group broadcast voice or a point-to-point private voice based on signaling messages and / or session control information from upper-layer applications. Accordingly, it selects the group public key or the corresponding terminal's point-to-point private key as the master key for encryption. This allows terminals holding the group public key to decrypt group broadcast voice in group call scenarios, and two terminals holding the point-to-point private key to decrypt the corresponding private voice in point-to-point private call scenarios within the same group. Other terminals that do not hold the group public key or the point-to-point private key can only receive incomprehensible voice.
11. The method according to claim 1, characterized in that, The master key is negotiated between relevant terminals through a key negotiation protocol at the time of each session establishment, and is updated during the session according to a predetermined number of frames or time intervals, combined with random numbers and / or session sequence numbers. The updated master key is used for time-domain scrambling and frequency-domain encryption subkey derivation of subsequent voice frames to improve the ability to resist replay attacks and brute-force attacks.
12. The end-to-end voice encryption method according to claim 1, characterized in that, The process of framing and segmenting, transforming, inverse frequency domain encryption, and inverse time domain scrambling of the encrypted voice signal transmitted through the multi-hop voice communication channel at the receiving end also includes missing detection and loss compensation processing, which includes: Based on the frame sequence number and / or timestamp and / or segment sequence number carried in the received encrypted voice data, missing voice frames and / or time-domain segments are detected to determine missing voice frames and / or missing time-domain segments. When missing voice frames and / or time-domain segments are determined, in order to maintain the frame index and / or segment index consistent with the sending end for subkey synchronization derivation, placeholder decryption time-domain data corresponding to the missing voice frames and / or time-domain segments is generated, and the placeholder decryption time-domain data is used in subsequent inverse time-domain scrambling and splicing output. The placeholder decryption time-domain data is obtained through at least one of the following strategies: performing time-domain interpolation on the recovered voice data before and after the missing position; masking the missing position and replacing it with a copy of the adjacent voice frame or adjacent time-domain segment and applying energy attenuation and / or window function smoothing; or predictively replacing the missing position based on the transform domain coefficients of the adjacent voice frames before performing the inverse transform.
13. The method according to claim 1, characterized in that, Before transmitting the encrypted voice signal to the multi-hop voice communication channel, a bandwidth control process is also included, which includes at least one of the following: low-pass and / or band-pass filtering of the encrypted voice signal to limit the target frequency band; The filtered signal is downsampled to reduce the sampling rate; and / or the transmitter bit rate and / or frame rate are adaptively adjusted based on network bandwidth, packet loss rate and / or latency jitter, the adjustment including adjusting the target bit rate of the speech codec and / or adjusting at least one of the number of time-domain segments N per frame, frame length and overlap length, to reduce the amount of transmitted data and maintain decryptability and recovery at the receiver.
14. An end-to-end encryption device for voice signals, characterized in that, The device includes a processor and a memory, wherein the memory stores program instructions that can be executed on the processor, and when the program instructions are executed by the processor, the processor performs the end-to-end voice encryption method for voice signals according to any one of claims 1 to 13.
15. A Bluetooth headset for performing the end-to-end voice encryption method for multi-hop lossy channels as described in any one of claims 1 to 13, characterized in that, The device includes a microphone, a speaker, a Bluetooth communication module, a processor, and a memory. The memory stores program instructions that can run on the processor. When the program instructions are executed by the processor, the processor is configured to: perform front-end noise reduction processing on the voice signal collected by the microphone, determine the call session mode, select a master key, and perform encryption processing; generate encrypted voice signals for each voice frame and send them to an external terminal device through the Bluetooth communication module; and / or receive the encrypted voice signals sent by the external terminal device from the Bluetooth communication module and perform a decryption process that is the opposite of the end-to-end voice encryption method to obtain a decrypted voice signal and play it through the speaker.
16. The Bluetooth headset according to claim 15, characterized in that, The voice processing module, which implements the front-end noise reduction and end-to-end voice encryption methods, is integrated with the Bluetooth communication module in the same audio processing chip.
Citation Information
Patent Citations
Vocoder penetration encryption and decryption method and system and voice encryption and decryption method
CN105071896A
Voice encryption method applied to narrow-band wireless digital communication system
CN103002406A
Voice encryption method and device for voice band compression system
CN105788602A