Encryption transmission method and device and electronic equipment

By combining voice feature analysis with a differentiated encryption strategy based on dynamic quantum key sharding, the problems of key update difficulty and insufficient synchronization in traditional encryption methods are solved, achieving efficient and secure encrypted transmission in voice calls.

CN120856431APending Publication Date: 2025-10-28CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511116801.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional encryption methods face challenges in key fragmentation and synchronization during voice calls, especially in quantum encryption technology, where key updates are difficult and security is insufficient.

Method used

By analyzing voice features in real time and integrating them with dynamic quantum key sharding, a differentiated encryption strategy is adopted. Newly generated quantum keys are used for encryption during active voice periods, the previous keys are reused during silent periods, and key updates are triggered at semantic boundaries to ensure security and synchronization.

Benefits of technology

It improves the efficiency of quantum key utilization and call security, realizes key sharding and synchronization, and improves the security and resource utilization efficiency of voice call systems under the threat of quantum computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856431A_ABST
    Figure CN120856431A_ABST
Patent Text Reader

Abstract

The invention relates to an encryption transmission method and device and electronic equipment. The method comprises the following steps: performing frame segmentation on an acquired voice signal to obtain voice frames; determining the voice state of each voice frame; the voice state comprises a silent frame and an active frame; encrypting the current active frame by adopting a first encryption mode, and encrypting the current silent frame by adopting a second encryption mode to obtain an encrypted voice frame; wherein the first encryption mode comprises the step of encrypting the current active frame by adopting the obtained key segment; the second encryption mode comprises multiplexing a key fragment used by the first encryption mode for encrypting the previous active frame; and packaging the encrypted voice frame into a data packet and outputting the data packet. By adopting the method, secret key fragmentation and synchronization in the voice communication process can be effectively realized in the quantum encryption process, so that the communication security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network technology and security, and in particular to an encrypted transmission method, apparatus and electronic device. Background Technology

[0002] With the development of digitalization, data security has become paramount, and encryption technology is widely used as a key means to ensure data security. Currently, encryption technology is a core means of ensuring communication security, preventing data from being eavesdropped on, tampered with, or leaked during transmission. However, traditional encryption methods suffer from poor security. Summary of the Invention

[0003] Therefore, it is necessary to provide an encrypted transmission method, device, and electronic device to address the aforementioned technical problems. This device can effectively achieve key fragmentation and synchronization during voice calls in the quantum encryption process, thereby improving the utilization rate of quantum keys and the security of calls.

[0004] In a first aspect, this application provides an encrypted transmission method applied to a voice transmitting end, comprising:

[0005] The acquired speech signal is segmented into frames to obtain individual speech frames;

[0006] Determine the speech state of each speech frame; speech state includes silent frames and active frames;

[0007] The current active frame is encrypted using a first encryption method, and the current silent frame is encrypted using a second encryption method to obtain an encrypted voice frame. The first encryption method includes encrypting the current active frame using an acquired key fragment, and the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame.

[0008] Encapsulate encrypted voice frames into data packets and output them.

[0009] In one embodiment, the frame type of each speech frame is determined, including:

[0010] Extract the speech signal energy features of each speech frame; the speech signal energy features include energy values;

[0011] Voice frames with energy values ​​higher than a preset threshold are identified as active frames, while voice frames with energy values ​​lower than a preset threshold are identified as silent frames.

[0012] In one embodiment, before performing frame segmentation on the acquired speech signal to obtain each speech frame, the method further includes:

[0013] In response to the synchronous acquisition of the quantum key stream with the voice receiver, the key pointer is initialized to the starting position; wherein, the quantum key stream includes key segments; the key pointer is used to identify the position of the key segment in the quantum key stream.

[0014] In one embodiment, the first encryption method includes: extracting a key segment from the quantum key stream based on the position pointed to by the current key pointer, and encrypting the current active frame using a symmetric encryption algorithm;

[0015] The second encryption method includes: reusing the key fragment used in the first encryption method to encrypt the previous active frame, and encrypting the current silent frame using a symmetric encryption algorithm.

[0016] In one embodiment, the voice state further includes semantic boundaries; the method also includes:

[0017] Extract the frequency domain features of the speech signal from each speech frame; the frequency domain features of the speech signal include the fundamental frequency;

[0018] When a semantic boundary in a speech frame is detected based on changes in the base frequency, the current key pointer is updated to obtain the updated key pointer.

[0019] In one embodiment, the method further includes: if no semantic boundary is detected in the voice frame, embedding the current key pointer into the redundant field of the corresponding encrypted voice frame.

[0020] In one embodiment, the method further includes: if a semantic boundary in a speech frame is detected, embedding the current key pointer and the updated key pointer into the redundant field of the corresponding encrypted speech frame.

[0021] Secondly, this application also provides an encrypted transmission method applied to a voice receiving end, the method comprising:

[0022] Receive data packets from the voice sender;

[0023] The encrypted voice frames in the data packet are decrypted using the obtained key fragment to obtain the decrypted voice frames.

[0024] In one embodiment, the step of obtaining a key fragment includes:

[0025] The data packet is parsed to extract the encrypted voice frame and the key pointer corresponding to the encrypted voice frame; the key pointer corresponding to the encrypted voice frame includes at least the current key pointer.

[0026] Based on the key pointer, the corresponding key segment is extracted from the local quantum key stream. Based on the key segment, the encrypted voice frame is decrypted using a symmetric decryption algorithm to obtain the decrypted voice frame.

[0027] In one embodiment, the key pointer corresponding to the encrypted voice frame further includes an updated key pointer; the method further includes:

[0028] Update the local key pointer based on the updated key pointer.

[0029] Thirdly, this application also provides an encrypted transmission device for use in a voice transmission end, comprising:

[0030] The segmentation unit is used to segment the acquired speech signal into frames to obtain each speech frame;

[0031] The state determination unit is used to determine the speech state of each speech frame; the speech state includes silent frames and active frames.

[0032] An encryption unit is used to encrypt the current active frame using a first encryption method and to encrypt the current silent frame using a second encryption method to obtain an encrypted voice frame; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame;

[0033] The data output unit is used to encapsulate encrypted voice frames into data packets and output them.

[0034] Fourthly, this application also provides an encrypted transmission device for use in a voice receiving end, the device comprising:

[0035] A data receiving unit is used to receive data packets from the voice transmitter.

[0036] The decryption unit is used to decrypt the encrypted voice frame in the data packet using the acquired key fragment, so as to obtain the decrypted voice frame.

[0037] Fifthly, this application also provides an encrypted transmission system, including a voice transmitter and a voice receiver;

[0038] The voice transmitting end is used to implement the steps in the method provided in the first aspect above;

[0039] The voice receiver is used to implement the steps in the method provided in the second aspect above.

[0040] Sixthly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the methods described above.

[0041] In a seventh aspect, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the methods described above.

[0042] Eighthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the methods described above.

[0043] The aforementioned encrypted transmission method, apparatus, electronic device, computer-readable storage medium, and computer program product, by segmenting the acquired voice signal into frames at the voice transmitting end and determining the voice state of each voice frame, encrypting each voice frame using different key fragments according to different voice states, and encapsulating the encrypted voice frames into data packets for output, thereby combining voice state with key management and employing differentiated encryption strategies based on voice state to improve transmission security; simultaneously, at the voice receiving end, the encrypted voice frames are decrypted using the acquired key, ensuring the correlation or synchronization between the decryption method at the voice receiving end and the encryption method at the voice encryption end, thereby improving transmission security. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating an encrypted transmission method applied to a voice transmitter in one embodiment.

[0046] Figure 2 This is a flowchart illustrating an encrypted transmission method applied to a voice receiver in one embodiment.

[0047] Figure 3 This is a structural block diagram of an encrypted transmission device applied to a voice transmitting end in one embodiment;

[0048] Figure 4 This is a structural block diagram of an encrypted transmission device applied to a voice receiver in one embodiment;

[0049] Figure 5 This is a schematic diagram of the workflow of an encrypted transmission system in one embodiment;

[0050] Figure 6 This is a schematic diagram illustrating the specific process of an encrypted transmission method in one embodiment;

[0051] Figure 7 This is a diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0053] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0054] Traditional quantum-key-based voice encryption technologies often employ fixed keys or periodically updated key strategies. Some techniques periodically sample and encode voice data at low rates in quantum networks, utilizing quantum key encryption and jitter buffers. Others introduce quantum encryption and decryption modules into the voice communication device, converting the voice signal into a single-photon signal for transmission. However, none of these technologies dynamically manage the key in accordance with the characteristics of the voice content. Therefore, traditional encryption methods face challenges such as difficulty in updating symmetric encryption keys and weak quantum resistance in asymmetric encryption. In particular, while quantum encryption technology offers high security, it still presents challenges in key fragmentation and synchronization during voice calls.

[0055] Based on the aforementioned traditional technologies, embodiments of this application provide an encrypted transmission method, apparatus, and electronic device. By integrating real-time analysis of voice features with dynamic quantum key fragmentation, differentiated encryption is implemented, achieving on-demand fragmentation and dynamic encryption of quantum keys. This effectively realizes key fragmentation and synchronization during voice calls, improving quantum key utilization efficiency and call security. More specifically, compared with existing technologies, this application tightly integrates voice content with key management, improving the dynamism and randomness of key updates compared to traditional fixed-period key update methods. A differentiated encryption strategy is adopted, using newly generated quantum keys for encryption during active voice periods and reusing previous keys during silent periods to save resources. Key updates are triggered at semantic boundaries to ensure security and synchronization. Before transmission, the key pointer is embedded in the redundant field of the voice frame, enabling key synchronization transmission without additional bandwidth. Furthermore, this application helps improve the security of voice call systems under quantum computing threats by improving the quantum key update mechanism in existing voice calls, thereby achieving more efficient and secure key scheduling and encrypted transmission, providing key technical support for building quantum-resistant voice communication systems.

[0056] For example, this application is applicable to voice call encryption scenarios, supporting two-party and multi-party voice calls. Optionally, the application scenarios of this application may also include: VoIP and 5G / 6G voice calls, smart home voice control, and video conferencing. It should be noted that the beneficial effects or technical problems solved by the embodiments of this application are not limited to this one, but may also be other implicit or related problems, which can be referred to in the description of the embodiments below.

[0057] Before introducing the specific embodiments of this application, the technical terms involved in this application will be explained:

[0058] Quantum Key Distribution (QKD) is a key generation and distribution technology based on the principles of quantum mechanics. Its core mechanism involves two communicating parties negotiating and generating a shared quantum key over a public channel by sending and measuring quantum states. QKD theoretically possesses unconditionally secure key distribution capabilities.

[0059] Energy value (speech signal energy characteristic): The energy value (Signal Energy, SE) is a key feature for measuring the signal strength of a speech frame. It is usually obtained by calculating the sum of squares or root mean square (RMS) value of the amplitude of the speech signal's time-domain waveform. Its physical meaning is the amount of energy carried by the speech signal per unit time.

[0060] Active frames: When the energy value of a speech frame is higher than a preset threshold, it indicates that the frame contains valid speech content (such as vowels and voiced consonants), and is called an "active frame", such as a clear pronunciation segment when speaking.

[0061] Silent frames: When the energy value is below the threshold, it indicates that the frame is background noise or speech gap, and is called a "silent frame", such as a pause in speech or breathing sound.

[0062] Zero-Crossing Rate (Speech Signal Temporal Feature): The zero-crossing rate (ZCR) refers to the number of times a speech signal crosses the zero level per unit time, and is an important parameter characterizing the temporal variation of a speech signal. It is calculated by counting the number of times a sample value in a frame of speech signal changes from positive to negative or vice versa. The zero-crossing rate plays a crucial role in speech abrupt change point identification: when a speech signal transitions from silence to articulation (such as the onset of a consonant) or changes in articulation mode, the temporal waveform of the signal changes drastically, and the zero-crossing rate increases significantly; while in vowel or smooth speech segments, the signal waveform is relatively flat, and the zero-crossing rate is low. Therefore, by monitoring abrupt changes in the zero-crossing rate, instantaneous change points such as the onset of consonants and syllable boundaries in speech can be accurately located, providing temporal basis for speech feature analysis and encryption strategy adjustment.

[0063] Fundamental Frequency (F0): The fundamental frequency (F0) is the lowest frequency component in a speech signal, corresponding to the basic frequency of vocal cord vibration, and is measured in Hertz (Hz). For human speech, the fundamental frequency determines the pitch: males typically have a fundamental frequency of 80-180Hz, females 160-250Hz, and children have higher frequencies. In semantic segment boundary detection, the variation pattern of the fundamental frequency is valuable: when there are pauses, semantic shifts, or emotional changes in the speech content, the vocal cord vibration frequency undergoes significant abrupt changes (such as a drop in pitch or a rise in pitch before a pause), resulting in a significant inflection point in the fundamental frequency curve; while within continuous semantic segments, the fundamental frequency changes relatively smoothly.

[0064] Key pointer: The key pointer is used to uniquely identify the currently used quantum key segment, and may include key parameters such as the key start position index and the key length.

[0065] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0066] In one embodiment, such as Figure 1 As shown, an encrypted transmission method is provided. Taking the application of this method to a voice transmitter as an example, it is understood that this method can also be applied to other terminals, servers, and systems including terminals and servers, and is implemented through the interaction between the terminal and the server. For example, when applied to a system including a voice transmitter and a voice receiver, it is implemented through the interaction between the voice receiver and the voice transmitter. The method includes the following steps.

[0067] Step 200: Perform frame segmentation on the acquired speech signal to obtain each speech frame.

[0068] The speech signal can be a speech stream. Upon receiving the speech signal, the speech transmitter can perform frame segmentation. Taking a speech stream as an example, the speech stream can be segmented according to a fixed time length (e.g., 20ms / frame), dividing the speech stream into multiple smaller segments, each segment being called a speech frame, for subsequent analysis.

[0069] Optionally, the voice transmitter divides the voice signal into frames according to a specific time length, resulting in individual voice frames. Dividing the voice signal into frames facilitates subsequent differentiated encryption of different parts of the voice signal, i.e., each voice frame, thereby improving encryption security. For example, a frame length of 20ms can be used for segmentation.

[0070] Step 300: Determine the speech state of each speech frame; the speech state includes silent frames and active frames.

[0071] In this context, an active frame refers to a speech frame with an energy value higher than a preset threshold, indicating that the speech frame contains valid speech content; a silent frame refers to a speech frame with an energy value lower than the preset threshold, indicating that the speech frame is background noise or a speech gap. It should be noted that the speech state refers to the state of the speech signal in a corresponding time frame. This embodiment of the application can identify the speech state and mark it as an active frame or a silent frame. Furthermore, the speech state in this embodiment of the application may include, but is not limited to, silent frames and active frames, and may also include semantic boundaries.

[0072] In this embodiment, the voice transmitter identifies the voice state, marks silent frames and active frames, and then dynamically schedules the key according to the voice state to achieve efficient use of the key.

[0073] Step 400: Encrypt the current active frame using a first encryption method and encrypt the current silent frame using a second encryption method to obtain an encrypted voice frame; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame.

[0074] Specifically, the key fragment can be a quantum key fragment. The voice transmitter can encrypt the current active frame using a first encryption method and the current silent frame using a second encryption method, thereby improving security by employing differentiated encryption strategies.

[0075] For example, the first encryption method may refer to encrypting the current active frame with the currently acquired key fragment when the current voice frame is determined to be an active frame, that is, when the current voice frame has a high energy value; the second encryption method may refer to encrypting the current silent frame with the same key fragment used to encrypt the active frame when the current voice frame is determined to be a silent frame, that is, when the current voice frame has a low energy value.

[0076] By employing a differentiated encryption strategy—that is, using newly generated key fragments for encryption during active voice periods and reusing the previous key during quiet periods—security is ensured while also saving key resources.

[0077] Step 500: Encapsulate the encrypted voice frame into a data packet and output it.

[0078] Specifically, the voice transmitter encapsulates the encrypted voice frame into a data packet and outputs the data packet to the voice receiver. In this embodiment, the data packet can also be called a transmission data packet. It should be noted that encapsulation refers to the process by which the voice transmitter encapsulates the encrypted voice frame into a transmission data packet. It is understood that the encapsulation can take other forms, and is not limited to the forms mentioned in the above embodiments, as long as it can achieve the function of obtaining the data packet.

[0079] Optionally, the voice transmitter encapsulates the current encrypted voice frame and the identification information or location information (such as a key pointer) of the corresponding key segment into a transmission data packet and sends it to the voice receiver (hereinafter referred to as the receiver) so that the receiver can decrypt it.

[0080] In the above-mentioned encrypted transmission method, differentiated encrypted transmission methods are adopted for voice frames with different voice characteristics. In particular, for high-energy voice frames, a key fragment of appropriate length is extracted from the current position to ensure the transmission security of critical voice information. Furthermore, the encrypted voice frame is encapsulated into a data packet for transmission, thereby effectively improving the security of the encrypted transmission method.

[0081] In one exemplary embodiment, determining the frame type of each speech frame may include the following steps:

[0082] Extract the speech signal energy features of each speech frame; the speech signal energy features include energy values;

[0083] Voice frames with energy values ​​higher than a preset threshold are identified as active frames, while voice frames with energy values ​​lower than a preset threshold are identified as silent frames.

[0084] Specifically, after acquiring the speech signal and performing frame segmentation at the speech transmitting end, key speech features are extracted from each segmented speech frame. For example, key speech features may include: speech signal energy features, speech signal time-domain features, and speech signal frequency-domain features. The speech signal energy features include at least an energy value. Based on the energy value of the speech frame, the speech state is identified, and the speech frame is marked as an active frame or a silent frame. The specific criteria are as follows: when the energy value of a speech frame is higher than a preset threshold, it indicates that the frame contains valid speech content (such as vowels and voiced consonants), and is called an "active frame," such as a clear pronunciation segment during speech; when the energy value is lower than the threshold, it indicates that the frame is background noise or a speech gap, and is called a "silent frame," such as a pause or breathing sound during speech. By detecting the speech signal energy features of the speech frame and determining and marking the speech frame type based on the energy value, the effect of accurately distinguishing the energy type of the speech frame can be achieved.

[0085] It is understood that in some embodiments, before performing frame segmentation on the acquired speech signal to obtain each speech frame, the following steps are included:

[0086] In response to the synchronous acquisition of the quantum key stream with the voice receiver, the key pointer is initialized to the starting position; wherein, the quantum key stream includes key segments; the key pointer is used to identify the position of the key segment in the quantum key stream.

[0087] Specifically, before initiating an encrypted call, key pre-allocation is performed. Both parties synchronously acquire a consistent quantum key stream using OKD or other quantum key generation mechanisms and initialize the key pointer to the starting position. Then, the session is established. Once the voice connection is established, the encrypted communication preparation phase begins. By ensuring that the voice sender and receiver synchronously acquire a consistent quantum key stream and initialize the key pointer to the starting position before the call, conditions are created for effectively implementing key fragmentation and synchronization during the subsequent voice call, improving quantum key utilization, and enhancing call security.

[0088] In an exemplary embodiment, the first encryption method includes: extracting a key segment from the quantum key stream based on the position pointed to by the current key pointer, and encrypting the current active frame using a symmetric encryption algorithm;

[0089] The second encryption method includes: reusing the key fragment used in the first encryption method to encrypt the previous active frame, and encrypting the current silent frame using a symmetric encryption algorithm.

[0090] Specifically, in this embodiment, the key pointer is used to uniquely identify the currently used quantum key segment, and may include key parameters such as the key starting position index and the key length.

[0091] The corresponding encryption strategy is executed based on the state of the voice frame: if it is an active frame, the first encryption method is used, which involves extracting a quantum key fragment of a specified length (e.g., 128 bits) from the current key pointer position and encrypting the voice frame using a symmetric encryption algorithm such as SM4. The key fragment used in the previous active frame is then reused for encryption. By employing different encryption methods for active and silent frames, real-time voice feature analysis is integrated with dynamic quantum key fragmentation, thus tightly combining voice content with key management. Compared to the traditional fixed-period key update method, this improves the dynamism and randomness of key updates. Furthermore, by adopting a differentiated encryption strategy—using newly generated key fragments for encryption during active voice periods and reusing the previous key during silent periods—key resources are saved while ensuring security.

[0092] Furthermore, in some embodiments of this application, the voice state also includes semantic boundaries; the method further includes:

[0093] Extract the frequency domain features of the speech signal from each speech frame; the frequency domain features of the speech signal include the fundamental frequency;

[0094] When a semantic boundary in a speech frame is detected based on changes in the base frequency, the current key pointer is updated to obtain the updated key pointer.

[0095] Specifically, this application embodiment, by continuously monitoring the fundamental frequency mutation rate (e.g., a fundamental frequency change exceeding 20% ​​between adjacent frames), can accurately identify semantic boundaries (e.g., pause positions) in speech, thereby triggering key updates and other encryption strategy adjustments to achieve dynamic security protection of "one semantic segment, one key". It should be noted that the frequency domain features of the speech signal in this application can also be other types of features, and this application is not limited in this regard.

[0096] Optionally, when a semantic boundary is detected, the key pointer is moved to the next available bit, the key pointer is updated, and a new key fragment is extracted. By triggering a key update at a semantic boundary, security and synchronization during encrypted transmission are ensured.

[0097] In one exemplary embodiment, if no semantic boundary is detected in the voice frame, the current key pointer is embedded into the redundant field of the corresponding encrypted voice frame.

[0098] Specifically, when no semantic boundary is detected in the voice frame, there is no need to update the key pointer. The current key pointer can be embedded in the voice frame redundancy field, thus completing the synchronous transmission of the key without additional bandwidth and saving communication resources.

[0099] In an exemplary embodiment, if a semantic boundary is detected in a speech frame, the current key pointer and the updated key pointer are embedded into the redundant field of the corresponding encrypted speech frame.

[0100] Specifically, when a semantic boundary is detected in a speech frame, the key pointer needs to be updated. Then, by embedding the key pointer into the speech frame redundancy field, the key can be transmitted synchronously without additional bandwidth, thus saving communication resources.

[0101] In one embodiment, such as Figure 2 As shown, an encrypted transmission method is provided. This embodiment illustrates the application of this method to a voice receiving end. It is understood that this method can also be applied to other terminals, servers, and systems including both terminals and servers, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0102] Step 600: Receive data packets from the voice sender.

[0103] In this embodiment, the data packet may be the data packet output by the aforementioned voice transmitter, which will not be described in detail here.

[0104] Step 700: Decrypt the encrypted voice frame in the data packet using the obtained key fragment to obtain the decrypted voice frame.

[0105] Optionally, the voice transmitter obtains a key fragment corresponding to the encrypted voice frame in the received data packet, and uses the key fragment to decrypt the encrypted voice frame in the data packet, thereby obtaining a decrypted voice frame. By receiving data packets from the voice transmitter and obtaining the key fragment to decrypt the encrypted voice frame in the data packet, dynamic decryption of the voice frame is achieved, realizing the synchronization of encryption and decryption.

[0106] In the above-mentioned encrypted transmission method applied to the voice receiver, the encrypted voice frames in the data packet are decrypted by receiving data packets from the voice sender and obtaining key fragments, thereby realizing the dynamic decryption of the voice frames and achieving the synchronization of encryption and decryption.

[0107] Furthermore, in some embodiments of this application, the step of obtaining a key fragment includes:

[0108] The data packet is parsed to extract the encrypted voice frame and the key pointer corresponding to the encrypted voice frame; the key pointer corresponding to the encrypted voice frame includes at least the current key pointer.

[0109] The key pointer uniquely identifies the currently used quantum key segment and may include key parameters such as the key start position index and the key length. Specifically, the voice receiver parses the received data packets to extract the key pointer corresponding to the current encrypted voice frame, the possible new key pointer, and the encrypted voice frame itself. By separating and recovering the key pointer information from the data packets, it is used for dynamic decryption and key updates of subsequent voice frames, thereby achieving synchronization of the encryption and decryption process and improving transmission security.

[0110] Based on the key pointer, the corresponding key segment is extracted from the local quantum key stream. Based on the key segment, the encrypted voice frame is decrypted using a symmetric decryption algorithm to obtain the decrypted voice frame.

[0111] Specifically, in this embodiment, based on the key pointer corresponding to the current encrypted voice frame, the corresponding key segment is located using the locally pre-stored quantum key stream. The voice frame is then decrypted using a symmetric decryption algorithm consistent with that used at the transmitting end, restoring the voice signal for playback. By locating the key segment based on the key pointer in the data packet and simultaneously using a symmetric decryption algorithm corresponding to the symmetric encryption algorithm at the voice transmitting end to decrypt the encrypted voice frame, the encryption and decryption processes are effectively synchronized, ensuring the security and synchronization of the encrypted voice transmission process.

[0112] Furthermore, in some embodiments of this application, the key pointer corresponding to the encrypted voice frame also includes an updated key pointer; the method further includes:

[0113] Update the local key pointer based on the updated key pointer.

[0114] Specifically, if the voice receiver parses a new key pointer during the data packet parsing process, the receiver will update the local key pointer to ensure that the new key is used in subsequent calls, thereby ensuring the synchronization of encryption and decryption and improving the security of the transmission process.

[0115] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0116] Based on the same inventive concept, this application also provides an encrypted transmission device for implementing the encrypted transmission method for a voice transmitter as described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the encrypted transmission device for a voice transmitter provided below can be found in the limitations of the encrypted transmission method for a voice transmitter described above, and will not be repeated here.

[0117] In one exemplary embodiment, such as Figure 3 As shown, an encrypted transmission device 110 is provided, applied to a voice transmission end, including: a segmentation unit 111, a status determination unit 112, an encryption unit 113, and a data output unit 114, wherein:

[0118] The segmentation unit 111 is used to segment the acquired speech signal into frames to obtain each speech frame.

[0119] The state determination unit 112 is used to determine the speech state of each speech frame; the speech state includes silent frames and active frames.

[0120] The encryption unit 113 is used to encrypt the current active frame using a first encryption method and to encrypt the current silent frame using a second encryption method to obtain an encrypted voice frame; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame.

[0121] The data output unit 114 is used to encapsulate encrypted voice frames into data packets and output them.

[0122] In one embodiment, the state determination unit 112 is used to extract the speech signal energy features of each speech frame; the speech signal energy features include energy values; speech frames with energy values ​​higher than a preset threshold are determined as active frames, and speech frames with energy values ​​lower than the preset threshold are determined as silent frames.

[0123] In one embodiment, the state determination unit 112 is configured to initialize the key pointer to the starting position in response to synchronously acquiring the quantum key stream with the voice receiver; wherein the quantum key stream includes key segments; and the key pointer is used to identify the position of the key segments in the quantum key stream.

[0124] In one embodiment, the encryption unit 113 is configured to extract a key segment from the quantum key stream based on the position pointed to by the current key pointer, and encrypt the current active frame using a symmetric encryption algorithm; it is also configured to reuse the key segment used by the first encryption method that encrypts the previous active frame, and encrypt the current silent frame using a symmetric encryption algorithm.

[0125] In one embodiment, the state determination unit 112 is used to extract the frequency domain features of the speech signal of each speech frame; the frequency domain features of the speech signal include the fundamental frequency; when a semantic boundary in the speech frame is detected based on the change of the fundamental frequency, the current key pointer is updated to obtain the updated key pointer.

[0126] In one embodiment, the data output unit 114 is configured to embed the current key pointer into the redundant field of the corresponding encrypted voice frame if no semantic boundary is detected in the voice frame. In another embodiment, the data output unit 114 is configured to embed both the current key pointer and the updated key pointer into the redundant field of the corresponding encrypted voice frame if a semantic boundary is detected in the voice frame.

[0127] Based on the same inventive concept, this application also provides an encrypted transmission device for implementing the encrypted transmission method for a voice receiver as described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations of one or more embodiments of the encrypted transmission device for a voice receiver provided below can be found in the limitations of the encrypted transmission method for a voice transmitter described above, and will not be repeated here.

[0128] In one exemplary embodiment, such as Figure 4 As shown, an encrypted transmission device 120 is provided, applied to a voice receiving end, including: a data receiving unit 121 and a decryption unit 122, wherein:

[0129] The data receiving unit 121 is used to receive data packets from the voice transmitting end;

[0130] The decryption unit 122 is used to decrypt the encrypted voice frame in the data packet using the acquired key fragment to obtain the decrypted voice frame.

[0131] In one embodiment, the decryption unit 122 is used to parse the data packet, extract the encrypted voice frame, and the key pointer corresponding to the encrypted voice frame; the key pointer corresponding to the encrypted voice frame includes at least the current key pointer; according to the key pointer, the corresponding key segment is extracted from the local quantum key stream, and according to the key segment, the encrypted voice frame is decrypted using a symmetric decryption algorithm to obtain the decrypted voice frame.

[0132] In one embodiment, the decryption unit 122 is used to update the local key pointer according to the updated key pointer.

[0133] Each module in the aforementioned encrypted transmission device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each module.

[0134] Based on the same inventive concept, this application also provides an encrypted transmission system for implementing the encrypted transmission method described above. The solution provided by this system is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more encrypted transmission system embodiments provided below can be found in the limitations of the encrypted transmission method described above, and will not be repeated here.

[0135] In one exemplary embodiment, an encrypted transmission system is provided, comprising: a voice transmitter and a voice receiver, wherein:

[0136] The voice transmitter is used to segment the acquired voice signal into frames to obtain each voice frame; determine the voice state of each voice frame; the voice state includes silent frames and active frames; encrypt the current active frame using a first encryption method and encrypt the current silent frame using a second encryption method to obtain encrypted voice frames; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used in the first encryption method that encrypted the previous active frame; and encapsulate the encrypted voice frames into data packets and output them.

[0137] In one embodiment, the voice transmitting end is used to extract the voice signal energy features of each voice frame; the voice signal energy features include energy values; voice frames with energy values ​​higher than a preset threshold are determined as active frames, and voice frames with energy values ​​lower than the preset threshold are determined as silent frames.

[0138] In one embodiment, the voice transmitter is configured to initialize the key pointer to the starting position in response to synchronously acquiring the quantum key stream with the voice receiver; wherein the quantum key stream includes key segments; and the key pointer is used to identify the position of the key segments in the quantum key stream.

[0139] In one embodiment, the voice transmitter is configured to extract a key segment from the quantum key stream based on the position pointed to by the current key pointer, encrypt the current active frame using a symmetric encryption algorithm, and reuse the key segment used in the first encryption method that encrypted the previous active frame to encrypt the current silent frame using a symmetric encryption algorithm.

[0140] In one embodiment, the voice transmitting end is used to extract the frequency domain features of the voice signal of each voice frame; the frequency domain features of the voice signal include the fundamental frequency; when the semantic boundary in the voice frame is detected based on the change of the fundamental frequency, the current key pointer is updated to obtain the updated key pointer.

[0141] In one embodiment, the voice transmitter is configured to embed the current key pointer into the redundant field of the corresponding encrypted voice frame if no semantic boundary is detected in the voice frame.

[0142] In one embodiment, the voice transmitter is configured to embed the current key pointer and the updated key pointer into the redundant field of the corresponding encrypted voice frame if a semantic boundary is detected in the voice frame.

[0143] The voice receiver is used to receive data packets from the voice sender; it decrypts the encrypted voice frames in the data packets using the obtained key fragment to obtain the decrypted voice frames.

[0144] In one embodiment, the voice transmitter is configured to parse data packets, extract encrypted voice frames, and the key pointer corresponding to the encrypted voice frames; the key pointer corresponding to the encrypted voice frames includes at least the current key pointer; extract the corresponding key fragment from the local quantum key stream according to the key pointer; and decrypt the encrypted voice frames using a symmetric decryption algorithm based on the key fragment to obtain decrypted voice frames.

[0145] In one embodiment, the voice transmitter is configured to update the local key pointer based on the updated key pointer.

[0146] Optionally, refer to Figure 5 In one embodiment, the voice transmitting end of the encrypted transmission system includes: a real-time voice feature analysis module, a quantum key dynamic sharding and encryption module, and a key-voice collaborative transmission module; the voice receiving end includes: a receiving end dynamic decryption module.

[0147] The real-time speech feature analysis module performs frame segmentation on the speech stream (e.g., 20ms / frame) and simultaneously extracts key speech features, such as:

[0148] Energy value: Distinguishes between active frames (high energy) and silent frames (low energy).

[0149] Zero crossover rate: Identifying abrupt changes in speech (such as the starting position of a consonant);

[0150] Fundamental frequency: Used to detect semantic segment boundaries in speech (such as sentence pauses).

[0151] Then, the speech is labeled with various features, which are divided into: active frames, silent frames, semantic boundaries, etc.

[0152] The quantum key dynamic fragmentation and encryption module implements dynamic fragmentation and encryption strategies for quantum keys based on real-time voice characteristics:

[0153] Active frame encryption: For high-energy voice frames, a key segment of appropriate length (e.g., 128 bits) is extracted from the current key pointer position and encrypted using a symmetric encryption algorithm (e.g., SM4) to ensure the security of critical voice information transmission.

[0154] Silent Frame Encryption: When a low-energy voice frame is detected, "key sleep mode" is enabled, reusing the previous key fragment for encryption to avoid consuming new quantum keys and save resources.

[0155] Semantic boundary triggering: When a semantic boundary is detected, the key pointer is moved to the next available bit and a new key fragment is extracted. At the same time, the key pointer is embedded in the redundant bit of the voice frame to ensure that the receiving end updates the decryption key synchronously.

[0156] The key-voice collaborative transmission module embeds the key pointer into the redundancy bits of the voice frame. Specifically, this is achieved by inserting it into the extended header field of the RTP protocol, transmitting key management signaling without increasing bandwidth. The receiving end dynamic decryption module separates and recovers the key pointer information from the voice data packet for dynamic decryption and key updates in subsequent voice frames, achieving synchronization of the encryption and decryption process.

[0157] Optionally, in one embodiment, reference is made to Figure 6 When implementing the encrypted transmission method in the encrypted transmission system, the following steps are performed: Key pre-allocation: Both parties in the call synchronously obtain a consistent quantum key stream through QKD or other quantum key generation mechanisms, and initialize the key pointer to the starting position.

[0158] Session establishment: Both parties complete the establishment of a voice call connection and enter the encrypted communication preparation stage.

[0159] Speech Feature Analysis and State Labeling: The speech transmitter acquires speech signals and performs frame segmentation (e.g., 20ms / frame). Key speech features, such as energy value, zero crossover rate, and fundamental frequency, are extracted from each frame. Based on these features, the speech state is identified, and active frames, silent frames, and semantic boundaries are labeled.

[0160] Dynamic key fragmentation and encryption processing: The corresponding encryption strategy is executed according to the state of the voice frame: If it is an active frame, a quantum key fragment of a specified length (such as 128 bits) is extracted from the current key pointer position and the voice frame is encrypted using a symmetric encryption algorithm (such as SM4); if it is a silent frame, the key fragment used by the previous active frame is reused for encryption; if a semantic boundary is detected, the key pointer is moved forward and the key pointer is updated.

[0161] Key pointer embedding and data transmission: The voice transmitter embeds the key pointer corresponding to the current encrypted voice frame, as well as the new key pointer that may be updated (which may be empty), into the redundant fields of the voice frame (such as the RTP extension header), and encapsulates it with the encrypted voice frame into a data packet before sending it to the receiver.

[0162] Data decryption and key update: The receiving end parses the received data packets, extracting the key pointer corresponding to the current encrypted voice frame, the possible new key pointer, and the encrypted voice frame itself. Then, based on the key pointer corresponding to the current encrypted voice frame, it locates the corresponding key segment using a locally stored quantum key stream and decrypts the voice frame using the same symmetric decryption algorithm as the sending end, restoring the voice signal for playback. If a new key pointer is found, the receiving end updates its local key pointer to ensure that the new key is used in subsequent calls.

[0163] End of call: The call ends when either party initiates a request to terminate the call.

[0164] In one exemplary embodiment, an electronic device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, this electronic device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an encrypted transmission method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.

[0165] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0166] In one exemplary embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0167] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method embodiments.

[0168] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the above-described method embodiments.

[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0171] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for encrypted transmission, characterized in that, Applied to a voice transmitting end, the method includes: The acquired speech signal is segmented into frames to obtain individual speech frames; The speech state of each of the aforementioned speech frames is determined; the speech state includes silent frames and active frames. The current active frame is encrypted using a first encryption method, and the current silent frame is encrypted using a second encryption method to obtain an encrypted voice frame; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame; The encrypted voice frames are encapsulated into data packets and output.

2. The encrypted transmission method according to claim 1, characterized in that, Determining the frame type of each of the aforementioned speech frames includes: Extract the speech signal energy features of each of the speech frames; the speech signal energy features include energy values; Voice frames with energy values ​​higher than a preset threshold are identified as active frames, while voice frames with energy values ​​lower than a preset threshold are identified as silent frames.

3. The encrypted transmission method according to claim 1, characterized in that, Before performing frame segmentation on the acquired speech signal to obtain each speech frame, the method further includes: In response to synchronously acquiring the quantum key stream with the voice receiver, the key pointer is initialized to the starting position; wherein, the quantum key stream includes the key segment; the key pointer is used to identify the position of the key segment in the quantum key stream.

4. The method according to claim 3, characterized in that, The first encryption method includes: extracting the key segment from the quantum key stream based on the position pointed to by the current key pointer, and encrypting the current active frame using a symmetric encryption algorithm; The second encryption method includes: reusing the key fragment used by the first encryption method to encrypt the previous active frame, and encrypting the current silent frame using a symmetric encryption algorithm.

5. The method according to claim 3, characterized in that, The speech state also includes semantic boundaries; the method further includes: Extract the frequency domain features of the speech signal from each of the speech frames; the frequency domain features of the speech signal include the fundamental frequency; When a semantic boundary in the speech frame is detected based on the change in the base frequency, the current key pointer is updated to obtain the updated key pointer.

6. The method according to claim 5, characterized in that, The method further includes: If no semantic boundary is detected in the voice frame, the current key pointer is embedded into the redundant field of the corresponding encrypted voice frame.

7. The method according to claim 5, characterized in that, The method further includes: If a semantic boundary is detected in the voice frame, the current key pointer and the updated key pointer are embedded into the redundant field of the corresponding encrypted voice frame.

8. A method for encrypted transmission, characterized in that, Applied to a voice receiving end, the method includes: Receive data packets from the voice sender; The encrypted voice frames in the data packet are decrypted using the obtained key fragment to obtain the decrypted voice frames.

9. The encrypted transmission method according to claim 8, characterized in that, The steps to obtain the key fragment include: The data packet is parsed to extract the encrypted voice frame and the key pointer corresponding to the encrypted voice frame; the key pointer corresponding to the encrypted voice frame includes at least the current key pointer. The corresponding key segment is extracted from the local quantum key stream according to the key pointer. The encrypted voice frame is then decrypted using a symmetric decryption algorithm based on the key segment to obtain the decrypted voice frame.

10. The encrypted transmission method according to claim 9, characterized in that, The key pointer corresponding to the encrypted voice frame also includes an updated key pointer; the method further includes: Update the local key pointer based on the updated key pointer.

11. An encrypted transmission device, characterized in that, The device, applied to a voice transmission terminal, includes: The segmentation unit is used to segment the acquired speech signal into frames to obtain each speech frame; A state determination unit is used to determine the voice state of each of the voice frames; the voice state includes silent frames and active frames. An encryption unit is used to encrypt the current active frame using a first encryption method and to encrypt the current silent frame using a second encryption method to obtain an encrypted voice frame; wherein, the first encryption method includes encrypting the current active frame using an acquired key fragment; the second encryption method includes reusing the key fragment used by the first encryption method to encrypt the previous active frame; The data output unit is used to encapsulate the encrypted voice frame into a data packet and output it.

12. An encrypted transmission device, characterized in that, The device, applied to a voice receiving end, includes: A data receiving unit is used to receive data packets from the voice transmitter. The decryption unit is used to decrypt the encrypted voice frame in the data packet using the acquired key fragment to obtain the decrypted voice frame.

13. An encrypted transmission system, characterized in that, Includes both the voice transmitter and the voice receiver; The voice transmitting end is used to implement the steps of the method according to any one of claims 1 to 7; The voice receiver is used to implement the steps of the method according to any one of claims 8 to 10.

14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Voice QoS optimizing and playing method, device and equipment facing end-to-end information source encryption and medium

    CN121690858A