A converged communication encryption method and system for multi-mode handheld terminals
By using a converged communication encryption method for multi-mode handheld terminals, voice frames are encrypted separately using narrowband and broadband links, generating hybrid ciphertext and transmitting it using a joint root key. This solves the problems of difficulty in encrypting high-definition voice and easy decryption of single channels in existing technologies, and achieves efficient and secure voice communication.
Patent Information
- Application Number
- CN202511793717.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-12-02
AI Technical Summary
Existing voice communication encryption methods use a single encryption algorithm, which leads to bandwidth bottlenecks, making it unable to support high-definition voice encryption. This results in a sharp drop in voice quality, and single-channel encryption is easily decrypted maliciously, reducing its practicality and stability.
A converged communication encryption method using multi-mode handheld terminals is adopted. Voice frames are encrypted through narrowband and broadband links respectively to generate hybrid ciphertext. A joint root key is used for transmission and decryption. The original voice stream is reconstructed by combining the voice fusion model.
It achieves deep encryption of high-definition voice, improves encryption stability and efficiency, avoids the risks of single-channel encryption, and ensures the integrity and security of voice.
Smart Images

Figure CN121240073B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication encryption technology, and in particular to a converged communication encryption method and system for multi-mode handheld terminals. Background Technology
[0002] In today's rapidly developing digital age, data security and privacy protection have become key concerns across various industries. With the increasing frequency and complexity of data interactions between enterprises and users, the application of data encryption technology is becoming more widespread. However, with the surge in data volume, traditional data encryption methods are gradually revealing serious problems in terms of computational efficiency. Especially in practical business applications, the encryption computational load generated by the real-time interaction of large amounts of data has become a bottleneck affecting data transmission efficiency and system performance. Existing voice communication encryption typically uses a single encryption algorithm to encrypt the communication voice content, which suffers from bandwidth bottlenecks and cannot support high-definition voice encryption. This leads to a precipitous drop in voice quality, and single-channel encryption is easily decrypted maliciously, resulting in data leakage and losses, thus reducing practicality and stability. Summary of the Invention
[0003] To address the problems mentioned above, this invention provides a converged communication encryption method and system for multi-mode handheld terminals to solve the problems mentioned in the background art. Existing voice communication encryption typically uses a single encryption algorithm to encrypt the voice content, which has a bandwidth bottleneck, cannot support high-definition voice encryption, resulting in a sharp drop in voice quality, and single-channel encryption is easily decrypted maliciously, leading to data leakage and losses, reducing practicality and stability.
[0004] A converged communication encryption method for a multi-mode handheld terminal includes the following steps:
[0005] Obtain voice service requests initiated by multi-mode handheld terminals, parse the voice service requests to determine the service sensitivity level;
[0006] Whether to activate the dual-link joint encryption strategy is determined based on the business sensitivity level. If so, the voice frame is encrypted through narrowband and broadband links respectively to obtain the mixed ciphertext.
[0007] A joint root key is generated based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal. The core frame ciphertext block in the hybrid ciphertext is transmitted through the narrowband channel, and the enhanced frame ciphertext block in the hybrid ciphertext is transmitted through the wideband channel.
[0008] At the receiving end, the core frame ciphertext block and the enhanced frame ciphertext block are decrypted and reassembled using the joint root key to obtain the original voice stream.
[0009] Preferably, the step of acquiring voice service requests initiated by multi-mode handheld terminals and parsing the voice service requests to determine the service sensitivity level includes:
[0010] The terminal's multi-mode communication management module listens for user-triggered voice communication requests and determines whether the user has manually triggered the emergency encryption button based on the voice communication request. If so, it is determined to be a highly sensitive service; otherwise, it is initially determined to be a low-sensitivity service.
[0011] Collect voice content emitted by multi-mode handheld terminals, analyze whether the voice content contains preset keywords, if so, determine it as a highly sensitive service, if not, determine it as a low-sensitivity service.
[0012] The GPS positioning parameters of the multi-mode handheld terminal are detected. Based on the GPS positioning parameters, it is determined whether the multi-mode handheld terminal is in a sensitive area. If it is, it is judged as a highly sensitive service. If not, it is judged as a low-sensitivity service after three checks.
[0013] The sensitivity level of a voice service request is determined based on the number of times the high-sensitivity service judgment mechanism for voice communication requests is triggered.
[0014] Preferably, the step of determining whether to activate the dual-link joint encryption strategy based on the service sensitivity level, and if so, encrypting the voice frame through both narrowband and broadband links to obtain hybrid ciphertext, includes:
[0015] Determine whether the voice communication request is a high-risk service based on the business sensitivity level. If it is, determine whether a dual-link joint encryption strategy needs to be activated. If not, determine whether a dual-link joint encryption strategy does not need to be activated and determine whether a dynamic preferred link encryption strategy should be activated.
[0016] The core speech frames in the speech content are extracted by frame extraction technology, and the original speech frames are encoded with high precision to generate speech enhancement frames.
[0017] The first ciphertext is obtained by encrypting the voice core frame with AES-256 through a narrowband link, and the second ciphertext is obtained by encrypting the voice enhancement frame with the national standard SM4 through a broadband link.
[0018] The first ciphertext and the second ciphertext are spatiotemporally interleaved and encapsulated to generate a mixed ciphertext.
[0019] Preferably, the step of generating a joint root key based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal, transmitting the core frame ciphertext block in the hybrid ciphertext through a narrowband channel, and transmitting the enhanced frame ciphertext block in the hybrid ciphertext through a wideband channel includes:
[0020] The narrowband timeslot number and broadband sequence number of the multi-mode handheld terminal are fused and concatenated into N bytes. The N bytes are then XORed with the temporary random number broadcast by the base station, and a joint root key is generated based on the result.
[0021] The core frame ciphertext block and the enhanced frame ciphertext block in the hybrid ciphertext are divided into blocks using data block segmentation technology;
[0022] Determine the transmission requirements for the ciphertext block, determine the expected transmission efficiency based on the transmission requirements, determine the channel evaluation index based on the expected transmission efficiency, and evaluate the channel quality of narrowband and wideband channels based on the channel evaluation index.
[0023] Based on the evaluation results, the channel parameters of the narrowband and wideband channels are adjusted. The core frame ciphertext block is transmitted through the adjusted narrowband channel, and the enhanced frame ciphertext block is transmitted through the adjusted wideband channel.
[0024] Preferably, the step of decrypting and reconstructing the core frame ciphertext block and the enhanced frame ciphertext block using a joint root key at the receiving end to obtain the original voice stream includes:
[0025] At the receiving end, the core frame ciphertext block and the enhanced frame ciphertext block are decrypted using the joint root key to obtain the decrypted data segment and determine the data tag of each decrypted data segment.
[0026] The core frame decryption blocks and the enhancement frame decryption blocks are sorted according to the data tags, and the sorting results are obtained. The core frame speech sequence and the enhancement frame speech sequence are obtained according to the sorting results.
[0027] The core frame speech sequence and the enhanced frame speech sequence are input into the preset speech fusion model. The speech fusion model extracts the basic speech skeleton from the core frame speech sequence and extracts high-frequency detail speech and background sound from the enhanced frame speech sequence.
[0028] The original speech stream is generated based on the basic speech skeleton, high-frequency detail speech, and background noise.
[0029] A converged communication encryption system for multi-mode handheld terminals, the system comprising:
[0030] The business sensitivity level determination module is used to obtain voice service requests initiated by multi-mode handheld terminals, parse the voice service requests, and determine the business sensitivity level.
[0031] The voice frame encryption module is used to determine whether to activate the dual-link joint encryption strategy based on the business sensitivity level. If so, the voice frame is encrypted through the narrowband link and the broadband link respectively to obtain the mixed ciphertext.
[0032] The ciphertext block transmission module is used to generate a joint root key based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal, transmit the core frame ciphertext block in the hybrid ciphertext through the narrowband channel, and transmit the enhanced frame ciphertext block in the hybrid ciphertext through the wideband channel.
[0033] The ciphertext block decryption module is used at the receiving end to decrypt and reassemble the core frame ciphertext block and the enhanced frame ciphertext block using a joint root key to obtain the original voice stream.
[0034] Preferably, the business sensitivity level determination module includes:
[0035] The business sensitivity attribute determination submodule is used to listen to the user-triggered voice communication requests through the terminal's multi-mode communication management module, and determine whether the user has manually triggered the emergency encryption button based on the voice communication request. If so, it is determined to be a highly sensitive business; otherwise, it is initially determined to be a low-sensitivity business.
[0036] The secondary judgment submodule for sensitive business attributes is used to collect voice content sent by multi-mode handheld terminals, analyze whether the voice content contains preset keywords, if so, it is judged as a highly sensitive business, if not, it is judged as a low sensitive business.
[0037] The business sensitivity attribute three-stage determination submodule is used to detect the GPS positioning parameters of the multi-mode handheld terminal and determine whether the multi-mode handheld terminal is in a sensitive area based on the GPS positioning parameters. If it is, it is determined to be a highly sensitive business; if not, it is determined to be a low sensitive business after three stages.
[0038] The Business Sensitivity Level Determination Submodule is used to determine the business sensitivity level corresponding to a voice service request based on the number of times the high-sensitivity business judgment mechanism for voice communication requests is triggered.
[0039] Preferably, the voice frame encryption module includes:
[0040] The encryption strategy determination submodule is used to determine whether a voice communication request is a high-risk service based on the service sensitivity level. If it is, it determines that a dual-link joint encryption strategy needs to be started. If not, it determines that a dual-link joint encryption strategy does not need to be started and determines that a dynamic preferred link encryption strategy should be started.
[0041] The speech frame extraction submodule is used to extract the core speech frames from the speech content using frame extraction technology, and to generate speech enhancement frames by high-precision encoding of the original speech frames.
[0042] The voice frame encryption submodule is used to encrypt the voice core frame using AES-256 through a narrowband link to obtain the first ciphertext, and to encrypt the voice enhancement frame using the national standard SM4 through a broadband link to obtain the second ciphertext.
[0043] The ciphertext encapsulation submodule is used to encapsulate the first ciphertext and the second ciphertext in a spatiotemporal interleaving to generate a hybrid ciphertext.
[0044] Preferably, the ciphertext block transmission module includes:
[0045] The joint root key generation submodule is used to fuse and concatenate the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal into N bytes, perform an XOR operation on the N bytes and the temporary random number broadcast by the base station, and generate a joint root key based on the operation result.
[0046] The ciphertext block processing submodule is used to process the core frame ciphertext block and the enhanced frame ciphertext block in the mixed ciphertext by using data block segmentation technology.
[0047] The channel quality assessment submodule is used to determine the transmission requirements for the ciphertext block, determine the expected transmission efficiency based on the transmission requirements, determine the channel assessment index based on the expected transmission efficiency, and perform channel quality assessment on narrowband and wideband channels based on the channel assessment index.
[0048] The ciphertext block transmission submodule is used to adjust the channel parameters of the narrowband channel and the wideband channel according to the evaluation results. It transmits the core frame ciphertext block through the adjusted narrowband channel and the enhanced frame ciphertext block through the adjusted wideband channel.
[0049] Preferably, the ciphertext block decryption module includes:
[0050] The ciphertext block decryption submodule is used at the receiving end to decrypt the core frame ciphertext block and the enhanced frame ciphertext block using the joint root key, obtain the decrypted data segment, and determine the data tag of each decrypted data segment;
[0051] The speech sequence acquisition submodule is used to sort the core frame decryption blocks and enhancement frame decryption blocks according to the data tags, obtain the sorting results, and obtain the core frame speech sequence and enhancement frame speech sequence based on the sorting results.
[0052] The speech parameter extraction submodule is used to input the core frame speech sequence and the enhanced frame speech sequence into the preset speech fusion model. The speech fusion model extracts the basic speech skeleton from the core frame speech sequence and extracts high-frequency detail speech and background sound from the enhanced frame speech sequence.
[0053] The speech stream generation submodule is used to generate the original speech stream based on the basic speech skeleton, high-frequency detail speech, and background noise.
[0054] Through the above-mentioned technical means, the present invention achieves the following beneficial effects:
[0055] By encrypting voice content through a dual-channel joint voice encryption strategy based on the business sensitivity level, deep encryption of high-definition voice can be achieved by encrypting the core voice frame and the enhanced voice frame separately through dual channels. This improves encryption stability and efficiency, ensures complete recording of the original sound, and avoids the risk of malicious decryption of single-channel encryption, thus improving security, reliability, and practicality.
[0056] By performing a triple-checking process on user-issued voice communication requests, it is possible to comprehensively and accurately determine whether a user's communication service is a highly sensitive service, thereby improving the sensitivity level of the service and enhancing the accuracy and reliability of the check.
[0057] By selecting encryption strategies based on business risk attributes, the appropriate encryption strategy can be flexibly chosen according to the risk status of the voice content, thereby improving encryption stability and reliability. Furthermore, encrypting voice frames through broadband and narrowband links can avoid the duplication of encrypted data, thus ensuring encryption efficiency and stability.
[0058] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0059] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0060] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0061] Figure 1 A flowchart illustrating the process of a converged communication encryption method for a multi-mode handheld terminal provided by the present invention;
[0062] Figure 2 Another flowchart of the converged communication encryption method for a multi-mode handheld terminal provided by the present invention;
[0063] Figure 3 This is a schematic diagram of the structure of a converged communication encryption system for a multi-mode handheld terminal provided by the present invention;
[0064] Figure 4 This is a schematic diagram of the structure of the service sensitivity level determination module in a converged communication encryption system for a multi-mode handheld terminal provided by the present invention. Detailed Implementation
[0065] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0066] In today's rapidly developing digital age, data security and privacy protection have become key concerns across various industries. With the increasing frequency and complexity of data interactions between enterprises and users, the application of data encryption technology is becoming more widespread. However, with the surge in data volume, traditional data encryption methods are gradually revealing serious problems in terms of computational efficiency. Especially in practical business applications, the encryption computational load generated by the real-time interaction of large amounts of data has become a bottleneck affecting data transmission efficiency and system performance. Existing voice communication encryption typically uses a single encryption algorithm to encrypt the communication voice content, resulting in bandwidth bottlenecks and an inability to support high-definition voice encryption. This leads to a sharp decline in voice quality, and single-channel encryption is easily decrypted maliciously, causing data leakage and losses, thus reducing practicality and stability. To address these issues, this embodiment discloses a converged communication encryption method for multi-mode handheld terminals.
[0067] A converged communication encryption method for multi-mode handheld terminals, such as Figure 1 As shown, it includes the following steps:
[0068] Step S101: Obtain the voice service request initiated by the multi-mode handheld terminal, and parse the voice service request to determine the service sensitivity level;
[0069] Step S102: Determine whether to activate the dual-link joint encryption strategy based on the business sensitivity level. If so, encrypt the voice frame through the narrowband link and the broadband link respectively to obtain the mixed ciphertext.
[0070] Step S103: Generate a joint root key based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal, transmit the core frame ciphertext block in the hybrid ciphertext through the narrowband channel, and transmit the enhanced frame ciphertext block in the hybrid ciphertext through the wideband channel.
[0071] Step S104: At the receiving end, the core frame ciphertext block and the enhanced frame ciphertext block are decrypted and reassembled using the joint root key to obtain the original voice stream.
[0072] The working principle of the above technical solution is as follows: The terminal listens for voice requests and first checks whether the user has pressed the "Emergency Encryption" physical or virtual button. If so, it is directly marked as highly sensitive. If not, real-time keyword recognition of the voice content is initiated (for example, a preset dictionary contains "confidential," "operation code," etc.). If the keyword is recognized, it is marked as highly sensitive. If it still cannot be determined, the terminal's GPS location is checked to see if it is within a preset sensitive area (such as within the electronic fence of a specific command center or research base). If so, it is marked as highly sensitive. Finally, the level is determined based on the number of triggered judgment mechanisms (for example, triggering ≥2 mechanisms indicates high sensitivity). If the service is determined to be highly sensitive, dual-link encryption is initiated. The voice processor separates the original voice stream: it obtains the voice core frame through low-pass filtering and pitch extraction (ensuring basic intelligibility); it performs high-definition encoding on the original voice (such as using an OPUS encoder) to generate an enhanced frame containing rich details. The narrowband link uses the computationally efficient and strong AES-256 encryption core frame to generate the first ciphertext; the broadband link uses the national standard SM4 encryption for the larger data volume of the enhanced frame to generate the second ciphertext. Subsequently, the two types of ciphertext segments are tagged (e.g., type, timestamp, CRC), and a random sequence generated based on the joint root key is filled into an interleaving matrix for spatiotemporal interleaving, outputting the final mixed ciphertext stream. The terminal concatenates its slot number in the narrowband network (e.g., slot identifier in a TDMA frame) with its temporary sequence number in the broadband network (e.g., PDCP layer sequence number). This concatenation result is XORed with the temporary random number broadcast by the current serving base station (updated per session), and the result is SHA-256 hashed to generate a 128-bit joint root key. Before transmission, the mixed ciphertext is divided into blocks. Simultaneously, the quality of the current narrowband and broadband channels is evaluated (e.g., based on SNR and BER). If the narrowband channel quality is acceptable, it is used to transmit the critical core frame ciphertext blocks; the broadband channel is responsible for transmitting the large data-volume enhanced frame ciphertext blocks. Channel parameters can be fine-tuned based on the evaluation results, such as increasing the overhead of error correction codes on unstable narrowband channels. The receiving end holds the same joint root key (synchronously generated using the same algorithm). It receives data blocks from two channels, decrypts them using a joint root key, and sorts and verifies their integrity based on data tags. The sorted core frame sequence and enhanced frame sequence are then fed into a speech fusion model. This model first reconstructs the basic waveform (skeleton) of the speech using the core frame sequence, then extracts high-frequency components and background noise from the enhanced frame sequence, and fuses the detailed information into the basic waveform through frequency domain alignment and weighted superposition algorithms, ultimately synthesizing a restored speech that is highly consistent with the original speech stream.
[0073] The beneficial effects of the above technical solution are as follows: By encrypting voice content through a dual-channel joint voice encryption strategy based on the business sensitivity level, deep encryption of high-definition voice can be achieved by encrypting the core voice frame and the enhanced voice frame separately through dual channels. This improves encryption stability and efficiency, ensures complete recording of the original sound, and avoids the risk of malicious decryption of single-channel encryption. This enhances security, reliability, and practicality. It solves the problems mentioned in existing technologies where voice communication encryption usually uses a single encryption algorithm to encrypt the communication voice content, resulting in bandwidth bottlenecks, inability to support high-definition voice encryption, a sharp drop in voice quality, and vulnerability of single-channel encryption to malicious decryption, leading to data leakage and losses, and reducing practicality and stability.
[0074] In one embodiment, obtaining voice service requests initiated by multi-mode handheld terminals and parsing the voice service requests to determine the service sensitivity level includes:
[0075] The terminal's multi-mode communication management module listens for user-triggered voice communication requests and determines whether the user has manually triggered the emergency encryption button based on the voice communication request. If so, it is determined to be a highly sensitive service; otherwise, it is initially determined to be a low-sensitivity service.
[0076] Collect voice content emitted by multi-mode handheld terminals, analyze whether the voice content contains preset keywords, if so, determine it as a highly sensitive service, if not, determine it as a low-sensitivity service.
[0077] The GPS positioning parameters of the multi-mode handheld terminal are detected. Based on the GPS positioning parameters, it is determined whether the multi-mode handheld terminal is in a sensitive area. If it is, it is judged as a highly sensitive service. If not, it is judged as a low-sensitivity service after three checks.
[0078] The sensitivity level of a voice service request is determined based on the number of times the high-sensitivity service judgment mechanism for voice communication requests is triggered.
[0079] The beneficial effects of the above technical solution are as follows: by performing a triple judgment on the voice communication request issued by the user, it is possible to comprehensively and accurately determine whether the user's communication service is a highly sensitive service and then determine its service sensitivity level, thereby improving the accuracy and reliability of the judgment.
[0080] In this embodiment, it also includes:
[0081] Detect whether the terminal is in an abnormal motion state and whether the terminal network is frequently switching, and obtain device status perception data;
[0082] Real-time analysis of ambient background noise collected by the terminal microphone; identification of specific dangerous acoustic events based on ambient background noise; acquisition of environmental perception data of the device.
[0083] By monitoring stress characteristics or heart rate changes in user voice through terminal sensors, quantitative data on user emotions can be obtained.
[0084] The device status perception data, the device's environment perception data, and the user's emotion quantification data are input into a preset lightweight neural network model to obtain a dynamic sensitivity score for voice services.
[0085] The sensitivity level of voice service requests is verified and alerted based on dynamic sensitivity scoring.
[0086] The beneficial effects of the above technical solution are as follows: single or static decision rules appear rigid when facing complex and ever-changing real-world application scenarios. By introducing more dimensions of contextual information and using machine learning models for fusion decision-making, the accuracy, adaptability, and foresight of the decision can be greatly improved.
[0087] In one embodiment, determining whether to activate the dual-link joint encryption strategy based on the service sensitivity level, and if so, encrypting the voice frame through both narrowband and broadband links to obtain the hybrid ciphertext, includes:
[0088] Determine whether the voice communication request is a high-risk service based on the business sensitivity level. If it is, determine whether a dual-link joint encryption strategy needs to be activated. If not, determine whether a dual-link joint encryption strategy does not need to be activated and determine whether a dynamic preferred link encryption strategy should be activated.
[0089] The core speech frames in the speech content are extracted by frame extraction technology, and the original speech frames are encoded with high precision to generate speech enhancement frames.
[0090] The first ciphertext is obtained by encrypting the voice core frame with AES-256 through a narrowband link, and the second ciphertext is obtained by encrypting the voice enhancement frame with the national standard SM4 through a broadband link.
[0091] The first ciphertext and the second ciphertext are spatiotemporally interleaved and encapsulated to generate a mixed ciphertext.
[0092] In this embodiment, spatiotemporal interleaving encapsulation means performing temporal and spatial domain data interleaving encapsulation on the first ciphertext and the second ciphertext;
[0093] In this embodiment, the speech core frame is represented as a speech frame that contains only speech content, while the speech enhancement frame is represented as a speech frame that contains both speech content and ambient speech.
[0094] The beneficial effects of the above technical solution are as follows: by selecting the encryption strategy according to the business risk attributes, the appropriate encryption strategy can be flexibly selected according to the risk status of the voice content, which improves the encryption stability and reliability. Furthermore, by encrypting voice frames through broadband and narrowband links, the duplication of encrypted data can be avoided, thereby ensuring encryption efficiency and stability.
[0095] In this embodiment, the first ciphertext and the second ciphertext are spatiotemporally interleaved and encapsulated to generate a hybrid ciphertext, including:
[0096] The first ciphertext is divided into multiple key data segments, and the second ciphertext is divided into multiple enhanced data segments;
[0097] Add voice data tags and integrity verification codes to each key data segment and enhanced data segment;
[0098] Obtain a random mapping of each key data segment and enhanced data segment based on the joint root key, and construct an interleaving matrix based on the random mapping;
[0099] Generate mixed ciphertext based on the interleaving matrix.
[0100] In one embodiment, such as Figure 2 As shown, the method of generating a joint root key based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal, transmitting the core frame ciphertext block in the hybrid ciphertext through the narrowband channel, and transmitting the enhanced frame ciphertext block in the hybrid ciphertext through the wideband channel includes:
[0101] Step S201: Merge and concatenate the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal into N bytes, XOR the N bytes with the temporary random number broadcast by the base station, and generate a joint root key based on the result of the operation.
[0102] Step S202: Divide the core frame ciphertext block and the enhanced frame ciphertext block in the hybrid ciphertext into blocks using data block segmentation technology;
[0103] Step S203: Determine the transmission requirements for the ciphertext block, determine the expected transmission efficiency based on the transmission requirements, determine the channel evaluation index based on the expected transmission efficiency, and evaluate the channel quality of the narrowband channel and the wideband channel based on the channel evaluation index.
[0104] Step S204: Adjust the channel parameters of the narrowband channel and the wideband channel according to the evaluation results. Transmit the core frame ciphertext block through the adjusted narrowband channel and the enhanced frame ciphertext block through the adjusted wideband channel.
[0105] In this embodiment, the joint root key is generated in specific software. The joint root key generated by the sending end through the software is synchronously generated in the same software at the receiving end. That is, the receiving end can obtain the joint root key generated by the sending end through the software without transmission.
[0106] In this embodiment, the terminal reads the timeslot number (e.g., value 5) allocated by the narrowband base station it is currently connected to and the sequence number (e.g., value 0xA1B2C3D4) allocated by the broadband network. These are concatenated into 16 bytes of data according to a preset format. The base station broadcasts a 16-byte temporary random number (valid for this session). The terminal performs a byte-by-byte XOR operation between the concatenated data and the temporary random number. The result of the XOR operation is applied to the SHA-256 hash function, and the first 128 bits of the hash value are used as the joint root key for this communication. This key remains unchanged during the session and is used for subsequent encapsulation and decryption. The sending end determines the transmission efficiency target based on the current network conditions (e.g., latency requirement <100ms, packet loss rate <1%). Accordingly, signal-to-noise ratio (SNR) and bit error rate (BER) are selected as core evaluation indicators. These two indicators for both narrowband and broadband channels are monitored in real time. If a decrease in SNR is detected in the narrowband channel, its transmit power is automatically increased, and it may switch to a more robust modulation scheme (such as downgrading from 16-QAM to QPSK) to ensure reliable transmission of the core frame ciphertext block. The wideband channel, based on its good quality, maintains a higher-order modulation scheme to transmit a large number of enhanced frame ciphertext blocks.
[0107] The beneficial effects of the above technical solution are as follows: generating the joint root key by XOR operation can ensure the compatibility of the root key with both wide and narrow bands, thereby ensuring the stability of decryption. Furthermore, by performing channel quality assessment, it can be ensured that both wide and narrow band channels meet the data transmission requirements, thus improving the stability of data transmission.
[0108] In this embodiment, the joint root key can also be generated in the following way:
[0109] Both communicating parties use the current communication link to send specific probe signals to each other. Both parties independently measure the channel characteristic data of the channel, which includes: channel state information and fast fluctuation physical layer parameters indicating received signal strength.
[0110] The two communicating parties quantize the extracted channel feature data to generate an initial key material;
[0111] The initial key material is combined with the narrowband timeslot number and wideband serial number of the multi-mode handheld terminal to perform key derivation function operations, and finally generate the session joint root key.
[0112] The beneficial effects of the above technical solution are as follows: even if the attacker obtains the device identifier and random number, they will not be able to generate the same key because they cannot know the unique instantaneous channel characteristics of the two communicating parties, thus achieving physical layer security enhancement and effectively resisting eavesdropping.
[0113] In one embodiment, the step of decrypting and reconstructing the core frame ciphertext block and the enhanced frame ciphertext block using a joint root key at the receiving end to obtain the original voice stream includes:
[0114] At the receiving end, the core frame ciphertext block and the enhanced frame ciphertext block are decrypted using the joint root key to obtain the decrypted data segment and determine the data tag of each decrypted data segment.
[0115] The core frame decryption blocks and the enhancement frame decryption blocks are sorted according to the data tags, and the sorting results are obtained. The core frame speech sequence and the enhancement frame speech sequence are obtained according to the sorting results.
[0116] The core frame speech sequence and the enhanced frame speech sequence are input into the preset speech fusion model. The speech fusion model extracts the basic speech skeleton from the core frame speech sequence and extracts high-frequency detail speech and background sound from the enhanced frame speech sequence.
[0117] The original speech stream is generated based on the basic speech skeleton, high-frequency detail speech, and background noise.
[0118] The beneficial effects of the above technical solution are as follows: by sorting the encrypted blocks based on data tags, speech frames can be effectively distinguished and sorted to avoid data mixing, while also ensuring the sorting accuracy of speech frames. Furthermore, by extracting the basic speech skeleton, high-frequency details, and background sounds through the speech fusion model, the user's call speech can be tracked and recorded in all aspects, ensuring the integrity and high quality of the speech stream.
[0119] In this embodiment, after generating the original speech stream, the following steps are also included:
[0120] Acoustic anomaly detection model is used to detect acoustic abrupt changes in the original speech stream to obtain acoustic abrupt change features, and multiple abrupt acoustic time period nodes are marked based on the acoustic abrupt change features.
[0121] Low-frequency acoustic signals are acquired at nodes of abnormal acoustic periods. The low-frequency acoustic signals are processed using windowing and framing techniques, and the short-time energy of each frame is calculated to obtain the short-time energy sequence.
[0122] Based on the short-time energy sequence, multiple abrupt energy points and time delay parameters between each abrupt energy point are determined, and the energy abrupt change pattern of each node in the anomalous acoustic period is determined based on the time delay parameters.
[0123] The abnormal acoustic information content of each abnormal acoustic time period node is determined by analyzing the energy mutation law, and the abnormal acoustic quantitative characterization of each abnormal acoustic time period node is determined based on the abnormal acoustic information content.
[0124] A set of multi-frequency acoustic wave test signals is generated based on the abnormal acoustic quantization characterization. The multi-frequency acoustic wave test signals are then propagated and analyzed in a simulated acoustic environment to construct the acoustic wave penetration characteristic matrix and the acoustic wave penetration characteristic matrix.
[0125] The type of abnormal acoustic feature is determined based on the acoustic wave penetration characteristic matrix and the acoustic wave penetration characteristic matrix.
[0126] By using speech activity detection and sound source separation technology, user speech acoustic features and environmental baseline acoustic features are separated from the original speech stream. Based on the distribution of the two in the signal-to-noise ratio spectrum, a minimum signal-to-noise ratio threshold in which user speech can be clearly identified is determined as the critical value of user voice.
[0127] Based on the user's voice threshold, the ideal acoustic features that the user's voice should have in a clean environment are estimated through a speech enhancement algorithm and used as the target acoustic features.
[0128] Based on the determined abnormal acoustic feature type, the standard acoustic feature corresponding to that type is called from the preset abnormal feature library, and the multidimensional acoustic feature difference vector between the standard acoustic feature and the target acoustic feature is determined.
[0129] Automatic speech recognition and natural language processing are performed on the original speech stream to obtain its contextual semantic features. The contextual semantic features and the multidimensional acoustic feature difference vector are input into a preset end-to-end speech generation model to generate missing speech segments for each abnormal acoustic time node that are semantically coherent with the context and have smooth transitions in acoustic features. These missing speech segments are then inserted into the original speech stream to generate a complete speech stream.
[0130] The completed speech stream will be output as the final speech stream.
[0131] In this embodiment, the acoustic anomaly detection model is represented as a hybrid model based on convolutional neural network (CNN) and long short-term memory network (LSTM) to extract spatiotemporal features from the speech spectrogram and identify abrupt changes that deviate from the normal speech pattern.
[0132] In this embodiment, low-frequency acoustic signals are represented as audio signals with a frequency range of 0Hz to 300Hz;
[0133] In this embodiment, the time delay parameter reflects the temporal distribution pattern of the energy mutation point. Based on the time delay parameter, the pattern category of energy mutation within the abnormal period (e.g., 'single pulse type', 'continuous oscillation type', 'random burst type') is determined by calculating its statistical characteristics (such as variance and skewness) or by inputting it into a pre-trained classifier. That is, the energy mutation law.
[0134] In this embodiment, the amount of anomalous acoustic information for each anomalous acoustic time period node is calculated by normalizing the entropy value based on the mode category and the short-time energy sequence. The higher the entropy value, the greater the randomness and uncertainty of the anomalous event.
[0135] In this embodiment, the anomalous acoustic quantization characterization is a multidimensional feature vector whose elements include, but are not limited to: the mode category, the normalized entropy value, the maximum value and the average value of the short-time energy sequence;
[0136] In this embodiment, the acoustic wave penetration characteristic matrix is used to characterize the attenuation and scattering characteristics of acoustic waves of different frequencies in the simulated environment, and the acoustic wave penetration characteristic matrix is used to characterize the reflection and diffraction characteristics of acoustic waves of different frequencies at the boundary of the simulated environment.
[0137] In this embodiment, the anomalous acoustic feature type is determined by matching the features of two matrices with a database of known anomalous types.
[0138] The beneficial effects of the above technical solution are as follows: it can actively identify and repair voice interruptions, distortions, and sudden noises caused by environmental interference, equipment failures, or transmission errors, ensuring that core voice information can be transmitted clearly and completely even in harsh communication environments. Furthermore, by repairing the original voice stream through context-aware methods based on contextual semantics and acoustic feature differences, it can ensure that the repaired voice is natural and coherent in both content and listening experience, while achieving a smooth transition with the user's original voice timbre and intonation, realizing seamless splicing at the acoustic level, guaranteeing voice quality, and improving practicality and stability.
[0139] In one embodiment, this embodiment also discloses a converged communication encryption system for a multi-mode handheld terminal, such as... Figure 3 As shown, the system includes:
[0140] The business sensitivity level determination module 301 is used to obtain voice service requests initiated by multi-mode handheld terminals, and to parse the voice service requests to determine the business sensitivity level.
[0141] The voice frame encryption module 302 is used to determine whether to start the dual-link joint encryption strategy based on the service sensitivity level. If so, the voice frame is encrypted through the narrowband link and the broadband link respectively to obtain the mixed ciphertext.
[0142] The ciphertext block transmission module 303 is used to generate a joint root key based on the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal, transmit the core frame ciphertext block in the hybrid ciphertext through the narrowband channel, and transmit the enhanced frame ciphertext block in the hybrid ciphertext through the wideband channel.
[0143] The ciphertext block decryption module 304 is used at the receiving end to decrypt and reassemble the core frame ciphertext block and the enhanced frame ciphertext block using a joint root key to obtain the original voice stream.
[0144] The working principle and beneficial effects of the above technical solution have been explained in the method embodiments, and will not be repeated here.
[0145] In one embodiment, such as Figure 4 As shown, the business sensitivity level determination module 301 includes:
[0146] The business sensitivity attribute determination submodule 3011 is used to listen to the user-triggered voice communication request through the terminal's multi-mode communication management module, and determine whether the user has manually triggered the emergency encryption button based on the voice communication request. If so, it is determined to be a highly sensitive business; if not, it is initially determined to be a low-sensitivity business.
[0147] The secondary judgment submodule 3012 for sensitive business attributes is used to collect the voice content sent by the multi-mode handheld terminal, and analyze whether the voice content contains preset keywords. If it does, it is judged as a highly sensitive business; if not, it is judged as a low-sensitivity business.
[0148] The business sensitivity attribute three-stage determination submodule 3013 is used to detect the GPS positioning parameters of the multi-mode handheld terminal and determine whether the multi-mode handheld terminal is in a sensitive area based on the GPS positioning parameters. If it is, it is determined to be a highly sensitive business; if not, it is determined to be a low sensitive business after three stages.
[0149] The business sensitivity level determination submodule 3014 is used to determine the business sensitivity level corresponding to a voice service request based on the number of times the high-sensitivity business judgment mechanism for voice communication requests is triggered.
[0150] In one embodiment, the voice frame encryption module includes:
[0151] The encryption strategy determination submodule is used to determine whether a voice communication request is a high-risk service based on the service sensitivity level. If it is, it determines that a dual-link joint encryption strategy needs to be started. If not, it determines that a dual-link joint encryption strategy does not need to be started and determines that a dynamic preferred link encryption strategy should be started.
[0152] The speech frame extraction submodule is used to extract the core speech frames from the speech content using frame extraction technology, and to generate speech enhancement frames by high-precision encoding of the original speech frames.
[0153] The voice frame encryption submodule is used to encrypt the voice core frame using AES-256 through a narrowband link to obtain the first ciphertext, and to encrypt the voice enhancement frame using the national standard SM4 through a broadband link to obtain the second ciphertext.
[0154] The ciphertext encapsulation submodule is used to encapsulate the first ciphertext and the second ciphertext in a spatiotemporal interleaving to generate a hybrid ciphertext.
[0155] In one embodiment, the ciphertext block transmission module includes:
[0156] The joint root key generation submodule is used to fuse and concatenate the narrowband timeslot number and wideband sequence number of the multi-mode handheld terminal into N bytes, perform an XOR operation on the N bytes and the temporary random number broadcast by the base station, and generate a joint root key based on the operation result.
[0157] The ciphertext block processing submodule is used to process the core frame ciphertext block and the enhanced frame ciphertext block in the mixed ciphertext by using data block segmentation technology.
[0158] The channel quality assessment submodule is used to determine the transmission requirements for the ciphertext block, determine the expected transmission efficiency based on the transmission requirements, determine the channel assessment index based on the expected transmission efficiency, and perform channel quality assessment on narrowband and wideband channels based on the channel assessment index.
[0159] The ciphertext block transmission submodule is used to adjust the channel parameters of the narrowband channel and the wideband channel according to the evaluation results. It transmits the core frame ciphertext block through the adjusted narrowband channel and the enhanced frame ciphertext block through the adjusted wideband channel.
[0160] In one embodiment, the ciphertext block decryption module includes:
[0161] The ciphertext block decryption submodule is used at the receiving end to decrypt the core frame ciphertext block and the enhanced frame ciphertext block using the joint root key, obtain the decrypted data segment, and determine the data tag of each decrypted data segment;
[0162] The speech sequence acquisition submodule is used to sort the core frame decryption blocks and enhancement frame decryption blocks according to the data tags, obtain the sorting results, and obtain the core frame speech sequence and enhancement frame speech sequence based on the sorting results.
[0163] The speech parameter extraction submodule is used to input the core frame speech sequence and the enhanced frame speech sequence into the preset speech fusion model. The speech fusion model extracts the basic speech skeleton from the core frame speech sequence and extracts high-frequency detail speech and background sound from the enhanced frame speech sequence.
[0164] The speech stream generation submodule is used to generate the original speech stream based on the basic speech skeleton, high-frequency detail speech, and background noise.
[0165] Those skilled in the art should understand that the "first" and "second" in this invention simply refer to different application stages.
[0166] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0167] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for convergent communication encryption of a multi-mode handheld terminal, characterized in that, The method comprises the following steps: obtaining a voice service request initiated by a multi-mode handheld terminal, analyzing the voice service request to determine a service sensitivity level; determining whether to start a dual-link joint encryption strategy based on the service sensitivity level, and if so, encrypting the voice frame through a narrowband link and a broadband link respectively to obtain mixed ciphertext; generating a joint root key based on the narrowband time slot number and the broadband sequence number of the multi-mode handheld terminal, transmitting the core frame ciphertext block in the mixed ciphertext through the narrowband channel, and transmitting the enhanced frame ciphertext block in the mixed ciphertext through the broadband channel; decrypting and recombining the core frame ciphertext block and the enhanced frame ciphertext block through the joint root key at the receiving end to obtain the original voice stream; The method of decrypting and recombining the core frame ciphertext block and the enhanced frame ciphertext block through the joint root key at the receiving end to obtain the original voice stream comprises: decrypting the core frame ciphertext block and the enhanced frame ciphertext block through the joint root key at the receiving end to obtain decrypted data segments, and determining the data tags of each decrypted data segment; sorting the core frame decrypted ciphertext block and the enhanced frame decrypted ciphertext block according to the data tags to obtain a sorting result, and obtaining the core frame voice sequence and the enhanced frame voice sequence according to the sorting result; inputting the core frame voice sequence and the enhanced frame voice sequence into a preset voice fusion model, extracting a basic voice skeleton from the core frame voice sequence and extracting high-frequency detail voice and background sound from the enhanced frame voice sequence through the voice fusion model; generating the original voice stream according to the basic voice skeleton, the high-frequency detail voice and the background sound; After the original voice stream is generated, the method further comprises: detecting the original voice stream through an acoustic anomaly detection model to obtain acoustic mutation features, and marking a plurality of abnormal acoustic period nodes based on the acoustic mutation features; collecting low-frequency acoustic signals at the abnormal acoustic period nodes, processing the low-frequency acoustic signals through a windowing and framing technique, and calculating the short-time energy of each frame to obtain a short-time energy sequence; determining a plurality of mutation energy points and time delay parameters between the mutation energy points according to the short-time energy sequence, and determining the energy mutation law of each abnormal acoustic period node according to the time delay parameters; analyzing the energy mutation law to determine the abnormal acoustic information amount of each abnormal acoustic period node, and determining the abnormal acoustic quantization representation of each abnormal acoustic period node based on the abnormal acoustic information amount; generating a group of multi-frequency sound wave test signals according to the abnormal acoustic quantization representation, propagating and analyzing the multi-frequency sound wave test signals in a simulated acoustic environment to construct a sound wave penetration characteristic matrix; determining an abnormal acoustic feature type based on the sound wave penetration characteristic matrix; separating user voice acoustic features and environmental baseline acoustic features from the original voice stream through a voice activity detection and sound source separation technology, determining a minimum signal-to-noise ratio threshold at which a user voice can be clearly recognized as a user sound line critical value according to the distribution of the two on a signal-to-noise ratio spectrum, and estimating the ideal acoustic features that a user voice should have in a pure environment as target acoustic features through a voice enhancement algorithm based on the user sound line critical value. According to the determined abnormal acoustic feature type, a standard acoustic feature corresponding to the type is called from a preset abnormal feature library, and a multi-dimensional acoustic feature difference vector between the standard acoustic feature and the target acoustic feature is determined; Automatic speech recognition and natural language processing are performed on the original speech stream to obtain context semantic features, the context semantic features and the multi-dimensional acoustic feature difference vector are input into a preset end-to-end speech generation model to generate a missing speech segment that is coherent in context semantics and smoothly transitions in acoustic features for each abnormal acoustic period node, and the missing speech segment is inserted into the original speech stream to generate a completed speech stream; The completed speech stream is output as a final speech stream.
2. The fusion communication encryption method for a multi-mode handheld terminal according to claim 1, characterized in that, The voice service request initiated by the multi-mode handheld terminal is acquired, and the voice service request is analyzed to determine the service sensitivity level, including: Listening to the voice communication request triggered by the user through the multi-mode communication management module of the terminal, determining whether the user manually triggers the emergency encryption button based on the voice communication request, if yes, determining as high-sensitive service, if not, initially determining as low-sensitive service; Acquiring the voice content sent by the multi-mode handheld terminal, analyzing whether the voice content contains a preset keyword, if yes, determining as high-sensitive service, if not, secondarily determining as low-sensitive service; Detecting the GPS positioning parameters of the multi-mode handheld terminal, determining whether the multi-mode handheld terminal is in a sensitive area according to the GPS positioning parameters, if yes, determining as high-sensitive service, if not, thirdly determining as low-sensitive service; Determining the service sensitivity level corresponding to the voice service request according to the number of high-sensitive service determination mechanisms triggered by the voice communication request.
3. The fusion communication encryption method for a multi-mode handheld terminal according to claim 1, characterized in that, If yes, the voice frame is encrypted through the narrowband link and the broadband link respectively to obtain the hybrid ciphertext, including: According to the service sensitivity level, determining whether the voice communication request is a high-risk service, if yes, determining that the dual-link joint encryption strategy needs to be started, if not, determining that the dual-link joint encryption strategy does not need to be started, and determining to start the dynamic link encryption strategy; Extracting the voice core frame in the voice content through frame extraction technology, and generating a voice enhanced frame by high-precision encoding the original voice frame; Encrypting the voice core frame through the narrowband link using AES-256 to obtain the first ciphertext, and encrypting the voice enhanced frame through the broadband link using the national standard SM4 to obtain the second ciphertext; The first ciphertext and the second ciphertext are spatio-temporally interleaved and packaged to generate a hybrid ciphertext.
4. The fusion communication encryption method for a multi-mode handheld terminal according to claim 1, characterized in that, The narrowband time slot number and the broadband sequence number of the multi-mode handheld terminal are fused and spliced into N bytes, and the N bytes are XORed with the temporary random number broadcast by the base station to generate a joint root key; The core frame ciphertext block and the enhanced frame ciphertext block in the hybrid ciphertext are processed by block division technology. Determine the transmission requirement of the ciphertext block, determine the expected transmission efficiency according to the transmission requirement, determine the channel evaluation index according to the expected transmission efficiency, and evaluate the channel quality of the narrowband channel and the broadband channel based on the channel evaluation index; Adjust the channel parameters of the narrowband channel and the broadband channel according to the evaluation result, transmit the core frame ciphertext block through the adjusted narrowband channel, and transmit the enhanced frame ciphertext block through the adjusted broadband channel.
5. A converged communication encryption system for a multi-mode handheld terminal, characterized by, The system comprises: A service sensitivity level determination module is configured to obtain a voice service request initiated by a multi-mode handheld terminal, analyze the voice service request, and determine a service sensitivity level; A voice frame encryption module is configured to determine whether to start a dual-link joint encryption strategy based on the service sensitivity level, and if so, encrypt voice frames through a narrowband link and a broadband link respectively to obtain mixed ciphertext; A ciphertext block transmission module is configured to generate a joint root key based on a narrowband time slot number and a broadband sequence number of the multi-mode handheld terminal, transmit core frame ciphertext blocks in the mixed ciphertext through a narrowband channel, and transmit enhanced frame ciphertext blocks in the mixed ciphertext through a broadband channel; A ciphertext block decryption module is configured to decrypt and recombine the core frame ciphertext blocks and the enhanced frame ciphertext blocks through the joint root key at a receiving end to obtain an original voice stream; The ciphertext block decryption module comprises: A ciphertext block decryption submodule is configured to decrypt the core frame ciphertext blocks and the enhanced frame ciphertext blocks through the joint root key at the receiving end to obtain decrypted data segments and determine data tags of each decrypted data segment; A voice sequence acquisition submodule is configured to sort the core frame decrypted ciphertext blocks and the enhanced frame decrypted ciphertext blocks according to the data tags, obtain a sorting result, and obtain core frame voice sequences and enhanced frame voice sequences according to the sorting result; A voice parameter extraction submodule is configured to input the core frame voice sequences and the enhanced frame voice sequences into a preset voice fusion model, extract a basic voice skeleton from the core frame voice sequences and high-frequency detail voice and background sound from the enhanced frame voice sequences through the voice fusion model; A voice stream generation submodule is configured to generate an original voice stream according to the basic voice skeleton and the high-frequency detail voice and the background sound; After the original voice stream is generated, the following steps are further included: Detect the original voice stream through an acoustic anomaly detection model to obtain acoustic mutation features, and mark a plurality of abnormal acoustic period nodes based on the acoustic mutation features; Collect low-frequency acoustic signals at the abnormal acoustic period nodes, process the low-frequency acoustic signals through a windowing and framing technique, calculate the short-time energy of each frame, and obtain a short-time energy sequence; Determine a plurality of mutation energy points and time delay parameters between the mutation energy points according to the short-time energy sequence, and determine the energy mutation law of each abnormal acoustic period node according to the time delay parameters; Analyze the energy mutation law to determine the abnormal acoustic information amount of each abnormal acoustic period node, and determine the abnormal acoustic quantization representation of each abnormal acoustic period node based on the abnormal acoustic information amount; Generate a group of multi-frequency sound wave test signals according to the abnormal acoustic quantization representation, propagate and analyze the multi-frequency sound wave test signals in a simulated acoustic environment to construct a sound wave penetration characteristic matrix; Determine the abnormal acoustic feature type based on the sound wave penetration characteristic matrix; The user voice acoustic feature and the environmental baseline acoustic feature are separated from the original voice stream through voice activity detection and sound source separation technology, a minimum signal-to-noise ratio threshold at which the user voice can be clearly recognized is determined according to the distribution of the two on the signal-to-noise ratio spectrum, and the minimum signal-to-noise ratio threshold is used as a user sound ray critical value; Based on the user sound ray critical value, an ideal acoustic feature that the user voice should have in a pure environment is estimated as a target acoustic feature through a voice enhancement algorithm; According to the determined abnormal acoustic feature type, a standard acoustic feature corresponding to the type is called from a preset abnormal feature library, and a multi-dimensional acoustic feature difference vector between the standard acoustic feature and the target acoustic feature is determined; The context semantic feature is obtained by performing automatic speech recognition and natural language processing on the original voice stream, and the context semantic feature and the multi-dimensional acoustic feature difference vector are input into a preset end-to-end voice generation model to generate a missing voice segment that is coherent in context semantics and smoothly transits in acoustic feature for each abnormal acoustic period node, and the missing voice segment is inserted into the original voice stream to generate a completed voice stream; The completed voice stream is output as a final voice stream.
6. The converged communication encryption system for multi-mode handheld terminals of claim 5, wherein, The service sensitive level determination module comprises: A service sensitive attribute judgment submodule is configured to listen to a voice communication request triggered by a user through a multi-mode communication management module of the terminal, determine whether the user manually triggers an emergency encryption button based on the voice communication request, determine that the service is a high-sensitive service if the user manually triggers the emergency encryption button, and preliminarily determine that the service is a low-sensitive service if the user does not manually trigger the emergency encryption button; A service sensitive attribute secondary judgment submodule is configured to collect voice content sent by the multi-mode handheld terminal, analyze whether the voice content contains a preset keyword, determine that the service is a high-sensitive service if the voice content contains the preset keyword, and secondarily determine that the service is a low-sensitive service if the voice content does not contain the preset keyword; A service sensitive attribute tertiary judgment submodule is configured to detect a GPS positioning parameter of the multi-mode handheld terminal, determine whether the multi-mode handheld terminal is located in a sensitive area according to the GPS positioning parameter, determine that the service is a high-sensitive service if the multi-mode handheld terminal is located in the sensitive area, and thirdly determine that the service is a low-sensitive service if the multi-mode handheld terminal is not located in the sensitive area; A service sensitive level determination submodule is configured to determine a service sensitive level corresponding to the voice communication request according to a high-sensitive service judgment mechanism triggering quantity of the voice communication request.
7. The converged communication encryption system for multi-mode handheld terminals of claim 5, wherein, The voice frame encryption module comprises: An encryption strategy determination submodule is configured to determine whether the voice communication request is a high-risk service according to the service sensitive level, determine that a dual-link joint encryption strategy needs to be started if the voice communication request is the high-risk service, and determine that the dual-link joint encryption strategy does not need to be started and a dynamic optimal link encryption strategy needs to be started if the voice communication request is not the high-risk service; A voice frame extraction submodule is configured to extract a voice core frame in the voice content through a frame extraction technology, and generate a voice enhancement frame by high-precision encoding of an original voice frame; A voice frame encryption submodule is configured to encrypt the voice core frame through a narrowband link to obtain a first ciphertext, and encrypt the voice enhancement frame through a wideband link to obtain a second ciphertext; A ciphertext packaging submodule is configured to perform time-space interleaving packaging on the first ciphertext and the second ciphertext to generate a mixed ciphertext.
8. The converged communication encryption system for multi-mode handheld terminals of claim 5, wherein, The ciphertext block transmission module comprises: The joint root key generation submodule is configured to fuse and splice the narrowband time slot number of the multi-mode handheld terminal and the wideband sequence number into N bytes, perform XOR operation on the N bytes and the temporary random number broadcast by the base station, and generate a joint root key according to the operation result; The ciphertext block processing submodule is configured to perform block processing on the core frame ciphertext block and the enhanced frame ciphertext block in the mixed ciphertext through a data block block technology; The channel quality evaluation submodule is configured to determine a transmission requirement for the ciphertext block, determine an expected transmission efficiency according to the transmission requirement, determine a channel evaluation index according to the expected transmission efficiency, and perform channel quality evaluation on the narrowband channel and the wideband channel based on the channel evaluation index; The ciphertext block transmission submodule is configured to adjust channel parameters of the narrowband channel and the wideband channel according to the evaluation result, transmit the core frame ciphertext block through the adjusted narrowband channel, and transmit the enhanced frame ciphertext block through the adjusted wideband channel.
Citation Information
Patent Citations
Wideband and narrowband integrated multi-connection trunking system and distribution method of transmission channels of same
CN104601316A
Secure data transmission method, terminal and multi-mode communication terminal
CN110166410A