A lightweight encryption and end-side sensitive content protection method for narrowband weak network

CN122205124BActive Publication Date: 2026-09-29BEIJING IACTIVE NETWORK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610343402.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-20
Publication Date
2026-09-29
Estimated Expiration
2046-03-20

AI Technical Summary

Technical Problem

[0005]鉴于此,本发明提出了一种用于窄带弱网的轻量加密与端侧敏感内容保护方法,旨在解决窄带弱网条件下,现有音视频通信在引入安全保护与敏感信息保护机制时易产生传输与处理开销增大、导致难以兼顾安全保护与稳定传输的问题

Benefits of technology

[0016]与现有技术相比,本发明的有益效果在于:通过在会话建立阶段进行双向认证并派生会话密钥与初始计数器参数,使发送端与接收端在SM4计数器模式下能够使用一致的密钥与计数器参数完成媒体码流的加密与解密;通过在发送端对每一视频帧在编码前进行端侧敏感内容检测并对满足预设敏感阈值的敏感区域执行像素级遮蔽,同时对对应原始视频帧进行缓存清除且不写入持久化存储,有利于降低原始敏感帧在终端侧的留存风险;并通过对加密后的密文块进行前向纠错编码并与冗余包封装发送,使接收端在丢包条件下可利用前向纠错解码恢复密文块集合并重组媒体码流,从而更适配窄带弱网链路下的传输需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122205124B_ABST
    Figure CN122205124B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of information security, and discloses a lightweight encryption and end-side sensitive content protection method for narrow-band weak networks, which comprises the following steps: performing bidirectional authentication and deriving session keys and initial counter parameters when a session is established between a sending terminal and a receiving terminal; performing end-side sensitive content detection on each frame of video before encoding at the sending terminal, performing real-time pixel masking on qualified areas, and clearing the original frame; dividing the desensitized frame and audio into blocks after encoding, and performing SM4 counter mode encryption; performing forward error correction encoding on the cipher block to generate a redundant packet, and encapsulating the redundant packet into a transmission unit for sending; and first recovering the cipher block by error correction at the receiving terminal, then decrypting, recombining and decoding, and finally outputting. The present application reduces bandwidth overhead through the fusion encoding of encryption and forward error correction, and realizes in-situ desensitization of sensitive content through real-time detection and masking at the end side, so as to balance transmission efficiency, data security and content compliance under the condition of narrow-band weak networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information security technology, and more specifically, to a lightweight encryption and edge-side sensitive content protection method for narrowband and weak network applications. Background Technology

[0002] With the development of mobile internet, the Internet of Things, and edge computing, audio and video communication has gradually expanded from traditional broadband networks to areas with weak cellular coverage, low-speed private networks, congested hotspots, and some narrowband access environments. In these environments, link bandwidth is limited, latency jitter is significant, packet loss rate is increased, and terminal-side computing power, storage, and energy consumption budgets are limited, which places more stringent requirements on audio and video services for transmission protocols, media encapsulation, and security mechanisms.

[0003] Existing audio and video communication systems typically need to provide data confidentiality protection during media transmission and, in certain scenarios, protect or control sensitive information in the audio and video content. Traditional security protection methods often employ common session negotiation and encryption encapsulation mechanisms to group media data and attach necessary protocol headers, verification fields, and control information; while content protection mechanisms often involve additional detection, annotation, reporting, or storage management processes in engineering implementation.

[0004] The above mechanisms are generally acceptable under broadband and stable network conditions, but under narrowband and weak network conditions, the additional byte overhead, handshake interaction overhead and processing overhead brought about by security and content protection will further squeeze the bandwidth budget available for media payload, and amplify the impact of packet loss and jitter on media continuity, making it difficult for the system to maintain stable real-time transmission while ensuring data security and content compliance. Summary of the Invention

[0005] In view of this, the present invention proposes a lightweight encryption and end-side sensitive content protection method for narrowband weak networks, aiming to solve the problem that existing audio and video communication under narrowband weak network conditions is prone to increased transmission and processing overhead when introducing security protection and sensitive information protection mechanisms, making it difficult to balance security protection and stable transmission.

[0006] In one aspect, the present invention proposes a lightweight encryption and end-side sensitive content protection method for narrowband weak networks, comprising: When the sending terminal and the receiving terminal establish a session connection, a two-way authentication is performed using an identifier cryptosystem, and a session key and initial counter parameters are derived based on the output of the two-way authentication. When acquiring audio and video data on the sending terminal side, sensitive content detection is performed on each video frame before encoding to obtain a set of sensitive regions; pixel-level masking is performed on sensitive regions in the set of sensitive regions whose detection confidence is greater than or equal to a preset sensitive threshold to obtain desensitized video frames, and the corresponding original video frames are cleared from the cache and not written to persistent storage. The desensitized video frames and audio data are encoded to obtain a media stream, and the media stream is divided into blocks aligned with the SM4 block length to obtain several plaintext blocks; Using the SM4 national cryptographic algorithm in counter mode, the session key and the initial counter parameters are used to encrypt several plaintext blocks to generate several corresponding ciphertext blocks; Forward error correction coding is performed on several ciphertext blocks to generate FEC redundancy packets, and several ciphertext blocks and FEC redundancy packets are encapsulated into transmission units and sent. The total bandwidth of the narrowband weak network link is less than or equal to a preset narrowband threshold, and the packet loss rate is greater than a preset packet loss threshold. On the receiving terminal side, the received transmission unit is forward-corrected and decoded to obtain a set of ciphertext blocks; the ciphertext block set is decrypted using the counter mode of the national cryptographic SM4 algorithm, using the session key and the initial counter parameters to obtain plaintext blocks and reassemble them into a media stream; the media stream is then decoded to obtain a video frame sequence and an audio sampling sequence.

[0007] Furthermore, when establishing a session connection between the sending terminal and the receiving terminal, a two-way authentication is performed using an identifier cryptosystem, and when deriving the session key and initial counter parameters based on the output of the two-way authentication, the process includes: The sending terminal generates a first random number and constructs a session establishment message, which includes a protocol version number, an encryption mode identifier, an FEC parameter set identifier, a media encoding parameter identifier, the first random number, and a sending terminal identifier, and sends it to the receiving terminal. After receiving the session establishment message, the receiving terminal generates a second random number, constructs a session confirmation message, which includes the second random number, the receiving terminal identifier, and a session parameter confirmation field, and sends it to the sending terminal. The sending terminal and the receiving terminal perform bidirectional authentication on the session establishment message and the session confirmation message respectively, and obtain bidirectional authentication output; The sending terminal identifier, receiving terminal identifier, first random number, second random number and the two-way authentication output are concatenated to form an input string. The input string is input into a key derivation function to generate key material, and the session key and the initial counter parameter are extracted from the key material. The sending terminal generates a key confirmation value and sends it to the receiving terminal. After verifying the key confirmation value, the receiving terminal generates an confirmation response and returns it to the sending terminal. The key confirmation value is obtained by calculating a message authentication code from the digests of the session establishment message and the session confirmation message. The message authentication code calculation uses the session key as the key.

[0008] Furthermore, when the sending terminal generates the first random number and constructs the session establishment message, it includes: Set a preset byte length for the first random number and generate it using a cryptographically secure random number generator; The capability field is set in the session establishment message. The capability field includes at least the set of supported FEC group sizes, the set of redundant packet numbers, the set of transmission unit header lengths, and the retransmission disable flag. The message establishment message for the session is configured with a message sequence number field and a timestamp field, and the message sequence number is maintained to be monotonically increasing on the sending terminal side. The fields of the session establishment message, excluding the authentication field, are concatenated to form a string to be authenticated. A first authentication value is calculated on the string to be authenticated using the private key of the identifier cryptosystem, and the first authentication value is written into the authentication field of the session establishment message.

[0009] Further, when concatenating the sending terminal identifier, receiving terminal identifier, the first random number, the second random number, and the two-way authentication output to form an input string, inputting the input string into a key derivation function to generate key material, and extracting the session key and the initial counter parameters from the key material, the process includes: A hash operation is performed on the input string to obtain a digest of a defined length; The determined length digest is input into the key derivation function to obtain the derived output, and the derived output is divided into an encryption key segment, an integrity key segment, and a counter seed segment; Use the encrypted key segment as the session key; The initial counter parameters are obtained by concatenating the counter seed segment with the session direction flag. When the sending or receiving end detects that the counter is about to roll back, it generates an updated count value, concatenates the updated count value with the determined length digest, inputs it into the key derivation function to obtain a new counter seed segment, and uses the new counter seed segment to update the initial counter parameters.

[0010] Furthermore, when performing end-side sensitive content detection on each acquired video frame before encoding, it includes: The video frames are scaled, pixel normalized, and channel arranged to obtain the detection input frames; The detection input frame is converted into a fixed-point representation and input into the end-side detection model, and a candidate region list and the detection confidence of the candidate regions are output. During continuous frame processing, the candidate region list of the previous frame is written into the volatile buffer, and the region position is extrapolated to the current frame to obtain the predicted region. When the predicted region overlaps with the candidate region, the overlapping regions are merged and the corresponding confidence records are updated.

[0011] Furthermore, when obtaining the set of sensitive regions, it includes: The candidate region list is overlapped and merged to output a merged region list. Each merged region records the rectangle coordinates, frame number and corresponding detection confidence. The list of merged regions is converted into a region index table, which uses the frame number as the key and the set of coordinates of the merged regions as the value. The region index table is written into volatile memory and bound to the current video frame to form the sensitive region set; When processing the next video frame, the memory space occupied by the region index table corresponding to the previous video frame is released.

[0012] Furthermore, when performing pixel-level masking processing on sensitive regions in the set of sensitive regions whose detection confidence is greater than or equal to a preset sensitivity threshold to obtain desensitized video frames, the process includes: Align the coordinates of each sensitive region to the boundary of the video coding macroblock to obtain the aligned masking window; A binary mask image is generated within the masking window, and the mask image is mapped to the luminance and chrominance components of the video frame; The occlusion window is divided into pixel sub-blocks of a preset size. A representative pixel value is calculated for each pixel sub-block and then filled back into all pixel positions of that pixel sub-block. The masked luminance and chrominance components are recombined to obtain the desensitized video frame.

[0013] Furthermore, when clearing the corresponding original video frame from the buffer without writing it to persistent storage, this includes: Write the original video frames into a volatile circular buffer; After obtaining the desensitized video frame, the memory area in the volatile circular buffer corresponding to the original video frame is overwritten and then the memory area is released. On the sending terminal side, the original video frames are prohibited from entering the file writing interface, system-level image cache interface, screen capture cache interface, and log disk writing interface.

[0014] Furthermore, when the media stream is divided into blocks of equal length to the SM4 packet length to obtain several plaintext blocks, it includes: The media stream is sequentially divided into block sequences, and a block number and a frame number are written for each block. When the last block is less than the length of an SM4 packet, padding bytes are appended to the end and the padding length field is written. A block header is generated for each block. The block header includes a session identifier, the block sequence number, the frame sequence number, and an original length field. The session identifier is obtained by concatenating the sending terminal identifier, the receiving terminal identifier, the first random number, and the second random number. When performing counter mode encryption, the initial counter parameter is concatenated with the block number to form a counter input, and the block number is incremented for different blocks to obtain different counter inputs.

[0015] Further, when performing forward error correction coding on several ciphertext blocks to generate an FEC redundancy packet, and encapsulating the several ciphertext blocks and the FEC redundancy packet into a transmission unit for transmission, the process includes: Divide a series of consecutive ciphertext blocks into FEC groups, and assign a group identifier and a block number within the group to each FEC group. For each FEC group, perform systematic erasure coding to generate a set of redundant packets corresponding to that FEC group. The number of redundant packets is determined by the link packet loss statistics read by the sender. A transmission unit header is generated for each ciphertext block and each redundant packet. The transmission unit header includes a session identifier, the group identifier, the block sequence number within the group, a payload type flag, and a timestamp field. The integrity key segment is used to calculate the message authentication code for the header and payload of the transmission unit, and the message authentication code is written into the verification field of the transmission unit before being sent.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: By performing two-way authentication and deriving session keys and initial counter parameters during the session establishment phase, the sending end and the receiving end can use consistent keys and counter parameters to complete the encryption and decryption of the media stream in SM4 counter mode; by performing end-side sensitive content detection on each video frame before encoding at the sending end and performing pixel-level masking on sensitive areas that meet the preset sensitivity threshold, while clearing the cache of the corresponding original video frame and not writing it to persistent storage, it is beneficial to reduce the risk of the original sensitive frames remaining on the terminal side; by performing forward error correction encoding on the encrypted ciphertext blocks and encapsulating them with redundant packets for transmission, the receiving end can use forward error correction decoding to recover the ciphertext block set and reassemble the media stream under packet loss conditions, thereby better adapting to the transmission requirements under narrowband weak network links. Attached Figure Description

[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart of a lightweight encryption and end-side sensitive content protection method for narrowband weak networks provided in an embodiment of the present invention. Detailed Implementation

[0018] Exemplary embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] See Figure 1 As shown, this application proposes a lightweight encryption and end-side sensitive content protection method for narrowband weak networks, including: S1: When the sending terminal and the receiving terminal establish a session connection, a two-way authentication is performed using an identifier cryptosystem, and a session key and initial counter parameters are derived based on the output of the two-way authentication. S2: When acquiring audio and video data on the sending terminal side, sensitive content detection is performed on each video frame before encoding to obtain a set of sensitive regions; pixel-level masking is performed on sensitive regions in the set of sensitive regions whose detection confidence is greater than or equal to a preset sensitive threshold to obtain desensitized video frames, and the corresponding original video frames are cleared from the buffer and not written to persistent storage. S3: Encode the desensitized video frames and audio data to obtain the media stream, and divide the media stream into blocks according to the SM4 block length to obtain several plaintext blocks; S4: The counter mode uses the national cryptographic SM4 algorithm to encrypt several plaintext blocks using the session key and initial counter parameters to generate corresponding ciphertext blocks; S5: Perform forward error correction coding on several ciphertext blocks to generate FEC redundant packets, and encapsulate several ciphertext blocks and FEC redundant packets into transmission units for transmission. The total bandwidth of the narrowband weak network link is less than or equal to the preset narrowband threshold, and the packet loss rate is greater than the preset packet loss threshold. S6: On the receiving terminal side, forward error correction decoding is performed on the received transmission unit to obtain a set of ciphertext blocks; the ciphertext block set is decrypted using the counter mode of the national cryptographic SM4 algorithm, using the session key and initial counter parameters to obtain plaintext blocks and reassemble them into a media stream; the media stream is then decoded to obtain a video frame sequence and an audio sampling sequence.

[0020] Specifically, before transmitting audio and video, the sending and receiving terminals establish a session connection. This session connection can be understood as an end-to-end communication context used to agree on the key parameters required for subsequent data encryption and decryption. During this stage, both parties use an identifier cryptosystem for two-way authentication. The identifier cryptosystem is a type of cryptosystem that uses the identification information of both communicating parties (such as device identifier, user identifier, or account identifier) ​​as elements related to the public key. Two-way authentication means that the sending terminal verifies the identity of the receiving terminal and the receiving terminal verifies the identity of the sending terminal. After authentication, both parties derive the session key and initial counter parameters from the output of the two-way authentication. The session key is used for symmetric encryption and decryption of media data during this session. The initial counter parameters are used to generate the counter input corresponding to each data block in counter mode. The initial counter parameters may include the counter start value, the session direction flag, and components associated with the data block sequence number, so that the sending and receiving terminals can use the same counter input to generate a consistent key stream for the same data block.

[0021] Subsequently, when acquiring audio and video data at the sending terminal, the video data is typically generated frame by frame, such as each frame captured by a camera. Before encoding the video frame, edge-side sensitive content detection is performed. Edge-side means that the detection is completed locally at the sending terminal rather than relying on network-side processing. The output of sensitive content detection is a set of sensitive regions, which describes the area that needs to be protected in the current video frame. It is usually expressed as region coordinates, region size, etc., and a detection confidence score is given for each region. The detection confidence score is a numerical measure of whether the region is classified as sensitive content, and its value can be a real number from 0 to 1 or an integer from 0 to 100. When the edge-side detection model uses binary classification output, the detection confidence score can be obtained from the sigmoid output of the last layer of the model. When the edge-side detection model uses multi-class classification output, the detection confidence score can be obtained from the probability value corresponding to the sensitive content category in the softmax output or mapped from that probability value. The preset sensitivity threshold is used as the boundary for determining the detection confidence level. It is defined as the set of sensitive regions that need to be masked when the detection confidence level reaches the threshold. The preset sensitivity threshold is determined in advance during the model release or version solidification stage. Specifically, it can be obtained by performing inference statistics on the end-side detection model on the validation dataset to obtain the detection status of candidate regions under different candidate thresholds, and selecting one from the candidate threshold set as the preset sensitivity threshold. The candidate threshold set can be a set of values ​​discrete in the interval of 0 to 1 with a step size or a set of values ​​discrete in the interval of 0 to 100 with a step size. The preset sensitivity threshold can be distributed as a model parameter along with the model file, or stored as a configuration parameter in the configuration area of ​​the sending terminal and read when the detection module loads the model. The configuration record can be associated with the model version number and model verification summary for storage, so that the terminal completes the consistency verification before reading the preset sensitivity threshold when starting the detection module.

[0022] Sensitive regions in the sensitive region set whose detection confidence is greater than or equal to a preset sensitivity threshold are subjected to pixel-level masking. Pixel-level masking is a processing method that directly replaces or rewrites regions at the image pixel level. For example, it can perform block mosaic, mean backfilling, or other masking operations on the pixels in the region, so that the output frame becomes a desensitized video frame, that is, frame data in which the sensitive regions have been masked at the image content level. After generating the desensitized video frame, the corresponding original video frame is cleared from the buffer and not written to persistent storage. The buffer can be understood as the buffer space in the terminal's memory used to temporarily store the acquired frames. Not writing to persistent storage means not writing to the local file system or other storage media that can be stored for a long time.

[0023] Next, the desensitized video frames and audio data are encoded to obtain the media stream. Encoding is the process of converting the original audio and video samples into a compressed stream. The media stream usually contains video encoded data, audio encoded data, and time-series information such as timestamps. Then, the media stream is divided into blocks aligned with the SM4 block length. SM4 is a block cipher algorithm, and its block length is the basic data block length processed by the algorithm. Aligned block division means that the continuous stream is divided into several plaintext blocks that match the block length, which facilitates subsequent unified encryption processing.

[0024] Subsequently, the plaintext block is encrypted using the counter mode of the national cryptographic SM4 algorithm. The counter mode is a working method that uses the session key and the counter input to generate a key stream and combine it with the plaintext block to obtain the ciphertext block. The sending end uses the session key and the initial counter parameters and combines them with the block number of the plaintext block to form the counter input for each block, generating a number of corresponding ciphertext blocks.

[0025] To adapt to narrowband and weak network conditions, forward error correction (FEC) redundancy packets are generated by forward error correction coding on the transmission side. Forward error correction is a coding method that adds redundant information to the data so that the receiver can still recover the original data block even if a certain number of data blocks are lost. The FEC redundancy packet is a redundant data unit generated by error correction coding. Subsequently, the ciphertext block and the redundancy packet are encapsulated into a transmission unit and sent. The transmission unit can be understood as the basic carrier of network transmission, which includes a header field and a payload field. The payload carries the ciphertext block or the redundancy packet, and the header can carry information such as session identifier, sequence number, and timestamp for reassembly and error correction. The total bandwidth of the narrowband and weak network link is limited to less than or equal to a preset narrowband threshold and the packet loss rate is greater than a preset packet loss threshold to characterize the link conditions.

[0026] After receiving the transmission unit, the receiving terminal first performs forward error correction decoding to obtain a set of ciphertext blocks. That is, it uses the received ciphertext blocks and FEC redundancy packets to complete error correction recovery and fill in the lost ciphertext blocks. Then, it uses the counter mode of the national cryptographic SM4 algorithm, using the session key and initial counter parameters to decrypt the set of ciphertext blocks to obtain plaintext blocks, and reassembles them into a media stream according to the block sequence number. Finally, it decodes the media stream, decoding the video part into a video frame sequence and the audio part into an audio sampling sequence. It can also complete the audio and video timing alignment according to the timestamp information in the media stream before outputting it. This forms a complete processing link from session establishment, end-side detection and masking, block encryption, FEC encapsulation and transmission to receiving-end error correction recovery, decryption and reassembly and decoding output.

[0027] In some embodiments of this application, when establishing a session connection between the sending terminal and the receiving terminal, a cryptographic system is used for two-way authentication, and when deriving the session key and initial counter parameters based on the output of the two-way authentication, the process includes: The sending terminal generates a first random number and constructs a session establishment message. The session establishment message includes a protocol version number, an encryption mode identifier, an FEC parameter set identifier, a media encoding parameter identifier, a first random number, and a sending terminal identifier, and sends it to the receiving terminal. After receiving the session establishment message, the receiving terminal generates a second random number, constructs a session confirmation message, which includes the second random number, the receiving terminal identifier, and the session parameter confirmation field, and sends it to the sending terminal. The sending terminal and the receiving terminal perform bidirectional authentication on the session establishment message and the session acknowledgment message respectively, and obtain the bidirectional authentication output; The sending terminal identifier, receiving terminal identifier, first random number, second random number and two-way authentication output are concatenated to form an input string. The input string is then input into the key derivation function to generate key material, and the session key and initial counter parameters are extracted from the key material. The sending terminal generates a key confirmation value and sends it to the receiving terminal. After verifying the key confirmation value, the receiving terminal generates an confirmation response and returns it to the sending terminal. The key confirmation value is obtained by calculating the message authentication code from the digests of the session establishment message and the session confirmation message. The message authentication code calculation uses the session key as the key.

[0028] Specifically, during the session connection establishment phase, the sending terminal first generates a first random number, which serves as the random factor for this session. This ensures that the key material inputs differ between different sessions, thereby avoiding the risk of different sessions deriving the same key. Subsequently, the sending terminal constructs a session establishment message and sends it to the receiving terminal. The protocol version number in the session establishment message indicates the message format and field semantics followed by both parties, facilitating compatibility parsing between different version implementations. The encryption mode identifier indicates the encryption mode used for subsequent media data, such as whether a counter mode is used and the corresponding algorithm family category. The FEC parameter set identifier indicates the forward error correction coding configuration set to be used by the sending end, such as the parameter set index corresponding to FEC group size, number of redundant packets, or erasure coding type. The media encoding parameter identifier indicates the encoding configuration of the media bitstream, such as the parameter set index of video encoding type, resolution level, frame rate level, or audio sampling rate level. The sending terminal identifier is used to identify the message initiator and serves as one of the authentication inputs under the identification cryptosystem.

[0029] After receiving the session establishment message, the receiving terminal generates a second random number and constructs a session confirmation message to return to the sending terminal. The second random number serves as a random factor introduced by the receiving end, ensuring that session randomness is provided jointly by both parties, thus reducing the risk of session security degradation caused by unilateral random number anomalies. The session confirmation message includes the receiving terminal identifier and a session parameter confirmation field. The session parameter confirmation field is used to confirm or select the capabilities and parameter sets carried in the session establishment message. For example, it confirms whether the protocol version number is supported, whether the encryption mode identifier is supported, and confirms the consistency of the FEC parameter set identifier and media encoding parameter identifier or selects an available parameter set identifier, thereby forming a consistent session parameter set between the two parties.

[0030] Subsequently, the sending terminal and the receiving terminal perform bidirectional authentication on the session establishment message and the session confirmation message respectively to obtain bidirectional authentication output. The bidirectional authentication output can be understood as the shared result obtained in the authentication calculation or the authentication material that can be consistently verified by both parties. It is at least used to prove that the other party's identifier and message content have not been tampered with and match the corresponding private key holder. In implementation, the sending terminal can perform authentication verification on the session confirmation message returned by the receiving terminal, and the receiving terminal can perform authentication verification on the session establishment message sent by the sending terminal. In the authentication calculation, the identifiers of both parties, random numbers and key parameter fields are included in the data to be authenticated to ensure that the authentication is bound to the parameters of this session.

[0031] After completing two-way authentication, the sending terminal identifier, receiving terminal identifier, first random number, second random number, and two-way authentication output are concatenated to form an input string. This concatenation involves stitching the elements together into a continuous bit or byte string according to a preset field order and encoding format. For example, the identifier field may use fixed-length or variable-length encoding, the random number may use a preset byte length, and the authentication output may use a preset length encoding. This input string is then fed into a key derivation function to generate key material. The key derivation function maps the input string to a derived output of the required length and possesses scalability and irreversibility, allowing the extraction of key fragments for different purposes from the same derived output. Subsequently, the session key and initial counter parameters are extracted from the key material. The session key is used for subsequent encryption and decryption in the SM4 counter mode, while the initial counter parameters determine the starting counting state of the counter mode and the basis for constructing the counter input. The initial counter parameters may include a counter start value and a session direction flag to distinguish the counter sequences used in the sending and receiving directions.

[0032] To ensure that both parties derive consistent session keys, the sending terminal generates a key confirmation value and sends it to the receiving terminal. The key confirmation value is obtained by calculating a message authentication code from the digests of the session establishment message and the session confirmation message. The digest can be a hash operation to obtain a fixed-length result by concatenating the key fields of the two types of messages in a preset order. The message authentication code calculation uses the session key as the key, so that only the party holding the same session key can calculate a consistent key confirmation value. After receiving the key confirmation value, the receiving terminal uses its own derived session key to perform message authentication code calculation on the same digest and compares and verifies it. After successful verification, it generates an confirmation response and returns it to the sending terminal, thus completing the closed-loop confirmation of parameter negotiation, identity authentication, key derivation, and key consistency during the session establishment phase.

[0033] In some embodiments of this application, when the sending terminal generates a first random number and constructs a session establishment message, it includes: Set a preset byte length for the first random number and generate it using a cryptographically secure random number generator; Set the capability field in the session establishment message. The capability field should include at least the set of supported FEC group sizes, the set of redundant packet quantities, the set of transmission unit header lengths, and the retransmission disable flag. Set the sequence number field and timestamp field for the session establishment message, and maintain the monotonically increasing sequence number on the sending terminal side; Concatenate all fields of the session establishment message except the authentication field to form the string to be authenticated. Calculate the first authentication value of the string to be authenticated using the private key of the identifier cryptosystem, and write the first authentication value into the authentication field of the session establishment message.

[0034] Specifically, when generating the first random number, the sending terminal first presets the byte length of the first random number. This preset byte length is used to constrain the value space and encoding format of the random number, so as to keep the field length consistent during message parsing and subsequent key derivation input string construction. Then, the sending terminal calls a cryptographically secure random number generator to generate the first random number. The cryptographically secure random number generator is a random source implementation that meets the requirements of unpredictability and statistical randomness, which makes the generated first random number difficult to predict even given a known output.

[0035] When constructing a session establishment message, the sending terminal sets a capability field in the message to declare the set of capabilities it can support. Among them, the set of supported FEC group sizes describes the forward error correction group granularity options that the sending terminal can adopt, such as the candidate set corresponding to the number of ciphertext blocks contained in each FEC group. The set of redundant packet numbers describes the option of the number of redundant packets that can be generated for each FEC group. The set of transmission unit header lengths describes the range of bytes occupied by the transmission unit header or the candidate length set that the sending terminal can support, so that the receiving terminal can select a matching header format when parsing and buffering. The retransmission disable flag is used to indicate that the sending terminal does not enable the acknowledgment-based retransmission mechanism in this session. This flag is used to make the receiving terminal clearly aware that the session adopts a transmission strategy that relies solely on forward error correction recovery when confirming session parameters, thereby avoiding state inconsistencies caused by waiting for retransmission acknowledgments on the receiving end side.

[0036] To improve the consistency of messages during transmission and processing in the link, the sending terminal also sets a message sequence number field and a timestamp field for the session establishment message. The message sequence number field is used to identify the sending order of the session establishment message on the sending end side. The sending terminal maintains the message sequence number in a monotonically increasing manner, meaning that the message sequence number is incremented by a preset step size every time a session establishment message or session-related control message is sent, so that the receiving end can distinguish between duplicate messages and out-of-order messages. The timestamp field is used to record the time information of the generation or sending of the session establishment message. It can be represented by a local clock count value, system time, or relative time count. It is used by the receiving end to identify the newness of the message in session management and is referenced in subsequent session timeout judgment.

[0037] To bind the session establishment message to the sending terminal's identity and message content, the sending terminal concatenates all fields in the session establishment message except for the authentication field to form a string to be authenticated. This string can be formed by concatenating fields in a preset order, such as the protocol version number, encryption mode identifier, FEC parameter set identifier, media encoding parameter identifier, first random number, sending terminal identifier, capability field, message sequence number field, and timestamp field. Variable-length fields are prefixed with a length or encoded using TLV to avoid parsing ambiguity. Then, the sending terminal uses its private key from the identifier cryptosystem to calculate the string to be authenticated. The first authentication value, which can be a digital signature or a verifiable authentication tag, is calculated by taking the sending terminal's private key and the string to be authenticated as input. This allows the receiving terminal to verify the first authentication value using the public key information corresponding to the sending terminal's identifier after receiving the session establishment message. Finally, the sending terminal writes the first authentication value into the authentication field of the session establishment message. This ensures that even if the session establishment message is intercepted or tampered with during transmission, it will still be recognized by the receiving terminal due to the authentication field failing verification, thus completing the expression and constraint of the integrity of the session establishment message content and the association of the sending terminal's identity.

[0038] In some embodiments of this application, when concatenating the sending terminal identifier, receiving terminal identifier, first random number, second random number, and two-way authentication output to form an input string, inputting the input string into a key derivation function to generate key material, and extracting the session key and initial counter parameters from the key material, the process includes: A hash operation is performed on the input string to obtain a digest of a fixed length; The derivation function is obtained by inputting a determined length digest into the key, and the derivation output is divided into an encryption key segment, an integrity key segment, and a counter seed segment. Use the encrypted key segment as the session key; The initial counter parameters are obtained by concatenating the counter seed segment with the session direction flag. When the sending or receiving end detects that the counter is about to roll back, it generates an updated count value, concatenates the updated count value with a defined length digest, inputs it into the key derivation function to obtain a new counter seed segment, and uses the new counter seed segment to update the initial counter parameters.

[0039] Specifically, after completing two-way authentication, the sending terminal and receiving terminal first concatenate the sending terminal identifier, receiving terminal identifier, first random number, second random number, and two-way authentication output according to a preset field order to form an input string. This input string is used to simultaneously bind the identities of the session participants, the randomness of the session, and the authentication result to the subsequent key derivation process. The sending terminal identifier and receiving terminal identifier are used to distinguish the two parties in the session and prevent different terminal combinations from generating the same derived input. The first random number and the second random number are used to provide a source of randomness for this session. The two-way authentication output is used to bind the derivation result to the authentication session, reducing the risk of the authentication material and the derived material becoming uncoordinated. In implementation, the input string can be formed by concatenating byte strings, and each field can be encoded using a defined rule. For example, the identifier field can be encoded using UTF-8 or binary encoding, the random number can be represented using a preset byte length, and the two-way authentication output can be encoded using a preset length or with a length prefix to avoid ambiguity at field boundaries.

[0040] Since the length of the input string may vary depending on the identifier length or authentication output format, a first cryptographic process is performed on the input string to facilitate subsequent key derivation function processing, resulting in a first intermediate value. This first intermediate value is data of a fixed length. Optionally, the first cryptographic process includes hash operations or cryptographic expansion function processing. As a preferred approach, the first cryptographic process is an SM3 hash operation. The hash operation can employ the national cryptographic hash algorithm or other hash algorithms with a fixed output length. Its fixed-length digest has the characteristics of a fixed byte length, sensitivity to input changes, and irreversibility, ensuring that the input for subsequent key derivation has a uniform length and good diffusivity.

[0041] The determined-length digest is then input into the key derivation function to obtain the derived output. Optionally, the key derivation function is the key derivation function KDF(Z,klen) specified in GB / T 32918.3-2016, where the hash function used by KDF is the SM3 cryptographic hash algorithm specified in GB / T32905-2016. The key derivation function is used to expand the fixed-length digest to the required total length of key material while maintaining statistical independence between different output segments. In implementation, it can adopt an iterative derivation method with a counter or a labeled expansion method, so that the same digest can stably produce the same derived output. The derived output is segmented to obtain at least two key segments. The key segments include at least an encryption key segment for encryption and a counter seed segment for counter construction. Optionally, the key segments also include an integrity key segment for integrity protection, where the encryption key segment is used as the key input for the symmetric encryption algorithm, the integrity key segment is used for integrity verification calculations such as message authentication codes, and the counter seed segment is used to construct the initial counter parameters for the counter mode.

[0042] After using the encryption key segment as the session key, the session key becomes the shared key for the sender and receiver to perform SM4 counter mode encryption and decryption within the same session. The initial counter parameters are obtained by concatenating the counter seed segment with the session direction flag, where the session direction flag is used to distinguish between the sending and receiving directions, ensuring that different counter sequences are used in opposite directions within the same session, and avoiding the risk of key stream reuse due to the reuse of the same counter input in bidirectional data streams under the same key. In implementation, the initial counter parameters can be obtained by concatenating the counter start value and the direction flag, or by generating the counter start value from the counter seed segment after formatting, and then writing the direction flag into preset bits to form a set of parameters that can be directly used to construct the counter input.

[0043] In counter mode, the counter input increments with the block sequence number. When the number of consecutively sent or received data blocks increases, there is a possibility that the counter space will be exhausted. Therefore, when the sending or receiving end detects that the counter is about to roll back, it is necessary to update the relevant counter parameters to avoid counter duplication. The detection that the counter is about to roll back can be achieved by comparing the remaining available count range between the current counter value and the maximum counter value. For example, when the remaining count range is less than the preset safety margin, an update is triggered.

[0044] When an update is triggered, an update count is generated. This update count can be an update counter maintained within the session, a rotation number, or a renegotiation flag, used to indicate a change in the key derivation input. The update count is a monotonically increasing rotation number within the session. When the sender triggers a rollback update, it increments the rotation number and writes it into a control message, sending it to the receiver. Upon receiving the rotation number, the receiver uses the same rotation number to generate a new counter seed and update the initial counter parameters. The update count is then concatenated with a defined-length digest and input into the key derivation function again to obtain a new counter seed. The concatenation method also uses a preset encoding rule to ensure consistency between the two parties. For example, the update count can be used as a prefix or suffix to be concatenated into the defined-length digest to form a new derivation input. Because the derivation input has changed, the output of the key derivation function changes accordingly, resulting in a new counter seed. The new counter seed is then used to update the initial counter parameters, causing subsequent counter inputs to continue increasing from the new starting state. As long as the sender and receiver execute the above update process under the same triggering condition and the same update count generation rule, the consistency of the counter parameter updates can be maintained, ensuring that subsequent encryption and decryption can still align with the same counter sequence.

[0045] In some embodiments of this application, when performing end-side sensitive content detection on each acquired video frame before encoding, the process includes: The video frames are scaled, pixel normalized, and channel arranged to obtain the detection input frames; The detection input frame is converted into a fixed-point representation and input into the edge detection model, and the output is a list of candidate regions and the detection confidence of the candidate regions. During continuous frame processing, the candidate region list of the previous frame is written into the volatile buffer, and the region position is extrapolated to the current frame to obtain the predicted region. When the predicted region overlaps with the candidate region, the overlapping regions are merged and the corresponding confidence records are updated.

[0046] Specifically, before performing edge-side sensitive content detection on each acquired video frame, the transmitting terminal preprocesses the video frame to meet the input specifications of the edge-side detection model. Size scaling transforms the original video frame from the acquisition resolution to the input resolution required by the detection model, for example, scaling from a higher resolution to the model's fixed input size or scaling to the upper limit size according to a preset ratio, to reduce inference computation and maintain consistent input dimensions. Pixel normalization transforms the scaled pixel values ​​according to preset rules, such as mapping integer pixels from 0 to 255 to floating-point pixels from 0 to 1, or performing mean offset and scale scaling on each channel to ensure the input data distribution remains consistent with the model training phase. Channel arrangement adjusts the channel order and storage layout of the image data, for example, converting the RGB order to the BGR order required by the model, or converting the storage format by height, width, and channel to a storage format by channel, height, and width, thus forming the detection input frame.

[0047] Optionally, the transmitting terminal converts the detection input frame into a fixed-point representation. To further reduce the computational overhead at the edge, the transmitting terminal converts the detection input frame into a fixed-point representation. Fixed-point representation maps floating-point values ​​to integer representations according to a preset quantization scale, such as using 8-bit or 16-bit fixed-point quantization. The normalized pixel values ​​are scaled and rounded according to the quantization ratio coefficient, allowing the edge detection model to perform inference on the integer operation path. Based on this, the transmitting terminal inputs the fixed-point represented detection input frame into the edge detection model for inference. The edge detection model can be a target detection or instance segmentation model. For example, the edge detection model can be a pruned and quantized object detection model. The model can be a YOLO series model or a lightweight segmentation model. Optionally, the edge detection model is obtained and updated from the management platform via secure OTA. The model file, preset sensitivity thresholds, and model version number are associated and stored, and integrity verification is performed during loading. Its output includes a candidate region list and the detection confidence score of each candidate region. The candidate region list describes the location and range of regions that may contain sensitive content, for example, represented by the coordinates of the top left corner, bottom right corner, or center point of a rectangle, along with width and height parameters. The detection confidence score indicates the degree of confidence that the candidate region belongs to the sensitive content category, facilitating subsequent screening and processing of candidate regions. In some embodiments, if the edge detection model fails to load or the integrity verification fails, the sensitive content detection module is stopped and an error status flag is generated. Optionally, the sending terminal performs full-frame masking processing on the current video frame before entering the encoding process, or terminates video transmission and returns an error code to the upper layer.

[0048] For continuous video frame processing, in order to reduce the detection fluctuations of repeated regions in adjacent frames and maintain the consistency of region tracking, the transmitting terminal writes the candidate region list of the previous frame into a volatile buffer. The volatile buffer is a temporary data area stored in memory, which can be overwritten or released after entering the next frame processing and is not persistently saved. When processing the current frame, the region position is extrapolated to obtain the predicted region based on the position and size information of the candidate region in the previous frame. The region position extrapolation is the process of calculating the region coordinates of the previous frame according to a preset motion model. For example, it is assumed that the region position changes little in a short period of time and the coordinates of the previous frame are directly used, or the coordinates of the current frame are calculated based on the coordinate changes between the previous frame and an earlier frame. The predicted region is used to represent the possible location range of sensitive regions in the current frame.

[0049] Then, the candidate region output by the current frame edge detection model is compared with the predicted region for overlap judgment. Overlap judgment can be achieved by calculating the intersection-union ratio of the two regions or judging the overlap of coordinate ranges. When the predicted region and the candidate region overlap, the overlapping regions are merged and the corresponding confidence records are updated. Merging can be represented by taking the union boundary of the overlapping regions or using the boundary of the region with higher confidence as the merging result. Updating the confidence records can be represented by retaining a larger confidence for the same merged region, performing a weighted average, or fusing the confidence according to preset rules, thereby forming a more stable candidate region representation for the current frame.

[0050] In some embodiments of this application, obtaining the set of sensitive regions includes: The candidate region list is overlapped and merged, and the merged region list is output. Each merged region records the rectangle coordinates, frame number and corresponding detection confidence. The list of merged regions is converted into a region index table, with the frame number as the key and the coordinate set of the merged regions as the value. Write the region index table into volatile memory and bind it to the current video frame to form a set of sensitive regions; When processing the next video frame, the memory space occupied by the region index table corresponding to the previous video frame is released.

[0051] Specifically, after the sending terminal outputs the candidate region list from the edge detection model, it needs to organize the candidate regions into a data structure that can be directly used for occlusion processing. Therefore, the candidate region list is first overlapped and merged to eliminate duplicate boxes or highly overlapping boxes generated by the same sensitive target during the detection process.

[0052] Optionally, the overlapping merging can be performed according to the preset regional overlap determination rules, such as calculating the intersection-union ratio of any two candidate regions or determining whether their rectangular coordinate ranges intersect. When the overlap condition is met, they are regarded as candidate descriptions of the same sensitive target. When merging, the bounding rectangle of each candidate region can be taken as the rectangular coordinates of the merged region, or the rectangular coordinates of the candidate regions with higher confidence can be retained and the boundary can be appropriately expanded so that the merged region can cover all overlapping candidate regions.

[0053] After merging, a list of merged regions is output. Each merged region records at least the rectangle coordinates, frame number, and corresponding detection confidence score. The rectangle coordinates are used to clarify the position and range of the region in the video frame, the frame number is used to identify the video frame to which the region belongs, and the detection confidence score is used to characterize the credibility of the merged region being judged as sensitive content. For a merged region obtained by merging multiple candidate regions, its detection confidence score can be the maximum confidence score of the merged candidate regions, the area-weighted average confidence score, or obtained according to a preset fusion rule, so as to maintain consistency with the subsequent sensitivity threshold determination.

[0054] The list of merged regions is then converted into a region index table. The region index table is a data organization format that facilitates fast lookup. It uses the frame number as the key, indicating which video frame the index item corresponds to; and the coordinate set of the merged regions as the value, representing the coordinate set of all merged regions corresponding to that frame. Each item in the coordinate set can contain rectangular coordinate parameters and the detection confidence associated with those rectangular coordinates. If necessary, it can also include a region number or region category identifier for subsequent processing. The region index table is written to volatile memory and bound to the current video frame, meaning that the region index table only exists within the processing lifecycle of the current video frame and is associated with the data flow process of the current video frame. For example, when the masking module reads the current video frame, it synchronously reads the coordinate set corresponding to the frame number, thus forming a sensitive region set. Logically, the sensitive region set consists of one or more sensitive regions corresponding to the current video frame, along with their coordinates and confidence information. In some embodiments, optionally, the sending terminal generates and reports an audit data packet after forming a sensitive area set: the sending terminal extracts audit fields from the sensitive area set, the audit fields including at least frame sequence number, timestamp, set of rectangular coordinates of the sensitive area, detection confidence, and sensitive category identifier; optionally, the audit fields also include session identifier, end-side detection model version identifier, and masking processing parameter identifier. The sending terminal serializes and encodes the audit fields to obtain audit plaintext, and encrypts the audit plaintext using an audit key to obtain audit ciphertext; wherein, the audit key may optionally be extracted from key materials, or derived from an integrity key segment. The sending terminal calculates a message authentication code for the audit ciphertext and writes it into the audit data packet verification field; optionally, the audit data packet is sent to the audit server through a control channel independent of media transmission, or encapsulated in a reserved extension field of the transmission unit and identified as an audit payload by a payload type flag.

[0055] To avoid sensitive area information residing on the terminal side for a long time and to reduce memory usage, the memory space occupied by the area index table corresponding to the previous video frame is released when the next video frame is processed. This release can be manifested by deleting the index entry corresponding to the previous frame number or overwriting it with an invalid value. This allows the sensitive area set to be updated as the video frame processing progresses, ensuring that only the necessary temporary area information within the current frame or preset window is retained, and that this area information is not written to persistent storage.

[0056] In some embodiments of this application, when performing pixel-level masking processing on sensitive regions in the sensitive region set whose detection confidence is greater than or equal to a preset sensitivity threshold to obtain desensitized video frames, the process includes: Align the coordinates of each sensitive region to the boundary of the video coding macroblock to obtain the aligned masking window; A binary mask is generated within the masking window, and the mask is mapped to the luminance and chrominance components of the video frame. The masking window is divided into pixel sub-blocks of a preset size. The representative pixel value of each pixel sub-block is calculated and backfilled to all pixel positions of that pixel sub-block. The masked luminance and chrominance components are recombined to obtain the desensitized video frame.

[0057] Specifically, after determining the set of sensitive regions, the sending terminal first performs pixel-level masking processing on the sensitive regions whose detection confidence is greater than or equal to a preset sensitivity threshold, so as to generate desensitized video frames that can enter the subsequent encoding process.

[0058] First, the coordinates of each sensitive region are aligned to the boundaries of the video coding macroblocks. Macroblock boundaries are the basic block boundaries used for transform, prediction, and quantization processes in video coding. Macroblock boundaries are the basic block boundaries used for block partitioning, prediction, and transform processes within the encoder. The macroblock size is determined by the encoder configuration. The alignment operation involves expanding the rectangular coordinates of the sensitive region outward or clipping them inward to the nearest macroblock boundary position, so that the coordinates of the upper left and lower right corners of the masking window fall on the macroblock boundary. This ensures that the masked region is consistent with the coding block structure, avoiding the increase in processing complexity or the introduction of additional boundary discontinuities caused by the masking edges crossing multiple coding blocks.

[0059] After alignment, an aligned masking window is obtained. The masking window is used to limit the scope of the masking process on the video frame. Subsequently, a binary mask is generated within the masking window. The binary mask is used to identify which pixel positions within the masking window need to be masked. It is usually represented by two values, 0 and 1. Positions with a value of 1 correspond to the pixel areas that need to be masked, and positions with a value of 0 correspond to areas that retain their original pixels. The mask can be generated by directly filling the rectangular area of ​​the masking window, or a more closely fitting mask can be generated within the rectangular area according to the shape of the sensitive area or edge expansion rules.

[0060] Since video frames are typically represented in a color space with separate luminance and chrominance components before encoding, such as the YUV format, where the luminance component represents luminance information and the chrominance component represents color information, mapping the mask to the luminance and chrominance components means aligning and scaling the same mask on the sampling grids of different components. For example, in the case of chrominance subsampling, the resolution of the chrominance component is lower than that of the luminance component. Therefore, the mask needs to be downsampled or its coordinates converted on the chrominance component according to the corresponding sampling ratio so that the mask covers the same spatial area on both components.

[0061] Next, the masking window is divided into pixel sub-blocks of a preset size. The preset size is used to limit the width and height of the sub-blocks, for example, using a block size of several pixels × several pixels, so that the masking process can be performed on a block-by-block basis. For each pixel sub-block, a representative pixel value is calculated and backfilled into all pixel positions of the pixel sub-block. The representative pixel value can be the mean, mode, median, or a pixel value selected according to a preset rule within the sub-block. Optionally, the representative pixel value is determined by the center pixel value of the sub-block or by a preset quantile of the pixel values ​​within the sub-block. Backfilling means writing the representative pixel value into all pixel positions within the sub-block, thereby forming a uniform pixel texture at the sub-block scale and achieving the masking effect on the original detail information. This backfilling process is applied to both the luminance and chrominance components, so that the masking is effective in both luminance and chrominance aspects simultaneously.

[0062] Finally, the masked luminance and chrominance components are recombined. The combination is carried out by combining the components according to their corresponding sampling relationships in the color space format into a single video frame data, such as reorganizing it into YUV planar or half-planar format data, thereby obtaining a desensitized video frame. This desensitized video frame has performed masking processing on sensitive areas that meet the threshold conditions at the pixel level and maintains the same input format as the subsequent encoder, so it can be directly entered into the media encoding and transmission processing link.

[0063] In some embodiments of this application, when the corresponding original video frame is cleared from the buffer and not written to persistent storage, the following steps are included: Write the original video frames into a volatile circular buffer; After obtaining the desensitized video frame, the memory area corresponding to the original video frame in the volatile circular buffer is overwritten and then the memory area is released. On the sending terminal side, raw video frames are prohibited from entering the file writing interface, system-level image cache interface, screen capture cache interface, and log disk writing interface.

[0064] Specifically, to ensure that the original video frames exist only as temporary data on the sending terminal side and do not form a data copy that can be retained for a long time, the sending terminal writes the acquired original video frames into a volatile circular buffer.

[0065] A volatile circular buffer is a circular buffer structure that resides in memory. Its storage space is allocated to a preset capacity during initialization and is reused in a circular manner by writing pointers to overwrite sequentially. Because the buffer is located in volatile memory, its contents are not retained when the device is powered off or the process ends. Furthermore, the circular structure ensures that older frames are automatically overwritten by subsequent new frames, thereby limiting the residence time of the original frames in memory.

[0066] After the original video frame is written to the circular buffer, it can be read by the detection module for on-side sensitive content detection and subsequent masking processing. When a corresponding desensitized video frame is generated based on the original video frame, the sending terminal overwrites and releases the memory area in the circular buffer corresponding to the original video frame. Overwriting refers to writing irrelevant data to the memory area according to a preset overwriting mode, such as writing all zeros, all ones, or pseudo-random byte sequences, to eliminate residual traces of the original pixel data in memory. Releasing the memory area means marking the memory area as reusable, so that the write pointer of the circular buffer can subsequently occupy the area to store new original frames, thereby terminating the reference to the original frame data at the memory management level.

[0067] To prevent the original video frames from being indirectly stored in other system paths, the sending terminal also prohibits the original video frames from entering the file write interface, system-level image cache interface, screen capture cache interface, and log disk persistence interface. The file write interface includes interface call paths that write frame data to the local file system, database, or other persistent media; the system-level image cache interface includes interface call paths that may cause image data to be cached, such as image caches, thumbnail caches, or media preview caches maintained by the operating system or graphics subsystem; the screen capture cache interface includes cache paths that may temporarily store image content during screen recording, screenshotting, or screen mirroring; and the log disk persistence interface includes interface paths that directly write frame content or binary data related to frame content to a log file. In some embodiments, optionally, the sending terminal sets up a security middleware or operating system access control policy to intercept and filter persistent storage access requests from the process to which the video capture module belongs, intercepting at least the file write interface, media library write interface, and log disk persistence interface; when a write request related to the original video frame data is detected, the write request is rejected and an error code is returned.

[0068] With the above constraints, the lifespan of the original video frame on the terminal side is limited to a short window of volatile memory, and it is cleared by overwrite and release reuse after the desensitized video frame is generated. At the same time, the original frame is prevented from being copied through persistent or system cache channels by the interface-level prohibition policy.

[0069] In some embodiments of this application, when the media stream is divided into blocks of length equal to the SM4 packet length to obtain several plaintext blocks, the following methods are included: The media stream is sequentially divided into block sequences, and a block number and a frame number are written for each block. When the last block is less than the length of an SM4 packet, padding bytes are appended to the end and the padding length field is written. A block header is generated for each block. The block header contains a session identifier, a block sequence number, a frame sequence number, and an original length field. The session identifier is obtained by concatenating the sending terminal identifier, the receiving terminal identifier, a first random number, and a second random number. When performing counter mode encryption, the initial counter parameter is concatenated with the block number to form the counter input, and the block number is incremented for different blocks to obtain different counter inputs.

[0070] Specifically, the media stream is sequentially segmented to obtain a block sequence. Sequential segmentation means that data is continuously retrieved from front to back according to the byte order of the media stream in memory. Each segment retrieved is a block with a length equal to the length of an SM4 packet. To facilitate reconstruction and positioning at the receiving end, each block is written with a block number and a frame number. The block number is used to identify the position of the block in the block sequence and keeps monotonically increasing. The frame number is used to identify the media frame or the encoded frame group to which the block belongs, so that the receiving end can recover the corresponding video frame sequence according to the frame boundaries after reconstructing the media stream.

[0071] If the number of bytes remaining at the end of the media stream is less than the length of one SM4 packet, padding bytes are appended to the end to make it the length of the SM4 packet, and written into the padding length field. The padding length field is used to indicate the number of appended padding bytes. After decryption and reassembly, the receiving end can use this field to remove the corresponding padding bytes from the end, thereby restoring the true byte length of the media stream.

[0072] A block header is generated for each block. The block header carries the identifier and length information related to the block. The block header includes at least a session identifier, a block sequence number, a frame sequence number, and an original length field. The session identifier is obtained by concatenating the sending terminal identifier, the receiving terminal identifier, a first random number, and a second random number. The concatenation method can be to concatenate the fields in a preset order and apply a deterministic encoding rule to each field. For example, the identifier field can use fixed-length encoding or encoding with a length prefix, and the random number can use encoding with a preset byte length to ensure that the construction results of the session identifier by the sending end and the receiving end are consistent. The original length field is used to indicate the number of valid plaintext bytes corresponding to the block. In particular, when there is padding at the end of the block, it is used in conjunction with the padding length field to express the true length boundary.

[0073] After block segmentation and header generation are completed, the counter mode encryption stage begins. During counter mode encryption, the initial counter parameters are concatenated with the block sequence number to form the counter input. The initial counter parameters provide the counter's starting state and direction information, while the block sequence number provides a block-level increment factor, ensuring that the counter inputs for different blocks are distinct. By incrementing the block sequence number for different blocks to obtain different counter inputs, it is clear that for each block processed, the next block sequence number is used in the counter input construction. Thus, with a fixed session key, a corresponding keystream is generated for each block, and block-level encryption is completed. When decrypting, the receiving end can align the keystream and recover the corresponding plaintext block by constructing the counter input using the same initial counter parameters and the same block sequence number.

[0074] In some embodiments of this application, when forward error correction coding is performed on several ciphertext blocks to generate an FEC redundancy packet, and the several ciphertext blocks and the FEC redundancy packet are encapsulated into a transmission unit for transmission, the process includes: Divide a series of consecutive ciphertext blocks into FEC groups, and assign a group identifier and a block number within the group to each FEC group. For each FEC group, perform systematic erasure coding to generate a set of redundant packets corresponding to that FEC group. The number of redundant packets is determined by the link packet loss statistics read by the sender. A transmission unit header is generated for each ciphertext block and each redundant packet. The transmission unit header includes a session identifier, a group identifier, a block sequence number within the group, a payload type flag, and a timestamp field. The integrity key segment is used to calculate the message authentication code for the transmission unit header and payload, and the message authentication code is written into the verification field of the transmission unit before being sent.

[0075] Specifically, after the transmitting terminal completes the block encryption of the media bitstream, it obtains several ciphertext blocks. To facilitate data recovery under packet loss conditions, the transmitting terminal divides several consecutive ciphertext blocks in the sending order into an FEC group and assigns a group identifier and a block sequence number within the group to the FEC group. The group identifier is used to distinguish different FEC groups, and can be generated, for example, by an incrementing group count value within the session. The block sequence number within the group is used to identify the position number of each ciphertext block within the FEC group, so that the receiving end can aggregate the blocks by group and perform erasure coding decoding after receiving the packets. Subsequently, each FEC group is systematically erasure coded. Systematic means that the encoded output simultaneously contains the original ciphertext block and redundant packets. The redundant packets are calculated from several ciphertext blocks within the group according to the erasure coded rules and are used to participate in the recovery when some ciphertext blocks are lost at the receiving end. The number of redundant packets is determined by the link packet loss statistics read by the sending end. In some embodiments, the link packet loss statistics are estimated by a sliding window. The sending end uses a preset time window or a preset number of FEC groups as the statistical window, records the number of transmission units sent within the statistical window and the number of missing indications fed back by the receiving end, and calculates the window packet loss rate. The sending end pre-stores a mapping table between the packet loss rate and the number of redundant packets, and determines the number of redundant packets by looking up the table according to the window packet loss rate. Optionally, when missing feedback is unavailable, the sender uses the previous effective window packet loss rate or a preset default value to determine the number of redundant packets. The link packet loss statistics can be obtained from the sliding window statistics maintained by the sender within the session. For example, the loss ratio of the most recent transmission units, ACK / feedback information (if present), or the packet loss estimate inferred by the sender can be used as input, and the number of redundant packets can be selected according to a preset mapping relationship. For example, fewer redundant packets can be selected when packet loss is low, and more redundant packets can be selected when packet loss is high. After erasure coding is completed, the sending terminal generates a transmission unit header for each ciphertext block and each redundant packet. The transmission unit header is used to carry necessary session and packet information, including at least a session identifier, a group identifier, a block sequence number within the group, a payload type flag, and a timestamp field. The session identifier is used to bind the transmission unit to the corresponding session, the payload type flag is used to distinguish whether the payload of the transmission unit is a ciphertext block or a redundant packet, and the timestamp field is used to record the time information of the generation or transmission of the transmission unit, which is convenient for the receiver to sort and correlate timing during reassembly and decoding.To prevent the header or payload of a transmission unit from being tampered with during transmission, the sending terminal uses an integrity key segment to calculate a message authentication code for the header and payload of the transmission unit. The integrity key segment is a key material derived within the session and used differently from the encryption key segment. The message authentication code can be implemented using an authentication code algorithm compatible with the national cryptographic system. Optionally, the message authentication code is calculated using the SM4-CMAC or SM3-based HMAC algorithm. For example, the header and payload are concatenated in a preset order and a check value is calculated. Then, the message authentication code is written into the check field of the transmission unit and sent. This allows the receiving end to use the same integrity key segment to recalculate the message authentication code for the received header and payload after receiving the transmission unit and compare it with the check field. This verifies the consistency between the header field and payload data of each transmission unit and aggregates ciphertext blocks and redundant packets under the organization of group identifier and block sequence number within the group, providing structured input for subsequent erasure coding decoding and ciphertext block recovery.

[0076] In some embodiments, the receiving end discards the transmission unit when the message authentication code verification of the transmission unit fails; if at least one ciphertext block is still missing after the erasure coding decoding of a certain FEC group, the data segment corresponding to the FEC group is discarded; optionally, the receiving end generates a loss notification message and sends it to the sending end, and the sending end adjusts the number of subsequent FEC redundancy packets or performs packet loss hiding processing accordingly.

[0077] In some embodiments, if two-way authentication fails or key confirmation value verification fails, the session establishment process is terminated and the session state is released; alternatively, an error code is returned to the upper-layer application and a session re-establishment is triggered.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A lightweight encryption and endpoint sensitive content protection method for narrowband weak networks, characterized in that, include: When the sending terminal and the receiving terminal establish a session connection, a two-way authentication is performed using an identifier cryptosystem, and a session key and initial counter parameters are derived based on the output of the two-way authentication. When acquiring audio and video data on the sending terminal side, sensitive content detection is performed on each video frame before encoding to obtain a set of sensitive regions; pixel-level masking is performed on sensitive regions in the set of sensitive regions whose detection confidence is greater than or equal to a preset sensitive threshold to obtain desensitized video frames, and the corresponding original video frames are cleared from the cache and not written to persistent storage. The desensitized video frames and audio data are encoded to obtain a media stream, and the media stream is divided into blocks according to the SM4 block length to obtain several plaintext blocks; Using the SM4 national cryptographic algorithm in counter mode, the session key and the initial counter parameters are used to encrypt several plaintext blocks to generate several corresponding ciphertext blocks; Forward error correction coding is performed on several ciphertext blocks to generate FEC redundancy packets, and several ciphertext blocks and FEC redundancy packets are encapsulated into transmission units and sent. The total bandwidth of the narrowband weak network link is less than or equal to a preset narrowband threshold. On the receiving terminal side, forward error correction decoding is performed on the received transmission unit to obtain a set of ciphertext blocks; The ciphertext block set is decrypted using the SM4 national cryptographic algorithm in counter mode, using the session key and the initial counter parameters to obtain plaintext blocks, which are then reassembled into a media stream. The media stream is then decoded to obtain a video frame sequence and an audio sampling sequence.

2. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 1, characterized in that, When establishing a session connection between the sending terminal and the receiving terminal, a two-way authentication method using an identifier cryptosystem is employed. This includes deriving the session key and initial counter parameters based on the output of the two-way authentication, and includes: The sending terminal generates a first random number and constructs a session establishment message, which includes a protocol version number, an encryption mode identifier, an FEC parameter set identifier, a media encoding parameter identifier, the first random number, and a sending terminal identifier, and sends it to the receiving terminal. After receiving the session establishment message, the receiving terminal generates a second random number, constructs a session confirmation message, which includes the second random number, the receiving terminal identifier, and a session parameter confirmation field, and sends it to the sending terminal. The sending terminal and the receiving terminal perform bidirectional authentication on the session establishment message and the session confirmation message respectively, and obtain bidirectional authentication output; The sending terminal identifier, receiving terminal identifier, first random number, second random number and the two-way authentication output are concatenated to form an input string. The input string is input into a key derivation function to generate key material, and the session key and the initial counter parameter are extracted from the key material. The sending terminal generates a key confirmation value and sends it to the receiving terminal. After verifying the key confirmation value, the receiving terminal generates an confirmation response and returns it to the sending terminal. The key confirmation value is obtained by calculating a message authentication code from the digests of the session establishment message and the session confirmation message. The message authentication code calculation uses the session key as the key.

3. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 2, characterized in that, When the sending terminal generates the first random number and constructs the session establishment message, it includes: Set a preset byte length for the first random number and generate it using a cryptographically secure random number generator; The capability field is set in the session establishment message. The capability field includes at least the set of supported FEC group sizes, the set of redundant packet numbers, the set of transmission unit header lengths, and the retransmission disable flag. The message establishment message for the session is configured with a message sequence number field and a timestamp field, and the message sequence number is maintained to be monotonically increasing on the sending terminal side. The fields of the session establishment message, excluding the authentication field, are concatenated to form a string to be authenticated. A first authentication value is calculated on the string to be authenticated using the private key of the identifier cryptosystem, and the first authentication value is written into the authentication field of the session establishment message.

4. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 3, characterized in that, When concatenating the sending terminal identifier, receiving terminal identifier, the first random number, the second random number, and the two-way authentication output to form an input string, inputting the input string into a key derivation function to generate key material, and extracting the session key and the initial counter parameters from the key material, the process includes: A hash operation is performed on the input string to obtain a digest of a defined length; The determined length digest is input into the key derivation function to obtain the derived output, and the derived output is divided into an encryption key segment, an integrity key segment, and a counter seed segment; Use the encrypted key segment as the session key; The initial counter parameters are obtained by concatenating the counter seed segment with the session direction flag. When the sending or receiving end detects that the counter is about to roll back, it generates an updated count value, concatenates the updated count value with the determined length digest, inputs it into the key derivation function to obtain a new counter seed segment, and uses the new counter seed segment to update the initial counter parameters.

5. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 4, characterized in that, When performing edge-side sensitive content detection on each acquired video frame before encoding, the process includes: The video frames are scaled, pixel normalized, and channel arranged to obtain the detection input frames; The detection input frame is converted into a fixed-point representation and input into the end-side detection model, and a candidate region list and the detection confidence of the candidate regions are output. During continuous frame processing, the candidate region list of the previous frame is written into the volatile buffer, and the region position is extrapolated for the current frame to obtain the predicted region. When the predicted region overlaps with the candidate region, the overlapping regions are merged and the corresponding confidence records are updated.

6. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 5, characterized in that, When obtaining the set of sensitive regions, it includes: The candidate region list is overlapped and merged to output a merged region list. Each merged region records the rectangle coordinates, frame number and corresponding detection confidence. The list of merged regions is converted into a region index table, which uses the frame number as the key and the set of coordinates of the merged regions as the value. The region index table is written into volatile memory and bound to the current video frame to form the sensitive region set; When processing the next video frame, the memory space occupied by the region index table corresponding to the previous video frame is released.

7. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 6, characterized in that, When performing pixel-level masking on sensitive regions in the set of sensitive regions whose detection confidence is greater than or equal to a preset sensitivity threshold to obtain desensitized video frames, the process includes: Align the coordinates of each sensitive region to the boundary of the video coding macroblock to obtain the aligned masking window; A binary mask image is generated within the masking window, and the mask image is mapped to the luminance and chrominance components of the video frame; The occlusion window is divided into pixel sub-blocks of a preset size. A representative pixel value is calculated for each pixel sub-block and then filled back into all pixel positions of that pixel sub-block. The masked luminance and chrominance components are recombined to obtain the desensitized video frame.

8. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 7, characterized in that, When the corresponding original video frame is cleared from the buffer without being written to persistent storage, this includes: Write the original video frames into a volatile circular buffer; After obtaining the desensitized video frame, the memory area in the volatile circular buffer corresponding to the original video frame is overwritten and then the memory area is released. On the sending terminal side, the original video frames are prohibited from entering the file writing interface, system-level image cache interface, screen capture cache interface, and log disk writing interface.

9. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 8, characterized in that, When the media stream is divided into blocks aligned with the SM4 packet length to obtain several plaintext blocks, it includes: The media stream is sequentially divided into block sequences, and a block number and a frame number are written for each block. When the last block is less than the length of an SM4 packet, padding bytes are appended to the end and the padding length field is written. A block header is generated for each block. The block header includes a session identifier, the block sequence number, the frame sequence number, and an original length field. The session identifier is obtained by concatenating the sending terminal identifier, the receiving terminal identifier, the first random number, and the second random number. When performing counter mode encryption, the initial counter parameter is concatenated with the block number to form a counter input, and the block number is incremented for different blocks to obtain different counter inputs.

10. The lightweight encryption and end-side sensitive content protection method for narrowband weak networks according to claim 9, characterized in that, When performing forward error correction coding on several ciphertext blocks to generate an FEC redundancy packet, and encapsulating the several ciphertext blocks and the FEC redundancy packet into a transmission unit for transmission, the process includes: Divide a series of consecutive ciphertext blocks into FEC groups, and assign a group identifier and a block number within the group to each FEC group. For each FEC group, perform systematic erasure coding to generate a set of redundant packets corresponding to that FEC group. The number of redundant packets is determined by the link packet loss statistics read by the sender. A transmission unit header is generated for each ciphertext block and each redundant packet. The transmission unit header includes a session identifier, the group identifier, the block sequence number within the group, a payload type flag, and a timestamp field. The integrity key segment is used to calculate the message authentication code for the header and payload of the transmission unit, and the message authentication code is written into the verification field of the transmission unit before being sent.

Citation Information

Patent Citations

  • Lightweight identity authentication and data security interaction method, system and equipment for protecting measurement and control device and medium

    CN121098618A

  • Data transmission method and system for meteorological satellite communication system

    CN121690338A