Multi-target tracking method and system for cross-border transmission encrypted video

By extracting motion features from encrypted video streams and constructing encrypted domain feature sets, and combining encrypted domain detectors and Kalman filtering algorithms, the problems of low efficiency and poor accuracy in target detection and tracking under encrypted video are solved. This achieves efficient and accurate target recognition and tracking under encrypted conditions, and is suitable for cross-border transmission and privacy protection.

CN121750896APending Publication Date: 2026-03-27SUN YAT SEN UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing video target detection and tracking methods mainly operate in the plaintext domain, which cannot efficiently and accurately achieve target detection and tracking in encrypted video, and may violate privacy protection regulations.

Method used

By extracting motion features from encrypted video streams, constructing encrypted domain feature sets, using encrypted domain detectors for multi-target detection, and performing tracking processing through Kalman filtering and Hungarian algorithms, multi-target trajectory output is achieved, avoiding the decryption of video content.

Benefits of technology

While maintaining video content encryption, real-time multi-target recognition and tracking of cross-border video transmission was achieved in the cloud, improving detection efficiency and accuracy, complying with privacy protection regulations, and ensuring data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750896A_ABST
    Figure CN121750896A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target tracking method and system for a cross-border transmission encrypted video, and relates to the technical field of encrypted video target tracking. Firstly, motion features are directly extracted from an encrypted video code stream, and an encrypted domain feature set is constructed, so that a cloud can obtain target motion related information on the premise of not recovering a video plaintext, and video content privacy and cross-border transmission security are ensured. Secondly, multi-target detection is executed based on the encryption domain feature set, effective target position identification can be realized under the condition of keeping the encryption structure unchanged, and the adaptability of target detection is improved. Furthermore, by tracking a detection result set and stably outputting a multi-target trajectory, a system can still obtain a continuous target motion trajectory under the condition that a video is not decrypted in the whole process, the reliability of cross-frame association is improved, and finally, efficient and accurate target detection and tracking can be realized in a video encryption state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of encrypted video target tracking technology, and in particular to a multi-target tracking method and system for cross-border transmission of encrypted video. Background Technology

[0002] With the rapid development of information technology and cloud computing, the application of video data in fields such as security monitoring, intelligent transportation, telemedicine, and social media is constantly expanding. To meet the needs of cross-regional data transmission and cloud storage, more and more videos are being uploaded to cloud servers for centralized management and analysis. At the same time, privacy protection and data security issues are becoming increasingly prominent. Video content often contains personal identity, behavioral characteristics, and scene information; therefore, encryption methods are usually used during cloud transmission and storage to prevent leakage.

[0003] Currently, encrypted video has become the mainstream form of cross-border data transmission and cloud-based analysis. However, while encryption ensures data security, it renders traditional plaintext video processing algorithms ineffective. Existing video target detection and tracking methods mainly operate in the plaintext domain, relying on pixel-level features or image semantic information. Once the video is encrypted, the algorithm cannot directly access the original frame content, thus failing to extract effective features, resulting in low efficiency and poor accuracy in detection and tracking tasks. Downloading the video first and then decrypting and analyzing it locally is not only inefficient but may also violate privacy regulations, further limiting the feasibility of cloud-based target tracking in encrypted videos. Summary of the Invention

[0004] Therefore, it is necessary to address the problem that existing technologies cannot achieve efficient and accurate target detection and tracking under video encryption conditions, and to provide a multi-target tracking method and system for cross-border transmission of encrypted video, which can realize real-time multi-target recognition and tracking of cross-border transmission video in the cloud while keeping the video content encrypted.

[0005] To achieve the above-mentioned technical effects, the technical solution of the present invention is as follows: A method for multi-target tracking of encrypted video transmitted across borders, comprising: S1. Receive encrypted video streams transmitted across borders; S2. Extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set; S3. Based on the encrypted domain feature set, perform multi-target detection to obtain a set of detection results; S4. Track and process the detection result set to obtain multi-target trajectories.

[0006] Preferably, the cross-border transmitted encrypted video stream is a video stream that has been compressed and encoded and then selectively encrypted in a format compatible manner.

[0007] Preferably, the encryption domain feature set includes macroblock-level features, subblock-level features, and transform coefficient-level features.

[0008] Preferably, the macroblock-level feature is the number of encoded bits of each macroblock in the encrypted video bitstream; the number of encoded bits of each macroblock The expression is as follows:

[0009] in, For macroblocks,

[0010] Preferably, the sub-block level feature is the fineness of the division of each sub-block in the encrypted video bitstream; the fineness of the division of each sub-block The expression is as follows:

[0011] in, For macroblock mapping functions, This refers to the set of sub-blocks in an encrypted video stream.

[0012] Preferably, the transform coefficient level feature is the number of non-zero coefficients in each sub-block of the encrypted video bitstream after undergoing discrete cosine transform; the number of non-zero coefficients in each sub-block after undergoing discrete cosine transform. The expression is as follows:

[0013] Where Count The number of non-zero high-frequency coefficients. If the luminance component is non-zero, the value is 1; otherwise, it is 0. as one The sub-block.

[0014] Preferably, the step of performing multi-target detection based on the encrypted domain feature set to obtain a detection result set includes: S31. Constructing a cryptographic domain detector ; S32. Input the encrypted domain feature set into the encrypted domain detector. Output the set of detection results The detection result set includes detection boxes and corresponding confidence scores; the detection result set The expression is as follows:

[0015] in Indicates the first Candidate bounding boxes, For the first The confidence scores of each candidate bounding box; the set of detection results For the encrypted domain detector For encrypted domain feature set The mapping; the expression for the mapping is as follows:

[0016] in, This is a set of features for the encrypted domain.

[0017] Preferably, the step of tracking the detection result set to obtain multi-target trajectories includes: S41. Obtain the target trajectory of the detection box; S42. Establish a Kalman filter state vector for the target trajectory; S43. Input the Kalman filter state vector into a preset constant-speed model and output the predicted target box; S44. Calculate the association cost between the detection box and the predicted target box in the detection result set; S45. Based on the association cost, the predicted target box and the detection box are matched to obtain the successfully matched multi-target trajectory.

[0018] Preferably, the associated cost The calculation formula is:

[0019] in,

[0020] This invention also provides a multi-target tracking system for cross-border transmission of encrypted video, comprising: The video stream receiving module is used to receive encrypted video streams transmitted across borders. The motion feature extraction module is used to extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set. A multi-target detection module is used to perform multi-target detection based on the encrypted domain feature set to obtain a set of detection results; The multi-target tracking module is used to track and process the detection result set to obtain the trajectory of multiple targets.

[0021] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This invention proposes a multi-target tracking method and system for cross-border transmission of encrypted video. First, by directly extracting motion features from the encrypted video stream and constructing an encrypted domain feature set, the cloud can obtain target motion-related information without recovering the plaintext video, ensuring video content privacy and cross-border transmission security. Second, by performing multi-target detection based on the encrypted domain feature set, effective target location identification can be achieved while maintaining the encrypted structure, improving the adaptability of target detection. Furthermore, by tracking the detection result set and stably outputting multi-target trajectories, the system can still obtain continuous target motion trajectories without decrypting the video throughout, improving the reliability of cross-frame correlation. Ultimately, efficient and accurate target detection and tracking can be achieved even under encrypted video conditions. Attached Figure Description

[0022] Figure 1 This is a flowchart of a multi-target tracking method for cross-border transmission of encrypted video in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the feature extraction of a multi-target tracking method for cross-border transmission of encrypted video in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating YOLOX target detection in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating target tracking using a Kalman filter and a Hungarian algorithm in an embodiment of the present invention; Figure 5 The figures shown are some of the experimental results in the embodiments of the present invention; Figure 6 This is a system block diagram of a multi-target tracking method for cross-border transmission of encrypted video in an embodiment of the present invention. Detailed Implementation

[0023] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. It is understandable to those skilled in the art that some well-known details may be omitted from the accompanying drawings; The positional relationships depicted in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent. To better illustrate this embodiment, some parts of the accompanying drawings may be omitted, enlarged, or reduced, and do not represent actual dimensions. The descriptions of directions such as "up" and "down" are not intended to limit this patent. To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.

[0024] Example 1 like Figure 1 As shown, this embodiment provides a multi-target tracking method for cross-border transmission of encrypted video, including: S1. Receive encrypted video streams transmitted across borders; In S1, the video capture terminal acquires video and compresses and encodes the original video frame by frame using standards such as H.264 / HEVC. Then, it performs selective encryption on the compressed video stream, generating an encrypted H.264 stream that still conforms to standard syntax. This encryption process maintains the stream structure, encrypting only certain syntax elements, such as intra-frame prediction modes, motion vector residuals, and residual coefficients. Both encoding and encryption use frames as the basic processing unit and NALs as the network transmission granularity. After completing the encoding and encryption of the current frame, the terminal immediately streams the corresponding NAL segment to the cloud without waiting for the entire GOP to complete, significantly reducing transmission latency between the terminal and the cloud. Encryption uses a terminal-side key and is performed in a secure execution environment (HSM / TEE). The key remains at the terminal and is not transmitted in the link. The cloud always receives a structurally valid but invisible encrypted stream, avoiding the risk of plaintext exposure during cross-border transmission from the source. The cloud server obtains the compressed and encoded video stream with format-compatible selective encryption from the video capture terminal via a cross-border network connection, and receives and caches the received encrypted video stream in real time for subsequent encrypted domain feature extraction processing.

[0025] Specifically, to perform format-compatible encryption on videos, existing video format-compatible encryption methods can be used, or the following methods can be employed: ① Select bits of parameters from the video bitstream that are used for encryption without destroying the video format, such as residual coefficients, intra-frame prediction mode, motion vector residuals, etc. ② Encrypt the selected bits using encryption algorithms such as AES-GCM and RC4; ③ Replace the encrypted bits with the original bits to obtain the encrypted video.

[0026] S2. Extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set; In S2, such as Figure 2 As shown, after receiving the encrypted video stream, the cloud server directly extracts motion features from the encrypted stream without decrypting the video content; by macroblock ( The three-scale encrypted-motion features (EMF) and its sub-block structure are directly extracted from the encrypted stream: ① Regarding the first Frame number macroblock Calculate the total number of bits used after macroblock encoding:

[0027] ② Regarding the first The first macroblock indivual sub-block The fineness of the encoder's division of this region is calculated:

[0028] Where mapping Pick:

[0029] ③ Regarding the first The first macroblock indivual sub-block Calculate the number of non-zero coefficients in the sub-block after DTC transformation:

[0030] Where Count The number of non-zero high-frequency coefficients. If the luminance component is non-zero, the value is 1; otherwise, it is 0. as one The sub-block.

[0031] S3. Based on the encrypted domain feature set, perform multi-target detection to obtain a set of detection results; In S3, based on the encryption domain features obtained in step (2) The cloud implements a Tracking by Detection process in the encrypted domain, such as... Figure 3 and Figure 4 As shown, it specifically includes: The anchor-free decoupled detection head of YOLOX is directly adapted to the EMF input: keeping the YOLOX backbone network and PAFPN structure unchanged, the input domain is changed from the original RGB image to a three-scale encrypted domain feature map (α / β / γ) derived from the encryption stream. YOLOX no longer processes appearance information such as texture / color, but learns the mapping relationship between the encrypted domain feature set and the target bounding box in the encrypted domain. Constructing a YOLOX-based cryptographic domain detector : Let the detector output the detection set :

[0032] in Indicates the first Candidate bounding boxes, Indicates the corresponding confidence level; the detection rule is defined as the encrypted domain detector. For encrypted domain feature set Mapping;

[0033] The In cryptographic domain features Training and fine-tuning were performed on it to achieve the desired results. It can generate high-confidence checks in the encrypted feature space.

[0034] S4. Track and process the detection result set to obtain multi-target trajectories.

[0035] In S4, association and tracking: With existing trajectory sets Perform association: Generate a set of detection boxes for each frame using a detector trained based on cryptographic domain features. Its confidence level; maintain an 8-dimensional Kalman filter state vector for each trajectory.

[0036] in Indicates the center position of the bounding box. Indicates the width and height of the bounding box. Indicates the velocity at the center position. Indicates the rate of change of width and height; A constant-velocity model is used for prediction and updating. The constant-velocity model assumes that the target moves at a constant speed over a short period, meaning that the target's position change is linearly predictable between two consecutive frames because the velocity of most objects (such as people and vehicles) does not change abruptly between frames. It is a linear motion model based on Kalman filtering, used to describe the smooth motion state of the target between adjacent frames. Its state vector includes the bounding box center coordinates, width, height, and its velocity components. The state transition equation of the constant-velocity model is:

[0037] in Let F be the process noise, and F be the state transition matrix of the constant-rate model.

[0038] The position of the detection box is determined by the position and velocity prediction of the previous frame:

[0039] The width and height of the detection frame are similar:

[0040] However, the velocity component remains unchanged:

[0041] The bounding box output by the YOLOX detector provides the observation vector:

[0042] The observation equation is:

[0043] in This is observation noise (introduced by the uncertainty of the detection rate). The observation matrix H is:

[0044] This indicates that the observations only include position and scale, excluding velocity. A set of bounding boxes is generated for each frame using a detector trained on cryptographic domain features. Its confidence level; Calculate the association cost between the detection boxes and the predicted target boxes in the detection box set; Based on the association cost, the predicted target box and the detection box are matched to obtain the successfully matched multi-target trajectory; Association cost is defined as:

[0045] The association was solved using the Hungarian algorithm, and a two-stage matching strategy was adopted: the first stage involved applying a high threshold between the tracked trajectory and the missing trajectory and the detection box. IoU cost matching; the second stage involves performing low-threshold matching again between unconfirmed trajectories and remaining detection boxes. IoU cost matching. If the confidence of the new detection box is greater than the initial threshold... Then initialize it as a new trajectory; if a trajectory is continuous If no detection box is matched within the frame, the target is considered to have left the field of view and its trajectory is removed from the trajectory set.

[0046] While continuously receiving encrypted streams, the cloud performs frame-by-frame reassembly and processing of arriving NAL units based on H.264 syntax information, such as POC / PTS. After processing, the cloud only returns semantic-level tracking metadata corresponding to that frame, including: bounding box position, trajectory ID, and frame-level FrameID from POC / PTS for edge synchronization, as well as optional timestamps or motion vectors. Throughout the process, the cloud does not return or store any plaintext images or video content. The edge device decrypts the video frame-by-frame locally using the key and accurately aligns the detection / tracking results with the local video frames based on the returned FrameID, achieving overlay display of bounding boxes / trajectory IDs for real-time visualization or forensics. Experiments show that on a single RTX3090 GPU, this system can achieve real-time encrypted domain multi-target tracking performance of approximately 25 FPS without video decryption, meeting the real-time requirements of practical deployments.

[0047] This embodiment proposes a secure identification technology for compressed video transmitted across borders. The core of the technical solution includes: after the video terminal device compresses and encodes the captured video, a format-compatible selective encryption strategy is used to encrypt parts of the bitstream, such as motion vector residuals and intra-frame prediction modes; the encrypted bitstream is uploaded to a cloud server in real time. The cloud directly processes the encrypted video bitstream without decryption, extracting features containing motion information for multi-target identification. Specifically, the cloud first parses macroblock-level motion information from the encrypted bitstream: locating macroblocks containing motion vector residuals and estimating the residual size based on codeword length to generate a frame-level motion residual map; simultaneously, it extracts multi-scale motion features such as macroblock size, macroblock segmentation depth, and non-zero DCT coefficients within 4×4 blocks. These features comprehensively reflect motion activities at different scales and serve as input for target detection. Building upon this foundation, a tracking-detection paradigm is introduced: the cloud uses a detector based on encrypted features to detect targets in each frame, then uses a Kalman filter to predict the motion state of each target, and performs multi-target association based on IOU distance using the Hungarian algorithm; a two-stage matching strategy is adopted, first matching high-confidence detections, then matching low-confidence detections, to improve tracking robustness in complex scenarios such as occlusion; trajectory lifecycle management includes state transitions such as creation, tracking, loss, and removal. The entire process supports a streaming mode of processing while transmitting, allowing the cloud to update detection results as it receives a portion of the video data. Finally, the cloud outputs the target's location information and motion trajectory recognition results, and returns this information to the terminal or user for local decryption and video display or further analysis.

[0048] This embodiment achieves efficient moving target recognition in the cloud without decrypting the video content, balancing security and practicality. Through multi-scale motion feature fusion, this embodiment achieves robust tracking of multiple targets in an encrypted domain, significantly outperforming methods relying solely on single features. Furthermore, this embodiment employs a compression domain processing method, reducing algorithm complexity and improving real-time performance, making it suitable for cross-border transmission and compatible applications across various device environments. This technology complies with privacy regulations such as GDPR and PIPL, ensuring that video content remains encrypted throughout transmission and analysis, preventing the leakage of user plaintext information and effectively protecting personal privacy.

[0049] Example 2 This embodiment demonstrates the results of the above embodiments.

[0050] like Figure 5 As shown, pedestrians in the video can be accurately located on encrypted video using our algorithm.

[0051]

[0052] Table 1 As shown in Table 1, our detection algorithm achieves high HOTA and MOTA on different datasets.

[0053] Example 3 like Figure 6 As shown, this embodiment provides a multi-target tracking system for cross-border transmission of encrypted video, including: The video stream receiving module is used to receive encrypted video streams transmitted across borders. The motion feature extraction module is used to extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set. A multi-target detection module is used to perform multi-target detection based on the encrypted domain feature set to obtain a set of detection results; The multi-target tracking module is used to track and process the detection result set to obtain the trajectory of multiple targets.

[0054] This embodiment proposes a multi-target tracking system for cross-border transmission of encrypted video. First, by directly extracting motion features from the encrypted video stream and constructing an encrypted domain feature set, the cloud can obtain target motion-related information without recovering the plaintext video, ensuring video content privacy and cross-border transmission security. Second, by performing multi-target detection based on the encrypted domain feature set, effective target location identification can be achieved while maintaining the encrypted structure, improving the adaptability of target detection. Furthermore, by tracking the detection result set and stably outputting multi-target trajectories, the system can still obtain continuous target motion trajectories without decrypting the video throughout, improving the reliability of cross-frame correlation. Ultimately, efficient and accurate target detection and tracking can be achieved even under encrypted video conditions.

Claims

1. A method for multi-target tracking of encrypted video transmitted across borders, characterized in that, include: S1. Receive encrypted video streams transmitted across borders; S2. Extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set; S3. Based on the encrypted domain feature set, perform multi-target detection to obtain a set of detection results; S4. Track and process the detection result set to obtain multi-target trajectories.

2. The multi-target tracking method for cross-border transmission of encrypted video according to claim 1, characterized in that, The cross-border transmitted encrypted video stream is a video stream that has been compressed and encoded and then selectively encrypted in a format compatible manner.

3. The multi-target tracking method for cross-border transmission of encrypted video according to claim 1, characterized in that, The encryption domain feature set includes macroblock-level features, subblock-level features, and transform coefficient-level features.

4. The multi-target tracking method for cross-border transmission of encrypted video according to claim 3, characterized in that, The macroblock-level feature is the number of encoded bits of each macroblock in the encrypted video bitstream; the number of encoded bits of each macroblock The expression is as follows: in, For macroblocks, Encoding function, This is a function for determining the length.

5. The multi-target tracking method for cross-border transmission of encrypted video according to claim 3, characterized in that, The sub-block level feature refers to the fineness of the division of each sub-block in the encrypted video bitstream; the fineness of the division of each sub-block The expression is as follows: in, For macroblock mapping functions, This refers to the set of sub-blocks in an encrypted video stream.

6. The multi-target tracking method for cross-border transmission of encrypted video according to claim 3, characterized in that, The transform coefficient level feature is the number of non-zero coefficients in each sub-block of the encrypted video bitstream after undergoing discrete cosine transform; the number of non-zero coefficients in each sub-block after undergoing discrete cosine transform. The expression is as follows: Where Count The number of non-zero high-frequency coefficients. If the luminance component is non-zero, the value is 1; otherwise, it is 0. as one The sub-block.

7. The multi-target tracking method for cross-border transmission of encrypted video according to claim 1, characterized in that, The step of performing multi-target detection based on the encrypted domain feature set to obtain a detection result set includes: S31. Constructing a cryptographic domain detector ; S32. Input the encrypted domain feature set into the encrypted domain detector. Output the set of detection results The detection result set includes detection boxes and corresponding confidence scores; the detection result set The expression is as follows: in Indicates the first Candidate bounding boxes, For the first The confidence scores of each candidate bounding box; the set of detection results For the encrypted domain detector For encrypted domain feature set The mapping; the expression for the mapping is as follows: in, This is a set of features for the encrypted domain.

8. The multi-target tracking method for cross-border transmission of encrypted video according to claim 7, characterized in that, The process of tracking and processing the detection result set to obtain multi-target trajectories includes: S41. Obtain the target trajectory of the detection box; S42. Establish a Kalman filter state vector for the target trajectory; S43. Input the Kalman filter state vector into a preset constant-speed model and output the predicted target box; S44. Calculate the association cost between the detection box and the predicted target box in the detection result set; S45. Based on the association cost, the predicted target box and the detection box are matched to obtain the successfully matched multi-target trajectory.

9. The multi-target tracking method for cross-border transmission of encrypted video according to claim 8, characterized in that, The associated cost The calculation formula is: in, To predict the intersection-union ratio of the target bounding box and the detection bounding box, To predict the target bounding box, This is the detection frame.

10. A multi-target tracking system for cross-border transmission of encrypted video, characterized in that, include: The video stream receiving module is used to receive encrypted video streams transmitted across borders. The motion feature extraction module is used to extract motion features from the cross-border transmitted encrypted video stream to obtain an encrypted domain feature set. A multi-target detection module is used to perform multi-target detection based on the encrypted domain feature set to obtain a set of detection results; The multi-target tracking module is used to track and process the detection result set to obtain the trajectory of multiple targets.