Streaming media multicast tamper-proofing method and device based on deep feature fingerprints
By combining SHA-1 hashing and deep feature vectors to generate a comprehensive fingerprint in the IPTV multicast system, the security and real-time issues of multicast video are solved, and efficient and reliable secure video distribution is achieved.
Patent Information
- Application Number
- CN202511239387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-25
AI Technical Summary
Existing IPTV multicast systems face problems such as weak collision resistance, high false alarm rate and insufficient real-time performance in terms of content security. Traditional SHA-1 digests are difficult to meet the real-time requirements of low latency and high concurrency, while deep feature vectors lack rigorous mathematical proof in terms of authenticity determination.
A deep feature fingerprint-based anti-tampering method is adopted. A two-stage fingerprint is generated on the live source side: the SHA-1 hash processing module and the lightweight convolutional neural network extract deep feature vectors, and combine them with preset weights to form a comprehensive digital fingerprint. The fingerprint is then compared twice on the terminal side, and Hamming distance and cosine similarity are used to determine whether the video content has been tampered with.
It achieves efficient and reliable secure video distribution, reduces false alarm rates, improves the security and real-time transmission performance of multicast video, and provides rapid response and visual traceability capabilities.
Smart Images

Figure CN121012969A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of streaming media tamper-proofing, and particularly relates to a streaming media multicast tamper-proofing method and device based on deep feature fingerprints. BACKGROUND
[0002] In recent years, with the continuous improvement of network bandwidth and the popularity of terminal equipment, IPTV (Internet Protocol Television) has become one of the important ways of video distribution. Its multicast technology can significantly reduce the bandwidth overhead when multiple users watch live at the same time, and is widely used in operator, enterprise and education scenarios. IPTV multicast mainly relies on IGMP (Internet Group Management Protocol) and PIM (Protocol Independent Multicast) protocols to realize one-to-many video stream transmission and improve network resource utilization efficiency. However, during the transmission of multicast streaming media across networks and operators, due to the complexity of routing nodes, network congestion and rewriting of intermediate devices, video content is easily tampered with or damaged during transmission, thereby affecting user experience and digital copyright protection. Existing content tamper-proofing technology is mainly based on classic hash algorithms such as SHA-1 and MD5, which calculate the digest of video PES (Packetized Elementary Stream) packets and embed them into the transport stream for video integrity verification at the receiving end. This method has the advantages of simple implementation and less dependence on hardware, but has three major pain points: first, the SHA-1 algorithm was designed by the NSA of the United States and standardized by NIST in 1995. In recent years, it has been proven to be insufficient in collision resistance. The Google / CWI "SHAttered" experiment first publicly demonstrated a collision attack that can generate different files with the same SHA-1 value under actual conditions, significantly increasing the security risk. Second, the "avalanche effect" of SHA-1 makes it extremely sensitive to any small changes, and in video transmission, normal compression and re-encoding, color difference adjustment, and network packet reassembly will introduce small differences, resulting in high false positive and false negative rates. Third, in the face of real-time high-definition streaming media, traditional SHA-1 calculation cannot fully utilize the parallel capabilities of modern CPUs / GPUs, and performance becomes a bottleneck, making it impossible to meet the real-time requirements of low latency and high concurrency. On the other hand, with the rapid development of deep learning in computer vision, feature extraction technology based on CNN (Convolutional Neural Network) can generate high-dimensional vector features that are highly sensitive to image or video frame semantic information and robust to natural perturbations. Such deep feature vectors have achieved significant results in video fingerprinting and forgery detection, and can capture the structural and semantic differences of video content, maintaining high similarity under slight compression noise, while showing sharp detection capabilities for malicious tampering. Although deep models have higher requirements for computing power and resources, through lightweight network structures and hardware acceleration, real-time processing requirements can be met to some extent. In addition, the dual-fingerprint mechanism that combines traditional hash and deep features can balance data integrity and semantic consistency, significantly improving the accuracy and robustness of tamper-proofing detection.
[0003] In summary, the current IPTV multicast system faces severe challenges in content security, and although the traditional SHA-1 digest can achieve integrity verification, it is difficult to meet the actual needs due to weak collision resistance, high false positive rate and insufficient real-time performance. While the deep feature vector has robustness, it lacks rigorous mathematical proof in authenticity determination. Therefore, a new tamper-proofing technology that combines the advantages of both is urgently needed to improve multicast video security, reduce false alarm rate and consider real-time transmission performance. SUMMARY
[0004] In view of the deficiencies of the prior art, the purpose of the application is to provide a streaming media multicast tamper-proofing method and device based on deep feature fingerprints, electronic equipment and storage medium, which can help operators to realize efficient and reliable video security distribution in complex network environment.
[0005] In a first aspect of the application, a streaming media multicast tamper-proofing method based on deep feature fingerprints is proposed, which is applied to a system including a live source side and a terminal side, comprising:
[0006] At the live source side, the original video stream collected is packaged by a TS multiplexer according to the PES format, and the PES packet is transmitted to the tamper-proofing generation module;
[0007] The tamper-proofing generation module analyzes the PES packet header in real time to extract the display timestamp, decoding timestamp and header information, and screens out the PES packet carrying I frame data after locating the I frame according to the NALU / VOP flag bit in the header information;
[0008] The SHA-1 hash processing module and the deep feature vector extraction submodule are called respectively to generate the first and second fingerprints;
[0009] The fusion fingerprint generation submodule merges the first and second fingerprints according to a preset weight to form a comprehensive digital fingerprint, and encapsulates the comprehensive digital fingerprint into a 188-byte special TS packet of PID, inserts the special TS packet before the corresponding I frame TS packet, and transmits the original video stream to the terminal side through IP multicast;
[0010] At the terminal side, the TS demultiplexer analyzes the received TS transport stream and distinguishes and separates the regular video PES packet and the embedded tamper-proofing special TS packet according to the PID field, and buffers the regular video PES packet and the embedded tamper-proofing special TS packet into the video buffer and the fingerprint buffer, respectively;
[0011] The tamper-proofing verification module reconstructs the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameters;
[0012] The fusion fingerprint generation submodule fuses the SHA-1 digest and the deep feature vector as a local comprehensive fingerprint with a preset weight consistent with the live source side;
[0013] The matching score is obtained by double comparison of the Hamming distance and the cosine similarity of the local comprehensive fingerprint and the comprehensive digital fingerprint generated and packaged from the fingerprint cache on the live source side;
[0014] According to the size of the matching score, it is judged whether the TS transport stream is tampered with.
[0015] Further, in the above-mentioned stream media multicast tamper-proofing method based on deep feature fingerprints, the first route fingerprint is generated by calling a SHA-1 hash processing module, comprising:
[0016] The SHA1 is calculated after splicing the PES packet data M carrying I frame data and the scrambling string k to generate the first route fingerprint.
[0017] Further, in the above-mentioned stream media multicast tamper-proofing method based on deep feature fingerprints, the second route fingerprint is generated by calling a deep feature vector extraction submodule, comprising:
[0018] The PES packet data corresponding to the I frame data is decoded into a pixel image, which is input into a lightweight convolutional neural network to extract a high-dimensional semantic feature vector through forward inference;
[0019] The high-dimensional semantic feature vector is subjected to L2 norm normalization to obtain the second route fingerprint.
[0020] Further, in the above-mentioned stream media multicast tamper-proofing method based on deep feature fingerprints, the fusion fingerprint generation submodule combines the first route fingerprint and the second route fingerprint to form a comprehensive digital fingerprint according to a preset weight, comprising:
[0021] The first route fingerprint and the second route fingerprint are weighted or concatenated according to the preset weights α and β to form the comprehensive digital fingerprint;
[0022] Wherein, α and β are initially set to 0.5, and are subsequently optimized according to actual conditions.
[0023] Further, in the above-mentioned stream media multicast tamper-proofing method based on deep feature fingerprints, the tamper-proofing verification module reconstructs the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameters, comprising:
[0024] The PES packet data carrying I frame data corresponding to the embedded anti-tamper special TS packet is read from the video cache in sequence;
[0025] Based on the scrambling code string k and the depth model parameters, performing SHA-1 digest operation on PES packet data carrying I frame data corresponding to the embedded tamper-proof private TS packet, to generate SHA-1 digest;
[0026] Decoding the PES packet data carrying I frame data corresponding to the embedded tamper-proof private TS packet into a pixel image;
[0027] Inputting the pixel image into a lightweight convolutional neural network to extract a high-level semantic feature vector;
[0028] Performing L2 normalization on the high-level semantic feature vector to obtain a depth feature vector.
[0029] Further, in the above-mentioned stream media multicast tamper-proofing method based on depth feature fingerprints, whether the TS transmission stream is tampered with is determined according to the size of the matching score, comprising:
[0030] Comparing the size of the matching score with a tolerance threshold defined by a pre-established uncertainty propagation model;
[0031] When the matching score is greater than the tolerance threshold defined by the pre-established uncertainty propagation model, it is determined that the TS transmission stream is tampered with.
[0032] Further, the above-mentioned stream media multicast tamper-proofing method based on depth feature fingerprints further comprises:
[0033] When the TS transmission stream is tampered with, sending alarm information to the upper layer application through an alarm interface and triggering remedial measures according to the operator's strategy;
[0034] The remedial measures at least include automatic retransmission, back-to-source verification or user prompting.
[0035] The second aspect of the present application also proposes a stream media multicast tamper-proofing device based on depth feature fingerprints, applied to a system comprising a live source side and a terminal side, comprising:
[0036] A packaging module: used for packaging the collected original video stream in PES format through a TS multiplexer at the live source side, and transmitting the PES packet into a tamper-proof generation module;
[0037] A tamper-proof generation module: used for real-time analyzing the PES packet header to extract the display timestamp, decoding timestamp and header information, and screening out the PES packet carrying I frame data after locating the I frame according to the NALU / VOP flag bit in the header information;
[0038] An SHA-1 hash processing module and a depth feature vector extraction submodule are used to generate first and second fingerprints respectively;
[0039] Fusion fingerprint generation submodule: for merging the first road fingerprint and the second road fingerprint according to a preset weight to form a comprehensive digital fingerprint, and packaging the comprehensive digital fingerprint into a PID 188-byte special TS package, and inserting the special TS package before the corresponding I frame TS package, and transmitting the original video stream through IP multicast to the terminal side;
[0040] Receiving module: for parsing the received TS transport stream by a TS demultiplexer at the terminal side, and distinguishing and separating the regular video PES package and the embedded anti-tampering special TS package according to the PID field, and buffering the regular video PES package and the embedded anti-tampering special TS package into a video buffer and a fingerprint buffer, respectively;
[0041] Anti-tampering checking module: for reconstructing the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameter;
[0042] Fusion fingerprint generation submodule: for fusing the SHA-1 digest and the deep feature vector into a local comprehensive fingerprint with a preset weight consistent with the live source side;
[0043] Double comparison module: for double comparing the Hamming distance and the cosine similarity of the local comprehensive fingerprint and the comprehensive digital fingerprint extracted from the fingerprint buffer and generated and packaged by the live source side to obtain a matching score;
[0044] Judgment module: for judging whether the TS transport stream is tampered with according to the size of the matching score.
[0045] The third aspect of the present application also proposes an electronic device, comprising: a processor and a memory;
[0046] The processor is used for executing any one of the deep feature fingerprint-based stream media multicast anti-tampering methods described above by calling the program or instruction stored in the memory.
[0047] The fourth aspect of the present application also proposes a computer readable storage medium, which stores a program or instruction, and the program or instruction makes a computer execute any one of the deep feature fingerprint-based stream media multicast anti-tampering methods described above.
[0048] The beneficial effects of the present application are as follows: by generating two-stage fingerprints for video PES package key frames (I frames) on the live source side: on the one hand, SHA-1 and scrambling string are used to generate a high-sensitivity binary digest, ensuring zero tolerance for any precise tampering behavior; on the other hand, a lightweight convolutional neural network is used to extract a high-dimensional semantic feature vector of the frame, providing fault tolerance support for harmless disturbances such as normal compression, color difference and format conversion; the two-stage fingerprints are transmitted to the terminal together with the TS stream after weighted fusion. At the terminal side, the present application realizes a double fingerprint reconstruction process completely synchronized with the live source side by upgrading the TS demultiplexer, independently calculates the SHA-1 digest and the deep feature vector in sequence, and generates a reconstructed fingerprint according to the same fusion weight. Subsequently, by weighted comprehensive comparison of Hamming distance and cosine similarity, combined with a preset threshold, it determines in real time whether the video content has been tampered with, and triggers an alarm to the upper layer application when an anomaly is detected, providing operators with fast response and visual tracing capabilities. BRIEF DESCRIPTION OF DRAWINGS
[0049] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:
[0050] Figure 1 A deep feature fingerprint-based streaming media multicast tamper-proofing method provided for an embodiment of the present application;
[0051] Figure 2 A method for generating a second path fingerprint provided for an embodiment of the present application;
[0052] Figure 3 A method for reconstructing SHA-1 digest and deep feature vector provided for an embodiment of the present application;
[0053] Figure 4 A method for determining whether a TS transport stream has been tampered with provided for an embodiment of the present application;
[0054] Figure 5 A deep feature fingerprint-based streaming media multicast tamper-proofing device provided for an embodiment of the present application;
[0055] Figure 6 A schematic block diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to better understand the technical solutions in the embodiments of the present application, the technical solutions of the present application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. It should be understood that these descriptions are only exemplary, but not used to limit the scope of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the scope of protection of the present application.
[0057] In addition, in the following description, the description of the known structures and technologies is omitted to avoid unnecessary confusion of the concepts disclosed in the present application.
[0058] In the description of the present application, the terms "first", "second", "third" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance. The terms "mounting", "connecting", "connecting" should be broadly understood, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0059] The exemplary embodiments will be described in detail herein, and the examples are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application. Instead, they are only examples of methods and systems consistent with some aspects of the present application as detailed in the appended claims.
[0060] The present application proposes a deep feature fingerprint-based streaming media multicast tamper-proofing method, device, electronic equipment and storage medium, which can help operators to realize efficient and reliable video security distribution in complex network environment.
[0061] Method embodiment
[0062] Figure 1 A deep feature fingerprint-based streaming media multicast tamper-proofing method is provided for the embodiments of the present application.
[0063] In a first aspect of the present application, a deep feature fingerprint-based streaming media multicast tamper-proofing method is proposed, which is applied to a system including a live source side and a terminal side, and combines Figure 1 , including nine steps S1 to S9:
[0064] S1: On the live source side, the original video stream collected is packaged in PES format by a TS multiplexer, and the PES packet is transmitted to the tamper-proof generation module.
[0065] Specifically, in the embodiment of the present application, the original video stream collected is input into a TS multiplexer on the live source side, and the video data is PES packaged according to the MPEG-TS standard; each 188-byte TS packet output by the TS multiplexer is sequentially transmitted to the tamper-proof generation module.
[0066] S2: The tamper-proof generation module analyzes the PES packet header in real time to extract the display timestamp, decoding timestamp and header information, and screens out the PES packet carrying I-frame data after locating the I-frame according to the NALU / VOP flag bit in the header information.
[0067] Specifically, in the embodiment of the present application, the frame type is determined according to the NALU or VOP flag bit in the header information, and the PES packet carrying I-frame data is screened out, and the I-frame data is key frame data.
[0068] S3: SHA-1 hash processing module and deep feature vector extraction submodule are called respectively to generate first and second fingerprints.
[0069] Specifically, in the embodiment of the present application, the method of calling SHA-1 hash processing module and deep feature vector extraction submodule respectively to generate first and second fingerprints f1 and v_norm is introduced in detail below.
[0070] S4: The fusion fingerprint generation submodule merges the first and second fingerprints according to the preset weight to form a comprehensive digital fingerprint, and encapsulates the comprehensive digital fingerprint as a 188-byte special TS packet of PID, and inserts the special TS packet before the corresponding I-frame TS packet, and transmits the original video stream to the terminal side through IP multicast.
[0071] Specifically, in the embodiment of the present application, the above two fingerprints are weighted or concatenated according to the preset weights a and β by the fusion fingerprint generation submodule to form a comprehensive digital fingerprint F_total, the comprehensive digital fingerprint is encapsulated as a special TS packet of PID 0x1FFF, the special TS packet of PID 0x1FFF is 188 bytes, the insufficient part is filled with FFH, and the special TS packet is inserted before the original I-frame TS packet, realizing synchronous multicast transmission of tamper-proof information and video data.
[0072] S5: On the terminal side, the TS demultiplexer analyzes the received TS transport stream and distinguishes and separates the regular video PES packet and the embedded tamper-proof special TS packet according to the PID field, and buffers the regular video PES packet and the embedded tamper-proof special TS packet into the video buffer and the fingerprint buffer respectively.
[0073] S6: The anti-tampering verification module reconstructs the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameters.
[0074] Specifically, in the embodiments of the present application, the method for the anti-tampering verification module to reconstruct the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameters is described in detail below.
[0075] S7: The SHA-1 digest and the deep feature vector are fused into a local comprehensive fingerprint with a preset weight consistent with the live source side.
[0076] Specifically, in the embodiments of the present application, the fusion fingerprint generation submodule fuses the SHA-1 digest f1' and the deep feature vector v_norm' into a local comprehensive fingerprint F_total' with preset weights a and β consistent with the live source side.
[0077] S8: The Hamming distance and the cosine similarity between the local comprehensive fingerprint and the comprehensive digital fingerprint generated and encapsulated by the live source side are obtained by double comparison, and a matching score is obtained.
[0078] Specifically, in the embodiments of the present application, the Hamming distance and the cosine similarity between the local comprehensive fingerprint F_total' and the comprehensive digital fingerprint F_total generated and encapsulated by the live source side are weighted and integrated to obtain a matching score S.
[0079] S9: Whether the TS transport stream is tampered with is determined according to the size of the matching score.
[0080] Specifically, in the embodiments of the present application, the method for determining whether the TS transport stream is tampered with according to the size of the matching score S is described in detail below.
[0081] Further, in the above-mentioned stream media multicast anti-tampering method based on deep feature fingerprints, the SHA-1 hash processing module is called to generate a first route fingerprint, which includes:
[0082] The PES packet data M carrying the I frame data is spliced with the scrambling string k, and SHA1 is calculated to generate a first route fingerprint.
[0083] Specifically, in the embodiments of the present application, the method for generating a first route fingerprint can be represented by the following formula:
[0084] f_sha1 = SHA1 (M + k)
[0085] M represents the input PES packet data, k represents a key for scrambling, that is, a scrambling string, the formula f_sha1 represents that the PES packet data M is spliced with the scrambling k, and then a SHA-1 hash algorithm is called to generate a fixed-length digital fingerprint, that is, a first path fingerprint, so as to obtain a 160-bit digest highly sensitive to any byte modification, generate a fixed-length digital fingerprint, that is, a first path fingerprint, the first path fingerprint strictly reflects the integrity of the input data, and any slight data change will cause a fingerprint difference, so it is suitable for preliminary judgment of whether the data is modified, and with the high sensitivity of SHA-1 and the semantic recognition of deep features, malicious tampering can be accurately captured in a real attack scene, and the false positive rate of normal disturbance can be significantly reduced.
[0086] It should be understood that the modular software modification method only needs to update the TS multiplexer and the TS demultiplexer, without large-scale hardware replacement, and is suitable for multiple brands and multiple types of IPTV platforms.
[0087] Figure 2 A method for generating a second path fingerprint is provided for the embodiments of the application.
[0088] Further, in the above-mentioned stream media multicast tamper-proofing method based on deep feature fingerprints, a deep feature vector extraction submodule is called to generate a second path fingerprint, combined with Figure 2 , including two steps S21 to S22:
[0089] S21: decode the PES packet data corresponding to the I frame data into a pixel image, input the lightweight convolutional neural network, and extract a high-dimensional semantic feature vector through forward reasoning;
[0090] S22: perform L2 norm normalization on the high-dimensional semantic feature vector to obtain a second path fingerprint.
[0091] Specifically, in the embodiments of the application, the PES packet data corresponding to the I frame data is decoded into a pixel image X, The feature is extracted through a pre-trained model lightweight convolutional neural network F(·), where F(X) represents the forward propagation process of the lightweight convolutional neural network, and each layer of convolution, activation, and pooling operations jointly construct a high-level semantic feature representation The lightweight convolutional neural network can be MobileNet or a customized lightweight CNN.
[0092] The extracted feature vector v is normalized to obtain a second path fingerprint, so as to ensure the consistency and robustness of feature comparison, and the commonly used L2 norm normalization processing is as follows:
[0093]
[0094] Here, by lightening the convolutional neural network structure and GPU / TPU acceleration, high-precision tamper-proof detection is realized while ensuring millisecond-level delay (about 35 ms / frame) and considerable throughput (30 frames / second).
[0095] Further, in the above-mentioned deep feature fingerprint-based streaming media multicast tamper-proofing method, the fusion fingerprint generation submodule combines the first fingerprint and the second fingerprint according to a preset weight to form a comprehensive digital fingerprint, comprising:
[0096] The first fingerprint and the second fingerprint are weighted or concatenated according to preset weights alpha and beta to form a comprehensive digital fingerprint.
[0097] Wherein, alpha and beta are initially set to 0.5, and are subsequently optimized according to actual conditions.
[0098] Specifically, in the embodiment of the present application, the digital fingerprint f1 calculated by SHA1 is fused with the normalized depth feature vector v_norm, and the fusion method is concatenation or weighted combination.
[0099] Figure 3 A method for reconstructing SHA-1 digest and depth feature vector is provided for the embodiment of the present application.
[0100] Further, in the above-mentioned deep feature fingerprint-based streaming media multicast tamper-proofing method, the tamper-proofing verification module reconstructs SHA-1 digest and depth feature vector based on the scrambling string and the depth model parameters, comprising five steps S31 to S35:
[0101] S31: Read the PES packet data carrying I-frame data corresponding to the embedded tamper-proofing dedicated TS packet from the video buffer in sequence.
[0102] Specifically, in the embodiment of the present application, the key frame PES data M' corresponding to the dedicated TS packet is read from the video buffer in sequence.
[0103] S32: Perform SHA-1 digest operation on the PES packet data carrying I-frame data corresponding to the embedded tamper-proofing dedicated TS packet based on the scrambling string k and the depth model parameters, to generate SHA-1 digest.
[0104] Specifically, in the embodiment of the present application, SHA-1 digest operation is performed on M' to generate real-time fingerprint f1.
[0105] S33: Decode the PES packet data carrying I-frame data corresponding to the embedded tamper-proofing dedicated TS packet into a pixel image.
[0106] Specifically, in the embodiment of the present application, the PES packet data carrying I-frame data corresponding to the embedded tamper-proofing dedicated TS packet is decoded into a pixel image X'.
[0107] S34: inputting the pixel image into the lightweight convolutional neural network to extract a high-level semantic feature vector.
[0108] Specifically, in the embodiment of the present application, the pixel image X' is input into the lightweight convolutional neural network F(·) to extract a high-level semantic feature vector v'.
[0109] S35: performing L2 normalization on the high-level semantic feature vector to obtain a deep feature vector.
[0110] Specifically, in the embodiment of the present application, L2 normalization is performed on the high-level semantic feature vector v' to obtain a deep feature vector v_norm'.
[0111] Figure 4 A method for determining whether a TS transmission stream is tampered is provided in the embodiment of the present application.
[0112] Further, in the above-mentioned stream media multicast anti-tampering method based on a deep feature fingerprint, whether the TS transmission stream is tampered is determined according to the size of the matching score, and the combination of the deep feature vector and the high-level semantic feature vector can improve the accuracy of the anti-tampering method. Figure 4 , comprising two steps of S41 to S42:
[0113] S41: comparing the size of the matching score with a tolerance threshold defined by a pre-established uncertainty propagation model;
[0114] S42: when the matching score is greater than the tolerance threshold defined by the pre-established uncertainty propagation model, it is determined that the TS transmission stream is tampered.
[0115] Specifically, in the embodiment of the present application, the matching score S is compared with the tolerance threshold T defined by the pre-established uncertainty propagation model, and when the matching score S is greater than T, it is determined that the TS transmission stream is possibly tampered.
[0116] Further, the above-mentioned stream media multicast anti-tampering method based on a deep feature fingerprint further comprises:
[0117] When the TS transmission stream is tampered, an alarm information is sent to the upper layer application through an alarm interface, and a remedial measure is triggered according to the operator strategy;
[0118] The remedial measure at least includes automatic retransmission, source verification or user prompt.
[0119] Specifically, in the embodiment of the present application, when the TS transmission stream is tampered, an alarm information is immediately sent to the upper layer application through an alarm interface, and an automatic retransmission, a source verification or a user prompt and other remedial measures can be triggered according to the operator strategy, so that an end-to-end, real-time and double-fingerprint anti-tampering verification from the receiving end to the application layer is realized.
[0120] Device embodiment
[0121] Figure 5 A deep feature fingerprint-based streaming media multicast tamper-proofing device is provided for the embodiment of the present application.
[0122] The second aspect of the present application also proposes a deep feature fingerprint-based streaming media multicast tamper-proofing device applied in a system including a live source side and a terminal side, combining Figure 5 , comprising:
[0123] The packing module 51 is configured to pack the collected original video stream in the PES format through a TS multiplexer at the live source side, and transmit the PES packet to the tamper-proofing generation module.
[0124] Specifically, in the embodiment of the present application, the packing module 51 inputs the collected original video stream into the TS multiplexer at the live source side, and packs the video data in the PES format according to the MPEG-TS standard; each 188-byte TS packet output by the TS multiplexer is sequentially transmitted to the tamper-proofing generation module.
[0125] The tamper-proofing generation module 52 is configured to analyze the PES packet header in real time to extract the display timestamp, decoding timestamp and header information, and screen out the PES packet carrying the I-frame data in combination with the NALU / VOP flag bit in the header information.
[0126] Specifically, in the embodiment of the present application, the tamper-proofing generation module 52 determines the frame type in combination with the NALU or VOP flag bit in the header information, and screens out the PES packet carrying the I-frame data.
[0127] The SHA-1 hash processing module 53 and the deep feature vector extraction submodule 54 are configured to generate the first fingerprint and the second fingerprint, respectively.
[0128] Specifically, in the embodiment of the present application, the SHA-1 hash processing module 53 and the deep feature vector extraction submodule 54 are respectively called to generate the first fingerprint f1 and the second fingerprint v_norm.
[0129] The fusion fingerprint generation submodule 55 is configured to merge the first fingerprint and the second fingerprint to form a comprehensive digital fingerprint according to a preset weight, encapsulate the comprehensive digital fingerprint as a 188-byte special TS packet of PID, insert the special TS packet before the corresponding I-frame TS packet, and transmit the original video stream and the special TS packet to the terminal side through IP multicast.
[0130] Specifically, in the embodiment of the present application, the two fingerprints are weighted or connected in series by the fusion fingerprint generation submodule 55 according to the preset weights a, b to form a comprehensive digital fingerprint F_total, the comprehensive digital fingerprint is packaged into a special TS package with PID 0x1FFF, the special TS package with PID 0x1FFF is 188 bytes, the insufficient part is filled with FFH, and the special TS package is inserted before the original I frame TS package to realize synchronous multicast transmission of the tamper-proof information and the video data.
[0131] The receiving module 56 is configured to parse the received TS transport stream by a TS demultiplexer, distinguish and separate the regular video PES package and the embedded tamper-proof special TS package according to the PID field, and buffer the regular video PES package and the embedded tamper-proof special TS package into a video buffer and a fingerprint buffer, respectively.
[0132] The tamper-proof checking module 57 is configured to reconstruct the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameter.
[0133] Specifically, in the embodiment of the present application, the tamper-proof checking module 57 reconstructs the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameter.
[0134] The fusion fingerprint generation submodule 55 is configured to fuse the SHA-1 digest and the deep feature vector into a local comprehensive fingerprint with the preset weight consistent with the live source side.
[0135] Specifically, in the embodiment of the present application, the fusion fingerprint generation submodule 55 fuses the SHA-1 digest f1' and the deep feature vector v_norm' into a local comprehensive fingerprint F_total' with the preset weights a, b consistent with the live source side.
[0136] The double comparison module 58 is configured to double compare the Hamming distance and the cosine similarity of the local comprehensive fingerprint and the comprehensive digital fingerprint generated and packaged by the live source side extracted from the fingerprint buffer to obtain a matching score.
[0137] Specifically, in the embodiment of the present application, the double comparison module 58 double compares the Hamming distance and the cosine similarity of the local comprehensive fingerprint F_total' and the comprehensive digital fingerprint F_total generated and packaged by the live source side extracted from the fingerprint buffer to obtain a matching score S by weighted synthesis.
[0138] The judgment module 59 is configured to judge whether the TS transport stream is tampered with according to the size of the matching score.
[0139] Specifically, in the embodiment of the present application, the judgment module 59 judges whether the TS transport stream is tampered with according to the size of the matching score S.
[0140] The third aspect of the present application further provides an electronic device, comprising: a processor and a memory;
[0141] The processor is configured to execute any one of the methods for deep feature fingerprint based streaming media multicast tamper prevention method as described above by invoking programs or instructions stored in the memory.
[0142] The fourth aspect of the present application further provides a computer readable storage medium storing programs or instructions, which cause a computer to execute any one of the methods for deep feature fingerprint based streaming media multicast tamper prevention method as described above.
[0143] Figure 6 is a schematic block diagram of an electronic device provided by an embodiment of the present application.
[0144] As shown in Figure 6 , the electronic device comprises at least one processor 601, at least one memory 602 and at least one communication interface 603. Each component in the electronic device is coupled together through a bus system 604. The communication interface 603 is configured to transmit information between the electronic device and external devices. It can be understood that the bus system 604 is configured to realize the connection communication between the components. In addition to the data bus, the bus system 604 also includes power bus, control bus and status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as bus system 604 in Figure 6 .
[0145] It can be understood that the memory 602 in the embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
[0146] In some embodiments, the memory 602 stores the following elements, executable units or data structures, or a subset of them, or an extended set of them: an operating system and an application program.
[0147] The operating system includes various system programs, such as framework layer, core library layer, driver layer, etc., for realizing various basic services and processing hardware-based tasks. The application program includes various application programs, such as media player (Media Player), browser (Browser), etc., for realizing various application services. The programs for implementing any one of the methods for deep feature fingerprint based streaming media multicast tamper prevention method provided by the embodiment of the present application can be included in the application program.
[0148] In the embodiments of the present application, the processor 601 calls the program or instruction stored in the memory 602, specifically, the program or instruction stored in the application program, and the processor 601 is configured to execute the steps of each embodiment of the deep feature fingerprint-based streaming media multicast tamper-proofing method provided by the embodiments of the present application.
[0149] At the live source side, the collected original video stream is packaged in PES format by a TS multiplexer, and the PES packet is transmitted to a tamper-proofing generation module;
[0150] The tamper-proofing generation module analyzes the PES packet header in real time to extract the display timestamp, decoding timestamp and header information, and screens out the PES packet carrying the I frame data after locating the I frame according to the NALU / VOP flag bit in the header information;
[0151] The SHA-1 hash processing module and the deep feature vector extraction submodule are called respectively to generate the first fingerprint and the second fingerprint;
[0152] The fusion fingerprint generation submodule merges the first fingerprint and the second fingerprint according to a preset weight to form a comprehensive digital fingerprint, encapsulates the comprehensive digital fingerprint into a 188-byte special TS packet of PID, inserts the special TS packet before the corresponding I frame TS packet, and transmits the original video stream and the special TS packet to the terminal side through IP multicast;
[0153] At the terminal side, the TS demultiplexer analyzes the received TS transmission stream, distinguishes and separates the regular video PES packet and the embedded tamper-proofing special TS packet according to the PID field, and buffers the regular video PES packet and the embedded tamper-proofing special TS packet into the video buffer and the fingerprint buffer, respectively;
[0154] The tamper-proofing checking module reconstructs the SHA-1 digest and the deep feature vector based on the scrambling string and the deep model parameters;
[0155] The SHA-1 digest and the deep feature vector are fused as a local comprehensive fingerprint with a preset weight consistent with the live source side;
[0156] The local comprehensive fingerprint and the comprehensive digital fingerprint generated and encapsulated by the live source side are double-compared to obtain a matching score according to the Hamming distance and the cosine similarity;
[0157] According to the size of the matching score, it is judged whether the TS transmission stream is tampered with.
[0158] Any method in the deep feature fingerprint-based streaming media multicast anti-tampering method provided in this embodiment of the invention can be applied to, or implemented by, processor 601. Processor 601 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 601 or by instructions in software form. Processor 601 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor.
[0159] The steps of any method in the streaming media multicast anti-tampering method based on deep feature fingerprinting provided in this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software units in the decoding processor. The software units can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 602, and processor 601 reads the information in memory 602 and combines it with hardware to complete the steps of the method.
[0160] Those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of the invention and form different embodiments.
[0161] Those skilled in the art will understand that the descriptions of the various embodiments have different focuses, and for parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention. All such modifications and variations fall within the scope defined by the appended claims. The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0163] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for preventing tampering in streaming media multicast based on deep feature fingerprints, characterized in that, Applied to systems including both the live streaming source side and the terminal side, including: On the live stream source side, the acquired raw video stream is packaged in PES format using a TS multiplexer, and the PES package is then passed to the anti-tampering generation module. The anti-tampering generation module parses the PES packet header in real time to extract the display timestamp, decoding timestamp and header information. After locating the I-frame by combining the NALU / VOP flag in the header information, it filters out the PES packets carrying I-frame data. The SHA-1 hash processing module and the deep feature vector extraction submodule are called respectively to generate the first fingerprint and the second fingerprint; The fingerprint generation submodule merges the first fingerprint and the second fingerprint according to a preset weight to form a comprehensive digital fingerprint, and encapsulates the comprehensive digital fingerprint into a 188-byte dedicated TS packet with a PID. The dedicated TS packet is inserted before the corresponding I-frame TS packet and transmitted to the terminal side via IP multicast along with the original video stream. On the terminal side, the TS demultiplexer parses the received TS transport stream and distinguishes and separates the regular video PES packets and the embedded anti-tampering special TS packets according to the PID field. The regular video PES packets and the embedded anti-tampering special TS packets are cached in the video buffer and the fingerprint buffer, respectively. The anti-tampering verification module reconstructs the SHA-1 digest and deep feature vector based on the scrambling string and deep model parameters; The fingerprint generation submodule fuses the SHA-1 digest and the deep feature vector into a local comprehensive fingerprint using a preset weight consistent with the live stream source side. A matching score is obtained by comparing the Hamming distance and cosine similarity between the local integrated fingerprint and the integrated digital fingerprint generated and encapsulated from the live source side extracted from the fingerprint cache. The TS transport stream is determined to have been tampered with based on the matching score.
2. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, The first fingerprint is generated by calling the SHA-1 hash processing module, including: The first fingerprint is generated by concatenating the PES packet data M carrying I-frame data with the scrambling string k and calculating SHA1.
3. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, The deep feature vector extraction submodule is invoked to generate the second fingerprint, including: The PES packet data corresponding to the I-frame data is decoded into pixel images, input into a lightweight convolutional neural network, and high-dimensional semantic feature vectors are extracted through forward inference. The high-dimensional semantic feature vector is normalized using the L2 norm to obtain the second fingerprint.
4. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, The fingerprint fusion generation submodule merges the first and second fingerprints according to preset weights to form a comprehensive digital fingerprint, including: A comprehensive digital fingerprint is formed by weighting or concatenating the first and second fingerprints according to preset weights α and β. α and β are initially set to 0.5 each, and will be adjusted later according to the actual situation.
5. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, The anti-tampering verification module reconstructs the SHA-1 digest and deep feature vector based on the scrambling string and deep model parameters, including: Read PES packet data carrying I-frame data sequentially from the video buffer and the embedded anti-tampering dedicated TS packet; Based on the scrambling string k and the depth model parameters, SHA-1 digest operation is performed on the PES packet data carrying I-frame data corresponding to the embedded anti-tampering dedicated TS packet to generate SHA-1 digest; Decode the PES packet data carrying I-frame data corresponding to the embedded tamper-proof dedicated TS packet into a pixel image; The pixel image is input into a lightweight convolutional neural network to extract high-level semantic feature vectors; L2 normalization is performed on the high-level semantic feature vector to obtain the deep feature vector.
6. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, Determining whether the TS transport stream has been tampered with based on the matching score includes: Compare the matching score with the tolerance threshold defined by the pre-established uncertainty propagation model; When the matching score is greater than the tolerance threshold defined by the pre-established uncertainty propagation model, it is determined that the TS transport stream has been tampered with.
7. The method for preventing tampering with streaming media multicast based on deep feature fingerprints according to claim 1, characterized in that, The method further includes: When the TS transport stream is tampered with, an alarm message is sent to the upper layer application through the alarm interface and remedial measures are triggered according to the operator's policy. The remedial measures include at least: automatic retransmission, origin verification, or user prompts.
8. A streaming media multicast anti-tampering device based on deep feature fingerprinting, characterized in that, Applied to systems including both the live streaming source side and the terminal side, including: Packaging module: Used on the live source side to package the captured raw video stream according to the PES format through the TS multiplexer, and pass the PES package to the anti-tampering generation module; Anti-tampering generation module: used to parse the PES packet header in real time to extract the display timestamp, decoding timestamp and header information, and after locating the I-frame by combining the NALU / VOP flag in the header information, filter out the PES packets carrying I-frame data; The SHA-1 hash processing module and the deep feature vector extraction submodule are used to generate the first fingerprint and the second fingerprint, respectively. The fingerprint generation submodule is used to merge the first fingerprint and the second fingerprint according to the preset weight to form a comprehensive digital fingerprint, and encapsulate the comprehensive digital fingerprint into a 188-byte dedicated TS packet of PID. The dedicated TS packet is inserted before the corresponding I-frame TS packet and transmitted to the terminal side via IP multicast with the original video stream. The receiving module is used on the terminal side. The TS demultiplexer parses the received TS transport stream and distinguishes and separates the regular video PES packets and the embedded anti-tamper special TS packets according to the PID field. The regular video PES packets and the embedded anti-tamper special TS packets are cached in the video buffer and the fingerprint buffer, respectively. Anti-tampering verification module: used to reconstruct SHA-1 digest and deep feature vector based on scrambling string and deep model parameters; Fingerprint generation submodule: used to fuse the SHA-1 digest and the deep feature vector into a local comprehensive fingerprint with preset weights consistent with those of the live stream source side; Dual comparison module: used to compare the Hamming distance and cosine similarity of the local integrated fingerprint and the integrated digital fingerprint generated and encapsulated from the live source side extracted from the fingerprint cache to obtain a matching score; Judgment module: used to determine whether the TS transport stream has been tampered with based on the magnitude of the matching score.
9. An electronic device, characterized in that, include: Processor and memory; The processor executes a streaming media multicast anti-tampering method based on deep feature fingerprints as described in any one of claims 1 to 7 by calling the program or instructions stored in the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to execute a streaming media multicast anti-tampering method as described in any one of claims 1 to 7.