Video key domain anti-counterfeiting and detection method and device, and computer program product

By steganizing the md5 code value and position information of the video keyframe to the DCT domain intermediate frequency, the problem of low anti-counterfeiting embedding and detection efficiency caused by high redundancy of video keyframes is solved, and efficient video anti-counterfeiting processing is achieved.

CN119963982APending Publication Date: 2025-05-09中国邮政储蓄银行股份有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510058256.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing technology has not conducted a coordinated analysis of video keyframes, resulting in high redundancy of keyframes and low anti-counterfeiting embedding and detection efficiency.

Method used

By extracting the md5 code value and position information of the video frame and steganography on the DCT domain intermediate frequency of the corresponding frame, an anti-counterfeiting information chain is formed to ensure the continuity and integrity of the information.

Benefits of technology

It effectively reduces the number of keyframes, improves the anti-counterfeiting embedding and detection efficiency, reduces the damage to video quality, and enhances the anti-counterfeiting ability of video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963982A_ABST
    Figure CN119963982A_ABST
Patent Text Reader

Abstract

The invention provides an anti-counterfeiting and detection method and device for a video key domain, and the method comprises the steps: a first extraction step: extracting a current key frame, and extracting an md5 code value and position information of the current key frame, a key frame set being a set formed by key frames extracted from all video frames; steganography: steganography the md5 code value and the position information of the current key frame to the intermediate frequency of the DCT domain of the corresponding frame; the next key frame of the current key frame is updated to be a new current key frame, the first extraction step and the steganography step are sequentially repeated for at least one time until the current key frame is the last key frame, the md5 code value and the position information of the current key frame are steganographically written to a tail frame, and the tail frame is the last video frame in all the video frames. The problem that in the prior art, key frames are not comprehensively analyzed, and the redundancy of the key frames is high, so that anti-counterfeiting embedding and anti-counterfeiting detection efficiency is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video anti-counterfeiting and detection technology, and in particular to an anti-counterfeiting and detection method for a key video domain, an anti-counterfeiting and detection device for a key video domain, a computer-readable storage medium, and a computer program product. Background Art

[0002] The bank has achieved round-the-clock monitoring of the bank through the video surveillance system to assist the bank's security management. As a result, a large amount of video information has been accumulated. During daily inspections, relevant video information can be provided for evidence collection. In order to prevent the supporting video from being intercepted and tampered with, the video needs to be processed for anti-counterfeiting.

[0003] The current mainstream anti-counterfeiting methods are mostly to embed the same anti-counterfeiting information into each frame of the video, so as to achieve the purpose of identifying whether the video frame has been tampered with, but it is not effective to identify whether the video frame has been intercepted, and because each frame needs to be embedded, it is time-consuming and has a great damage to the video. According to the characteristics of bank video monitoring and detection, in the provided forensic video, most of the useful key domain information is key frames, which are very few compared to the number of frames occupied by the entire forensic video, that is, the current mainstream anti-counterfeiting strategy has caused too much redundant information.

[0004] In order to protect the video from malicious operations such as tampering and deletion, it is often necessary to perform anti-counterfeiting processing on the video to ensure the correctness of the video. Video anti-counterfeiting information is usually divided into two categories. Image steganography is a technology used to hide information, in which secret information is embedded in the image so that only specific people can read it. The purpose of image steganography technology is to protect the confidentiality of data. In contrast, image watermarking technology is a method used for anti-counterfeiting, copyright protection, and image authentication. A watermark is a type of information added to an image so that the image can still be identified when it is copied or modified, and can be used to prove its ownership or source. Based on the above situation, the current anti-counterfeiting information embedding steps are divided into the following types: the first is to traverse the video frame, embed the hidden anti-counterfeiting information or anti-counterfeiting watermark directly into the inconspicuous area of ​​the surface pixels of the video frame, so as to achieve the purpose of anti-counterfeiting. This method is simple and easy to implement, but the robustness is poor; the second is to traverse the video frame, first perform transform domain processing (such as frequency domain, wavelet domain) on the video frame, and embed the hidden anti-counterfeiting information or anti-counterfeiting watermark into the transform domain of the video frame. This method is more complex, but has better robustness and hiding; the third is to create an independent anti-counterfeiting frame, and add a variety of anti-counterfeiting information to the anti-counterfeiting frame according to its own requirements, and then directly insert the anti-counterfeiting frame at a random or specific position in the video. This method has a lot of anti-counterfeiting information and is easy to implement, but the robustness is poor and it is easy to affect the normal video perception. The key information of bank security video is often the embodiment of abnormal situations, concentrated on a small number of frames, most of the video information is non-key information, and a few are key frames. Both of these focus on tampering anti-counterfeiting, and there is less anti-counterfeiting for interception of key frame information; and the protection of key frame information is not strong enough.

[0005] In video anti-counterfeiting, anti-counterfeiting information is usually added to the video frame. When the video is modified, the anti-counterfeiting information is destroyed. Therefore, the anti-counterfeiting information ensures the effectiveness of detecting frame tampering, but the information cannot ensure the effective detection of video frame interception. If video anti-counterfeiting adds anti-counterfeiting information to each frame, it ignores the characteristics of the video itself and has a high overhead. The key frame extraction method based on color, shape, lens, etc. extracts a large amount of information and lacks the necessary feature optimization, resulting in low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection; and the key frames are not comprehensively analyzed, resulting in high redundancy of key frames. Summary of the invention

[0006] The main purpose of the present application is to provide a method for anti-counterfeiting and detection of key video domains, an anti-counterfeiting and detection device for key video domains, a computer-readable storage medium and a computer program product, so as to at least solve the problem in the prior art that no overall analysis of key frames is performed, and the key frame redundancy is high, resulting in low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection.

[0007] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for anti-counterfeiting and detection of a key domain of a video is provided, comprising: a first extraction step, extracting a current key frame, and extracting an MD5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; a steganographic step, steganographically writing the MD5 code value and the position information of the current key frame to a DCT domain intermediate frequency of a corresponding frame, wherein the corresponding frame is the next key frame of the current key frame in the key frame set; updating the next key frame of the current key frame to a new current key frame, and repeating the first extraction step and the steganographic step in sequence at least once until the current key frame is the last key frame, and steganographically writing the MD5 code value and the position information of the current key frame to a tail frame, wherein the tail frame is the last video frame of all the video frames.

[0008] Optionally, before extracting the current key frame, the method also includes: extracting image color moments and edge operators corresponding to all the video frames; multiplying the image color moments and edge operators of each video frame with corresponding weights to form fusion features corresponding to the video frames; extracting principal components from the fusion features of each video frame using a principal component analysis method to reduce the dimension of the fusion features to obtain reduced dimension features corresponding to each video frame; performing cluster analysis based on all the reduced dimension features using a K-means clustering algorithm to obtain the key frame set.

[0009] Optionally, extracting the image color moments and edge operators corresponding to all the video frames includes: a second extraction step of extracting the first-order moment, the second-order moment and the third-order moment in the current video frame; a composition step of using the first-order moment, the second-order moment and the third-order moment to compose the image color moment of the current video frame; a smoothing step of using a Gaussian filter to smooth the corresponding current video frame according to the image color moment to obtain a corresponding smoothed video frame; a calculation step of using a first-order partial derivative finite difference to calculate the gradient amplitude and gradient direction of the smoothed video frame; a suppression step of performing a non-maximum suppression operation on the smoothed video frame according to the gradient amplitude and the gradient direction to obtain the amplitude of each pixel in the smoothed video frame; a detection step of using a double threshold algorithm to detect the amplitude of each pixel in the smoothed video frame to generate the edge operator of the current video frame; and repeating the second extraction step, the composition step, the smoothing step, the calculation step, the suppression step and the detection step in sequence at least once until the image color moments and the edge operators corresponding to all the video frames are extracted.

[0010] Optionally, the principal component analysis method is used to extract the principal component of the fusion feature of each video frame to reduce the dimension of the fusion feature to obtain the reduced dimension feature corresponding to each video frame, including: solving the covariance matrix corresponding to the fusion feature of each video frame according to a first formula, and the first formula is m represents the total number of video frames, x i represents the i-th video frame, Represents the mean vector of the fusion features of all the video frames, C represents the covariance matrix; the eigenvalue decomposition method is used to calculate the eigenvalues ​​and eigenvectors of each covariance matrix, one eigenvalue corresponds to one eigenvector; the eigenvectors corresponding to the first k eigenvalues ​​are selected in descending order to form a matrix, and the dimensionality reduction features corresponding to each video frame are obtained.

[0011] Optionally, after extracting the current key frame and extracting the md5 code value and position information of the current key frame, the method further includes: when the current key frame is the first key frame in the key frame set, extracting the total number of key frames and the position information of the next key frame of the current key frame; and steganographically writing the md5 code value of the current key frame, the position information, the total number of key frames and the position information of the next key frame of the current key frame to the DCT domain intermediate frequency of the header frame, the header frame being the first video frame among all the video frames.

[0012] Optionally, the md5 code value and the position information of the current key frame are steganographically written into the DCT domain intermediate frequency of the corresponding frame, including: performing a discrete cosine transform on the corresponding frame to obtain the transformed DCT domain data of the corresponding frame; and steganographically writing the md5 code value and the position information of the current key frame into a specified position of the DCT domain intermediate frequency of the corresponding frame according to the DCT domain data.

[0013] Optionally, after the md5 code value and the position information of the current key frame are steganographically written to the tail frame, the method further comprises: an inverse transformation step, performing an inverse discrete cosine transform on the steganographic corresponding frame, extracting the md5 code value and the position information steganographically written in the steganographic corresponding frame, the steganographic corresponding frame is any video frame in a steganographic key frame set, the steganographic key frame set is a set formed by all the video frames in which the md5 code value and the position information of the key frame are steganographically written, the steganographic key frame set comprises a head frame, all the key frames and the tail frame, the head frame being the first video frame among all the video frames; a searching step, searching the corresponding key frame according to the position information, obtaining a target key frame, and extracting the md5 code value of the target key frame; a comparing step, comparing the md5 code value steganographically written in the steganographic corresponding frame with the md5 code value of the target key frame; and determining Step, when the md5 code value of the steganographic corresponding frame is inconsistent with the md5 code value of the target key frame, determine that the target key frame of the video has been tampered with; repeating step, when the md5 code value of the steganographic corresponding frame is consistent with the md5 code value of the target key frame, update the target key frame to a new steganographic corresponding frame, and repeat the inverse transformation step, the search step, the comparison step, the determination step and the repetition step at least once in sequence until all the key frames complete the video detection; obtain the detection number, the detection number is the number of the key frames participating in the video detection; when the detection number is less than the total number of key frames, determine that the key frame of the video has been intercepted; when the detection number is equal to the total number of key frames and there is no tampering problem, determine that the video is in a safe state.

[0014] According to another aspect of the present application, an anti-counterfeiting and detection device for a key domain of a video is provided, the device comprising: a first extraction unit, used to perform a first extraction step, extract a current key frame, and extract an md5 code value and position information of the current key frame, the current key frame is any key frame in a key frame set, the key frame set is a set formed by the key frames extracted from all video frames; a first steganography unit, used to perform a steganography step, steganographically write the md5 code value and the position information of the current key frame to a DCT domain intermediate frequency of a corresponding frame, the corresponding frame is the next key frame of the current key frame in the key frame set; a first repetition unit, used to update the next key frame of the current key frame to a new current key frame, and repeat the first extraction step and the steganography step in sequence at least once, until the current key frame is the last key frame, and steganographically write the md5 code value and the position information of the current key frame to a tail frame, the tail frame is the last video frame of all the video frames.

[0015] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described.

[0016] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, any one of the methods described above is implemented.

[0017] Applying the technical solution of the present application, in the anti-counterfeiting and detection method of the key domain of the video, first, in the first extraction step, the current key frame is extracted, and the md5 code value and position information of the current key frame are extracted, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; then, in the steganographic step, the md5 code value and the position information of the current key frame are steganographically written to the DCT domain intermediate frequency of the corresponding frame, and the corresponding frame is the next key frame of the current key frame in the key frame set; finally, the next key frame of the current key frame is updated to the new current key frame, and the first extraction step and the steganographic step are repeated at least once in sequence until the current key frame is the last key frame, and the md5 code value and the position information of the current key frame are steganographically written to the tail frame, and the tail frame is the last video frame of all the video frames. This application adopts the md5 code value of the key frame to be steganographically written into the image DCT domain. Considering the key frame extraction efficiency, key frame redundancy filtering and interception anti-counterfeiting, the md5 code value of the previous key frame is sequentially steganographically written into the image DCT domain intermediate frequency of the next key frame (embedded into the rear X tail of the floating point of the specified position group), and finally the md5 code value of the last key frame is steganographically written into the tail frame, which also reduces the damage of the video itself to the excessive embedded information. This application solves the problem that the key frames are not comprehensively analyzed in the prior art, and the high redundancy of the key frames leads to low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A hardware structure block diagram of a mobile terminal for executing an anti-counterfeiting and detection method for a key video domain provided in an embodiment of the present application is shown;

[0019] Figure 2 A schematic flow chart of a method for anti-counterfeiting and detecting a key domain of a video provided in accordance with an embodiment of the present application is shown;

[0020] Figure 3 A schematic diagram of a process of steganography of anti-counterfeiting information in a key domain of a video provided according to an embodiment of the present application is shown;

[0021] Figure 4 A schematic diagram of a process for detecting anti-counterfeiting information in a key domain of a video provided according to an embodiment of the present application is shown;

[0022] Figure 5 A structural block diagram of an anti-counterfeiting and detection device for a key domain of a video provided according to an embodiment of the present application is shown.

[0023] The above drawings include the following reference numerals:

[0024] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION

[0025] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] As introduced in the background technology, in the prior art, if anti-counterfeiting information is added to each frame of the video, the characteristics of the video itself are ignored, and the cost is relatively high. However, the key frame extraction method based on color, shape, lens, etc. extracts a large amount of information, lacks the necessary feature optimization, and the efficiency of anti-counterfeiting embedding and anti-counterfeiting detection is low; and the key frames are not comprehensively analyzed, and the key frame redundancy is high. In order to solve the problem that the key frames are not comprehensively analyzed in the prior art, the key frame redundancy is high, resulting in low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection, the embodiments of the present application provide a video key domain anti-counterfeiting and detection method, video key domain anti-counterfeiting and detection device, computer-readable storage medium and computer program product.

[0029] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0030] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal of a method for anti-counterfeiting and detecting a key video domain according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0031] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as computer programs corresponding to the anti-counterfeiting and detection methods of the key domain of the video in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, the above-mentioned method is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific examples of the network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0032] In this embodiment, a method for anti-counterfeiting and detection of key video domains running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0033] Figure 2 FIG. 1 is a flow chart of a method for anti-counterfeiting and detecting a key domain of a video according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0034] Step S201, the first extraction step, extracts the current key frame, and extracts the md5 code value and position information of the current key frame, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames.

[0035] Specifically, a set of key frames is determined from a video frame sequence through a preset feature analysis and clustering algorithm. A key frame is selected as the current key frame, and the MD5 code value of the current key frame is calculated, which usually involves hashing the binary representation of the key frame pixel value to obtain a hash code of a fixed length as the unique identifier of the key frame. The position information of the current key frame in the video sequence is recorded, which can be a frame index or a timestamp. By extracting the MD5 code value and position information of the key frame, key anti-counterfeiting identification and positioning clues are provided for subsequent steganography and detection. The MD5 code value can ensure the integrity and invariance of the key frame, while the position information can be used to locate and restore the key frame in the video, which is crucial for detecting whether the key frame has been intercepted and tampered with. In large-scale video surveillance systems, the accurate and timely extraction of this information is the basis for achieving effective video anti-counterfeiting and forensic support.

[0036] Step S202, a steganographic step, steganographically writes the md5 code value of the current key frame and the position information onto the DCT domain intermediate frequency of the corresponding frame, where the corresponding frame is the next key frame of the current key frame in the key frame set.

[0037] Specifically, the next key frame is subjected to discrete cosine transform (DCT) and converted to the DCT domain. Then, a specific position group of the DCT domain intermediate frequency (usually an area with less visual impact on the image) is selected, and the md5 code value and position information of the current key frame are encoded into a format suitable for embedding, such as a binary string. Next, the encoded information is embedded into the low bits of the selected DCT coefficients, and the information is hidden by modifying these bits of the DCT coefficients. Finally, the DCT domain data is inversely transformed to obtain a video frame embedded with anti-counterfeiting information. By embedding anti-counterfeiting information in the DCT domain intermediate frequency, the md5 code value and position information of the key frame can be hidden in the video while ensuring the video quality, as a key basis for subsequent detection and recovery. This method avoids the visual distortion caused by embedding information directly in image pixels, reduces the risk of accidental destruction of anti-counterfeiting information, and improves the robustness and security of the anti-counterfeiting mechanism. For large-scale video surveillance systems, this embedding method can achieve effective anti-counterfeiting of video key frames without significantly increasing storage and transmission overhead, thereby enhancing the overall security and reliability of the system.

[0038] Step S203, updating the next key frame of the current key frame as the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, and steganographically writing the md5 code value and the position information of the current key frame to the tail frame, which is the last video frame among all the video frames.

[0039] Specifically, the next key frame is then updated to the new current key frame, and the first extraction step and the steganographic step are repeated until all key frames are processed. When the last key frame is processed, the md5 code value and position information of the key frame are extracted and steganographically written to the tail frame of the video sequence as the termination mark of the entire video anti-counterfeiting information chain. The construction of the anti-counterfeiting information chain is realized, and the information chain not only connects all key frames, but also ensures the continuity and integrity of the information. This mechanism can effectively detect the interception and tampering of key frames, and provides strong support for the anti-counterfeiting of video surveillance systems. During the detection process, the key frames can be quickly located and verified through the serialized key frame md5 code value and position information, which is of great significance for improving the security of the video surveillance system and the legal effect of video evidence. In addition, the process and automation of information chain construction reduce the need for manual intervention, improve the efficiency and consistency of video anti-counterfeiting processing, and are suitable for real-time processing of continuous video streams and batch processing of large-scale video data.

[0040] In this embodiment, first, in the first extraction step, the current key frame is extracted, and the md5 code value and position information of the current key frame are extracted, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; then, in the steganographic step, the md5 code value and the position information of the current key frame are steganographically written to the DCT domain intermediate frequency of the corresponding frame, and the corresponding frame is the next key frame of the current key frame in the key frame set; finally, the next key frame of the current key frame is updated to the new current key frame, and the first extraction step and the steganographic step are repeated at least once in sequence until the current key frame is the last key frame, and the md5 code value and the position information of the current key frame are steganographically written to the tail frame, and the tail frame is the last video frame of all the video frames. This application adopts the md5 code value of the key frame to be steganographically written into the image DCT domain. Considering the key frame extraction efficiency, key frame redundancy filtering and interception anti-counterfeiting, the md5 code value of the previous key frame is sequentially steganographically written into the image DCT domain intermediate frequency of the next key frame (embedded into the rear X tail of the floating point of the specified position group), and finally the md5 code value of the last key frame is steganographically written into the tail frame, which also reduces the damage of the video itself to the excessive embedded information. This application solves the problem that the key frames are not comprehensively analyzed in the prior art, and the high redundancy of the key frames leads to low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection.

[0041] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the anti-counterfeiting and detection method of the key domain of the video of the present application will be described in detail below in combination with specific embodiments.

[0042] In order to reduce the number of key frames and improve the efficiency of key frame detection and anti-counterfeiting information embedding, in an optional implementation manner, before the above step S201, the method further includes:

[0043] Step S301, extracting image color moments and edge operators corresponding to all the above video frames;

[0044] Step S302, forming a fusion feature corresponding to the video frame according to the multiplication of the image color moment of each video frame and the edge operator with the corresponding weight;

[0045] Step S303, extracting principal components from the fusion features of each of the video frames using a principal component analysis method to reduce the dimension of the fusion features to obtain reduced-dimensional features corresponding to each of the video frames;

[0046] Step S304: Perform cluster analysis using a K-means clustering algorithm based on all of the above dimensionality reduction features to obtain the above key frame set.

[0047] In the above embodiment, if Figure 3 As shown, for each video frame, the first-order moment, second-order moment and third-order moment are first calculated to obtain the image color moment. This usually involves traversing each pixel in the frame and calculating its mean, variance and skewness in the RGB color space to form a feature vector that describes the color distribution of the image. Then, an edge detection algorithm, such as Canny edge detection, is applied to each video frame to obtain an edge operator, which involves using a Gaussian filter to smooth the image to remove noise, then calculating the gradient amplitude and direction of the image, performing non-maximum suppression and double threshold detection, and finally generating a binary image to mark the edge area in the video frame. Extracting the image color moment and edge operator of the video frame can capture the color distribution and structural information of the image in the video, providing a basis for subsequent feature fusion and key frame detection. The weights of the image color moment and edge operator are calculated, which are usually determined based on prior knowledge or through experiments to reflect the relative importance of the two features in key frame detection. Then, the image color moment and edge operator feature vector of each video frame are multiplied by their respective weights to obtain a weighted feature vector, as shown in the formula: X = [F color *w1,F canny *w2] constitutes the fusion feature X of the video frame, where w1 and w2 are the fusion weights, F color Represents the image color moment, F cannyRepresents the edge operator. Step S301: Extract the image color moments and edge operators corresponding to all the above video frames. Machine learning methods, such as support vector machines (SVM) or neural networks, can be used to automatically learn the weights of image color moments and edge operators to dynamically adapt to the feature importance of different video contents. In addition, in addition to simple addition, other feature fusion methods, such as feature splicing, feature cascading or deep feature extraction, can be tried to further enhance the expressiveness and discrimination of fused features. By fusing the image color moment and the edge operator, the comprehensive features of the video frame can be obtained, which not only contain color information but also reflect the structural characteristics of the image. The composition of the fused features provides a more comprehensive and rich content description for subsequent dimensionality reduction and clustering analysis. The fused feature vectors of all video frames are combined into a matrix, and then PCA (principal component analysis) is performed on this matrix. This first involves calculating the mean vector and covariance matrix of the fused features, and then performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and eigenvectors. According to the size of the eigenvalues, the first few principal components are selected, and these principal components constitute a low-dimensional space. Finally, the fused feature vector of each video frame is projected onto this low-dimensional space to obtain the reduced-dimensional feature vector. Through PCA dimensionality reduction, the redundancy and noise in the fused features can be removed, and the most representative and discriminative feature dimensions can be extracted, providing an optimized data representation for the selection of key frames. The reduced-dimensional features not only reduce the computational workload and storage requirements of data processing, but also improve the accuracy and efficiency of subsequent clustering and key frame detection. The K-means clustering algorithm is applied to perform cluster analysis on the reduced-dimensional features of all video frames. This usually starts with randomly selecting K initial cluster centers, and then assigning the reduced-dimensional feature vector of each video frame to the nearest cluster center to form K clusters. Next, the center of each cluster is recalculated, and the cluster assignment process is iterated until the clustering results converge. When the similarity between the video frame and the constructed k-means cluster class β reaches the threshold M, it is considered to be a key frame keyp, keyp=pdis(β,p)>M, and all key frames form a key frame set. Through K-means clustering, the dimensionality reduction features of all video frames can be divided into several clusters, each cluster represents a type or state of image content in the video. Selecting key frames from each cluster can ensure that the key information points of the video are covered, while reducing the number of key frames and improving the efficiency of key frame detection and anti-counterfeiting information embedding. The results of cluster analysis provide a basis for the positioning and identification of key frames for the subsequent construction of the anti-counterfeiting information chain.

[0048] In order to provide comprehensive data support for subsequent key frame recognition and embedding of anti-counterfeiting information and improve key frame recognition, in an optional implementation, the above step S301 includes:

[0049] Step S3011, a second extraction step, extracting the first-order moment, the second-order moment and the third-order moment in the current video frame;

[0050] Step S3012, a composition step, using the first-order moment, the second-order moment and the third-order moment to compose the image color moment of the current video frame;

[0051] Step S3013, a smoothing step, using a Gaussian filter to smooth the corresponding current video frame according to the image color moment to obtain a corresponding smoothed video frame;

[0052] Step S3014, a calculation step, using first-order partial derivative finite difference to calculate the gradient amplitude and gradient direction of the smoothed video frame;

[0053] Step S3015, a suppression step, performing a non-maximum suppression operation on the smoothed video frame according to the gradient amplitude and the gradient direction, to obtain the amplitude of each pixel in the smoothed video frame;

[0054] Step S3016, a detection step, using a double threshold algorithm to detect the amplitude of each pixel in the smoothed video frame to generate the edge operator of the current video frame;

[0055] Step S3017, repeating the second extraction step, the composition step, the smoothing step, the calculation step, the suppression step and the detection step at least once in sequence until the image color moments and the edge operators corresponding to all the video frames are extracted.

[0056] In the above embodiment, a frame is defined as h, and each color component of each pixel in the frame is defined as h. i,j , since color information is mainly distributed in low-order moments, the first-order moment, second-order moment and third-order moment are sufficient to express the color distribution of the image. The first-order moment is The second moment is The third moment is The first three order color matrices are used to form a 9-dimensional histogram matrix, that is, the color moment vector is: F color =[μ1,σ1,S1,μ2,σ2,S2,μ3,σ3,S3], which is the color moment of the above image. Use Gaussian filter K to smooth the original image and achieve denoising, where Through convolution, we can get g = K * F color , let g be the smoothed image, that is, the smoothed video frame mentioned above, * represents convolution. For the smoothed g, the gradient amplitude G and gradient direction θ are calculated using the first-order partial derivative finite difference, where, θ=atan 2 (G y ,G x ), G x ,Gy Represent the partial derivatives of the image along the x and y directions respectively. The non-maximum suppression operation is performed on the edge of the image by obtaining the gradient amplitude and gradient direction. The calculation formula is as follows: M(g) = w*M(gα) + (1-w)*M(gβ), where w = distinct(g,gα) / distince(gα,gβ), w is the weight parameter, gα and gβ are different pixels in the video frame, M() represents the amplitude of the pixel, and distinct is the Euclidean distance between the two points. Use the double threshold algorithm to detect the edge operator F canny , the calculation formula is as follows: canny =Q(M(g)), the Q() method selects candidate boundaries based on two thresholds minVal and maxVal, detects and filters the pixel amplitude M(g), and obtains the edge operator of the video frame. By repeating the above steps, the image color moment and edge operator of each frame in the video can be extracted, providing comprehensive data support for the subsequent key frame recognition and anti-counterfeiting information embedding. It ensures that each frame of the entire video is fully analyzed and processed, thereby improving the anti-counterfeiting ability of the entire video.

[0057] In order to remove redundant information and noise in the key frame, in an optional implementation manner, the above step S303 includes:

[0058] Step S3031, solving the covariance matrix corresponding to the fusion features of each of the video frames according to the first formula, the first formula is: m represents the total number of the above video frames, x i represents the i-th video frame above, represents the mean vector of the above fusion features of all the above video frames, and C represents the above covariance matrix;

[0059] Step S3032, using an eigenvalue decomposition method to calculate the eigenvalues ​​and eigenvectors of each of the above covariance matrices, where one eigenvalue corresponds to one eigenvector;

[0060] Step S3033, selecting the eigenvectors corresponding to the first k eigenvalues ​​in descending order of the eigenvalues ​​to form a matrix, and obtaining the dimensionality reduction features corresponding to each of the video frames.

[0061] In the above embodiment, first, all key frames are traversed, the fusion feature vector of each key frame is calculated, and the mean vector of all fusion features is obtained. Then, according to the first formula, the fusion feature vector of each key frame is centered, and then the centered vector is multiplied by its transpose, that is, This operation is repeated for all video frames, and finally all results are added and divided by (m-1) to obtain the covariance matrix C.

[0062] Specifically, in order to improve computational efficiency, a parallel computing framework, such as OpenMP or CUDA, can be used to accelerate the calculation process of the covariance matrix. An online learning mechanism can be introduced to dynamically adjust the weights of the feature vectors and the calculation method of the covariance matrix to adapt to changes in different types of videos and environments. Using the incremental calculation method of PCA, the covariance matrix and principal components can be gradually updated when key frames flow in, avoiding memory and computing bottlenecks caused by processing a large amount of data at one time. The covariance matrix reflects the statistical correlation in the fused feature vector, provides a quantitative description of the relationship between the dimensions in the feature space, and improves the accuracy and efficiency of key frame recognition.

[0063] The calculated covariance matrix (C) is subjected to eigenvalue decomposition, which is a linear algebra operation that reveals the intrinsic structure of a matrix by solving the eigenvalues ​​and corresponding eigenvectors of the matrix. The result of eigenvalue decomposition is a set of eigenvalues ​​and corresponding unit eigenvectors, where the magnitude of the eigenvalue represents the importance of the eigenvector in the original data. More efficient eigenvalue decomposition algorithms, such as the QR algorithm or Arnoldi iteration, can be used to reduce computing time and resource consumption. At the same time, approximate calculation methods, such as randomized SVD or random projection, can be used to quickly estimate the eigenvalues ​​and eigenvectors of the covariance matrix, which is particularly helpful for processing large data sets. In addition, the distribution characteristics of the eigenvalues ​​can be used, such as using the explained variance ratio of PCA to determine the cutoff point of the eigenvalue decomposition and optimize the feature selection process.

[0064] According to the obtained eigenvalues ​​and eigenvectors, the first (k) eigenvectors with the largest eigenvalues ​​are selected, and these (k) eigenvectors constitute a dimensionality reduction matrix. An adaptive (k) value selection algorithm can be used to dynamically adjust the value of (k) according to the structure of the data and the evaluation of the dimensionality reduction effect. In addition, sparse coding technology can be introduced to sparsely represent the eigenvectors to further reduce the dimension and computational complexity of the dimensionality reduction features. Other dimensionality reduction techniques, such as t-SNE or Autoencoder, can also be combined to perform secondary dimensionality reduction on the data after PCA dimensionality reduction to adapt to more complex nonlinear data structures. By selecting the first (k) eigenvectors with the largest eigenvalues ​​for feature dimensionality reduction, redundant information and noise in the data can be removed, and the most critical and discriminative feature dimensions can be retained. The most representative and discriminative dimensionality reduction features can be extracted from the original fused features. This not only reduces the computational workload and storage requirements for subsequent processing, but also improves the accuracy and robustness of key frame recognition. The extraction of dimensionality reduction features makes data processing more efficient while maintaining the intrinsic structure and information of the data, providing a more optimized and refined feature representation for key frame detection and embedding of anti-counterfeiting information.

[0065] It should be noted that dimensionality reduction processing not only improves computational efficiency, but also enhances the accuracy and robustness of key frame recognition, which is of great significance for processing large-scale video data and realizing real-time video surveillance. Through reasonable solution expansion, such as parallel computing, online learning and adaptive feature selection, the practicality and flexibility of this method can be further improved, making it more adaptable to different scenarios and needs. This video anti-counterfeiting method based on frequency domain and statistical feature analysis is not only suitable for the field of bank security, but can also be widely used in other scenarios that require video data security and data integrity.

[0066] In order to enhance the anti-counterfeiting capability of the video, in an optional implementation manner, after the above step S201, the method further includes:

[0067] Step S401, when the current key frame is the first key frame in the key frame set, extracting the total number of key frames and the position information of the next key frame of the current key frame;

[0068] Step S402, stego-write the md5 code value of the current key frame, the position information, the total number of key frames and the position information of the next key frame of the current key frame to the DCT domain intermediate frequency of the header frame, where the header frame is the first video frame among all the video frames.

[0069] In the above embodiment, after determining the key frame set, the first key frame is identified as the starting point of the key frame sequence, and the total number of key frames in the key frame set is counted. A more efficient data structure, such as a hash table or a tree structure, can be used to store key frame information and positions for quick search and access. A structured framework for video anti-counterfeiting information is provided to ensure the integrity and coherence of subsequent key frame information. By recording the total number and position information of key frames, subsequent key frames can be quickly located, thereby improving the efficiency and accuracy of anti-counterfeiting detection. The head frame is subjected to discrete cosine transform (DCT) to convert it from the time domain to the frequency domain. An intermediate frequency region is determined in the DCT domain for embedding anti-counterfeiting information. The intermediate frequency region is selected because it has a small impact on visual quality and can provide sufficient embedding capacity. The md5 code value of the first key frame, the total number of key frames, and the position information of the next key frame are encoded into a compact number or string. The encoded information is steganographically written to the determined intermediate frequency region, and the information is embedded while minimizing the impact on image quality by adjusting the DCT coefficient. Of course, more complex encoding methods can also be used, such as differential encoding or adaptive encoding, to adjust the strength and position of information embedding according to the difference between frames to enhance the robustness of anti-counterfeiting information. Cryptography techniques, such as encryption or digital signatures, can also be used to protect anti-counterfeiting information such as md5 code values, increasing the difficulty of tampering with information. An adaptive embedding algorithm can also be implemented to dynamically adjust the density and depth of information embedding according to the clarity and content complexity of the video, so as to find a balance between anti-counterfeiting effect and video quality. By embedding the md5 code value, position information and total number of key frames in the DCT domain of the header frame, a hidden and difficult-to-detect information carrier is provided for video anti-counterfeiting, enhancing the anti-counterfeiting ability of the video. When the video is intercepted or tampered with, detecting the anti-counterfeiting information in the DCT domain of the header frame can quickly determine the integrity of the video and provide clues for locating the key frames.

[0070] In order to enhance the concealment and security of video anti-counterfeiting, in an optional implementation manner, the above step S202 includes:

[0071] Step S2021, performing discrete cosine transform on the corresponding frame to obtain DCT domain data of the corresponding frame after the transform;

[0072] Step S2022: Steganographically write the md5 code value and the position information of the current key frame to a designated position of the DCT domain intermediate frequency of the corresponding frame according to the DCT domain data.

[0073] In the above embodiment, a discrete cosine transform (DCT) is performed on the next key frame of the current key frame. Considering that different frequency bands have different effects on image perception, the DCT domain data can be further analyzed and optimized, and the most suitable frequency band can be selected for information hiding. Through DCT transformation, the image can be converted to the frequency domain, and the frequency domain characteristics can be used to hide and embed information, effectively reducing the impact of information embedding on the visual quality of the image. The acquisition of DCT domain data provides the possibility for subsequent steganographic anti-counterfeiting information in the intermediate frequency region, and also facilitates the compression and transmission of the video, because the DCT coefficient can be used for video encoding, and the steganographic information will not destroy the compression performance of the video. Determine a specific position group of the intermediate frequency in the DCT domain, and the selection of this position should take into account the minimum impact on the visual quality of the image and the concealment of the information. Encode the md5 code value and position information to be embedded into binary data, and then embed the binary data into the low bits (for example, the last few bits) of the selected DCT coefficients, and hide the information by modifying these low bits of the DCT coefficients. In order to maintain the natural appearance and compression performance of the image, the amount of embedded information should be moderate to avoid significant changes in the coefficient values. The embedded information can be encrypted to prevent unauthorized detection and tampering. Of course, the machine learning model can also be used to predict the best embedding position. By analyzing the distribution and correlation of the DCT coefficients, the intermediate frequency position that is most suitable for embedding information can be automatically selected to maximize the concealment and security of the information. By writing the md5 code value and position information at the intermediate frequency position in the DCT domain, the embedding of anti-counterfeiting information is achieved, and this embedding method is not easy to be detected and has high concealment and security. The selection of the intermediate frequency position ensures that the information embedding will not significantly damage the visual quality and compression performance of the image, so it can achieve the dual purposes of video anti-counterfeiting and video transmission. This embedding method also enhances the detection ability of key frame interception and tampering, because the md5 code value and position information are the unique identifiers of the key frame. If this information does not meet expectations during the detection process, the integrity problem of the video key frame can be immediately identified.

[0074] Specifically, the intermediate frequency position in the DCT domain is an ideal choice for embedding information because it has little visual impact on the image. By embedding information in the frequency domain, the obvious visual effects that may be caused by embedding information directly in the image space are avoided, further enhancing the concealment and anti-counterfeiting effects. Without sacrificing image quality, by embedding the md5 code value and position information in the intermediate frequency region, it can be ensured that even if the video is compressed or transmitted, the anti-counterfeiting information can still be fully restored and detected. This information embedding method based on frequency domain characteristics is not only suitable for video, but also for anti-counterfeiting processing of other multimedia data such as images and audio, and has broad application prospects and practical value. In addition, through the expansion of solutions such as encryption and machine learning prediction, the concealment and security of information can be further enhanced, making the anti-counterfeiting mechanism more difficult to crack, thereby better meeting the needs of high-security applications.

[0075] Specifically, DCT is a mathematical transformation that converts an image from the spatial domain to the frequency domain, which can concentrate the energy of the image in the low-frequency part and disperse the noise or detail information in the high-frequency part. Through DCT transformation, the original image is converted into a series of DCT coefficients, which constitute the DCT domain data.

[0076] In order to improve the efficiency and accuracy of video anti-counterfeiting detection, in an optional implementation, after the above step S203, the method further includes:

[0077] Step S501, an inverse transformation step, performing an inverse discrete cosine transform on the steganographic corresponding frame, extracting the md5 code value and position information steganographically written in the steganographic corresponding frame, the steganographic corresponding frame is any video frame in a steganographic key frame set, the steganographic key frame set is a set formed by all the video frames in which the md5 code value and position information of the key frame are steganographically written, the steganographic key frame set includes a head frame, all the key frames and the tail frame, and the head frame is the first video frame among all the video frames;

[0078] Step S502, a search step, searches for the corresponding key frame according to the position information, obtains a target key frame, and extracts the md5 code value of the target key frame;

[0079] Step S503, a comparison step, comparing the md5 code value of the steganographic corresponding frame with the md5 code value of the target key frame;

[0080] Step S504, a determination step, in which, when the md5 code value of the steganographic corresponding frame is inconsistent with the md5 code value of the target key frame, it is determined that the target key frame of the video has been tampered with;

[0081] Step S505, repeating the steps, when the md5 code value of the steganographic corresponding frame is consistent with the md5 code value of the target key frame, updating the target key frame to the new steganographic corresponding frame, and repeating the inverse transformation step, the search step, the comparison step, the determination step and the repetition step at least once in sequence, until all the key frames complete the video detection;

[0082] Step S506, obtaining the detection quantity, where the detection quantity is the number of the key frames involved in the video detection;

[0083] Step S507, when the detection quantity is less than the total quantity of key frames, it is determined that the key frames of the video have been intercepted;

[0084] Step S508: When the detection quantity is equal to the total quantity of key frames and there is no tampering problem, it is determined that the video is in a safe state.

[0085] In the above embodiment, if Figure 4As shown, first, locate the steganographic corresponding frame in the steganographic key frame set, that is, start from the tail frame, traverse to the first key frame and the head frame, perform the inverse DCT transform (IDCT) on each steganographic corresponding frame, and restore it to the spatial domain. Then, in the restored frame, extract the embedded md5 code value and position information from the pre-agreed specific position group of the DCT domain intermediate frequency, which is usually encoded in the low bit of the DCT coefficient. Finally, decode the extracted information to obtain the md5 code value and the position information of the previous key frame. The inverse transformation step can recover the steganographic md5 code value and position information from the steganographic corresponding frame, providing a basis for the subsequent key frame positioning, md5 code value detection and verification of the anti-counterfeiting mechanism. This process is the first step of video anti-counterfeiting detection, which ensures that the starting point of the anti-counterfeiting information chain and the anti-counterfeiting mark of each key frame are accurately obtained from the video sequence. Use the position information extracted from the steganographic corresponding frame to locate the corresponding previous key frame in the video sequence, and regard the key frame as the target key frame. Next, the md5 code value of the target key frame is calculated as its real md5 code value for subsequent comparison and verification. The search step locates the previous key frame through the position information and calculates its md5 code value, providing a reference for the real md5 code value for subsequent comparison and judgment. This process ensures the accurate identification of each node in the anti-counterfeiting information chain and helps to quickly locate the tampered or intercepted key frame. The md5 code value extracted from the corresponding steganographic frame is compared with the real md5 code value calculated by the target key frame to check the matching degree of the two. If the md5 code value does not match in the comparison step, the tampering of the key frame is reported immediately, which can quickly identify the tampering of the video key frame, provide an immediate alarm for video security monitoring, and help to take real-time measures to deal with possible security threats. Of course, in order to improve security, a multiple judgment mechanism can be designed, such as using multiple independent md5 code values ​​or using a more complex hash algorithm for cross-validation to ensure the accuracy of anti-counterfeiting detection. If the md5 code value matches, continue to use the previous key frame of the target key frame as the new steganographic corresponding frame, repeat the above inverse transformation step, the above search step, the above comparison step, the above determination step and the above repetition step until all key frames are detected, ensuring the integrity verification of all key frames. During the detection process, the number of key frames detected is recorded as the detection number. Compare the detection number and the total number of key frames. If the detection number is less than the total number of key frames, it means that a key frame has been intercepted during the transmission or storage process. By comparing the detection number and the total number of key frames, it is possible to identify the situation where the video key frames are intercepted, which helps to discover potential security threats during video transmission and storage. After completing the detection of all key frames, if the detection number is consistent with the total number of key frames and the md5 code values ​​of all key frames match, the video is determined to be safe.

[0086] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0087] The embodiment of the present application also provides an anti-counterfeiting and detection device for a key domain of a video. It should be noted that the anti-counterfeiting and detection device for a key domain of a video in the embodiment of the present application can be used to execute the anti-counterfeiting and detection method for a key domain of a video provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and those that have been described will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0088] The following is an introduction to the anti-counterfeiting and detection device for the key domain of a video provided in an embodiment of the present application.

[0089] Figure 5 is a structural block diagram of the video key domain anti-counterfeiting and detection device according to an embodiment of the present application. Figure 5 As shown, the device comprises:

[0090] The first extraction unit 10 is used to perform the first extraction step, extract the current key frame, and extract the md5 code value and position information of the current key frame. The current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames.

[0091] Specifically, a set of key frames is determined from a video frame sequence through a preset feature analysis and clustering algorithm. A key frame is selected as the current key frame, and the MD5 code value of the current key frame is calculated, which usually involves hashing the binary representation of the key frame pixel value to obtain a hash code of a fixed length as the unique identifier of the key frame. The position information of the current key frame in the video sequence is recorded, which can be a frame index or a timestamp. By extracting the MD5 code value and position information of the key frame, key anti-counterfeiting identification and positioning clues are provided for subsequent steganography and detection. The MD5 code value can ensure the integrity and invariance of the key frame, while the position information can be used to locate and restore the key frame in the video, which is crucial for detecting whether the key frame has been intercepted and tampered with. In large-scale video surveillance systems, the accurate and timely extraction of this information is the basis for achieving effective video anti-counterfeiting and forensic support.

[0092] The first steganographic unit 20 is used to perform a steganographic step to steganographically write the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame, and the corresponding frame is the next key frame of the current key frame in the key frame set.

[0093] Specifically, the next key frame is subjected to discrete cosine transform (DCT) and converted to the DCT domain. Then, a specific position group of the intermediate frequency in the DCT domain (usually an area with less visual impact on the image) is selected, and the md5 code value and position information of the current key frame are encoded into a format suitable for embedding, such as a binary string. Next, the encoded information is embedded into the low bits of the selected DCT coefficients, and the information is hidden by modifying these bits of the DCT coefficients. Finally, the DCT domain data is inversely transformed to obtain a video frame embedded with anti-counterfeiting information. By embedding anti-counterfeiting information in the intermediate frequency of the DCT domain, the md5 code value and position information of the key frame can be hidden in the video while ensuring the video quality, which serves as the key basis for subsequent detection and recovery. This method avoids the visual distortion caused by embedding information directly in the image pixels, reduces the risk of accidental destruction of the anti-counterfeiting information, and improves the robustness and security of the anti-counterfeiting mechanism. For large-scale video surveillance systems, this embedding method can achieve effective anti-counterfeiting of video key frames without significantly increasing storage and transmission overhead, thereby enhancing the overall security and reliability of the system.

[0094] The first repeating unit 20 is used to update the next key frame of the above-mentioned current key frame to the new above-mentioned current key frame, and repeat the above-mentioned first extraction step and the above-mentioned steganographic step at least once in sequence, until the above-mentioned current key frame is the last above-mentioned key frame, and the above-mentioned md5 code value and the above-mentioned position information of the above-mentioned current key frame are steganographically written to the tail frame, and the above-mentioned tail frame is the last above-mentioned video frame among all the above-mentioned video frames.

[0095] Specifically, the next key frame is then updated to the new current key frame, and the first extraction step and the steganographic step are repeated until all key frames are processed. When the last key frame is processed, the md5 code value and position information of the key frame are extracted and steganographically written to the tail frame of the video sequence as the termination mark of the entire video anti-counterfeiting information chain. The construction of the anti-counterfeiting information chain is realized, and the information chain not only connects all key frames, but also ensures the continuity and integrity of the information. This mechanism can effectively detect the interception and tampering of key frames, and provides strong support for the anti-counterfeiting of video surveillance systems. During the detection process, the key frames can be quickly located and verified through the serialized key frame md5 code value and position information, which is of great significance for improving the security of the video surveillance system and the legal effect of video evidence. In addition, the process and automation of information chain construction reduce the need for manual intervention, improve the efficiency and consistency of video anti-counterfeiting processing, and are suitable for real-time processing of continuous video streams and batch processing of large-scale video data.

[0096] In this embodiment, the first extraction unit is used to perform the first extraction step, extract the current key frame, and extract the md5 code value and position information of the current key frame, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; the first steganography unit is used to perform the steganography step, steganographically write the md5 code value and the position information of the current key frame to the DCT domain intermediate frequency of the corresponding frame, and the corresponding frame is the next key frame of the current key frame in the key frame set; the first repetition unit is used to update the next key frame of the current key frame to the new current key frame, and repeat the first extraction step and the steganography step at least once in sequence until the current key frame is the last key frame, and steganographically write the md5 code value and the position information of the current key frame to the tail frame, and the tail frame is the last video frame of all the video frames. This application adopts the md5 code value of the key frame to be steganographically written into the image DCT domain. Considering the key frame extraction efficiency, key frame redundancy filtering and interception anti-counterfeiting, the md5 code value of the previous key frame is sequentially steganographically written into the image DCT domain intermediate frequency of the next key frame (such as embedded into the rear X tail of the floating point of the specified position group), and finally the md5 code value of the last key frame is steganographically written into the tail frame, which also reduces the damage to the video itself caused by too much embedded information. This application solves the problem that the key frames are not comprehensively analyzed in the prior art, and the high redundancy of the key frames leads to low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection.

[0097] In order to reduce the number of key frames and improve the efficiency of key frame detection and anti-counterfeiting information embedding, in an optional implementation manner, the device further includes:

[0098] A second extraction unit, used for extracting image color moments and edge operators corresponding to all the above video frames before extracting the current key frame;

[0099] A calculation unit, used for forming a fusion feature corresponding to the video frame according to the multiplication of the image color moment of each video frame and the edge operator with a corresponding weight;

[0100] A dimension reduction unit, used for extracting principal components from the fusion features of each of the video frames by using a principal component analysis method, so as to reduce the dimension of the fusion features and obtain reduced dimension features corresponding to each of the video frames;

[0101] The clustering unit is used to perform clustering analysis using a K-means clustering algorithm according to all of the above-mentioned dimensionality reduction features to obtain the above-mentioned key frame set.

[0102] In the above embodiment, if Figure 3 As shown, for each video frame, the first-order moment, second-order moment and third-order moment are first calculated to obtain the image color moment. This usually involves traversing each pixel in the frame and calculating its mean, variance and skewness in the RGB color space to form a feature vector that describes the color distribution of the image. Then, an edge detection algorithm, such as Canny edge detection, is applied to each video frame to obtain an edge operator, which involves using a Gaussian filter to smooth the image to remove noise, then calculating the gradient amplitude and direction of the image, performing non-maximum suppression and double threshold detection, and finally generating a binary image to mark the edge area in the video frame. Extracting the image color moment and edge operator of the video frame can capture the color distribution and structural information of the image in the video, providing a basis for subsequent feature fusion and key frame detection. The weights of the image color moment and edge operator are calculated, which are usually determined based on prior knowledge or through experiments to reflect the relative importance of the two features in key frame detection. Then, the image color moment and edge operator feature vector of each video frame are multiplied by their respective weights to obtain a weighted feature vector, as shown in the formula: X = [F color *w1,F canny *w2] constitutes the fusion feature X of the video frame, where w1 and w2 are the fusion weights, F color Represents the image color moment, F cannyRepresents the edge operator. Step S301: Extract the image color moments and edge operators corresponding to all the above video frames. Machine learning methods, such as support vector machines (SVM) or neural networks, can be used to automatically learn the weights of image color moments and edge operators to dynamically adapt to the feature importance of different video contents. In addition, in addition to simple addition, other feature fusion methods, such as feature splicing, feature cascading or deep feature extraction, can be tried to further enhance the expressiveness and discrimination of fused features. By fusing the image color moment and the edge operator, the comprehensive features of the video frame can be obtained, which not only contain color information but also reflect the structural characteristics of the image. The composition of the fused features provides a more comprehensive and rich content description for subsequent dimensionality reduction and clustering analysis. The fused feature vectors of all video frames are combined into a matrix, and then PCA (principal component analysis) is performed on this matrix. This first involves calculating the mean vector and covariance matrix of the fused features, and then performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and eigenvectors. According to the size of the eigenvalues, the first few principal components are selected, and these principal components constitute a low-dimensional space. Finally, the fused feature vector of each video frame is projected onto this low-dimensional space to obtain the reduced-dimensional feature vector. Through PCA dimensionality reduction, the redundancy and noise in the fused features can be removed, and the most representative and discriminative feature dimensions can be extracted, providing an optimized data representation for the selection of key frames. The reduced-dimensional features not only reduce the computational workload and storage requirements of data processing, but also improve the accuracy and efficiency of subsequent clustering and key frame detection. The K-means clustering algorithm is applied to perform cluster analysis on the reduced-dimensional features of all video frames. This usually starts with randomly selecting K initial cluster centers, and then assigning the reduced-dimensional feature vector of each video frame to the nearest cluster center to form K clusters. Next, the center of each cluster is recalculated, and the cluster assignment process is iterated until the clustering results converge. When the similarity between the video frame and the constructed k-means cluster class β reaches the threshold M, it is considered to be a key frame keyp, keyp=pdis(β,p)>M, and all key frames form a key frame set. Through K-means clustering, the dimensionality reduction features of all video frames can be divided into several clusters, each cluster represents a type or state of image content in the video. Selecting key frames from each cluster can ensure that the key information points of the video are covered, while reducing the number of key frames and improving the efficiency of key frame detection and anti-counterfeiting information embedding. The results of cluster analysis provide a basis for the positioning and identification of key frames for the subsequent construction of the anti-counterfeiting information chain.

[0103] In order to provide comprehensive data support for subsequent key frame recognition and embedding of anti-counterfeiting information and improve key frame recognition, in an optional implementation, the second extraction unit includes:

[0104] An extraction module, used to perform a second extraction step to extract a first-order moment, a second-order moment and a third-order moment in the current video frame;

[0105] A composition module, used to execute the composition step, using the first-order moment, the second-order moment and the third-order moment to compose the image color moment of the current video frame;

[0106] A smoothing module is used to perform a smoothing step, and use a Gaussian filter to smooth the corresponding current video frame according to the image color moment to obtain a corresponding smoothed video frame;

[0107] A first calculation module is used to execute the calculation step, and calculate the gradient amplitude and gradient direction of the smoothed video frame by using first-order partial derivative finite difference;

[0108] A suppression module, used to execute the suppression step, perform a non-maximum suppression operation on the smoothed video frame according to the gradient amplitude and the gradient direction, and obtain the amplitude of each pixel in the smoothed video frame;

[0109] A detection unit, configured to perform the detection step, using a double threshold algorithm to detect the amplitude of each pixel in the smoothed video frame, and generate the edge operator of the current video frame;

[0110] A repetition module is used to sequentially repeat the second extraction step, the composition step, the smoothing step, the calculation step, the suppression step and the detection step at least once until the image color moments and the edge operators corresponding to all the video frames are extracted.

[0111] In the above embodiment, a frame is defined as h, and each color component of each pixel in the frame is defined as h. i,j , since color information is mainly distributed in low-order moments, the first-order moment, second-order moment and third-order moment are sufficient to express the color distribution of the image. The first-order moment is The second moment is The third moment is The first three order color matrices are used to form a 9-dimensional histogram matrix, that is, the color moment vector is: F color =[μ1,σ1,S1,μ2,σ2,S2,μ3,σ3,S3], which is the color moment of the above image. Use Gaussian filter K to smooth the original image and achieve denoising, where Through convolution, we can get g = K * F color , let g be the smoothed image, that is, the smoothed video frame mentioned above, * represents convolution. For the smoothed g, the gradient amplitude G and gradient direction θ are calculated using the first-order partial derivative finite difference, where, θ=atan 2 (G y ,G x ), G x ,Gy Represent the partial derivatives of the image along the x and y directions respectively. The non-maximum suppression operation is performed on the edge of the image by obtaining the gradient amplitude and gradient direction. The calculation formula is as follows: M(g) = w*M(gα) + (1-w)*M(gβ), where w = distinct(g,gα) / distince(gα,gβ), w is the weight parameter, gα and gβ are different pixels in the video frame, M() represents the amplitude of the pixel, and distinct is the Euclidean distance between the two points. Use the double threshold algorithm to detect the edge operator F canny , the calculation formula is as follows: canny =Q(M(g)), the Q() method selects candidate boundaries based on two thresholds minVal and maxVal, detects and filters the pixel amplitude M(g), and obtains the edge operator of the video frame. By repeating the above steps, the image color moment and edge operator of each frame in the video can be extracted, providing comprehensive data support for the subsequent key frame recognition and anti-counterfeiting information embedding. It ensures that each frame of the entire video is fully analyzed and processed, thereby improving the anti-counterfeiting ability of the entire video.

[0112] In order to remove redundant information and noise in the key frame, in an optional implementation manner, the dimensionality reduction unit includes:

[0113] A solution module is used to solve the covariance matrix corresponding to the above fusion features of each of the above video frames according to the first formula. The above first formula is: m represents the total number of the above video frames, x i represents the i-th video frame above, represents the mean vector of the above fusion features of all the above video frames, and C represents the above covariance matrix;

[0114] A second calculation module is used to calculate the eigenvalues ​​and eigenvectors of each of the above covariance matrices by using an eigenvalue decomposition method, wherein one eigenvalue corresponds to one eigenvector;

[0115] The selection module is used to select the eigenvectors corresponding to the first k eigenvalues ​​in the order of the eigenvalues ​​from large to small to form a matrix, so as to obtain the dimensionality reduction features corresponding to each of the video frames.

[0116] In the above embodiment, first, all key frames are traversed, the fusion feature vector of each key frame is calculated, and the mean vector of all fusion features is obtained. Then, according to the first formula, the fusion feature vector of each key frame is centered, and then the centered vector is multiplied by its transpose, that is, This operation is repeated for all video frames, and finally all results are added and divided by (m-1) to obtain the covariance matrix C.

[0117] Specifically, in order to improve computational efficiency, a parallel computing framework, such as OpenMP or CUDA, can be used to accelerate the calculation process of the covariance matrix. An online learning mechanism can be introduced to dynamically adjust the weights of the feature vectors and the calculation method of the covariance matrix to adapt to changes in different types of videos and environments. Using the incremental calculation method of PCA, the covariance matrix and principal components can be gradually updated when key frames flow in, avoiding memory and computing bottlenecks caused by processing a large amount of data at one time. The covariance matrix reflects the statistical correlation in the fused feature vector, provides a quantitative description of the relationship between the dimensions in the feature space, and improves the accuracy and efficiency of key frame recognition.

[0118] The calculated covariance matrix (C) is subjected to eigenvalue decomposition, which is a linear algebra operation that reveals the intrinsic structure of a matrix by solving the eigenvalues ​​and corresponding eigenvectors of the matrix. The result of eigenvalue decomposition is a set of eigenvalues ​​and corresponding unit eigenvectors, where the magnitude of the eigenvalue represents the importance of the eigenvector in the original data. More efficient eigenvalue decomposition algorithms, such as the QR algorithm or Arnoldi iteration, can be used to reduce computing time and resource consumption. At the same time, approximate calculation methods, such as randomized SVD or random projection, can be used to quickly estimate the eigenvalues ​​and eigenvectors of the covariance matrix, which is particularly helpful for processing large data sets. In addition, the distribution characteristics of the eigenvalues ​​can be used, such as using the explained variance ratio of PCA to determine the cutoff point of the eigenvalue decomposition and optimize the feature selection process.

[0119] According to the obtained eigenvalues ​​and eigenvectors, the first (k) eigenvectors with the largest eigenvalues ​​are selected, and these (k) eigenvectors constitute a dimensionality reduction matrix. An adaptive (k) value selection algorithm can be used to dynamically adjust the value of (k) according to the structure of the data and the evaluation of the dimensionality reduction effect. In addition, sparse coding technology can be introduced to sparsely represent the eigenvectors to further reduce the dimension and computational complexity of the dimensionality reduction features. Other dimensionality reduction techniques, such as t-SNE or Autoencoder, can also be combined to perform secondary dimensionality reduction on the data after PCA dimensionality reduction to adapt to more complex nonlinear data structures. By selecting the first (k) eigenvectors with the largest eigenvalues ​​for feature dimensionality reduction, redundant information and noise in the data can be removed, and the most critical and discriminative feature dimensions can be retained. The most representative and discriminative dimensionality reduction features can be extracted from the original fused features. This not only reduces the computational workload and storage requirements for subsequent processing, but also improves the accuracy and robustness of key frame recognition. The extraction of dimensionality reduction features makes data processing more efficient while maintaining the intrinsic structure and information of the data, providing a more optimized and refined feature representation for key frame detection and embedding of anti-counterfeiting information.

[0120] Dimensionality reduction processing not only improves computational efficiency, but also enhances the accuracy and robustness of key frame recognition, which is of great significance for processing large-scale video data and realizing real-time video surveillance. Through reasonable solution expansion, such as parallel computing, online learning and adaptive feature selection, the practicality and flexibility of this method can be further improved, making it more adaptable to different scenarios and needs. This video anti-counterfeiting method based on frequency domain and statistical feature analysis is not only suitable for the field of bank security, but can also be widely used in other scenarios that require video data security and data integrity.

[0121] In order to enhance the anti-counterfeiting capability of the video, in an optional implementation manner, the device further includes:

[0122] A third extraction unit is used to extract the current key frame, and extract the md5 code value and position information of the current key frame, and then, if the current key frame is the first key frame in the key frame set, extract the total number of key frames and the position information of the next key frame of the current key frame;

[0123] The second steganographic unit is used to steganographically write the md5 code value of the current key frame, the position information, the total number of key frames and the position information of the next key frame of the current key frame onto the DCT domain intermediate frequency of the header frame, where the header frame is the first video frame among all the video frames.

[0124] In the above embodiment, after determining the key frame set, the first key frame is identified as the starting point of the key frame sequence, and the total number of key frames in the key frame set is counted. A more efficient data structure, such as a hash table or a tree structure, can be used to store key frame information and positions for quick search and access. A structured framework for video anti-counterfeiting information is provided to ensure the integrity and coherence of subsequent key frame information. By recording the total number and position information of key frames, subsequent key frames can be quickly located, thereby improving the efficiency and accuracy of anti-counterfeiting detection. The head frame is subjected to discrete cosine transform (DCT) to convert it from the time domain to the frequency domain. An intermediate frequency region is determined in the DCT domain for embedding anti-counterfeiting information. The intermediate frequency region is selected because it has a small impact on visual quality and can provide sufficient embedding capacity. The md5 code value of the first key frame, the total number of key frames, and the position information of the next key frame are encoded into a compact number or string. The encoded information is steganographically written to the determined intermediate frequency region, and the information is embedded while minimizing the impact on image quality by adjusting the DCT coefficient. Of course, more complex encoding methods can also be used, such as differential encoding or adaptive encoding, to adjust the strength and position of information embedding according to the difference between frames to enhance the robustness of anti-counterfeiting information. Cryptography techniques, such as encryption or digital signatures, can also be used to protect anti-counterfeiting information such as md5 code values, increasing the difficulty of tampering with information. An adaptive embedding algorithm can also be implemented to dynamically adjust the density and depth of information embedding according to the clarity and content complexity of the video, so as to find a balance between anti-counterfeiting effect and video quality. By embedding the md5 code value, position information and total number of key frames in the DCT domain of the header frame, a hidden and difficult-to-detect information carrier is provided for video anti-counterfeiting, enhancing the anti-counterfeiting ability of the video. When the video is intercepted or tampered with, detecting the anti-counterfeiting information in the DCT domain of the header frame can quickly determine the integrity of the video and provide clues for locating the key frames.

[0125] In order to enhance the concealment and security of video anti-counterfeiting, in an optional implementation manner, the first steganographic unit includes:

[0126] A cosine transform module, used for performing discrete cosine transform on the corresponding frame to obtain DCT domain data of the corresponding frame after the transform;

[0127] The steganographic module is used to steganographically write the md5 code value and the position information of the current key frame to the designated position of the DCT domain intermediate frequency of the corresponding frame according to the DCT domain data.

[0128] In the above embodiment, a discrete cosine transform (DCT) is performed on the next key frame of the current key frame. Considering that different frequency bands have different effects on image perception, the DCT domain data can be further analyzed and optimized, and the most suitable frequency band can be selected for information hiding. Through DCT transformation, the image can be converted to the frequency domain, and the frequency domain characteristics can be used to hide and embed information, effectively reducing the impact of information embedding on the visual quality of the image. The acquisition of DCT domain data provides the possibility for subsequent steganographic anti-counterfeiting information in the intermediate frequency region, and also facilitates the compression and transmission of the video, because the DCT coefficient can be used for video encoding, and the steganographic information will not destroy the compression performance of the video. Determine a specific position group of the intermediate frequency in the DCT domain, and the selection of this position should take into account the minimum impact on the visual quality of the image and the concealment of the information. Encode the md5 code value and position information to be embedded into binary data, and then embed the binary data into the low bits (for example, the last few bits) of the selected DCT coefficients, and hide the information by modifying these low bits of the DCT coefficients. In order to maintain the natural appearance and compression performance of the image, the amount of embedded information should be moderate to avoid significant changes in the coefficient values. The embedded information can be encrypted to prevent unauthorized detection and tampering. Of course, the machine learning model can also be used to predict the best embedding position. By analyzing the distribution and correlation of the DCT coefficients, the intermediate frequency position that is most suitable for embedding information can be automatically selected to maximize the concealment and security of the information. By writing the md5 code value and position information at the intermediate frequency position in the DCT domain, the embedding of anti-counterfeiting information is achieved, and this embedding method is not easy to be detected and has high concealment and security. The selection of the intermediate frequency position ensures that the information embedding will not significantly damage the visual quality and compression performance of the image, so it can achieve the dual purposes of video anti-counterfeiting and video transmission. This embedding method also enhances the detection ability of key frame interception and tampering, because the md5 code value and position information are the unique identifiers of the key frame. If this information does not meet expectations during the detection process, the integrity problem of the video key frame can be immediately identified.

[0129] Specifically, the intermediate frequency position in the DCT domain is an ideal choice for embedding information because it has little visual impact on the image. By embedding information in the frequency domain, the obvious visual effects that may be caused by embedding information directly in the image space are avoided, further enhancing the concealment and anti-counterfeiting effects. Without sacrificing image quality, by embedding the md5 code value and position information in the intermediate frequency region, it can be ensured that even if the video is compressed or transmitted, the anti-counterfeiting information can still be fully restored and detected. This information embedding method based on frequency domain characteristics is not only suitable for video, but also for anti-counterfeiting processing of other multimedia data such as images and audio, and has broad application prospects and practical value. In addition, through the expansion of solutions such as encryption and machine learning prediction, the concealment and security of information can be further enhanced, making the anti-counterfeiting mechanism more difficult to crack, thereby better meeting the needs of high-security applications.

[0130] Specifically, DCT is a mathematical transformation that converts an image from the spatial domain to the frequency domain, which can concentrate the energy of the image in the low-frequency part and disperse the noise or detail information in the high-frequency part. Through DCT transformation, the original image is converted into a series of DCT coefficients, which constitute the DCT domain data.

[0131] In order to improve the efficiency and accuracy of video anti-counterfeiting detection, in an optional implementation manner, the device further includes:

[0132] an inverse transformation unit, configured to, after the md5 code value and the position information of the current key frame are steganographically written to the tail frame, perform an inverse transformation step, perform an inverse discrete cosine transform on the steganographic corresponding frame, and extract the md5 code value and the position information steganographically written in the steganographic corresponding frame, wherein the steganographic corresponding frame is any one video frame in a steganographic key frame set, wherein the steganographic key frame set is a set formed by all the video frames in which the md5 code value and the position information of the key frame are steganographically written, and wherein the steganographic key frame set includes a head frame, all the key frames and the tail frame, wherein the head frame is the first video frame among all the video frames;

[0133] A search unit, used to perform the search step, search for the corresponding key frame according to the position information, obtain the target key frame, and extract the md5 code value of the target key frame;

[0134] A comparison unit, used to perform a comparison step, comparing the md5 code value of the steganographic corresponding frame with the md5 code value of the target key frame;

[0135] The first determination unit is used to perform a determination step, and when the md5 code value of the steganographic corresponding frame is inconsistent with the md5 code value of the target key frame, determine that the target key frame of the video has been tampered with;

[0136] The second repeating unit is used to execute the repeating step, when the md5 code value of the steganographic corresponding frame is consistent with the md5 code value of the target key frame, update the target key frame to the new steganographic corresponding frame, and repeat the inverse transformation step, the search step, the comparison step, the determination step and the repeating step at least once in sequence until all the key frames complete the video detection;

[0137] An acquisition unit, used to acquire a detection quantity, where the detection quantity is the number of the key frames involved in the video detection;

[0138] A second determination unit is configured to determine that the key frame of the video has a problem of being intercepted when the detection quantity is less than the total number of key frames;

[0139] The third determination unit is used to determine that the video is in a safe state when the detection quantity is equal to the total quantity of the key frames and the tampering problem does not exist.

[0140] In the above embodiment, if Figure 4As shown, first, locate the steganographic corresponding frame in the steganographic key frame set, that is, start from the tail frame, traverse to the first key frame and the head frame, perform the inverse DCT transform (IDCT) on each steganographic corresponding frame, and restore it to the spatial domain. Then, in the restored frame, extract the embedded md5 code value and position information from the pre-agreed specific position group of the DCT domain intermediate frequency, which is usually encoded in the low bit of the DCT coefficient. Finally, decode the extracted information to obtain the md5 code value and the position information of the previous key frame. The inverse transformation step can recover the steganographic md5 code value and position information from the steganographic corresponding frame, providing a basis for the subsequent key frame positioning, md5 code value detection and verification of the anti-counterfeiting mechanism. This process is the first step of video anti-counterfeiting detection, which ensures that the starting point of the anti-counterfeiting information chain and the anti-counterfeiting mark of each key frame are accurately obtained from the video sequence. Use the position information extracted from the steganographic corresponding frame to locate the corresponding previous key frame in the video sequence, and regard the key frame as the target key frame. Next, the md5 code value of the target key frame is calculated as its real md5 code value for subsequent comparison and verification. The search step locates the previous key frame through the position information and calculates its md5 code value, providing a reference for the real md5 code value for subsequent comparison and judgment. This process ensures the accurate identification of each node in the anti-counterfeiting information chain and helps to quickly locate the tampered or intercepted key frame. The md5 code value extracted from the corresponding steganographic frame is compared with the real md5 code value calculated by the target key frame to check the matching degree of the two. If the md5 code value does not match in the comparison step, the tampering of the key frame is reported immediately, which can quickly identify the tampering of the video key frame, provide an immediate alarm for video security monitoring, and help to take real-time measures to deal with possible security threats. Of course, in order to improve security, a multiple judgment mechanism can be designed, such as using multiple independent md5 code values ​​or using a more complex hash algorithm for cross-validation to ensure the accuracy of anti-counterfeiting detection. If the md5 code value matches, continue to use the previous key frame of the target key frame as the new steganographic corresponding frame, repeat the above inverse transformation step, the above search step, the above comparison step, the above determination step and the above repetition step until all key frames are detected, ensuring the integrity verification of all key frames. During the detection process, the number of key frames detected is recorded as the detection number. Compare the detection number and the total number of key frames. If the detection number is less than the total number of key frames, it means that a key frame has been intercepted during the transmission or storage process. By comparing the detection number and the total number of key frames, it is possible to identify the situation where the video key frames are intercepted, which helps to discover potential security threats during video transmission and storage. After completing the detection of all key frames, if the detection number is consistent with the total number of key frames and the md5 code values ​​of all key frames match, the video is determined to be safe.

[0141] The anti-counterfeiting and detection device for the key domain of the video includes a processor and a memory. The first extraction unit, the first steganographic unit, the first repetition unit, etc. are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions. The modules are all located in the same processor; or, the modules are located in different processors in any combination.

[0142] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters can be adjusted to solve the problem that the key frames are not comprehensively analyzed in the prior art, and the key frame redundancy is high, resulting in low anti-counterfeiting embedding and anti-counterfeiting detection efficiency.

[0143] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0144] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the anti-counterfeiting and detection method for the key domain of the video.

[0145] Specifically, the anti-counterfeiting and detection methods of the key video domains include:

[0146] Step S201, a first extraction step, extracting a current key frame, and extracting an md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames;

[0147] Step S202, a steganographic step, steganographically writing the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame, the corresponding frame being the next key frame of the current key frame in the key frame set;

[0148] Step S203, updating the next key frame of the current key frame as the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, and steganographically writing the md5 code value and the position information of the current key frame to the tail frame, which is the last video frame among all the video frames.

[0149] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein the anti-counterfeiting and detection method of the key domain of the video is executed when the program is running.

[0150] Specifically, the anti-counterfeiting and detection methods of the key video domains include:

[0151] Step S201, a first extraction step, extracting a current key frame, and extracting an md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames;

[0152] Step S202, a steganographic step, steganographically writing the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame, the corresponding frame being the next key frame of the current key frame in the key frame set;

[0153] Step S203, updating the next key frame of the current key frame as the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, and steganographically writing the md5 code value and the position information of the current key frame to the tail frame, which is the last video frame among all the video frames.

[0154] An embodiment of the present invention provides a video anti-counterfeiting detection system, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:

[0155] Step S201, a first extraction step, extracting a current key frame, and extracting an md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames;

[0156] Step S202, a steganographic step, steganographically writing the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame, the corresponding frame being the next key frame of the current key frame in the key frame set;

[0157] Step S203, updating the next key frame of the current key frame as the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, and steganographically writing the md5 code value and the position information of the current key frame to the tail frame, which is the last video frame among all the video frames.

[0158] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing at least the following method steps:

[0159] Step S201, a first extraction step, extracting a current key frame, and extracting an md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames;

[0160] Step S202, a steganographic step, steganographically writing the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame, the corresponding frame being the next key frame of the current key frame in the key frame set;

[0161] Step S203, updating the next key frame of the current key frame as the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, and steganographically writing the md5 code value and the position information of the current key frame to the tail frame, which is the last video frame among all the video frames.

[0162] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0163] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0164] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0165] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0167] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0168] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0169] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0170] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0171] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0172] 1) The anti-counterfeiting and detection method of the video key domain of the present application, first, the first extraction step, extracting the current key frame, and extracting the md5 code value and position information of the current key frame, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; then, the steganographic step, steganographically writing the md5 code value and the position information of the current key frame to the DCT domain intermediate frequency of the corresponding frame, the corresponding frame is the next key frame of the current key frame in the key frame set; finally, updating the next key frame of the current key frame to the new current key frame, and repeating the first extraction step and the steganographic step at least once in sequence, until the current key frame is the last key frame, the md5 code value and the position information of the current key frame are steganographically written to the tail frame, and the tail frame is the last video frame of all the video frames. This application adopts the md5 code value of the key frame to be steganographically written into the image DCT domain. Considering the key frame extraction efficiency, key frame redundancy filtering and interception anti-counterfeiting, the md5 code value of the previous key frame is sequentially steganographically written into the image DCT domain intermediate frequency of the next key frame (such as embedded into the rear X tail of the floating point of the specified position group), and finally the md5 code value of the last key frame is steganographically written into the tail frame, which also reduces the damage to the video itself caused by too much embedded information. This application solves the problem that the key frames are not comprehensively analyzed in the prior art, and the high redundancy of the key frames leads to low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection.

[0173] 2) The anti-counterfeiting and detection device of the video key domain of the present application, the first extraction unit, is used to perform the first extraction step, extract the current key frame, and extract the md5 code value and position information of the current key frame, the current key frame is any key frame in the key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; the first steganography unit, is used to perform the steganography step, and steganography the md5 code value and the position information of the current key frame to the DCT domain intermediate frequency of the corresponding frame, and the corresponding frame is the next key frame of the current key frame in the key frame set; the first repetition unit, is used to update the next key frame of the current key frame to the new current key frame, and repeat the first extraction step and the steganography step in sequence at least once, until the current key frame is the last key frame, and the md5 code value and the position information of the current key frame are steganographically written to the tail frame, and the tail frame is the last video frame among all the video frames. This application adopts the md5 code value of the key frame to be steganographically written into the image DCT domain. Considering the key frame extraction efficiency, key frame redundancy filtering and interception anti-counterfeiting, the md5 code value of the previous key frame is sequentially steganographically written into the image DCT domain intermediate frequency of the next key frame (such as embedded into the rear X tail of the floating point of the specified position group), and finally the md5 code value of the last key frame is steganographically written into the tail frame, which also reduces the damage to the video itself caused by too much embedded information. This application solves the problem that the key frames are not comprehensively analyzed in the prior art, and the high redundancy of the key frames leads to low efficiency of anti-counterfeiting embedding and anti-counterfeiting detection.

[0174] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for anti-counterfeiting and detection of key video domains, characterized in that: include: The first extraction step is to extract the current key frame, and extract the md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; A steganographic step of steganographically writing the md5 code value and the position information of the current key frame to the DCT domain intermediate frequency of a corresponding frame, wherein the corresponding frame is the next key frame of the current key frame in the key frame set; Update the next key frame of the current key frame to the new current key frame, and repeat the first extraction step and the steganographic step in sequence at least once, until the current key frame is the last key frame, and steganographically write the md5 code value and the position information of the current key frame to the tail frame, and the tail frame is the last video frame among all the video frames.

2. The method according to claim 1, characterized in that: Before extracting the current key frame, the method further includes: Extracting image color moments and edge operators corresponding to all the video frames; Composing a fusion feature corresponding to the video frame by multiplying the image color moment and the edge operator of each video frame with a corresponding weight; Extracting principal components from the fusion features of each of the video frames using a principal component analysis method to reduce the dimension of the fusion features and obtain reduced-dimensional features corresponding to each of the video frames; A K-means clustering algorithm is used to perform clustering analysis based on all of the dimensionality reduction features to obtain the key frame set.

3. The method according to claim 2, characterized in that Extracting image color moments and edge operators corresponding to all the video frames, including: The second extraction step extracts the first-order moment, the second-order moment and the third-order moment in the current video frame; A composition step, using the first-order moment, the second-order moment and the third-order moment to compose the image color moment of the current video frame; A smoothing step, using a Gaussian filter to smooth the corresponding current video frame according to the image color moment to obtain a corresponding smoothed video frame; A calculation step, using first-order partial derivative finite difference to calculate the gradient amplitude and gradient direction of the smoothed video frame; a suppression step of performing a non-maximum suppression operation on the smoothed video frame according to the gradient amplitude and the gradient direction to obtain the amplitude of each pixel in the smoothed video frame; A detection step, using a double threshold algorithm to detect the amplitude of each pixel in the smoothed video frame to generate the edge operator of the current video frame; The second extraction step, the composition step, the smoothing step, the calculation step, the suppression step and the detection step are sequentially repeated at least once until the image color moments and the edge operators corresponding to all the video frames are extracted.

4. The method according to claim 2, characterized in that: The principal component analysis method is used to extract the principal component of the fusion feature of each video frame to reduce the dimension of the fusion feature to obtain the reduced dimension feature corresponding to each video frame, including: The covariance matrix corresponding to the fusion features of each video frame is solved according to the first formula, and the first formula is: m represents the total number of video frames, x i represents the i-th video frame, represents the mean vector of the fusion features of all the video frames, and C represents the covariance matrix; Using an eigenvalue decomposition method to calculate the eigenvalues ​​and eigenvectors of each of the covariance matrices, one eigenvalue corresponds to one eigenvector; The eigenvectors corresponding to the first k eigenvalues ​​are selected in descending order of the eigenvalues ​​to form a matrix, so as to obtain the dimensionality reduction features corresponding to each of the video frames.

5. The method according to claim 1, characterized in that After extracting the current key frame, and extracting the md5 code value and position information of the current key frame, the method further includes: When the current key frame is the first key frame in the key frame set, extracting the total number of key frames and the position information of the next key frame of the current key frame; The md5 code value of the current key frame, the position information, the total number of key frames and the position information of the next key frame of the current key frame are steganographically written to the DCT domain intermediate frequency of the header frame, and the header frame is the first video frame among all the video frames.

6. The method according to claim 1, characterized in that Steganographically writing the md5 code value of the current key frame and the position information to the DCT domain intermediate frequency of the corresponding frame includes: Performing discrete cosine transform on the corresponding frame to obtain DCT domain data of the corresponding frame after transformation; The md5 code value and the position information of the current key frame are steganographically written to a designated position of the DCT domain intermediate frequency of the corresponding frame according to the DCT domain data.

7. The method according to claim 1, characterized in that After the md5 code value and the position information of the current key frame are stego-written to the tail frame, the method further includes: an inverse transformation step, performing an inverse discrete cosine transform on the steganographic corresponding frame, extracting the md5 code value and position information steganographically stored in the steganographic corresponding frame, wherein the steganographic corresponding frame is any video frame in a steganographic key frame set, wherein the steganographic key frame set is a set formed by all the video frames in which the md5 code value and position information of the key frame are steganographically stored, wherein the steganographic key frame set includes a head frame, all the key frames and the tail frame, and the head frame is the first video frame among all the video frames; A search step, searching for the corresponding key frame according to the position information, obtaining a target key frame, and extracting the md5 code value of the target key frame; A comparison step, comparing the md5 code value of the steganographic corresponding frame with the md5 code value of the target key frame; A determination step, in which, when the md5 code value of the steganographic corresponding frame is inconsistent with the md5 code value of the target key frame, it is determined that the target key frame of the video has been tampered with; Repeating step, when the md5 code value of the steganographic corresponding frame is consistent with the md5 code value of the target key frame, updating the target key frame to the new steganographic corresponding frame, and repeating the inverse transformation step, the search step, the comparison step, the determination step and the repeating step at least once in sequence until all the key frames complete the video detection; Obtaining a detection quantity, where the detection quantity is the number of the key frames involved in the video detection; When the number of detected frames is less than the total number of key frames, determining that the key frames of the video have a problem of being intercepted; When the detection quantity is equal to the total quantity of key frames and there is no tampering problem, it is determined that the video is in a safe state.

8. An anti-counterfeiting and detection device for a key video domain, characterized in that: The device comprises: A first extraction unit is used to perform a first extraction step, extract a current key frame, and extract an md5 code value and position information of the current key frame, wherein the current key frame is any key frame in a key frame set, and the key frame set is a set formed by the key frames extracted from all video frames; A first steganographic unit is used to perform a steganographic step to steganographically write the md5 code value and the position information of the current key frame to the DCT domain intermediate frequency of a corresponding frame, wherein the corresponding frame is the next key frame of the current key frame in the key frame set; The first repeating unit is used to update the next key frame of the current key frame to the new current key frame, and repeat the first extraction step and the steganographic step at least once in sequence until the current key frame is the last key frame, and the md5 code value and the position information of the current key frame are steganographically written to the tail frame, and the tail frame is the last video frame among all the video frames.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Deep learning-based image abnormal tampering automatic detection method, storage medium and equipment

    CN120544014A

  • Deep learning-based image abnormal tampering automatic detection method, storage medium and equipment

    CN120544014B

  • Media stream tampering monitoring method and device and medium

    CN121056661A