A stable video source identification method

By combining unsampled dual-tree complex wavelet transform and Bayesian adaptive search algorithm with affine transformation model, the problem of sensor pattern noise extraction and restoration in stable video is solved, and the accuracy and speed of video source recognition are improved.

CN117408965BActive Publication Date: 2025-09-26GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311366164.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-20
Publication Date
2025-09-26
Estimated Expiration
2043-10-20

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively extracting and recovering camera sensor pattern noise from stabilized videos, resulting in a decrease in source recognition performance, especially when the distortion between frames has a significant impact during the video stabilization process.

Method used

Unsampled dual-tree complex wavelet transform and Bayesian adaptive search algorithm are used in combination with affine transformation model to extract and restore sensor pattern noise and use peak correlation energy for matching and recognition.

Benefits of technology

It improves the accuracy and matching speed of stable video source identification, reduces the interference caused by video stabilization, and enhances the ability to extract sensor pattern noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117408965B_ABST
    Figure CN117408965B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying the source of a stable video, comprising the following steps: extracting sensor pattern noise from multiple reference stable videos; extracting sensor pattern noise from a stable video to be detected; restoring the sensor pattern noise of the stable video to be detected using an affine transformation model combined with a Bayesian adaptive direct search algorithm; and calculating the correlation between the sensor pattern noise of the reference stable video and the restored sensor pattern noise of the test stable video using peak correlation energy. If the peak correlation energy value is greater than or equal to a preset threshold, it is determined that the stable video to be detected is from the camera that recorded the reference video; conversely, the stable video to be detected is not from the camera that recorded the reference stable video. The present invention can extract more sufficient sensor pattern noise from the stable video, effectively improve the interference caused by video stabilization, and enhance the accuracy and matching speed of stable video source identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of stable video evidence collection, and in particular to a stable video source identification method based on unsampled dual-tree complex wavelet transform and Bayesian adaptive search. Background Art

[0002] With the increasing prevalence of electronic imaging devices such as smartphones and tablets, digital forensics—using photos or videos to identify the source of the device capturing the image—has become increasingly important, as the resulting identification can serve as valuable evidence in court proceedings. Current methods for identifying source devices primarily rely on photo response non-uniformity (PRNU) noise left by camera image sensors in digital multimedia. PRNU noise is unique to each imaging device, acting like a "fingerprint," and has been successfully used to identify and authenticate the origin of digital images.

[0003] PRNU is caused by the uneven response of the photosensitive material in the imaging sensor to light, a response that varies from camera to camera. However, various types of non-unique artifacts may be introduced during camera imaging. For example, there is CFA color interpolation noise and quantization noise from A / D conversion. In addition, when various correction processes are performed, noise from the image correction algorithm is introduced. Furthermore, when the resulting digital multimedia is set to a specific file format (e.g., JPEG, AVI, MOV, MP4, etc.) in the camera, noise from compression encoding increases. These processing steps may interfere with the inherent PRNU pattern. In particular, when the video is stabilized, the PRNU pattern may be distorted from frame to frame, significantly degrading source identification performance. Therefore, how to effectively extract and recover PRNU from stabilized video is crucial for source identification in forensic research.

[0004] As researchers at home and abroad have conducted more in-depth research on sensor pattern noise extraction, they have found that sensor pattern noise is a relatively weak signal with lower energy than the original content of the image. To solve this problem, researchers have proposed a filtering algorithm that can effectively extract sensor pattern noise from images. For example: Lukas J, Fridrich J, Goljan M. Digital camera identification from sensor pattern noise [J]. IEEE Transactions on Information Forensics and Security, 2006, 1 (2): 205-214. (Lukas J, Fridrich J, Goljan M. Digital camera identification from sensor pattern noise [J]. IEEE Transactions on Information Forensics and Security, 2006, 1 (2): 205-214.), the authors creatively proposed using a wavelet decomposition filtering method to extract pattern noise from images and use it as the camera's "fingerprint" to identify the source camera. For example: Chen M, Fridrich J, Goljan M. Digital imaging sensor identification (further study) [C] / / Security, steganography, and watermarking of multimedia contents IX. SPIE, 2007, 6505: 258-270. (Chen M, Fridrich J, Goljan M. Digital image sensor identification (further study) [C] / / Security, steganography and watermarking of multimedia contents IX. SPIE, 2007, 6505: 258-270.), Fridrich et al. believe that the source camera identification task is a joint estimation and detection problem. They extract the sensor output by simplifying the model to obtain the sensor pattern noise, and then derive the maximum likelihood estimator of the PRNU. Another example: Li CT. Source camera identification using enhanced sensor pattern noise[J]. IEEE Transactions on Information Forensics and Security, 2010, 5(2): 280-287. (Li CT. Source camera identification based on enhanced sensor pattern noise[J]. IEEE Transactions on Information Forensics and Security, 2010, 5(2): 280-287.), it is proposed that scene details will seriously contaminate the PRNU extracted from the image. When the signal in the PRNU is strong, the reliability of this component is low and therefore it should be attenuated. The authors proposed a method to enhance the PRNU by assigning a weight factor inversely proportional to the size of the PRNU component, thereby improving the recognition rate of the device. For example: Chen M, Fridrich J, Goljan M, et al. Determining image origin and integrity using sensor noise [J]. IEEE Transactions on information forensics and security, 2008, 3 (1): 74-90. (Chen M, Fridrich J, Goljan M, et al. Determining image origin and integrity using sensor noise [J]. IEEE Transactions on information forensics and security, 2008, 3 (1): 74-90), proposed a unified framework for forensic methods applicable to compressed images, which was tested on familiar image datasets and proved to be robust. For example: Quan Y, Li C T. On addressing the impact of ISO speed upon PRNU and forgery detection[J]. IEEE Transactions on Information Forensics and Security, 2020, 16: 190-202. (Quan Y, Li C T. On addressing the impact of ISO speed upon PRNU and forgery detection[J]. IEEE Transactions on Information Forensics and Security, 2020, 16: 190-202), Quan et al. proposed a method called content-based ISO speed inference (CINFISOS), which confirmed through analysis and experiments that the correlation between the sensor pattern noise of an image and its reference PRNU does not depend not only on the content but also on the camera sensitivity setting. For example: Mufti SN, Khan S. Blind Image Clustering Based on PRNU With Reduced Computational Complexity[C] / / 2022International Conference on IT and Industrial Technologies(ICIT).IEEE,2022:1-6.(Mufti SN, Khan S. Blind Image Clustering Based on PRNU With Reduced Computational Complexity[C] / / 2022International Conference on IT and Industrial Technologies(ICIT).IEEE,2022:1-6.), a two-stage clustering method based on camera fingerprint was adopted in the study of grouping images, and the proposed model has higher efficiency and smaller estimation error.

[0005] In fact, it is feasible to extend image-based PRNU methods for source camera identification to video. However, in the article: Van Houten W, Geradts Z. Source video camera identification for multiplycompressed videos originating from YouTube[J]. Digital Investigation, 2009, 6(1-2):48-60. (Van Houten W, Geradts Z. Source video camera identification for multiplycompressed videos originating from YouTube[J]. Digital Investigation, 2009, 6(1-2):48-60.), the authors mentioned that video-based PRNU estimation is more challenging than image-based PRNU due to the low resolution of video frames caused by video compression. In existing video coding standards (such as H.264, MPEG series, or later versions), a group of pictures (GOP) consists of intra-coded pictures (I-frames), predictive coded pictures (P-frames), and bidirectional interpolated predicted frames (B-frames). Several studies have proposed techniques to mitigate the impact of video compression on PRNU analysis. For example, in the literature: Altinisik E, Tasdemir K, Sencar H T. Extracting PRNU noise from H.264 coded videos[C] / / 2018 26th European signal processing conference(EUSIPCO). IEEE, 2018: 1367-1371. (Altinisik E, Tasdemir K, Sencar H T. Extracting PRNU noise from H.264 coded videos[C] / / 2018 26th European signal processing conference(EUSIPCO). IEEE, 2018: 1367-1371.), strategies such as removing post-decoding artifacts or selectively utilizing macroblocks are proposed to improve the accuracy of PRNU estimation by alleviating the adverse effects of the H.264 and H.265 video compression standards.Literature: Su K, Tian N, Pan Q. Multimedia source identification using an improved weight photo response non-uniformity noise extraction model in short compressed videos[J]. Forensic Science International: Digital Investigation, 2022, 42: 301473. (Su K, Tian N, Pan Q. Multimedia source identification using an improved weight photo response non-uniformity noise extraction model in short compressed videos[J]. International Forensic Science: Digital Investigation, 2022, 42: 301473), Su et al. solved the further compression problem caused by videos uploaded to social networking sites by proposing a weighted PRNU extraction model based on variance-stabilized transform multi-scale iterative least squares filtering. These technologies are beneficial to short video identification.

[0006] In addition to video compression, electronic image stabilization (EIS) is another key post-processing step that affects source camera identification. The purpose of EIS is to compensate for motion blur caused by camera shake. Specifically, video stabilization technology may introduce additional noise in the captured video frames and complicate the PRNU analysis process. Therefore, identifying the source camera in this case becomes more complicated. As mentioned in the literature: T,Brolund P,NorellK.Identifying camcorders using noise patterns from video clips recorded withimage stabilisation[C] / / 2011 7th International Symposium on Image and SignalProcessing and Analysis(ISPA).IEEE,2011:668-671.( T, Brolund P, Norell K. Noise pattern recognition in video clips recorded using image stabilization cameras[C] / / 2011 7th International Symposium on Image and Signal Processing and Analysis (ISPA). IEEE, 2011: 668-671. ), The authors address the problem of identifying stabilized video sources after image stabilization processing by considering the effect of post-processing electronic image stabilization compensation. For example: Taspinar S, Mohanty M, Memon N. Source camera attribution using stabilized video[C] / / 2016IEEE InternationalWorkshop on Information Forensics and Security(WIFS).IEEE, 2016:1-6. (Taspinar S, Mohanty M, Memon N. Source camera attribution using stabilized video[C] / / 2016IEEE InternationalWorkshop on Information Forensics and Security(WIFS).IEEE, 2016:1-6.) proposed a method to detect whether a video has been stabilized by estimating the peak-to-correlation energy (PCE) value between the same set of video frames. The method also includes compensating the stabilized video through rotation and translation transformation in the later process, but the brute force search method leads to high computational cost. Reference: Mandelli S, Bestagini P, Verdoliva L, et al. Facing device attribution problem for stabilized video sequences [J]. IEEE Transactions on Information Forensics and Security, 2019, 15: 14-27. (Mandelli S, Bestagini P, Verdoliva L, et al. Facing device attribution problem for stabilized video sequences [J]. IEEE Transactions on Information Forensics and Security, 2019, 15: 14-27.) explores a "hybrid" source camera identification method that identifies the source camera through image-to-video and video-to-video methods. In addition to translation and rotation transformations, the authors also consider scaling geometric transformations. A particle swarm search algorithm is used instead of a brute force search.Literature: Mandelli S, Argenti F, Bestagini P, et al. A modified Fourier-Mellin approach for source device identification on stabilized videos[C] / / 2020 IEEE International Conference on Image Processing(ICIP). IEEE, 2020: 1266-1270. (Mandelli S, Argenti F, Bestagini P, et al. A modified Fourier-Mellin approach for source device identification on stabilized videos[C] / / 2020 IEEE International Conference on Image Processing(ICIP). IEEE, 2020: 1266-1270.), Sara et al. proposed a modified Fourier Mellin (MFM) transform algorithm, which aims to reduce the computation time by searching for rotation and scaling parameters in the frequency domain of PRNU registration. In the paper: Bellavia F, Fanfani M, Colombo C, et al. Experience with electronic image stabilization and PRNU through scenecontent image registration[J]. Pattern Recognition Letters, 2021, 145: 8-15. (Bellavia F, Fanfani M, Colombo C, et al. Experience with electronic image stabilization and PRNU through scenecontent image registration[J]. Pattern Recognition Letters, 2021, 145: 8-15.), the authors discussed content-based image registration as an alternative to traditional PRNU-based methods, which match keypoint descriptors of image scene content to find the conversion between two different formats. However, it requires reacquiring images of static scenes, which may be more practical in some scenarios.

[0007] Removing camera motion (e.g., shake or vibration) is often necessary to obtain high-quality video, which can be achieved through geometric transformations. Aligning the PRNU estimates for each frame is crucial for recovering the camera's unique PRNU, particularly for PRNU-based video source recognition. Therefore, obtaining accurate and reliable PRNU estimates is crucial for improving the matching recognition rate in the subsequent recognition process. For these reasons, in order to extract more sensor pattern noise from stabilized video and improve the recognition rate and matching speed of stabilized video, it is necessary to develop a stabilized video source recognition algorithm for videos that have been stabilized. Summary of the Invention

[0008] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a stable video source identification method, which can extract more sufficient sensor pattern noise from stable video, effectively improve the interference caused by video anti-shake, and improve the accuracy and matching speed of stable video source identification.

[0009] To achieve the above objectives, the technical solutions provided by the present invention are:

[0010] A stable video source identification method comprises the following steps:

[0011] S1, extract sensor pattern noise from multiple reference stabilized videos;

[0012] S2, extracting sensor pattern noise from the stable video to be detected;

[0013] S3, using the affine transformation model combined with the Bayesian adaptive direct search algorithm to restore the sensor pattern noise of the stable video to be detected;

[0014] S4. Calculate the correlation between the sensor pattern noise of the reference stabilized video and the sensor pattern noise of the test stabilized video after recovery in step S3 using peak correlation energy. If the PCE value is greater than or equal to a preset threshold, it is determined that the stabilized video to be tested comes from the camera that recorded the reference video. Otherwise, the stabilized video to be tested does not come from the camera that recorded the reference stabilized video.

[0015] Furthermore, in steps S1 and S2, extracting the corresponding sensor pattern noise includes:

[0016] A1. Convert the video back into bitstream data, intervene in the decoding process, and output all video frames before the codec loop filter;

[0017] A2. Only the key frame intra-coded pictures are selected, and each video frame is subjected to undecimated dual-tree complex wavelet transform and decomposed into J-layer wavelet coefficients to obtain detailed information in all directions.

[0018] A3. Use a method based on minimum mean square error estimation to perform spatial adaptive denoising on the wavelet coefficients to obtain denoised wavelet coefficients;

[0019] A4. Reconstruct the denoised wavelet coefficients using an inverse undecimated dual-tree complex wavelet transform, thereby achieving perfect reconstruction and obtaining a denoised video frame; subtract the input video frame from the denoised video frame to obtain the sensor pattern noise of each video frame.

[0020] Furthermore, in step A2, the specific process of the unsampled dual-tree complex wavelet decomposition transform is:

[0021] The undecimated dual-tree complex wavelet transform decomposes and reconstructs the signal by downsampling only the first layer but not the remaining layers. The first tree generates the real part and the second generates the imaginary part. The sampling frequency of the filters in the two trees is the same, and the delay between them is one sampling interval. The binary decimation of the first layer in the imaginary part tree just samples the sample values ​​lost by the binary decimation in the real part tree.

[0022] Furthermore, in step A3, the wavelet coefficients are spatially adaptively denoised using a method based on minimum mean square error estimation. The formula for obtaining the denoised wavelet coefficients is as follows:

[0023]

[0024] In formula (1), W in Expressed as the wavelet coefficient before filtering, W out Expressed as the wavelet coefficient after filtering, the noise variance is estimated as The subband variance is

[0025] Furthermore, step S3 includes:

[0026] Apply a geometric transformation to each acquired test frame and select an affine transformation model for processing, which represents the inverse coordinate mapping between the stabilized test video frame and the original frame pixels:

[0027]

[0028] Where (x′, y′) is the coordinate of the stabilized pixel, and (x, y) is the coordinate of the original pixel after restoration. T represents the affine transformation matrix, the parameter s in the matrix represents the scaling ratio, θ represents the rotation angle, and c x Indicates horizontal shift, c y Indicates vertical displacement;

[0029] The Bayesian adaptive direct search algorithm searches at multiple scales, divides the search space using the idea of ​​grid partitioning, and combines global search with local search to find the optimal solution.

[0030] Furthermore, Bayesian adaptive direct search includes two steps: Search stage and Poll stage;

[0031] Search stage:

[0032] An active search strategy based on Gaussian processes and Bayesian optimization is used to minimize the number of objective function evaluations, effectively explore uncertain regions and develop high-predictive-value regions, quickly finding the global optimal solution:

[0033]

[0034] Before the i-th iteration, the optimal solution of the objective function PCE is expressed as x i , that is, {s,θ,c}, the calculated feasible solution set is defined as X i ; The set of all search points is defined as M i ,in is the size parameter of the search grid, D is the grid direction set, and z is a full-rank positive integer matrix;

[0035] The Search stage is divided into four steps:

[0036] Step 1: x i Construct grid cells for search centers;

[0037] Step 2: Calculate the objective values ​​of the finite grid points near the constructed grid elements and find a feasible solution to the improved objective function;

[0038] Step 3: If a feasible solution that improves the objective function is found, the search is successful; at this time, the grid center is moved to that position and the grid size parameter is increased in the i+1th iteration step.

[0039] Step 4: Failure to find a feasible solution to improve the objective function indicates that the search process of the iterative step has failed; then turn to the next optimization step, the poll stage, and then reduce the grid size parameter in the i+1th iteration step

[0040] Poll stage:

[0041] When the search phase fails to find a feasible solution, the polling phase begins. Polling searches for a feasible solution by evaluating different directions at the grid points, with the step size adjusted based on success or failure. The core idea is to narrow the area of ​​improvement.

[0042]

[0043] Among them, a set Q consisting of directions with larger density is defined i , matrix D i It consists of column elements concentrated in the grid direction;

[0044] The parameters control the size of the filtering box. The polling phase is performed in the area around the current best solution. If a feasible solution is successfully found, the next polling phase will be continued. If a feasible solution fails, the parameters will be reduced to avoid falling into a finite set. The update method is used to avoid falling into a finite set during the polling phase and increase the probability of finding the optimal direction.

[0045] Furthermore, in step S4, the peak correlation energy (PCE) is expressed as formula (13):

[0046]

[0047] where ρ is the normalized cross-correlation between the reference pattern noise and the pattern noise of the test image, s is the mapping of all elements in ρ, Ω represents a small neighborhood centered at ρ(0,0), and |Ω| represents the number of elements in Ω.

[0048] Furthermore, in order to align the extracted pattern noise with the reference pattern noise, an affine transformation W is performed on the pattern noise of each test frame so that the PCE value between the transformed pattern noise and the camera reference pattern noise is maximized:

[0049]

[0050] Among them, W test , I test represents the pattern noise of the test video frame and the original video frame, K represents the estimated reference video fingerprint, and p represents the maximum PCE value between the test pattern noise and the reference video pattern noise when the search parameters are {s, θ, c}; the maximization problem in Equation (14) is solved using the proposed Bayesian adaptive direct search algorithm;

[0051] Finally, for the N video frames tested, take the highest PCE value P in the frame set max As the final result of this test video and the reference video:

[0052]

[0053] Compared with the existing technology, the principles and advantages of this solution are as follows:

[0054] 1. To improve the interference caused by the loop filter in the codec, during the encoding and decoding process, the video frame is extracted before entering the loop filter module to retain more information about the sensor pattern noise.

[0055] 2. The introduced UDTCWT provides accurate translation invariance, improved directional selectivity, and complex sub-bands, while avoiding the over-complete representation problem of the Undecimated Discrete Wavelet Transform (UDWT). This enables the UDTCWT to better capture detailed information in the image, making it more effective for extracting sensor pattern noise in stabilized videos.

[0056] 3. Perform geometric transformation on the extracted pattern noise to match the reference pattern noise, eliminate the impact of the global frame jitter stabilization operation, and improve the recognition rate of the source video.

[0057] 4. A Bayesian adaptive direct search algorithm is used to search for parameters in the sensor pattern noise recovery process. This algorithm automatically adapts to the characteristics of the search space and the objective function, avoiding ineffective searches and redundant computations, thereby improving search efficiency. It also has the ability to handle noise and uncertainty. By establishing a probabilistic model and conducting random exploration, it reduces the impact of noise and increases the ability to explore the global optimal solution.

[0058] 5. This solution retains more sensor pattern noise for the input video frame, uses an affine transformation model and combines it with a Bayesian adaptive summary search algorithm to restore the tested sensor pattern noise. Therefore, this solution has higher recognition effect, robustness and faster computing speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the services required for use in the embodiments or the prior art descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 This is a principle flow chart of a stable video source identification method of the present invention;

[0061] Figure 2 It is a decomposition diagram based on undecimated dual-tree complex wavelet transform. DETAILED DESCRIPTION

[0062] The present invention will be further described below in conjunction with specific embodiments:

[0063] like Figure 1 As shown, the method for identifying a stable video source according to this embodiment includes the following steps:

[0064] S1, extract sensor pattern noise from multiple reference stabilized videos;

[0065] S2, extracting sensor pattern noise from the stable video to be detected;

[0066] In the above steps S1 and S2, extracting the corresponding sensor pattern noise includes:

[0067] A1. Convert the video back into bitstream data, intervene in the decoding process, and output all video frames before the codec loop filter;

[0068] A2. Only the key frame intra-coded pictures are selected, and each video frame is subjected to undecimated dual-tree complex wavelet transform and decomposed into J-layer wavelet coefficients to obtain detailed information in all directions.

[0069] Specifically, if Figure 2 As shown in Figure 2, the specific process of unsampled dual-tree complex wavelet decomposition transform is:

[0070] Only the first layer is downsampled, while other layers are not downsampled, which significantly reduces over-completion while maintaining exact translation invariance, and also provides improved directional selectivity and complex sub-bands;

[0071] The filters in each stage are based on the filters used in the DTCWT, and any set of perfectly reconstructed biorthogonal filters can be used as the first stage filters. The filters in tree(g00(n),g01(n)) and tree(g00(n),g01(n)) are exactly the same, but offset by one sample. The first stage high-pass filter of the UDTCWT can be represented by g.

[0072]

[0073]

[0074]

[0075] in are the real and imaginary outputs, respectively, which indicates that the same filter is used, but offset by one sample to generate each component. The definition of each complex coefficient y[n] is shown in equation (3).

[0076] Then, after shifting the input signal to the right by one sampling point, a new complex output y can be obtained. d [n], which is defined by Equations 4, 5, and 6. This is equivalent to the original untranslated real and imaginary parts The sum of is shown in Equation 6:

[0077]

[0078]

[0079]

[0080] The sub-band energy is obtained by calculating the sum of the squares of the signals in each sub-band. The sub-band energy is defined as:

[0081]

[0082]

[0083] In UDTCWT, the filter response for each subband is exactly the same, differing only by the phase offset. Therefore, a shift in the input signal only changes the coefficient values ​​in the output subband by the phase offset, without affecting their energy. This means that the energy remains constant for any offset, resulting in perfect shift invariance. If no downsampling occurs at subsequent levels, the lower-scale subbands will also exhibit perfect shift invariance.

[0084] A3. Using a method based on minimum mean square error (MMSE) estimation to perform spatial adaptive denoising on the wavelet coefficients to obtain denoised wavelet coefficients;

[0085] The formula for minimum mean square error is as follows:

[0086]

[0087] In formula (9), W in Expressed as the wavelet coefficient before filtering, W out Expressed as the wavelet coefficient after filtering, the noise variance is estimated as The subband variance is

[0088] A4. Reconstruct the denoised wavelet coefficients using an inverse undecimated dual-tree complex wavelet transform, thereby achieving perfect reconstruction and obtaining a denoised video frame; subtract the input video frame from the denoised video frame to obtain the sensor pattern noise of each video frame.

[0089] S3. Use the affine transformation model combined with the Bayesian adaptive direct search algorithm to restore the sensor pattern noise of the stable video that needs to be detected.

[0090] The affine transformation model is shown in formula (10):

[0091] Apply geometric transformation to each acquired test frame and select affine transformation model for processing. Use formula (10) to model the inverse coordinate mapping between the stabilized test video frame and the original frame pixels:

[0092]

[0093] Where (x′, y′) is the coordinate of the stabilized pixel, and (x, y) is the coordinate of the original pixel after restoration. T represents the affine transformation matrix, the parameter s in the matrix represents the scaling ratio, θ represents the rotation angle, and c x Indicates horizontal shift, c y Indicates vertical shift.

[0094] Bayesian Adaptive Direct Search (BADS) is a global-local random search algorithm that combines the performance of MADS with the BO search performed by a GP agent. BADS uses multi-scale search and gridding to combine global and local search to find the optimal solution. The algorithm consists of two steps: the Search stage and the Poll stage.

[0095] 1. Search stage

[0096] An active search strategy based on Gaussian processes and Bayesian optimization is used to minimize the number of objective function evaluations, effectively explore uncertain regions and develop high-predictive-value regions, quickly finding the global optimal solution:

[0097]

[0098] Before the i-th iteration, the optimal solution of the objective function PCE is expressed as x i , that is, {s,θ,c}, the calculated feasible solution set is defined as X i The set of all search points is defined as M i ,in is the size parameter of the search grid, D is the set of grid directions, and z is a full-rank positive integer matrix.

[0099] The Search stage is mainly divided into four steps:

[0100] Step 1: Take x i Construct grid cells for the search center.

[0101] Step 2: Calculate the objective values ​​of the finite grid points near the constructed grid elements and find a feasible solution to the improved objective function.

[0102] Step 3: If a feasible solution that improves the objective function is found, the search is successful. At this time, the grid center is moved to that position and the grid size parameter is increased in the i+1th iteration step.

[0103] Step 4: Failure to find a feasible solution to improve the objective function indicates that the search process of the iterative step has failed. Then turn to the next optimization step, the poll stage, and then reduce the grid size parameter in the i+1th iteration step.

[0104] Poll stage

[0105] When the search phase fails to find a feasible solution, the polling phase begins. Polling searches for a feasible solution by evaluating different directions at the grid points, adjusting the step size based on success or failure. The core idea is to narrow the area of ​​improvement.

[0106]

[0107] Here, a set Q consisting of directions with larger density is defined. i , matrix D i Consists of column elements concentrated in the grid direction.

[0108] Parameters control the size of the filtering box. The polling phase operates within the region surrounding the current optimal solution. Successful discovery of a feasible solution leads to the next polling phase. Failure to find a feasible solution leads to a reduction in parameters to avoid being trapped in a finite set. An update method avoids this trap during the polling phase, increasing the probability of finding the optimal direction. The BADS algorithm extends the global pattern search optimization strategy by generating a set of filtering points, resulting in rapid convergence and efficient local optimization capabilities.

[0109] S4. Calculate the correlation between the sensor pattern noise of the reference stabilized video and the sensor pattern noise of the test stabilized video after recovery in step S3 using peak correlation energy. If the PCE value is greater than or equal to a preset threshold, it is determined that the stabilized video to be tested comes from the camera that recorded the reference video. Otherwise, the stabilized video to be tested does not come from the camera that recorded the reference stabilized video.

[0110] Specifically, in this step, the peak correlation energy (PCE) is as shown in formula (13):

[0111]

[0112] where ρ is the normalized cross-correlation between the reference pattern noise and the pattern noise of the test image, s is the mapping of all elements in ρ, Ω represents a small neighborhood centered at ρ(0,0), and |Ω| represents the number of elements in Ω.

[0113] In order to align the extracted pattern noise with the reference pattern noise, we perform an affine transformation (W) on the pattern noise of each test frame so that the PCE value between the transformed pattern noise and the camera reference pattern noise is maximized:

[0114]

[0115] Among them, W test , I testDenotes the pattern noise of the test video frame and the original video frame, K denotes the estimated reference video fingerprint, and p denotes the maximum PCE value between the test pattern noise and the reference video pattern noise when the search parameters are {s, θ, c}. The maximization problem in Equation (14) can be solved using the proposed Bayesian adaptive direct search algorithm.

[0116] Finally, for the N video frames tested, take the highest PCE value P in the frame set max As the final result of this test video and the reference video.

[0117]

[0118] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, any changes made based on the shape and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for identifying a stable video source, characterized in that: The following steps are involved: S1, extract sensor pattern noise from multiple reference stabilized videos; S2, extracting sensor pattern noise from the stable video to be detected; S3, using the affine transformation model combined with the Bayesian adaptive direct search algorithm to restore the sensor pattern noise of the stable video to be detected; S4. Calculate, using peak correlation energy, the correlation between the sensor pattern noise of the reference stabilized video and the sensor pattern noise of the test stabilized video recovered in step S3. If the peak correlation energy value is greater than or equal to a preset threshold, determine that the stabilized video to be tested is from the camera that recorded the reference stabilized video. On the contrary, the stabilized video to be detected is not from the camera that recorded the reference stabilized video; Step S3 includes: Apply a geometric transformation to each acquired test frame and select an affine transformation model for processing, which represents the inverse coordinate mapping between the stabilized test video frame and the original frame pixels: Where (x′, y′) is the coordinate of the stabilized pixel, (x, y) is the coordinate of the original pixel after restoration; T represents the affine transformation matrix, the parameter s in the matrix represents the scaling ratio, θ represents the rotation angle, and c x Indicates horizontal shift, c y Indicates vertical displacement; The Bayesian adaptive direct search algorithm searches at multiple scales, divides the search space using the idea of ​​grid partitioning, and combines global search with local search to find the optimal solution.

2. A stable video source identification method according to claim 1, characterized in that: In steps S1 and S2, extracting the corresponding sensor pattern noise includes: A1. Convert the video back into bitstream data, intervene in the decoding process, and output all video frames before the codec loop filter; A2. Only the key frame intra-coded pictures are selected, and each video frame is subjected to undecimated dual-tree complex wavelet transform and decomposed into J-layer wavelet coefficients to obtain detailed information in all directions. A3. Use a method based on minimum mean square error estimation to perform spatial adaptive denoising on the wavelet coefficients to obtain denoised wavelet coefficients; A4. Reconstruct the denoised wavelet coefficients using an inverse undecimated dual-tree complex wavelet transform, thereby achieving perfect reconstruction and obtaining a denoised video frame; subtract the input video frame from the denoised video frame to obtain the sensor pattern noise of each video frame.

3. A stable video source identification method according to claim 2, characterized in that: In step A2, the specific process of the unsampled dual-tree complex wavelet decomposition transform is: The undecimated dual-tree complex wavelet transform decomposes and reconstructs the signal by downsampling only the first layer but not the remaining layers. The first tree generates the real part and the second generates the imaginary part. The sampling frequency of the filters in the two trees is the same, and the delay between them is one sampling interval. The binary decimation of the first layer in the imaginary part tree just samples the sample values ​​lost by the binary decimation in the real part tree.

4. The method for identifying a stable video source according to claim 2, wherein: In step A3, the wavelet coefficients are spatially adaptively denoised using a method based on minimum mean square error estimation. The formula for obtaining the denoised wavelet coefficients is as follows: In formula (9), W in Expressed as the wavelet coefficient before filtering, W out Expressed as the wavelet coefficient after filtering, the noise variance is estimated as The subband variance is 5. The method for identifying a stable video source according to claim 1, wherein: Bayesian adaptive direct search consists of two steps: Search stage and Poll stage; Search stage: An active search strategy based on Gaussian processes and Bayesian optimization is used to minimize the number of objective function evaluations, effectively explore uncertain regions and develop high-predictive-value regions, quickly finding the global optimal solution: Before the i-th iteration, the optimal solution of the objective function is expressed as x i , that is, {s,θ,c}, the calculated feasible solution set is defined as X i ; The set of all search points is defined as M i ,in is the size parameter of the search grid, D is the grid direction set, and z is a full-rank positive integer matrix; The Search stage is divided into four steps: Step 1: x i Construct grid cells for search centers; Step 2: Calculate the objective values ​​of the finite grid points near the constructed grid elements and find a feasible solution to the improved objective function; Step 3: If a feasible solution to the improved objective function is found, the search is successful; at this time, the grid center is moved to the feasible solution position and the grid size parameter is increased in the i+1th iteration step. Step 4: Failure to find a feasible solution to improve the objective function indicates that the search process of the iterative step has failed; then turn to the next optimization step, the poll stage, and then reduce the grid size parameter in the i+1th iteration step Poll stage: When the search phase fails to find a feasible solution, the polling phase begins. Polling searches for a feasible solution by evaluating different directions at the grid points, with the step size adjusted based on success or failure. The core idea is to narrow the area of ​​improvement. Among them, a set Q consisting of directions with larger density is defined i , matrix D i It consists of column elements concentrated in the grid direction; The parameters control the size of the filtering box. The polling phase is performed in the area around the current best solution. If a feasible solution is successfully found, the next polling phase will be continued. If a feasible solution fails, the parameters will be reduced to avoid falling into a finite set. The update method is used to avoid falling into a finite set during the polling phase and increase the probability of finding the optimal direction.

6. The method for identifying a stable video source according to claim 1, wherein: In step S4, the peak correlation energy PCE is shown in formula (13): Among them, ρ (0,0) is the normalized cross-correlation value between the reference pattern noise and the pattern noise of the test image at coordinate (0,0), τ is the mapping of all elements in ρ, Ω represents a small neighborhood centered at (0,0), and |Ω| represents the number of elements in Ω.

7. A stable video source identification method according to claim 6, characterized in that: In order to align the extracted pattern noise with the reference pattern noise, an affine transformation Γ(W) is performed on the pattern noise of each test frame so that the PCE value between the transformed pattern noise and the camera reference pattern noise is maximized: Among them, W test , I test represents the pattern noise of the test video frame and the original video frame, K represents the estimated reference video fingerprint, and p represents the maximum PCE value between the test pattern noise and the reference video pattern noise when the search parameters are {s, θ, c}; the maximization problem in Equation (14) is solved using the proposed Bayesian adaptive direct search algorithm; Finally, for the N video frames tested, take the highest PCE value P in the frame set max As the final result of this test video and the reference video:

Citation Information

Patent Citations

  • Short compressed video source identification method

    CN115550686A

  • Image stabilization techniques for video surveillance systems

    WO2014075022A1