Monitoring data tamper-proofing method, device and storage medium
Patent Information
- Application Number
- CN202610721225.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-05-25
AI Technical Summary
[0004]本申请实施例提供了一种监控数据防篡改方法、设备及存储介质,可以解决现有技术存在计算量大、成本高,或需专用硬件,难以适应普通监控设备简单、实时性的需求的问题
本申请实施例提供的监控数据防篡改方法,通过在监控数据录制过程中,对每一视频帧同步采集一个成像传感器的热噪声片段,降低了设备成本和部署难度。计算视频帧的图像活动度量值,以及热噪声片段的噪声能量度量值,计算量小,可实时与视频录制同步执行,适应普通监控设备的简单实时需求。将图像活动度量值与噪声能量度量值进行算术组合,得到帧的校验特征值,并与视频帧关联存储,难以伪造,且存储开销低。获取待验证视频帧,并根据待验证视频帧生成验证特征值,在验证特征值与已存储的校验特征值偏差超出预设范围时,确定监控数据存在篡改,实现了低成本、低复杂度、高实时性的监控数据完整性保护。
Smart Images

Figure CN122289913B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of surveillance data processing technology, and in particular relates to methods, devices and storage media for preventing surveillance data from being tampered with. Background Technology
[0002] Video surveillance data holds significant evidentiary value in fields such as security, law enforcement, and transportation. However, with the widespread use of video editing software, tampering with surveillance videos (such as deleting, replacing, or inserting frames, or modifying the content) has become easy and difficult to detect. Therefore, effective methods are needed to detect whether surveillance data has been tampered with to ensure its integrity.
[0003] Currently, commonly used anti-tampering methods mainly include digital signatures, digital watermarks, and verification based on external sensors. Digital signatures are computationally intensive and require secure key management; digital watermarks may affect image quality and are easily removed; environmental audio-based methods require microphone hardware, but many surveillance cameras are not equipped with audio acquisition modules. Furthermore, in the field of image forensics, there has been research on detecting static image tampering using sensor noise (such as non-uniformity of light response and thermal noise), but these methods are geared towards offline forensics, requiring complex noise modeling or multiple sampling, and cannot generate verification features in real time during video recording. They are also difficult to resist temporal attacks such as frame deletion and insertion targeting the video stream. Existing methods are either computationally intensive and costly, or require dedicated hardware, making them unsuitable for the simple and real-time requirements of ordinary surveillance equipment. Summary of the Invention
[0004] This application provides a method, device, and storage medium for preventing tampering of monitoring data, which can solve the problems of existing technologies having large computational loads, high costs, or requiring dedicated hardware, making it difficult to meet the simple and real-time needs of ordinary monitoring equipment.
[0005] In a first aspect, embodiments of this application provide a method for preventing tampering with monitoring data, including: During the monitoring data recording process, a thermal noise segment from an imaging sensor is synchronously acquired for each video frame; wherein, the thermal noise segment is taken from a pixel block with flat texture in the video frame; Calculate the image activity metric of the video frame and the noise energy metric of the thermal noise segment; The image activity metric and the noise energy metric are arithmetically combined to obtain the frame's verification feature value, which is then associated with and stored with the video frame. The system acquires a video frame to be verified and generates a verification feature value based on the video frame. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the monitoring data has been tampered with.
[0006] The technical solutions described in this application embodiment have at least the following technical effects: The monitoring data anti-tampering method provided in this application reduces equipment costs and deployment difficulty by synchronously acquiring a thermal noise segment from an imaging sensor for each video frame during the monitoring data recording process. Calculating the image activity metric of the video frame and the noise energy metric of the thermal noise segment involves minimal computation and can be performed in real-time synchronously with video recording, meeting the simple real-time needs of ordinary monitoring equipment. The image activity metric and noise energy metric are arithmetically combined to obtain the frame's verification feature value, which is then associated with and stored with the video frame, making it difficult to forge and resulting in low storage overhead. The method acquires the video frame to be verified and generates a verification feature value based on it. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it determines that the monitoring data has been tampered with, achieving low-cost, low-complexity, and high-real-time monitoring data integrity protection.
[0007] Secondly, embodiments of this application provide a monitoring data anti-tampering system, comprising: The acquisition unit is used to synchronously acquire a thermal noise segment from an imaging sensor for each video frame during the monitoring data recording process; wherein the thermal noise segment is taken from a pixel block with flat texture in the video frame. A measurement unit is used to calculate the image activity measurement value of the video frame and the noise energy measurement value of the thermal noise segment; The association unit is used to perform an arithmetic combination of the image activity metric and the noise energy metric to obtain the frame verification feature value, and store it in association with the video frame; The detection unit is used to acquire the video frame to be verified and generate a verification feature value based on the video frame to be verified. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the monitoring data has been tampered with.
[0008] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any of the foregoing aspects.
[0009] Fourthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to perform the method described in any of the above aspects.
[0010] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the above aspects, and will not be repeated here. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a method for preventing tampering with monitoring data provided in an embodiment of this application; Figure 2 This is a schematic diagram of a thermal noise segment in a monitoring data anti-tampering method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a monitoring data anti-tampering system provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0016] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determination" or "if the described condition or event is detected" may be interpreted, depending on the context, as "once determination," "in response to determination," "once the described condition or event is detected," or "in response to the detection of the described condition or event."
[0017] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0018] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0019] Video surveillance data holds significant evidentiary value in fields such as security, law enforcement, and transportation. However, with the widespread use of video editing software, tampering with surveillance videos (such as deleting, replacing, or inserting frames, or modifying the content) has become easy and difficult to detect. Therefore, effective methods are needed to detect whether surveillance data has been tampered with to ensure its integrity.
[0020] Currently, commonly used anti-tampering methods mainly include digital signatures, digital watermarks, and verification based on external sensors.
[0021] Digital signature methods calculate and encrypt a hash value for each frame or the entire video during recording, and then recalculate and compare the hash during verification. This method offers high security but is computationally intensive, requiring additional encryption hardware or secure storage keys, placing a heavy burden on embedded monitoring devices. Digital watermarking embeds information into the pixels or transform domain of video frames, but this can affect image quality, and attackers can bypass it through watermark removal or rewriting attacks.
[0022] Other methods utilize the synchronization between external signals in the recording environment and the video for verification. For example, ambient audio can be captured simultaneously, and a checksum can be generated by combining audio energy characteristics with video frame activity metrics, which is then compared during verification. These methods do not rely on encryption and have relatively low computational requirements, but they require microphone hardware, and many surveillance cameras (especially outdoor or low-cost models) are not equipped with audio capture modules; furthermore, ambient audio is susceptible to background noise interference and it is difficult to guarantee strict synchronization with video frames.
[0023] In the field of image forensics, researchers have proposed using sensor noise (such as photoresponse nonuniformity PRNU, thermal noise, etc.) to detect whether static images have been tampered with. However, these methods are mainly aimed at offline forensics of static images, requiring complex noise modeling (such as wavelet denoising, multi-frame averaging to extract PRNU), high computational overhead, and cannot effectively utilize the temporal continuity of video. In addition, they cannot generate verification features for subsequent verification in real time during video recording, nor do they have the ability to resist frame-level attacks against video streams (such as frame deletion, insertion, or dynamic position attacks).
[0024] Therefore, there is an urgent need for a method to prevent tampering of monitoring data that does not rely on additional hardware, has low computational requirements, can be synchronized with the video recording process, and effectively utilizes the physical randomness of sensors.
[0025] To address the aforementioned issues, this application provides a method, device, and storage medium for preventing tampering of monitoring data. In this method, by synchronously acquiring a thermal noise segment from an imaging sensor for each video frame during the monitoring data recording process, equipment costs and deployment complexity are reduced. The calculation of image activity metrics for the video frame and noise energy metrics for the thermal noise segment involves minimal computation and can be performed synchronously with video recording in real time, meeting the simple real-time needs of ordinary monitoring equipment. An arithmetic combination of the image activity metrics and noise energy metrics yields a frame verification feature value, which is then associated with and stored with the video frame, making it difficult to forge and resulting in low storage overhead. A video frame to be verified is acquired, and a verification feature value is generated based on the frame. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the monitoring data has been tampered with, achieving low-cost, low-complexity, and high-real-time monitoring data integrity protection.
[0026] The monitoring data anti-tampering method provided in this application embodiment can be applied to electronic devices. In this case, the electronic device is the executing subject of the monitoring data anti-tampering method provided in this application embodiment. This application embodiment does not impose any restrictions on the specific type of electronic device.
[0027] For example, electronic devices can be ultra-mobile personal computers (UMPCs), netbooks, desktop computers, computers, laptops, communication equipment, computing devices, satellite wireless equipment, etc.
[0028] To better understand the monitoring data anti-tampering method provided in the embodiments of this application, the specific implementation process of the monitoring data anti-tampering method provided in the embodiments of this application will be described by way of example below.
[0029] Figure 1This illustration shows a schematic flowchart of a monitoring data anti-tampering method provided in an embodiment of this application. The monitoring data anti-tampering method includes: During the monitoring data recording process, the S100 synchronously acquires a thermal noise segment from the imaging sensor for each video frame; the thermal noise segment is taken from a pixel block with flat texture in the video frame.
[0030] It is understandable that each frame of an image output by a surveillance device (such as a network camera, analog camera, dashcam, or mobile phone camera) during video recording can be recorded as a video frame. Please refer to [link / reference]. Figure 2 Thermal noise in imaging sensors refers to the random electrical signals generated by the thermal motion of electrons within pixels during the photoelectric conversion process of image sensors (such as CMOS or CCD), manifesting as tiny random fluctuations in pixel values. Thermal noise is an inherent physical characteristic of the sensor and can be extracted from video frames without additional hardware. Thermal noise segments are taken from flat pixel blocks within the video frame. Flat pixel blocks refer to local areas in the image with small grayscale variations and lacking edges and details (such as the sky, walls, ground, and vignetting). Within these areas, pixel value variations are primarily driven by sensor noise, rather than scene content. Synchronous acquisition means that for each frame, before encoding or storage, thermal noise segments are extracted immediately from that frame to ensure a strict correspondence between the noise and the video frame. A thermal noise segment is essentially a set of raw pixel values extracted from this flat pixel block.
[0031] Common anti-tampering methods in existing technologies, such as digital watermarking or ambient audio synchronization, often require additional hardware (microphones) or complex encryption calculations, making it impossible for ordinary surveillance cameras to achieve real-time anti-tampering at low cost. This application utilizes the inherent thermal noise of the imaging sensor itself. Thermal noise is a random pixel fluctuation caused by the thermal motion of electrons within the sensor. Each camera has a different thermal noise pattern that is difficult to replicate. No additional hardware is required, resulting in extremely low computational cost. Furthermore, the thermal noise is independent of the video content, making it difficult to forge. It is worth noting that the flat pixel blocks are not fixed; this application will dynamically select their positions to prevent attackers from knowing and preserving this area in advance.
[0032] Thermal noise can be extracted from each frame of the original image (RAW or YUV format) during the preprocessing stage of the video encoder to avoid noise degradation caused by compression or noise reduction processing. Alternatively, it can be extracted from a fixed region of the video frame (e.g., the top-left 32×32 pixel block), but it is preferable to dynamically select flat regions based on the content to enhance security.
[0033] In one possible implementation, S100, during the monitoring data recording process, synchronously acquires a thermal noise segment from the imaging sensor for each video frame, including: S110: Divide the video frame into multiple first candidate blocks, and select several second candidate blocks with the best flatness from the multiple first candidate blocks based on the Laplacian response value of each pixel in the video frame.
[0034] The first candidate block is a set of rectangular regions uniformly divided into video frames of a fixed size (e.g., 32×32, 64×64 pixels), serving as potential sources of thermal noise. The Laplacian response is a commonly used second-order differential operator in image processing, used to detect the degree of grayscale difference between a pixel and its neighbors. For each pixel, a larger absolute value of the Laplacian response indicates that the pixel is located in an edge or textured region; a response value close to zero indicates that the pixel is in a flat region. By calculating the sum of the absolute values of the Laplacian responses of all pixels within each first candidate block, the flatness of the block can be quantified (the smaller the sum, the flatter the block). The selected second candidate blocks are several first candidate blocks with optimal flatness (i.e., the smallest sum) (e.g., selecting the first four flattest blocks), used for subsequent random selection to minimize interference from scene content in the extracted thermal noise segments. The number of blocks K can be 16, 32, or 64, and the block size can be the same or adaptively adjusted according to the video resolution. The Laplacian response value can be quickly calculated using a 3×3 convolution kernel (e.g., [[0,1,0],[1,-4,1],[0,1,0]]). Alternatively, other edge detection operators (such as Sobel and Canny) can be used instead of Laplacian, but Laplacian offers advantages such as isotropy and computational simplicity, making it suitable for low-cost monitoring scenarios.
[0035] Optionally, in step S110, the video frame is divided into multiple first candidate blocks, and based on the Laplacian response value of each pixel in the video frame, several second candidate blocks with the best flatness are selected from the multiple first candidate blocks, including: S111, divide the video frame into K first candidate blocks and calculate the Laplacian operator response value of each pixel in the video frame; where K≥16.
[0036] As can be understood, K is a preset integer parameter representing the number of blocks after uniformly dividing the video frame horizontally and vertically. K ≥ 16 ensures that the size of each candidate block is small enough (e.g., for a 1920×1080 video, each block is approximately 120×67 pixels when K=16), resulting in relatively simple textures within the blocks. Calculating the Laplacian operator response value for each pixel typically employs a discrete convolution operation. The Laplacian operator response value is used not only to filter flat blocks but also for subsequent calculations of image activity metrics.
[0037] K can be 16, 25, 32, or 64. For higher resolution videos (such as 4K), K can be increased appropriately to maintain a reasonable block size. Laplacian calculation can be performed using a fast convolution algorithm or hardware acceleration (such as GPU or DSP), or the Laplacian response value of each pixel can be pre-calculated and stored for reuse in subsequent steps.
[0038] S112, for each first candidate block, accumulate the absolute values of the Laplacian responses of all pixels within the first candidate block as a flatness index of the first candidate block; where, the smaller the accumulated value, the flatter the first candidate block is.
[0039] As we understand it, the flatness metric is a scalar used to measure the texture complexity of an image patch. Accumulating the absolute values instead of the original values (which can be positive or negative) avoids cancellation and accurately reflects edge strength. A smaller flatness metric indicates fewer overall edges and weaker texture within the patch, making it more suitable for extracting thermal noise (because scene content contributes less to pixel values, while noise accounts for a relatively high proportion). The flatness metric can also be viewed as a "measure of local image activity." The accumulated sum can be used directly as the flatness metric without normalization (because all patches are the same size). Variance or standard deviation can also be used as the flatness metric, but Laplacian accumulation is computationally less expensive.
[0040] S120: Select a target block from all second candidate blocks based on pseudo-random rules, extract the original pixel values of the target block as thermal noise segments of the video frame, and record the position information of the selected target block.
[0041] It is understandable that after selecting L of the flattest blocks (second candidate blocks), if the block with the best flatness is always chosen, the position remains fixed (because the flattest block may be the same in adjacent frames), and an attacker can still selectively retain that block. To further enhance security, this application introduces a pseudo-random rule, randomly selecting one block from the L candidate blocks as the target block each frame. Pseudo-random means that the selection result appears random, but is actually determined by a reproducible seed (such as the frame number), so the same selection can be regenerated during verification. Extracting the original pixel values of the target block (without noise reduction, sharpening, etc.) preserves the purest thermal noise. Recording position information (such as row index and column index) is for quick location during verification, avoiding recalculation of flatness and random selection (although recalculation is possible, recording speeds up the process). Pseudo-random selection makes it impossible for attackers to predict which block will be selected. Attackers cannot predict which block will be selected and must retain all candidate blocks simultaneously to ensure they are not detected, which greatly increases the cost of tampering. At the same time, since the selection process is reproducible, the same position can be restored without additional information during the verification phase. A pseudo-random number generator can use the `rand()` function from the C standard library, with the seed set to the current frame's PTS (display timestamp). Alternatively, a cryptographically secure pseudo-random number generator (such as ChaCha20) can be used to enhance unpredictability. Location information can be stored along with the verification feature value or separately in the metadata.
[0042] Optionally, in step S120, a target block is selected from all second candidate blocks based on a pseudo-random rule, the original pixel values of the target block are extracted as thermal noise segments of the video frame, and the location information of the selected target block is recorded, including: S121, using the temporal identifier of the video frame as the seed of the pseudo-random number generator, a target block is selected from L second candidate blocks with uniform probability and pseudo-randomness.
[0043] As we understand it, a timing identifier is a unique value that identifies the order or time of video frames. It can include frame sequence numbers (incrementing integers starting from 1), recording timestamps (such as Unix milliseconds), PTS (display timestamp), or combinations thereof. Using a timing identifier as a random seed makes the selection of each frame independent and reproducible: using the same timing identifier during verification yields the same random sequence, thus locating the same target block. Uniform probability pseudo-randomization means that each second candidate block has an equal probability of being selected (1 / L), avoiding bias towards certain specific positions and improving security. L is the number of second candidate blocks (e.g., L=2, 4, or 8). L=4 can be chosen, meaning one block is randomly selected with equal probability from the four flattest blocks. A portion of the hash value of the video frame (such as MD5) can also be used as a seed to further enhance randomness. The seed can also be combined with a device-unique key (such as the device serial number) to prevent cross-device copying attacks.
[0044] S122, extract the original pixel values of all pixels in the target block to form a thermal noise segment of the video frame, and record the index coordinates of the target block in the video frame.
[0045] As can be understood, raw pixel values refer to pixel values directly output from the sensor or after minimal processing (such as black level correction and linearization), typically without non-linear processing such as noise reduction, sharpening, or color enhancement. Thermal noise segments are essentially two-dimensional arrays or one-dimensional sequences composed of these raw pixel values. Index coordinates are used to uniquely identify the position of a target block within the entire video frame; for example, (row, col) represents the row and column number of the block, or (top-left_x, top-left_y, width, height) represents the absolute position. The purpose of recording index coordinates is to allow the thermal noise segment of the frame to be re-extracted according to the same rules during verification. Index coordinates can be stored as two bytes of row and column numbers (e.g., uint8_t type, supporting up to 256 rows / columns). In another implementation, the absolute coordinates (x, y) of the top-left pixel of the block can be stored. Index coordinates are stored in association with verification feature values, for example, as part of the verification feature value metadata.
[0046] S200 calculates the image activity metric of the video frame and the noise energy metric of the thermal noise segment.
[0047] The image activity metric is a scalar used to describe the complexity or richness of detail of the entire video frame; a larger value indicates richer image texture and more edges. The noise energy metric describes the intensity of random fluctuations in pixel values within a thermal noise segment; a larger value indicates higher noise energy. These two metrics reflect the physical characteristics of the video frame from different perspectives: the image activity metric depends on the scene content, while the noise energy metric depends on the sensor state (temperature, gain). Combining them can create a difficult-to-forge verification feature because tampering operations typically change both the content and the noise distribution simultaneously. The image activity metric can be the sum of the absolute values of the full-image Laplacian response. The noise energy metric can be the sum of the squares of the pixel values within a thermal noise segment. The image activity metric can also be the sum of Sobel gradients or image entropy.
[0048] In one possible implementation, S200, calculating the image activity metric of the video frame and the noise energy metric of the thermal noise segment, including: S210, calculate the sum of the Laplacian operator response values of all pixels in the video frame, and use the sum of the absolute values of the Laplacian responses of all pixels as the first intermediate value.
[0049] It is understandable that the first intermediate value is denoted as ,in A represents the Laplacian response at pixel position (x, y), summed over all pixels in the entire video frame. The first intermediate value is sensitive to image details: if the video content is replaced, modified, or a scene change occurs, A will change significantly. Meanwhile, since the Laplacian is a second derivative, A is relatively insensitive to changes in illumination (e.g., brightness changes from day to night), but sensitive to changes in focus, lens smudges, etc. Therefore, when combined with thermal noise, it can distinguish between natural scene changes and malicious tampering. A is an important component for measuring image activity. Image boundaries can be ignored when calculating the Laplacian response (boundary pixels are filled with mirror or zero). Downsampling of video frames before calculation can reduce computational cost, but this may sacrifice detection sensitivity.
[0050] S220, calculate the square of the pixel value of each pixel in the thermal noise segment, and use the sum of the squares of all pixels in the thermal noise segment as the second intermediate value.
[0051] The second intermediate value can be denoted as B = Σ(p_i)², where p_i is the original pixel value of the i-th pixel within the thermal noise segment (usually an integer between 0 and 255 or 0 and 4095), and the summation covers all pixels within the segment. The sum of squares reflects the total energy of the noise segment, effectively removing the power measurement from the pixel symbol. Since thermal noise is approximately Gaussian white noise, its energy is related to factors such as sensor gain, temperature, and exposure time. During normal recording, the B value changes smoothly between consecutive frames; if a frame is replaced or altered (even if the pixel value change is small), the B value will jump abruptly. B serves as a component of the noise energy measurement. If the range of original pixel values is large (e.g., 10-bit), a 64-bit integer summation can be used to prevent overflow. Alternatively, the sum of the absolute values of the pixel values can be used instead of the sum of squares to reduce computation, but the sum of squares is more sensitive to large values and provides better detection results.
[0052] S230: Calculate the cross-correlation value between the Laplacian operator response value and the pixel value for each pixel in the thermal noise segment. Then, weight and combine the first intermediate value, the second intermediate value, and the cross-correlation value to obtain the final image activity metric and noise energy metric.
[0053] The cross-correlation value can be understood as C = (1 / N) * Σ(L_i * p_i), where L_i is the Laplacian response value at the i-th pixel position within the thermal noise segment (pre-calculated by S111), p_i is the original pixel value of that pixel, and N is the total number of pixels within the segment. The cross-correlation value reflects the statistical dependence between image edges (represented by the Laplacian response) and the original pixel value (containing thermal noise). In real, untampered videos, since thermal noise and image content are independent, the correlation between L_i and p_i is weak, and C is close to zero; however, tampering operations (such as copy-paste, splicing, and recompression) may disrupt this independence, causing C to deviate significantly from zero. Weighted combination refers to multiplying A, B, and C by preset weight coefficients (such as α, β, γ) and then summing them to obtain the final image activity metric (such as αA + βC) and noise energy metric (such as γB + δC). Setting the weight coefficients can adjust the contribution of each component to the final feature, improving the detection sensitivity for specific types of tampering. The values can be set as α=0.6, β=0.4, γ=0.7, and δ=0.3. The weighting coefficients can also be obtained through statistical learning on a batch of normal video samples.
[0054] Optionally, in step S230, the cross-correlation value between the Laplacian operator response value and the pixel value of each pixel in the thermal noise segment is calculated. The first intermediate value, the second intermediate value, and the cross-correlation value are then weighted and combined to obtain the final image activity metric and noise energy metric, including: S231, traverse each pixel within the thermal noise segment and calculate the cross-correlation value between the Laplacian operator response value and the pixel value of the thermal noise segment.
[0055] Traversal, as we understand it, refers to accessing each pixel location within the thermal noise segment sequentially (e.g., row-major or column-major), performing the same operation for each location. Calculating the cross-correlation requires simultaneously obtaining the Laplacian response value of the pixel (already calculated and stored in S111) and the original pixel value, then summing the products. Since the thermal noise segment size is typically small (e.g., 32×32=1024 pixels), traversal overhead is extremely low. Traversal can be combined with the sum of squares in S220 into the same loop to reduce memory accesses and loop counts. A single loop can be used to simultaneously accumulate the sum of squares and the sum of products. If the thermal noise segment size is large (e.g., 128×128), parallel acceleration can be achieved using SIMD instructions (e.g., AVX2).
[0056] For example, in S231, traversing each pixel within the thermal noise segment, the cross-correlation value between the Laplacian operator response value and the pixel value of the thermal noise segment is calculated, including: S2311, during the traversal, the product of the Laplacian operator response value and the pixel value corresponding to each pixel is accumulated. After the traversal is completed, the sum of the products is divided by the number of pixels in the thermal noise segment to obtain the cross-correlation value.
[0057] It can be understood that the cumulative sum Sum = Σ(L_i*p_i). The cross-correlation value C = Sum / N, which is the normalized mean of the product. Normalization can eliminate the influence of different thermal noise segment sizes, making C comparable across different frames and devices. If normalization is not performed and Sum is used directly as the cross-correlation value, the segment size must be fixed (e.g., all frames use thermal noise segments of the same size), otherwise the baseline of C will change with the size. Normalization is usually preferred to improve robustness. N is known in advance (e.g., fixed block size), and a division can be performed directly after traversal. To reduce floating-point operations, C can be magnified by a certain factor and stored as an integer.
[0058] S232, multiply the first intermediate value by the first weighting coefficient and sum it with the cross-correlation value multiplied by the second weighting coefficient to obtain the image activity metric.
[0059] The image activity metric can be understood as w1*A + w2*C. Here, A is the sum of the absolute values of the Laplacian of the entire image (reflecting global texture), and C is the local cross-correlation value (reflecting the independence of edges from noise). By weighted summation, the information from both can be combined: if an attacker only modifies flat areas while preserving edge areas, A may not change significantly, but C will change significantly due to the altered correlation between the Laplacian response of flat areas and pixel values, thus being detected. The first weighting coefficient w1 and the second weighting coefficient w2 are preset constants, ranging from 0 to 1. w1 = 0.8, w2 = 0.2. The weights can be dynamically adjusted based on the video content; for example, increasing the weight of w2 in low-texture scenes to adapt to different types of monitoring scenarios (such as indoor / outdoor switching, changes in lighting, etc.).
[0060] S233, multiply the second intermediate value by the third weighting coefficient and sum it with the cross-correlation value multiplied by the fourth weighting coefficient to obtain the noise energy metric.
[0061] The noise energy metric can be understood as w3*B + w4*C. Here, B is the sum of squares of pixels within the thermal noise segment (reflecting noise energy), and C is the cross-correlation value (reflecting edge-noise independence). In normal video, the trends of B and C are relatively consistent; tampering may cause B to change while C remains unchanged (e.g., modifying pixel values without disrupting edge-noise independence), or C to change abruptly while B remains flat (e.g., inserting an edge into a flat area). By weighting and combining these values separately, we can distinguish between normal scene changes (where B and C change synchronously) and malicious tampering (where B and C change asynchronously). The third weighting coefficient w3 and the fourth weighting coefficient w4 are preset constants: w3 = 0.9, w4 = 0.1. The weights can be adaptively adjusted based on the sensor model or ambient temperature.
[0062] S300 performs an arithmetic combination of the image activity metric and the noise energy metric to obtain the frame's verification feature value, and stores it in association with the video frame.
[0063] It's understandable that arithmetic combinations can be addition, multiplication, logarithmic operations, or combinations thereof, with the aim of fusing two metrics into a single scalar value as the integrity fingerprint of the frame. The verification feature value is the final stored data used for subsequent verification. Associated storage refers to saving the verification feature value along with the corresponding video frame (or the video frame index), which can be stored in the video file's metadata, a separate file, or embedded in the user data field of the video stream. During verification, the pre-stored verification feature value is read from the storage medium and compared with the features calculated in real-time for the frame to be verified. The image activity metric and noise energy metric can be normalized using the maximum value of the first N frames, then multiplied and the natural logarithm taken. Alternatively, they can be directly added or weighted summed. Associated storage can use H.264 / HEVC SEI (Supplemental Enhancement Information) messages or a separate database.
[0064] S400: Obtain the video frame to be verified and generate a verification feature value based on the video frame to be verified. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the video frame to be verified has been tampered with.
[0065] It can be understood that the video frame to be verified refers to a specific frame (which can be all frames or a sampled frame) in a video file whose integrity needs to be checked. The verification feature value is a value obtained using the exact same calculation method as the check feature value (including thermal noise segment extraction, image activity measurement calculation, noise energy measurement calculation, and arithmetic combination). The deviation can be an absolute difference, relative difference, or Euclidean distance, etc. The preset range is a threshold or dynamic boundary (e.g., the upper and lower limits calculated by an exponentially weighted moving average). When the deviation exceeds the range, it is determined that the frame or adjacent frames have been tampered with. Verification can be performed frame by frame, or by sampling verification at fixed frame intervals to reduce computational load. The preset range can be set to a fixed threshold (e.g., 5%), or a dynamic update algorithm can be used (e.g., calculating the mean and standard deviation using feature values from the previous few frames). When the deviation exceeds the range, the deviation pattern can be further analyzed to distinguish between tampering and scene switching. It is worth noting that scene switching (such as switching from indoor to outdoor) can cause significant changes in feature values, but scene switching is legal. Therefore, subsequent steps (such as S430 or using first-order differential judgment) are needed to distinguish between tampering and scene switching. Ultimately, this achieves low-cost, low-complexity, and high-real-time anti-tampering detection for ordinary monitoring equipment.
[0066] In one possible implementation, S400, the video frame to be verified is acquired, and a verification feature value is generated based on the video frame to be verified. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the video frame to be verified has been tampered with, including: S410: Obtain the video frame to be verified and extract the corresponding thermal noise segment from the video frame to be verified.
[0067] It's understandable that the extraction method for the video frames to be verified is the same as during recording, such as decoding the raw image data of each frame from the video file. The corresponding thermal noise segment refers to the pixel block extracted from that frame according to the same rules (including dividing candidate blocks, flatness filtering, and pseudo-random block selection). Since the index coordinates or seed of each frame's thermal noise segment are recorded during recording, the selection process can be reproduced during verification to obtain the same thermal noise segment. If the original seed or coordinates cannot be obtained, they can also be obtained by recalculating the flatness and pseudo-random selection (requiring the seed to be consistent with that during recording). The stored index coordinates of each frame can be read from the video metadata to directly locate the thermal noise segment. If the coordinates are not stored, the same temporal identifier is used as the seed to recalculate the block selection process.
[0068] S420, Generate verification feature values based on the video frames and thermal noise segments of the video to be verified; wherein, the verification feature values are calculated in the same way as the check feature values.
[0069] It is understandable that the verification feature values use the same algorithm as the recording stage, meaning the verification and validation feature values are calculated in the same way: calculating the image activity metric of the video frame to be verified, the noise energy metric of the thermal noise segment, and all arithmetic combinations. Ensuring algorithm consistency is a prerequisite for effective comparison. After the verification feature values are generated, they are used for comparison with the pre-stored validation feature values. Parameters from the recording process (such as weighting coefficients, K values, L values, block sizes, etc.) can also be stored as metadata, and these parameters are read during verification to ensure consistency. In another implementation, these parameters can be preset in the system, and the same set of preset parameters is used for both recording and verification.
[0070] S430: Compare the difference between the verification feature value and the stored verification feature value as the deviation. When the deviation exceeds a preset threshold, it is determined that the video frame to be verified has been tampered with.
[0071] It is understandable that the difference can be an absolute difference, a relative difference, or the square of the difference. The preset threshold can be a fixed value (e.g., 0.1) or a boundary dynamically calculated based on historical data (e.g., mean ± 3 standard deviations). When the deviation exceeds the threshold, it indicates that the features of the current frame are inconsistent with the features at the time of original recording, and it is highly likely that the frame or adjacent frames (because inter-frame predictive coding may propagate tampering) have been modified. It is worth noting that if the deviation of a frame exceeds the range, it is necessary to further analyze its adjacent frames: if only a single frame is abnormal while the frames before and after are normal, it is likely that the frame has been deleted or replaced; if multiple consecutive frames are abnormal and the changes are gradual, it may be a scene switch. The threshold can be set to the upper and lower 3σ range of the exponentially weighted moving average of the verification feature value sequence. Multiple thresholds can be used for graded alarms (e.g., a warning for minor deviations, and a determination of tampering for severe deviations). If tampering is determined, an alarm can be triggered, a log can be recorded, or the video can be prevented from being used as evidence. This achieves real-time, low-overhead anti-tampering detection for ordinary surveillance cameras without increasing any hardware costs, and has good industrial application value.
[0072] Corresponding to the monitoring data anti-tampering method in the above embodiments, this application also provides a monitoring data anti-tampering system, in which each unit can implement each step of the monitoring data anti-tampering method. Figure 3 The diagram shows a structural block diagram of the monitoring data anti-tampering system provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0073] Reference Figure 3 The monitoring data anti-tampering system includes: The acquisition unit is used to synchronously acquire a thermal noise segment from an imaging sensor for each video frame during the monitoring data recording process; wherein the thermal noise segment is taken from a pixel block with flat texture in the video frame. A measurement unit is used to calculate the image activity measurement value of the video frame and the noise energy measurement value of the thermal noise segment; The association unit is used to perform an arithmetic combination of the image activity metric and the noise energy metric to obtain the frame verification feature value, and to store it in association with the video frame; The detection unit is used to acquire the video frame to be verified and generate a verification feature value based on the video frame to be verified. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the monitoring data has been tampered with.
[0074] It should be noted that the information interaction and execution process between the above systems / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0075] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit module can exist physically separately, or two or more unit modules can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0076] This application also provides an electronic device. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 4 Only one is shown in the image), at least one memory 61 ( Figure 4 (Only one is shown in the image) and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, it causes the electronic device 6 to perform the steps in any of the above-described embodiments of the monitoring data anti-tampering method, or causes the electronic device 6 to perform the functions of each unit in the above-described system embodiments.
[0077] For example, the computer program 62 may be divided into one or more units, which are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 62 in the electronic device 6.
[0078] Electronic device 6 can be a computing device or terminal device such as a desktop computer, laptop, handheld computer, or cloud server. This electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.
[0079] The processor 60 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0080] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may be an external storage device of the electronic device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 6. Furthermore, the memory 61 may include both internal and external storage units of the electronic device 6. The memory 61 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0081] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0082] This application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the steps in any of the above-described method embodiments.
[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions...
[0084] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0085] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0086] In the embodiments provided in this application, it should be understood that the disclosed monitoring data anti-tampering system / electronic device and method can be implemented in other ways. For example, the monitoring data anti-tampering system / electronic device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0087] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0088] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for preventing tampering with monitoring data, characterized in that, include: During the monitoring data recording process, a thermal noise segment from the imaging sensor is synchronously acquired for each video frame; wherein, the thermal noise segment is taken from a pixel block with flat texture in the video frame; Calculate the image activity metric of the video frame and the noise energy metric of the thermal noise segment; The image activity metric and the noise energy metric are arithmetically combined to obtain the frame's verification feature value, which is then associated with and stored with the video frame. The video frame to be verified is obtained, and a verification feature value is generated based on the video frame to be verified. When the deviation between the verification feature value and the stored verification feature value exceeds a preset range, it is determined that the video frame to be verified has been tampered with. The calculation of the image activity metric of the video frame and the noise energy metric of the thermal noise segment includes: Calculate the sum of the Laplacian operator response values of all pixels in the video frame, and use the sum of the absolute values of the Laplacian responses of all pixels as the first intermediate value; Calculate the square of the pixel value of each pixel within the thermal noise segment, and use the sum of the squares of all pixels within the thermal noise segment as the second intermediate value; The cross-correlation value between the Laplacian operator response value and the pixel value of each pixel in the thermal noise segment is calculated. The first intermediate value, the second intermediate value and the cross-correlation value are weighted and combined to obtain the final image activity metric and noise energy metric. The calculation of the cross-correlation value between the Laplacian operator response value and the pixel value of each pixel in the thermal noise segment, and the weighted combination of the first intermediate value, the second intermediate value and the cross-correlation value to obtain the final image activity metric and noise energy metric, includes: Traverse each pixel within the thermal noise segment and calculate the cross-correlation value between the Laplacian operator response value and the pixel value of the thermal noise segment; The image activity metric is obtained by multiplying the first intermediate value by the first weighting coefficient and then summing it with the cross-correlation value multiplied by the second weighting coefficient. The noise energy metric is obtained by multiplying the second intermediate value by the third weighting coefficient and then summing it with the cross-correlation value multiplied by the fourth weighting coefficient. The step of traversing each pixel within the thermal noise segment and calculating the cross-correlation value between the Laplacian operator response value and the pixel value of the thermal noise segment includes: During the traversal, the product of the Laplacian operator response value corresponding to each pixel and the pixel value is accumulated. After the traversal is completed, the sum of the accumulated products is divided by the number of pixels in the thermal noise segment to obtain the cross-correlation value.
2. The monitoring data anti-tampering method as described in claim 1, characterized in that, During the monitoring data recording process, a thermal noise segment from the imaging sensor is synchronously acquired for each video frame, including: The video frame is divided into multiple first candidate blocks. Based on the Laplacian response value of each pixel in the video frame, several second candidate blocks with the best flatness are selected from the multiple first candidate blocks. From all the second candidate blocks, a target block is selected based on a pseudo-random rule. The original pixel values of the target block are extracted as thermal noise segments of the video frame, and the position information of the selected target block is recorded.
3. The monitoring data anti-tampering method as described in claim 2, characterized in that, The step of dividing the video frame into multiple first candidate blocks, and selecting several second candidate blocks with optimal flatness from the multiple first candidate blocks based on the Laplacian response value of each pixel in the video frame, includes: The video frame is divided into K first candidate blocks, and the Laplacian operator response value of each pixel in the video frame is calculated; where K≥16; For each first candidate block, the absolute values of the Laplacian responses of all pixels within the first candidate block are summed to serve as the flatness index of the first candidate block; where the smaller the summed value, the flatter the first candidate block is. The L first candidate blocks with the smallest flatness index are selected as the second candidate blocks; where L≥2.
4. The method for preventing tampering with monitoring data as described in claim 2, characterized in that, The step of selecting a target block from all the second candidate blocks based on a pseudo-random rule, extracting the original pixel values of the target block as a thermal noise segment of the video frame, and recording the location information of the selected target block includes: Using the temporal identifier of the video frame as the seed of the pseudo-random number generator, a target block is selected from L second candidate blocks with uniform probability and pseudo-randomness. Extract the original pixel values of all pixels within the target block to form a thermal noise segment of the video frame, and record the index coordinates of the target block in the video frame.
5. The method for preventing tampering with monitoring data as described in claim 1, characterized in that, The process of acquiring the video frame to be verified, generating a verification feature value based on the video frame to be verified, and determining that the video frame to be verified has been tampered with when the deviation between the verification feature value and the stored verification feature value exceeds a preset range includes: Obtain video frames of the video to be verified, and extract corresponding thermal noise segments from the video frames to be verified; Verification feature values are generated based on the video frames and thermal noise segments of the video to be verified; wherein, the verification feature values are calculated in the same way as the check feature values; The difference between the verification feature value and the stored verification feature value is used as the deviation. When the deviation exceeds a preset threshold, it is determined that the video frame to be verified has been tampered with.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Recorded video authenticity detection system and method
CN121442056A
Realtime media provenance verification system
WO2025081164A1