A Document Image Tampering Detection Method and System Based on Dual-Branch Feature Fusion

Through the method of dual-branch feature fusion, high-frequency local features are extracted by combining the detail enhancement module and the local encoder, and global encoder extracts global features, and fuses through the hybrid attention module. Finally, the tampering area is located through the decoder, which solves the shortcomings in detecting and positioning the document image tampering area in the prior art, and achieves more efficient tampering area positioning and document image authenticity identification.

CN119904879BActive Publication Date: 2025-06-24BEIJING XINYING INNOVATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510377020.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and locate tampered areas in document images, especially in the interaction and fusion of image detail information and global structural information.

Method used

The dual-branch feature fusion method is adopted to extract the high-frequency local features of the image through the detail enhancement module and the local encoder, extract the global features in combination with the global encoder, and fuse it through the hybrid attention module, and finally position the tampering area through the decoder.

Benefits of technology

It significantly improves the accuracy and quality of tampering area positioning, enhances the authenticity identification ability of document images, reduces detailed information loss and avoids interference from irrelevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904879B_ABST
    Figure CN119904879B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting document image tampering based on dual-branch feature fusion, which relates to the technical field of image processing, and includes the steps of: obtaining an image to be detected; constructing a tampering detection and localization model, including a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module; inputting the image to be detected into the tampering detection and localization model, one branch extracts the high-frequency information of the image to be detected through the detail enhancement module and extracts the local features of the high-frequency information through the local encoder, and the other branch extracts the global features of the image to be detected by using the global encoder. Finally, the local features and global features are fused through the hybrid attention module, and the fusion result is decoded by the decoder to locate the tampering area. The document image tampering model proposed by the present invention combines the features of the image domain and the frequency domain, improving the detection accuracy and performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method and system for detecting document image tampering based on dual-branch feature fusion. Background Art

[0002] With the development of the computer industry and the imaging field, media information covering images is almost everywhere in the Internet and the real world. Images have become an important carrier and dissemination channel for information transmission and sharing. In recent years, there have been phenomena where lawbreakers use fake document images to spread rumors, fabricate false news, and illegally obtain economic benefits.

[0003] Most text tampering methods in document images can be roughly divided into three types: (1) splicing, where a region in one image is copied and pasted into another image; (2) copy-move, that is, moving the spatial position of an object within an image; (3) generation, where the original region of an image is replaced with other content. To promote the application of image forensics technology in a wider range of fields, algorithms for detecting document image tampering have gradually emerged. To eliminate the impact of illegal fake images on all parties, it is necessary to research and invent detection algorithms to authenticate the authenticity of document images. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for detecting document image tampering based on dual-branch feature fusion in view of the deficiencies of the above-mentioned existing technologies, so as to solve the problems in the existing technologies.

[0005] The present invention specifically provides the following technical solutions:

[0006] A method for detecting document image tampering based on dual-branch feature fusion includes the following steps:

[0007] Obtain the image to be detected;

[0008] Construct a tampering detection and localization model, where the tampering detection and localization model includes a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module;

[0009] Input the image to be detected into the trained tampering detection and localization model. One branch extracts the high-frequency information of the image to be detected through the detail enhancement module and extracts the local features in the high-frequency information through the local encoder. The other branch extracts the global features of the image to be detected using the global encoder. The local features and global features are fused through the hybrid attention module, and the fusion result is decoded through the decoder, and the tampering area is located through the decoding result.

[0010] Preferably, in the construction of the tampering detection and localization model, the detail enhancement module extracts high-frequency information in the image to be detected through discrete cosine transform coefficient quantization and Bayar convolution operations on the image to be detected, and performs initial enhancement for the local encoder; the local encoder is improved based on the ConvNeXt network, the global encoder adopts the ViT structure, and the hybrid attention module includes a reinforced self-attention module and an adaptive cross-attention module.

[0011] Preferably, the local encoder is improved based on the ConvNeXt network, including:

[0012] Design a detail branch network block based on ConvNeXt-B to extract features from the image. The network block has 4 layers, and the number of channels and the number of blocks are respectively: , ; Every 4 layers of the network is a stage, and there are four stages in total. The output dimensions of each stage are 128, 256, 512, and 1024 respectively.

[0013] Preferably, in the global encoder adopting the ViT structure, the construction process of the global encoder includes:

[0014] Represent the input image as , and represent the true label as , where h and w correspond to the height and width of the image respectively;

[0015] Pad the height and width of the image to and , and use as a constant, and pass to the windowed ViT encoder with 12 layers to obtain the global encoder; one complete global attention block is retained every 3 layers; specifically expressed as:

[0016] ;

[0017] Among them, represents the ViT encoder, represents the encoded feature map.

[0018] Preferably, the specific expression of the hybrid attention module is:

[0019] ;

[0020] ;

[0021] Among them, is the dual-branch main fusion feature, is the final fusion feature, is an Enhanced self-attention block, is an Adaptive cross-attention block, is a Decoder blending module.

[0022] Preferably, before inputting the image to be detected into the trained tampering detection and localization model, it further includes:

[0023] Training the tampering detection and localization model through backpropagation of the loss function to obtain the trained tampering detection and localization model.

[0024] The present invention provides a document image tampering detection system based on dual-branch feature fusion, including:

[0025] An acquisition module for obtaining the image to be detected;

[0026] A model construction module for the tampering detection and localization model, which includes a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module;

[0027] A detection module for inputting the image to be detected into the trained tampering detection and localization model. One branch extracts the high-frequency information of the image to be detected through the detail enhancement module and extracts the local features in the high-frequency information through the local encoder, while the other branch uses the global encoder to extract the global features of the image to be detected. The local features and global features are fused through the hybrid attention module, and the fusion result is decoded through the decoder, and the tampering area is located through the decoding result.

[0028] The present invention provides a computer device, including a memory and a processor. When the program stored in the memory is executed by the processor, the processor executes the steps of the above-mentioned document image tampering detection method based on dual-branch feature fusion.

[0029] The present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned document image tampering detection method based on dual-branch feature fusion are implemented.

[0030] Compared with the prior art, the present invention has the following remarkable advantages:

[0031] The present invention designs a tampering detection and localization model, which uses two branches to extract the local and global parts of the image to be detected respectively. One branch extracts the high-frequency information of the image to be detected through a detail enhancement module and extracts the local features in the high-frequency information through a local encoder, which focuses on the detailed information of the image. The other branch uses a global encoder to extract the global features of the image to be detected for capturing the overall structural information, so as to achieve a comprehensive analysis of the overall information of the image and the details of the text area, promote the interactive fusion of feature information at different levels, effectively avoid the interference of irrelevant information while reducing the loss of detailed information, significantly enhance the representation ability of the key text content area, improve the quality of the tampering area localization map, and facilitate the authenticity identification of document images. Description of the Drawings

[0032] Figure 1 is the overall step flowchart of a document image tampering detection method based on dual-branch feature fusion provided by the present invention;

[0033] Figure 2 is the step flowchart of making and obtaining the document tampering image to be detected provided by the present invention;

[0034] Figure 3 is the overall network model framework diagram provided by the present invention;

[0035] Figure 4 is the structural schematic diagram of the enhanced self-attention module provided by the present invention;

[0036] Figure 5 is the structural schematic diagram of the adaptive cross-attention module provided by the present invention;

[0037] Figure 6 is the tampering area detection effect diagram of the embodiment of the present invention on document images of different tampering types; among them, Figure 6 (a1) of Figure 6 (a2) of Figure 6 (a3) of Figure 6 (b1) of Figure 6 (b2) of Figure 6 (b3) of Figure 6 (c1) of Figure 6 (c2) of Figure 6 (c3) of Specific Embodiments

[0038] The following will clearly and completely describe the technical solutions of the embodiments of the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0039] Please refer to Figure 1 , which shows a flowchart of the steps of a document image tampering detection method based on dual-branch feature fusion provided by an embodiment of the present invention. The method includes the following steps:

[0040] Step S1: Obtain the image to be detected.

[0041] Figure 2 As shown, the embodiments of the present invention give the work process of making and obtaining the dataset of the present invention.

[0042] Obtain a document image dataset. The source dataset covers invoices, magazines, chat records, PDF images, and medical prescriptions; use existing text recognition technology to extract text and mark positions, tamper with similar samples for high-confidence text matching, and randomly augment low-confidence images. The tampered samples also undergo noise addition and style adjustment operations.

[0043] First, use existing text recognition technology to extract text information in the image, retrieve and locate the text content in the image, and record the marker information of the text position. Set a threshold for text detection. If the confidence threshold is exceeded, perform similarity matching based on color difference and recognized content, and select (similar ones for subsequent) tampering operations / methods.

[0044] For images with the confidence of the recognized text lower than the preset threshold, random image processing operations are taken to augment the dataset (random tampering method). These operations include random jitter, cropping, compression, etc., aiming to increase the diversity of data through randomness. For images exceeding the threshold, pairwise production is carried out, that is, pairwise text tampering samples are made and saved. To further increase the difficulty of tampering detection and the robustness of the model, noise is added and the image style is adjusted in these samples (post-processing). These operations can effectively improve the generalization ability of the model, making it perform more stably when dealing with different scenarios or changing image features.

[0045] Step S2: Construct a tampering detection and localization model. The tampering detection and localization model includes a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module.

[0046] The detail enhancement module extracts high-frequency information from the image to be detected through the operations of discrete cosine transform coefficient quantization (mapping) and Bayar convolution on the image to be detected, and then adds a convolution operation to expand the detail information to perform the initial enhancement for the local encoder.

[0047] The local encoder is improved based on the ConvNeXt network, the global encoder adopts the ViT (Vision Transformer) structure, and the hybrid attention module consists of a reinforced self-attention module and an adaptive cross-attention module.

[0048] The local encoder is improved based on the ConvNeXt network, including:

[0049] Design a detail branch network block based on ConvNeXt-B to extract features from the image. The network block has 4 layers, and the number of channels and the number of blocks in each layer are respectively: , ; Every 4 layers of the network form a stage, and there are four stages in total. The output dimensions of each stage are 128, 256, 512, and 1024 respectively.

[0050] The global encoder adopts the ViT (Vision Transformer) structure, including:

[0051] Represent the input image as , represent the true label as , where h and w correspond to the height and width of the image respectively; Perform positional embedding on the image, and pad the height and width of the image to and , and use as a constant, and pass to the windowed ViT encoder with 12 layers to obtain the global encoder; One complete global attention block is reserved every 3 layers; Specifically expressed as:

[0052] ;

[0053] Among them, represents the ViT encoder, represents the encoded feature map.

[0054] The hybrid attention module consists of a reinforced self-attention module and an adaptive cross-attention module, including:

[0055] The specific expression of the hybrid attention module is:

[0056] ;

[0057] ;

[0058] Among them, is a double-branch main fusion feature, is the final fusion feature, is an enhanced self-attention block, is an adaptive cross-attention block, is a decoder blending module.

[0059] Step S3: Input the image to be detected into the trained tampering detection and localization model. One branch extracts the high-frequency information of the image to be detected through the detail enhancement module and extracts the local features of the high-frequency information through the local encoder. The other branch uses the global encoder to extract the global features of the image to be detected. Finally, the local features and global features are fused through the hybrid attention module, and the fusion result is decoded through the decoder, and the tampered area is located through the decoding result.

[0060] As Figure 3 shown, the local encoder focuses on extracting information at the image detail level, such as high-resolution features like texture and edges; the global encoder focuses on the overall structure and semantic information. The fused features are input into the enhanced self-attention module (as Figure 4 shown), further optimizing the feature representation. In the enhanced self-attention module, in order to balance computational efficiency and model performance, a pair of asymmetric complementary convolutions (with convolution kernels of 1×3 and 3×1 respectively) are specially designed. By decomposing the two-dimensional convolution operation, it can efficiently extract the directional clues and texture features in the image. This design of asymmetric convolution not only retains the receptive field of the convolutional layer but also effectively reduces the computational complexity, while improving the model's ability to capture subtle forgery traces.

[0061] In addition, the enhanced self-attention module focuses on the edges and detail features of the forged area. To further improve the model's attention ability to the target area, an adaptive cross-attention module is added on this basis as another branch of the hybrid attention mechanism (as Figure 5 shown). This mechanism enables the model to balance between accuracy and efficiency through adaptive weight allocation of local and global information, thereby more accurately capturing and locating the forged area and significantly improving the network performance. Several implementation diagrams in the specific implementation are as Figure 6 shown.

[0062] Before inputting the image to be detected into the tampering detection and localization model, it also includes:

[0063] The tampering detection and localization model is trained by backpropagation through the loss function to obtain a trained tampering detection and localization model. The result of classifying the final fused features after upsampling by the decoder is compared with the ground truth label using binary cross-entropy and Dice constraints.

[0064] The fused multi-level features are used as the input to the decoder for the prediction mask. The decoder consists of three decoder blocks, each containing an upsampling operation and multiple convolutional layers to gradually restore the spatial resolution and extract higher-level semantic information. The size of the feature map is gradually restored to the original image size through the upsampling operation, and the features are further refined by the convolutional layers in each decoder block. In the final decoder output, the classification result, i.e., the prediction mask, is obtained.

[0065] Specifically, the prediction mask is compared with the ground truth label, and the classification result is constrained by calculating the binary cross-entropy loss and Dice loss. Among them, the binary cross-entropy loss is used to measure the pixel-level difference between the classification result and the ground truth label, while the Dice loss is used to optimize the shape similarity and overlap rate between the target region and the ground truth label. Finally, by minimizing the loss function, the model parameters are continuously updated using backpropagation, thereby improving the accuracy of the model for image tampering detection and localization. To improve the detection effect, the present invention designs a decoder to decode the dual-branch fused features, and adopts a multi-stage loss supervision strategy to optimize the detection results of the tampered regions layer by layer, significantly improving the quality of the tampered region localization map.

[0066] Based on the above method, the present invention provides a document image tampering detection system based on dual-branch feature fusion, including: an acquisition module, a model construction module, and a detection module.

[0067] Among them, the acquisition module is used to obtain the image to be detected; the model construction module is used for the tampering detection and localization model, and the tampering detection and localization model includes a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module; the detection module is used to input the image to be detected into the trained tampering detection and localization model. One branch extracts the high-frequency information of the image to be detected through the detail enhancement module and extracts the local features in the high-frequency information through the local encoder, and the other branch uses the global encoder to extract the global features of the image to be detected. The local features and global features are fused through the hybrid attention module, and the fused result is decoded by the decoder to locate the tampered region through the decoding result.

[0068] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of a document image tampering detection method based on dual-branch feature fusion.

[0069] According to the disclosed embodiments, a computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.

[0070] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a document image tampering detection method based on dual-branch feature fusion are implemented.

[0071] According to the disclosed embodiments, the storage medium can be a non-volatile computer-readable storage medium, for example, it can include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0072] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A document image tampering detection method based on dual-branch feature fusion, characterized in that: include: Acquire the image to be detected; Constructing a tampering detection and localization model, wherein the tampering detection and localization model includes a detail enhancement module, a local encoder, a global encoder, and a hybrid attention module; The image to be detected is input into a trained tampering detection and positioning model, one branch extracts high-frequency information of the image to be detected through a detail enhancement module, and extracts local features in the high-frequency information through a local encoder, and the other branch uses a global encoder to extract global features of the image to be detected, and fuses the local features and the global features through a hybrid attention module, decodes the fusion result through a decoder, and locates the tampered area through the decoding result; In the tampering detection and positioning model, the detail enhancement module extracts high-frequency information in the image to be detected by performing discrete cosine transform coefficient quantization and Bayar convolution operations on the image to be detected, and performs the initial enhancement for the local encoder; the local encoder is improved based on the ConvNeXt network, the global encoder adopts the ViT structure, and the hybrid attention module includes an enhanced self-attention module and an adaptive cross-attention module.

2. A document image tampering detection method based on dual-branch feature fusion as claimed in claim 1, characterized in that: The local encoder is improved based on the ConvNeXt network, including: Based on ConvNeXt-B, a detailed branch network block is designed to extract features from the image. The network block has 4 layers. The number of channels C and blocks B in each layer are C = (128, 256, 512, 1024) and B = (3, 3, 27, 3). Every 4 layers of the network are a stage, with a total of four stages. The dimensions of the final output of each stage are 128, 256, 512 and 1024 respectively.

3. The document image tampering detection method based on dual-branch feature fusion as claimed in claim 1, characterized in that: The global encoder adopts the ViT structure, and the global encoder construction process includes: Denote the input image as X∈R 1×h×w , denote the true label as M∈R 1×h×w , where h and w correspond to the height and width of the image respectively; Fill the image's height and width to X P ∈R 3×H×W and M P ∈R 1×H×W , and set H = W = 1024 as a constant, and set X P Pass it to the 12-layer windowed ViT encoder to obtain a global encoder, where every 3 layers retain a complete global attention block; specifically expressed as: Among them, V represents the ViT encoder, G e Represents the encoded feature map.

4. The document image tampering detection method based on dual-branch feature fusion as claimed in claim 1, characterized in that: The hybrid attention module F h The specific expression is: F h =ESA(F f )+ACA(F f ); F o =DBM(F h ); Among them, F f is the main fusion feature of the two branches, F o is the final fusion feature, ESA is the enhanced self-attention block, ACA is the adaptive cross attention block, and DBM is the decoding fusion module.

5. The document image tampering detection method based on dual-branch feature fusion as claimed in claim 1, characterized in that: Before inputting the image to be detected into the trained tampering detection and positioning model, the method further includes: The tampering detection and positioning model is trained through back propagation of the loss function to obtain a trained tampering detection and positioning model.

6. A document image tampering detection system based on dual-branch feature fusion, characterized in that: include: An acquisition module, used for acquiring an image to be detected; A model building module, used to build a tampering detection and positioning model, wherein the tampering detection and positioning model includes a detail enhancement module, a local encoder, a global encoder and a hybrid attention module; A detection module is used to input the image to be detected into a trained tampering detection and positioning model, one branch extracts high-frequency information of the image to be detected through a detail enhancement module, and extracts local features in the high-frequency information through a local encoder, and the other branch uses a global encoder to extract global features of the image to be detected, and fuses the local features and the global features through a hybrid attention module, decodes the fusion result through a decoder, and locates the tampered area through the decoding result; In the tampering detection and positioning model, the detail enhancement module extracts high-frequency information in the image to be detected by performing discrete cosine transform coefficient quantization and Bayar convolution operations on the image to be detected, and performs the initial enhancement for the local encoder; the local encoder is improved based on the ConvNeXt network, the global encoder adopts the ViT structure, and the hybrid attention module includes an enhanced self-attention module and an adaptive cross-attention module.

7. A computer device, characterized in that: It includes a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a document image tampering detection method based on dual-branch feature fusion as described in any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a document image tampering detection method based on dual-branch feature fusion as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Document image tampering detection and classification method based on double-domain and multi-scale network

    CN117314714A

  • Text image tampering detection method and device

    CN119027788A