Ticket forgery detection method and system fused with wavelet frequency domain DCT, and storage medium
By integrating wavelet frequency domain enhancement and DCT-guided attention detection methods, the problem of tampering detection in high-text-density document images is solved, achieving high-precision and stable tampering detection that is adaptable to complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TOBACCO ZHEJIANG IND CO LTD
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack sensitivity to text details, have poor environmental robustness, and fail to fully utilize frequency domain features when dealing with document images with high text density and regular structure, resulting in low tamper detection accuracy and weak generalization ability.
A detection method that combines wavelet frequency domain enhancement and DCT-guided attention is adopted. By acquiring multi-scale spatial and frequency domain features, feature fusion and wavelet pattern feature enhancement are performed to achieve pixel-level classification.
It improves the accuracy and stability of document image tampering detection, maintains stable feature extraction capabilities in complex environments, and enhances the accuracy and adaptability of detection.
Smart Images

Figure CN121884086A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ticket forgery detection, and more specifically to a ticket forgery detection method that integrates wavelet frequency domain enhancement and DCT-guided attention. Background Technology
[0002] In recent years, with the popularization of digital image processing technology and the acceleration of the digitalization of certificates, the forgery and tampering of images of various licenses, business licenses, and other certificates have become increasingly prominent, seriously disrupting market order and threatening public safety. Taking the tobacco monopoly certificate as an example, its high text density, structured layout, and key information areas (such as validity period and seal) make it difficult for traditional tampering detection methods based on natural scene images to effectively address the issue. Therefore, developing specialized anti-counterfeiting detection technologies for certificate images is of crucial theoretical and practical significance for improving the accuracy and efficiency of certificate verification.
[0003] Existing image tampering detection methods mainly rely on deep learning-based semantic segmentation models, such as encoder-decoder structures like U-Net and DeepLab, which extract RGB spatial features through convolutional neural networks to locate tampered regions. While some methods attempt to introduce frequency domain information (such as DCT and wavelet transform) to capture compression artifacts and texture anomalies, they still have significant limitations in practical applications: First, there is a lack of pixel-level labeled data for ID card images, resulting in insufficient model generalization ability; second, they are not sensitive to subtle tampering in text regions (such as font and boundary modifications), leading to frequent missed and false detections; third, they are difficult to adapt to complex shooting environments such as low light and reflections, exhibiting poor robustness; and fourth, the modeling of frequency domain artifacts is insufficient, failing to effectively integrate spatial and frequency domain cues.
[0004] To address the aforementioned shortcomings, there is an urgent need to construct a document image tampering detection scheme that integrates multimodal features and attention mechanisms. This scheme would enhance the sensitivity to subtle tampering in text regions, the adaptability to complex environments, and the ability to detect frequency domain artifacts, thus providing reliable technical support for intelligent verification of document authenticity. Summary of the Invention
[0005] The purpose of this invention is to provide a ticket forgery detection method, system, and storage medium that integrates wavelet frequency domain enhancement and DCT-guided attention, in order to solve the technical problems of low tampering detection accuracy and weak generalization ability in the face of specific document images with high text density and regular structure, due to lack of dedicated data, insensitivity to text details, poor environmental robustness, and insufficient utilization of frequency domain features.
[0006] To achieve the above objectives, this invention provides a ticket forgery detection method that integrates wavelet frequency domain and DCT, characterized in that the detection method includes: Acquire the image to be detected; RGB features are extracted from the image to be detected to obtain multi-scale spatial domain features; DCT feature extraction is performed on the image to be detected to obtain multi-scale frequency domain features; The multi-scale spatial domain features and multi-scale frequency domain features are fused to obtain a multi-scale fused feature map. The multi-scale fused feature map is decomposed in the frequency domain using a wavelet pattern feature enhancement module to obtain global fused features; The global fusion features are classified at the pixel level to obtain detection results.
[0007] Optionally, RGB feature extraction is performed on the image to be detected to obtain multi-scale spatial domain features, including: The first branch feature map and the second branch feature map are obtained based on the image to be detected; According to formula (1), feature enhancement is performed on the first branch feature map and the second branch feature map to obtain the first branch enhanced feature map and the second branch enhanced feature map: (1) in, For the first The output feature map of the s-th branch at the 1st level. For hierarchical indexes, For the convolution operation corresponding to the s-th branch, For upsampling or downsampling, Indexes to all branches except the current target branch s; The third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; The third branch feature map is enhanced using formula (1) to obtain the third branch enhanced feature map; The fourth branch feature map is obtained based on the first branch enhanced feature map, the second branch enhanced feature map, and the third branch enhanced feature map; The fourth branch feature map is enhanced using formula (1) to obtain the fourth branch enhanced feature map.
[0008] Optionally, RGB feature extraction is performed on the image to be detected to obtain multi-scale spatial domain features, including: Multi-scale spatial domain features are obtained according to formulas (2) to (5): (2) (3) (4) (5) in, For horizontal edge response, For vertical edge response, These are the pixel values of the normalized image. For total edge features, This is the final RGB feature map after edge enhancement of the i-th branch. The final output feature map obtained by fusing features from the i-th branch is... This represents the corresponding number of branches.
[0009] Optionally, DCT feature extraction is performed on the image to be detected to obtain multi-scale frequency domain features, including: The frequency coefficients of each pixel block in the image to be detected are obtained according to formulas (6) to (7): (6) (7) in, These are the transformed DCT coefficients. For horizontal frequency indexing, For vertical frequency indexing, For horizontal basis functions, For vertical basis functions, These are the normalization coefficients; DCT feature maps are obtained based on the frequency coefficients; The DCT feature map is convolved to obtain the first branch feature map and the second branch feature map. Feature enhancement is performed on the first branch feature map and the second branch feature map to obtain a first branch enhanced feature map and a second branch enhanced feature map; The third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; The third branch feature map is enhanced to obtain a third branch enhanced feature map; Multi-scale frequency domain features are obtained according to formulas (8) to (9): (8) (9) in, The feature map output by the convolutional layer. Convolutional layers specifically designed for processing DCT features. The input feature map is the residual connection. This is the final output feature map after residual connection.
[0010] Optionally, the multi-scale spatial domain features and multi-scale frequency domain features are fused to obtain a multi-scale fused feature map, including: Features with the same spatial resolution in the RGB and DCT branches are paired to form feature map groups of different scales; According to formulas (10) to (11), feature maps of different scales of channel enhancement are obtained: (10) (11) in, These are the feature maps at different scales. For global average pooling, For the weights of the first fully connected layer, It is the ReLU activation function. For the weights of the second fully connected layer, It is the Sigmoid activation function. This is a channel attention weight map. For channel-by-channel multiplication, Feature map set after channel attention weighting; Based on formulas (12) to (13), obtain feature maps of different scales for spatial augmentation: (12) (13) in, Average pooling is performed along the channel dimension. Max pooling for the channel dimension, This is a spatial attention weight map. For the final attention-weighted feature map set; The final attention-weighted feature map group is spliced and cross-modal information fusion is performed to obtain a multi-scale fused feature map.
[0011] Optionally, the multi-scale fused feature map is decomposed in the frequency domain using a wavelet pattern feature enhancement module to obtain global fused features, including: The multi-scale fused feature map is decomposed into wavelet-style frequency domain according to formulas (14) to (15): (14) (15) in, The multi-scale fused feature map, for Average pooling, Low-frequency characteristics after downsampling The low-frequency features are recovered by upsampling. It is a high-frequency residual characteristic; Enhanced frequency domain features are obtained based on the low-frequency features and high-frequency residual features.
[0012] Optionally, the global fusion features are classified at the pixel level to obtain detection results, including: Pixel-level classification probabilities are obtained using formulas (16) and (17): (16) (17) in, For global fusion features, for Convolution class header, The original classification score. For the target category parameter, To sum the index parameters, The score for category c, The score for category k, Let be the probability that a pixel belongs to category c.
[0013] Optionally, the detection method further includes performance evaluation of the model, including: Obtain the intersection-union ratio according to formula (18): (18) in, For intersection, union, and comparison, For a real example, As a false positive example, This is a false negative example; The accuracy is obtained using formula (19): (19) in, For accuracy; The recall rate is obtained according to formula (20); (20) in, Recall rate; Obtain the F1 score according to formula (21): ,(twenty one) in, The score is the F1 score. The area under the curve can be obtained using formulas (22) to (24): ,(twenty two) ,(twenty three) (twenty four) in, For a true negative example, For the true rate, For false positives, The area under the curve is denoted as .
[0014] On the other hand, the present invention also provides a ticket forgery detection method that integrates wavelet frequency domain enhancement and DCT guided attention, the detection system including a processor configured to perform the method as described above.
[0015] In another aspect, the present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores instructions that, when executed by a processor, implement any of the methods described above.
[0016] The beneficial effects of this invention are: This invention achieves efficient, accurate, and adaptive detection of tampered regions in document images by fusing multimodal features and an attention mechanism. Specifically, it includes: This invention possesses the ability to simultaneously extract spatial and frequency domain features using a two-stream network, and also achieves multi-scale feature fusion based on a cross-modal attention mechanism. By combining spatial feature perception with frequency domain anomaly detection, it realizes the functions of detail preservation and artifact recognition in document image tampering detection.
[0017] This invention enhances model robustness and improves detection stability in complex scenes by using a wavelet pattern feature enhancement module. Furthermore, the method of this invention outperforms traditional tamper detection models in both detection accuracy and generalization ability. Traditional detection methods are insensitive to text regions and subtle tampering in document images, and their performance significantly degrades under complex shooting conditions such as low light and reflections. This invention, through multi-scale frequency domain decomposition and feature reconstruction, enables the model to maintain stable feature extraction capabilities under different imaging conditions, greatly enhancing its practical application value.
[0018] Because the present invention completes the deep fusion and optimization of multimodal features during the training phase, the computational overhead of the dual-stream network and attention mechanism is controllable and does not require additional preprocessing modules during actual deployment. At the same time, it supports incremental learning of new tampering methods, can continuously adapt to the development and changes of document forgery technology, and can gradually improve the accuracy, adaptability and reliability of document anti-counterfeiting detection system. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart of a ticket forgery detection method that integrates wavelet frequency domain enhancement and DCT-guided attention according to an embodiment of the present invention; Figure 2 This is a flowchart of RGB feature extraction according to one embodiment of the present invention; Figure 3 This is a flowchart of DCT feature extraction according to an embodiment of the present invention; Figure 4 This is a pixel-level detection map of the model according to one embodiment of the present invention; Figure 5 This is a comparison chart of test results for the model according to one embodiment of the present invention. Detailed Implementation
[0020] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0021] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0022] This embodiment uses the datasets from the Ali Tianchi Real Image Tampering Detection Competition and the third phase of the Ali Tianchi Competition as references to further illustrate the method of the present invention.
[0023] like Figure 1 The diagram shows a flowchart of a ticket forgery detection method fusing wavelet frequency domain enhancement and DCT-guided attention according to an embodiment of the present invention. Figure 1 The detection method may include the following steps: In step S10, the image to be detected is acquired; In step S11, RGB features are extracted from the image to be detected to obtain multi-scale spatial domain features; In step S12, DCT feature extraction is performed on the image to be detected to obtain multi-scale frequency domain features; In step S13, the multi-scale spatial domain features and multi-scale frequency domain features are fused to obtain a multi-scale fused feature map. In step S14, the multi-scale fused feature map is decomposed in the frequency domain through the wavelet pattern feature enhancement module to obtain global fused features; In step S15, the global fusion features are classified at the pixel level to obtain the detection results.
[0024] In such Figure 1 In the method shown, step S10 is used to acquire the image to be detected. Specifically, in this embodiment, the original images may be screened to remove images whose subjects are simulated objects or non-realistic scenes, images with overly bright background colors, images that may introduce model interference, and non-realistic images containing a large amount of drawing or illustration content. This targeted data cleaning effectively ensures the authenticity and purity of the training dataset, establishing a reliable data foundation for subsequent feature extraction and tamper detection. Furthermore, this screening process can adjust the screening criteria according to actual needs, further enhancing the model's adaptability in specific application scenarios.
[0025] In the construction of the filtered model, each module follows standard convolution calculation specifications, where the convolutional layers employ... Convolution kernel, classification head is Normalization is performed using Batch Normalization (BN), and the activation function is ReLU.
[0026] The calculation process of BN is shown in formulas (1) to (2): (1) (2) in, The original input value, This is the standardized value after normalizing the original output value. This represents the mean of the mini-batch. The variance of the mini-batch. It is a constant, usually 1. To effectively prevent mathematical errors such as dividing by zero, These are learnable parameters, i.e., scaling factors. The learnable parameter is the offset. This is the final output of the BN layer, i.e., the value passed to subsequent layers (such as ReLU).
[0027] The ReLU activation function is defined as shown in formula (3): (3) in, The output value of the ReLU function is the final activation value of the neuron, which represents the neuron's response to the input. The response intensity For threshold comparison, take the larger function, if ,but ,if ,but .
[0028] Step S11 is used to extract RGB features from the image to be detected to obtain multi-scale spatial domain features. In this embodiment, the specific method for obtaining multi-scale spatial domain features in step S11 can be of various forms known to those skilled in the art. In one example of the present invention, step S11 may include, for example... Figure 2 The steps shown are described. Figure 2 In this context, step S11 may include: In step S20, the first branch feature map and the second branch feature map are obtained based on the image to be detected; In step S21, feature enhancement is performed on the first branch feature map and the second branch feature map to obtain the first branch enhanced feature map and the second branch enhanced feature map; In step S22, the third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; In step S23, feature enhancement is performed on the third branch feature map to obtain the third branch enhanced feature map; In step S24, a fourth branch feature map is obtained based on the first branch enhanced feature map, the second branch enhanced feature map, and the third branch enhanced feature map; In step S25, feature enhancement is performed on the fourth branch feature map to obtain the fourth branch enhanced feature map; In step S26, multi-scale spatial domain features are obtained based on the first branch enhanced feature map, the second branch enhanced feature map, the third branch enhanced feature map, and the fourth branch enhanced feature map.
[0029] In such Figure 2 The method shown may involve normalizing the image to be detected. Specifically, in this embodiment, the original RGB image may be normalized according to formula (4): (4) in, To normalize the image at position At that location, the pixel value of the c-th color channel. For the original image at position At that location, the pixel value of the c-th color channel. Let be the mean of the c-th color channel. Let be the standard deviation of the c-th color channel.
[0030] Next, the normalized image to be detected is downsampled to obtain a preliminary convolutional feature map. Specifically, in this embodiment, downsampling and feature extraction can be performed on the normalized input image. This stage sequentially passes through four BasicBlock residual modules, each containing convolution, BN, ReLU, and residual connections. Although a single module outputs 64 channels, channel dilation at the end of the stage results in a final output of 256 channels. The final preliminary convolutional feature map is obtained with dimensions [batch_size, 256, 128, 128].
[0031] Step S20 is used to obtain a first branch feature map and a second branch feature map by performing a branch splitting operation on the initial convolutional feature map. Specifically, in this embodiment, the initial convolutional feature map may be split through a branch splitting operation to generate two new feature maps with different resolutions: a high-resolution branch, i.e., the first branch, through a stride of 1. The input is processed by convolution to obtain a 48-channel, size-invariant feature map [batch_size,48,128,128]; the medium-resolution branch, i.e., the second branch, is processed by a step size of 2. Convolution is downsampled to obtain a 96-channel feature map with half the size [batch_size,96,64,64].
[0032] Step S21 is used to perform feature enhancement on the first branch feature map and the second branch feature map to obtain a first branch enhanced feature map and a second branch enhanced feature map. Specifically, in this embodiment, the first branch enhanced feature map and the second branch enhanced feature map can be obtained according to formula (5): (5) in, For the first The output feature map of the s-th branch at the 1st level. For hierarchical indexes, For the convolution operation corresponding to the s-th branch, For upsampling or downsampling, This provides the indices for all branches except the current target branch s. The first branch receives upsampled features from the second branch, and the second branch receives downsampled features from the first branch. This process ensures that each branch retains its own features while incorporating complementary information from another scale, resulting in a more discriminative enhanced feature map. This keeps the output dimensions of the first and second branches unchanged, but their content is enriched by cross-branch interactions.
[0033] Steps S22 to S25 are used to progressively construct and enhance the third and fourth branches, forming a multi-resolution feature map set for information fusion. Specifically, in this embodiment, the network can process the two branches in parallel. First, a third branch is added by downsampling the second branch. Second, while the three branches are extracting features respectively, they continuously perform cross-resolution feature exchange and fusion according to formula (5). Finally, the third branch is further downsampled to add a fourth branch, and finally, a parallel feature processing architecture containing four different resolutions is constructed according to formula (5), that is, an enhanced four-branch feature map set.
[0034] Step S26 is used to perform edge enhancement on the normalized RGB image and the enhanced four-branch feature map set to obtain the final RGB feature map. Specifically, in this embodiment, the final edge-enhanced RGB feature map can be obtained according to formulas (6) to (9): (6) (7) (8) (9) in, For horizontal edge response, For vertical edge response, For total edge features, The final RGB feature map for edge enhancement. To fuse all layers, the final output feature map is obtained by fusing the features from each branch. In the final parallel processing stage of the network, the four branches perform deep feature extraction on feature maps of different resolutions. Specifically, the first branch processes 48 channels... Feature map of dimensions; second branch processes 96 channels, Feature map of dimensions; third branch processes 192 channels. Feature map of dimensions; fourth branch processes 384 channels, Feature maps of different sizes. Each branch further enhances the feature representation by stacking BasicBlock modules, and finally outputs four feature maps of different scales in a synchronized manner, with dimensions of [batch_size,48,128,128], [batch_size,96,64,64], [batch_size,192,32,32] and [batch_size,384,16,16], respectively.
[0035] Step S12 is used to perform DCT feature extraction on the image to be detected to obtain multi-scale frequency domain features. In this embodiment, the specific method for obtaining multi-scale spatial domain features in step S12 can be of various forms known to those skilled in the art. In one example of the present invention, step S12 may include, for example... Figure 3 The steps shown are described. Figure 3 In this context, step S12 may include: In step S30, the image to be detected is divided into blocks to obtain the frequency coefficients of each pixel block; In step S31, the DCT feature map is obtained based on the frequency coefficients; In step S32, the DCT feature map is processed by convolution to obtain the first branch feature map and the second branch feature map; In step S33, feature enhancement is performed on the first branch feature map and the second branch feature map to obtain the first branch enhanced feature map and the second branch enhanced feature map; In step S34, the third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; In step S35, feature enhancement is performed on the third branch feature map to obtain the third branch enhanced feature map; In step S36, multi-scale frequency domain features are obtained based on the first branch enhanced feature map, the second branch enhanced feature map, and the third branch enhanced feature map.
[0036] In such Figure 3 In the method shown, step S30 is used to perform image block processing on the image to be detected to obtain the frequency coefficient of each pixel block. Specifically, in this embodiment, the block-sized data can be obtained according to formulas (10) to (11). Pixel block set: (10) (11) in, For the luminance channel image, This represents the pixel value of the red channel. This represents the pixel value of the green channel. This represents the pixel value of the blue channel. For the coordinates place Pixel blocks, For block row indexes, For block column indexes, Let be the range of row coordinates of the i-th block in the image. Let be the range of column coordinates of the j-th block in the image.
[0037] Secondly, each The pixel block is obtained according to formulas (12) to (13). Frequency coefficient matrix: (12) (13) in, These are the transformed DCT coefficients, i.e., the coefficients of each pixel block. Frequency coefficient, For horizontal frequency indexing, For vertical frequency indexing, For horizontal basis functions, For vertical basis functions, This is the normalization coefficient.
[0038] Step S31 is used to obtain the DCT feature map based on the frequency coefficients. Specifically, in this embodiment, it may be based on... The frequency coefficient matrix is used to obtain the three-dimensional binary volume according to formula (14): (14) in, The position in the k-th binary channel The value, Let k be the k-th binarization threshold, where k is the threshold channel index.
[0039] Secondly, the input 3D binary volumetric representation of DCT features (assuming an initial channel count of 512) is enhanced using dilated convolution to capture multi-scale contextual information. Secondly, the pre-processed DCT feature map is then processed through... Convolution is used to reduce the dimensionality of features and obtain the JPEG block structure features through rearrangement mapping according to formula (15): (15) in, It is a block structure feature. For block row indexes, For block column indexes, For channel indexing, This is the dimensionality-reduced DCT feature map. The vertical coordinates of the dimensionality-reduced DCT feature map are shown below. The horizontal coordinates of the dimensionality-reduced DCT feature map. This is the line offset within the block. This represents the column offset within the block. Finally, deep feature extraction is performed using four stacked BasicBlock modules. Each BasicBlock module contains convolution, batch normalization, and a ReLU activation function, using a... The convolutional layer performs channel compression, reducing the number of channels in the feature map from 512 to 96, and finally outputs the DCT feature map extracted from the deep layer, with specific dimensions of [batch_size, 96, 64, 64].
[0040] Step S32 is used to obtain a first branch feature map and a second branch feature map by performing a convolution operation on the DCT feature map. Specifically, in this embodiment, it can be a branch expansion operation performed on the 96-channel DCT feature map after deep extraction. This is done by performing a convolution operation with a stride of 2. Convolution is used for downsampling, increasing the number of channels to 192 and compressing the spatial size to... After BN and ReLU activation, a second branch feature map is formed, while the original 96-channel feature is retained as the first branch, resulting in two branches: [batch_size,96,64,64] and [batch_size,192,32,32].
[0041] Step S33 is used to perform feature enhancement on the first branch feature map and the second branch feature map to obtain a first branch enhanced feature map and a second branch enhanced feature map. Specifically, in this embodiment, it can be the first branch (96 channels, ) and the second branch (channel 192, Deep features are extracted using stacked BasicBlocks. During this process, cross-resolution feature fusion is continuously performed between the two branches: the features of the second branch are upsampled and added to the features of the first branch, while the features of the first branch are downsampled and fused with the features of the second branch. According to the enhancement processing of formula (5), two enhanced feature maps with different resolutions are finally output, with their dimensions maintained as [batch_size,96,64,64] and [batch_size,192,32,32].
[0042] Steps S34 and S35 are used to create a third branch feature map through downsampling and enhance it through cross-resolution feature fusion to obtain a third branch enhanced feature map. Specifically, in this embodiment, it can be based on the enhanced feature map of the second branch with 192 channels, enhanced by a step size of 2. Convolutional downsampling expands the number of channels to 384, compressing the spatial size to... The third branch feature map [batch_size,384,16,16] is generated. Next, while each of the three branches extracts features through BasicBlock stacking, cross-resolution feature fusion is performed according to formula (5). That is, the third branch receives downsampled features from the first two branches, and its features are also upsampled to the first two branches for fusion. Finally, three enhanced feature maps are output, with dimensions of [batch_size,96,64,64], [batch_size,192,32,32], and [batch_size,384,16,16], respectively.
[0043] Step S36 is used to obtain multi-scale frequency domain features based on the first branch enhanced feature map, the second branch enhanced feature map, and the third branch enhanced feature map. Specifically, in this embodiment, the multi-scale frequency domain features can be obtained according to formulas (16) to (17): (16) (17) in, The feature map output by the convolutional layer. Convolutional layers specifically designed for processing DCT features. The input feature map is the residual connection. This is the final output feature map after residual convolution. Since the residual convolution enhancement operation is applied separately to each established DCT feature branch and does not change the feature map dimension, the multi-scale frequency domain features and the three enhanced features... Figure 1 The values are [batch_size,96,64,64], [batch_size,192,32,32], and [batch_size,384,16,16].
[0044] Step S13 is used to fuse multi-scale spatial domain features with multi-scale frequency domain features to obtain a multi-scale fused feature map. Specifically, in this embodiment, features with the same spatial resolution in the RGB and DCT branches can be paired first to form feature map groups of different scales. Then, feature map groups of different scales for channel enhancement are obtained according to formulas (18) to (19): (18) (19) in, These are feature maps at different scales. For global average pooling, For the weights of the first fully connected layer, It is the ReLU activation function. For the weights of the second fully connected layer, It is the Sigmoid activation function. This is a channel attention weight map. For channel-by-channel multiplication, Feature map set after channel attention weighting; Finally, feature maps of different scales of spatial enhancement are obtained according to formulas (20) to (21): (20) ,(twenty one) in, Average pooling is performed along the channel dimension. Max pooling for the channel dimension, This is a spatial attention weight map. This results in the final attention-weighted feature map set. Finally, the final attention-weighted feature map set is concatenated and cross-modal information is fused to obtain a multi-scale fused feature map, which is [batch_size,96,64,64], [batch_size,192,32,32], and [batch_size,384,16,16].
[0045] Step S14 is used to perform frequency domain decomposition on the multi-scale fused feature map through the wavelet pattern feature enhancement module to obtain global fused features. Specifically, in this embodiment, the multi-scale fused feature map may first be input into the wavelet pattern feature enhancement module and wavelet pattern frequency domain decomposition may be performed according to formulas (22) to (23). Wavelet pattern frequency domain decomposition is performed on the fused multi-scale features according to formulas (22) to (23): ,(twenty two) ,(twenty three) in, For the fused multi-scale features, for Average pooling, Low-frequency characteristics after downsampling The low-frequency features are recovered by upsampling. These are high-frequency residual features. Each feature map is obtained through... Low-frequency components are extracted using average pooling, and then high-frequency residual components are obtained by upsampling from the original features and subtracting the low-frequency components. Next, these high- and low-frequency components are convolved and their output channels are uniformly mapped to 256 dimensions to generate a set of enhanced multi-scale frequency domain features with output sizes of: [batch_size, 256, 128, 128], [batch_size, 256, 64, 64], [batch_size, 256, 32, 32], and [batch_size, 256, 16, 16].
[0046] All 256-channel feature maps at all scales are upsampled to a 128×128 resolution, resulting in four enhanced feature maps of uniform size [batch_size, 256, 128, 128]. Finally, these four upsampled feature maps are concatenated along the channel dimension to fuse their multi-scale information, ultimately generating a globally fused feature map with significantly increased channel dimensions, outputting in dimensions [batch_size, 1024, 128, 128].
[0047] Step S15 is used to perform pixel-level classification on the globally fused features to obtain detection results. Specifically, in this embodiment, a classification head can be used to perform this classification through a... The convolutional layer processes globally fused features. The core of this convolutional operation is mapping the rich 1024-dimensional channel features of each pixel into a two-dimensional raw classification score. Its output dimension is [batch_size, 2, 128, 128]. Based on this, pixel-level classification is performed on the global fusion features, and the pixel-level classification probability is obtained according to formulas (24) to (25): ,(twenty four) (25) in, For global fusion features, for Convolution class header, The original classification score. For the target category parameter, To sum the index parameters, The score for category c, The score for category k, Let be the probability that a pixel belongs to class c. k=1 and k=2 corresponds to the normal region and the tampered region respectively, and the final output shape is Pixel-level detection images, such as Figure 4As shown, this is to achieve accurate detection of tampered areas at high resolution.
[0048] Furthermore, the model's performance can be evaluated based on pixel-level classification probabilities. Specifically, in this embodiment, the intersection-union ratio can be obtained according to formula (26): (26) in, Intersection over Union (IoU) measures the degree of overlap between the predicted and actual regions. For a real example, As a false positive example, This is a false negative example; The accuracy is obtained using formula (27): (27) in, Precision rate, which is the proportion of correctly predicted areas within the tampered region; The recall rate is obtained according to formula (28); (28) in, Recall rate, which is the proportion of genuinely tampered areas that are correctly detected; Obtain the F1 score according to formula (29): (29) in, The score is the F1 score. The area under the curve can be obtained using formulas (30) to (32): (30) (31) (32) in, For a true negative example, For the true rate, For false positives, The area under the curve (AUC) represents the overall classification performance of the model at different thresholds. The closer the AUC value is to 1, the better the model's performance. Essentially, it is the area under the ROC curve, which is composed of the true positive rate (FPR) and false positive rate (TPR) at different thresholds. By continuously changing the classification threshold, the model can obtain a series of (FPR, TPR) coordinate points, which ultimately form the ROC curve.
[0049] Based on the pixel-level detection results obtained in step S15, a comprehensive performance evaluation was conducted on the test set using five metrics: intersection-over-union ratio (IoU), precision, recall, F1 score, and AUC. Specific test results are as follows: Figure 5 As shown.
[0050] The test results demonstrate the quantitative results of the proposed method and the comparative method Cat-Net on five key metrics. For example... Figure 5 As shown, our proposed method outperforms Cat-Net across all metrics, specifically: AUC of 0.9624, F1 score of 0.7223, IOU of 0.7672, precision of 0.7192, and recall of 0.7255. This comprehensive lead in all metrics validates the effectiveness and superiority of our proposed model in image tampering detection.
[0051] Furthermore, hyperparameter tuning of the model can be performed based on evaluation metrics. Specifically, in this embodiment, the input image size can be uniformly set to 512×512 to maintain compatibility with mainstream visual models and obtain sufficient detail representation capabilities. The batch size is set to 4 based on the GPU memory capacity, and gradient stability is observed during training, with fine-tuning within the range of 4-8. The optimizer uses SGD with an initial learning rate of 0.005, employing a segmented decay strategy, i.e., adjusting by a factor of 0.1 at approximately 30%, 60%, and 80% of the training progress. The initial weight decay value is 0.0005, dynamically adjusted between 0.0003 and 0.001 based on overfitting. For data augmentation, while maintaining random horizontal flipping, a multi-scale training strategy of 0.8-1.2 times is used to improve the model's scale robustness. The number of training epochs is adjusted between 150-300 epochs based on convergence, the downsampling rate is fixed at 1 to preserve complete spatial information, and the number of data loading threads is set to 4-16 based on the number of CPU cores. The entire tuning process was guided by validation set metrics, focusing on adjusting core parameters such as learning rate, batch size, number of training epochs, weight decay, and multi-scale range to ensure that the model achieves the optimal balance between convergence speed and generalization ability.
[0052] On the other hand, the present invention also provides a ticket forgery detection system that integrates wavelet frequency domain enhancement and DCT-guided attention, the system including a processor configured to perform any of the control methods described above.
[0053] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the acquisition methods described above.
[0054] The beneficial effects of this invention are: This invention achieves efficient, accurate, and adaptive detection of tampered regions in document images by fusing multimodal features and an attention mechanism. Specifically, it includes: This invention possesses the ability to simultaneously extract spatial and frequency domain features using a two-stream network, and also achieves multi-scale feature fusion based on a cross-modal attention mechanism. By combining spatial feature perception with frequency domain anomaly detection, it realizes the functions of detail preservation and artifact recognition in document image tampering detection.
[0055] This invention enhances model robustness and improves detection stability in complex scenes by using a wavelet pattern feature enhancement module. Furthermore, the method of this invention outperforms traditional tamper detection models in both detection accuracy and generalization ability. Traditional detection methods are insensitive to text regions and subtle tampering in document images, and their performance significantly degrades under complex shooting conditions such as low light and reflections. This invention, through multi-scale frequency domain decomposition and feature reconstruction, enables the model to maintain stable feature extraction capabilities under different imaging conditions, greatly enhancing its practical application value.
[0056] Because the present invention completes the deep fusion and optimization of multimodal features during the training phase, the computational overhead of the dual-stream network and attention mechanism is controllable and does not require additional preprocessing modules during actual deployment. At the same time, it supports incremental learning of new tampering methods, can continuously adapt to the development and changes of document forgery technology, and can gradually improve the accuracy, adaptability and reliability of document anti-counterfeiting detection system.
[0057] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0058] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0059] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0060] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0061] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0062] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0063] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0064] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0065] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for detecting ticket forgery that integrates wavelet frequency domain and DCT, characterized in that, The detection method includes: Acquire the image to be detected; RGB features are extracted from the image to be detected to obtain multi-scale spatial domain features; DCT feature extraction is performed on the image to be detected to obtain multi-scale frequency domain features; The multi-scale spatial domain features and multi-scale frequency domain features are fused to obtain a multi-scale fused feature map. The multi-scale fused feature map is decomposed in the frequency domain using a wavelet pattern feature enhancement module to obtain global fused features; The global fusion features are classified at the pixel level to obtain detection results.
2. The detection method according to claim 1, characterized in that, RGB feature extraction is performed on the image to be detected to obtain multi-scale spatial domain features, including: The first branch feature map and the second branch feature map are obtained based on the image to be detected; According to formula (1), feature enhancement is performed on the first branch feature map and the second branch feature map to obtain the first branch enhanced feature map and the second branch enhanced feature map: ,(1) in, For the first The output feature map of the s-th branch at the 1st level. For hierarchical indexes, For the convolution operation corresponding to the s-th branch, For upsampling or downsampling, Indexes to all branches except the current target branch s; The third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; The third branch feature map is enhanced using formula (1) to obtain the third branch enhanced feature map; The fourth branch feature map is obtained based on the first branch enhanced feature map, the second branch enhanced feature map, and the third branch enhanced feature map; The fourth branch feature map is enhanced using formula (1) to obtain the fourth branch enhanced feature map.
3. The detection method according to claim 2, characterized in that, RGB feature extraction is performed on the image to be detected to obtain multi-scale spatial domain features, including: Multi-scale spatial domain features are obtained according to formulas (2) to (5): ,(2) ,(3) ,(4) ,(5) in, For horizontal edge response, For vertical edge response, These are the pixel values of the normalized image. For total edge features, This is the final RGB feature map after edge enhancement of the i-th branch. The final output feature map obtained by fusing features from the i-th branch is... This represents the corresponding number of branches.
4. The detection method according to claim 1, characterized in that, DCT feature extraction is performed on the image to be detected to obtain multi-scale frequency domain features, including: The frequency coefficients of each pixel block in the image to be detected are obtained according to formulas (6) to (7): ,(6) ,(7) in, These are the transformed DCT coefficients. For horizontal frequency indexing, For vertical frequency indexing, For horizontal basis functions, For vertical basis functions, These are the normalization coefficients; DCT feature maps are obtained based on the frequency coefficients; The DCT feature map is convolved to obtain the first branch feature map and the second branch feature map. Feature enhancement is performed on the first branch feature map and the second branch feature map to obtain a first branch enhanced feature map and a second branch enhanced feature map; The third branch feature map is obtained based on the first branch enhanced feature map and the second branch enhanced feature map; The third branch feature map is enhanced to obtain a third branch enhanced feature map; Multi-scale frequency domain features are obtained according to formulas (8) to (9): ,(8) ,(9) in, The feature map output by the convolutional layer. Convolutional layers specifically designed for processing DCT features. The input feature map is the residual connection. This is the final output feature map after residual connection.
5. The detection method according to claim 1, characterized in that, The multi-scale spatial domain features and multi-scale frequency domain features are fused to obtain a multi-scale fused feature map, including: Features with the same spatial resolution in the RGB and DCT branches are paired to form feature map groups of different scales; According to formulas (10) to (11), feature maps of different scales of channel enhancement are obtained: ,(10) ,(11) in, These are the feature maps at different scales. For global average pooling, For the weights of the first fully connected layer, It is the ReLU activation function. For the weights of the second fully connected layer, It is the Sigmoid activation function. This is a channel attention weight map. For channel-by-channel multiplication, Feature map set after channel attention weighting; Based on formulas (12) to (13), obtain feature maps of different scales for spatial augmentation: ,(12) ,(13) in, Average pooling is performed along the channel dimension. Max pooling for the channel dimension, This is a spatial attention weight map. For the final attention-weighted feature map set; The final attention-weighted feature map group is spliced and cross-modal information fusion is performed to obtain a multi-scale fused feature map.
6. The detection method according to claim 1, characterized in that, The multi-scale fused feature map is decomposed in the frequency domain using a wavelet pattern feature enhancement module to obtain global fused features, including: The multi-scale fused feature map is decomposed into wavelet-style frequency domain according to formulas (14) to (15): ,(14) ,(15) in, The multi-scale fused feature map, for Average pooling, Low-frequency characteristics after downsampling The low-frequency features are recovered by upsampling. It is a high-frequency residual characteristic; Enhanced frequency domain features are obtained based on the low-frequency features and high-frequency residual features.
7. The detection method according to claim 1, characterized in that, The global fusion features are classified at the pixel level to obtain detection results, including: Pixel-level classification probabilities are obtained using formulas (16) and (17): ,(16) ,(17) in, For global fusion features, for Convolution class header, The original classification score. For the target category parameter, To sum the index parameters, The score for category c, The score for category k, Let be the probability that a pixel belongs to category c.
8. The detection method according to claim 1, characterized in that, The detection method also includes performance evaluation of the model, including: Obtain the intersection-union ratio according to formula (18): ,(18) in, For intersection, union, and comparison, For a real example, This is a false positive example. This is a false negative example; The accuracy is obtained using formula (19): ,(19) in, For accuracy; The recall rate is obtained according to formula (20); ,(20) in, Recall rate; Obtain the F1 score according to formula (21): ,(21) in, The score is the F1 score. The area under the curve can be obtained using formulas (22) to (24): ,(22) ,(23) (24) in, For a true negative example, For the true rate, For false positives, The area under the curve is denoted as .
9. A ticket forgery detection system integrating wavelet frequency domain and DCT, characterized in that, The detection system includes a processor configured to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.