Deep forgery detection method based on firefly optimization algorithm and locality sensitive hashing

Through the combination of firefly optimization algorithm and local sensitive hashing, document feature weights are optimized and a forged sensitive hash model is constructed, which solves the problem of identifying complex forged images in the existing technology, and efficient and adaptive document forgery detection is achieved, which reduces the false alarm rate and missed alarm rate, and adapts to the diversity forgery mode.

CN120339813APending Publication Date: 2025-07-18SUZHOU HUANLONG NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510407382.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing document forgery detection methods have poor recognition effects when facing complex or high-fidelity forgery images, and are highly computationally complex, which is difficult to meet the real-time detection needs of large-scale document data. It lacks dynamic perception and adaptability of forgery technology, making it difficult to deal with rapid evolution and diversity scenarios.

Method used

The deep forgery detection method based on firefly optimization algorithm and locally sensitive hash is adopted. The document feature weight is optimized through the firefly optimization algorithm, and a locally sensitive hash matching model is constructed. Combined with cross-scale forgery adversarial fitness function and forgery sensitive hash function family, a step-by-step enhanced forgery recognition process from initial screening to confirmation is realized.

Benefits of technology

It significantly reduces the false alarm rate and omission rate, improves the recognition ability of complex forged documents, provides high interpretability and visual support, adapts to the diverse forgery mode, and improves the clustering speed and detection rate of forgery samples in large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339813A_ABST
    Figure CN120339813A_ABST
Patent Text Reader

Abstract

The invention discloses a deep forgery detection method based on a firefly optimization algorithm and locality sensitive hashing. The deep forgery detection method comprises the following steps: S1, converting collected document data into a standard format image document data set; s2, generating a preprocessed image document data set; s3, forming a document feature vector set; s4, obtaining an optimal feature weight combination; s5, generating an optimized document feature vector set; s6, constructing a locality sensitive Hash matching model based on the optimized document feature vector set, and mapping the document feature vectors into a low-dimensional space through a preset Hash function family to form a plurality of Hash buckets; and S7, carrying out similarity matching on the document data in each hash bucket, and outputting a final document data high-forgery detection result and a corresponding document forgery confidence score by adopting structural similarity comparison and image training comparison. According to the method, a step-by-step enhanced counterfeit identification process from preliminary screening to confirmation is realized, and the false alarm rate and the missing report rate are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forgery detection, and in particular, to a deep forgery detection method based on a firefly optimization algorithm and local sensitive hashing. Background Art

[0002] With the development and application of digital technologies, document data plays a core role in the fields of financial bills, document recognition, legal documents, and copyright protection. However, the accompanying information security risks are becoming increasingly serious. In particular, the continuous evolution of document image forgery technologies has posed great challenges to traditional document authenticity verification mechanisms. Currently, forgery technologies have evolved from simple image clipping to highly complex and sophisticated forgery methods based on AI generation, image cloning, region stitching, and texture migration. These technologies can often bypass traditional detection means, seriously threatening data credibility and business security.

[0003] Existing document forgery detection methods mainly include image processing methods based on artificial rules, image classification methods based on deep learning, and detection methods based on feature template comparison. Artificial rule methods usually achieve forgery judgment by detecting edges and analyzing low-level visual features such as color histograms. Although effective for simple forgeries, they have poor recognition effects on images with complex structures or locally high-fidelity forgeries. Deep learning methods, although having strong feature extraction capabilities, often rely on a large number of labeled forgery samples for training, facing problems such as insufficient training data coverage, lagging model updates, and poor adaptability to new forgery means. At the same time, their model structures are complex and the computational costs are high, making it difficult to meet the real-time detection requirements in the environment of massive document data.

[0004] In addition, in the big data document environment, how to efficiently retrieve suspected forgery samples from millions of documents has always been a key point that is difficult to break through in the existing technology. Current retrieval technologies mostly rely on brute-force matching, full-text feature comparison, or fixed feature indexing, facing problems such as high computational complexity, slow response speed, and high false matching rate. They are particularly inefficient when dealing with non-uniform format and unstructured image documents. At the same time, existing feature extraction methods do not fully consider the key regions, layout structures, and text-graphic fusion information inside the documents, resulting in limited expression capabilities for forgery information. In addition, existing methods mostly adopt static model parameter configurations and lack the ability of dynamic perception and adaptive adjustment for forgery features, making it difficult to cope with the rapid evolution of forgery technologies and the adaptation requirements of diverse scenarios.

[0005] In summary, there is an urgent need for a detection method that combines fast processing capabilities, adaptive optimization capabilities, and sensitivity to forgery features to achieve intelligent recognition and efficient matching of highly forged document data. Summary of the Invention

[0006] An object of the present invention is to propose a deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing. The present invention realizes a gradually enhanced forgery recognition process from preliminary screening to confirmation, not only significantly reducing the false positive rate and false negative rate, but also providing high interpretability and visualization support for subsequent manual review or system automatic decision-making.

[0007] A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to an embodiment of the present invention includes the following steps:

[0008] S1. Collect the document data to be detected, and convert the collected document data into a standard format image document dataset;

[0009] S2. Preprocess the standard format image document dataset to generate a preprocessed image document dataset;

[0010] S3. Extract features from the preprocessed image document dataset to form a document feature vector set;

[0011] S4. Apply the firefly optimization algorithm to globally search the document feature vector set, optimize the document feature weights, and obtain the optimal feature weight combination;

[0012] S5. Use the optimal feature weight combination to adjust the document feature vector set to generate an optimized document feature vector set;

[0013] S6. Build a locality-sensitive hashing matching model based on the optimized document feature vector set, map the document feature vectors into a low-dimensional space through a preset hash function family, and form multiple hash buckets;

[0014] S7. Perform similarity matching on the document data in each hash bucket, screen out the candidate document data that highly matches the known forged document features, and further perform multi-level forgery confirmation on the screened candidate document data, using structural similarity comparison and image recognition comparison, and output the final highly forged detection result of the document data and the corresponding document forgery confidence score.

[0015] Optionally, the S1 includes the following steps:

[0016] S11. Collect the original document data to be detected. The original document data includes image format documents and non-image format documents. Each original document data consists of the original file format of the document, timestamp information, and the document source identification code;

[0017] S12. Identify the format of the collected original document data. For non-image format documents that do not meet the image processing requirements, call the format conversion module to convert the non-image format documents into image format documents with a unified standard. During the format conversion process, maintain the integrity of the document content structure and page layout, so that the converted image format documents retain the visual information corresponding to the original documents;

[0018] S13. Perform standard size adjustment operations on the image format documents, uniformly scale all image format documents to the set height and width, maintain the consistency of the image aspect ratio during the size adjustment process, and perform size alignment;

[0019] S14. Combine the image format documents that have completed format conversion and size standardization with the corresponding timestamp information and source identification codes to form a standard format image document dataset D std 。

[0020] Optionally, the S2 includes the following steps:

[0021] S21. Perform image noise removal processing on each image document in the standard format image document dataset D std Perform a smoothing operation on the image pixel matrix of the image document. After the noise removal processing, a noise suppression image matrix is obtained;

[0022] S22. Perform edge enhancement processing on the noise suppression image matrix to enhance the texture and contour features of the image edge region, and obtain an edge enhancement image matrix;

[0023] S23. Perform size normalization processing on the edge enhancement image matrix to obtain a normalized image matrix;

[0024] S24. Perform grayscale processing on the normalized image matrix. If the image document is a color image, convert the image document into a single-channel grayscale image to obtain a grayscale image matrix and combine the grayscale image matrix with its corresponding timestamp information t i and source identification code s i to form a preprocessed image document dataset where N represents the total number of image documents in the preprocessed image document dataset.

[0025] Optionally, the S3 includes the following steps:

[0026] S31. Extract texture features from each grayscale image matrix in the preprocessed image document dataset Calculate the gray-level co-occurrence matrix of the local region of the image and extract texture statistics, including contrast, energy, entropy, and uniformity, to form a texture feature vector

[0027] S32. Apply an edge detection operator to the grayscale image matrix to extract the main edge contour regions in the image, and perform quantization encoding on the edge direction, edge density, and edge continuity to obtain an edge contour feature vector

[0028] S33. Based on the pixel arrangement rule of the image, extract the structure features of the grayscale image matrix Calculate the average gray value, standard deviation, and position centroid coordinates of each region of the structure features, and construct a structure feature vector

[0029] S34. Select the position of the key region of the page as the region of interest in the grayscale image matrix According to the preset layout template, extract the key text blocks or seal regions containing forgery traces in the image, and calculate the relative position, gray scale statistics, and texture complexity of the key text blocks or seal regions containing forgery traces to construct a key region distribution feature vector

[0030] S35. Perform a feature-level splicing operation on the texture feature vector edge contour feature vector structure feature vector and the key region distribution feature vector to obtain the final feature vector of the i-th image document, and construct a set of the final feature vectors of all image document samples to form a document feature vector set

[0031] Optionally, the S4 includes the following steps:

[0032] S41. Initialize the parameters of the firefly population, set the initial population size as M, the maximum number of iterations as T max , the initial light intensity as I0, the light intensity attenuation factor as γ, the initial attractiveness as β0, and use the document feature vector set as the initial search space of the firefly algorithm;

[0033] S42. Construct a variable-structure firefly network and adaptively divide the firefly population into subgroups;

[0034] S43. For each divided subgroup, define a cross-scale forgery adversarial fitness function for firefly individuals. The cross-scale forgery adversarial fitness function combines the position of firefly individuals with the cross-scale enhancement and robustness resistance of the high forgery risk features of the document:

[0035]

[0036] Among them, Denote the cross-scale forgery adversarial fitness value of the j-th firefly individual in the c-th subgroup. is the corresponding document feature weight vector. Denote the enhanced value of the forgery risk feature sensitivity under the current weight. Denote the robustness resistance ability of the forgery risk feature under cross-scale perturbation, where α and λ are weight factors.

[0037] S44. Use the cross-scale forgery adversarial fitness value as the basis for updating the light intensity. The light intensity of each firefly individual is adjusted exponentially according to the light intensity fitness value. The higher the fitness value, the slower the light intensity decays, retaining a stronger attraction ability to guide other firefly individuals to approach the current firefly individual. For firefly individuals with a low light intensity fitness value, the light intensity drops rapidly, and the attraction weakens during the search process, promoting the rapid convergence of the global optimal solution.

[0038] S45. Calculate the mutual attraction degree between firefly individuals within the subgroup based on the difference in light intensity of firefly individuals:

[0039]

[0040] Among them, Denote the attraction degree of the j-th firefly individual to the k-th firefly individual in the c-th subgroup. Denote the Euclidean distance between the positions of the j-th and k-th firefly individuals. is the difference degree of fitness between the two firefly individuals, making the movement between firefly individuals tend to the firefly individual with more forgery adversarial advantages.

[0041] S46. Adopt a firefly individual position update strategy that fuses cross-scale forgery adversarial gradient and attraction degree information to perform dynamic optimization of the firefly individual position. Define the updated firefly individual position as:

[0042]

[0043] Among them, and respectively denote the updated document feature weight position and the pre-updated document feature weight position of the j-th firefly individual in the c-th subgroup. η is the cross-scale forgery adversarial gradient guiding coefficient. Denote the gradient vector of the cross-scale forgery adversarial fitness function at the current position. ξ is a random perturbation factor, and rand is a random vector.

[0044] Compared with traditional algorithms that only move based on attractiveness, the present invention introduces a cross-scale forgery-sensitive gradient term, which can directly move towards the most discriminative feature direction, significantly improving the feature learning efficiency in forgery detection tasks; while the random perturbation term prevents getting stuck in local optima, and this structure is particularly effective in dealing with difficult-to-distinguish boundary forgeries.

[0045] S47. At the end of each iteration, re-perform subgroup division based on the currently updated set of document feature vectors, adjust the subgroup structure in real time, and repeat steps S42 to S46 until the maximum number of iterations T is reached. max , and finally output the document feature weight vector corresponding to the firefly individual with the highest cross-scale forgery adversarial fitness value as the optimal feature weight combination W for highly forged detection of document data. * 。

[0046] Optionally, the subgroup division is specifically based on the distribution of the set of document feature vectors in the high-dimensional space, and an hierarchical clustering method based on the Euclidean distance between feature vectors is used to adaptively divide the firefly population to generate multiple subgroups. Each subgroup aggregates firefly individuals with similar features. When the Euclidean distance between the feature vectors of any two image format document samples is less than the dynamically set Euclidean distance threshold, they are grouped into the same subgroup, so that each subgroup has common forgery risk features internally.

[0047] Optionally, S5 includes the following steps:

[0048] S51. Obtain the optimal feature weight combination for highly forged detection of document data finally output, and perform a weight mapping adjustment operation on the set of document feature vectors, using the optimal feature weight combination W * to perform per-dimension weighted adjustment on each document feature vector to generate an optimized document feature vector

[0049]

[0050] where, represents the optimal weight value corresponding to the D-th feature dimension in the optimal feature weight combination, and f i,D represents the original document feature vector value of the i-th document data in the D-th dimension, and D is the total dimension of the document feature vector;

[0051] S53. Store all the weighted-adjusted document feature vectors in sequence to form a new set of document feature vectors as the optimized set of document feature vectors The optimized set of document feature vectors performs weight strengthening processing on the feature vectors of each document sample based on the forgery detection optimization goal while maintaining the original order of the document samples.

[0052] Optionally, S6 includes the following steps:

[0053] S61. Based on the optimized document feature vector set Construct a family of locality-sensitive hashing functions H that incorporates the sensitivity weights of forgery features fake :

[0054]

[0055] where represents the optimized document feature vector set The hash value after being mapped by the forgery-sensitive weight, a fake is a projection vector randomly generated from the standard normal distribution, with the same dimension D as the dimension of the optimized document feature vector. b fake represents an offset value randomly generated within the interval [0, r fake , r fake is the dynamic forgery feature sensitivity hashing quantization scale parameter. The symbol represents the Hadamard element-wise product operation, reflecting the feature sensitivity guidance of the optimal feature weight combination W * for the hash mapping process represents the set of real numbers consisting of all real vectors of dimension D is the set of positive real numbers;

[0056] S62. For the highly forged document detection task, construct a hash function adaptive selection mechanism. According to the forgery feature sensitive distribution state of the optimized document feature vector set in the low-dimensional space, dynamically adjust the number and parameter configuration of the locality-sensitive hashing function family H fake in the hash functions, so that the hash functions are concentratedly distributed in the sensitive areas where highly forged document features tend to gather, and obtain the forgery-sensitive hash codes of the optimized document features

[0057] S63. Based on the set of forgery-sensitive hash codes of all optimized document features Construct a forgery feature adaptive multi-level hash index structure to form a set of forgery-sensitive hash buckets B fake , and each forgery-sensitive hash bucket is automatically indexed according to the forgery risk level of the forgery-sensitive hash code, forming a multi-level hash index structure from high to low.

[0058] Optionally, S7 includes the following steps:

[0059] S71. For each hash bucket in the set of forgery-sensitive hash buckets Perform local similarity matching operations on the optimized document feature vectors within, based on the forged sensitive hash codes in the optimized document feature vector set Calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library:

[0060]

[0061] where Sim i,k is the similarity between the i-th document to be detected and the k-th forged sample, and represent the optimized feature vectors of the document to be detected and the forged sample respectively. If the feature cosine similarity Sim i,k ≥θ sim , where θ sim is the forged similarity threshold, then the i-th document is determined as a candidate forged sample;

[0062] S72. Incorporate the document data samples that meet the forged similarity threshold into the candidate document data set D cand , perform structural similarity comparison on each candidate document sample in the candidate document data set, extract the grayscale image matrix of the candidate document and the image structure information of the forged sample, and calculate the structural similarity index SSIM:

[0063]

[0064] where x is the grayscale image matrix block of the current document sample to be detected, y is the corresponding forged sample image matrix block, μ x , μ y are the local means of the images respectively, is the variance, σ xy is the covariance, and C1, C2 are stable constants to prevent the denominator from being zero; when the structural similarity index SSIM(x, y)≥θ ssim , where θ ssim is the image structure similarity threshold, it is determined that the image structure has high structural similarity;

[0065] S74. Based on the structural similarity comparison, classify and identify the key regions in the candidate document image and extract the image features, evaluate whether there are forged features such as tampering, cloning, splicing, or image enhancement in the key regions of the candidate document image, and compare and verify them with the features in the forged template library;

[0066] S75. Define the document forgery confidence score C i :

[0067] C i =ω1·Simi +ω2·SSIM i ;

[0068] where ω1 and ω2 represent weights, and SSIM i is the structural similarity index score, and Sim i is the feature cosine similarity score;

[0069] S76. Output the final highly forged detection result of the document data according to the forgery confidence score C i and the set detection result determination rule.

[0070] Optionally, the final highly forged detection result of the document data includes:

[0071] Normal document data: When the forgery confidence score C i < 0.4, there is no forgery risk;

[0072] Suspicious document data: When the forgery confidence score 0.4 ≤ C i < 0.7, there is a forgery trend, and further manual review is recommended;

[0073] High-risk forged document data: When the forgery confidence score C i ≥ 0.7, it is highly suspected of forgery, marked as a detection positive and a comparison report is output.

[0074] The beneficial effects of the present invention are:

[0075] (1) Based on the traditional firefly optimization algorithm, the present invention designs a variable-structure firefly network based on the dynamic restructuring of the document structure, which can divide dynamic subgroups according to the clustering structure of document features, improving the modeling ability of the optimization algorithm for the local forgery feature distribution. At the same time, a cross-scale forgery adversarial fitness function combining forgery enhancement and robust resistance ability is defined, and the weight optimization of the forgery-sensitive direction is introduced through the gradient guidance mechanism, so that the finally obtained feature weight combination can not only highlight the forgery traces but also improve the stability against complex interference, thereby enhancing the adaptability of the entire detection system to diverse forgery patterns.

[0076] (2) During the construction of local sensitive hashing, the present invention introduces the forgery sensitivity weights of feature dimensions, and fuses the optimized feature vector and the weight vector dimension by dimension through the Hadamard product to form a forgery sensitivity hash function family, making the hash mapping process more capable of focusing on forgery features. At the same time, the system dynamically constructs a multi-level index structure according to the risk level of the hash code, maps the document feature vector to the structured forgery risk hierarchical space, greatly improving the clustering speed and detection rate of high-risk forged samples in large-scale data, and overcoming the problem of insufficient matching ability of traditional hashing methods for complex structural forged samples.

[0077] (3) The present invention designs a two - level forgery determination mechanism. First, candidate documents are screened through the optimized feature cosine similarity. Then, the structural similarity index comparison and the forgery pattern analysis of key regions are performed on the candidate samples. Finally, a forgery confidence score is generated based on the similarity fusion score, effectively integrating the global feature similarity and the fine discrimination ability of the local image structure, realizing a gradually enhanced forgery recognition process from preliminary screening to confirmation. This not only significantly reduces the false positive rate and false negative rate, but also provides high interpretability and visualization support for subsequent manual review or system automatic decision - making. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0079] Figure 1 is a flowchart of a deep - fake detection method based on the firefly optimization algorithm and locality - sensitive hashing proposed by the present invention;

[0080] Figure 2 is a schematic diagram of the dynamic update of firefly individuals under the forgery - resistant fitness function of a deep - fake detection method based on the firefly optimization algorithm and locality - sensitive hashing proposed by the present invention;

[0081] Figure 3 is a flowchart of locality - sensitive hashing mapping and hash bucket construction based on forgery - sensitive weights of a deep - fake detection method based on the firefly optimization algorithm and locality - sensitive hashing proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0083] Refer to Figures 1-3 , a deep - fake detection method based on the firefly optimization algorithm and locality - sensitive hashing, includes the following steps:

[0084] S1. Collect the document data to be detected and convert the collected document data into a standard - format image document dataset;

[0085] S2. Pre - process the standard - format image document dataset to generate a pre - processed image document dataset;

[0086] S3. Extract features from the pre - processed image document dataset to form a document feature vector set;

[0087] S4. Apply the firefly optimization algorithm to globally search the document feature vector set, optimize the document feature weights, and obtain the optimal feature weight combination;

[0088] S5. Use the optimal feature weight combination to adjust the document feature vector set and generate an optimized document feature vector set;

[0089] S6. Build a locality-sensitive hashing matching model based on the optimized document feature vector set, map the document feature vectors into a low-dimensional space through a preset hash function family, and form multiple hash buckets;

[0090] S7. Perform similarity matching on the document data in each hash bucket, screen out candidate document data that highly matches the known forged document features, and further perform multi-level forgery confirmation on the screened candidate document data. Use structural similarity comparison and image recognition comparison to output the final highly forged detection result of the document data and the corresponding document forgery confidence score.

[0091] In this embodiment, S1 includes the following steps:

[0092] S11. Collect the original document data to be detected. The original document data includes image format documents and non-image format documents. Each piece of original document data consists of the original file format of the document, timestamp information, and document source identification code;

[0093] S12. Perform format recognition on the collected original document data. For non-image format documents that do not meet the image processing requirements, call the format conversion module to convert the non-image format documents into image format documents with a unified standard. The format conversion process maintains the integrity of the document content structure and page layout, so that the converted image format documents retain the visual information corresponding to the original documents;

[0094] S13. Perform standard size adjustment operations on the image format documents, uniformly scale all image format documents to the set height and width, maintain the consistency of the image aspect ratio during the size adjustment process, and perform size alignment;

[0095] S14. Combine the image format documents that have completed format conversion and size standardization with the corresponding timestamp information and source identification code to form a standard format image document dataset D std .

[0096] In this embodiment, S2 includes the following steps:

[0097] S21. Perform image noise removal processing on each image document in the standard format image document dataset D std and perform a smoothing operation on the image pixel matrix of the image document. After the noise removal processing, a noise suppression image matrix is obtained;

[0098] S22. Perform edge enhancement processing on the noise-suppressed image matrix to enhance the texture and contour features of the image edge region, obtaining an edge-enhanced image matrix;

[0099] S23. Perform size normalization processing on the edge-enhanced image matrix to obtain a normalized image matrix;

[0100] S24. Perform grayscale processing on the normalized image matrix. If the image document is a color image, convert the image document into a single-channel grayscale image to obtain a grayscale image matrix and combine the grayscale image matrix with its corresponding timestamp information t i and the source identification code s i to form a preprocessed image document dataset where N represents the total number of image documents in the preprocessed image document dataset.

[0101] In this embodiment, S3 includes the following steps:

[0102] S31. For each grayscale image matrix in the preprocessed image document dataset perform texture feature extraction, calculate the gray-level co-occurrence matrix of the local region of the image and extract texture statistics, including contrast, energy, entropy, and homogeneity, to form a texture feature vector

[0103] S32. Apply an edge detection operator to the grayscale image matrix to extract the main edge contour region in the image, and perform quantization coding on the edge direction, edge density, and edge continuity to obtain an edge contour feature vector

[0104] S33. Based on the pixel arrangement rule of the image, extract the structural features of the grayscale image matrix calculate the gray-level mean value, standard deviation, and position centroid coordinates of each region of the structural features, and construct a structural feature vector

[0105] S34. Select the position of the key region of the page as the region of interest in the grayscale image matrix extract the key text block or seal region containing forgery traces in the image according to the preset layout template, calculate the relative position, gray-level statistics, and texture complexity of the key text block or seal region containing forgery traces, and construct a key region distribution feature vector

[0106] S35. The texture feature vector the edge contour feature vector the structural feature vector With the eigenvector of the key area distribution Perform feature-level splicing operation to obtain the final eigenvector of the i-th image document, and construct a set of the final eigenvectors of all image document samples to form a document eigenvector set

[0107] In this embodiment, S4 includes the following steps:

[0108] S41. Initialize the parameters of the firefly population, set the initial population size as M, the maximum number of iterations as T max , the initial light intensity as I0, the light intensity attenuation factor as γ, the initial attractiveness as β0, and use the document eigenvector set as the initial search space of the firefly algorithm;

[0109] S42. Construct a variable-structure firefly network to adaptively divide the firefly population into subgroups;

[0110] S43. For each divided subgroup, define the cross-scale forgery adversarial fitness function of the firefly individual. The cross-scale forgery adversarial fitness function combines the cross-scale enhancement and robustness resistance of the firefly individual position and the high forgery risk features of the document features:

[0111]

[0112] Among them, represents the cross-scale forgery adversarial fitness value of the j-th firefly individual in the c-th subgroup, is the corresponding document feature weight vector, represents the enhanced value of the forgery risk feature sensitivity under the current weight, represents the robustness resistance ability of the forgery risk feature under cross-scale perturbation, and α, λ are weight factors;

[0113] The formula of S43 is used to evaluate the feature weight vector represented by the j-th firefly individual in the c-th subgroup The overall performance score between "forgery sensitivity enhancement" and "robustness maintenance" is used as the fitness value of this individual. The higher the fitness value, the more helpful the feature combination is to expose potential document forgery risks.

[0114] Traditional fitness functions only consider classification accuracy or error minimization, ignoring the dual requirements of "forgery feature enhancement" and "anti-perturbation". The present invention constructs a unified evaluation function by linearly combining forgery sensitivity and feature robustness with adjustable parameters, ensuring that during the optimization process, forgery traces can be captured and the model is prevented from overfitting to pseudo signals, which is applicable to dealing with forgery scenarios with strong concealment such as "micro-forgery" and "boundary fusion".

[0115] For measuring the current feature weight combination The ability to magnify the difference between forged samples and normal samples, in order to obtain Based on the current weight combination, this embodiment performs weighted processing on all known labeled forged samples and normal samples in the document feature vector set respectively to obtain new weighted feature vectors. Then, the system statistically calculates the weighted feature means and the degree of feature difference of the forged samples and normal samples respectively, and then calculates the separation degree between the forged samples and normal samples. When the difference in feature means between the weighted forged samples and normal samples is larger and the intra-class consistency is higher (i.e., the volatility is smaller), it indicates that the current weight combination is more conducive to effectively distinguishing forged features from normal documents, and the sensitivity of forgery detection is also higher. At this time, the value of this index is also higher. Therefore, It can reflect the ability of the current feature weight combination in enhancing the separability of forgery detection, and is a key metric for measuring whether the difference in forged features is magnified during the optimization process.

[0116] For evaluating the anti-interference ability of the current feature weight combination against the common interferences (scanning blur, rotation, compression, scaling) in actual forged documents, to prevent the system from misjudging due to slight deformations or processing means, in order to obtain In this embodiment, multiple variant versions with interferences are generated for each forged sample, and different degrees of brightness adjustment, edge blurring, or scaling and rotation are performed. After the perturbed samples and the original samples are weighted and processed together, the system evaluates the consistency of their feature manifestations. If, under the current weight combination, the features of the perturbed version and the original version remain highly similar, it indicates that the weight combination has strong robustness and can effectively resist common image interferences and reduce false alarms. Therefore, It reflects the ability of the system to maintain stable recognition while not losing detection sensitivity, ensuring that feature optimization not only improves the discrimination degree but also has the ability to resist forgery variations.

[0117] and Form a "sensitivity enhancement + stable recognition" dual-objective mechanism:

[0118] Sensitivity enhancement is responsible for improving the ability of the detection system to magnify forged features;

[0119] Robustness resistance is responsible for ensuring that the detection system still maintains a stable performance under different forgery operations. By introducing these two dimensions as components of the fitness function in the firefly optimization algorithm, the system achieves a dynamic balance between accuracy and stability, and at the same time has the ability to adapt to new forgery methods, greatly improving the practicality and reliability of highly forged document data detection.

[0120] S44. Use the cross-scale forgery resistance fitness value as the basis for updating the light intensity. The light intensity of each firefly individual is adjusted exponentially according to the light intensity fitness value. The higher the fitness value, the slower the light intensity decays, retaining a stronger attraction ability to guide other firefly individuals to approach the current firefly individual. For firefly individuals with a low light intensity fitness value, the light intensity decreases rapidly, weakening the attraction during the search process and promoting the rapid convergence of the global optimal solution;

[0121] S45. Based on the difference in the light intensity of firefly individuals, calculate the mutual attraction degree between firefly individuals within the subgroup:

[0122]

[0123] Among them, represents the attraction degree of the j-th firefly individual to the k-th firefly individual within the c-th subgroup, represents the Euclidean distance between the positions of the j-th and k-th firefly individuals, is the difference degree of fitness between two firefly individuals, making the movement between firefly individuals tend to the firefly individual with more forgery resistance advantages;

[0124] The formula in S45 is used to measure the intensity of the attraction of the j-th firefly individual by the k-th individual in the c-th subgroup. The higher the attraction degree, the easier it is for this individual to approach the better solution, accelerating the convergence process.

[0125] Innovatively introduce the fitness difference as an attraction enhancement term, breaking the inertia of the traditional model that only drives movement by distance. In the actual scenario, even if the feature weights are close in distance, but if the response capabilities to forged features are very different, their aggregation still needs to be avoided; conversely, individuals with a slightly farther distance but better fitness should enhance the attraction. This mechanism can guide the fireflies to aggregate in the direction of stronger forgery discrimination ability, effectively improving the group search efficiency and discrimination stability.

[0126] The core meaning is that in traditional firefly optimization, the attraction is mainly controlled by the position distance and does not necessarily consider the strength of individual fitness; in the present invention, the fitness f(W) of each individual is a cross-scale forgery resistance function, representing the comprehensive performance of the feature weight combination in forgery detection;

[0127] If the fitness of a firefly is much higher than the current individual, its "attraction intensity" should be greater; so the absolute value term of the fitness difference is introduced here to strengthen the tendency to approach high-fitness individuals.

[0128] This can make the feature weight update not only "approach" but also preferentially approach the direction that is "more effective" and sensitive and robust to forged information.

[0129] In document forgery detection, the feature space is very complex, presenting the following challenges: the distances between different documents may be small (similar formats), but their "forgery risk differences" are large; if the search is guided only by the "feature space distance", it is easy to fall into local optima and difficult to jump out of the forgery-insensitive area;

[0130] The current formula can make the optimization search tend to the weight direction with better forgery detection effect by introducing the "fitness difference" as an additional gravitational factor, ultimately enhancing the robustness and globality of the overall optimization search.

[0131] S46. Adopt a firefly individual position update strategy that fuses cross-scale forgery adversarial gradient and attractiveness information to dynamically optimize the positions of firefly individuals. Define the updated position of a firefly individual as:

[0132]

[0133] where and respectively represent the position of the document feature weight after update and before update of the j-th firefly individual in the c-th subgroup. η is the cross-scale forgery adversarial gradient guiding coefficient, represents the gradient vector of the cross-scale forgery adversarial fitness function at the current position. ξ is a random perturbation factor, and rand is a random vector;

[0134] Compared with the traditional algorithm that moves only relying on attractiveness, the present invention introduces a cross-scale forgery-sensitive gradient term, which can directly move towards the most discriminative feature direction, significantly improving the feature learning efficiency in forgery detection tasks; while the random perturbation term prevents falling into local optima, and this structure is especially effective when dealing with difficult-to-distinguish boundary forgeries.

[0135] is the introduced local gradient optimization term, which is derived from the cross-scale forgery adversarial fitness function of the current individual, and the direction points to the fastest direction to improve fitness.

[0136] The traditional firefly algorithm is only guided by attractiveness and is prone to falling into local optima. This term directly introduces the directional knowledge in the task scenario to jump out of the trap and strengthen the rapid approximation to the forgery-sensitive area.

[0137] S47. At the end of each iteration, re-perform subgroup division based on the currently updated set of document feature vectors, adjust the subgroup structure in real time, and repeat steps S42 to S46 until the maximum number of iterations T is reached max , and finally output the document feature weight vector corresponding to the firefly individual with the highest cross-scale forgery adversarial fitness value as the optimal feature weight combination W for highly forged detection of document data * .

[0138] In this embodiment, the subgroup division is specifically based on the distribution of the document feature vector set in the high-dimensional space. A hierarchical clustering method based on the Euclidean distance between feature vectors is used to adaptively divide the firefly population, generating multiple subgroups. Each subgroup aggregates firefly individuals with similar features. When the Euclidean distance between the feature vectors of any two image format document samples is less than the dynamically set Euclidean distance threshold, they are classified into the same subgroup, so that each subgroup has common forgery risk features internally.

[0139] Based on the traditional firefly optimization algorithm, this embodiment designs a variable-structure firefly network based on the dynamic restructuring of the document structure, which can divide dynamic subgroups according to the clustering structure of document features, improving the modeling ability of the optimization algorithm for the local forgery feature distribution. At the same time, a cross-scale forgery adversarial fitness function that combines forgery enhancement and robust resistance ability is defined, and the weight optimization of the forgery-sensitive direction is introduced through the gradient guidance mechanism, so that the finally obtained feature weight combination can not only highlight the forgery traces but also improve the stability against complex interference, thus enhancing the adaptability of the entire detection system to diverse forgery patterns.

[0140] In this embodiment, S5 includes the following steps:

[0141] S51. Obtain the optimal feature weight combination for the highly forged detection of the finally output document data, and perform a weight mapping adjustment operation on the document feature vector set. Use the optimal feature weight combination W * Perform per-dimensional weighted adjustment on each document feature vector to generate an optimized document feature vector

[0142]

[0143] where represents the optimal weight value corresponding to the D-th feature dimension in the optimal feature weight combination, and f i,D represents the original document feature vector value of the i-th document data in the D-th dimension, and D is the total dimension of the document feature vector;

[0144] S53. Store all the weighted-adjusted document feature vectors in sequence to form a new document feature vector set as the optimized document feature vector set The optimized document feature vector set performs weight strengthening processing on the feature vectors of each document sample based on the forgery detection optimization target while keeping the order of the original document samples unchanged.

[0145] In this embodiment, S6 includes the following steps:

[0146] S61. Based on the optimized document feature vector set Construct a family of locality-sensitive hashing functions H that incorporates the sensitivity weights of forgery features fake :

[0147]

[0148] where represents the optimized document feature vector set the hash value after being mapped by the forgery-sensitive weight, a fake is a projection vector randomly generated from the standard normal distribution, with the same dimension D as the optimized document feature vector dimension, b fake represents an offset value randomly generated within the interval [0, r fake , r fake is the dynamic forgery feature sensitivity hashing quantization scale parameter, the symbol represents the Hadamard element-wise product operation, reflecting the optimal feature weight combination W * the feature sensitivity guidance for the hashing mapping process represents the set of real numbers consisting of all real vectors of dimension D, is the set of positive real numbers;

[0149] The formula linearly maps the optimized document feature vector after the fusion weight perturbation to generate a low-dimensional hash value for fast clustering and approximate matching.

[0150] Traditional LSH ignores the importance differences of feature dimensions, while the present invention magnifies the important feature dimensions by introducing W * operations, enabling the forgery-sensitive dimensions to dominate the mapping process, thereby clustering samples that may be forged in the same hash bucket, greatly improving the recall rate and matching efficiency, and being suitable for complex document environments with sparse forged samples and subtle feature changes.

[0151] At the same time, it does not simply apply traditional LSH, but introduces a task relevance feature weighting mechanism in the original LSH framework, making the hashing mapping process more sensitive to forgery detection and more adaptable to the forgery feature distribution, with the following advantages:

[0152] Introduce the optimal feature weights to improve the clustering ability of the hash function for forgery features;

[0153] Construct a function family with adjustable hash scales to enhance the self-adaptability of feature distributions;

[0154] Retain the original low-complexity advantage of LSH, suitable for efficient processing of large-scale document samples.

[0155] S62. For the highly forged detection task of document data, a hash function adaptive selection mechanism is constructed. According to the forged feature sensitive distribution state of the optimized document feature vector set in the low-dimensional space, the number and parameter configuration of the local sensitive hash function family H in it are dynamically adjusted, so that the hash functions are concentratedly distributed in the sensitive areas where highly forged document features are likely to gather, and the forged sensitive hash codes of the optimized document feature vectors are obtained. fake In the process, the number and parameter configuration of the hash functions in the local sensitive hash function family H are adjusted according to the forged feature sensitive distribution state of the optimized document feature vector set in the low-dimensional space, so that the hash functions are concentratedly distributed in the sensitive areas where highly forged document features are likely to gather, and the forged sensitive hash codes of the optimized document feature vectors are obtained.

[0156] S63. Based on the set of forged sensitive hash codes of all optimized document feature vectors A forged feature adaptive multi-level hash index structure is constructed to form a set of forged sensitive hash buckets B fake Each forged sensitive hash bucket is automatically indexed at multiple levels according to the forged risk level of the forged sensitive hash code, forming a multi-level hash index structure from high to low.

[0157] In this embodiment, a feature dimension forgery sensitivity weight is introduced in the construction process of local sensitive hashing. Through the Hadamard product, the optimized feature vector is fused with the weight vector dimension by dimension to form a forgery sensitive hash function family, making the hash mapping process more capable of focusing on forgery features. At the same time, the system dynamically constructs a multi-level index structure according to the risk level of the hash code, mapping the document feature vector to the structured forgery risk hierarchical space, greatly improving the clustering speed and detection rate of high-risk forged samples in large-scale data, and overcoming the problem of insufficient matching ability of traditional hash methods for complex structural forged samples.

[0158] In this embodiment, S7 includes the following steps:

[0159] S71. Perform a local similarity matching operation on the optimized document feature vectors in each hash bucket in the set of forged sensitive hash buckets, and calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library based on the forged sensitive hash codes in the set of optimized document feature vectors: In the set of forged sensitive hash buckets, perform a local similarity matching operation on the optimized document feature vectors in each hash bucket, and calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library based on the forged sensitive hash codes in the set of optimized document feature vectors: Calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library based on the forged sensitive hash codes in the set of optimized document feature vectors:

[0160]

[0161] Where Sim i,k is the similarity between the i-th document to be detected and the k-th forged sample, and represent the optimized feature vectors of the document to be detected and the forged sample respectively. If the feature cosine similarity Sim i,k ≥θ sim , where θ sim is the forged similarity threshold, then the i-th document is determined as a candidate forged sample;

[0162] S72. Incorporate the document data samples that meet the forgery similarity threshold into the candidate document data set D cand , perform a structural similarity comparison on each candidate document sample in the candidate document data set, extract the structural information of the grayscale image matrix of the candidate document and the forgery sample image, and calculate the structural similarity index SSIM:

[0163]

[0164] where x is the grayscale image matrix block of the current document sample to be detected, y is the corresponding forgery sample image matrix block, and μ x , μ y are the local means of the images respectively, is the variance, σ xy is the covariance, and C1 and C2 are stability constants to prevent the denominator from being zero; when the structural similarity index SSIM(x, y) ≥ θ ssim , where θ ssim is the image structure similarity threshold, it is determined that the image structure has high structural similarity;

[0165] S74. On the basis of the structural similarity comparison, classify and identify the key regions in the candidate document image and extract image features, evaluate whether there are forgery features such as tampering, cloning, splicing or image enhancement in the key regions of the candidate document image, and compare and verify them with the features in the forgery template library;

[0166] S75. Define the document forgery confidence score C i :

[0167] C i = ω1·Sim i + ω2·SSIM i ;

[0168] where ω1 and ω2 represent weights, SSIM i is the structural similarity index score, and Sim i is the feature cosine similarity score;

[0169] S76. Output the final highly forged detection result of the document data according to the forgery confidence score C i and the set detection result determination rule.

[0170] In this embodiment, the final highly forged detection result of the document data includes:

[0171] Normal document data: When the forgery confidence score C i < 0.4, there is no forgery risk;

[0172] Suspicious document data: When the forgery confidence score 0.4 ≤ C i <0.7, there is a forgery trend, and further manual review is recommended;

[0173] High-risk forgery document data: When the forgery confidence score C i ≥0.7, it is highly suspected of forgery, marked as a detection positive and an alignment report is output.

[0174] This embodiment designs a two-level forgery determination mechanism. First, candidate documents are screened through the optimized feature cosine similarity, then the candidate samples are subjected to structural similarity index comparison and key area forgery pattern analysis, and finally a forgery confidence score is generated based on the similarity fusion score, effectively integrating the global feature similarity and the fine discrimination ability of the local image structure, realizing a gradually enhanced forgery recognition process from preliminary screening to confirmation, not only significantly reducing the false alarm rate and missed alarm rate, but also providing high interpretability and visualization support for subsequent manual review or system automatic decision-making.

[0175] Example 1:

[0176] In a digital publishing copyright review in April 2024, a provincial publishing review unit deployed the method based on the present invention to identify possible tampered or spliced forged image content in book scans.

[0177] At 10:35 on April 13, 2024, the system received document data with the upload number "BOOK-FJ-20240413-0734". The original format of the document was PDF, the upload source identification code was SRC-804210, and the upload timestamp was 1681372405 (UTC+8).

[0178] Step 1: Standardized preprocessing

[0179] The system first identifies that the document is in PDF format, automatically calls the format conversion module to convert it to the standard image format PNG, while maintaining the integrity of the page structure and text layout. After size normalization processing, the system uniformly scales the image to a 1024×768 pixel grayscale image matrix and numbers it

[0180] Step 2: Feature extraction and optimization

[0181] After performing feature extraction on the image matrix the system generates an initial feature vector with a dimension of 128. Subsequently, the system inputs the feature vector into a variable structure firefly network, and through subgroup division, it is classified into the 3rd subgroup C3. The initial light intensity I0 = 0.9, and the maximum number of iteration rounds is set to 50.

[0182] At the 24th iteration, the individual weight vector reaches the maximum cross-scale forgery adversarial fitness f = 0.823, and the optimal weight combination W * is used to perform weighted adjustment to generate an optimized feature vector

[0183] Step 3: Hash encoding and hash bucket clustering

[0184] The system inputs the optimized feature vector into the forgery-sensitive hash function family H fake Under the action of 8 groups of dynamic hash functions, the generated hash code is:

[0185]

[0186] The encoding is matched to the forgery-sensitive hash bucket The hash bucket is the "bucket with high structural local consistency risk level".

[0187] Step 4: Similarity matching and threshold triggering

[0188] The system performs local feature comparison between the document and the document numbered "FAKE-TEMPLATE-00456" in the forgery sample library, and obtains the feature cosine similarity:

[0189] Sim 237,00456 = 0.894;

[0190] Structural similarity index:

[0191]

[0192] Since the document satisfies:

[0193] Sim i,k ≥θ sim = 0.85;

[0194] SSIM(x, y)≥θ ssim = 0.80;

[0195] The system determines it as a candidate forged document and adds it to the candidate set D cand .

[0196] Step 5: Key area forgery enhancement detection and risk score output

[0197] The system further detects the red seal area in the lower left corner of the copyright page in this document. Within the spatial coordinate region (x = 342 - 397, y = 612 - 640), image reconstruction traces are detected (local texture spectrum imbalance, pseudo-edge texture aggregation index = 0.612), and it shows a 92.1% contour similarity with the forged seal template number "STAMP-RED-0031".

[0198] The system finally outputs the forged confidence score:

[0199] C 237 = 0.6·0.894 + 0.4·0.846 = 0.875:

[0200] Since C 237 ≥ 0.7, the system automatically marks it as a "high-risk forged document" and outputs the following report content:

[0201] Detection number: DETECT-20240413-00087;

[0202] Determination level: High-risk forgery;

[0203] Similar sample: FAKE-TEMPLATE-00456;

[0204] Risk area location: Page 2 of the copyright page, red seal area;

[0205] Feature matching method: Multi-channel forgery-sensitive hash index + SSIM comparison;

[0206] Output confidence score: 0.875;

[0207] Suggested handling action: Enter the manual review system and restrict the approval process of this publication record;

[0208] Meanwhile, the system writes the detection record into the blockchain traceability platform and sends a warning to the document reviewer "User ID: Reviewer-A059" through the internal message gateway, and completes the entire process within 17 seconds.

[0209] Example 1 truly demonstrates the processing flow of the present invention in the scenario of large-scale document forgery detection, clarifies the key indicators and response chain from standardized processing, feature optimization, hash mapping, similarity comparison, local forgery analysis to decision output. The system not only makes a detection judgment with high accuracy, but also outputs a clear, traceable and reviewable judgment result, with extremely high practical value and deployability.

[0210] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing, characterized in that It includes the following steps: S1. Collect the document data to be detected and convert the collected document data into a standard format image document dataset; S2. Preprocess the standard format image document dataset to generate a preprocessed image document dataset; S3. Extract features from the preprocessed image document dataset to form a document feature vector set; S4. Apply the firefly optimization algorithm to globally search the document feature vector set, optimize the document feature weights, and obtain the optimal feature weight combination; S5. Use the optimal feature weight combination to adjust the document feature vector set to generate an optimized document feature vector set; S6. Build a local sensitive hashing matching model based on the optimized document feature vector set, map the document feature vectors into a low-dimensional space through a preset hash function family, and form multiple hash buckets; S7. Perform similarity matching on the document data in each hash bucket, screen out the candidate document data that highly matches the known forged document features, and further perform multi-level forgery confirmation on the screened candidate document data. Adopt structural similarity comparison and image recognition comparison to output the final highly forged detection result of the document data and the corresponding document forgery confidence score.

2. The method for deepfake detection based on the firefly optimization algorithm and locality-sensitive hashing according to claim 1, wherein The S1 includes the following steps: S11. Collect the original document data to be detected. The original document data includes image format documents and non-image format documents. Each piece of original document data consists of the original text format of the document, timestamp information, and document source identification code; S12. Perform format recognition on the collected original document data. For non-image format documents that do not meet the image processing requirements, call the format conversion module to convert the non-image format documents into uniformly standard image format documents. The format conversion process maintains the integrity of the document content structure and page layout, so that the converted image format documents retain the visual information corresponding to the original documents; S13. Perform standard size adjustment operations on the image format documents, uniformly scale all image format documents to the set height and width, maintain the consistency of the image aspect ratio during the size adjustment process, and perform size alignment; S14. Combine the image format document that has completed format conversion and size standardization with the corresponding timestamp information and source identification code to form the standard format image document dataset D std .

3. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 1, characterized in that The S2 includes the following steps: S21. Perform image noise removal processing on each image document in the standard format image document training set D std and perform a smoothing operation on the image pixel matrix of the image document. After the noise removal processing, a noise suppression image matrix is obtained; S22. Perform edge enhancement processing on the noise-suppressed image matrix to enhance the texture and contour features of the image edge region and obtain an edge-enhanced image matrix; S23. Perform size normalization processing on the edge-enhanced image matrix to obtain a normalized image matrix; S24. Gray-scale the normalized image matrix. If the image document is a color image, convert the image document into a single-channel gray-scale image to obtain a gray-scale image matrix And combine the gray-scale image matrix with its corresponding timestamp information t i And the source identification code s i To form a preprocessed image document data set Among them, N represents the total number of image documents in the preprocessed image document data set.

4. The deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 3, characterized in that, The S3 includes the following steps: S31. For each grayscale image matrix in the preprocessed image document dataset perform texture feature extraction, calculate the gray-level co-occurrence matrix of the local region of the image and extract texture statistics, including contrast, energy, entropy, and homogeneity, to form a texture feature vector S32. For the grayscale image matrix Apply an edge detection operator to extract the main edge contour region in the image, and perform quantization encoding on the edge direction, edge density, and edge continuity to obtain an edge contour feature vector S33. Extract the grayscale image matrix based on the arrangement rule of image pixels Extract the structural features, calculate the grayscale mean, standard deviation, and centroid coordinates of each region of the structural features, and construct a structural feature vector S34. Select the position of the key area of the page as the region of interest in the grayscale image matrix Extract the key text blocks or seal areas containing forgery traces in the image according to the preset layout template, calculate the relative positions, grayscale statistics and texture complexities of the key text blocks or seal areas containing forgery traces, and construct the key area distribution feature vector S35. Combine the texture feature vector with the edge contour feature vector and the structure feature vector through feature-level concatenation operation with the key area distribution feature vector to obtain the final feature vector of the i-th image document, and construct a set of the final feature vectors of all image document samples to form a document feature vector set 5. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 4, characterized in that, The S4 includes the following steps: S41. Initialize the parameters of the firefly population, set the initial population size as M, the maximum number of iterations as T max , the initial light intensity as I0, the light intensity attenuation factor as γ, the initial attractiveness as β0, and use the document feature vector set as the initial search space of the firefly algorithm; S42. Build a variable structure firefly network to adaptively divide the firefly population into subgroups; S43. For each subgroup divided, define the cross-scale forgery adversarial fitness function of the firefly individuals. The cross-scale forgery adversarial fitness function combines the cross-scale enhancement and robustness resistance of the firefly individual positions and the highly forged risk features of the document features: Among them, represents the cross-scale forgery adversarial fitness value of the j-th firefly individual in the c-th subgroup, is the corresponding document feature weight vector, represents the forgery risk feature sensitivity enhancement value under the current weight, that is, the "forgery enhancement ability", represents the forgery risk feature robustness resistance ability under cross-scale perturbation, that is, the "robustness ability", and α and λ are weight factors; S44. The cross-scale forgery adversarial fitness value is used as the basis for updating the light intensity. The light intensity of each firefly individual is adjusted exponentially according to the light intensity fitness value. The higher the fitness value, the slower the light intensity attenuation, retaining a stronger attraction ability to guide other firefly individuals to approach the current firefly individual. For firefly individuals with a low light intensity fitness value, the light intensity decreases rapidly, and the attraction weakens during the search process, promoting the rapid convergence of the global optimal solution; S45. Based on the difference in the light intensity of firefly individuals, calculate the mutual attraction degree between firefly individuals within the subgroup: Among them, represents the attractiveness of the j-th firefly individual in the c-th subgroup to the k-th firefly individual, represents the Euclidean distance between the positions of the j-th and k-th firefly individuals, is the degree of difference in fitness between two firefly individuals, making the movement between firefly individuals tend to the firefly individuals with more forgery resistance advantages; S46. Adopt a firefly individual position update strategy that fuses cross-scale forgery adversarial gradients and attraction degree information to dynamically optimize the positions of firefly individuals. Define the updated position of a firefly individual as: Among them, and respectively represent the updated document feature weight position and the pre-updated document feature weight position of the j-th firefly individual in the c-th subgroup. η is the cross-scale forgery adversarial gradient guidance coefficient. represents the gradient vector of the cross-scale forgery adversarial fitness function at the current position. ξ is the random perturbation factor, and rand is the random vector. S47. At the end of each iteration, subgroup division is re - performed based on the currently updated set of document feature vectors, the subgroup structure is adjusted in real - time, and steps S42 to S46 are repeatedly executed until the maximum number of iterations T is reached max , and finally, the document feature weight vector corresponding to the firefly individual with the highest cross - scale forgery - resistant fitness value is output as the optimal feature weight combination W for highly forged detection of document data * .

6. The deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 5, characterized in that The subgroup division is specifically based on the distribution of the document feature vector set in the high-dimensional space. An hierarchical clustering method based on the Euclidean distance between feature vectors is used to adaptively divide the firefly population to generate multiple subgroups. Each subgroup aggregates firefly individuals with similar features. When the Euclidean distance between the feature vectors of any two image format document samples is less than the dynamically set Euclidean distance threshold, they are classified into the same subgroup, so that each subgroup has common forgery risk features internally.

7. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 5, characterized in that, The said S5 includes the following steps: S51. Obtain the optimal feature weight combination for highly forged detection of the finally output document data, and perform a weight mapping adjustment operation on the document feature vector set, and adopt the optimal feature weight combination W * Perform per-dimensional weighted adjustment on each document feature vector to generate an optimized document feature vector Among them, represents the optimal weight value corresponding to the D-th feature dimension in the optimal feature weight combination, and f i,D represents the original document feature vector value of the i-th document data in the D-th dimension, where D is the total dimension of the document feature vector; S53. Store all the weighted and adjusted document feature vectors in sequence to form a new set of document feature vectors, which serves as the optimized set of document feature vectors. On the basis of keeping the original document sample order unchanged, the optimized set of document feature vectors performs weight strengthening processing on the feature vectors of each document sample based on the forgery detection optimization objective.

8. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 7, characterized in that The said S6 includes the following steps: S61. Based on the optimized set of document feature vectors Construct a family of locality-sensitive hashing functions H that incorporates the sensitivity weights of forgery features fake : Among them, represents the optimized document feature vector set the hash value after being mapped by the forged sensitive weight, a fake is a projection vector randomly generated from the standard normal distribution, with the same dimension D as the optimized document feature vector, b fake represents the offset value randomly generated within the interval [0, r fake , r fake is the dynamic forged feature sensitivity hash quantization scale parameter, symbol represents the Hadamard element-wise product operation, reflecting the optimal feature weight combination W * guides the feature sensitivity of the hash mapping process, represents the real number set composed of all real vectors with dimension D, is the set of positive real numbers; S62. For the highly forged detection task of document data, a hash function adaptive selection mechanism is constructed. According to the forged feature sensitive distribution state of the optimized document feature vector set in the low-dimensional space, the number and parameter configuration of the local sensitive hash function family H in fake are dynamically adjusted, so that the hash functions are concentratedly distributed in the sensitive areas where highly forged document features are likely to gather, and the forged sensitive hash codes of the optimized document feature vectors are obtained. fake ​ Forgery-sensitive hash code set based on all optimized document feature vectors Construct a forgery feature adaptive multi-level hash index structure to form a forgery-sensitive hash bucket set B fake , and each forgery-sensitive hash bucket is automatically hierarchically indexed according to the forgery risk level of the forgery-sensitive hash code, forming a multi-level hash index structure from high to low.

9. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 8, characterized in that The said S7 includes the following steps: S71. Perform a local similarity matching operation on the optimized document feature vectors in each hash bucket in the forged sensitive hash bucket set, and calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library based on the forged sensitive hash codes in the optimized document feature vector set within the optimized document feature vectors in each hash bucket Calculate the feature cosine similarity between the current document feature vector to be detected and the vectors in the forged document feature library Among them, Sim i,k is the similarity between the i-th document to be detected and the k-th forged sample, and respectively represent the optimized feature vectors of the document to be detected and the forged sample. If the feature cosine similarity Sim i,k ≥θ sim , where θ sim is the forged similarity threshold, then the i-th document is determined to be a candidate forged sample; S72. Incorporate the document data samples that meet the forgery similarity threshold into the candidate document data set D cand , perform structural similarity comparison on each candidate document sample in the candidate document data set, extract the structural information of the grayscale image matrix of the candidate document and the forgery sample image, and calculate the structural similarity index SSIM: where x is the grayscale image matrix block of the current document sample to be detected, y is the corresponding forged sample image matrix block, and μ x , μ y are the local means of the images respectively, is the variance, σ xy is the covariance, and C1 and C2 are stability constants to prevent the denominator from being zero; when the structural similarity index SSIM(x, y) ≥ θ ssim , where θ ssim is the image structural similarity threshold, it is determined that the image structure has high structural similarity; S74. On the basis of structural similarity comparison, classify and identify key regions in the candidate document image and extract image features, evaluate whether there are forgery features such as tampering, cloning, splicing, or image enhancement in the key regions of the candidate document image, and compare and verify them with the features in the forgery template library; S75. Define the confidence score C of document forgery based on the comprehensive structural similarity index and the characteristic cosine similarity i : C i = ω1·Sim i + ω2·SSIM i ; Among them, ω1 and ω2 represent weights, and SSIM i is the structural similarity index score, and Sim i is the feature cosine similarity score; S76. Output the final detection result of highly forged document data according to the forged confidence score C i Output the final detection result of highly forged document data according to the set detection result determination rule 10. A deepfake detection method based on the firefly optimization algorithm and locality-sensitive hashing according to claim 9, characterized in that, The final highly forged detection result of the document data includes: Normal document training: When the forgery confidence score C i < 0.4, there is no forgery risk; Suspicious document data: When the forgery confidence score 0.4 ≤ C i < 0.7, there is a forgery trend, and further manual review is recommended; High-risk forged document data: When the forgery confidence score C i ≥ 0.7, it is highly suspected of forgery, marked as a positive detection, and a comparison report is output.

Citation Information

Cited By

  • Credit material repeated detection method and device based on image fingerprints

    CN120876079A

  • Multi-modal power data retrieval method and device and medium

    CN120929662A