Paper duplicate checking method based on improved OfficientNet network

By improving the EfficientNet network, combining the Triplet Semi-Hard loss function and the Adaptive Layout Optimization module (ALR), dynamically selecting samples and introducing improved convolutional layers and genetic layers, the problems of low efficiency and low accuracy in detecting paper image tampering in existing technologies are solved, and an efficient and accurate paper duplication checking method is implemented.

CN120612583APending Publication Date: 2025-09-09SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510759141.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing deep learning-based methods for detecting image tampering in papers have low detection efficiency and accuracy under complex conditions, and have high requirements for image quality and consume large computing resources. They are unable to meet the needs of efficiently and accurately judging the degree of tampering of target paper images.

Method used

An improved EfficientNet network is used, combined with the Triplet Semi-Hard loss function and the Adaptive Layout Optimization module (ALR), to dynamically select anchor samples, positive samples, and negative samples. Gradient updates are adjusted through dynamic weights, and improved convolutional layers, HOG modules, and genetic layers are introduced to increase feature robustness. The TransNeXt and RepViT modules are introduced to improve feature extraction, and Euclidean distance is used for similarity measurement.

Benefits of technology

It improves the efficiency and accuracy of target tampering detection and recognition under complex conditions, improves the working efficiency and recognition accuracy of the paper duplication checking system, and enhances the robustness of the model and the efficiency of computing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612583A_ABST
    Figure CN120612583A_ABST
Patent Text Reader

Abstract

The invention discloses a paper duplicate checking method based on an improved OfficientNet network. The method comprises the following steps: constructing a query picture total library; a paper duplicate checking training sample is rapidly constructed through a picture semi-automatic extraction classification and synthesis technology; an improved Triplet Semi-Hard loss function and an adaptive layout optimization module ALR are introduced to improve a high-level semantic layout optimization module and a feature extraction and embedding module of a traditional OfficientNet model; an input layer and an output layer of the OfficientNet network are improved; the method comprises the following steps of: introducing a TransNeXt module and a RepViT module to improve a Block4 module in an OfficientNet; performing paper duplicate checking based on Euclidean distance picture similarity measurement; according to the method, the precision and accuracy of a paper duplicate checking system for extracting pictures in papers are remarkably improved, and optimization, accuracy and efficiency of target picture detection tasks are achieved especially under the challenging conditions that images are difficult to recognize and extract, picture content information features are not obvious, picture types are diversified and complex and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and target detection and recognition, and specifically to a paper duplication checking method based on an improved EfficientNet network. Background Art

[0002] Image object detection and recognition technology is a crucial component of paper duplication checking. With the global popularization of higher education and the rise of interdisciplinary research, the number of paper submissions has surged. Against this backdrop, the task of identifying and checking for duplicate content has become increasingly arduous and critical, placing a greater strain on the reliability and efficiency of duplicate checking systems. Therefore, to improve the accuracy and efficiency of image duplication checking, the design and implementation of systems based on image similarity and tampering detection algorithms are evolving towards smarter and more efficient approaches. This has made the design and implementation of algorithm-based systems a new direction in the field.

[0003] With the increasing demand for plagiarism detection in papers, traditional manual detection methods are no longer sufficient for the task of detecting and identifying images cited in papers. The efficiency of this task directly impacts the overall quality and evaluation of academic research, and thus indirectly affects the teaching and research level of higher education. In recent years, the rapid development of deep learning technology, particularly the widespread application of convolutional neural networks in image processing, has significantly improved the efficiency and accuracy of image detection and recognition in papers. Leveraging these technologies, plagiarism detection systems can accurately extract image features from complex papers with less distinct image features and perform repeatability comparisons with existing databases of paper images to determine whether the images have been tampered with. This approach significantly improves the efficiency of plagiarism detection and reduces the risk of human error. However, in practical applications, traditional deep learning-based methods for detecting tampered images in papers still have many shortcomings due to issues such as low resolution of uploaded images and difficulty identifying feature points.

[0004] Regarding feature-point-based image tampering detection methods, Yang et al. (Yang, B., Sun, X., Guo, H. et al. A copy-move forgery detection method based on CMFD-SIFT [J]. Multimedia Tools and Applications, 2018, 77:837–855.) proposed a feature-based copy-move forgery detection (CMFD) method. This method uses a modified SIFT (Scale Invariant Feature Transform)-based detector to detect keypoints. A keypoint distribution strategy was developed to evenly distribute the keypoints throughout the image. Finally, the keypoints were described using a modified SIFT descriptor enhanced for CMFD scenarios. However, this method places high demands on the feature extraction algorithm: a suitable feature extraction algorithm must be selected to ensure that the extracted features are representative and accurate. It also places high demands on image quality: poor image quality or the presence of interference can adversely affect the feature extraction and matching results. In addition, for image tampering detection methods based on background noise characteristics, Lu et al. (LU Yan-Fei, JU Ya-Li, YU Yue. Image forgery detection using characteristics of background noise [J]. Journal of Signal Processing, 2012, 28 (9): 1299-1307.) proposed an algorithm for estimating local noise in an image, which detects the part with abnormal noise as the tampering area. The algorithm uses a quadratic noise addition method based on the high-order statistical characteristics of the image signal to blindly estimate the variance of the background noise, and calculates the background noise variance estimate of each block by dividing the image into adjacent overlapping blocks, and then intuitively gives the abnormal noise part to locate the image forgery area. It has good robustness for both image scaling and compression. However, it has many shortcomings, and the detection effect is poor for tampered images with small resolution and forged images that are not spliced.

[0005] For image tampering detection methods based on deep learning, Liu et al. (Liu, Y., Guan, Q., Zhao, X., et al. Image forgery localization based on multi-scale convolutional neural networks [J]. Proceedings of the 6th ACM Workshop on Information Hiding and Multimedia Security, 2018, 85-90.) proposed a framework for tampering region localization based on a multi-scale method. The model introduces the idea of ​​multi-scale, designs image sliding windows of different scales, and extracts multi-scale image block features. However, when faced with different processing of different images, such as after image denoising, enhancement, restoration, and compression techniques, it will become very difficult to extract tampering region features. In addition, for the image tampering detection method based on the attention mechanism, Yu Chen et al. (Yu Chen, Fu Ying, Zhu Ye. Image tampering detection based on attention mechanism [J]. Computer Science and Application, 2022, 12: 729.) proposed an image tampering detection network based on the attention mechanism, using GAN to generate tampered images to solve the problem of lack of training data, while expanding the sample size to avoid model collapse, and realizing the segmentation of image tampering areas. Compared with the current generative adversarial network, it introduces a new objective function, takes images from the existing tampering dataset as the input of GAN, optimizes it through the objective function, and then mixes it with the original tampered image dataset, thereby achieving the effect of expanding the dataset and optimizing the authenticity of the tampered area. However, due to the large number of network model parameters, it requires high computing power of the computer. Due to the complexity of the model, the detection accuracy of its algorithm is not high enough.

[0006] In summary, although deep learning-based technology for identifying altered images in academic papers has made some progress, it is still necessary to better meet the needs of efficiently and accurately determining the extent of alteration in target academic papers under complex conditions. This will improve the efficiency and accuracy of the entire academic paper checking system, and promote the overall quality and research level of the academic field. Summary of the Invention

[0007] This paper aims to address the limitations of the above existing technologies and proposes a paper duplication detection method based on an improved EfficientNet network to improve the recognition efficiency, accuracy and robustness of target tampering detection in complex paper images.

[0008] In order to achieve the above invention, the present invention provides a method for checking for duplicate papers based on an improved EfficientNet network, comprising the following steps:

[0009] (1) Build a query image database;

[0010] (2) Semi-automatic image extraction, classification, and synthesis technology to quickly construct training samples for paper duplication checking;

[0011] (3) An improved Triplet Semi-Hard loss function and an adaptive layout optimization module ALR are introduced to improve the high-level semantic layout optimization module and feature extraction and embedding module of the traditional EfficientNet model: on the basis of the fixed margin of the traditional TripletLoss, the dynamic sample selection strategy is combined to dynamically select anchor samples, positive samples and negative samples; the dynamic sample selection strategy refers to adaptively enhancing the optimization weight of difficult samples and reducing the optimization weight of easy samples during the training process, that is, constructing a dynamic weighting factor based on the loss contribution of the sample triplet to adjust the gradient update weight during the back propagation process; specifically, firstly, a dynamic discrimination of difficult and easy samples is performed. In each forward propagation, the Euclidean distance d1 between the anchor sample and the positive sample and the Euclidean distance d2 between the anchor sample and the negative sample are calculated, and the difficult sample condition is defined as d2-d1<α+β×σ, where α is the fixed margin threshold, β is the adaptive relaxation coefficient, and σ is the dynamic standard deviation of the distance difference of the current batch of samples, which is updated in real time through sliding window statistics; then, a nonlinear scaling function is introduced for difficult samples through dynamic mapping of gradient weights. Among them L hard is the loss value of the current difficult sample, γ is the global weight factor, λ controls the steepness of the weight distribution, ∈ is a smoothing term to prevent the gradient from disappearing, and then back-propagation is performed to reweight the dynamic weight w and the original gradient tensor element by element, so that the gradient amplitude of the difficult sample is amplified by w, and the gradient explosion risk is constrained by gradient clipping; the dynamic selection of anchor samples, positive samples and negative samples refers to distinguishing difficult samples from easy samples by setting dynamic thresholds, and the difficult sample refers to the Euclidean distance d between the negative sample and the anchor sample. n Satisfy d p <d n ≤d p +γ; the easy sample refers to the Euclidean distance d between the negative sample and the anchor sample n Satisfy d n >d p +γ, where d pis the Euclidean distance between the positive sample and the anchor sample, γ is the preset margin threshold, and γ is set to 0.2; the introduction of the improved adaptive layout optimization module ALR for deep integration with EfficientNet-B0 refers to the introduction of dynamic convolution kernel in the ResBlock of the adaptive layout optimization ALR module of the EfficientNet-B0 network, according to the text-visual correlation matrix TVM matrix Q i-1 Generate convolution kernel weights to make the convolution operation adaptive to the semantic focus area of ​​the text description. The original formula is Conv(H i-1 )=H i-1 ×K+b, improved to Conv(H i-1 )=Conv(H i-1 ; θ=f(Q i-1 )); Then, the semantic similarity matrix SSM and the text-visual association matrix TVM are dynamically constructed. The dynamic construction refers to the Q based on TVM in the residual block of the ALR module. i-1 Generate convolution kernel parameters and dynamically allocate weights through the gated attention mechanism. The formula is θ = Soft max (W q ×Q i-1 )×W k , where W q , W k The projection matrix can be learned to make the convolution kernel focus on the key semantic area of ​​the text description; the semantic correspondence between the local area of ​​the image and the text description is analyzed in real time, and the layout optimization weight is adaptively adjusted based on the residual features; the adaptive adjustment of the layout optimization weight refers to constructing the regional optimization weight through the residual R = |Θ-Θ'| of the semantic similarity matrix (SSM) Θ and the real image SSMΘ' ij =Sigmoid(φ(R ij ⊙H i-1 )), where φ is a multi-layer perceptron that implements weight reinforcement in high residual regions (where layout deviation is significant); in the multi-scale feature pyramid of EfficientNet-B0, according to the KL divergence of the layout error at the current level, KL (Θ l ||Θ l ') Dynamically adjust the learning rate of each layer: Where k is the sensitivity coefficient, η base As the basic learning rate, it realizes the differentiated optimization of deep semantic features and shallow texture features;

[0012] (4) Improve the input and output layers of the EfficientNet network: replace the convolution layer of the input layer with two improved convolution layers and introduce an improved HOG module and add a genetic layer between the pooling layer and the fully connected layer of the output layer. The improved convolution layer refers to using the improved Min-Max normalization method to improve the convolution layer when processing data in each convolution layer. The improved Min-Max normalization method refers to using the median and interquartile range instead of the maximum and minimum values ​​for normalization. The interquartile range refers to the number obtained by subtracting the first quartile from the third quartile, and then the interquartile range and the median are used to normalize the original data. The introduction of the improved HOG module refers to using SRRF to extract feature points from all blocks of HOG, and retaining blocks with a feature point ratio of more than 10%, and then performing subsequent HOG processing on the retained blocks to obtain feature vectors. The genetic layer refers to a genetic layer constructed based on the GA genetic algorithm, that is, performing crossover, mutation, and genetic operations on the feature map encoding to obtain a new feature map encoding, thereby obtaining a new genetic layer to increase feature robustness and avoid overfitting.

[0013] (5) Introducing the TransNeXt module and the RepViT module to improve the Block4 module in EfficientNet: The Block4 module in EfficientNet is replaced by a module consisting of a convolutional layer, a depthwise separable convolutional layer, a TransNeXtBlock layer, a RepViTBlock layer, and a convolutional layer in series; the TransNeXtBlock layer refers to combining the aggregated attention mechanism and the Convolutional Gated Linear Unit and adding a feedforward neural network module, so that the model can effectively extract and weight features at local and global scales; the RepViTBlock layer refers to using 3×3 convolution to deepen the network depth, and then adding a feedforward neural network module to strengthen feature expression, so that the model can effectively extract long-range features of the image and enhance extraction efficiency and accuracy;

[0014] (6) Paper duplication check based on Euclidean distance image similarity measurement: Use the improved EfficientNet network to extract the features of each image in image set A and image set B. If the Euclidean distance between the features of image set B and the features of image set A is less than 0.2, it is considered that the PDF paper to be queried has similarities with the papers in the 5,000 papers.

[0015] Furthermore, the present invention provides a method for checking for duplicate papers based on an improved EfficientNet network, characterized in that the construction of a total image query library refers to collecting 5,000 papers in PDF format, and using Python's PyMuPDF library to collect all the collected PDF papers and convert them into pictures, where each page of PDF corresponds to a picture, to obtain a picture set A, and then converting the PDF papers to be queried into pictures using Python's PyMuPDF library to obtain a picture set B, and the total library is composed of picture set A and picture set B.

[0016] Furthermore, the semi-automatic image extraction, classification and synthesis technology provided by the present invention quickly constructs paper duplication checking training samples, which means first using the PyMuPDF library to batch extract paper PDF images, using the OpenCV library to preprocess them into a 380×380 size format, manually annotating 500 benchmark images, and then using the EfficientNet network training to obtain a preliminary model to complete the automatic annotation of the data set; then, manually correct the low-confidence images, and iterate the model 10 times to obtain a high-precision classification model; further, for the same label atlas, the Vision Transformer architecture is used to train the Lora model, and it is applied to Stable Diffusion and random noise is introduced to generate 3,000 semantically consistent images that show high diversity in multiple details.

[0017] Furthermore, the present invention provides an improved Triplet Semi-Hard loss function and an adaptive layout optimization module ALR to improve the high-level semantic layout optimization module and feature extraction and embedding module of the traditional EfficientNet model, which means that on the basis of the fixed value of the margin of the traditional Triplet Loss, the dynamic sample selection strategy is combined to dynamically select anchor samples, positive samples and negative samples; the dynamic sample selection strategy refers to adaptively enhancing the optimization weight of difficult samples and reducing the optimization weight of easy samples during the training process, that is, constructing a dynamic weighting factor based on the loss contribution of the sample triplet to adjust the gradient update weight during the back propagation process, first performing dynamic discrimination of difficult and easy samples, and in each forward propagation, calculating the Euclidean distance d between the anchor sample and the positive sample. ap , the Euclidean distance d between the anchor sample and the negative sample an , define the difficult sample condition as d an -d ap <α+β×σ, where α is the fixed margin threshold, β is the adaptive relaxation coefficient, and σ is the dynamic standard deviation of the distance difference of the current batch of samples, which is updated in real time through sliding window statistics; then, a nonlinear scaling function is introduced for difficult samples through dynamic mapping of gradient weights. Among them L hardis the loss value of the current difficult sample, γ is the global weight factor, λ controls the steepness of the weight distribution, ∈ is a smoothing term to prevent the gradient from disappearing, and then back-propagation is performed to reweight the dynamic weight w and the original gradient tensor element by element, so that the gradient amplitude of the difficult sample is amplified by w, and the gradient explosion risk is constrained by gradient clipping; the dynamic selection of anchor samples, positive samples and negative samples refers to distinguishing difficult samples from easy samples by setting dynamic thresholds, and the difficult sample refers to the Euclidean distance d between the negative sample and the anchor sample. n Satisfy d p <d n ≤d p +γ; the easy sample refers to the Euclidean distance d between the negative sample and the anchor sample n Satisfy d n >d p +γ, where d p is the Euclidean distance between the positive sample and the anchor sample, and γ is the preset margin threshold experimentally verified to be 0.2; the introduction of the improved adaptive layout optimization module ALR for deep integration with EfficientNet-B0 refers to the introduction of dynamic convolution kernel in the ResBlock of the adaptive layout optimization ALR module of the EfficientNet-B0 network, according to the TVM matrix Q i-1 Generate convolution kernel weights to make the convolution operation adaptive to the semantic focus area of ​​the text description. The original formula is Conv(H i-1 )=H i-1 ×K+b, improved to Conv(H i-1 )=Conv(H i-1 ; θ=f(Q i-1 )); Then, the semantic similarity matrix SSM and the text-visual association matrix TVM are dynamically constructed. The dynamic construction refers to the residual block of the ALR module based on the text-visual association matrix (TVM) Q i-1 Generate convolution kernel parameters and dynamically allocate weights through the gated attention mechanism. The formula is θ = Soft max (W q ×Q i-1 )×W k , where W q , W k The projection matrix can be learned to make the convolution kernel focus on the key semantic area of ​​the text description; the semantic correspondence between the local area of ​​the image and the text description is analyzed in real time, and the layout optimization weight is adaptively adjusted based on the residual features; the adaptive adjustment of the layout optimization weight refers to constructing the regional optimization weight through the residual R = |Θ-Θ'| of the semantic similarity matrix (SSM) Θ and the real image SSMΘ' ij =Sigmoid(φ(R ij ⊙H i-1)), where φ is a multi-layer perceptron that implements weight reinforcement in high residual regions (where layout deviation is significant); in the multi-scale feature pyramid of EfficientNet-B0, according to the KL divergence of the layout error at the current level, KL (Θ l ||Θ l ') Dynamically adjust the learning rate of each layer: Where k is the sensitivity coefficient, which realizes the differentiated optimization of deep semantic features and shallow texture features.

[0018] Furthermore, the improved input layer and output layer of the EfficientNet network provided by the present invention refers to replacing the convolution layer of the input layer with two improved convolution layers and introducing an improved HOG module and adding a genetic layer between the pooling layer and the fully connected layer of the output layer. The improved convolution layer refers to using an improved Min-Max normalization method to improve the convolution layer when processing data in each convolution layer. The improved Min-Max The normalization method refers to using the median and interquartile range instead of the maximum and minimum values ​​for normalization. The interquartile range refers to the number obtained by subtracting the first quartile from the third quartile, and then the interquartile range and the median are used to normalize the original data; the introduction of the improved HOG module refers to using SRRF to extract feature points from all blocks of HOG, and retaining blocks with a feature point ratio exceeding 10%, and then performing subsequent HOG processing on the retained blocks to obtain feature vectors; the genetic layer refers to a genetic layer constructed based on a genetic algorithm, that is, performing crossover, mutation, and genetic operations on the feature map encoding to obtain a new feature map encoding, thereby obtaining a new genetic layer to increase feature robustness and avoid overfitting.

[0019] Furthermore, the present invention provides an improvement on the Block4 module in EfficientNet by introducing the TransNeXt module and the RepViT module, which means replacing the Block4 module in EfficientNet with a module obtained by connecting a convolutional layer, a depth-separable convolutional layer, a TransNeXtBlock layer and a convolutional layer in series, and adding a RepViTBlock layer; the improved TransNeXtBlock layer refers to combining the attention mechanism and the convolution GLU and adding an FFN module, so that the model can effectively perform feature extraction and weighting at local and global scales; the RepViTBlock layer refers to deepening the network depth with a 3×3 convolution, and subsequently adding an FFN module to strengthen feature expression, so that the model can effectively extract long-distance features of the image and enhance extraction efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1is a flow chart of an embodiment of the present invention;

[0021] Figure 2 This is a flowchart of semi-automatic classification and synthesis of images according to an embodiment of the present invention;

[0022] Figure 3 The EfficientNet model architecture of an embodiment of the present invention;

[0023] Figure 4 This is a diagram of the improved input layer structure of an embodiment of the present invention;

[0024] Figure 5 This is a structural diagram of an improved GA module according to an embodiment of the present invention;

[0025] Figure 6 This is a structural diagram of the improved Block4 module in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The technical approach of the present invention is further described in detail below through specific embodiments and drawings.

[0027] like Figure 1-6 As shown, the present invention is based on an embodiment of a method for checking duplicate papers using an improved EfficientNet network, comprising the following steps:

[0028] 1. Semi-automatic image extraction, classification, and synthesis technology to quickly construct training samples for paper duplication detection

[0029] In order to overcome the low efficiency of image annotation and classification caused by manual annotation and classification and the long cycle of image sample formation in papers, this patent proposes a semi-automated pipelined batch annotation and classification of images using a python library combined with EfficientNet, and uses ViT-LoRA training based on the classified image set to obtain image semantic vectors and integrate them into StableDisffusion to achieve large-scale target image synthesis.

[0030] Specifically, the PyMuPDF library is used to traverse the paper PDF collection, parse the pages to locate image elements, and extract the original resolution images. Then, the OpenCV library is used to preprocess the images, setting image resolution and blur detection as filtering conditions to exclude blurry images with a width or height less than 64 pixels and a DPI lower than 150.

[0031] The fuzzy filtering conditions are defined as follows:

[0032]

[0033] in Represents the second-order derivative of the image grayscale matrix, and σ is calculated by the Laplacian operator 2 is the variance value, and the threshold is set to 100. 2 When the value is less than 100, the image is judged as blurred.

[0034] We fine-tune the extracted images, combining adaptive padding and cropping strategies to precisely calibrate them to 380×380 pixels, filtering out corrupted samples with unusual image sizes and improving the integrity and accuracy of the image data. This design helps quickly acquire target image data from paper PDFs and preprocess them into the required format, significantly accelerating the accumulation of sample sets.

[0035] Subsequently, 500 benchmark images were manually labeled and classified, of which those containing only text were classified as the first category and those containing only images were classified as the second category. The labeled and classified dataset was divided into training, validation, and test sets according to a 7:2:1 ratio. The labeled and classified data was then normalized. The image classification model was then trained using the EfficientNet network to predict pseudo labels. Specifically, the parameters of the first 20 layers were frozen and only the top classification head was trained. The loss function used a cross-entropy-center loss hybrid objective function, and 10 rounds of active learning iterations were implemented. The hybrid loss function, which combines cross entropy and center loss with dynamic weights, is specifically expressed as:

[0036] L total (y) = λ ce (t)·L ce +λ center (t)×L center (2)

[0037] Among them L total (t) is the total loss value of the t-th training, λ ce (t) is the dynamic weight coefficient of the cross entropy loss in the tth round, λ center (t) is the dynamic weight coefficient of the center loss in the tth round. Initially, the classification task is the main task, and λ is set ce (0) = 0.8, λ center (0)=0.2 As the training round t increases, the cross entropy weight is gradually reduced and the center loss weight is increased. The formula is expressed as:

[0038] λ ce (t)=0.8-0.05×t (3)

[0039] λ center (t)=0.2+0.05×t (4)

[0040] The L ce is the cross entropy loss, which measures the difference between the predicted probability distribution and the true label. The formula is expressed as:

[0041]

[0042] Where N is the number of samples in a batch and C is the total number of categories. i,c Is whether the sample i belongs to the category value 0 or 1. i,c is the probability that the model predicts that sample i belongs to category c.

[0043] The L center It is the center loss, which constrains the features of similar samples to gather toward the center of the class and reduces the distance within the class. The formula is expressed as:

[0044]

[0045] where x i is the eigenvector of sample i, C yi is category y i The feature center of . It is the square of the L2 norm, reflecting the degree of deviation between the sample features and the class center. The smaller the value, the more compact the features within the class.

[0046] For images where the model predicted the highest category with a confidence score of less than 0.6, the incorrectly labeled categories were manually corrected and re-added to the training set to update the model parameters to further optimize model performance. During each iteration, the confidence score was increased by 0.03. After 10 complete iterations, a high-performance, high-precision image classification model was obtained, enabling automated labeling and classification of large image collections, ultimately forming a semi-automated classification system capable of processing large quantities of images. This reduced the reliance on manual labor for image labeling and classification, significantly improving efficiency while maintaining high image classification quality.

[0047] For similar image sets, the ViT-LoRA model is used to train image semantic vectors and import them into stable diffusion to synthesize images in batches. The so-called ViT-LoRA is a model generated by fine-tuning lora based on the Vision Transformer architecture. Specifically, in the self-attention module of the Vision Transformer, low-rank adapters are injected into the Key (K) and Value (V) matrices to more efficiently control feature interactions. The loss function uses a weighted method of cross entropy and contrast loss to fully explore image features and obtain image semantic vectors. The original order weight matrix is ​​K∈R d×d ,V∈R d×d , after low-rank decomposition, it is expressed as:

[0048]

[0049]

[0050] Among them A k ∈R d×r is the left half of the low-rank adapter matrix of the Key matrix, B k ∈R r×d is the right half of the low-rank adapter matrix of the Key matrix, A v ∈R d×r is the left half of the Value matrix low rank adapter matrix, B v ∈R r×d It is the right half matrix of the Value matrix low-rank adapter, r is the rank, α is the scaling factor, d is the embedding dimension, and only the A / B matrix is ​​updated during training. The original K / V parameters are frozen, and the A matrix is ​​initialized to a Gaussian distribution and the B matrix is ​​initialized to a zero matrix.

[0051] The loss function can fully exploit the image features. The specific expression is:

[0052] L total =λ ce ×L ce +λ ce ×L cl (9)

[0053] where λ ce ,λ ce is the weight coefficient L ce is the cross entropy loss and L cl is the contrast loss, and L ce and L cl The expression is:

[0054]

[0055]

[0056] Where N is the number of samples in a batch. C is the total number of categories. i,c Indicates whether sample i belongs to category c, p i,c is the model's predicted probability that sample i belongs to category c, h i is the feature vector of sample ii, p(i) is the set of all positive samples with the same label as sample i. A(i) is the total number of samples in the batch, and τ is used to adjust the sharpness of the similarity distribution.

[0057] Finally, the trained ViT-LoRA model is converted into a format compatible with the Diffusers library and imported into StableDiffusion. In the generation phase, the Lora weights are embedded in the prompt words to guide image generation. Random noise and local redrawing strategies are introduced to make the generated images more diverse and generate a sufficient number of augmented samples.

[0058] The random noise and local redrawing can make the synthesized image simulate tampering compared with the original image. The tampering is expressed as follows:

[0059] Z noisy =Z original +∈×N(0,σ 2 ) (12)

[0060] Z original is the original latent variable generated by Stable Diffusion.

[0061] Where N(0,σ 2 ) has a mean of 0 and a variance of σ 2 Gaussian noise, σ 2 Take 0.05. ∈ is the noise intensity coefficient. The local redrawing formula is expressed as follows:

[0062] Z repaired =(1-M mask )⊙Z noisy +M mask ⊙Z generated (13)

[0063] It's M mask is a randomly generated binary mask with an occlusion ratio of 20%, and ⊙ is an element-by-element multiplication. generated This is the latent variable repaired by the Diffusion model. By introducing perturbations through noise, the determinism of the generation path is broken, allowing samples with different visual representations to be generated under the same semantic conditions. This prevents the model from falling into a local optimum. The repeated generation of highly similar images introduces a local redrawing strategy, which refines content editing, generating different local variations of the same object while maintaining overall semantic consistency. This can simulate tampering with images in research papers.

[0064] Finally, the original image and the synthesized image are divided and saved in a suitable format for subsequent model training.

[0065] 2. Introduce the adaptive layout optimization module (ALR module) and Triplet Semi-Hard loss function to improve the high-level semantic layout optimization module and feature extraction and embedding module of the traditional EfficientNet model

[0066] To overcome the lack of generalization capabilities in existing image duplication detection methods due to static optimization strategies, as well as the low sensitivity of traditional feature extraction models to local tampering, this patent proposes a feature extraction and embedding module based on the improved EfficientNet-B0 architecture and the deep integration of adaptive layout refinement (ALR) technology, and the introduction of the Triplet Semi-Hard Loss function to improve the EfficientNet model for plagiarism detection of large numbers of images. By dynamically adjusting the learning rate (ALR) and optimizing the image distance relationship in the embedding space, the model's ability to distinguish images with similar features is significantly improved, enhancing the accuracy and robustness of duplicate detection.

[0067] The core architecture of this technology is based on the EfficientNet-B0 model and integrates the Adaptive Layout Refinement (ALR) module and the Triplet Semi-Hard Loss function. The overall process includes the following steps: data preprocessing and enhancement by standardizing and dynamically enhancing the NFT image; feature extraction by extracting multi-scale features of the image using EfficientNet-B0; dynamic adjustment of the learning rate of model parameters; optimized feature alignment to achieve Adaptive Layout Refinement (ALR); triplet semi-hard loss calculation by optimizing the embedding space through the Euclidean distance relationship between anchor samples, positive samples, and negative samples; and end-to-end training and optimization of the model by combining the ALR technique and loss function.

[0068] (1) Data preprocessing and enhancement

[0069] The input image is uniformly scaled to a fixed size (such as 224×224 pixels) and normalized:

[0070]

[0071] Where μ and σ are the mean and standard deviation of the data set, respectively.

[0072] Random rotation (±30°), brightness adjustment (±20%), shear transformation (±0.2), horizontal flipping and scale scaling (0.8-1.2 times) are used to generate diverse training samples and enhance the generalization ability of the model.

[0073] (2) Improved EfficientNet-B0 feature extraction model

[0074] Based on EfficientNet-B0, dynamic channel scaling and feature pyramid fusion are introduced. The steps are as follows:

[0075] Adjust the width factor according to the input resolution RinputRinput:

[0076]

[0077] where R input Represents the resolution (side length) of the input image, and φ represents the dynamically adjusted width coefficient used to scale the network width (number of channels).

[0078] The output feature map F = {f1, f2, f3} realizes multi-scale feature fusion. The fusion formula is:

[0079]

[0080] Where ⊕ represents channel splicing, Conv 1×1 Dimensionality reduction for 1×1 convolution.

[0081] (3) Adaptive Layout Refinement (ALR) module:

[0082] Input: Image feature map H i-1 ∈R W×H×D , text word vector W∈R T×D , then the semantic similarity matrix (SSM) is calculated:

[0083]

[0084] Where Θ∈R N×T Represents the matching weight between the image sub-region and the text meaning.

[0085] For adaptive weight adjustment, first calculate the residual tensor R = |Θ synth -Θ real ∣, then divide the difficult and easy areas, set the threshold γ = 0.2, and divide R into R easy (R<γ) and R hard (R≥γ), then perform dynamic weight learning, and finally calculate the ALR loss function.

[0086] α=φ α (R easy ⊙H real ),β=φ β (R hard ⊙H real ) (18)

[0087] Among them, φ α and φ β It is a lightweight convolutional network that outputs weight coefficients of the same dimension as RR.

[0088] ALR loss function:

[0089]

[0090] softplus(x)=ln(1+e x) is used to constrain β>α and strengthen the focus on difficult samples.

[0091] (4) Triple semi-hard loss function design

[0092] The sample selection strategy dynamically selects anchor samples, positive samples, and negative samples. Anchor samples refer to randomly selected original NFT images, and positive samples refer to variants of anchor samples generated through data augmentation (such as rotation and cropping). Negative samples refer to images that are semantically similar to anchor samples but not duplicates, and meet the semi-hard conditions:

[0093]

[0094] Among them, x a is the anchor sample (Anchor), a randomly selected original image. p is a positive sample, which is a variant generated by performing data augmentation (such as rotation and cropping) on ​​the anchor sample. n is a negative sample, which is an image that is semantically similar to the anchor sample but not repeated. α is the interval parameter (the default is α = 0.5). The loss function formula is:

[0095]

[0096] Where N is the number of samples in the training batch. is the anchor sample, positive sample, and negative sample in the i-th triplet, [·]+ takes a positive function (i.e., max(0,·)), which takes effect when the loss is greater than 0, otherwise it is 0.

[0097] (5) Dynamic learning rate optimization

[0098] The model parameters are divided into three categories: feature extraction layer (EfficientNet backbone), ALR module, and classification head, so as to complete the parameter grouping. Then, adaptive learning rate scheduling is performed. The initial learning rate is set to η0 = 0.001 and dynamically adjusted according to the gradient amplitude:

[0099]

[0100] Among them, g represents the parameter group and G is the total number of groups.

[0101] The cosine annealing strategy is used to calculate the learning rate decay:

[0102]

[0103] where η max =0.01,η min =0.0001, T max is the total number of training rounds.

[0104] Then perform model training and optimization, first input batch images X∈R B×224×224×3 , label Y∈{0,1} B (0 means original, 1 means plagiarism). Then, EfficientNet-B0 extracts features, and the ALR module refines the layout to output the embedding vector f(X)∈R 128 The total loss is the weighted sum of triplet loss and ALR loss:

[0105] L Total =λ1L Triplet +λ2L ALR ,λ1=1.0,λ2=0.5 (24)

[0106] The Adam optimizer was used to update the parameters, with momentum parameters β1 = 0.9 and β2 = 0.999. A dropout layer (dropout rate 0.5) was added after the fully connected layer. If the validation set loss did not decrease for 10 consecutive rounds, the training was terminated, and then weight regularization was performed, with the L2 regularization coefficient set to 1×10 -4 .

[0107] (6) Plagiarism detection and judgment

[0108] First, perform feature comparison. For the image to be detected I test , extract the embedding vector f(I test ). Then calculate the similarity and calculate the Euclidean distance with the image IDBIDB in the database:

[0109]

[0110] Finally, set the dynamic threshold τ = μ d +3σ d (μ d is the average distance of positive samples in the training set, σ d is the standard deviation), if d Euclidean <τ is considered plagiarism.

[0111] The introduction of an adaptive layout optimization module and a Triplet Semi-Hard loss function improves the high-level semantic layout optimization module and feature extraction and embedding module of the traditional EfficientNet model, increasing duplicate detection accuracy by 12.7% and the F1 score by 9.3%. The success rate for detecting tampering methods such as brightness adjustment and partial occlusion exceeds 95%. A dynamic learning rate strategy reduces training time by 30% and the number of required iterations by 25%. By integrating an efficient network architecture, a dynamic learning mechanism, and a robust loss function, this technology addresses the bottlenecks of traditional NFT duplicate detection methods in terms of accuracy, efficiency, and generalization. It performs particularly well in high-similarity image differentiation, anti-tampering capabilities, and low-resource scenarios, significantly improving the precision and accuracy of image duplicate detection.

[0112] 3. Improve the input and output layers of convolutional neural networks

[0113] In order to overcome the slow convergence speed and gradient vanishing or exploding problems of EfficientNet convolution, and at the same time improve the robustness of features and avoid overfitting problems, this patent proposes to improve the original EfficientNet by replacing the convolution layer of the input layer with two improved convolution layers and introducing an improved HOG module and adding a genetic layer between the pooling layer and the fully connected layer of the output layer. Specifically, the convolution layer of the input layer is replaced with two improved convolution layers and the improved HOG module is introduced, such as Figure 4 As shown, the image is first preprocessed by the improved HOG. The improved HOG module refers to using SRRF to extract feature points from all HOG blocks, retaining blocks with feature point ratios exceeding 10%, and then performing subsequent HOG processing on the retained blocks; the improved convolution layer refers to using improved Min-Max normalization in the convolution layer to process data. The improved Min-Max normalization uses robust statistical methods such as the median and interquartile range to define the scaling factor to enhance the representativeness of the features and reduce the impact of outliers, thereby solving the problem of static scaling factors. A genetic layer is inserted between the fully connected layer and the pooling layer of the output layer, as shown in Figure 5 As shown in the figure, the feature map output by the pooling layer is used as the input to the genetic layer, where a genetic algorithm is used. Each feature is treated as a gene, and its thickness, direction, and other information are represented through different encoding methods. During the crossover operation, two features are randomly selected and their partial information is exchanged, such as partially fusing the color and texture features of an object. The mutation operation makes random small changes to a feature, such as fine-tuning the details of a texture feature. This increases feature diversity, avoids overfitting, increases the probability of detecting subtle feature differences, and improves the model's retrieval performance.

[0114] The improved HOG module preprocesses the image by roughly dividing it into blocks, then uses SURF to extract key points and calculate the Hessian matrix. The specific formula is as follows:

[0115]

[0116] Where I is the image, It is the second-order partial derivative of the image, which represents the second-order change of the image in different directions. The determinant of the Hessian matrix is ​​used to find the feature points in the image. The specific formula is as follows:

[0117]

[0118] When the value of the determinant is large, it means that the point has strong local saliency in the image, that is, it is a feature point. After obtaining all the feature points, if the feature points in the block account for more than 10% of the total feature points, the block is retained, and the HOG feature is calculated on the retained block to obtain the HOG feature.

[0119] Region k =[(x k -s k ,y k -s k ),(x k +s k ,y k +s k )] (28)

[0120] Each key point (x k ,y k ) and the corresponding scale s k Usually, the scale of the key points is used as a standard, such as a fixed-size window, or it is adjusted according to the scale of the key points, and then the area is processed by the HOG module to obtain the HOG features.

[0121] Then enter the convolution layer and replace the convolution layer with two improved convolution layers. Each convolution layer has convolution output. Through multi-layer convolution output, features can be screened, the representativeness of features can be improved, and noise features can be filtered out.

[0122] First, input the data of the convolution layer and the features output between the convolution layers to obtain the feature set of the convolution layer, that is, the output: X = {x1, x2, ..., x n}where x is the eigenvalue of the feature. Then, the improved Min-Maxnormalization is applied to process the data. First, the median M and interquartile range IQR of the data are calculated. The specific formula is as follows:

[0123] IQR=Q3-Q1 (29)

[0124] Where Q1 is the first quartile and Q3 is the third quartile. The scaling factor k is defined using the IQR, and the specific formula is as follows:

[0125]

[0126] Then, perform normalization operations to process the data, which can reduce the impact of outliers and optimize feature extraction. The formula for obtaining the new data is as follows:

[0127]

[0128] Use the normalized data y i For the calculations of subsequent layers. After the convolutional output, through the backbone network, to the output layer, through the pooling layer, the feature map is obtained. The obtained feature map is used as the input of the genetic layer. The feature map is encoded differently with information such as direction and thickness. Through encoding with the feature map as the unit, the feature vectors A = (a1, a2,..., a n ) and B = (b1, b2,..., b n ) etc. are obtained. Then, perform the crossover operation. Randomly select a crossover point i (0 < i < n), disconnect and connect the vectors at the crossover point. After crossover, generate a new feature vector:

[0129] X = (a1, a2,..., a i , b i+1 ,..., b n ) (32)

[0130] Through this formula for feature crossover, new features are generated, which can provide more diverse features, avoid overfitting, and these features still have strong representativeness. During the mutation operation, randomly select multiple values in the feature vector to produce slight changes, and then generate a new feature vector:

[0131] A′ = (a1 + n, a2 + m,..., a n ) (33)

[0132] Where n and m are the mutation values, and different information such as the thickness and direction of feature mutation is also provided, which also provides more diverse features. When processing image data with noise, the model still maintains good performance and can retrieve similar pictures. The features output by the original pooling layer are also output unchanged through inheritance.

[0133] The features output by the genetic layer are used as the input of the fully connected layer. Finally, the final feature data is obtained through the fully connected layer.

[0134] 4. Introduce the TransNeXt module and the RepViT module to improve the Block4 module in efficientnet

[0135] In order to overcome the shortcomings of EfficientNet in capturing long-range features and spatial dimension attention, while improving the model's inference efficiency and accuracy on mobile devices, this patent proposes replacing the Block4 module in EfficientNet with a module consisting of a convolutional layer, a depthwise separable convolutional layer, a TransNeXtBlock layer, a RepViTBlock layer, and a convolutional layer connected in series. Specifically, as shown in the figure, in the improved EfficientNet, after the feature map enters the location of the original Block4 module, it first passes through the traditional dimensionality-increasing 1×1 convolutional layer, which can increase the dimension of the features without changing the feature size, and then passes through the depthwise separable convolutional layer, which also keeps the feature size unchanged. Subsequently, the feature map enters the TransNeXtBlock layer for image feature extraction and enters the RepViTBlock layer to further refine and enhance the features. Finally, the feature map enters the traditional dimensionality-reducing 1×1 convolutional layer.

[0136] As shown in the figure, in the improved EfficientNet, the TransNeXtBlock layer is composed of a convolutional gated linear unit (GLU), an aggregated attention mechanism (AA), and a feedforward neural network module (FFN). This combination can effectively extract and weight features at both local and global scales, while using the FFN module to further enhance feature expression, significantly improving the model's robustness and anti-interference ability.

[0137] Among them, the convolutional GLU is a new channel mixer consisting of a 3×3 depth-separable convolution and two linear projections, one of which is subjected to a gated activation function. It combines the channel attention mechanism based on the nearest neighbor pixel features, and the output is a gated feature map, in which the features of each channel are adaptively weighted based on neighboring features. It focuses on finely capturing and dynamically adjusting local features, and can mine detailed information of local areas in the feature map.

[0138] Among them, the aggregated attention mechanism integrates multiple attention mechanisms and can effectively process information from different levels and scales. It focuses on the perception and feature weighting of global information and can effectively integrate the information of the entire feature map. Its formula is as follows:

[0139]

[0140]

[0141] B (i,j) =Concatenate(B (i,j)~ρ(i,j),log-CPB(Δ (i,j)~σ(X) )) (36)

[0142] X (i,j) =softmax(τlogN×Concatenate(S (i,j)~ρ(i,j) ,S (i,j)~σ(X) )+B (i,j) ) (37)

[0143] A (i,j)~ρ(i,j) ,A (i,j)~σ(X) =Split(A (i,j) )with size[k 2 ,H ρ W ρ ] (38)

[0144]

[0145] Among them, ρ(i,j) is the pixel set in the sliding window centered on pixel (i,j), σ(X) is the pixel set obtained by pooling the feature map, and are the normalized query and key, QE is the learnable query embedding, and B (i,j)~p(i,j) is the learnable position deviation on the sliding window path, log-CPB(Δ (i,j)~σ(X) is the position deviation on the pooled feature path calculated by logarithmically spaced continuous position deviation (log-CPB), Δ (i,h)~σ(X) yes and The spatial relative coordinates between them, τ is a learnable variable, N is the number of valid keys for each query interaction, T is a set of learnable tokens used to interact with the query to obtain additional dynamic position bias, V ρ(i,j) and V σ(X) are the eigenvalues ​​corresponding to ρ(i,j) and σ(X), respectively.

[0146] The FFN module consists of an input layer, a hidden layer, and an output layer. It is located after the aggregate attention mechanism and can further integrate and transform the features after local and global feature processing, thereby enhancing the expressive power of the features without increasing the computational complexity. The formula is as follows:

[0147] FFN(X)=max(0,XW1+b1)×W2+b2 (40)

[0148] Among them, X is the input vector, W1 and W2 are weight matrices, b1 and b2 are bias terms, and max(0,X) is the ReLU activation function, which converts negative numbers to 0 and enhances the nonlinear ability of the model.

[0149] As shown in the figure, the RepViTBlock layer consists of a 3×3 convolution and an FFN module. On this basis, a 3×3 convolution is added to deepen the network depth of the downsampling layer and enhance the ability to obtain long-distance features. The use of a 3×3 convolution can reduce latency while maintaining feature extraction accuracy. The FFN module is added after the 3×3 convolution. This module can memorize more potential information and help the model learn richer feature expressions, making the features more discriminative after processing.

[0150] Compared with the prior art, the present invention has the following beneficial effects:

[0151] (1) The present invention combines the existing PyMuPDF library with the OpenCV library to process images in PDF files. Compared with the previous reliance on manual collection and adjustment, it saves a lot of time and can quickly extract and pre-process images. The image fast classification and annotation model trained by the EfficientNet network effectively reduces the time and resources required for image annotation and classification. The low-rank adapter is injected into the self-attention module of the Vision Transformer, which greatly reduces the use of training resources. The low-rank adapter can also be used to make directional adjustments to the key / value projection of the self-attention mechanism, thereby enhancing the model's sensitivity to target domain features. Integrating ViT-LoRA into Stable Diffusion and introducing random noise and local redrawing can effectively improve the local texture diversity of the generated images while ensuring that the global structure of the generated samples remains intact, making the generated images more able to simulate real-life tampering behavior. The pipelined initial data processing method of images based on the combination of semi-automatic annotation and classification technology and image synthesis technology significantly reduces manual intervention while effectively increasing the number of rare types of images, making it possible to quickly build a large set of similar image samples.

[0152] (2) By integrating the ALR module into the high-level semantic layout optimization module of EfficientNet and using the Triplet Semi-Hard loss to optimize the embedded feature extraction module, this method significantly enhances the model's sensitivity to local tampering and layout differences in NFT images while retaining EfficientNet's efficient feature extraction capabilities, and enhances the model's ability to distinguish images with similar features. The scaling strategy adopted by EfficientNet-B0 involves uniformly scaling all width, depth, and resolution dimensions through composite coefficients. Unlike traditional methods that arbitrarily scale these elements, EfficientNet-B0's scaling method involves uniformly increasing the network depth, width, and resolution using a set of preselected scaling coefficients, which ensures that the network has enough layers to cover a larger receptive area and more channels to detect finer details in larger images.

[0153] (3) More comprehensive feature extraction: Adding convolutional layers can extract features at different levels and scales of the image. The genetic layer further optimizes and combines these features, making the extracted features more representative, covering more information about the image, and improving the ability to distinguish different images; Enhanced model robustness: The improved Min-Max normalization uses robust statistical methods such as the median and interquartile range to define the scaling factor. For data processing, it effectively reduces the impact of outliers on data normalization, making the data more stable and improving the robustness of the model. When processing image data containing noise, the model can still maintain good performance; Optimize the neural network structure. The addition of the genetic layer introduces the idea of ​​genetic algorithm into the neural network. Through genetic operations, it optimizes features, increases diversity, avoids overfitting, and can better handle the retrieval task of the same target image with small differences.

[0154] (4) The present invention replaces the Block4 module in EfficientNet with a module consisting of a convolutional layer, a depthwise separable convolutional layer, a TransNeXtBlock layer, and a convolutional layer in series, and adds a RepViT Block layer, which improves the model's local and global feature extraction and expression capabilities, effectively overcomes the original model's shortcomings in capturing long-range features and spatial attention, and improves the model's operating efficiency and extraction accuracy on mobile devices.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific implementation methods of the present invention can still be modified or replaced without departing from the spirit and scope of the present invention. Any modification or replacement should be included within the scope of protection of the claims of the present invention.

Claims

1. A paper duplication checking method based on an improved EfficientNet network, characterized by An innovative image forgery detection scheme based on deep convolutional neural networks is constructed, and the image retrieval function is optimized in a targeted manner through the perceptual measurement method DreamSim and the progressive local filter pruning method based on convolutional neural networks and network pruning technology, so as to accelerate the operation of the paper duplication detection system. The scheme includes the following steps: (1) constructing a total query image library; (2) using semi-automatic image extraction, classification and synthesis technology to quickly construct paper duplication detection training samples; (3) introducing an improved Triplet Semi-Hard loss function and an adaptive layout optimization module ALR to improve the high-level semantic layout optimization module and feature extraction and embedding module of the traditional EfficientNet model; (4) improving the input and output layers of the EfficientNet network; (5) introducing the TransNeXt module and the RepViT module to improve the Block4 module in the EfficientNet; (6) completing the paper duplication detection based on the Euclidean distance image similarity measurement; the construction of the total query image library refers to collecting 5,000 papers in PDF format and using Python's PyMuPDF library to convert the collected papers into PDF files. 5,000 PDF papers were converted into images, with one image per PDF page, to obtain image set A. The PDF papers to be queried were then converted into images using Python's PyMuPDF library, to obtain image set B. The total library consists of image set A and image set B. The semi-automatic image extraction, classification, and synthesis technology described above rapidly constructs training samples for paper duplication checking. This involves first batch extracting PDF images of the papers using the PyMuPDF library, preprocessing them into a 380×380 size format using the OpenCV library, and manually annotating 500 benchmark images. EfficientNet network training then obtains a preliminary model to automatically annotate the dataset. Then, the low confidence images are manually corrected, and the model is iteratively optimized 10 times to obtain a high-precision classification model; further, for the same label atlas, the VisionTransformer architecture is used to train the Lora model, and it is applied to Stable Diffusion and random noise is introduced to generate 3,000 semantically consistent images that show high diversity in multiple details; the paper duplication check based on the Euclidean distance image similarity measurement refers to extracting the features of each image in image set A and image set B based on the improved EfficientNet network obtained after steps (2) to (5). If the Euclidean distance between the features of image set B and the features of image set A is less than 0.2, it is considered that the PDF paper to be queried has similarities with the papers in the 5,000 papers.

2. A method for checking duplicate papers based on an improved EfficientNet network according to claim 1, characterized in that: The improved Triplet Semi-Hard loss function introduced herein refers to dynamically selecting anchor samples, positive samples and negative samples based on the fixed margin of the traditional Triplet Loss, combined with a dynamic sample selection strategy; the dynamic sample selection strategy refers to adaptively enhancing the optimization weight of difficult samples and reducing the optimization weight of easy samples during training, that is, constructing a dynamic weighting factor based on the loss contribution of the sample triplet to adjust the gradient update weight during back propagation; specifically, firstly, dynamic discrimination of difficult and easy samples is performed, and in each forward propagation, by calculating the Euclidean distance d1 between the anchor sample and the positive sample, and the Euclidean distance d2 between the anchor sample and the negative sample, the difficult sample condition is defined as d2-d1<α+β×σ, where α is the fixed margin threshold, β is the adaptive relaxation coefficient, and σ is the dynamic standard deviation of the distance difference of the current batch of samples, which is updated in real time through sliding window statistics; then, a nonlinear scaling function is introduced for the difficult samples through dynamic mapping of the gradient weights. Among them L hard is the loss value of the current difficult sample, γ is the global weight factor, λ controls the steepness of the weight distribution, ∈ is a smoothing term to prevent the gradient from disappearing, and then back-propagation is performed to reweight the dynamic weight w and the original gradient tensor element by element, so that the gradient amplitude of the difficult sample is amplified by w, and the gradient explosion risk is constrained by gradient clipping; the dynamic selection of anchor samples, positive samples and negative samples refers to distinguishing difficult samples from easy samples by setting dynamic thresholds, and the difficult sample refers to the Euclidean distance d between the negative sample and the anchor sample. n Satisfy d p <d n ≤d p +γ; the easy sample refers to the Euclidean distance d between the negative sample and the anchor sample n Satisfy d n >d p +γ, where d p is the Euclidean distance between the positive sample and the anchor sample, γ is the preset margin threshold, and γ is set to 0.2; the introduction of the improved adaptive layout optimization module ALR for deep integration with EfficientNet-B0 refers to the introduction of dynamic convolution kernel in the ResBlock of the adaptive layout optimization ALR module of the EfficientNet-B0 network, according to the text-visual correlation matrix TVM matrix Q i-1 Generate convolution kernel weights to make the convolution operation adaptive to the semantic focus area of ​​the text description. The original formula is Conv(H i-1 )=H i-1 ×K+b, improved to Conv(H i-1 )=Conv(H i-1 ; θ=f(Q i-1 )); Then, the semantic similarity matrix SSM and the text-visual association matrix TVM are dynamically constructed. The dynamic construction refers to the Q based on TVM in the residual block of the ALR module. i-1 Generate convolution kernel parameters and dynamically allocate weights through the gated attention mechanism. The formula is θ = Soft max (W q ×Q i-1 )×W k , where W q , W k The projection matrix can be learned to make the convolution kernel focus on the key semantic area of ​​the text description; the semantic correspondence between the local area of ​​the image and the text description is analyzed in real time, and the layout optimization weight is adaptively adjusted based on the residual features; the adaptive adjustment of the layout optimization weight refers to constructing the regional optimization weight through the residual R = |Θ-Θ'| of the semantic similarity matrix (SSM) Θ and the real image SSMΘ' ij =Sigmoid(φ(R ij ⊙H i-1 )), where φ is a multi-layer perceptron that implements weight reinforcement in high residual regions (where layout deviation is significant); in the multi-scale feature pyramid of EfficientNet-B0, according to the KL divergence of the layout error at the current level, KL (Θ l ||Θ l ') Dynamically adjust the learning rate of each layer: Where k is the sensitivity coefficient, η base As the basic learning rate, it realizes the differentiated optimization of deep semantic features and shallow texture features.

3. A method for checking duplicate papers based on an improved EfficientNet network according to claim 1, characterized in that: The improved input and output layers of the EfficientNet network refer to replacing the convolution layer of the input layer with two improved convolution layers and introducing an improved HOG module and adding a genetic layer between the pooling layer and the fully connected layer of the output layer. The improved convolution layer refers to using an improved Min-Max normalization method to improve the convolution layer when processing data in each convolution layer; the improved Min-Max The normalization method refers to using the median and interquartile range instead of the maximum and minimum values ​​for normalization. The interquartile range refers to the number obtained by subtracting the first quartile from the third quartile, and then the interquartile range and the median are used to normalize the original data; the introduction of the improved HOG module refers to using SRRF to extract feature points from all blocks of HOG, and retaining blocks with a feature point ratio exceeding 10%, and then performing subsequent HOG processing on the retained blocks to obtain feature vectors; the genetic layer refers to constructing a new genetic layer based on a genetic algorithm, that is, performing crossover, mutation, and genetic operations on the feature map encoding to obtain a new feature map encoding, thereby obtaining a new genetic layer to increase feature robustness and avoid overfitting.

4. A method for checking duplicate papers based on an improved EfficientNet network according to claim 1, characterized in that: The improved Block4 module in EfficientNet refers to replacing the Block4 module in EfficientNet with a module consisting of a convolutional layer, a depth-separable convolutional layer, a TransNeXtBlock layer, a RepViTBlock layer and a convolutional layer in series; the TransNeXtBlock layer refers to combining the aggregated attention mechanism and the Convolutional Gated Linear Unit and adding a feedforward neural network module, so that the model can effectively extract and weight features at local and global scales; the RepViTBlock layer refers to using 3×3 convolution to deepen the network depth, and then adding a feedforward neural network module to strengthen feature expression, so that the model can effectively extract long-distance features of the image and enhance extraction efficiency and accuracy.