End-to-End Dual-Stream Seal Automatic Verification Method Based on Deep Learning
Through the end-to-end double-stream seal automatic verification method of deep learning, the seal image enhancement network and the dual-stream image classification network are used to solve the problems of artificial dependence and insufficient computer vision accuracy in seal verification, and efficient and accurate identification of seal authenticity is achieved, which is suitable for seal image verification in complex backgrounds.
Patent Information
- Application Number
- CN202411340629.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-09-25
AI Technical Summary
The existing seal verification methods rely on manual judgment to be time-consuming and labor-intensive, and are easily affected by subjective factors. In addition, computer vision technology has low accuracy in identifying seal authenticity and false recognition in the context of small differences and complex backgrounds, and is insufficient robustness, making it difficult to meet the needs of high precision and high efficiency.
The end-to-end double-stream seal automatic verification method based on deep learning is adopted, and image enhancement processing and registration correction are performed through the seal image enhancement network, and feature extraction and classification are combined with the dual-stream image classification network to generate seal verification information to achieve authenticity and false identification of seals.
It improves the efficiency and accuracy of authenticity identification of seals, has strong anti-interference ability, and is suitable for seal image verification tasks in various complex backgrounds.
Smart Images

Figure CN119339184B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a verification method, in particular to an end-to-end dual-stream seal automatic verification method based on deep learning. Background Art
[0002] As an indispensable credit voucher in business operations and administrative management, seals have legal effect and play a crucial role in ensuring the legality of transactions and the validity of documents.
[0003] Traditional seal verification methods mostly rely on the judgment of professionals, which is not only time-consuming and laborious, but also easily affected by subjective factors and human errors. In addition, traditional methods lack immediacy and automation capabilities, and therefore cannot efficiently and accurately meet the needs of seal authenticity identification. The emergence of computer vision technology has provided a revolutionary solution for seal automatic verification.
[0004] When using computer vision technology for seal verification, deep learning and image processing algorithms can be used to achieve seal image enhancement and authenticity identification. Specifically, computer vision technology can remove the noise interference of the seal image and accurately correct the image, so as to efficiently and accurately verify the authenticity of the seal and ensure the authenticity and legality of the document. When using computer vision technology for seal verification, not only can the speed and accuracy of seal authenticity verification be improved, but also the labor cost can be reduced and the interference of human factors can be eliminated.
[0005] In addition, when using computer vision for seal verification, it can also meet the needs of various application scenarios, such as the needs of seal verification in intelligent contract management and automated document processing. Therefore, computer vision technology has brought new opportunities and prospects for the research of seal automatic verification, and is expected to promote greater progress in this field.
[0006] The application with the publication number CN117765561A discloses a method for authenticating the authenticity of seal images based on a Siamese feature extraction neural network. Specifically, a Siamese model is designed according to the characteristics of seal images to make it applicable to the authenticity identification of seal images. Specifically, the technical core of this application includes: obtaining a seal image to be tested and a corresponding reference image, where the reference image is a genuine seal image; registering and aligning the obtained images; extracting features from the registered and aligned seal image to be tested and the reference image through a Siamese-based feature extraction neural network to obtain the deep learning features of the image to be tested and the reference image. Among them, the Siamese-based feature extraction neural network includes at least two feature extraction sub-networks connected to each other at the intermediate layer through an SE-Block structure, and the feature extraction sub-network is a DenseNet network structure; based on the deep learning features of the seal image to be tested and the reference image extracted, a similarity metric is performed; based on the similarity, the authenticity of the seal image to be tested is identified. The present invention can perform real-time intelligent authentication on the seal image to be tested, determine whether it is a genuine seal, and achieve fast and effective identification of the authenticity of the seal.
[0007] However, when this application identifies genuine and fake seal images with minor differences, there are problems of low accuracy and insufficient robustness; especially for the seal image authentication task with complex background interference, the DenseNet network used has low accuracy and is difficult to meet the actual application requirements.
[0008] In summary, the existing seal verification methods mainly have the following deficiencies: the manual verification process is cumbersome and highly subjective, and is not suitable for large-scale applications; traditional computer vision techniques have insufficient recognition accuracy for minor differences when identifying genuine and fake seals, and cannot meet the requirements of high-precision verification; although existing deep learning methods have made certain progress in the field of seal verification, they still face the problem of insufficient robustness, are difficult to accurately verify seals in complex scenarios, and are easily affected by factors such as background interference, the size and direction of the stamping pressure. Summary of the Invention
[0009] The object of the present invention is to overcome the deficiencies existing in the prior art and provide an end-to-end dual-stream seal automatic verification method based on deep learning, which can efficiently and accurately identify the authenticity of seals, and also has strong anti-interference ability, and is applicable to the authenticity verification tasks of seal images with various complex backgrounds.
[0010] According to the technical solution provided by the present invention, an end-to-end dual-stream seal automatic verification method based on deep learning, the seal automatic verification method includes:
[0011] Obtain the seal image to be verified, and extract the corresponding reference seal image from the seal verification library based on the seal name of the seal image to be verified, where the reference seal image is at least consistent with the seal name of the seal image to be verified;
[0012] Load the obtained seal image to be verified into a pre-constructed seal verification model to verify the seal image to be verified by using the seal verification model, where,
[0013] The seal verification model at least includes a seal image enhancement network and a two-stream image classification network;
[0014] When verifying the seal image to be verified, first perform image enhancement processing by using the seal image enhancement network to generate a seal image to be inspected after the image enhancement processing;
[0015] Overlay the seal image to be inspected with the above-extracted reference seal image to generate an overlaid seal image, and use the two-stream image classification network to perform feature extraction and classification processing on the overlaid seal image to generate seal verification information after the feature extraction and classification processing;
[0016] Based on the seal verification information, determine the authenticity of the seal image to be verified.
[0017] When performing image enhancement processing on the seal image to be verified by using the seal image enhancement network, it includes image denoising processing and / or seal correction and registration processing, where,
[0018] When the image enhancement processing of the seal image to be verified includes image denoising processing and seal correction processing, the seal image enhancement network includes an enhancement network input processing module, a denoiser, and a registration alignment module connected in sequence to perform image denoising processing and seal correction and registration processing on the seal image to be verified in sequence by using the seal image enhancement network, where,
[0019] Use the enhancement network input processing module to load the seal image to be verified into the denoiser;
[0020] Use the denoiser to perform image denoising processing on the seal image to be verified to generate a denoised seal image to be verified after the image denoising processing;
[0021] Use the registration alignment module to perform seal correction and registration processing on the denoised seal image to be verified to generate a seal image to be inspected after the seal correction and registration processing.
[0022] The denoiser includes a multi-stage residual dense connection module for denoising the seal image to be verified, an upsampling module for increasing the image resolution, and an image quality enhancement module for improving the image quality, where,
[0023] The multi-level residual dense connection module, the upsampling module, and the image quality enhancement module are connected in sequence. The multi-level residual dense connection module is connected to the output end of the enhancement network input processing module, and the image quality enhancement module is adaptively connected to the registration alignment module;
[0024] The multi-level residual dense connection module includes a feature enhancement basic unit and a dense connection convolutional layer. Among them, the feature enhancement basic unit includes several sequentially connected feature enhancement basic modules;
[0025] The feature enhancement basic unit is connected to the dense connection convolutional layer. The output of the dense connection convolutional layer and the output of the enhancement network input processing module are both adaptively connected to the dense connection splicer, and the output end of the dense connection splicer is adaptively connected to the upsampling module.
[0026] For any feature enhancement basic module, it includes a dense block unit group and a basic module residual scaling unit adaptively connected to the dense block unit group. Among them, the dense block unit group includes several sequentially connected dense block units,
[0027] The dense block unit includes a first branch unit, a second branch unit, and a dense block unit internal splicer for channel splicing of the first branch unit and the second branch unit;
[0028] The first branch unit includes a dense block and a dense block residual scaling unit adaptively connected to the dense block;
[0029] The input of the dense block in the first branch unit is directly loaded into the dense block unit internal splicer through the second branch unit;
[0030] For two adjacent dense block units, along the serial connection direction of the dense block units, the output end of the dense block unit internal splicer of the previous dense block unit is adaptively connected to the first branch unit and the second branch unit of the subsequent dense block unit.
[0031] When the registration alignment module performs stamp correction and registration processing on the denoised stamp image to be verified, it includes:
[0032] Extract the pattern center point of the denoised stamp image to be verified;
[0033] Generate registration control point A and registration control point B based on the pattern center point of the denoised stamp image to be verified and the key contour points of the central pattern of the denoised stamp image to be verified;
[0034] Generate the correction and registration angle of the denoised stamp image to be verified based on the pattern center point, configuration control point A, and registration control point B,
[0035] Among them, the correction and registration angle is:
[0036]
[0037] Among them, θ jz is the correction registration angle, and θ B is the polar coordinate angle of the registration control point B, and θ A is the polar coordinate angle of the registration control point A;
[0038] Based on the correction registration angle, the denoised seal image to be verified is rotated to the target state, and a seal image to be inspected is generated after rotating to the target state.
[0039] The dual-stream image classification network includes a dual-stream backbone network and a classification network output module connected in sequence. Among them,
[0040] The dual-stream backbone network includes an EfficientNet B0 module for global feature and local feature extraction, an SVIT module for global dependency extraction, and a backbone network channel splicer for channel splicing.
[0041] For the superimposed seal image, the superimposed seal global feature and the superimposed seal local feature are extracted through the EfficientNet B0 module, and the superimposed seal global dependency is extracted through the SVIT module;
[0042] The superimposed seal global feature, the superimposed seal local feature, and the superimposed seal global dependency are spliced by channels through the backbone network channel splicer, and are classified and recognized and output through the classification network output module, so as to obtain seal verification information after the classification and recognition output.
[0043] The EfficientNet B0 module includes a first convolutional layer for feature extraction, an automatic inverted residual bottleneck convolutional unit group, a second convolutional layer for feature extraction, a pooling layer, and a global extraction fully connected layer connected in sequence. Among them,
[0044] When extracting the superimposed seal global feature and the superimposed seal local feature of the superimposed seal image, the superimposed seal image is sequentially processed by the first convolutional layer for feature extraction, the automatic inverted residual bottleneck convolutional unit group, the second convolutional layer for feature extraction, the pooling layer, and the global extraction fully connected layer, and the superimposed seal global feature and the superimposed seal local feature are output through the global extraction fully connected layer;
[0045] The automatic inverted residual bottleneck convolutional unit group includes a number of automatic inverted residual bottleneck convolutional units, and the automatic inverted residual bottleneck convolutional units are connected in series in sequence.
[0046] The automatic inverted residual bottleneck convolutional basic unit includes a first convolutional layer inside the bottleneck convolution, a depthwise separable convolutional layer, a squeeze-and-excitation module, a second convolutional layer inside the bottleneck convolution, a bottleneck convolution Dropout layer, and a bottleneck convolution adder. Among them,
[0047] The input end of the first convolutional layer in the bottleneck convolution is connected to the first addition end of the bottleneck convolution adder to serve as the input end of the automatic inverted residual bottleneck convolution unit;
[0048] The output of the first convolutional layer in the bottleneck convolution is loaded into the depthwise separable convolutional layer after batch normalization processing;
[0049] The output of the depthwise separable convolutional layer is loaded into the squeeze-and-excitation module after batch normalization processing. The output end of the squeeze-and-excitation module is connected to the second addition end of the bottleneck convolution adder through the second convolutional layer and the bottleneck convolution Dropout layer in the bottleneck convolution;
[0050] The output end of the bottleneck convolution adder serves as the output end of the automatic inverted residual bottleneck convolution basic unit where it is located.
[0051] The SVIT module includes a block embeddor, a positional embeddor, a Transformer encoder, and a module layer normalization unit connected in sequence, where
[0052] When extracting the global dependencies of the superimposed seal image, the superimposed seal image is processed by the block embeddor, the positional embeddor, the Transformer encoder, and the module layer normalization unit in sequence, where
[0053] The block embeddor is used to divide the superimposed seal image to generate several seal image blocks after division, and at the same time, the information of the seal image is converted into a high-dimensional embedding representation to achieve the linear projection of the seal image blocks;
[0054] The positional embeddor is used to determine the position information of each seal image block, and all the seal image blocks are connected in sequence to form an image block feature sequence;
[0055] The Transformer encoder processes the image block feature sequence and generates the global dependencies of the superimposed seal through the module layer normalization unit.
[0056] The Transformer encoder includes several Transformer encoding units connected in series in sequence, where
[0057] For any two serially connected Transformer encoding units, along the serial connection direction of the Transformer encoding units, the output end of the previous Transformer encoding unit is connected to the input end of the next Transformer encoding unit;
[0058] For any Transformer encoding unit, it includes a multi-head attention mechanism unit and a multi-layer perceptron module;
[0059] The input end of the multi-head attention mechanism unit serves as the input end of the corresponding Transformer encoding unit, and the multi-head attention mechanism unit is adaptively connected to the multi-layer perceptron module;
[0060] The input end and the output end of the multi-head attention mechanism unit are respectively connected to the corresponding input ends of the first adder of the encoder, and the output end of the first adder of the encoder is connected to the input end of the multi-layer perceptron module.
[0061] The output end of the first adder of the encoder and the output end of the multi-layer perceptron module are connected to the corresponding input ends of the second adder of the encoder, and the output end of the second adder of the encoder serves as the output end of the corresponding Transformer encoding unit.
[0062] Advantages of the present invention: A seal verification model is pre-constructed, and the seal verification model is used to verify the seal image to be verified. When verifying the seal image to be verified, first, an image enhancement process is performed using a seal image enhancement network to generate a seal image to be inspected after the image enhancement process; the seal image to be inspected is superimposed on the above-mentioned extracted reference seal image to generate a superimposed seal image, and a two-stream image classification network is used to perform feature extraction and classification processing on the superimposed seal image to generate seal verification information after the feature extraction and classification processing; based on the seal verification information, the authenticity of the seal image to be verified is determined, which can efficiently and accurately identify the authenticity of the seal, and also has strong anti-interference ability, and is applicable to the authenticity verification task of seal images with various complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic flowchart of an embodiment of the automatic seal verification of the present invention.
[0064] Figure 2 It is a structural block diagram of an embodiment of the seal image enhancement network of the present invention.
[0065] Figure 3 It is a structural block diagram of an embodiment of the feature enhancement basic module of the present invention.
[0066] Figure 4 It is a structural block diagram of an embodiment of the two-stream image classification network of the present invention.
[0067] Figure 5 It is a structural block diagram of an embodiment of the automatic inverted residual bottleneck convolution unit of the present invention.
[0068] Figure 6 It is a schematic flowchart of an embodiment of the working process of the SVIT module of the present invention.
[0069] Figure 7This is a structural block diagram of an embodiment of the Transformer encoder of the present invention.
[0070] Figure 8 This is a schematic diagram of an embodiment for training and generating a seal image enhancement network of the present invention.
[0071] Figure 9 This is a schematic diagram of an embodiment for training and generating a two-stream image classification network of the present invention. Detailed implementation manners
[0072] The present invention will be further described below in conjunction with specific drawings and embodiments.
[0073] In order to be able to efficiently and accurately identify the authenticity of seals and apply to the authenticity verification task of seal images in various complex backgrounds, the present invention provides an end-to-end two-stream seal automatic verification method based on deep learning. Specifically, the seal automatic verification method includes:
[0074] Obtain the seal image to be verified, and extract the corresponding reference seal image in the seal verification library based on the seal name of the seal image to be verified, wherein the reference seal image is at least consistent with the seal name of the seal image to be verified;
[0075] Load the obtained seal image to be verified into a pre-constructed seal verification model to verify the seal image to be verified by using the seal verification model, wherein,
[0076] The seal verification model at least includes a seal image enhancement network and a two-stream image classification network;
[0077] When verifying the seal image to be verified, first perform image enhancement processing by using the seal image enhancement network to generate a to-be-inspected seal image after the image enhancement processing;
[0078] Overlay the to-be-inspected seal image with the above-extracted reference seal image to generate an overlaid seal image, and perform feature extraction and classification processing on the overlaid seal image by using the two-stream image classification network to generate seal verification information after the feature extraction and classification processing;
[0079] Based on the seal verification information, determine the authenticity of the seal image to be verified.
[0080] Figure 1FIG. 0 shows a schematic diagram of an embodiment of the present invention for seal image verification. In the figure, when performing seal image verification, it is necessary to first obtain the seal image to be verified. Here, the seal image to be verified is the seal image whose authenticity is to be identified. The seal image to be verified can be obtained by using existing common methods. For example, first perform a stamping operation on the seal image to be verified, and then obtain it by scanning, or directly perform image extraction and other methods to obtain the seal image to be verified. The specific method for obtaining the seal image to be verified can be selected according to needs, and no further examples will be given here.
[0081] After obtaining the seal image to be verified, common technical means in the technical field can be used to obtain the name of the seal image to be verified. For example, the seal name of the seal image to be verified can be obtained based on the method of Optical Character Recognition (OCR). Specifically, when using the OCR method, the position of the text can be located in the seal image to be verified and the text content can be recognized, so that the seal name of the seal image to be verified can be obtained after recognizing the text content. The method of using the OCR method to determine the seal name of the seal image to be verified can be consistent with the existing method. Of course, other methods can also be used to obtain the seal name of the seal image to be verified. The specific method for obtaining the seal name of the seal image to be verified can be selected according to needs, and no further examples will be given here.
[0082] It should be noted that when the present invention verifies the authenticity of the seal image to be verified, a reference seal image is required. Therefore, a seal verification library can be constructed according to the actual application scenario. One or more reference seal images are pre-stored in the seal verification library, and the corresponding reference seal image can be extracted from the seal verification library by the seal name. It can be seen from this that the reference seal image is at least consistent with the seal name of the seal image to be verified. It can be understood that the reference seal image can be a standard seal image, and the standard seal image is a clean, clear and complete seal image, and the reference seal image generally corresponds to the real seal, that is, the corresponding reference seal image can be made from the real seal. It can be seen from this that when the corresponding reference seal image cannot be extracted from the seal verification library according to the seal name of the seal image to be verified, the authenticity verification of the seal image to be verified cannot be performed. At this time, the entire seal authenticity verification process will be exited. Specifically, when implementing, the reference seal image can be made by using existing common methods and stored in the seal verification library.
[0083] From Figure 1It can be known that when verifying the seal image to be verified, it is necessary to load the obtained seal image to be verified into a pre-constructed seal verification model, so as to use the seal verification model to verify the seal image to be verified. In an embodiment of the present invention, the seal verification model at least includes a seal image enhancement network and a two-stream image classification network; specifically, when the seal verification model includes a seal image enhancement network and a two-stream classification network, the image enhancement process is first performed using the seal image enhancement network to generate a seal image to be inspected after the image enhancement process. The method and process of using the seal image enhancement network to perform image enhancement on the seal image to be verified can be referred to the following description, that is, the method of generating the seal image to be inspected can be referred to the following description.
[0084] Overlay the seal image to be inspected with the above-extracted reference seal image to generate an overlaid seal image. Specifically, when generating the overlaid seal image, the corresponding pixel values of the seal image to be inspected and the reference seal image can be overlaid according to the corresponding ratio. For example, a feasible overlay method is: I = λ 1 I 1 +λ 2 I 2 , where λ 1 and λ 2 are overlay coefficients, I 1 is the pixel value of the seal image to be inspected, and I 2 is the pixel value of the reference seal image. Generally, the sum of the overlay coefficient λ 1 and the overlay coefficient λ 2 is 1. Generally, the overlay coefficient λ 1 and the overlay coefficient λ 2 can both take 0.5. Of course, other overlay methods can also be used, and the specific overlay methods used will not be listed one by one, as long as the overlaid seal image can be generated.
[0085] After obtaining the overlaid seal, use the two-stream image classification network to perform feature extraction and classification processing on the overlaid seal image to generate seal verification information after the feature extraction and classification processing; at this time, based on the seal verification information, determine the authenticity of the seal image to be verified. Among them, the seal verification information can be a true / false binary classification result. Among them, when the true / false binary classification result is 1, it indicates that the seal image to be verified is genuine; when the true / false binary classification result is 0, it indicates that the seal image to be verified is fake. Specifically, determining the authenticity of the seal image to be verified specifically refers to whether the seal image to be verified is consistent with the reference seal image. When the seal image to be verified is consistent with the reference seal image, it indicates that the seal image to be verified is genuine, otherwise, it indicates that the seal image to be verified is fake. The seal image to be verified is consistent with the reference seal image, specifically referring to that the seal image to be verified and the reference seal image have the same features.
[0086] As can be seen from the above description, the seal verification model can be used to verify the seal image to be verified. When using the seal verification model for verification, the end-to-end verification of the seal image to be verified is achieved. Thus, it can be known that the authenticity of the seal can be efficiently and accurately identified, and it is applicable to the authenticity verification tasks of seal images with various complex backgrounds.
[0087] In an embodiment of the present invention, when using the seal image enhancement network to perform image enhancement processing on the seal image to be verified, it includes image denoising processing and / or seal correction and registration processing. Among them,
[0088] When the image enhancement processing of the seal image to be verified includes image denoising processing and seal correction processing, the seal image enhancement network includes an enhancement network input processing module, a denoiser, and a registration and alignment module connected in sequence, so as to use the seal image enhancement network to perform image denoising processing and seal correction and registration processing on the seal image to be verified in sequence. Among them,
[0089] Use the enhancement network input processing module to load the seal image to be verified into the denoiser;
[0090] Use the denoiser to perform image denoising processing on the seal image to be verified, so as to generate a denoised seal image to be verified after the image denoising processing;
[0091] Use the registration and alignment module to perform seal correction and registration processing on the denoised seal image to be verified, so as to generate a seal image to be inspected after the seal correction and registration processing.
[0092] As can be seen from the above description, the seal image enhancement network can perform image enhancement processing on the seal image to be verified. Specifically, in implementation, the image enhancement processing can be image denoising processing and / or seal correction and registration processing. Taking the seal image enhancement network to perform image denoising processing and seal correction processing on the seal image to be verified in sequence as an example, the situation of the seal image enhancement network will be specifically described below.
[0093] When using the seal image enhancement network to perform image denoising processing and seal correction processing on the seal image to be verified in sequence, in an embodiment of the present invention, the seal image enhancement network includes an enhancement network input processing module, a denoiser, and a registration and alignment module connected in sequence. Specifically, the enhancement network input processing module can receive the seal image to be verified, perform input processing on the received seal image to be verified and then load it into the denoiser, so as to use the denoiser to perform image denoising processing, and finally use the registration and alignment module to perform seal correction and registration processing. The enhancement network input processing module, the denoiser, and the registration and alignment module in the seal image enhancement network will be specifically described below.
[0094] In one embodiment of the present invention, the denoiser includes a multi-level residual dense connection module for denoising the seal image to be verified, an upsampling module for increasing the image resolution, and an image quality enhancement module for improving the image quality. Among them,
[0095] The multi-level residual dense connection module, the upsampling module, and the image quality enhancement module are connected in sequence. The multi-level residual dense connection module is connected to the output end of the enhancement network input processing module, and the image quality enhancement module is adaptively connected to the registration alignment module;
[0096] The multi-level residual dense connection module includes a feature enhancement basic unit and a dense connection convolutional layer. Among them, the feature enhancement basic unit includes several sequentially connected feature enhancement basic modules;
[0097] The feature enhancement basic unit is connected to the dense connection convolutional layer. The output of the dense connection convolutional layer and the output of the enhancement network input processing module are both adaptively connected to the dense connection splicer, and the output end of the dense connection splicer is adaptively connected to the upsampling module.
[0098] Figure 2 An embodiment is shown in which the enhancement network input processing module in the seal image enhancement network uses a convolutional layer. Figure 2 Among them, Conv1 is the convolutional layer used by the enhancement network input processing module. Specifically, the enhancement network input processing module can use a convolutional layer with a convolution kernel of 3×3, a stride of 1, and a padding size of 1. Therefore, the input processing performed by the enhancement network input processing module on the seal image to be verified is the convolution operation performed on the seal image to be verified using the convolutional layer.
[0099] In specific implementation, the denoiser may include a multi-level residual dense connection module, an upsampling module, and an image quality enhancement module connected in sequence. Figure 2 Among them, the part between the upsampling module and the enhancement network input processing module is the multi-level residual dense connection module. The multi-level residual dense connection module can be used to perform denoising processing on the seal image to be verified. The upsampling module can be used to increase the resolution of the seal image after denoising processing, and the image quality enhancement module can be used to improve the image quality of the seal image processed by the upsampling module.
[0100] In one embodiment of the present invention, the multi-level residual dense connection module may include a feature enhancement processing basic unit and a dense connection convolutional layer. Figure 2 Among them, the basic module (feature enhancement) is the feature enhancement processing basic unit, and Conv2 is the dense connection convolutional layer. The dense connection convolutional layer can use a convolutional layer with a convolution kernel size of 3×3. Figure 2In it, the multi-level residual dense connection module further includes a dense connection splicer. Cat1 is the dense connection splicer. The output end of the dense connection convolutional layer is connected to the dense connection splicer. In addition, in addition to being connected to the input end of the feature enhancement basic unit, the enhanced network input processing module is also connected to the dense connection splicer, so as to use the dense connection splicer to perform channel splicing on the output of the enhanced network input processing module and the output of the dense connection convolutional layer. The output end of the dense connection splicer is adaptively connected to the upsampling module to achieve the adaptive connection between the multi-level residual dense connection module and the upsampling module.
[0101] During specific implementation, the dense connection splicer can perform an element-level addition operation on two tensors with the same shape ([number of batch images, channel dimension, feature map height, feature map width]) output by the convolutional layer Conv1 and the convolutional layer Conv2 to achieve channel splicing; the splicers in the present invention all adopt the same splicing operation, and the splicing operation methods adopted by other splicers can refer to the description here.
[0102] The feature enhancement basic unit includes a plurality of sequentially connected feature enhancement basic modules. Figure 2 In it, the "8×" of the basic module (feature enhancement) means that the feature enhancement basic unit includes 8 sequentially connected feature enhancement basic modules, that is Figure 2 An embodiment in which the feature enhancement basic unit includes 8 sequentially connected feature enhancement basic modules is shown in. Of course, the number of sequentially connected feature enhancement basic modules in the feature enhancement basic unit can also be selected according to actual needs, which will not be elaborated here.
[0103] It can be understood that when the number of feature enhancement basic modules increases, the seal image enhancement network can capture more detailed features and noise patterns when processing images. More feature enhancement basic modules mean more levels and deeper feature extraction capabilities, so as to better learn and identify the noise features in the seal image, making the denoising effect more significant.
[0104] After increasing the number of feature enhancement basic modules, it will also increase the number of parameters and computational complexity of the seal image enhancement network, increase the training and inference time. A small number of feature enhancement basic modules are suitable for application scenarios with limited resources and can provide a basic denoising effect at a lower computational cost; a large number of feature enhancement basic modules are suitable for scenarios with high requirements for image denoising effect and sufficient computational resources and can provide a better denoising effect at a higher computational cost. During specific implementation, the number of feature enhancement basic modules in the feature enhancement basic unit is preferably 8.
[0105] Generally, multiple sequentially connected feature enhancement basic modules can adopt the same structural form. Figure 3 An embodiment of the feature enhancement basic module is shown in. The following combinesFigure 3 Specific explanations are given to the feature enhancement basic module.
[0106] In an embodiment of the present invention, for any feature enhancement basic module, it includes a dense block unit group and a basic module residual scaling unit adaptively connected to the dense block unit group. Among them, the dense block unit group includes several sequentially connected dense block units.
[0107] The dense block unit includes a first branch unit, a second branch unit, and a dense block unit internal splicer for channel splicing of the first branch unit and the second branch unit.
[0108] The first branch unit includes a dense block and a dense block residual scaling unit adaptively connected to the dense block. The output end of the dense block residual scaling unit is connected to the dense block unit internal splicer.
[0109] The input of the dense block in the first branch unit is directly loaded into the dense block unit internal splicer through the second branch unit.
[0110] For two adjacent dense block units, along the connection direction of the dense block units, the output end of the dense block unit internal splicer of the previous dense block unit is adaptively connected to the first branch unit and the second branch unit of the subsequent dense block unit.
[0111] In specific implementation, the feature enhancement basic module may include a dense block unit group and a basic module residual scaling unit. The dense block unit group includes several sequentially connected dense block units. Figure 3 An embodiment is shown in which the dense block unit group includes 3 sequentially connected dense block units. Among them, each dense block unit includes a first branch unit, a second branch unit, and a dense block unit internal splicer. The first branch unit includes a dense block and a dense block residual scaling unit adaptively connected to the dense block. The dense block is connected to an input end of the dense block unit internal splicer through the dense block residual scaling unit. The input end of the dense block in the first branch unit is also directly connected to another input end of the dense block unit internal splicer through the second branch unit to realize channel splicing in the dense block unit internal splicer.
[0112] Figure 3Among them, MJ1, MJ2, and MJ3 are the three dense blocks. Among them, the dense block unit where MJ1 is located is the first dense block unit, the dense block unit where MJ2 is located is the second dense block unit, and the dense block unit where MJ3 is located is the third dense block unit. Specifically, β1, β2, and β3 respectively correspond to the three dense block residual scaling units. Among them, the dense residual scaling unit corresponding to β1 is the dense residual scaling unit within the first dense block unit, the dense residual scaling unit corresponding to β2 is the dense residual scaling unit within the second dense block unit, and the dense residual scaling unit corresponding to β3 is the dense residual scaling unit within the third dense block unit. Cat2, Cat3, and Cat4 respectively correspond to the splicers within the dense block units. Among them, the splicer within the dense unit corresponding to Cat2 is the splicer within the dense block unit within the first dense block unit, the splicer within the dense unit corresponding to Cat3 is the splicer within the dense block unit within the second dense block unit, and the splicer within the dense unit corresponding to Cat4 is the splicer within the dense block unit within the third dense block unit.
[0113] As can be seen from the above description, Figure 3 Within the feature enhancement basic module shown in, the first dense block unit includes the dense block MJ1, the dense block residual scaling unit β1, and the splicer Cat2 within the dense block unit. The second dense block unit includes the dense block MJ2, the dense block residual scaling unit β2, and the splicer Cat3 within the dense block unit. The third dense block unit includes the dense block MJ3, the dense block residual scaling unit β3, and the splicer Cat4 within the dense block unit.
[0114] As can be seen from the above description, within the feature enhancement basic module, the dense block units are connected in series in sequence. The direction of the connection of the dense block units is Figure 3 the direction from the first dense block unit to the third dense unit in. For two adjacent dense block units, they can be divided into the previous dense block unit and the subsequent dense block unit. For example, when the first dense block unit is connected in series with the second dense block unit, the first dense block unit is the previous dense block unit, and the second dense block unit forms the subsequent dense block unit. Similarly, when the second dense block unit is connected in series with the third dense block unit, the second dense block unit forms the previous dense block unit, and the third dense block unit forms the subsequent dense block unit. Therefore, according to the connection situation of the dense block units, the previous dense block unit and the corresponding subsequent dense block unit can be determined.
[0115] When cascaded, the output end of the splicer in the dense unit of the previous dense block unit is adaptively connected to the first branch unit and the second branch unit of the next dense block unit. For example, when the first dense block unit is cascaded with the second dense block unit, the output end of the splicer Cat2 in the dense block unit is connected to the input end of the dense block MJ3. As can be seen from the above description, the output end of the splicer Cat2 in the dense block unit is also connected to the input end of the splicer Cat3 in the dense block unit. The connection conditions during the cascading of other dense block units can refer to the description here, and no further examples will be given one by one.
[0116] Figure 3 An embodiment schematic diagram of the basic module residual scaling unit is also shown in the figure. In the figure, the basic module residual scaling unit includes a module residual scaling module and a splicer within the module unit. Among them, β4 is the module residual scaling module, and Cat5 is the splicer within the module unit. The output end of the module residual scaling module is connected to one input end of the splicer within the module unit; when the dense block unit group is connected to the basic module residual scaling unit, the output end of the dense block unit group is connected to the input end of the module residual scaling module, and the input end of the dense block unit group is also connected to the other input end of the splicer within the module unit.
[0117] During specific implementation, the dense block residual scaling units β1, β2, β3, and the module residual scaling module β4 can adopt the same residual scaling parameters. Generally, the corresponding residual scaling parameters of the dense block residual scaling units β1, β2, β3, and the module residual scaling module β4 can be set to 0.2.
[0118] Figure 3 In the figure, the output end of the splicer within the dense block unit in the third dense block unit can form the output end of the dense block unit group, and the input end of the dense block in the first dense block unit corresponds to the input end of the dense block unit group. Therefore, the output end of the splicer Cat4 in the dense block unit is connected to the input end of the module residual scaling module β4, and the input end of the dense block MJ1 is also connected to the other input end of the splicer Cat5 within the module unit. During specific implementation, the output end of the splicer Cat5 within the module unit forms the output end of the feature enhancement basic module, and the input end of the dense block MJ1 forms the input end of the feature enhancement basic module. When multiple feature enhancement basic modules are cascaded in sequence, the output end of the previous feature enhancement basic module is connected to the input end of the next feature enhancement basic module. When multiple feature enhancement basic modules are cascaded in sequence, the situations of the previous feature enhancement basic module and the next feature enhancement basic module can refer to the corresponding descriptions of the previous dense block unit and the next dense block unit above, and will not be elaborated here.
[0119] In specific implementation, the dense blocks within the feature enhancement basic module generally adopt the same form. For example, the dense blocks MJ1 to MJ3 mentioned above can adopt the same form. The dense block includes a first type of sub-unit group of the dense block and at least one second type of sub-unit of the dense block. The first type of sub-unit group of the dense block is adaptively connected to the second type of sub-unit of the dense block. The first type of sub-unit group of the dense block includes several sequentially connected first type of sub-units of the dense block. The first type of sub-unit of the dense block includes a first type of sub-unit convolutional layer and a first type of sub-unit LReLU layer adaptively connected to the first type of sub-unit convolutional layer. The second type of sub-unit of the dense block includes a second type of sub-unit convolutional layer.
[0120] Specifically, the first type of sub-unit convolutional layer and the second type of sub-unit convolutional layer can adopt convolutional layers with a convolution kernel of 3×3, a stride of 1, and a padding size of 1. The first type of sub-unit LReLU layer can adopt default parameters.
[0121] Figure 3 An embodiment of the dense block is also shown in the figure. In the figure, an embodiment of the first type of sub-unit group of the dense block includes four sequentially connected first type of sub-units of the dense block. Among them, Conv6 is the first type of sub-unit convolutional layer within the first first type of sub-unit of the dense block, LR3 is the first type of sub-unit LReLU layer within the first first type of sub-unit of the dense block, Conv7 is the first type of sub-unit convolutional layer within the second first type of sub-unit of the dense block, LR4 is the first type of sub-unit LReLU layer within the second first type of sub-unit of the dense block, Conv8 is the first type of sub-unit convolutional layer within the third first type of sub-unit of the dense block, LR5 is the first type of sub-unit LReLU layer within the third first type of sub-unit of the dense block, Conv9 is the first type of sub-unit convolutional layer within the fourth first type of sub-unit of the dense block, and LR6 is the first type of sub-unit LReLU layer within the fourth first type of sub-unit of the dense block.
[0122] Specifically, for the input loaded into the dense block, the input passes through a series of convolutional layers. The number of input channels of each convolutional layer gradually increases, while the number of output channels remains the same. For example, the number of input channels of the first type of sub-unit convolutional layer Conv6, the first type of sub-unit convolutional layer Conv7, the first type of sub-unit convolutional layer Conv8, and the first type of sub-unit convolutional layer Conv9 are 64, 128, 192, and 256 in sequence, and the number of output channels is 64 for all.
[0123] When connected in series, the first type of sub-unit LReLU layer LR3 is connected to the first type of sub-unit convolutional layer Conv7, the first type of sub-unit LReLU layer LR4 is connected to the first type of sub-unit convolutional layer Conv8, and the first type of sub-unit LReLU layer LR5 is connected to the first type of sub-unit convolutional layer Conv9.
[0124] Figure 3Among them, Conv10 is the second - type sub - unit convolutional layer. When the first - type sub - unit group of the dense block is adaptively connected to the second - type sub - unit of the dense block, the LReLU layer LR6 of the first - type sub - unit is connected to the convolutional layer Conv10 of the second - type sub - unit. The input end of the convolutional layer Conv6 of the first - type sub - unit forms the input end of the dense block, and the output end of the convolutional layer Conv10 of the second - type sub - unit forms the output end of the dense block.
[0125] Furthermore, Figure 2 An embodiment of the up - sampling module is also shown in [the figure]. The up - sampling module may include an up - sampling convolutional layer and an up - sampling LeakyReLU layer. The up - sampling convolutional layer is connected to the up - sampling LeakyReLU layer. In the figure, Conv3 is the up - sampling convolutional layer, and the up - sampling convolutional layer can be a convolutional layer with a convolution kernel of 3×3; LRU1 is the up - sampling LeakyReLU layer. The up - sampling module is connected to the output end of the dense connection splicer Cat1 through the up - sampling convolutional layer, and the up - sampling module is connected to the image quality enhancement module through the up - sampling LeakyReLU layer.
[0126] Furthermore, Figure 2 An embodiment of the image quality enhancement module is also shown in [the figure]. The image quality enhancement module includes an image quality enhancement first convolutional layer, an image quality enhancement LeakyReLU layer, and an image quality enhancement second convolutional layer. The image quality enhancement first convolutional layer, the image quality enhancement LeakyReLU layer, and the image quality enhancement second convolutional layer are connected in sequence. In the figure, Conv4 is the image quality enhancement first convolutional layer, LRU2 is the image quality enhancement LeakyReLU layer, and Conv5 is the image quality enhancement second convolutional layer. Both the image quality enhancement first convolutional layer and the image quality enhancement second convolutional layer can be convolutional layers with a convolution kernel size of 3×3. The image quality enhancement module is connected to the up - sampling module through the image quality enhancement first convolutional layer and is adaptively connected to the registration alignment module through the image quality enhancement second convolutional layer.
[0127] For the above - mentioned denoiser, through residual connection and dense connection, the multi - level residual dense connection module enables the stamp image enhancement network to capture more image details and noise features when denoising. The residual connection can better learn the distribution of noise and the method of denoising, while the dense connection ensures the efficient transmission of features and avoids information loss. The up - sampling module formed by combining convolution and non - linear activation functions can improve the resolution of the image after denoising by the multi - level residual dense connection module, and at the same time retain and enhance the detail information of the image. The image quality enhancement module constructed based on convolutional layers and non - linear activation functions, where the convolutional layer is responsible for extracting high - level features and the activation function provides non - linear transformation, enhancing the contrast and details of the image, making the denoised image clearer and cleaner.
[0128] In an embodiment of the present invention, when the registration alignment module performs stamp correction and registration processing on the denoised stamp image to be verified, it includes:
[0129] Extract the pattern center point of the denoised stamp image to be verified;
[0130] Generate registration control point A and registration control point B based on the pattern center point of the denoised stamp image to be verified and the key contour points of the central pattern of the denoised stamp image to be verified;
[0131] Generate the correction and registration angle of the denoised stamp image to be verified based on the pattern center point, configuration control point A, and registration control point B,
[0132] wherein, the correction and registration angle is:
[0133]
[0134] wherein, θ jz is the correction and registration angle, θ B is the polar coordinate angle of registration control point B, θ A is the polar coordinate angle of registration control point A;
[0135] Rotate the denoised stamp image to be verified to the target state based on the correction and registration angle, so as to generate the stamp image to be inspected after rotating to the target state.
[0136] Figure 2 The GeomPix registration module in is the registration alignment module. When performing stamp correction and registration processing, it is necessary to first determine the center point and the stamp control points. Thereafter, the correction and registration angle is determined based on the determined center point and the stamp control points, and the correction and registration processing is implemented based on the determined correction and registration angle. The specific processes of extracting the center point and registration control points A and B will be described in detail below.
[0137] Generally, there is a stamp pattern in the central area of the stamp. The stamp pattern is mostly a five-pointed star. Of course, the pattern in the central area of the stamp can also be other shapes. Here, the process of extracting the central area and registration control points A and B will be described by taking the five-pointed star pattern as an example.
[0138] As can be seen from the above description, the stamp image to be verified will be processed by the enhancement network input processing module and the denoiser to obtain the denoised stamp image to be verified. For the central area of the denoised stamp image to be verified, which includes the contour points of several stamp patterns, the contour points of all stamp patterns are subjected to the existing common circle fitting method, and the center point of the stamp pattern can be determined, that is, the pattern center point can be determined. Specifically, the method and process of using circle fitting to determine the pattern center point can be the same as the existing ones.
[0139] After determining the pattern center point of the seal pattern, calculate the Euclidean distances between the pattern center point and all contour points, so as to determine the key contour points of the seal pattern based on the calculated Euclidean distances; specifically, since the seal pattern is a pentagram, therefore, the key contour points are the five vertices of the pentagram. When determining the key contour points, the five contour points corresponding to the five largest Euclidean distances from the pattern center point are configured as the key contour points of the central pattern, and the five key contour points of the central pattern form a key contour point group.
[0140] Since there are five key contour points of the central pattern, therefore, when determining the registration control points A and B based on the pattern center point and the key contour points of the central pattern, it includes:
[0141] Arbitrarily select two key contour points of the central pattern from the key contour point group, the two selected key contour points of the central pattern and the pattern center point form a fan-shaped area, and count the pixel values of the fan-shaped area;
[0142] Repeat the above operation of arbitrarily selecting two key contour points of the central pattern from the key contour point group and counting the pixel values, and determine the fan-shaped area with the smallest counted pixel value;
[0143] Configure the two key contour points of the central pattern corresponding to the fan-shaped area with the smallest pixel value as the registration control point A and the registration control point B.
[0144] From the above description, it can be seen that the registration control point A and the registration control point B are two key points of the central pattern of the seal pattern, that is, two vertices of the pentagram. After determining the registration control point A and the registration control point B, the coordinates of the registration control point A and the registration control point B in the pixel coordinate system of the seal image to be verified can be determined. Specifically, the coordinates of the registration control point A are (x A , y A ), and the coordinates of the registration control point B are (x B , y A ). After that, the corresponding polar coordinate angles of the registration control point A and the registration control point B can be obtained. Specifically,
[0145]
[0146] wherein, θ B is the polar coordinate angle of the registration control point B, and θ A is the polar coordinate angle of the registration control point A.
[0147] After obtaining the polar coordinate angle θ B of the registration control point B and the polar coordinate angle θ A of the registration control point A, the correction registration angle can be determined. Among them, the correction registration angle is:
[0148]
[0149] During specific implementation, the corrected registration angle θ is obtained. jz After that, based on the corrected registration angle θ jz , the denoised seal image to be verified is rotated to the target state, so as to generate a seal image to be inspected after being rotated to the target state. The seal image to be inspected is the Figure 2 image output by the registration alignment module in
[0150] In an embodiment of the present invention, the dual-stream image classification network includes a dual-stream backbone network and a classification network output module connected in sequence, wherein
[0151] the dual-stream backbone network includes an EfficientNet B0 module for global feature and local feature extraction, an SVIT module for global dependency extraction, and a backbone network channel splicer for channel splicing.
[0152] For the superimposed seal image, the superimposed seal global feature and the superimposed seal local feature are extracted through the EfficientNet B0 module, and the superimposed seal global dependency is extracted through the SVIT module.
[0153] The superimposed seal global feature, the superimposed seal local feature, and the superimposed seal global dependency are channel-spliced through the backbone network channel splicer, and are classified and recognized through the classification network output module, so as to obtain seal verification information after classification and recognition output.
[0154] Specifically, the dual-stream image classification network includes a dual-stream backbone network and a classification network output module. The dual-stream backbone network includes an EfficientNet B0 module, an SVIT module, and a backbone network channel splicer. Figure 4 An embodiment of the dual-stream backbone network is shown in. In the figure, EfficientNet B0 is the EfficientNet B0 module of the dual-stream backbone network, SVIT is the SVIT module, and Cat6 is the backbone network channel splicer.
[0155] When using the dual-stream image classification network to perform feature extraction and classification processing on the superimposed seal image, the superimposed seal image global feature and the superimposed seal local feature are extracted through the EfficientNet B0 module, and the global dependency of the superimposed seal image is extracted through the SVIT module. The superimposed seal global feature, the superimposed seal local feature, and the superimposed seal global dependency are channel-spliced through the backbone network channel splicer, and are classified and recognized through the classification network output module, so as to obtain seal verification information after classification and recognition output. The following is combined with Figure 5An explanation is given for the specific details of the EfficientNet B0 module, SVIT (Streamlining Vision Transformer) module, and classification network output module of the present invention.
[0156] In an embodiment of the present invention, the EfficientNet B0 module includes a first convolutional layer for feature extraction, an automatic inverted residual bottleneck convolutional unit group, a second convolutional layer for feature extraction, a pooling layer, and a global extraction fully connected layer, which are connected in sequence. Among them,
[0157] When extracting the global feature and local feature of the superimposed seal image, the superimposed seal image is processed successively through the first convolutional layer for feature extraction, the automatic inverted residual bottleneck convolutional unit group, the second convolutional layer for feature extraction, the pooling layer, and the global extraction fully connected layer, and the global feature and local feature of the superimposed seal image are output through the global extraction fully connected layer;
[0158] The automatic inverted residual bottleneck convolutional unit group includes a number of automatic inverted residual bottleneck convolutional units, and the automatic inverted residual bottleneck convolutional units are connected in series in sequence.
[0159] Figure 4 Among them, Conv11 is the first convolutional layer for feature extraction, and the first convolutional layer for feature extraction can adopt a convolutional layer with a kernel size of 3×3. Conv12 is the second convolutional layer for feature extraction, and the second convolutional layer for feature extraction can adopt a convolutional layer with a kernel size of 1×1. The part between the first convolutional layer for feature extraction and the second convolutional layer for feature extraction is the automatic inverted residual bottleneck convolutional unit group. Figure 4 Among them, QN1 is the global feature extraction fully connected layer, and the layer between the global feature extraction fully connected layer and the second convolutional layer for feature extraction is the pooling layer.
[0160] Specifically, the first convolutional layer for feature extraction can adopt a convolutional layer with a kernel of 3×3, a stride of 2, and a padding size of 1, so as to quickly reduce the spatial resolution of the superimposed seal image and effectively extract the low-level features of the superimposed seal image. The automatic inverted residual bottleneck convolutional unit group gradually extracts higher-level features of the image, and retains important local features and global features. The second convolutional layer for feature extraction adopts a convolutional layer with a kernel of 1×1 and a stride of 1 to further integrate the feature map output by the automatic inverted residual bottleneck convolutional unit group and extract higher-level features. The pooling layer is used to downsample the feature map, reduce the spatial dimension of the feature map, retain important feature information, and at the same time reduce the computational complexity. The global extraction fully connected layer flattens the pooled feature map into a one-dimensional vector.
[0161] In specific implementation, the automatic inverted residual bottleneck convolution unit group includes several automatic inverted residual bottleneck convolution units, and the automatic inverted residual bottleneck convolution units are connected in series in sequence. Figure 4 An embodiment of the automatic inverted residual bottleneck convolution unit group is shown in Figure 4 . In the figure, the automatic inverted residual bottleneck unit group includes 7 automatic inverted residual bottleneck convolution units connected in series in sequence. Among them, MBC1 is the first automatic inverted residual bottleneck convolution unit, MBC2 is the second automatic inverted residual bottleneck convolution unit, MBC3 is the third automatic inverted residual bottleneck convolution unit, MBC4 is the fourth automatic inverted residual bottleneck convolution unit, MBC5 is the fifth automatic inverted residual bottleneck convolution unit, MBC6 is the sixth automatic inverted residual bottleneck convolution unit, and MBC7 is the seventh automatic inverted residual bottleneck convolution unit.
[0162] In specific implementation, the automatic inverted residual bottleneck convolution unit includes at least one automatic inverted residual splicing convolution basic unit. Figure 4 In Figure 4 , the automatic inverted residual bottleneck convolution unit MBC1 includes one automatic inverted residual splicing convolution basic unit, the automatic inverted residual bottleneck convolution unit MBC2 includes 2 automatic inverted residual splicing convolution basic units, the automatic inverted residual bottleneck convolution unit MBC3 includes 2 automatic inverted residual splicing convolution basic units, the automatic inverted residual bottleneck convolution unit MBC4 includes 3 automatic inverted residual splicing convolution basic units, the automatic inverted residual bottleneck convolution unit MBC5 includes 2 automatic inverted residual splicing convolution basic units, the automatic inverted residual bottleneck convolution unit MBC6 includes 4 automatic inverted residual splicing convolution basic units, and the automatic inverted residual bottleneck convolution unit MBC7 includes 1 automatic inverted residual splicing convolution basic unit; among them, when the automatic inverted residual bottleneck convolution unit includes multiple automatic inverted residual splicing convolution basic units, the multiple automatic inverted residual splicing convolution basic units are connected in series in sequence. For example, the automatic inverted residual bottleneck convolution unit MBC6 includes 4 automatic inverted residual splicing convolution basic units, then in the automatic inverted residual bottleneck convolution unit MBC6, the 4 automatic inverted residual splicing convolution basic units are connected in series in sequence. Other situations can refer to the description here and will not be elaborated further.
[0163] In an embodiment of the present invention, the automatic inverted residual bottleneck convolution basic unit includes a first convolution layer in the bottleneck convolution, a depthwise separable convolution layer, a squeeze-and-excitation module, a second convolution layer in the bottleneck convolution, a bottleneck convolution Dropout layer, and a bottleneck convolution adder, where
[0164] The input end of the first convolution layer in the bottleneck convolution and the first addition end of the bottleneck convolution adder are connected to each other to serve as the input end of the automatic inverted residual bottleneck convolution unit.
[0165] The output of the first convolutional layer in the bottleneck convolution is loaded into the depthwise separable convolutional layer after batch normalization;
[0166] The output of the depthwise separable convolutional layer is loaded into the squeeze-and-excitation module after batch normalization. The output end of the squeeze-and-excitation module is connected to the second addition end of the bottleneck convolution adder through the second convolutional layer and the bottleneck convolution Dropout layer in the bottleneck convolution;
[0167] The output end of the bottleneck convolution adder serves as the output end of the automatic inverted residual bottleneck convolution basic unit where it is located.
[0168] Specifically, the automatic inverted residual bottleneck convolution basic unit may include the first convolutional layer in the bottleneck convolution, the depthwise separable convolutional layer, the squeeze-and-excitation module, the second convolutional layer in the bottleneck convolution, the bottleneck convolution Dropout layer, and the bottleneck convolution adder. Figure 5 An embodiment of the automatic inverted residual bottleneck convolution basic unit is shown in. In the figure, Conv14 is the first convolutional layer in the bottleneck convolution, SDConv is the depthwise separable convolutional layer, and the depthwise separable convolutional layer can adopt a convolutional layer with a kernel size of k×k. SE is the squeeze-and-excitation module, Conv15 is the second convolutional layer in the bottleneck convolution, DP3 is the bottleneck convolution Dropout layer, and AD1 is the bottleneck convolution adder. Both the first convolutional layer and the second convolutional layer in the bottleneck convolution can adopt convolutional layers with a kernel size of 1×1.
[0169] In specific implementation, the first convolutional layer in the bottleneck convolution can adopt a convolutional layer with a kernel of 1×1 and a stride of 1. By means of the first convolutional layer in the bottleneck convolution, the number of channels of the feature map can be increased, enabling subsequent depth convolutions to capture more features. The depthwise separable convolutional layer is used to extract complex local features. By decomposing the standard convolution into a depth convolution and a pointwise convolution, the computational amount and the number of parameters are significantly reduced while important feature information is retained. The squeeze-and-excitation module is used to adaptively adjust the channel weights of the feature map to enhance important features and suppress irrelevant features. The second convolutional layer in the bottleneck convolution compresses the number of channels of the feature map output by the squeeze-and-excitation module. The bottleneck convolution Dropout layer randomly sets the values of some feature maps to zero, reducing the model's dependence on specific features and enhancing the generalization ability of the model. The bottleneck convolution adder adds the input feature map of the module to the feature map processed through operations such as the expansion convolution, the depth convolution, and the squeeze-and-excitation module, helping to maintain the original feature information, solve the problem of vanishing gradients, and accelerate the training process.
[0170] For Figure 4 the automatic inverted residual bottleneck convolution unit in, the k×k convolutional layers and the ratios adopted in the depthwise separable convolutional layer in each automatic inverted residual bottleneck convolution unit are different. For example, Figure 4In it, k3×3 in the automatic inverted residual bottleneck convolution unit MBC1 means that the depthwise separable convolution layer in the automatic inverted residual bottleneck convolution unit MBC1 uses a 3×3 convolution layer; k3×3 in the automatic inverted residual bottleneck convolution unit MBC2 means that there are two automatic inverted residual splicing convolution basic units in the automatic inverted residual bottleneck convolution unit MBC2, and each automatic inverted residual splicing convolution basic unit is composed of 6 stacked depthwise separable convolution layers, where each depthwise separable convolution layer uses a 3×3 convolution layer. For other cases, reference can be made to the description here, and no further examples will be given one by one.
[0171] In one embodiment of the present invention, the SVIT module includes a block embedder, a position embedder, a Transformer encoder, and a module layer normalization unit connected in sequence, where
[0172] When extracting the global dependencies of the superimposed seal image, the superimposed seal image is processed by a block embedder, a position embedder, a Transformer encoder, and a module layer normalization unit in sequence, where
[0173] The block embedder is used to divide the superimposed seal image to generate several seal image blocks after division, and at the same time, the information of the seal image is converted into a high-dimensional embedding representation to realize the linear projection of the seal image blocks;
[0174] The position embedder is used to determine the position information of each seal image block, and all the seal image blocks are connected in sequence to form an image block feature sequence;
[0175] The Transformer encoder is used to process the image block feature sequence, and the superimposed seal global dependencies are generated by the module layer normalization unit.
[0176] Figure 4 An embodiment of the SVIT module is also shown in. In the figure, the SVIT module includes a block embedder, a position embedder, a Transformer encoder, and a module layer normalization unit connected in sequence, where the block embedder is used to divide the superimposed seal image to generate several seal image blocks; the position embedder is used to determine the position information of each seal image block, and all the seal image blocks are connected in sequence to form an image block feature sequence; the Transformer encoder is used to process the image block feature sequence, and the superimposed seal global dependencies are generated by the module layer normalization unit.
[0177] Figure 4An embodiment of the block embedder is shown. The block embedder includes a block embedding convolutional layer and a rearrangement operation layer. The block embedding convolutional layer is connected to the rearrangement operation layer. The input end of the block embedding convolutional layer serves as the input end of the block embedder, and the output end of the rearrangement operation layer serves as the output end of the block embedder. In the figure, Conv13 is the block embedding convolutional layer. The block embedding convolutional layer can adopt a convolutional layer with a convolutional kernel size of 4×4 and a stride of 4. It divides the superimposed seal image into multiple non-overlapping image blocks, that is, obtains multiple seal image blocks. At the same time, each seal image block is mapped to a high-dimensional feature space through the block embedding convolutional layer, converting the information of the superimposed seal image into a high-dimensional embedding representation, enhancing the expression ability of features, that is, realizing the linear projection of the seal image block. The rearrangement operation layer flattens the feature map output by the block embedding convolutional layer, that is, flattens the seal image block, converting the spatial structure into a sequence form to meet the requirements of the Transformer encoder for subsequent encoding processing. The rearrangement operation layer can adopt the commonly used existing form.
[0178] For the position embedder, it includes a token mechanism module and a position embedding module. Among them, when the position embedder is connected to the block embedder, the output end of the rearrangement operation layer is connected to the input end of the token mechanism module. The token mechanism module is connected to the position embedding module, and the position embedding module is connected to the Transformer encoder. Specifically, the token mechanism module is a learnable parameter used to capture the global information of the feature map. It is concatenated with the seal image block during the forward propagation of the SVIT module to guide the SVIT module to perform the authenticity classification task of the seal image to be verified. The position embedding module provides position information for each seal image block to ensure that the SVIT module can utilize the spatial structure of the image for effective feature learning.
[0179] In an embodiment of the present invention, the Transformer encoder includes a plurality of sequentially cascaded Transformer encoding units, where
[0180] For any two cascaded Transformer encoding units, along the cascading direction of the Transformer encoding units, the output end of the previous Transformer encoding unit is connected to the input end of the next Transformer encoding unit;
[0181] For any Transformer encoding unit, it includes a multi-head attention mechanism unit and a multi-layer perceptron module;
[0182] The input end of the multi-head attention mechanism unit serves as the input end of the Transformer encoding unit where it is located. The multi-head attention mechanism unit is adaptively connected to the multi-layer perceptron module;
[0183] The input end of the multi-head attention mechanism unit and the output end of the multi-head attention mechanism unit are respectively connected to the corresponding input ends of the first adder of the encoder. The output end of the first adder of the encoder is connected to the input end of the multi-layer perceptron module.
[0184] The output end of the first adder of the encoder and the output end of the multi-layer perceptron module are connected to the corresponding input ends of the second adder of the encoder. The output end of the second adder of the encoder is connected as the output end of the corresponding Transformer encoding unit.
[0185] Specifically, the Transformer encoder may include a number of sequentially connected Transformer encoding units. Figure 4 An embodiment of the Transformer encoder is shown in the figure. In the figure, the Transformer encoder may include 6 sequentially connected Transformer encoding units, that is Figure 4 The "6×" in the Transformer encoder in the figure represents an embodiment in which the Transformer encoder includes 6 sequentially connected Transformer encoding units. In addition, the number of Transformer encoding units can be selected according to actual needs. For example, the number of Transformer encoding units can be more or less than 6, and the number of Transformer encoding units should be based on meeting actual application requirements.
[0186] Specifically, increasing the number of Transformer encoding units usually enhances the feature expression ability of the SVIT module. Each Transformer encoding unit performs feature transformation through the self-attention mechanism and the feed-forward neural network, and can capture more complex context information and high-level feature representations. As the number of Transformer encoding units increases, the model can learn more details and global dependencies, thereby improving the ability to understand and process image data.
[0187] Figure 7 A structural block diagram of an embodiment of the Transformer encoder is shown in the figure. In the figure, each Transformer encoding unit includes a multi-head attention mechanism unit, a first adder of the encoder, a multi-layer perceptron module (Multilayer Perceptron, MLP), and a second adder of the encoder. The multi-head attention mechanism unit can adopt the existing common form. The multi-layer perceptron module includes a first linear layer of the perceptron, a GELU layer, a second linear layer of the perceptron, and a dropout layer of the perceptron.
[0188] Figure 7An embodiment of the multi-layer perceptron module is shown. In the figure, LN1 is the first linear layer of the perceptron, LN2 is the second linear layer of the perceptron, and DP1 is the Dropout layer of the perceptron. The input end of the first linear layer of the perceptron serves as the input end of the multi-layer perceptron module, and the Dropout layer of the perceptron serves as the output end of the multi-layer perceptron module. Therefore, when the output end of the first adder of the encoder is connected to the input end of the multi-layer perceptron module, it is the output end of the first adder of the encoder connected to the input end of the first linear layer of the perceptron; when the output end of the multi-layer perceptron module is connected to the corresponding input end of the second adder of the encoder, specifically, it means the output end of the Dropout layer of the perceptron is connected to the corresponding input end of the second adder of the encoder.
[0189] Figure 6 FIG. shows a schematic diagram of an embodiment of the operation of the SVIT module of the present invention. From the above description, it can be seen that the superimposed seal image will be processed sequentially by the block embedder, position embedder, Transformer encoder, and module layer normalization layer. Figure 4 and Figure 6 In FIG., CG1 is the module layer normalization layer.
[0190] Specifically, the block embedder in the block embedder divides the superimposed seal image. The input image is divided into multiple non-overlapping image blocks. Each image block is mapped to a high-dimensional feature space through a convolutional layer, converting the information of the input image into a high-dimensional embedding representation. Finally, through the rearrangement operation layer, the image blocks output by the block embedder convolutional layer are converted into a sequence format. Among them, the rearrangement operation layer converts the image blocks output by the block embedder convolutional layer into a sequence format, which mainly realizes the flattening process and mainly meets the effective processing and analysis of the Transformer encoder.
[0191] Specifically, when the block embedder divides the superimposed seal image into blocks, the superimposed seal image is segmented into N 4x4 seal image patches (Patches) to provide a finer-grained feature representation, making it easier to capture the subtle features of the superimposed seal image. Therefore, Figure 6 The block-based image representation in FIG. means that the superimposed seal image is divided into N seal image patches by the convolutional layer of the block embedder. Among them, the number N of the seal image patches is determined by the size of the superimposed seal image and the convolutional kernel size of the block embedder convolutional layer.
[0192] Figure 6Flattening of the seal image blocks in it specifically means that each seal image block is flattened into a one-dimensional vector by using the rearrangement operation layer. For each 4×4 seal image block, the length of the flattened vector is 4×4×3 = 48. The vector of each flattened seal image block is converted into an embedding vector with a fixed dimension (128). Specifically, the tensor shape output by the block embedding convolutional layer is [16, 128, 100, 100] ([number of batch images, channel dimension, number of image blocks (in the height direction), number of image blocks (in the width direction)]). The flattening and arrangement of the seal image blocks become [16, 10000, 128] to be converted into a three-dimensional tensor format suitable for the Transformer encoder. As can be seen from the above description, the tensor shape output by the block embedding convolutional layer is a four-dimensional vector, and the rearrangement operation layer is used to convert the four-dimensional vector into a three-dimensional vector to meet the requirements of the Transformer encoder.
[0193] Each seal image block is mapped to a high-dimensional space through the linear projection operation of the block embedding convolutional layer, and a seal image block can be represented by a Token (each Token represents a local area of the superimposed seal image). At the same time, in order to retain the spatial structure information of the superimposed seal image, position encoding vectors are performed through the position embedder to represent the position information of each Token. The Tokens represented by all seal image blocks are connected in sequence to form an image block feature sequence.
[0194] As can be seen from the above description, the token mechanism is a learnable parameter. Through the token mechanism, class embedding can be performed. Specifically, a classification vector (tensor shape [16, 1, 128]) is added to the one-dimensional vector generated by the rearrangement operation layer to represent the overall information of the image. After adding the learnable class embedding, the shape becomes [16, 10001, 128], that is, a seal block embedding vector is formed.
[0195] Position embedding is used to add position information to the seal block embedding vector to assist in understanding the position of the seal image blocks in the superimposed seal image, thereby capturing the relative relationship and global context between the seal image blocks. It completely matches the shape of the input block embedding vector, ensuring that the position embedding can be correctly added to each seal image block and classification token. After the position embedding process, the shape of the seal block embedding vector is still [16, 10001, 128]. Thus, it can be seen that an image block feature sequence can be generated through the position embedder.
[0196] The sequence of image patch features generated by the position embedder is fed into the Transformer encoder. Each Transformer encoder consists of a multi-head self-attention layer and a multi-layer perceptron module. Among them, the multi-head self-attention layer is used to capture the global dependencies in the sequence of image patch features. The Tokens in the sequence of image patch features interact through the self-attention mechanism to obtain rich context information to determine the importance of the Tokens in the entire overlaid seal image.
[0197] The multi-layer perceptron module is used to further non-linearly transform and process the features in the local feature space. In the multi-layer perceptron module, high-dimensional mapping (HMLP) is adopted, and the hyperparameter γ is introduced. The input feature map x is mapped from m dimensions to a higher-dimensional feature space of m * γ. Considering the increase in the model's computational amount and without loss of generality, γ = 4 is selected. The calculation method of the multi-layer perceptron module is as follows:
[0198]
[0199] Among them, W 1 and b 1 represent the weight matrix and bias vector of the first linear layer of the perceptron, W 2 and b 2 represent the weight matrix and bias vector of the second linear layer of the perceptron. x is the feature map input to the first linear layer of the perceptron, and GELU is the operation processed using the GELU layer.
[0200] It can be understood that when high-dimensional mapping is introduced, the dimension of the weight matrix of the linear transformation changes from m dimensions to m * γ dimensions, and the bias vector changes from 1 dimension to γ dimensions, enabling the SVIT module to have greater freedom and stronger representation ability during the feature learning process. Thus, it can better adapt to different data distributions and feature patterns, helping the SVIT module learn more discriminative and generalizable feature representations.
[0201] The classification network output module at least includes a fully-connected layer of the classification network. Figure 4 An embodiment of the classification network output module is also shown in the figure. In the figure, QN2 is the fully-connected layer of the classification network. The fully-connected layer of the classification network is connected to the output end of the backbone network channel splicer. Through the fully-connected layer of the classification network, the classification information of the authenticity of the seal can be output, where the classification information of the authenticity of the seal is the seal verification information.
[0202] As can be seen from the above description, the seal verification model at least includes a seal image enhancement network and a two-stream image classification network. In order to form the seal verification model, generally, a seal image enhancement basic network and a two-stream image classification basic network need to be constructed respectively. After training the seal image enhancement basic network, a seal image enhancement network can be generated, and after training the two-stream image classification basic network, a two-stream image classification network can be generated. In an embodiment of the present invention, after constructing the seal image enhancement basic network and the two-stream image classification basic network, the seal image enhancement basic network and the two-stream image classification basic network are trained respectively. After training the seal image enhancement network and the two-stream image classification network, the seal image enhancement network and the two-stream image classification network are then connected in series to form a seal verification model.
[0203] The following specifically describes the process of training the seal image enhancement basic network and the two-stream image classification basic network and generating the seal image enhancement network and the two-stream image classification network respectively.
[0204] It can be understood that the constructed seal image enhancement basic network should be consistent with the seal image enhancement network. For the situation of the seal image enhancement basic network, reference can be made to the above description of the seal image enhancement network. In order to meet the requirements for training and generating the seal image enhancement network, a discriminator should also be set in the seal image enhancement basic network, and the discriminator is connected to the denoiser.
[0205] When training the seal image enhancement basic network, it is necessary to first construct a seal verification training dataset. Among them, the seal verification training dataset can be constructed based on 9 sets of seal bodies, with a total of 18 seal bodies. The 9 sets of seal bodies can be 6 sets of circular seals and 3 sets of oval seals. Specifically, 1 set of seal bodies includes two seals. The two seals in 1 set of seal bodies can have the same shape and text, such as both seals in two sets of seal bodies are circular seals or both are oval seals, but the details of the two seals are different, such as there are deviations in the seal patterns and text of the two seals. The deviations can be width, text content, and / or position, etc. The seal pattern can be the five-pointed star mentioned above. Of course, the number of seal bodies used to construct the seal verification training dataset can be selected according to actual needs, as long as it can meet the training requirements.
[0206] The seal verification training data set includes a number of network training seal images, which can be obtained by manually stamping the above-mentioned seal body on a paper document and then scanning it. The constructed seal verification training data set may contain 739 network training seal images. The 739 network training seal images may include noisy seal images and their corresponding clean seal images. Among them, the noisy seal images generally refer to seal images with natural text interference and computer-generated optical character interference. It should be noted that the network training seal images formed by a group of seal bodies have the same seal image name, and the acquisition method of the seal image name can refer to the above description.
[0207] When training the basic network for seal image enhancement, the Adam optimizer is adopted and combined with the momentum method and adaptive learning rate adjustment, which can effectively handle the problems of high-dimensional parameter space and sparse gradients. Figure 8 Figure 5 shows a schematic diagram of an embodiment when training the basic network for seal image enhancement. Figure 8 In this figure, the denoiser and the discriminator work in parallel and are indirectly connected through the adversarial loss. The denoiser generates the training denoised image, and the discriminator receives the training denoised image generated by the denoiser and the clean seal image and tries to distinguish between the two. Figure 8 In this figure, the seal training sample is the noisy seal image. The seal training sample is loaded into the denoiser, and the corresponding adversarial loss and perceptual loss are calculated. The discriminator discriminates the training denoised image based on the clean seal image and calculates the corresponding discriminative loss.
[0208] In specific implementation, by minimizing the adversarial loss and the perceptual loss, the weights of the denoiser are updated to make the generated training denoised image as close as possible to the clean seal image, so that the discriminator is difficult to distinguish the generated training denoised image and the clean seal image. By optimizing the adversarial loss, the discrimination ability of the discriminator is improved, and the classification accuracy when identifying the clean seal image and the training denoised image is enhanced.
[0209] Specifically, the calculation methods of the adversarial loss, the perceptual loss, and the discriminative loss can be consistent with the existing ones. During training, the training state of the basic network for seal image enhancement is determined by the discriminative loss. If the discriminative loss is stable and tends to a lower value and no longer significantly decreases or increases, it can be considered that the training of the basic network for seal image enhancement reaches the target state. At this time, the seal image enhancement network can be generated based on the basic network for seal image enhancement.
[0210] In an embodiment of the present invention, the constructed two-stream image classification basic network can refer to the description of the above two-stream image classification network. In order to meet the requirements of training to generate the two-stream image classification network, Figure 4An embodiment is shown in which the dual-stream image classification basic network further includes a data distribution adapter and a gradient reversal layer. The data distribution adapter includes a number of adapter unit bodies connected in sequence. Figure 4 It is shown that the data distribution adapter may include 4 adapter unit bodies connected in sequence. Specifically, for any adapter unit body, it includes an adapter linear layer, a batch normalization unit, a ReLU unit, and an adapter Dropout layer connected in sequence. Figure 4 In this case, LN3 is the adapter linear layer and DP2 is the adapter Dropout layer.
[0211] Figure 9 An embodiment schematic diagram of training the dual-stream image classification basic network is shown. Among them, the above-mentioned seal verification training data set can be used to train the dual-stream image classification basic network. During training, due to the significant difference in the data distribution of 9 pairs of seal bodies, it may lead to insufficient generalization ability of the dual-stream image classification network when encountering new seal images. To solve this problem, a dual-stream network training sample set and a dual-stream network test set can be constructed based on the seal verification training data set. Among them, the network training seal images corresponding to a pair of seal bodies can be used as the dual-stream network training and test set, and the network training seal images corresponding to the remaining pairs of seal bodies can be used to construct the first dual-stream network training sample set and the second dual-stream network training sample set. Both the first dual-stream network training sample set and the second dual-stream network training sample set include at least one set of network training seal images of oval seals, and the number of seal bodies included in the first dual-stream network training sample set and the second dual-stream network training sample set is preferably the same. When using 9 groups of seal bodies to construct the seal verification training data set, then both the first dual-stream network training sample set and the second dual-stream network training sample set include all the network training seal images of 4 groups of seal bodies.
[0212] In an embodiment of the present invention, a training method of multi-task learning is adopted to train the dual-stream image classification basic network and the data distribution adapter simultaneously. Through this strategy, the dual-stream image classification network can not only learn the feature differences of seals in the authenticity classification task, but also capture the common features between different seals through the adversarial training of the data distribution adapter, thereby improving the generalization ability of the dual-stream image classification network and enabling it to maintain good performance on unseen seal images.
[0213] In order to meet the above training method for multi-task learning, it is necessary to label the first sample set for training the dual-stream network and the second sample set for training the dual-stream network. For example, the labels of all network training seal images in the first sample set for training the dual-stream network can be set to "0", and the labels of all network training seal images in the second sample set for training the dual-stream network can be set to "1". Specifically, after labeling the first sample set for training the dual-stream network and the second sample set for training the dual-stream network, using the first sample set for training the dual-stream network and the second sample set for training the dual-stream network, the feature extractor of the dual-stream image classification basic network is adversarially trained through a data distribution adapter. The feature extractor in the dual-stream image classification basic network is composed of an EfficientNet B0 module and an SVIT module, which recognizes and learns the common features between different seals, enhancing the generalization ability of the model.
[0214] When training the dual-stream image classification basic network, in one embodiment of the present invention, the corresponding network training seal images in the first sample set for training the dual-stream network and the second sample set for training the dual-stream network are input into the dual-stream image classification basic network for training, and then the trained dual-stream image classification basic network is tested using the test set for training the dual-stream network. At this time, one epoch of training is completed. In addition, the Adam optimizer is used for the network training of the dual-stream image classification basic network, achieving efficient feature learning and recognition.
[0215] Specifically, in each epoch iterative training, for all network training seal images in the first sample set for training the dual-stream network, each time two network training seal images with the same seal name are taken as a first sample image pair, and the two network training seal images in the first sample image are respectively used as the classification network training seal image and the classification network reference seal image, and image preprocessing and image superposition processing are performed on the classification network training seal image and the classification network reference seal image to form the first set of training samples for the classification network, and the formed first set of training samples for the classification network is input into the dual-stream image classification basic network. After that, the true / false classification output information of the corresponding first sample image pair is output through the classification network output module, and the prediction information of the first sample image pair is output through the data distribution adapter.
[0216] Specifically, all network training seal images in the first sample set for training the dual-stream network can form multiple first sample image pairs. It can be understood that the two network training seal images in multiple first sample image pairs should not be exactly the same. When the constructed first sample image pair is exactly the same as the already formed first sample image pair, it indicates that all network training seal images in the current first sample set for training the dual-stream network have been utilized. At this time, the training of the dual-stream image classification basic network using the first sample set for training the dual-stream network should be stopped.
[0217] In each epoch iteration training, all the network training seal images in the second sample set are trained in the two-stream network. Each time, two network training seal images are taken as a second sample image pair. After that, the second sample image pair can be used to form the second set of training samples for the classification network. The second set of training samples for the classification network is input into the two-stream image classification basic network, and the prediction information of the second sample image pair is output through the data distribution adapter. Specifically, the situation of the second sample image pair and the second set of training samples for the classification network can refer to the corresponding descriptions of the first sample image pair and the first set of training samples for the classification network above; it can be understood that based on multiple second sample image pairs, multiple corresponding prediction information of the second sample image pairs can be obtained. That is, different from the first sample set trained by the two-stream network, when training the two-stream image classification basic network in the second sample set trained by the two-stream network, only the prediction information of the corresponding second sample image pair passing through the data distribution adapter can be output.
[0218] Through the above-mentioned true / false classification output information of the first sample image pair, the prediction information of the first sample image pair passing through the data distribution adapter, and the prediction information of the second sample image pair passing through the data distribution adapter, the corresponding loss function value after the current epoch iteration training can be calculated. In an embodiment of the present invention, the loss function can adopt VerifiLoss. Among them, when the loss function adopts the loss function VerifiLoss, there is:
[0219]
[0220] Among them, A_labels is the first label value of the first set of training samples for the classification network, A_output is the prediction information of the first sample image pair of the first set of training samples for the classification network passing through the data distribution adapter; B_labels is the first label value of the second set of training samples for the classification network, B_output is the prediction information of the second sample image pair of the second set of training samples for the classification network passing through the data distribution adapter; s_output is the true / false classification output information of the first sample image pair of the first set of training samples for the classification network; s_label is the second label value of the first set of training samples for the classification network.
[0221] As can be seen from the above description, since the labels of all the network training seal images in the first sample set for the two-stream network training can be "0", the first label value A_labels of the training samples in the first set of the classification network can be "0". When the first label value A_labels of the first sample image pair can be 0, the label value B_labels of the second sample image pair should be 1. The second label value s_label of the training samples in the first set of the classification network is related to the two training network seal images that form the training samples in the first set of the classification network. If the two training network seal images are from the same seal body, the second label value s_label of the training samples in the first set of the classification network can be 1; otherwise, the second label value s_label of the training samples in the first set of the classification network is 0.
[0222] In specific implementation, when calculating the loss function VerifiLoss, by using the second label value s_label of the training samples in the first set of the classification network and the true / false classification output information s_output of the first sample image pair in the first set of the classification network, the generated two-stream image classification network can recognize and learn the feature differences between genuine and fake seals, so as to achieve seal true / false classification.
[0223] When calculating the loss function VerifiLoss, the following takes the calculation of BCELoss(A_output, A_labels) as an example to illustrate the calculation process. Specifically:
[0224]
[0225] where M is the number of training samples in the first set of the classification network, y i is the first label value A_labels of the current training sample in the first set of the classification network, and p i is the prediction information A_output of the first sample image pair of the current training sample in the first set of the classification network after passing through the data distribution adapter.
[0226] It should be noted that when calculating BCELoss(A_output, A_labels) above, using y i to replace the first label value A_labels and using p i the prediction information A_output of the first sample image pair after passing through the data distribution adapter is only for the convenience of expression. The calculation of the above loss function VerifiLoss can refer to the description here and will not be elaborated one by one.
[0227] As can be seen from the above description, the basic network for dual-stream image classification may include a gradient reversal layer and a data distribution adapter. During each epoch of iterative training, in the forward propagation process, the gradient reversal layer directly passes the input features as an identity mapping; while in the backward propagation process, the gradient reversal layer multiplies it by a negative scalar coefficient α to reverse the gradient, so as to apply a negative gradient when optimizing the feature extractor composed of the EfficientNet B0 module and the SVIT module. At this time, it can prompt the feature extractor to ignore the features of specific seals during the training process and focus on the features irrelevant to the task of classifying the authenticity of seals, thereby weakening the influence of specific features of different seals and enhancing the generalization ability.
[0228] In addition, the linear layer in the data distribution adapter performs a linear transformation on the input data through a weight matrix and a bias vector, mapping the input features to a new space. The batch normalization unit normalizes each batch of data so that the output has zero mean and unit variance, thereby accelerating the training process and improving the stability of the basic network for dual-stream image classification. The ReLU unit in the data distribution adapter introduces non-linearity, enabling the basic network for dual-stream image classification to capture more complex feature relationships. The adapter Dropout layer in the data distribution adapter randomly discards a part of neurons with a probability of 50% during the training process, thereby effectively enhancing the generalization ability of the model and reducing the risk of overfitting.
[0229] It can be understood that during the training process, it is necessary to repeat the process of constructing the first sample set for dual-stream network training, the second sample set for dual-stream network training, and performing one epoch of iterative training, that is, generally, it is necessary to perform multiple epochs of iterative training on the basic network for dual-stream image classification. After each epoch of iterative training, through the loss function value of the loss function VerifiLoss, backward propagation is performed to update the network parameters of the basic network for dual-stream image classification.
[0230] During the process of performing multiple epochs of iterative training, by recording and observing the change of the loss function VerifiLoss during the training process, when the change of the loss function VerifiLoss is stable and tends to a lower value, and no longer significantly decreases or increases, it can be considered that the training target state is reached for the basic network for dual-stream image classification. At this time, the corresponding dual-stream image classification network can be generated according to the basic network for dual-stream image classification that reaches the target state.
[0231] In addition, in the later stage of training, averaging the network parameters of the two-stream image classification basic network in the last few epochs can smooth the parameter changes, reduce the fluctuations caused by noise and randomness, and reduce the risk of overfitting, thereby further enhancing the generalization ability of the two-stream image classification basic network. For example, when training for 3 epochs, the loss function VerifiLoss tends to stabilize. Therefore, the network parameters of the two-stream image classification network are determined by integrating the network parameters of the two-stream image classification basic network in the last 3 epochs. At this time, the overall performance of the two-stream image classification basic network on the two-stream network training and test set can be improved.
[0232] From the above description, when the network parameters of the two-stream image classification network are determined by integrating the network parameters of the two-stream image classification basic network in the last 3 epochs, we have:
[0233]
[0234] where n is the total number of epoch iterations for training the two-stream image classification basic network, current param is the network parameter of the two-stream image classification basic network in the i-th epoch iteration, and average param is the averaged network parameter. It can be understood that the network parameters of the two-stream image classification basic network are the necessary network parameters for forming the two-stream image classification network. For example, some bias parameters and weight parameters mentioned above. The situation of the necessary network parameters can be selected according to actual needs and will not be listed here.
[0235] In addition, Figure 9 in, the preprocessing of the classification network training seal image and the classification network reference seal image specifically refers to resizing the corresponding image sizes of the classification network training seal image and the classification network reference seal image to 400×400; when the corresponding images of the classification network training seal image and the classification network reference seal image are resized to 400×400, the corresponding sizes of the seal image to be verified and the reference seal image should also be 400×400. Figure 9 The image overlay method in can refer to the above description and will not be elaborated here.
[0236] Specifically, when implementing, for the two-stream image classification network trained based on the two-stream image classification basic network, it can be evaluated using the two-stream network training and test set. The evaluation indicators for the two-stream network training and test set can include precision and recall. The evaluation methods and purposes using precision and recall can be the same as those in the prior art and will not be elaborated one by one here.
Claims
1. An end-to-end dual-stream seal automatic verification method based on deep learning, characterized in that: The automatic seal verification method comprises: Acquire a seal image to be verified, and extract a corresponding reference seal image from a seal verification database based on the seal name of the seal image to be verified, wherein the reference seal image is at least consistent with the seal name of the seal image to be verified; The acquired seal image to be verified is loaded into a pre-built seal verification model, so as to verify the seal image to be verified using the seal verification model, wherein: The seal verification model at least includes a seal image enhancement network and a two-stream image classification network; When verifying the seal image to be verified, the seal image enhancement network is first used to perform image enhancement processing to generate the seal image to be verified after the image enhancement processing; The seal image enhancement network comprises an enhancement network input processing module, a denoiser and a registration and alignment module connected in sequence; The denoiser comprises a multi-level residual dense connection module for denoising the stamp image to be verified, an up-sampling module for increasing the image resolution and an image quality enhancement module for improving the image quality. Superimposing the seal image to be inspected with the reference seal image extracted above to generate a superimposed seal image, and performing feature extraction and classification processing on the superimposed seal image using a two-stream image classification network to generate seal verification information after the feature extraction and classification processing; Based on the seal verification information, determining the authenticity of the seal image to be verified; The dual-stream image classification network includes a dual-stream backbone network and a classification network output module connected in sequence, wherein: The dual-stream backbone network includes an EfficientNet B0 module for global feature local feature extraction, an SVIT module for global dependency extraction, and a backbone network channel splicer for channel splicing. For the superimposed seal image, the local features and global features of the superimposed seal are extracted through the EfficientNet B0 module, and the global dependency of the superimposed seal image is extracted through the SVIT module; The local features of the superimposed seals, the global features of the superimposed seals, and the global dependencies of the superimposed seals are channel-joined through the backbone network channel joiner, and the classification network output module is used to perform classification recognition output, so as to obtain the seal verification information after the classification recognition output; The SVIT module includes a block embedder, a position embedder, a Transformer encoder, and a module layer normalization unit connected in sequence.
2. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 1 is characterized in that: When the seal image enhancement network is used to perform image enhancement processing on the seal image to be verified, it includes image denoising processing and / or seal correction and registration processing, wherein: When the image enhancement processing of the seal image to be verified includes image denoising processing and seal correction processing, the seal image enhancement network is used to sequentially perform image denoising processing and seal correction and registration processing on the seal image to be verified, wherein: The seal image to be verified is loaded into the denoiser using the enhanced network input processing module; Performing image denoising processing on the seal image to be verified by using a denoiser, so as to generate a denoised seal image to be verified after the image denoising processing; The registration and alignment module is used to perform seal correction and registration processing on the denoised seal image to be verified, so as to generate the seal image to be verified after the seal correction and registration processing.
3. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 2 is characterized in that: The multi-level residual dense connection module, the upsampling module and the image quality enhancement module are connected in sequence, the multi-level residual dense connection module is connected to the output end of the enhancement network input processing module, and the image quality enhancement module is adaptively connected to the registration alignment module; The multi-level residual dense connection module includes a feature enhancement basic unit and a densely connected convolutional layer, wherein the feature enhancement basic unit includes a plurality of feature enhancement basic modules connected in series in sequence; The feature enhancement basic unit is connected to the densely connected convolutional layer, the output of the densely connected convolutional layer and the output of the enhanced network input processing module are both adaptively connected to the densely connected splicer, and the output end of the densely connected splicer is adaptively connected to the upsampling module.
4. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 3 is characterized in that: Any feature enhancement basic module includes a dense block unit group and a basic module residual scaling unit adaptively connected to the dense block unit group, wherein the dense block unit group includes a plurality of dense block units connected in series in sequence, The dense block unit includes a first branch unit, a second branch unit, and a dense block unit intra-splicing device for performing channel splicing on the first branch unit and the second branch unit; The first branch unit includes a dense block and a dense block residual scaling unit adaptively connected to the dense block; The input of the dense block in the first branch unit is directly loaded into the splicer in the dense block unit via the second branch unit; For two adjacent dense block units, along the serial connection direction of the dense block units, the output end of the splicer in the dense block unit of the previous dense block unit is adaptively connected to the first branch unit and the second branch unit in the next dense block unit.
5. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 2 is characterized in that: When the registration and alignment module performs the seal correction and registration processing on the denoised seal image to be verified, it includes: Extracting the center point of the pattern of the stamp image to be verified after denoising; Generate a registration control point A and a registration control point B based on the pattern center point of the denoised seal image to be verified and the key contour point of the central pattern of the denoised seal image to be verified; Based on the pattern center point, configuration control point A and registration control point B, the correction registration angle of the stamp image to be verified after denoising is generated. Wherein, the correction registration angle is: Among them, θ jz is the correction registration angle, θ B is the polar coordinate angle of the registration control point B, θ A is the polar coordinate angle of the registration control point A; Based on the correction registration angle, the denoised seal image to be verified is rotated to a target state, so as to generate the seal image to be verified after being rotated to the target state.
6. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 1 is characterized in that: The EfficientNet B0 module includes a first convolutional layer for feature extraction, an automatic inverted residual bottleneck convolution unit group, a second convolutional layer for feature extraction, a pooling layer, and a global extraction fully connected layer, which are connected in sequence. When extracting the superimposed seal global features and the superimposed seal local features of the superimposed seal image, the superimposed seal image is processed sequentially by a feature extraction first convolution layer, an automatic inverted residual bottleneck convolution unit group, a feature extraction second convolution layer, a pooling layer, and a global extraction fully connected layer, and the global extraction fully connected layer outputs the superimposed seal global features and the superimposed seal local features; The automatic inverted residual bottleneck convolution unit group includes a plurality of automatic inverted residual bottleneck convolution units, and the automatic inverted residual bottleneck convolution units are connected in series in sequence.
7. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 6 is characterized in that: The automatic inverted residual bottleneck convolution basic unit includes a first convolution layer in the bottleneck convolution, a depth-separable convolution layer, a compression excitation module, a second convolution layer in the bottleneck convolution, a bottleneck convolution Dropout layer and a bottleneck convolution adder, wherein: An input end of the first convolution layer in the bottleneck convolution and a first adding end of the bottleneck convolution adder are connected to each other to serve as an input end of the automatic inversion residual bottleneck convolution unit; The output of the first convolutional layer in the bottleneck convolution is batch normalized and then loaded into the depthwise separable convolutional layer; The output of the depthwise separable convolutional layer is loaded into the compression excitation module after batch normalization, and the output end of the compression excitation module is connected to the second adding end of the bottleneck convolution adder through the second convolutional layer in the bottleneck convolution and the bottleneck convolution Dropout layer; The output end of the bottleneck convolution adder is used as the output end of the automatic inversion residual bottleneck convolution basic unit.
8. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 6 is characterized in that: When extracting the global dependency of the superimposed seal image, the superimposed seal image is processed by the block embedder, the position embedder, the Transformer encoder, and the module layer normalization unit in sequence, where: The superimposed seal image is divided by using a block embedder to generate a number of seal image blocks after the division, and the information of the seal image is converted into a high-dimensional embedding representation to realize the linear projection of the seal image block; The position embedder is used to determine the position information of each seal image block, and all the seal image blocks are connected in sequence to form an image block feature sequence; The image block feature sequence is processed using the Transformer encoder, and the global dependency of the superimposed seal is generated through the module layer normalization unit.
9. The end-to-end dual-stream seal automatic verification method based on deep learning according to claim 6 is characterized in that: The Transformer encoder includes a plurality of Transformer encoding units connected in series, wherein: For any two serially connected Transformer coding units, along the serial connection direction of the Transformer coding units, the output end of the previous Transformer coding unit is connected to the input end of the next Transformer coding unit; For any Transformer encoding unit, including multi-head attention mechanism unit and multi-layer perceptron module; The input end of the multi-head attention mechanism unit is used as the input end of the Transformer encoding unit, and the multi-head attention mechanism unit is adaptively connected to the multi-layer perceptron module; The input end and the output end of the multi-head attention mechanism unit are respectively connected to the corresponding input end of the first adder of the encoder, and the output end of the first adder of the encoder is connected to the input end of the multi-layer perceptron module. The output end of the first adder of the encoder and the output end of the multilayer perceptron module are connected to the corresponding input end of the second adder of the encoder, and the output end of the second adder of the encoder is connected as the output end of the Transformer encoding unit.
Citation Information
Patent Citations
Seal image authenticity identification method and device based on deep learning
CN117765561A
Image processing method, terminal and computer readable storage medium
CN112396638A
Seal returning verification method and device, equipment and storage medium
CN115880710A