Credential Anti-counterfeiting detection

By combining a multimodal feature encoder and classifier for voucher images, the problem of insufficient accuracy in voucher image anti-counterfeiting detection in existing technologies is solved, achieving efficient identification and risk reduction of counterfeit voucher images.

WO2026060783A1PCT designated stage Publication Date: 2026-03-26ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing anti-counterfeiting detection solutions for voucher images have limited accuracy when faced with counterfeit voucher images. They are unable to effectively identify tampering traces in RGB format counterfeit voucher images, thus failing to protect the rights and interests of all parties.

Method used

By combining the voucher image in the spatial domain with the corresponding voucher noise feature map or voucher spectrum map, and using a voucher anti-counterfeiting detection model with a multimodal feature encoder and classifier, the voucher image is processed collaboratively to extract and fuse feature data from multiple modalities to generate accurate anti-counterfeiting detection results.

Benefits of technology

This improves the accuracy and effectiveness of anti-counterfeiting detection of vouchers, reduces the risks posed by counterfeit voucher images, and protects the rights and interests of all parties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129099_26032026_PF_FP_ABST
    Figure CN2024129099_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present description are a credential anti-counterfeiting detection method and apparatus, a computer-readable storage medium, an electronic device, and a computer program product. The credential anti-counterfeiting detection method may comprise: acquiring a first credential image obtained by performing image acquisition on a target credential of a user; processing the first credential image, and taking a generated credential noise feature map or credential spectrum map as a second credential image; and using a credential anti-counterfeiting detection model to perform collaborative anti-counterfeiting detection on multi-modal credential images such as the first credential image and the second credential image, so as to obtain an anti-counterfeiting detection result for the first credential image.
Need to check novelty before this filing date? Find Prior Art

Description

Certificate forgery detection TECHNICAL FIELD

[0001] The present specification relates to the technical field of image processing, and in particular to certificate forgery detection. BACKGROUND

[0002] With the rapid development of science and technology, the threshold for image forgery is also getting lower and lower. In the process of obtaining services, some users may submit forged certificate images to service providers for personal purposes. The behavior of these users submitting forged certificate images may involve risks such as fraud, illegal activities, etc., thereby causing potential risks and losses to society, enterprises or others. At present, service providers usually arrange staff to carefully detect and identify the authenticity of the certificate images submitted by users. Some service providers also begin to use deep learning algorithms to train models to analyze the entire certificate image submitted by the user to distinguish between real certificate images and forged certificate images. However, the above-mentioned solutions have limited accuracy of the forgery detection results generated for the certificate images, which may not guarantee the rights and interests of all parties.

[0003] Therefore, how to improve the accuracy of the forgery detection results for the certificate images of users has become a technical problem to be solved.

[0004] SUMMARY

[0005] The embodiments of the present specification provide a certificate forgery detection method and device, a computer readable storage medium, an electronic device and a computer program product, which perform forgery detection by combining a certificate image in a spatial domain and a multi-modal certificate image such as a corresponding certificate noise feature map or a certificate spectrum map, thereby improving the accuracy and effectiveness of the certificate forgery detection scheme.

[0006] The certificate forgery detection method provided by the embodiments of the present specification comprises: obtaining a first certificate image obtained by image acquisition on a target certificate of a user; processing the first certificate image to obtain a second certificate image, the second certificate image comprising: any one of a certificate noise feature map and a certificate spectrum map; performing forgery detection processing on the first certificate image and the second certificate image by using a certificate forgery detection model to obtain a first forgery detection result for the first certificate image output by the certificate forgery detection model; wherein the certificate forgery detection model comprises a multi-modal feature encoder and a classifier connected in sequence.

[0007] The embodiment of the present specification further provides a certificate anti-counterfeiting detection device, comprising: a first acquisition module configured to acquire a first certificate image obtained by image collection on a target certificate of a user; a second acquisition module configured to process the first certificate image to obtain a second certificate image, wherein the second certificate image comprises any one of a certificate noise feature map and a certificate spectrum map; and an anti-counterfeiting detection module configured to perform anti-counterfeiting detection processing on the first certificate image and the second certificate image by using a certificate anti-counterfeiting detection model to obtain a first anti-counterfeiting detection result of the first certificate image output by the certificate anti-counterfeiting detection model; wherein the certificate anti-counterfeiting detection model comprises a multi-modal feature encoder and a classifier connected in sequence.

[0008] The embodiment of the present specification further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the above method.

[0009] The embodiment of the present specification further provides an electronic device, comprising a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the steps of the above method.

[0010] The embodiment of the present specification further provides a computer program product having at least one instruction stored thereon, wherein the at least one instruction is executed by a processor to implement the steps of the above method.

[0011] In the embodiment of the present specification, when performing anti-counterfeiting detection on the first certificate image of the target certificate of the user in the spatial domain, the first certificate image can be processed to generate a certificate noise feature map or a certificate spectrum map as a second certificate image. Since the certificate noise feature map and the certificate spectrum map can reflect the counterfeiting traces in the first certificate image, the certificate anti-counterfeiting detection model can be used to cooperatively process the first certificate image and the second certificate image and other multi-modal certificate images, which is beneficial to improve the accuracy and effectiveness of the anti-counterfeiting detection result of the first certificate image, reduce the risk of using counterfeit certificate images by the user, and protect the interests of all parties. BRIEF DESCRIPTION OF DRAWINGS

[0012] FIG. 1 is a schematic diagram of an application scenario of a certificate anti-counterfeiting detection scheme according to an embodiment of the present specification;

[0013] FIG. 2 is a flowchart of a certificate anti-counterfeiting detection method according to an embodiment of the present specification;

[0014] FIG. 3 is a structural diagram of a certificate anti-counterfeiting detection model according to an embodiment of the present specification;

[0015] FIG. 4 is a structural diagram of another certificate anti-counterfeiting detection model according to an embodiment of the present specification;

[0016] FIG. 5 is a structural schematic diagram of another certificate anti-counterfeiting detection model provided by an embodiment of the present specification;

[0017] FIG. 6 is a structural schematic diagram of still another certificate anti-counterfeiting detection model provided by an embodiment of the present specification;

[0018] FIG. 7 is a structural schematic diagram of a certificate anti-counterfeiting detection device provided by an embodiment of the present specification;

[0019] FIG. 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0020] For the purpose, technical solutions and advantages of the present specification to be clearer, the technical solutions of the present specification will be described clearly and completely below in conjunction with the specific embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by a person of ordinary skill in the art without any creative work, fall within the scope of protection of the present specification.

[0021] In the process of providing services to users, there are cases where users submit counterfeit certificate images to achieve personal purposes. For example, users may forge identity cards, educational certificates, income certificates, etc. for fraudulent purposes to obtain loans, job opportunities. There are also users who may forge personal address certificates, professional certificates, etc. for personal privacy protection or other reasons to protect their privacy data. In addition, some users may forge identity cards or impersonate others to engage in illegal criminal activities or evade supervision. In general, service providers need to accurately identify the certificate images submitted by users in order to protect the rights and interests of all parties and the stable operation of the business.

[0022] Currently, service providers usually arrange staff to carefully detect and identify the authenticity of the certificate images submitted by users, and some service providers have begun to use deep learning algorithms to train models to detect the authenticity of RGB format certificate images. However, people have begun to use some auxiliary means (such as tamper repair) to fine-tune the counterfeit certificate images, making the tampering traces in the RGB format counterfeit certificate images not intuitive, to improve the tampering concealment of the RGB format counterfeit certificate images, thus easily leading to the existing certificate anti-counterfeiting detection scheme being unable to accurately identify the RGB format counterfeit certificate images.

[0023] Please refer to FIG. 1, which is an application scenario schematic diagram of a certificate anti-counterfeiting detection scheme provided by an embodiment of the present specification.

[0024] As shown in FIG. 1, in the process of obtaining a service, a user can use a user device 101 to send a first credential image 103 obtained by image acquisition of a target credential of the user to a device 102 of a service provider. When the service provider uses the device 102 to perform a credential anti-forgery detection on the first credential image 103, the service provider can first process the first credential image 103 to obtain a credential noise feature map or a credential frequency spectrum map as a second credential image.

[0025] The device 102 can also be deployed with a pre-trained credential anti-forgery detection model 104, and the credential anti-forgery detection model 104 can include a multi-modal feature encoder and a classifier connected in sequence, so that the credential anti-forgery detection model 104 can be used to perform an anti-forgery detection on the first credential image 103 and the second credential image, and obtain a first anti-forgery detection result of the first credential image 103 with high accuracy output by the credential anti-forgery detection model 104.

[0026] Please refer to FIG. 2, which is a flowchart of a credential anti-forgery detection method provided by an embodiment of the present specification. The execution subject of the flowchart can be a device for performing a credential anti-forgery detection, or an application program for performing a credential anti-forgery detection deployed at the device. The following will be a detailed description of the flowchart shown in FIG. 2. The credential anti-forgery detection method can specifically include the following steps 202 to 206.

[0027] Step 202: Obtain a first credential image obtained by image acquisition of a target credential of a user.

[0028] In the embodiment of the present specification, the first credential image of the user can be an image in a spatial domain of a user credential that needs to be detected, for example, an image in an RGB color mode, a grayscale image, etc. In actual applications, the type of the target credential can be various, and the type of the corresponding first credential image can also be various, for example, the first credential image can include but is not limited to an identity credential image, an educational credential image, a payment result credential image, an asset credential image, a real estate credential image, etc., without specific limitation.

[0029] Step 204: Process the first credential image to obtain a second credential image, the second credential image including any one of a credential noise feature map and a credential frequency spectrum map.

[0030] In the embodiments of the present specification, since the noise features between different regions in the forged certificate image are often mismatched when the partial image region in the certificate image is cut, moved, modified, or covered or filled with other images, the certificate noise feature map corresponding to the first certificate image can be used as the second certificate image to perform certificate anti-counterfeiting detection in combination with the second certificate image. The certificate noise feature map can refer to a feature map obtained by performing noise feature extraction processing on the first certificate image.

[0031] In addition, since the compression rounds experienced by different regions in the forged certificate image usually differ, the certificate image in the frequency domain can reflect the difference in the compression rounds, and therefore the certificate frequency spectrum corresponding to the first certificate image can be used as the second certificate image to perform certificate anti-counterfeiting detection in combination with the second certificate image. The certificate frequency spectrum can refer to a frequency spectrum in the frequency domain obtained by performing frequency domain conversion processing on the first certificate image.

[0032] In step 206, the first certificate image and the second certificate image are subjected to anti-counterfeiting detection processing by using the certificate anti-counterfeiting detection model, to obtain a first anti-counterfeiting detection result for the first certificate image output by the certificate anti-counterfeiting detection model. The certificate anti-counterfeiting detection model includes a multi-modal feature encoder and a classifier connected in sequence.

[0033] In the embodiments of the present specification, the certificate anti-counterfeiting detection model can be a classification model trained in advance by using certificate samples carrying label data reflecting whether the certificate samples belong to trusted certificate samples. Specifically, the multi-modal feature encoder in the certificate anti-counterfeiting detection model can perform feature extraction processing and feature fusion processing on the received first certificate image and second certificate image, to send the generated certificate image encoding data to the classifier, so that the classifier performs classification processing on the certificate image encoding data to obtain a first anti-counterfeiting detection result that can be used to reflect whether the first certificate image belongs to a trusted certificate image, which is convenient and fast.

[0034] The method in FIG. 2 can process the first certificate image when performing anti-counterfeiting detection on the first certificate image of the target certificate of the user in the spatial domain, to generate a certificate noise feature map or a certificate frequency spectrum as a second certificate image. Since the certificate noise feature map and the certificate frequency spectrum can reflect the forged traces in the first certificate image, the certificate anti-counterfeiting detection model can be used to cooperatively process the first certificate image and the second certificate image and other multi-modal certificate images, which is conducive to improving the accuracy and effectiveness of the anti-counterfeiting detection result generated for the first certificate image, reducing the risk brought by the use of forged certificate images by the user, and protecting the interests of all parties.

[0035] Based on the method in FIG. 2, the embodiments of the present specification further provide some specific implementation solutions of the method, which are described as follows.

[0036] For the convenience of understanding, the structure and working principle of the certificate anti-counterfeiting detection model are explained and described herein.

[0037] Optionally, the multi-modal feature encoder can include a first high-resolution network module, a second high-resolution network module, and a feature fusion module; the first high-resolution network module and the second high-resolution network module can be connected with the feature fusion module respectively, and the feature fusion module can be connected with the classifier.

[0038] The first high-resolution network module can be configured to perform feature extraction processing on the first certificate image to obtain a plurality of first certificate image feature data of different scales.

[0039] The second high-resolution network module can be configured to perform feature extraction processing on the second certificate image to obtain a plurality of second certificate image feature data of different scales.

[0040] The feature fusion module can be configured to perform feature fusion processing on the plurality of first certificate image feature data of different scales and the plurality of second certificate image feature data of different scales to obtain third certificate image feature data.

[0041] The classifier can be configured to perform classification processing based on the third certificate image feature data to obtain a first anti-counterfeiting detection result reflecting whether the first certificate image belongs to a trusted certificate image.

[0042] In the embodiments of the present specification, the high-resolution network (HRNet) can run each workflow for processing images of different resolutions in parallel, and enhance the final feature representation by fusing the extracted image feature maps of different resolutions. This design can enable the network to maintain high resolution while also having the powerful strength of deep learning, so the high-resolution network can be used to build a multi-modal feature encoder. Specifically, the multi-modal feature encoder can include a first high-resolution network and a second high-resolution network to perform feature extraction processing on the first certificate image and the second certificate image respectively, which is conducive to ensuring the accuracy of the extracted first certificate image feature data and second certificate image feature data.

[0043] In the embodiments of the present specification, the credential anti-counterfeiting detection model can also include a feature fusion module to perform feature fusion processing on the first credential image feature data and the second credential image feature data, and input the third credential image feature data obtained by feature fusion into a classifier (Classifier), so that the classifier can determine the probability value of the first credential image belonging to a trusted / counterfeit credential image, and further obtain a first anti-counterfeiting detection result for reflecting whether the first credential image belongs to a trusted credential image.

[0044] Optionally, the first high-resolution network module included in the multi-modal feature encoder can be an RGB Stream module in the compressed artifact tracking network model, and the second high-resolution network module can be a DCT Stream module in the compressed artifact tracking network model.

[0045] The RGB Stream module can be configured to perform feature extraction processing on the first credential image to obtain a plurality of first credential image feature data of different scales.

[0046] The DCT Stream module can be configured to perform feature extraction processing on the second credential image to obtain a plurality of second credential image feature data of different scales.

[0047] In the embodiments of the present specification, the compressed artifact tracking network model (CAT-Net) is a neural network model suitable for detecting and locating splicing regions in an image. The compressed artifact tracking network can generally include an RGB Stream module and a DCT Stream module built based on a high-resolution network (HRNet), wherein the RGB Stream module can perform feature extraction processing on an image in the spatial domain to learn visual artifact features, and the DCT Stream module can perform feature extraction processing on a frequency spectrum graph in the frequency domain to learn compressed artifact features, thereby learning image forgery features in multiple modalities such as the spatial domain and the frequency domain to accurately identify and segment the forged regions in the image.

[0048] In actual applications, since the input data of the DCT Stream module is a frequency spectrum graph, and the frequency spectrum graph and the noise feature graph both belong to two-dimensional data, and the formats of the two can also be consistent, when the noise feature graph is input into the DCT Stream module, the DCT Stream module can perform feature extraction processing and analysis on the noise feature graph to learn the image tampering features reflected in the noise feature graph.

[0049] Based on this, the RGB Stream module and the DCT Stream module in the compressed artifact tracking network model (CAT-Net) can be used as the first high-resolution network module and the second high-resolution network module in the multi-modal feature encoder, respectively, to perform feature extraction processing on the first credential image and the second credential image, which is beneficial to improving the accuracy of the extracted first credential image feature data and second credential image feature data.

[0050] In the embodiments of the present specification, there can be various implementation manners of the feature fusion module included in the multi-modal feature encoder. For ease of understanding, this is explained and described.

[0051] In implementation manner one, the feature fusion module included in the multi-modal feature encoder can be a Fusion Stage module in the compressed artifact tracking network model. The Fusion Stage module can be used to perform feature fusion processing on the plurality of scale-different first credential image feature data and the plurality of scale-different second credential image feature data, to obtain the third credential image feature data.

[0052] In the embodiments of the present specification, since the first high-resolution network module and the second high-resolution network module in the multi-modal feature encoder can be implemented by using the RGB Stream module and the DCT Stream module in the compressed artifact tracking network model (CAT-Net), and the Fusion Stage module included in the compressed artifact tracking network model (CAT-Net) has the function of performing feature fusion processing on the respective feature data output by the RGB Stream module and the DCT Stream module, therefore, the Fusion Stage module can be used as the feature fusion module in the multi-modal feature encoder, so as to perform feature fusion processing on the first credential image feature data and the second credential image feature data output by the first high-resolution network module and the second high-resolution network module in the multi-modal feature encoder. It can be understood that at this time, the structure of the multi-modal feature encoder and the compressed artifact tracking network model can be consistent, which is not only beneficial to simplifying the building process of the credential anti-counterfeiting detection model, but also beneficial to guaranteeing the accuracy of the third credential image feature data extracted by the multi-modal feature encoder.

[0053] For ease of understanding, FIG. 3 is a structural schematic diagram of a certificate anti-counterfeiting detection model provided by an embodiment of the present specification, as shown in FIG. 3, the multi-modal feature encoder in the certificate anti-counterfeiting detection model can include: a first high-resolution network module 31, a second high-resolution network module 32, and a feature fusion module 33. The above three modules can be in turn: the RGB Stream module, the DCT Stream module, and the Fusion Stage module in the CAT-Net. At this time, each output layer of the first high-resolution network module 31 and each output layer of the second high-resolution network module 32 can be connected to the feature fusion module 33, and the output layer of the feature fusion module 33 can be connected to the classifier 34.

[0054] In implementation manner two, the feature fusion module included in the multi-modal feature encoder can include: a first Transformer sub-module, a second Transformer sub-module, a Class Attention sub-module, and a multi-scale feature fusion sub-module. The multi-scale feature fusion sub-module can be a branch network in the Fusion Stage module in the compression artifact tracking network model for performing feature fusion processing on multi-scale feature maps.

[0055] Each output layer of the first high-resolution network module except the output layer for outputting the first certificate image feature data of the largest scale can be connected to the first Transformer sub-module, each output layer of the second high-resolution network module can be connected to the second Transformer sub-module, the first Transformer sub-module and the second Transformer sub-module can be connected to the Class Attention sub-module, the output layer for outputting the first certificate image feature data of the largest scale and the Class Attention sub-module can be connected to the multi-scale feature fusion sub-module, and the multi-scale feature fusion sub-module can also be connected to the classifier.

[0056] The first Transformer sub-module can be used for performing feature fusion processing on the first certificate image feature data of other scales except the largest scale to obtain fourth certificate image feature data.

[0057] The second Transformer sub-module can be used for performing feature fusion processing on the second certificate image feature data of each scale to obtain fifth certificate image feature data.

[0058] The Class Attention sub-module can be used for performing feature fusion processing on the fourth certificate image feature data and the fifth certificate image feature data to obtain sixth certificate image feature data.

[0059] The multi-scale feature fusion sub-module can be used for feature fusion processing on the first credential image feature data of the largest scale and the sixth credential image feature data, to obtain the third credential image feature data.

[0060] In the embodiments of the present specification, the first Transformer sub-module and the second Transformer sub-module can both be implemented by using a Transformer model, or can be implemented by adaptively modifying the Transformer model according to actual needs, and no specific limitation is made in this regard. By using the first Transformer sub-module to perform feature fusion processing on the first credential image feature data of other scales except the largest scale, fourth credential image feature data with good accuracy can be obtained. In addition, by using the second Transformer sub-module to perform feature fusion processing on the second credential image feature data of each scale, fifth credential image feature data with good accuracy can be obtained.

[0061] Class Attention (CA) is an attention mechanism for visual Transformer in CaiT, which aims to extract information from a group of processed patches. Therefore, in order to improve the feature fusion effect, the feature fusion module can also contain a Class Attention sub-module, so as to use the Class Attention sub-module to perform a feature fusion processing on the fourth credential image feature data and the fifth credential image feature data output by the first Transformer sub-module and the second Transformer sub-module, to obtain the sixth credential image feature data.

[0062] In addition, since the first credential image feature data of the largest scale and the sixth credential image feature data also need to be subjected to a feature fusion processing, in order to obtain the fusion of all the first credential image feature data and all the second credential image feature data, the feature fusion module can also contain a multi-scale feature fusion sub-module, so as to generate the third credential image feature data according to the first credential image feature data and the sixth credential image feature data. In actual application, the multi-scale feature fusion sub-module can be implemented by using the branch network in the Fusion Stage module in the compression artifact tracking network model for feature fusion processing on multi-scale feature maps, which is convenient and fast. Of course, the multi-scale feature fusion sub-module can also be implemented by other network structures (such as a multi-head attention module, a fully connected layer, etc.), and no specific limitation is made in this regard.

[0063] For ease of understanding, FIG. 4 is a structural schematic diagram of another certificate anti-counterfeiting detection model provided by the embodiments of the present specification, as shown in FIG. 4, the multi-modal feature encoder in the certificate anti-counterfeiting detection model can include: a first high-resolution network module 41, a second high-resolution network module 42, and a feature fusion module 43. The feature fusion module 43 can include: a first Transformer sub-module 401, a second Transformer sub-module 402, a Class Attention sub-module 403, and a multi-scale feature fusion sub-module 404. It is worth noting that the output layer of the first high-resolution network module 41 for outputting the first certificate image feature data with the largest scale and the output layer of the Class Attention sub-module 403 can be connected with different network layers at the multi-scale feature fusion sub-module 404, of course, they can also be connected with the same network layer at the multi-scale feature fusion sub-module 404, and no specific limitation is made to this.

[0064] In the embodiments of the present specification, the implementation mode of the classifier in the certificate anti-counterfeiting detection model can also be various. For ease of understanding, this is explained and described.

[0065] Implementation mode one, the classifier can include: a first classification head, a second classification head, and a third classification head.

[0066] The output layer of the feature fusion module can be connected with the first classification head, the output layer of the second high-resolution network module for outputting the second certificate image feature data with the largest scale can be connected with the second classification head, and the first classification head and the second classification head can be connected with the third classification head; wherein the output layer of the feature fusion module can be the output layer of the Fusion Stage module or the output layer of the multi-scale feature fusion sub-module.

[0067] The first classification head can be used for classification processing on the third certificate image feature data to obtain a first classification probability value.

[0068] The second classification head can be used for classification processing on the second certificate image feature data with the largest scale to obtain a second classification probability value.

[0069] The third classification head can be used for classification processing on the first classification probability value and the second classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first certificate image belongs to a trusted certificate image.

[0070] In the embodiments of the present specification, the classification head can refer to a part in the network model that is specially used to map the learned feature representation to a category prediction, which is usually located at the end of the network and immediately follows the feature extraction layer. Therefore, the classifier in the credential anti-counterfeiting detection model can be built by using the classification head.

[0071] Since the third credential image feature data is obtained by fusion processing on all feature data extracted from the first credential image and the second credential image, the image information contained in the third credential image feature data is more comprehensive, so the first classification head can be set to perform classification processing according to the third credential image feature data to obtain the first classification probability value that the first credential image belongs to a trusted / counterfeit credential image.

[0072] Since the second credential image feature data with the largest scale output by the second high-resolution network module can better reflect the image counterfeiting traces extracted from the credential noise feature map or the credential spectrum map, the second classification head can be set to perform classification processing according to the second credential image feature data with the largest scale to obtain the second classification probability value that the first credential image belongs to a trusted / counterfeit credential image.

[0073] In addition, the third classification head can also be set to perform classification processing on the first classification probability value and the second classification probability value to obtain the first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image, which is not only conducive to making full use of the extracted credential image feature data, but also conducive to improving the accuracy of the generated first anti-counterfeiting detection result.

[0074] In the second implementation manner, the credential anti-counterfeiting detection model can further include an edge perception module in the boundary guidance network model, and the classifier can further include a fourth classification head.

[0075] The output layer of the first high-resolution network module for outputting the first credential image feature data with the largest scale can also be connected with the input layer of the edge perception module, the output layer of the edge perception module can be connected with the fourth classification head in the classifier, and the fourth classification head can also be connected with the third classification head.

[0076] The edge perception module can be used to perform feature extraction processing on the first credential image feature data with the largest scale to obtain seventh credential image feature data.

[0077] The fourth classification head can be used to perform classification processing on the seventh credential image feature data to obtain a third classification probability value.

[0078] The third classification head can be specifically used for performing classification processing on the first classification probability value, the second classification probability value and the third classification probability value, to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

[0079] In the embodiments of the present specification, a boundary guided network (BGNet) can improve the performance of camouflage target detection through edge semantics. An edge-aware module (EAM) of the boundary guided network can mine object-related edge semantics from low-level features containing local edge details and high-level features containing global position information under object boundary supervision to guide feature learning and enhance boundary representation.

[0080] Based on this, an edge-aware module (EAM) can be set in the credential anti-counterfeiting detection model. The edge-aware module is used to process the credential image visual feature data (for example, the first credential image feature data with the largest scale) to detect abnormal edge traces existing in the counterfeit credential image caused by image tampering operations.

[0081] In addition, since the seventh credential image feature data output by the edge-aware module can reflect image counterfeiting traces, a fourth classification head can also be set to perform classification processing on the seventh credential image feature data to obtain a third classification probability value of whether the first credential image belongs to a trusted / counterfeit credential image.

[0082] At this time, the third classification head can perform classification processing according to the first classification probability value, the second classification probability value and the third classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image, which is not only conducive to making full use of the extracted credential image feature data, but also conducive to improving the accuracy of the generated first anti-counterfeiting detection result.

[0083] In actual applications, the convolution generalized mean pooling layer (ConvGem) can capture the spatial correlation of global pixels, thereby being able to reduce false positives caused by noise. Based on this, the above-mentioned first classification head to the fourth classification head can be implemented by using the convolution generalized mean pooling layer (ConvGem), thereby being conducive to improving the accuracy of the first anti-counterfeiting detection result of the first credential image generated by the classifier. Of course, the above-mentioned classification head can also be implemented by using a fully connected layer or other network structures, which is not limited in particular.

[0084] For ease of understanding, the specific structure of the classifier 34 in the credential anti-counterfeiting detection model provided in FIG. 3 is also shown, and it can be seen that the first classification head, the second classification head, and the fourth classification head in the classifier 34 can be connected with the output layer of the feature fusion module 33, the output layer of the second high-resolution network module 32 for outputting the second credential image feature data of the largest scale, and the output layer of the edge perception module 35, respectively. Moreover, the output ends of the first classification head, the second classification head, and the fourth classification head can also be connected with the third classification head, and details are not described herein.

[0085] The specific structure of the classifier 44 in the credential anti-counterfeiting detection model provided in FIG. 4 is also shown, and it can be seen that the first classification head, the second classification head, and the fourth classification head in the classifier 44 can be connected with the output layer of the feature fusion module 43, the output layer of the second high-resolution network module 42 for outputting the second credential image feature data of the largest scale, and the output layer of the edge perception module 45, respectively. Moreover, the output ends of the first classification head, the second classification head, and the fourth classification head can also be connected with the third classification head, and details are not described herein.

[0086] In actual application, only the first classification head, the third classification head, and the fourth classification head can be provided, and the second classification head is not provided, which is good in flexibility, and details are not limited herein.

[0087] In a third implementation manner, the credential anti-counterfeiting detection model can further include a confidence decoder and a weighted pooling module in the TruFor model; and the classifier can include a forgery detector in the TruFor model.

[0088] Each output layer at the first high-resolution network module and each output layer at the second high-resolution network module can be connected with the confidence decoder, the confidence decoder and the feature fusion module can be connected with the weighted pooling module, and the weighted pooling module can be connected with the forgery detector.

[0089] The confidence decoder can be configured to perform feature extraction processing and feature fusion processing on the first credential image feature data of each scale and the second credential image feature data of each scale to obtain eighth credential image feature data.

[0090] The weighted pooling module can be configured to determine a target weight required for weighted pooling of the third credential image feature data based on the eighth credential image feature data, and perform weighted pooling processing on the third credential image feature data according to the target weight to obtain ninth credential image feature data.

[0091] The forgery detector can be used for detection processing on the ninth certificate image feature data, to obtain a first anti-forgery detection result reflecting whether the first certificate image belongs to a trusted certificate image.

[0092] In the embodiments of the present specification, TruFor is a model for image forgery detection and positioning. The TruFor model can perform forgery detection according to the visual features of the RGB image and the noise features extracted from the RGB image, to obtain a score representing the overall authenticity of the image from the TruFor model.

[0093] In actual applications, since the spectrum graph and the noise feature graph both belong to two-dimensional data, and the formats of the two can also be consistent, the spectrum graph can also be used to replace the noise feature graph required to be input into the TruFor model, so that the TruFor model can perform feature extraction and fusion on the spectrum graph and the RGB image, to learn the image tampering features in the spectrum graph and the RGB image, thereby obtaining a score representing the overall authenticity of the image with high accuracy. Based on this, the TruFor architecture can also be used to build a certificate anti-forgery detection model.

[0094] Specifically, the first high-resolution network module and the second high-resolution network module in the multi-modal feature encoder in the certificate anti-forgery detection model can be regarded as the encoder (Ecoder) in the TruFor model, and the feature fusion model in the multi-modal feature encoder can be regarded as the anomaly decoder (Anomaly Decoder) in the TruFor model. In addition, since the functions required to be implemented by the classifier in the certificate anti-forgery detection model and the forgery detector (Forgery Detector) in the TruFor model can be consistent, the forgery detector in the TruFor model can be set as the classifier in the certificate anti-forgery detection model. In addition, the certificate anti-forgery detection model also needs to contain the confidence decoder (Confidence Decoder) and the weighted pooling module (Pooling) and other key structures in the TruFor model, so that the certificate anti-forgery detection model built has consistency with the TruFor model architecture, which is beneficial to improve the accuracy of the certificate anti-forgery detection result generated by the certificate anti-forgery detection model built.

[0095] For ease of understanding, the specific structure of the certificate anti-forgery detection model is explained and described in combination with FIG. 5 and FIG. 6.

[0096] As shown in FIG. 5, the multi-modal feature encoder in the certificate anti-fake detection model can include a first high-resolution network module 51, a second high-resolution network module 52, and a feature fusion module 53, and the certificate anti-fake detection model can further include a confidence decoder 54 and a weighted pooling module 55 in the TruFor model. Moreover, the classifier 56 of the certificate anti-fake detection model can include a forgery detector in the TruFor model. The connection relationship between the various model structures is also shown in FIG. 5, which will not be described here.

[0097] As shown in FIG. 6, the multi-modal feature encoder in the certificate anti-fake detection model can include a first high-resolution network module 61, a second high-resolution network module 62, and a feature fusion module 63, and the certificate anti-fake detection model can further include a confidence decoder 64 and a weighted pooling module 65 in the TruFor model. Moreover, the classifier 66 of the certificate anti-fake detection model can include a forgery detector in the TruFor model. The connection relationship between the various model structures is also shown in FIG. 6, which will not be described here.

[0098] In an implementation form, the classifier can further include a fifth classification head, a sixth classification head, and a seventh classification head.

[0099] The output layer of the feature fusion module can be connected to the fifth classification head, the output layer of the confidence decoder can be connected to the sixth classification head, and the output layers of the fifth classification head, the sixth classification head, and the forgery detector can be connected to the seventh classification head. The output layer of the feature fusion module can be the output layer of the Fusion Stage module or the output layer of the multi-scale feature fusion sub-module.

[0100] The forgery detector can be configured to perform detection processing on the ninth certificate image feature data to obtain a fourth classification probability value.

[0101] The fifth classification head can be configured to perform classification processing on the third certificate image feature data to obtain a fifth classification probability value.

[0102] The sixth classification head can be configured to perform classification processing on the eighth certificate image feature data to obtain a sixth classification probability value.

[0103] The seventh classification head can be configured to perform classification processing on the fourth classification probability value, the fifth classification probability value, and the sixth classification probability value to obtain a first anti-fake detection result reflecting whether the first certificate image belongs to a trusted certificate image.

[0104] In the embodiments of the present specification, since the forgery detector outputs a score representing the overall authenticity of the image, the score output by the forgery detector through detection processing on the ninth credential image feature data can be directly used as the fourth classification probability value reflecting whether the first credential image is a forged / credible credential image.

[0105] Since the output of the Fusion Stage module or the branch network of the Fusion Stage module for feature fusion processing on multi-scale feature maps is a segmentation result for a forged region in the image, when the output layer of the feature fusion module is the output layer of the Fusion Stage module, or the output layer of the branch network of the Fusion Stage module for feature fusion processing on multi-scale feature maps (i.e., the output layer of the multi-scale feature fusion submodule), the third credential image feature data output by the feature fusion module can reflect the forged region in the image. Therefore, in order to make full use of the data and further improve the result of the credential forgery detection, a fifth classification head can be set to perform classification processing on the third credential image feature data, so as to obtain a fifth classification probability value reflecting whether the first credential image is a forged / credible credential image.

[0106] Since the confidence map output by the confidence decoder can reflect the forged region that may exist in the credential image, and the real forged region and the random anomaly can be distinguished by analyzing the confidence map, in order to make full use of the data and further improve the result of the credential forgery detection, a sixth classification head can be set to perform classification processing on the eighth credential image feature data output by the confidence decoder, so as to obtain a sixth classification probability value reflecting whether the first credential image is a forged / credible credential image.

[0107] Subsequently, the seventh classification head can be used to perform classification processing on the fourth classification probability value output by the forgery detector, the fifth classification probability value and the sixth classification probability value, which is beneficial to improve the accuracy of the first forgery detection result obtained for reflecting whether the first credential image is a credible credential image.

[0108] In a fifth implementation manner, the credential forgery detection model can further include an edge perception module in the boundary guidance network model; and the classifier can further include an eighth classification head.

[0109] The output layer of the first high-resolution network module for outputting the first credential image feature data with the largest scale can be further connected to an input layer of the edge perception module, an output layer of the edge perception module can be connected to the eighth classification head in the classifier, and the eighth classification head can be further connected to the seventh classification head.

[0110] The edge-aware module can be configured to perform feature extraction processing on the first credential image feature data with the largest scale to obtain tenth credential image feature data.

[0111] The eighth classification head can be configured to perform classification processing on the tenth credential image feature data to obtain a seventh classification probability value.

[0112] The seventh classification head can be specifically configured to perform classification processing on the fourth classification probability value, the fifth classification probability value, the sixth classification probability value and the seventh classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

[0113] For ease of understanding, the specific structure and connection mode of the classifier 56 in the credential anti-counterfeiting detection model are also shown in FIG. 5. As can be seen, the counterfeit detector, the fifth classification head, the sixth classification head and the eighth classification head in the classifier 56 can be respectively connected with the output layer of the weighted pooling module 55, the output layer of the feature fusion module 53, the output layer of the confidence decoder 54 and the output layer of the edge-aware module 57. The seventh classification head can be connected with each output end of the counterfeit detector, the fifth classification head, the sixth classification head and the eighth classification head, and no further description is made.

[0114] The specific structure and connection mode of the classifier 66 in the credential anti-counterfeiting detection model are also shown in FIG. 6. As can be seen, the counterfeit detector, the fifth classification head, the sixth classification head and the eighth classification head in the classifier 66 can be respectively connected with the output layer of the weighted pooling module 65, the output layer of the feature fusion module 63, the output layer of the confidence decoder 64 and the output layer of the edge-aware module 67. The seventh classification head can be connected with each output end of the counterfeit detector, the fifth classification head, the sixth classification head and the eighth classification head, and no further description is made.

[0115] In the embodiments of the present specification, the processing on the first credential image to obtain a second credential image can include performing frequency domain transformation processing on the first credential image to obtain the credential frequency spectrum; and the frequency domain transformation processing includes discrete cosine transformation processing.

[0116] In the embodiments of the present specification, since the user usually uses an image compression algorithm based on Discrete Cosine Transform (DCT) in the process of forging the credential image, and the Discrete Cosine Transform coefficients of the image can often reflect the difference in the compression rounds performed for different regions of the image, the first credential image can be subjected to Discrete Cosine Transform, and the obtained Discrete Cosine Transform coefficients can be taken as the credential spectrum graph. By combining the Discrete Cosine Transform coefficients of the first credential image for anti-counterfeiting detection processing, the accuracy of the first anti-counterfeiting detection result of the generated first credential image can be improved.

[0117] Of course, a frequency domain transformation processing mode such as Fourier transform, Fast Fourier Transform (FFT), wavelet transform, frequency domain filtering, etc. can also be used to generate the credential spectrum graph corresponding to the first credential image, which is not specifically limited.

[0118] In the embodiments of the present specification, the credential anti-counterfeiting detection model can further include a noise feature extraction model.

[0119] The noise feature extraction model can be any one of a Noiseprint extractor, a Spatial Rich Model, and a Bayar Convolution model; and the noise feature extraction model can be connected with the second high-resolution network module. The noise feature extraction model is configured to process the first credential image to obtain the credential noise feature map.

[0120] In the embodiments of the present specification, the Noiseprint extractor can extract a noise-sensitive fingerprint from an RGB image. The Spatial Rich Model (SRM) can be used to suppress image content and highlight image tampering traces and noise. The Bayar Convolution model is usually used as a noise extractor to adaptively learn image tampering traces from image data. Therefore, any one of the above three models can be used to extract the credential noise feature map from the first credential image, which is convenient and fast.

[0121] Optionally, the noise feature extraction model in the credential anti-counterfeiting detection model can include at least two models of a Noiseprint extractor, a Spatial Rich Model, and a Bayar Convolution model, and a multi-head attention module; the at least two models are connected with the multi-head attention module, and the multi-head attention module is further connected with the second high-resolution network module.

[0122] The at least two models can be used to process the first credential image to obtain each initial credential noise feature data.

[0123] The multi-head attention module can be used for feature fusion processing on the respective initial credential noise feature data, to obtain the credential noise feature map.

[0124] In the embodiments of the present specification, a plurality of models can also be used to extract an initial credential noise feature data of the first credential image, and a multi-head attention (Multi-Head Attention) module is used for feature fusion processing on each initial credential noise feature data, so as to obtain a credential noise feature map with better accuracy as the second credential image, which is beneficial to improve the accuracy of the anti-fake detection result generated based on the second credential image.

[0125] Currently, since the key information region in the credential image contains the most important credential information, users often tamper with the key information region in the credential image. However, the existing credential anti-fake detection scheme does not perform key analysis and processing on the key information region of the credential image, so it is easy to generate false credential anti-fake detection results due to the interference of other regions in the credential image, which affects the accuracy and effectiveness of the counterfeit credential detection scheme.

[0126] In actual application, since the key information region in the trusted credential may need to be located at a preset position, and the text in the key information region may need to use a preset font, and the text in the key information region may also need to conform to a preset text alignment format, so that during credential anti-fake detection, a format detection strategy required to be executed on the key information region in the credential image can be summarized in advance in combination with expert experience, so as to generate a final credential anti-fake detection result according to the format detection result.

[0127] Based on this, the method shown in FIG. 2 can also include determining the position information of the key information region in the first credential image.

[0128] Based on the position information of the key information region, it is detected whether the key information region is located at a preset position in the first credential image, and / or whether the text data in the key information region conforms to a preset text alignment format, and / or whether the text data in the key information region uses a preset font, to obtain a second anti-fake detection result for the first credential image.

[0129] According to the first anti-fake detection result and the second anti-fake detection result, a comprehensive anti-fake detection result for the first credential image is generated.

[0130] In the embodiments of the present specification, the first credential image can generally include a region where key information is located and a region where non-key information is located. The region where key information is located can include key information in the first credential image. In actual applications, the categories of key information included in different types of credential images can differ, and the regions where the same type of key information is located in different types of credential images can also differ, which are not specifically limited.

[0131] For ease of understanding, an example is given. For example, assuming that the first credential image is an identity credential image, the region where key information is located in the first credential image can be a region including a user face image and a region including text identity information of the user; or, assuming that the first credential image is a payment result credential image, the region where key information is located in the first credential image can be a region including text or icons reflecting whether the payment is successful and a region including payment amount, payee, and other information, which are not specifically limited. The region in the first credential image other than the region where key information is located can belong to the region where non-key information is located, which is not described herein.

[0132] In actual applications, the first credential image can be input into a target detection model to obtain position information of the region where key information is located in the first credential image output by the target detection model. The target detection model can be a deep learning model for object detection, which can be used to predict the position and category of an object in an image. Based on this, the region where key information is located in the first credential image can be taken as a detection object, and a trusted credential image consistent with the credential category to which the first credential image belongs can be used as a training sample in advance, and the position information of the region where key information is located in the trusted credential image can be taken as label data, so that the target detection model can be trained to accurately identify the position information of the region where key information is located in the first credential image, which is convenient and efficient.

[0133] Alternatively, the first credential image can be subjected to correction processing and scale transformation processing to obtain a target size of a corrected credential image; the target size is the size of a credential template of the same type of credential to which the target credential belongs. Thus, the preset position information of the preset region where key information is located in the credential template set in advance can be determined as the position information of the region where key information is located in the corrected credential image. Subsequently, various types of format detection related to the region where key information is located need to be performed on the corrected credential image, thereby ensuring the accuracy of the second anti-counterfeiting detection result generated for the first credential image.

[0134] In the embodiments of the present specification, after detecting the region where the key information in the first credential image according to the preset format detection strategy, a second anti-counterfeiting detection result reflecting whether the first credential image belongs to a real / fake credential image can be obtained. When generating a comprehensive anti-counterfeiting detection result in combination with the first anti-counterfeiting detection result and the second anti-counterfeiting detection result of the first credential image, the comprehensive anti-counterfeiting detection result can be caused to represent that the first credential image should belong to a fake credential image when any one of the anti-counterfeiting detection results reflects that the first credential image should belong to a fake credential image. Alternatively, the first anti-counterfeiting detection result and the second anti-counterfeiting detection result can be further determined to correspond to a fake risk score, and a comprehensive fake risk score can be generated according to each fake risk score, so that the comprehensive anti-counterfeiting detection result is caused to represent that the first credential image should belong to a fake credential image when the comprehensive fake risk score reaches a threshold, which is not limited herein.

[0135] Optionally, the detecting, based on the position information of the region where the key information is located, whether the region where the key information is located is located at a preset position in the first credential image can include: determining, based on the position information of the region where the key information is located, whether a first distance between a left side of the region where the key information is located and a longitudinal center line of the first credential image is equal to a second distance between a right side of the region where the key information is located and the longitudinal center line; or determining, based on the position information of the region where the key information is located, whether a third distance between the left side of the region where the key information is located and a left side of the first credential image is equal to a fourth distance between a right side of the region where the key information is located and a right side of the first credential image.

[0136] In the embodiments of the present specification, when the distances between the left and right side lines of the region where the key information is located and the corresponding left and right side lines of the first credential image are equal, or when the distances between the left and right side lines of the region where the key information is located and the longitudinal center line of the first credential image are equal, it can be considered that the region where the key information is located is located at a position (i.e., a preset position) symmetrical to the longitudinal center line of the first credential image, and at this time, it can be considered that the possibility of tampering by the user for the region where the key information is located is small, so that a second anti-counterfeiting detection result reflecting that the first credential image does not belong to a fake credential image can be generated, or the second anti-counterfeiting detection result can be generated in combination with the detection results of other detection items, which is not limited herein.

[0137] Optionally, the detecting, based on the position information of the region where the key information is located, whether the text data in the region where the key information is located conforms to a preset text alignment format can include: determining, based on the position information of the region where the key information is located, position information of each text data in the region where the key information is located.

[0138] According to the position information of the text data, it is determined whether the vertical coordinates of the bottom edges of the text data in the same row are consistent, and / or, whether the vertical coordinates of the top edges of the text data in the same row are consistent, and / or, whether the horizontal coordinates of the left edges of the text data in the same column are consistent, and / or, whether the horizontal coordinates of the right edges of the text data in the same column are consistent.

[0139] In the embodiments of the present specification, the preset text alignment format required to be met by the text data in the region where the key information is located can be in-line alignment, column alignment, etc., which is not specifically limited.

[0140] When detecting the preset text alignment format, the position information of each text data in the region where the key information is located can be determined, wherein the position information of each text data can be used to determine the range of the horizontal and vertical coordinates of the text data. Based on this, it can be determined whether the maximum or minimum value of the vertical coordinates of the text data in the same row is the same, and if so, it can be considered that the text data in the region where the key information is located is specifically in-line alignment format (the preset text alignment format). Similarly, it can be determined whether the maximum or minimum value of the horizontal coordinates of the text data in the same column is the same, and if so, it can be considered that the text data in the region where the key information is located is specifically column alignment format (the preset text alignment format). At this time, it can be considered that the possibility of user tampering with the region where the key information is located is small, and therefore, a second anti-forgery detection result reflecting that the first credential image does not belong to a forged credential image can be generated, or the second anti-forgery detection result can also be generated in combination with the detection results of other detection items, which is not specifically limited.

[0141] Optionally, the detection of whether the text data in the region where the key information is located adopts a preset font based on the position information of the region where the key information is located can include: inputting a credential sub-image in the region where the key information is located in the first credential image into a font detection model to obtain a detection result output by the font detection model and used to reflect a target font adopted by the text data in the region where the key information is located; wherein the font detection model is obtained by pre-training a classification model using image samples carrying label data, and the label data is used to reflect the font category to which the text data in the image sample belongs.

[0142] Based on the detection result, it is judged whether the target font is consistent with the preset font. Alternatively, based on the position information of the region where the key information is located, the credential sub-image in the region where the key information is located in the first credential image is input into a font feature extraction model to obtain target font feature data of the text data in the region where the key information is located extracted by the font feature extraction model; wherein the font feature extraction model comprises: a self-encoder model trained in advance by using an image sample containing text data of a preset font.

[0143] It is judged whether the target font feature data is consistent with the reference font feature data of the preset font; wherein the reference font feature data is font feature data extracted by the font feature extraction model for a trusted credential image consistent with the credential category to which the first credential image belongs.

[0144] In the embodiments of the present specification, the preset font required by the text data in the region where the key information is located can be set according to actual conditions, for example, the preset font can be Kai Ti, Song Ti, and Hei Ti, etc., which is not limited. By using the pre-trained font detection model or font feature extraction model, it is convenient and fast to identify whether the text data in the region where the key information is located uses the preset font, and the accuracy is good.

[0145] In actual application, after it is determined that the text data in the region where the key information is located in the first credential image uses the preset font, a second anti-forgery detection result reflecting that the first credential image does not belong to a forged credential image can be directly generated, or the second anti-forgery detection result can also be generated in combination with the detection results of other detection items, which is not limited.

[0146] Since the metadata information of the image can include shooting time, location, camera information, image modification information, etc., the image can be subjected to anti-forgery detection based on the metadata of the image to verify the authenticity of the image.

[0147] Based on this, the method described in FIG. 2 can also include obtaining an exchangeable image file of the first credential image. It is judged whether the exchangeable image file contains a forged image keyword to obtain a target judgment result.

[0148] Correspondingly, the generation of the comprehensive anti-forgery detection result for the first credential image according to the first anti-forgery detection result and the second anti-forgery detection result can include: generating a comprehensive anti-forgery detection result for the first credential image according to the first anti-forgery detection result, the second anti-forgery detection result and the target judgment result.

[0149] In the embodiments of the present specification, the exchangeable image file (EXIF information) can record the attribute information and shooting data of the digital photo; the EXIF can be attached to the JPEG, TIFF, RIFF and the like, and add the content related to the shooting information of the digital camera and the version information of the index map or image processing software. Based on this, it can be identified in advance that after the user tampers with the credential image, the text reflecting the user's tampering with the credential image behavior carried in the exchangeable image file of the counterfeit credential image will be used as a counterfeit image keyword. If the target judgment result indicates that the exchangeable image file of the first credential image contains the counterfeit image keyword, it can be considered that the first credential image is more likely to belong to the counterfeit credential image. Therefore, by combining the target judgment result and the first anti-counterfeiting detection result and the second anti-counterfeiting detection result, a comprehensive anti-counterfeiting detection result for the first credential image is generated, which is beneficial to improve the accuracy of the comprehensive anti-counterfeiting detection result.

[0150] Based on the same idea, the present specification also provides a device corresponding to the above method. Please refer to FIG. 7, which is a structural schematic diagram of a credential anti-counterfeiting detection device provided by the present specification. As shown in FIG. 7, the credential anti-counterfeiting detection device 7 can be realized by software, hardware or a combination of the two to become all or part of an electronic device. According to some embodiments, the credential anti-counterfeiting detection device 7 can include a first acquisition module 71, a second acquisition module 72, and an anti-counterfeiting detection module 73.

[0151] The first acquisition module 71 can be used to acquire a first credential image obtained by image collection of a target credential of a user.

[0152] The second acquisition module 72 can be used to process the first credential image to obtain a second credential image, the second credential image including any one of a credential noise feature map and a credential spectrum map.

[0153] The anti-counterfeiting detection module 73 can be used to perform anti-counterfeiting detection processing on the first credential image and the second credential image by using a credential anti-counterfeiting detection model to obtain a first anti-counterfeiting detection result for the first credential image output by the credential anti-counterfeiting detection model; wherein the credential anti-counterfeiting detection model includes a multi-modal feature encoder and a classifier connected in sequence.

[0154] Optionally, the multi-modal feature encoder can include a first high-resolution network module, a second high-resolution network module, and a feature fusion module; the first high-resolution network module and the second high-resolution network module are connected with the feature fusion module respectively, and the feature fusion module is connected with the classifier.

[0155] The first high-resolution network module can be configured to perform feature extraction on the first credential image to obtain a plurality of first credential image feature data of different scales. The second high-resolution network module can be configured to perform feature extraction on the second credential image to obtain a plurality of second credential image feature data of different scales. The feature fusion module can be configured to perform feature fusion on the plurality of first credential image feature data of different scales and the plurality of second credential image feature data of different scales to obtain third credential image feature data. The classifier can be configured to perform classification based on the third credential image feature data to obtain a first anti-counterfeiting detection result reflecting whether the first credential image is a trusted credential image.

[0156] Optionally, the first high-resolution network module can be an RGB Stream module in a compact trace tracking network model; and the second high-resolution network module can be a DCT Stream module in the compact trace tracking network model.

[0157] The RGB Stream module can be configured to perform feature extraction on the first credential image to obtain a plurality of first credential image feature data of different scales. The DCT Stream module can be configured to perform feature extraction on the second credential image to obtain a plurality of second credential image feature data of different scales.

[0158] Optionally, the feature fusion module can be a Fusion Stage module in the compact trace tracking network model. The Fusion Stage module can be configured to perform feature fusion on the plurality of first credential image feature data of different scales and the plurality of second credential image feature data of different scales to obtain third credential image feature data.

[0159] Optionally, the feature fusion module can include a first Transformer sub-module, a second Transformer sub-module, a Class Attention sub-module, and a multi-scale feature fusion sub-module. The multi-scale feature fusion sub-module can be a branch network in the Fusion Stage module of the compression artifact tracking network model for performing feature fusion processing on multi-scale feature maps. Each output layer of the first high-resolution network module except for an output layer for outputting the first credential image feature data of the largest scale can be connected with the first Transformer sub-module. Each output layer of the second high-resolution network module can be connected with the second Transformer sub-module. The first Transformer sub-module and the second Transformer sub-module can be connected with the Class Attention sub-module. The output layer for outputting the first credential image feature data of the largest scale and the Class Attention sub-module can be connected with the multi-scale feature fusion sub-module. The multi-scale feature fusion sub-module is further connected with the classifier.

[0160] The first Transformer sub-module can be used for performing feature fusion processing on the first credential image feature data of other scales except for the largest scale to obtain fourth credential image feature data. The second Transformer sub-module can be used for performing feature fusion processing on the second credential image feature data of each scale to obtain fifth credential image feature data. The Class Attention sub-module can be used for performing feature fusion processing on the fourth credential image feature data and the fifth credential image feature data to obtain sixth credential image feature data. The multi-scale feature fusion sub-module can be used for performing feature fusion processing on the first credential image feature data of the largest scale and the sixth credential image feature data to obtain the third credential image feature data.

[0161] Optionally, the classifier can include a first classification head, a second classification head, and a third classification head.

[0162] The output layer of the feature fusion module can be connected with the first classification head. The output layer of the second high-resolution network module for outputting the second credential image feature data of the largest scale can be connected with the second classification head. The first classification head and the second classification head can be connected with the third classification head. The output layer of the feature fusion module can be an output layer of the Fusion Stage module or an output layer of the multi-scale feature fusion sub-module.

[0163] The first classification head can be used for classification processing on the third credential image feature data to obtain a first classification probability value. The second classification head can be used for classification processing on the second credential image feature data with the largest scale to obtain a second classification probability value. The third classification head can be used for classification processing on the first classification probability value and the second classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

[0164] Optionally, the credential anti-counterfeiting detection model can further include an edge perception module in the boundary guidance network model; the classifier can further include a fourth classification head. The output layer of the first high-resolution network module for outputting the first credential image feature data with the largest scale can be further connected with an input layer of the edge perception module, an output layer of the edge perception module can be connected with the fourth classification head in the classifier, and the fourth classification head can be further connected with the third classification head.

[0165] The edge perception module can be used for feature extraction processing on the first credential image feature data with the largest scale to obtain seventh credential image feature data. The fourth classification head can be used for classification processing on the seventh credential image feature data to obtain a third classification probability value. The third classification head can be specifically used for classification processing on the first classification probability value, the second classification probability value and the third classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

[0166] Optionally, the credential anti-counterfeiting detection model can further include a confidence decoder and a weighted pooling module in a TruFor model; and the classifier can include a forgery detector in the TruFor model.

[0167] Each output layer of the first high-resolution network module can be further connected with each output layer of the second high-resolution network module, and the confidence decoder and the feature fusion module can be connected with the weighted pooling module, and the weighted pooling module can be connected with the forgery detector.

[0168] The confidence decoder can be configured to perform feature extraction and feature fusion on the first credential image feature data of each scale and the second credential image feature data of each scale to obtain eighth credential image feature data. The weighted pooling module can be configured to determine a target weight required for weighted pooling of the third credential image feature data based on the eighth credential image feature data, and perform weighted pooling processing on the third credential image feature data according to the target weight to obtain ninth credential image feature data. The forgery detector can be configured to perform detection processing on the ninth credential image feature data to obtain a first anti-forgery detection result reflecting whether the first credential image is a trusted credential image.

[0169] Optionally, the classifier can further include a fifth classification head, a sixth classification head, and a seventh classification head.

[0170] The output layer of the feature fusion module can be connected to the fifth classification head, the output layer of the confidence decoder can be connected to the sixth classification head, and the output layers of the fifth classification head, the sixth classification head, and the forgery detector can be connected to the seventh classification head. The output layer of the feature fusion module can be the output layer of the Fusion Stage module or the output layer of the multi-scale feature fusion submodule.

[0171] The forgery detector can be configured to perform detection processing on the ninth credential image feature data to obtain a fourth classification probability value. The fifth classification head can be configured to perform classification processing on the third credential image feature data to obtain a fifth classification probability value. The sixth classification head can be configured to perform classification processing on the eighth credential image feature data to obtain a sixth classification probability value. The seventh classification head can be configured to perform classification processing on the fourth classification probability value, the fifth classification probability value, and the sixth classification probability value to obtain a first anti-forgery detection result reflecting whether the first credential image is a trusted credential image.

[0172] Optionally, the credential anti-forgery detection model can further include an edge perception module in the boundary guidance network model. The classifier can further include an eighth classification head. The output layer of the first high-resolution network module for outputting the first credential image feature data of the largest scale can be connected to the input layer of the edge perception module. The output layer of the edge perception module can be connected to the eighth classification head in the classifier, and the eighth classification head can be connected to the seventh classification head.

[0173] The edge perception module can be configured to perform feature extraction on the first credential image feature data with the largest scale to obtain tenth credential image feature data. The eighth classification head can be configured to perform classification on the tenth credential image feature data to obtain a seventh classification probability value. The seventh classification head can be configured to perform classification on the fourth classification probability value, the fifth classification probability value, the sixth classification probability value, and the seventh classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image is a trusted credential image.

[0174] Optionally, the second acquisition module can be configured to perform frequency domain transformation on the first credential image to obtain the credential frequency spectrum, wherein the frequency domain transformation can include discrete cosine transformation.

[0175] Optionally, the credential anti-counterfeiting detection model can further include a noise feature extraction model.

[0176] The noise feature extraction model can be any one of a Noiseprint extractor, a spatially rich model, and a Bayar convolution model. The noise feature extraction model is connected to the second high-resolution network module. The noise feature extraction model can be configured to process the first credential image to obtain the credential noise feature map.

[0177] The noise feature extraction model can include at least two models selected from the group consisting of a Noiseprint extractor, a spatially rich model, and a Bayar convolution model, and a multi-head attention module. The at least two models are connected to the multi-head attention module, and the multi-head attention module is connected to the second high-resolution network module.

[0178] The at least two models can be configured to process the first credential image to obtain respective initial credential noise feature data. The multi-head attention module can be configured to perform feature fusion on the respective initial credential noise feature data to obtain the credential noise feature map.

[0179] Optionally, the apparatus in FIG. 7 can further include:

[0180] A position determination module configured to determine position information of a region in the first credential image where key information is located.

[0181] The format detection module is configured to detect, based on the position information of the region where the key information is located, whether the region where the key information is located is located at a preset position in the first certificate image, and / or whether the text data in the region where the key information is located conforms to a preset text alignment format, and / or whether the text data in the region where the key information is located adopts a preset font, to obtain a second anti-counterfeiting detection result for the first certificate image.

[0182] The comprehensive anti-counterfeiting detection result generation module is configured to generate a comprehensive anti-counterfeiting detection result for the first certificate image according to the first anti-counterfeiting detection result and the second anti-counterfeiting detection result.

[0183] Optionally, the format detection module can be specifically configured to determine, based on the position information of the region where the key information is located, whether a first distance between a left side of the region where the key information is located and a vertical center line of the first certificate image is equal to a second distance between a right side of the region where the key information is located and the vertical center line. Alternatively, the format detection module can be specifically configured to determine, based on the position information of the region where the key information is located, whether a third distance between the left side of the region where the key information is located and a left side of the first certificate image is equal to a fourth distance between the right side of the region where the key information is located and a right side of the first certificate image.

[0184] Optionally, the format detection module can be specifically configured to determine, based on the position information of the region where the key information is located, position information of each text data in the region where the key information is located. According to the position information of the text data, it is determined whether the vertical coordinates of the bottom sides of the text data located in the same row are consistent, and / or whether the vertical coordinates of the top sides of the text data located in the same row are consistent, and / or whether the horizontal coordinates of the left sides of the text data located in the same column are consistent, and / or whether the horizontal coordinates of the right sides of the text data located in the same column are consistent.

[0185] Optionally, the format detection module can be specifically configured to input, based on the position information of the region where the key information is located, a certificate sub-image in the region where the key information is located in the first certificate image into a font detection model to obtain a detection result output by the font detection model and used for reflecting a target font adopted by the text data in the region where the key information is located. The font detection model is obtained by pre-training a classification model by using image samples carrying label data, and the label data is used for reflecting a font category to which the text data in the image sample belongs. Based on the detection result, it is determined whether the target font is consistent with the preset font.

[0186] Optionally, the format detection module can be specifically configured to: based on position information of the region where the key information is located, input a voucher sub-image in the region where the key information is located in the first voucher image into a font feature extraction model to obtain target font feature data of text data in the region where the key information is located extracted by the font feature extraction model; the font feature extraction model comprises a self-encoder model pre-trained by using an image sample containing text data of a preset font; and determine whether the target font feature data is consistent with reference font feature data of the preset font; the reference font feature data is font feature data extracted by the font feature extraction model for a trusted voucher image of a voucher category consistent with the first voucher image.

[0187] Optionally, the apparatus shown in FIG. 7 can further include a third acquisition module configured to acquire an exchangeable image file of the first voucher image; a judgment module configured to determine whether the exchangeable image file contains a counterfeit image keyword to obtain a target judgment result; and the comprehensive anti-counterfeiting detection result generation module can be specifically configured to generate a comprehensive anti-counterfeiting detection result for the first voucher image according to the first anti-counterfeiting detection result, the second anti-counterfeiting detection result and the target judgment result.

[0188] The embodiments of the present specification also provide a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the voucher anti-counterfeiting detection method shown in FIG. 2, and the specific implementation process can be referred to the specific description of the related embodiments of the voucher anti-counterfeiting detection method, which will not be repeated here.

[0189] The present specification also provides a computer program product storing at least one instruction, the at least one instruction being loaded and executed by the processor to implement the voucher anti-counterfeiting detection method shown in FIG. 2, and the specific implementation process can be referred to the specific description of the related embodiments of the voucher anti-counterfeiting detection method, which will not be repeated here.

[0190] The embodiments of the present specification also provide a structure diagram of an electronic device shown in FIG. 8. As shown in FIG. 8, at the hardware level, the electronic device can include a processor, an internal bus, a network interface, a memory and a non-volatile memory, and of course can also include other hardware required by the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to implement the voucher anti-counterfeiting detection method shown in FIG. 2, and the specific implementation process can be referred to the specific description of the related embodiments of the voucher anti-counterfeiting detection method, which will not be repeated here.

[0191] Of course, besides the software implementation, the present specification does not exclude other implementation manners, such as a logic device or a combination of software and hardware, that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.

[0192] Each embodiment in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments. Especially, for the computer readable storage medium, the computer program product and the electronic device shown in FIG. 8, since they are basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment.

[0193] In the 1990s, it was quite obvious to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structures of diodes, transistors, switches, etc.) or in software (improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flow into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented by hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it, rather than by asking a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented by "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in the above-mentioned hardware description languages and programmed into an integrated circuit, a hardware circuit implementing the logical method flow can be easily obtained.

[0194] The controller can be implemented in any suitable way, e.g. the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, e.g. software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of controllers include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to being implemented in pure computer readable program code, the controller can equally well be implemented to perform the same functions using logic gates, switches, an application specific integrated circuit, a programmable logic controller and an embedded microcontroller, etc. by means of a logic programming of the method steps. The controller can thus be considered a hardware component, and the means comprised therein for performing the various functions can be considered structures within the hardware component. Alternatively, or even additionally, the means for performing the various functions can be considered both software modules implementing the method and structures within the hardware component.

[0195] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0196] For the sake of description, the above apparatuses are described in various units by functions respectively. Of course, the functions of the units can be implemented in the same or multiple software and / or hardware in implementing the present specification.

[0197] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0198] The specification is presented with reference to flow diagrams and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing element or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.

[0199] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flow diagrams and / or block diagrams block or blocks.

[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.

[0201] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0202] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, for storing instructions and data used by one or more of the components of the computing device. The memory can further include non-volatile memory, such as read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or non-volatile random access memory (NVRAM) for storing instructions and data used by one or more of the components of the computing device. For example, the memory can include a hard disk or other magnetic, optical, electromagnetic, solid state, or other device used for storage. The memory can be volatile and / or non-volatile memory. The memory can include both volatile and non-volatile memory. The memory can be internal to the computing device and / or external to the computing device.

[0203] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0204] It should also be noted that the terms "comprising", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a list of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0205] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0206] The present specification can be described in the general context of computer-executable instructions, such as program modules, executed by computers. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media, including storage devices.

[0207] The various embodiments described in this specification are described using a numbering of embodiments approach: these are each individually integrated contributions pertaining to different but related aspects of the description. Each of the various embodiments can stand on its own, and each can be combined with the subject matter of other embodiments to produce further embodiments. Where appropriate, therefore, the contents of the specification can be regarded as being incorporated by reference, including the description, drawings, claims, abstract and the like.

[0208] The above description is embodied in the form of embodiments only and is not intended to limit the present specification. The present specification can be variously changed and modified by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification should be included in the scope of the claims of the present specification.

Claims

1. A method for certificate anti-counterfeiting detection, the method comprising: obtaining a first certificate image obtained by image acquisition on a target certificate of a user; processing the first certificate image to obtain a second certificate image, the second certificate image comprising any one of a certificate noise feature map and a certificate spectrum map; performing anti-counterfeiting detection processing on the first certificate image and the second certificate image by using a certificate anti-counterfeiting detection model to obtain a first anti-counterfeiting detection result output by the certificate anti-counterfeiting detection model for the first certificate image; wherein the certificate anti-counterfeiting detection model comprises a multi-modal feature encoder and a classifier connected in sequence.

2. The method of claim 1, the multi-modal feature encoder comprising: a first high-resolution network module, a second high-resolution network module, and a feature fusion module; the first high-resolution network module and the second high-resolution network module are connected with the feature fusion module, and the feature fusion module is connected with the classifier; wherein the first high-resolution network module is configured to perform feature extraction processing on the first certificate image to obtain a plurality of first certificate image feature data of different scales; the second high-resolution network module is configured to perform feature extraction processing on the second certificate image to obtain a plurality of second certificate image feature data of different scales; the feature fusion module is configured to perform feature fusion processing on the plurality of first certificate image feature data of different scales and the plurality of second certificate image feature data of different scales to obtain third certificate image feature data; the classifier is configured to perform classification processing based on the third certificate image feature data to obtain the first anti-counterfeiting detection result reflecting whether the first certificate image belongs to a trusted certificate image.

3. The method of claim 2, wherein the first high-resolution network module is an RGB Stream module in a compact trace tracking network model; and the second high-resolution network module is a DCT Stream module in the compact trace tracking network model. the RGB Stream module is configured to perform feature extraction processing on the first certificate image to obtain a plurality of first certificate image feature data of different scales; the DCT Stream module is configured to perform feature extraction processing on the second certificate image to obtain a plurality of second certificate image feature data of different scales.

4. The method of claim 2, wherein the feature fusion module is a Fusion Stage module in a compact trace tracking network model; the Fusion Stage module is configured to perform feature fusion processing on the plurality of first certificate image feature data of different scales and the plurality of second certificate image feature data of different scales to obtain the third certificate image feature data; or The feature fusion module comprises: The first Transformer submodule, the second Transformer submodule, the Class Attention submodule, and the multi-scale feature fusion submodule are branch networks in a Fusion Stage module of the compression artifact tracking network model and are used for feature fusion processing on multi-scale feature maps; Each output layer of the first high-resolution network module except an output layer used for outputting the first credential image feature data of the largest scale is connected with the first Transformer submodule, each output layer of the second high-resolution network module is connected with the second Transformer submodule, the first Transformer submodule and the second Transformer submodule are connected with the Class Attention submodule, the output layer used for outputting the first credential image feature data of the largest scale and the Class Attention submodule are connected with the multi-scale feature fusion submodule, and the multi-scale feature fusion submodule is further connected with the classifier; The first Transformer submodule is used for feature fusion processing on the first credential image feature data of other scales except the largest scale to obtain fourth credential image feature data. The second Transformer submodule is used for feature fusion processing on the second credential image feature data of each scale to obtain fifth credential image feature data. The Class Attention submodule is used for feature fusion processing on the fourth credential image feature data and the fifth credential image feature data to obtain sixth credential image feature data. The multi-scale feature fusion submodule is used for feature fusion processing on the first credential image feature data of the largest scale and the sixth credential image feature data to obtain the third credential image feature data.

5. The method of claim 4, the classifier comprising: The first classification head, the second classification head, and the third classification head; An output layer of the feature fusion module is connected with the first classification head, an output layer of the second high-resolution network module used for outputting the second credential image feature data of the largest scale is connected with the second classification head, and the first classification head and the second classification head are connected with the third classification head; wherein the output layer of the feature fusion module is an output layer of the Fusion Stage module or an output layer of the multi-scale feature fusion submodule; The first classification head is used for classification processing on the third credential image feature data to obtain a first classification probability value. The second classification head is used for classification processing on the second credential image feature data of the largest scale to obtain a second classification probability value. The third classification head is used for classification processing on the first classification probability value and the second classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image is a trusted credential image.

6. The method of claim 5, the credential anti-spoofing detection model further comprising: An edge perception module in a boundary guided network model; The classifier further includes a fourth classification head; An output layer of the first high-resolution network module for outputting the first credential image feature data of the largest scale is further connected with an input layer of the edge perception module, an output layer of the edge perception module is connected with the fourth classification head in the classifier, and the fourth classification head is further connected with the third classification head; The edge perception module is configured to perform feature extraction processing on the first credential image feature data of the largest scale to obtain seventh credential image feature data. The fourth classification head is configured to perform classification processing on the seventh credential image feature data to obtain a third classification probability value. The third classification head is specifically configured to perform classification processing on the first classification probability value, the second classification probability value, and the third classification probability value to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

7. The method of claim 4, the credential anti-spoofing detection model further comprising: A confidence decoder and a weighted pooling module in a TruFor model; The classifier includes a forgery detector in the TruFor model; Each output layer of the first high-resolution network module and each output layer of the second high-resolution network module are further connected with the confidence decoder, the confidence decoder and the feature fusion module are connected with the weighted pooling module, and the weighted pooling module is connected with the forgery detector; The confidence decoder is configured to perform feature extraction processing and feature fusion processing on the first credential image feature data of each scale and the second credential image feature data of each scale to obtain eighth credential image feature data. The weighted pooling module is configured to determine a target weight required for weighted pooling of the third credential image feature data based on the eighth credential image feature data, perform weighted pooling processing on the third credential image feature data according to the target weight to obtain ninth credential image feature data, and output the ninth credential image feature data. The forgery detector is configured to perform detection processing on the ninth credential image feature data to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

8. The method of claim 7, the classifier further comprising: A fifth classification head, a sixth classification head, and a seventh classification head; An output layer of the feature fusion module is further connected with the fifth classification head, an output layer of the confidence decoder is further connected with the sixth classification head, and output layers of the fifth classification head, the sixth classification head, and the forgery detector are connected with the seventh classification head; the output layer of the feature fusion module is an output layer of the Fusion Stage module or an output layer of the multi-scale feature fusion submodule; The forgery detector is specifically configured to perform detection processing on the ninth credential image feature data to obtain a fourth classification probability value. The fifth classification head is configured to perform classification processing on the third credential image feature data to obtain a fifth classification probability value. The sixth classification head is configured to perform classification processing on the eighth credential image feature data to obtain a sixth classification probability value. The seventh classification head is configured to perform classification processing on the fourth classification probability value, the fifth classification probability value, and the sixth classification probability value, to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image. The edge perception module in the boundary guidance network model; the classifier further includes: an eighth classification head; 9. The method of claim 8, the credential anti-spoofing detection model further comprising: The output layer of the first high-resolution network module configured to output the first credential image feature data with the largest scale is further connected with an input layer of the edge perception module, an output layer of the edge perception module is connected with the eighth classification head in the classifier, and the eighth classification head is further connected with the seventh classification head; The edge perception module is configured to perform feature extraction processing on the first credential image feature data with the largest scale, to obtain tenth credential image feature data; The eighth classification head is configured to perform classification processing on the tenth credential image feature data, to obtain a seventh classification probability value; The seventh classification head is specifically configured to perform classification processing on the fourth classification probability value, the fifth classification probability value, the sixth classification probability value, and the seventh classification probability value, to obtain a first anti-counterfeiting detection result reflecting whether the first credential image belongs to a trusted credential image.

10. The method of any one of claims 1-9, wherein the processing the first credential image to obtain a second credential image comprises: performing frequency domain transformation processing on the first credential image to obtain the credential frequency spectrum; wherein the frequency domain transformation processing comprises discrete cosine transformation processing. a noise feature extraction model; 11. The method of any one of claims 2-9, the credential anti-spoofing detection model further comprising: The noise feature extraction model is any one of a Noiseprint extractor, a spatially rich model, and a Bayar convolution model; the noise feature extraction model is connected with the second high-resolution network module; The noise feature extraction model is configured to process the first credential image to obtain the credential noise feature map; or The noise feature extraction model includes at least two models of a Noiseprint extractor, a spatially rich model, and a Bayar convolution model, and a multi-head attention module; the at least two models are connected with the multi-head attention module, and the multi-head attention module is further connected with the second high-resolution network module; The at least two models are configured to process the first credential image to obtain respective initial credential noise feature data; The multi-head attention module is configured to perform feature fusion processing on the respective initial credential noise feature data to obtain the credential noise feature map.

12. The method of any one of claims 1-9, further comprising: determining position information of a region in which key information in the first credential image is located; ​ detect, based on the position information of the region where the key information is located, whether the region where the key information is located is located at a preset position within the first credential image, and / or whether text data within the region where the key information is located conforms to a preset text alignment format, and / or whether the text data within the region where the key information is located uses a preset font, to obtain a second anti-counterfeiting detection result for the first credential image; generate, according to the first anti-counterfeiting detection result and the second anti-counterfeiting detection result, a comprehensive anti-counterfeiting detection result for the first credential image.

13. The method of claim 12, wherein the detecting, based on the position information of the region where the key information is located, whether the region where the key information is located is located at a preset position within the first credential image comprises: determining, based on the position information of the region where the key information is located, whether a first distance between a left side of the region where the key information is located and a vertical center line of the first credential image is equal to a second distance between a right side of the region where the key information is located and the vertical center line; or determining, based on the position information of the region where the key information is located, whether a third distance between the left side of the region where the key information is located and a left side of the first credential image is equal to a fourth distance between a right side of the region where the key information is located and a right side of the first credential image.

14. The method of claim 12, wherein the detecting, based on the position information of the region where the key information is located, whether text data within the region where the key information is located conforms to a preset text alignment format comprises: determining, based on the position information of the region where the key information is located, position information of each piece of text data within the region where the key information is located; determining, according to the position information of the text data, whether vertical coordinates of bottom sides of the text data located in the same row are consistent, and / or whether vertical coordinates of top sides of the text data located in the same row are consistent, and / or whether horizontal coordinates of left sides of the text data located in the same column are consistent, and / or whether horizontal coordinates of right sides of the text data located in the same column are consistent.

15. The method of claim 12, wherein the detecting, based on the position information of the region where the key information is located, whether text data within the region where the key information is located uses a preset font comprises: inputting, based on the position information of the region where the key information is located, a credential sub-image within the region where the key information is located in the first credential image into a font detection model, to obtain a detection result output by the font detection model and used to reflect a target font used by the text data within the region where the key information is located; wherein the font detection model is obtained by pre-training a classification model using image samples carrying label data, and the label data is used to reflect a font category to which text data in the image sample belongs; determining, based on the detection result, whether the target font is consistent with the preset font. ​ 16. The method of claim 12, wherein the detecting, based on the location information of the region where the key information is located, whether the text data in the region where the key information is located is in the preset font, comprises: inputting, based on the location information of the region where the key information is located, a credential sub-image in the region where the key information is located in the first credential image into a font feature extraction model to obtain target font feature data of the text data in the region where the key information is located extracted by the font feature extraction model; wherein the font feature extraction model comprises a self-encoder model pre-trained using image samples containing text data in the preset font; determining whether the target font feature data is consistent with reference font feature data of the preset font; wherein the reference font feature data is font feature data extracted by the font feature extraction model for a trusted credential image consistent with a credential category to which the first credential image belongs.

17. The method of claim 12, further comprising: obtaining an exchangeable image file of the first credential image; determining whether the exchangeable image file contains a counterfeit image keyword to obtain a target determination result; wherein the generating, based on the first anti-counterfeiting detection result and the second anti-counterfeiting detection result, a comprehensive anti-counterfeiting detection result for the first credential image, comprises: generating, based on the first anti-counterfeiting detection result, the second anti-counterfeiting detection result, and the target determination result, a comprehensive anti-counterfeiting detection result for the first credential image.

18. A credential anti-counterfeiting detection device, comprising: a first obtaining module configured to obtain a first credential image obtained by image collection of a target credential of a user; a second obtaining module configured to obtain a second credential image by processing the first credential image, the second credential image comprising any one of a credential noise feature map and a credential spectrum map; an anti-counterfeiting detection module configured to perform anti-counterfeiting detection processing on the first credential image and the second credential image by using a credential anti-counterfeiting detection model to obtain a first anti-counterfeiting detection result for the first credential image output by the credential anti-counterfeiting detection model; wherein the credential anti-counterfeiting detection model comprises a multi-modal feature encoder and a classifier connected in sequence.

19. A computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method of any one of claims 1-17.

20. An electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the steps of the method of any one of claims 1-17.

21. A computer program product having at least one instruction stored thereon, wherein the at least one instruction is executed by a processor to implement the steps of the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Face detection method, and training method and device of face detection model

    CN111461089A

  • Image tampering area detection method based on multiple features and convolutional neural network

    CN111553916A

  • Certificate image copying recognition method and device, electronic equipment and storage medium

    CN111767828A

  • Anti-counterfeiting human face detection method and system, electronic device, program and medium

    WO2018166515A1