Image detection

By combining detection methods that extract features from multiple image types and regions, the problem of identifying forged images has been solved, improving the accuracy and efficiency of image detection and ensuring the security of identity authentication.

WO2026060781A1PCT designated stage Publication Date: 2026-03-26ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and intercept forged images generated by generative artificial intelligence, especially during identity authentication, which threatens user interests and cybersecurity.

Method used

Two image detection methods are used: one is based on multiple types of image feature extraction, and the other is based on feature extraction from multiple image regions. The results of the two methods are combined to determine whether the image passes the detection.

Benefits of technology

It significantly improves the strength and accuracy of image detection, effectively identifying and blocking counterfeit documents and fake facial images, protecting user interests and network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128986_26032026_PF_FP_ABST
    Figure CN2024128986_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image detection method and a related device. The method comprises: acquiring a target image to be detected (S201); respectively extracting preset multiple types of image features from said target image, and performing image detection on said target image on the basis of the extracted multiple types of image features to obtain a first detection result (S202); respectively extracting image features from multiple image regions comprised in said target image, and performing image detection on said target image on the basis of the extracted image features corresponding to the multiple image regions to obtain a second detection result (S203); and in response to at least one of the first detection result and the second detection result indicating that said target image fails the image detection, determining that said target image fails the image detection (S204).
Need to check novelty before this filing date? Find Prior Art

Description

Image detection TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of image detection, and in particular to image detection. BACKGROUND

[0002] AI-Generated Content (AIGC) is a type of artificial intelligence technology that can generate various types of content according to given inputs or instructions, such as text, images, audio, and video, etc., to bring convenience to people's life and work. However, AI images generated by AIGC technology, such as face images and certificate images, etc., are sometimes maliciously used in the identity authentication process of various application programs to falsely identify real users for identity authentication, seriously endangering the interests and network security of users.

[0003] Moreover, with the rapid development of AIGC technology, face images and certificate images generated by AIGC technology are more difficult to be identified, therefore, how to improve the accuracy of image detection is a problem to be solved.

[0004] SUMMARY

[0005] Therefore, one or more embodiments of the present specification provide an image detection method and related equipment.

[0006] In a first aspect, the present specification provides an image detection method, the method comprising: obtaining a target image to be detected; extracting a plurality of types of image features of a preset from the target image respectively, and performing image detection on the target image based on the plurality of types of image features extracted to obtain a first detection result; and extracting image features from a plurality of image regions contained in the target image respectively, and performing image detection on the target image based on the image features corresponding to the plurality of image regions extracted to obtain a second detection result; in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, determining that the target image fails the image detection.

[0007] In a second aspect, the present specification provides an image detection device, comprising: an acquisition unit configured to acquire a target image to be detected; a multi-type feature extraction unit configured to extract a plurality of types of preset image features from the target image respectively, and perform image detection on the target image based on the plurality of types of extracted image features to obtain a first detection result; a multi-region feature extraction unit configured to extract image features from a plurality of image regions included in the target image respectively, and perform image detection on the target image based on the image features corresponding to the plurality of image regions to obtain a second detection result; and an image detection unit configured to determine that the target image fails the image detection in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection.

[0008] Correspondingly, the present specification also provides an electronic device, comprising: a memory and a processor; the memory stores a computer program / instruction executable by the processor; and the processor executes the computer program / instruction to perform the image detection method of the first aspect.

[0009] Correspondingly, the present specification also provides a computer readable storage medium, which stores a computer program / instruction, and the computer program / instruction is executed by a processor to perform the image detection method of the first aspect.

[0010] Correspondingly, the present specification also provides a computer program product, which comprises a computer program / instruction, and the computer program / instruction is executed by a processor to perform the image detection method of the first aspect.

[0011] In summary, the present application can first acquire a target image to be detected, then extract a plurality of types of image features from the target image, and perform image detection on the target image based on the extracted plurality of types of image features to obtain a first detection result. In addition, image features can also be extracted from a plurality of image regions included in the target image, and image detection can be performed on the target image based on the extracted image features corresponding to the plurality of image regions to obtain a second detection result. Further, in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, it can be determined that the target image fails the image detection. In this way, the present application performs image detection on the target image to be detected by two different image detection methods, one of which is based on a plurality of different types of image features extracted from the target image, and the other of which is based on image features extracted from a plurality of different image regions of the target image. By combining the detection results of the above two image detection methods, as long as at least one of the detection results indicates that the target image fails the image detection, it can be determined that the target image fails the image detection, thereby greatly improving the strength and accuracy of image detection. BRIEF DESCRIPTION OF DRAWINGS

[0012] FIG. 1 is a schematic diagram of a system architecture according to an example embodiment;

[0013] FIG. 2 is a schematic diagram of an image detection method according to an example embodiment;

[0014] FIG. 3 is a schematic diagram of an image detection method based on multi-type image feature extraction according to an example embodiment;

[0015] FIG. 4 is a schematic diagram of an image detection method based on multi-region image feature extraction according to an example embodiment;

[0016] FIG. 5 is a schematic diagram of knowledge distillation according to an example embodiment;

[0017] FIG. 6 is a schematic diagram of an image detection device according to an example embodiment;

[0018] FIG. 7 is a schematic diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0019] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals represent like elements, unless the context of use indicates otherwise. The following description of exemplary embodiments is not representative of all embodiments consistent with one or more aspects of the present specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of the present specification as detailed in the appended claims.

[0020] It should be noted that the steps of the methods performed by the respective methods in other embodiments are not necessarily performed in the order shown and described in the present specification. In some other embodiments, the steps included in the methods can be more or less than those described in the present specification. In addition, a single step described in the present specification can be divided into multiple steps for description in other embodiments; and multiple steps described in the present specification can be combined into a single step for description in other embodiments.

[0021] It should be noted that "multiple" as described in the present application refers to two or more.

[0022] In addition, the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0023] First, some terms in the present specification are explained and described to facilitate understanding by those skilled in the art.

[0024] (1) Electronic Know Your Customer (eKYC) is a technology that verifies the identity of a user through electronic means. eKYC can quickly and effectively identify various types of identity cards around the world, automatically extract key identity information and complete identity verification, and further input the extracted identity information into a database or corresponding system, thereby replacing the tedious manual input operation process. In addition, eKYC can also perform biometric verification, such as real-time collection of facial images through a camera, detection and recognition of facial expressions and movements during the interactive process, to determine whether the user is a real person, effectively prevent identity fraud, and ensure the authenticity and reliability of the user, etc., which will not be described in detail here.

[0025] As mentioned above, with the rapid development of AIGC technology, AIGC generated face images and certificate images and the like can be maliciously used in electronic identity authentication processes, seriously endangering the interests and network security of users. Among them, the certificate image generated by AIGC can be a certificate image containing a forged certificate (such as a forged personal identity certificate, a driver's license, a passport, etc.) generated by AIGC technology, and the like, which is not specifically limited in the present specification.

[0026] Based on this, the present specification provides a technical solution that can detect images through two different image detection methods at the same time, improving image detection strength and accuracy.

[0027] In implementation, first, the target image to be detected can be obtained, then the preset multiple types of image features can be extracted from the target image, and the target image can be detected based on the extracted multiple types of image features to obtain a first detection result. In addition, image features can also be extracted from multiple image regions contained in the target image, and the target image can be detected based on the extracted image features corresponding to the multiple image regions to obtain a second detection result. Further, in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, it can be determined that the target image fails the image detection.

[0028] In the above technical solution, the present application detects the target image to be detected through two different image detection methods, one of which is based on multiple different types of image features extracted from the target image, and the other of which is based on image features extracted from multiple different image regions of the target image. In this way, by combining the detection results of the above two image detection methods, as long as at least one of the detection results indicates that the target image fails the image detection, it can be determined that the target image fails the image detection, thereby greatly improving the strength and accuracy of image detection.

[0029] For example, taking a certificate image containing a user's certificate as an example, through the image detection method provided by the present application, the certificate image generated by AIGC technology can bypass one of the image detection methods, but as long as the other image detection method detects that the user's certificate in the certificate image is a fake certificate, it can be determined that it is a fake certificate, thereby effectively intercepting AIGC fake certificates (i.e. fake certificates) in identity verification, protecting the interests and network security of users.

[0030] Referring to FIG. 1, FIG. 1 is a schematic diagram of a system architecture provided in an example embodiment. As shown in FIG. 1, the system can include an electronic device 100, an electronic device 200a, an electronic device 200b, an electronic device 200c, and the like. Among them, the electronic device 200a, the electronic device 200b, and the electronic device 200c can respectively run a corresponding client with an identity verification function, for example, an eKYC client, or any other possible client with an eKYC function, for example, a bank client and a medical service client, and the like, which are not specifically limited in the specification. The electronic device 100 can run a server corresponding to the client with the identity verification function, for example, an eKYC server, for example, a bank server and a medical service server, and the like, which are not specifically limited in the specification. In an example embodiment, the electronic device 100 and the electronic device 200a, the electronic device 200b, and the electronic device 200c can establish a communication connection through any possible way, which is not specifically limited in the specification.

[0031] In the following, the image detection method provided in the present application will be described taking the electronic device 100 and the electronic device 200a as an example.

[0032] In an example embodiment, the user can upload a target image to be detected to the client of the electronic device 200a, and further, the client in the electronic device 200a can send the target image to the electronic device 100, specifically to the server in the electronic device 100.

[0033] Among them, the target image to be detected can be a certificate image containing a user certificate photographed by the user through the electronic device 200a, or a face image containing a user face, and the like, which are not specifically limited in the specification. For example, the user certificate can be a personal identity certificate, a driver's license, a passport, and the like, which are not specifically limited in the specification.

[0034] In an example embodiment, after receiving the target image sent by the client, the server in the electronic device 100 can extract a plurality of types of image features preset from the target image, and perform image detection on the target image based on the plurality of types of image features extracted, to obtain a first detection result, which can be referred to the description of the corresponding embodiment of FIG. 2 below, and will not be expanded here. Among them, the first detection result can be used to indicate that the target image passes / fails the image detection.

[0035] In an illustrated implementation, the server in the electronic device 100 can further extract image features from the plurality of image regions contained in the target image respectively, and perform image detection on the target image based on the extracted image features corresponding to the plurality of image regions to obtain a second detection result. For details, reference can be made to the description of the corresponding embodiment of FIG. 2 below, which will not be elaborated here. The second detection result can be used to indicate that the target image passes / fails the image detection.

[0036] Further, the server can determine a final detection result based on the first detection result and the second detection result. In an illustrated implementation, in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, the server can determine that the final detection result is that the target image fails the image detection.

[0037] In an illustrated implementation, taking the target image as a certificate image containing a user certificate as an example, performing image detection on the target image can include performing anti-counterfeiting detection on the user certificate contained in the target image; correspondingly, the first detection result and the second detection result can be used to indicate whether the user certificate contained in the target image is a counterfeit certificate; correspondingly, in response to at least one of the first detection result and the second detection result indicating that the user certificate contained in the target image is a counterfeit certificate, the server can determine that the final detection result is that the user certificate contained in the target image is a counterfeit certificate.

[0038] In an illustrated implementation, taking the target image as a face image containing a user face as an example, performing image detection on the target image can include performing face anti-counterfeiting detection on the target image; correspondingly, the first detection result and the second detection result can be used to indicate that the target image is a counterfeit face image; correspondingly, in response to at least one of the first detection result and the second detection result indicating that the target image is a counterfeit face image, the server can determine that the final detection result is that the target image is a counterfeit face image.

[0039] Further, the server can send the determined final detection result to the client in the electronic device 200a, and correspondingly, the client can receive the final detection result and output corresponding prompt information to the user based on the final detection result, such as outputting prompt information of “certificate detection fails” or “certificate detection passes” to the user, etc., which will not be limited in the present specification.

[0040] In an illustrated implementation, the server can also directly send the first detection result and the second detection result to the electronic device 200a. Correspondingly, the client can determine the final detection result based on the received first detection result and second detection result. For example, in response to at least one of the received first detection result and second detection result indicating that the user certificate contained in the target image is a fake certificate, the client can determine that the user certificate contained in the target image is a fake certificate.

[0041] It should be noted that in some possible implementations, the client can also perform the above-mentioned method steps of image detection. For example, after obtaining the target image to be detected, the client in the electronic device 200a can directly extract the preset multiple types of image features from the target image, and perform image detection on the target image based on the extracted multiple types of image features to obtain the first detection result, and extract image features from the multiple image regions contained in the target image, and perform image detection on the target image based on the extracted image features corresponding to the multiple image regions to obtain the second detection result. Further, the client can determine the final detection result based on the first detection result and the second detection result, and the like, which will not be described here.

[0042] Alternatively, in some possible implementations, the client and the server can also perform one of the two image detection methods respectively. For example, the client can extract the preset multiple types of image features from the target image, and perform image detection on the target image based on the extracted multiple types of image features to obtain the first detection result; and the server can extract image features from the multiple image regions contained in the target image, and perform image detection on the target image based on the extracted image features corresponding to the multiple image regions to obtain the second detection result, and the like, which will not be limited in the present specification.

[0043] So far, the present application detects the target image to be detected through two different image detection methods at the same time, which improves the image detection intensity and accuracy. Moreover, it should be emphasized that the present application only needs one target image when performing image detection, without the need to collect multiple images for repeated detection, thereby improving the detection efficiency and cost. Taking the anti-fake detection of certificate images as an example, the certificate images generated by AIGC technology can bypass one of the image detection methods, but as long as the other image detection method detects that the user certificate in the certificate image is a fake certificate, it can be determined that it is a fake certificate, thereby effectively intercepting AIGC fake certificates in identity verification, and maintaining user interests and network security.

[0044] In an illustrated embodiment, the electronic devices 200a, 200b and 200c can be smart wearable devices, smart phones, tablet computers, notebook computers and desktop computers with the above functions, and the like, and the present specification does not make specific limitations thereto; the electronic device 100 can be a desktop computer, a server, or a server cluster or a cloud computing center composed of multiple servers with the above functions, and the like, and the present specification does not make specific limitations thereto.

[0045] Referring to FIG. 2, FIG. 2 is a flowchart of an image detection method according to an example embodiment. The method can be applied to the system architecture shown in FIG. 1, and can be applied to any of the electronic devices shown in FIG. 1. As shown in FIG. 2, the method can specifically include the following steps S201-S204.

[0046] In step S201, a target image to be detected is obtained.

[0047] First, a target image to be detected is obtained. The target image can be an identification image containing a user's identification, or can be a face image containing a user's face, and the like, and the present specification does not make specific limitations thereto. In an illustrated embodiment, step S201 can specifically refer to the description of the corresponding embodiment of FIG. 1 above, which will not be repeated here.

[0048] In step S202, a plurality of types of image features are extracted from the target image, and image detection is performed on the target image based on the extracted plurality of types of image features to obtain a first detection result.

[0049] Further, after obtaining the target image to be detected, two different image detection methods can be used to perform image detection on the target image and obtain corresponding detection results.

[0050] In an illustrated embodiment, the two different image detection methods can include image detection based on multi-type image feature extraction and image detection based on multi-region image feature extraction.

[0051] Next, first, the image detection based on multi-type image feature extraction is described.

[0052] In an illustrated embodiment, a plurality of types of image features can be extracted from the target image based on a plurality of types of image feature extraction methods, and further image detection can be performed on the target image based on the extracted plurality of types of image features to obtain a first detection result.

[0053] It should be noted that the above-mentioned multiple types of image features are not particularly limited in the present specification. In an illustrated embodiment, the multiple types of image features can include any of the following: differential features of a target image, frequency domain features, and enhanced features based on Exact Feature Distribution Mixing (EFDMix), etc.

[0054] In an illustrated embodiment, when performing image detection on a target image based on the extracted multiple types of image features, the extracted multiple types of image features can be fused to obtain corresponding first fused features, and then the target image is detected based on the first fused features.

[0055] For example, the specific manner of the above-mentioned feature fusion can be feature concatenation, feature addition, etc., which is not particularly limited in the present specification.

[0056] It should be noted that the specific implementation of the above-mentioned image detection method based on multiple types of image features is not particularly limited in the present specification. In an illustrated embodiment, the above-mentioned image detection method can be implemented by a trained model. Specifically, a target image to be detected can be input into a pre-trained first model, the first model extracts a plurality of types of image features from the target image, and then performs image detection on the target image based on the extracted multiple types of image features, and outputs a first detection result.

[0057] In an illustrated embodiment, please refer to FIG. 3, which is a flowchart of an image detection method based on multiple types of image features according to an example embodiment. As shown in FIG. 3, the first model can be a basic model structure with multi-modal and multi-branch parallel input. Specifically, the first model can include a center difference convolutional network, a Discrete Cosine Transform (DCT) network, and an EFDMix network.

[0058] As shown in FIG. 3, the target image can be input into a center difference convolution network of the first model, and the target image is subjected to center difference convolution processing by the center difference convolution network to obtain difference features of the target image. In parallel, the target image (which can be a copy of the target image) can be input into a DCT transformation network of the first model, and the target image is subjected to discrete cosine transformation processing by the DCT transformation network to obtain frequency domain features of the target image. In parallel, the target image (which can be a copy of the target image) can be input into an EFDMix network of the first model, and the target image is subjected to EFDMix-based image enhancement processing by the EFDMix network to obtain enhanced features of the target image.

[0059] Further, as shown in FIG. 3, the first model can further include a feature fusion module. Correspondingly, the feature fusion module can perform feature fusion on the difference features, the frequency domain features, and the enhanced features of the target image extracted by the plurality of parallel feature extraction branches respectively, thereby obtaining corresponding first fusion features.

[0060] Further, as shown in FIG. 3, the first model can further include a binary classification network. Correspondingly, the feature fusion module can input the first fusion features obtained by fusion into the binary classification network, and the binary classification network performs image detection on the target image based on the first fusion features, for example, determines whether the user certificate contained in the target image is a fake certificate, or for example, determines whether the target image is a fake face image, etc.

[0061] It should be noted that the training process of the first model is not particularly limited in the present specification. In an illustrative embodiment, for the training of the binary classification network, considering that AIGC fake certificates may still have many unknown developments in the future, a single center constraint (One Class Contrastive Loss Function, OCCL) can be introduced as an auxiliary loss function on the basis of the conventional binary classification cross-entropy loss function. The basic principle of the single center constraint is to cluster the features of known real certificates as a single center, and the features of various fake certificates are learned as a divergent distribution away from the single center. Based on this, when a new AIGC fake certificate appears in the future, as long as the features of the AIGC fake certificate are detected to be far away from the center, the model will determine it as a fake certificate and be intercepted.

[0062] In step S203, image features are extracted from a plurality of image regions contained in the target image respectively, and image detection is performed on the target image based on the extracted image features corresponding to the plurality of image regions to obtain a second detection result.

[0063] In the following, image detection based on multi-region image feature extraction will be described.

[0064] In an illustrated embodiment, image features can be extracted from a plurality of image regions included in a target image respectively, and image detection can be performed on the target image based on the extracted image features corresponding to the plurality of image regions, to obtain a second detection result.

[0065] It is to be noted that the plurality of image regions and the image features corresponding to the plurality of image regions are not particularly limited in the present specification.

[0066] In an illustrated embodiment, taking a target image as an example of a user certificate image containing a user certificate, the plurality of image regions can include a full image region of the target image, a certificate region corresponding to the user certificate contained in the target image, and a background region other than the certificate region in the full image region. Correspondingly, the image features corresponding to the plurality of image regions can include full image features, certificate features, and background features; wherein the full image features (or referred to as full image representation) can include an embedding vector for representing image information contained in the full image region, the certificate features can include an embedding vector for representing image information contained in the certificate region, and the background features can include an embedding vector for representing image information contained in the background region.

[0067] In an illustrated embodiment, taking a target image as an example of a user certificate image containing a user certificate, the plurality of image regions can include a full image region of the target image, a certificate region corresponding to the user certificate contained in the target image, and a background region other than the certificate region in the full image region. Correspondingly, the image features corresponding to the plurality of image regions can include full image features, certificate features, and background features; wherein the full image features (or referred to as full image representation) can include an embedding vector for representing image information contained in the full image region, the certificate features can include an embedding vector for representing image information contained in the certificate region, and the background features can include an embedding vector for representing image information contained in the background region.

[0068] In an illustrated embodiment, compared with the background region and the background features, the certificate region (or the face region) and the certificate features (or the face features) can also be referred to as the foreground region and the foreground features, which are not particularly limited in the present specification.

[0069] In an illustrated embodiment, when performing image detection on the target image based on the extracted image features corresponding to the plurality of image regions, the extracted image features corresponding to the plurality of image regions can be fused first, to obtain corresponding second fused features, and then the target image can be detected based on the second fused features.

[0070] Exemplarily, the specific way of the above-mentioned feature fusion can be feature concatenation, or feature addition, etc., which are not particularly limited in the present specification.

[0071] It should be noted that compared with the certificate image of the real certificate, the content of the background region in the AIGC forged certificate image is often limited or even single. Based on this, the present application can accurately detect whether it is a forged certificate based on the background features in the certificate image. In addition, the present application also considers the whole image features and certificate features in the image, which can more comprehensively and accurately detect the forged certificate. It should be understood that the AIGC forged face image is the same, which will not be described here.

[0072] It should be noted that the present application does not particularly limit the specific implementation of the above-mentioned image detection method based on multi-region image feature extraction. In an illustrative embodiment, the above-mentioned image detection method can be implemented by a trained model. Specifically, the target image to be detected can be input into a pre-trained second model, the second model can extract image features from multiple image regions contained in the target image respectively, and the target image can be detected based on the extracted image features corresponding to the multiple image regions, and then the second detection result can be output.

[0073] In an illustrative embodiment, please refer to FIG. 4, which is a flowchart of an image detection method based on multi-region image feature extraction according to an exemplary embodiment. As shown in FIG. 4, the second model can specifically include a ViT (Vision Transformer) model and a SAM (Segmentation-agnostic Mask) model. In an illustrative embodiment, still taking the target image containing a user certificate as an example, the target image can be first input into the ViT model, and the ViT model can extract features of the target image to obtain whole image features of the target image. Further, the whole image features can be input into the SAM model, and the SAM model can segment the whole image features to obtain certificate features corresponding to the certificate region and background features corresponding to the background region.

[0074] In the following, the image detection based on multi-region image feature extraction will be described in combination with the specific structure of the ViT model and the SAM model shown in FIG. 4.

[0075] As shown in FIG. 4, the ViT model can specifically include a linear transformation layer, a position encoding, and a Transformer encoder.

[0076] In an illustrative embodiment, as shown in FIG. 4, the input target image can be first segmented into multiple patches, each patch can be regarded as a “token”, so as to realize the conversion of the target image into sequence data for subsequent processing. As shown in FIG. 4, the size of each patch obtained by segmentation is consistent, for example, each patch is 16x16 pixels.

[0077] Further, as shown in FIG. 4, a linear transformation (usually a fully connected layer) can be performed on the plurality of patches to obtain an embedding vector corresponding to each patch. In addition, as shown in FIG. 4, in order to preserve the spatial position information of each patch, a position encoding can also be added to the embedding vector of each patch, which helps the model to understand the relative position of each patch.

[0078] Further, as shown in FIG. 4, the embedding vectors of the plurality of patches can be input into a Transformer encoder, which can include a Multi-Head Self-Attention module and a Feed-Forward Neural Network (FFN), etc., to further extract the image features of the target image, and finally output the image embedding of the target image.

[0079] As shown in FIG. 4, the SAM model can be connected with an object detector (Grounding Dino).

[0080] In an illustrated embodiment, as shown in FIG. 4, first, the target image and the prompt text (for example, “detect the identity card in the image”) for prompting to detect the object as the identity region can be input into the object detector, and the object detector can detect the identity region from the target image according to the prompt text and output a box corresponding to the identity region.

[0081] Further, as shown in FIG. 4, the image embedding output by the Transformer encoder and the box corresponding to the identity region output by the Grounding Dino can be input into the SAM model, and the SAM model can segment the image embedding based on the box corresponding to the identity region, thereby obtaining the identity feature corresponding to the identity region and the background feature corresponding to the background region.

[0082] In an illustrated embodiment, as shown in FIG. 4, the SAM model can specifically include a Mask decoder and a Prompt encoder. Specifically, as shown in FIG. 4, since the Mask decoder of the SAM model can only process vectors, the box corresponding to the identity region output by the Grounding Dino can be first input into the Prompt encoder of the SAM model, and the Prompt encoder can convert the box into a corresponding vector.

[0083] Further, as shown in FIG. 4, the Prompt encoder can input the vector of the bounding box corresponding to the document region into the Mask decoder of the SAM model, and in addition, the full image representation output by the Transformer encoder can also be input into the Mask decoder of the SAM model. Correspondingly, the Mask decoder can segment the full image representation based on the full image representation and the vector of the bounding box corresponding to the document region, so as to obtain the document feature corresponding to the document region and the background feature (background embedding) corresponding to the background region.

[0084] Wherein, when segmenting the document feature and the background feature, the Mask decoder can specifically realize segmentation by generating a mask corresponding to different regions (i.e. the document region and the background region), which will not be expanded here.

[0085] Alternatively, in some possible embodiments, the SAM model can not be used, but a segmentation module can be separately arranged, which can directly segment the full image representation according to the bounding box of the document region output by the Grounding Dino. For example, the segmentation module can first restore a feature map, such as a feature map, based on the full image representation, each block (i.e. token) in the feature map corresponding to its position information in the image, and then use an Intersection over Union (IoU) calculation method to divide the tokens whose positions fall within the bounding box to the document region and the tokens whose positions fall outside the bounding box to the background region, so as to obtain the corresponding document feature and background feature, etc., which will not be limited in the present specification.

[0086] Further, as shown in FIG. 4, the second model can further include a feature fusion module, and correspondingly, the feature fusion module can fuse the image features (full image feature, document feature and background feature) corresponding to the plurality of image regions extracted to obtain a corresponding second fusion feature.

[0087] Further, as shown in FIG. 4, the second model can further include a feature fusion module, and correspondingly, the feature fusion module can fuse the image features (full image feature, document feature and background feature) corresponding to the plurality of image regions extracted to obtain a corresponding second fusion feature.

[0088] In an illustrated embodiment, the feature comparison module can include fully connected layers, which are not specifically limited in the present description.

[0089] In an illustrated embodiment, the fake ID feature library can be constructed based on the multi-region fusion features of existing ID images containing fake IDs. In an illustrated embodiment, if an image is detected as a fake ID image by the image detection method described in step S202, or is manually reviewed as a fake ID image, then the image features can be extracted from the multiple image regions contained in the image based on the multi-region image feature extraction method described in step S203, the image features corresponding to the multiple image regions are fused, and the fused features are input into the fake ID feature library, thereby constructing the fake ID feature library.

[0090] In an illustrated embodiment, when determining whether the user ID contained in the target image is a fake ID based on the comparison result, it can specifically include: if the similarity between the second fusion feature and any of the multiple features contained in the fake ID feature library is greater than a preset threshold (e.g., 80% or 85%), then it can be determined that the user ID contained in the target image is a fake ID; otherwise, it is determined that the user ID contained in the target image is a real ID.

[0091] In an illustrated embodiment, to improve the efficiency of feature comparison, the multiple features contained in the fake ID feature library can also be K-means clustered to obtain K feature clusters; wherein the K feature clusters can correspond to K center features, and the center feature corresponding to each feature cluster can be used to represent the average feature of all features contained in the feature cluster; K is an integer greater than 1. Accordingly, if the similarity between the second fusion feature and any of the K center features is greater than a preset threshold, then it can be determined that the user ID contained in the target image is a fake ID; otherwise, it is determined that the user ID contained in the target image is a real ID.

[0092] Referring to FIG. 5, FIG. 5 is a schematic diagram of knowledge distillation provided by an exemplary embodiment. As shown in FIG. 5, considering that the original ViT-based image encoder is very large and it takes a long time to output image embedding, in order to improve the timeliness of model deployment, the MobileSAM model is referred to for image feature extraction, the coupling optimization between the image encoder and the prompt encoder in the SAM model is resolved, and the heavy ViT-based (large) image encoder is distilled into a light ViT-based (small) image encoder. As shown in FIG. 5, for the Mask decoder in the SAM model, the worker can autonomously select whether to adjust according to actual conditions and needs, which is not specifically limited in the specification.

[0093] It should be noted that the training process of the second model is not particularly limited in the specification. In an exemplary embodiment, the model can be supervised trained based on a large number of training samples containing real certificates and fake certificates and their corresponding sample labels. In an exemplary embodiment, considering that the training samples of fake certificates that can be obtained are less, the model can be trained by using softmax and focal loss as a loss function together to solve the problem of unbalanced number of training samples of real and fake classes and improve the performance of the model on the minority class.

[0094] In addition, the extraction of certificate features and background features can be separately supervised trained.

[0095] For example, taking the training of the certificate feature extraction part as an example, first, the training sample is input into the model structure shown in FIG. 4, and the mask decoder (or a separately set segmentation module) outputs the corresponding certificate feature. Then, the background area in the training sample is masked to cover the background area (equivalent to cutting off the background area from the training sample), and the masked training sample is input into the model structure shown in FIG. 4, at this time, the image feature output by the ViT model is the certificate feature of the certificate area in the training sample. Further, the image feature output by the ViT model can be compared with the certificate feature output by the mask decoder, so as to realize the supervised training of the certificate feature extraction part.

[0096] Exemplarily, taking the training of the background feature extraction part as an example, first, the training sample is input into the model structure shown in FIG. 4, and the mask decoder (or a separately arranged segmentation module) outputs the corresponding background feature. Then, the certificate region in the training sample is subjected to mask processing to cover the certificate region (equivalent to cropping the certificate region from the training sample), and the training sample after mask processing is input into the model structure shown in FIG. 4, at this time, the full-image feature output by the ViT model is the background feature of the background region in the training sample. Further, the full-image feature output by the ViT model can be compared with the background feature output by the mask decoder, so as to realize the supervised training of the background feature extraction part.

[0097] In step S204, in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, it is determined that the target image fails the image detection.

[0098] Further, in response to at least one of the first detection result obtained in the above step S202 and the second detection result obtained in step S203 indicating that the target image fails the image detection, it is determined that the target image fails the image detection; otherwise, if both the first detection result and the second detection result indicate that the target image passes the image detection, it is determined that the target image passes the image detection.

[0099] Exemplarily, taking the target image as a certificate image containing a user certificate as an example, if at least one of the first detection result and the second detection result indicates that the user certificate contained in the target image is a fake certificate, it is determined that the user certificate contained in the target image is a fake certificate.

[0100] Exemplarily, taking the target image as a face image containing a user face as an example, if at least one of the first detection result and the second detection result indicates that the target image is a fake face image, it is determined that the target image is a fake face image, rather than a face image of a real user.

[0101] Corresponding to the method flow implementation, the embodiment of the present specification also provides an image detection device. Please refer to FIG. 6, which is a structural schematic diagram of an image detection device provided by an exemplary embodiment. The device 60 can be applied to the system architecture shown in FIG. 1, and specifically can be applied to any electronic device shown in FIG. 1. As shown in FIG. 6, the device 60 comprises:

[0102] The acquisition unit 601 is configured to acquire a target image to be detected.

[0103] The multi-type feature extraction unit 602 is configured to extract a plurality of types of preset image features from the target image respectively, and perform image detection on the target image based on the plurality of types of extracted image features, to obtain a first detection result.

[0104] The multi-region feature extraction unit 603 is configured to extract image features from a plurality of image regions included in the target image respectively, and perform image detection on the target image based on the image features corresponding to the plurality of image regions, to obtain a second detection result.

[0105] The image detection unit 604 is configured to determine that the target image fails the image detection, in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection.

[0106] In an illustrated embodiment, the target image is a certificate image including a user certificate, and the image detection includes anti-forgery detection on the user certificate included in the target image; and the image detection unit 604 is specifically configured to determine that the user certificate included in the target image is a forged certificate, in response to at least one of the first detection result and the second detection result indicating that the user certificate included in the target image is a forged certificate.

[0107] In an illustrated embodiment, the multi-type feature extraction unit 602 is specifically configured to input the target image into a first model pre-trained, extract a plurality of types of preset image features from the target image by the first model respectively, and perform image detection on the target image based on the plurality of types of extracted image features.

[0108] In an illustrated embodiment, the plurality of types of image features include differential features of the target image; the first model includes a center difference convolution network; and the multi-type feature extraction unit 602 is specifically configured to input the target image into the center difference convolution network included in the first model, perform center difference convolution processing on the target image by the center difference convolution network, and obtain the differential features of the target image.

[0109] In an illustrated embodiment, the plurality of types of image features include frequency domain features of the target image; the first model includes a discrete cosine transform network; and the multi-type feature extraction unit 602 is specifically configured to input the target image into the discrete cosine transform network included in the first model, perform discrete cosine transform processing on the target image by the discrete cosine transform network, and obtain the frequency domain features of the target image.

[0110] In an illustrated implementation, the multiple types of image features include enhanced features based on a feature distribution mixture EFDMix, and the first model includes an EFDMix network; and the multi-type feature extraction unit 602 is specifically configured to: input the target image into the EFDMix network included in the first model, and perform EFDMix-based image enhancement processing on the target image by the EFDMix network to obtain enhanced features of the target image.

[0111] In an illustrated implementation, the first model includes a binary classification network; and the multi-type feature extraction unit 602 is specifically configured to: perform feature fusion on the extracted multiple types of image features to obtain corresponding first fused features; input the first fused features into the binary classification network included in the first model, and determine, by the binary classification network based on the first fused features, whether the user certificate contained in the target image is a fake certificate.

[0112] In an illustrated implementation, a loss function used by the binary classification network in a training process includes a binary cross-entropy loss function and an one-center constraint OCCL loss function; and the OCCL loss function is used to aggregate, in a feature space, training samples in the training sample set of the binary classification network whose sample labels are real certificates.

[0113] In an illustrated implementation, the multi-region feature extraction unit 603 is specifically configured to: input the target image into a second model trained in advance, extract image features from multiple image regions contained in the target image by the second model respectively, and perform image detection on the target image based on the extracted image features corresponding to the multiple image regions.

[0114] In an illustrated implementation, the multiple image regions include: a full-image region of the target image, a certificate region corresponding to a user certificate contained in the target image, and a background region other than the certificate region in the full-image region; and the image features corresponding to the multiple image regions include: full-image features, certificate features, and background features; wherein the full-image features include an embedding vector for representing image information contained in the full-image region, the certificate features include an embedding vector for representing image information contained in the certificate region, and the background features include an embedding vector for representing image information contained in the background region.

[0115] In an illustrated implementation, the second model comprises a ViT model and a SAM model; and the multi-region feature extraction unit 603 is specifically configured to: input the target image into the ViT model, perform feature extraction on the target image by the ViT model, and obtain a full-image feature of the target image; and input the full-image feature into the SAM model, perform segmentation on the full-image feature by the SAM model, and obtain a document feature corresponding to the document region and a background feature corresponding to the background region.

[0116] In an illustrated implementation, the second model further comprises an object detector; and the multi-region feature extraction unit 603 is specifically configured to: input the target image and prompt text for prompting a detection object as the document region into the object detector, detect the document region from the target image according to the prompt text by the object detector, and output a bounding box corresponding to the document region; and input the full-image feature and the bounding box into the SAM model, and perform segmentation on the full-image feature based on the bounding box by the SAM model.

[0117] In an illustrated implementation, the multi-region feature extraction unit 603 is specifically configured to: perform feature fusion on the extracted image features corresponding to the plurality of image regions to obtain corresponding second fusion features; compare the second fusion features with a plurality of features contained in a pre-constructed fake document feature library, and determine whether the user document contained in the target image is a fake document based on a comparison result; and each feature contained in the fake document feature library is a fusion feature obtained by performing feature fusion on image features corresponding to a plurality of image regions contained in a document image containing a fake document.

[0118] In an illustrated implementation, the multi-region feature extraction unit 603 is specifically configured to: if a similarity between the second fusion feature and any one of the plurality of features contained in the fake document feature library is greater than a preset threshold, determine that the user document contained in the target image is a fake document.

[0119] In an illustrated implementation, the multi-region feature extraction unit 603 is specifically configured to: perform K-means clustering on the plurality of features contained in the fake document feature library to obtain K feature clusters; wherein the K feature clusters correspond to K center features, and each center feature corresponding to a feature cluster is used to represent an average feature of all features contained in the feature cluster; K is an integer greater than 1; and if a similarity between the second fusion feature and any one of the K center features is greater than a preset threshold, determine that the user document contained in the target image is a fake document.

[0120] The implementation process of the functions and roles of each unit in the apparatus 60 is specifically described in the above embodiments, and will not be repeated here. It should be understood that the apparatus 60 can be implemented by software, or by hardware or a combination of software and hardware. For example, the software implementation, as a logical device, is formed by reading the corresponding computer program instructions into the memory by the processor (CPU) of the device. From the hardware level, in addition to the CPU and the memory, the device where the apparatus is located usually also includes other hardware such as a chip for wireless signal transceiving and / or other hardware such as a board for implementing network communication function.

[0121] The apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical modules, i.e., they can be located in one place or distributed on multiple network modules. Some or all of the units or modules can be selected to achieve the purposes of the schemes of the present specification according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0122] The apparatus, units and modules described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, an in-vehicle computer or a combination of any of these devices.

[0123] Corresponding to the above method embodiments, the embodiments of the present specification also provide an electronic device. Please refer to FIG. 7, which is a structural schematic diagram of an electronic device provided in an exemplary embodiment. The electronic device can be any electronic device in the system architecture shown in FIG. 1. As shown in FIG. 7, the electronic device includes a processor 1001 and a memory 1002, and can further include an input device 1004 (such as a keyboard, etc.) and an output device 1005 (such as a display, etc.). The processor 1001, the memory 1002, the input device 1004 and the output device 1005 can be connected by a bus or other means. As shown in FIG. 7, the memory 1002 includes a computer readable storage medium 1003 which stores a computer program capable of being run by the processor 1001. The processor 1001 can be a CPU, a microprocessor, or an integrated circuit for controlling the execution of the above method embodiments. When the processor 1001 runs the stored computer program, it can execute each step of the image detection method in the embodiments of the present specification.

[0124] The detailed description of each step of the image detection method is described above, and will not be repeated here.

[0125] Corresponding to the method embodiments described above, the embodiments of the present specification also provide a computer readable storage medium, and the storage medium stores computer programs, and the computer programs execute each step of the image detection method in the embodiments of the present specification when run by a processor. For details, please refer to the description of the above embodiments, which will not be repeated here.

[0126] The above only describes the preferred embodiments of the present specification and does not limit the present specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the protection scope of the present specification.

[0127] In a typical configuration, a terminal device includes one or more CPUs, input / output interfaces, network interfaces, and memories.

[0128] The memory can include non-permanent memory in a computer readable medium, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer readable medium.

[0129] The computer readable medium includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data.

[0130] Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition in this paper, computer readable medium does not include transitory computer readable medium such as modulated data signal and carrier wave.

[0131] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0132] Those skilled in the art will appreciate that embodiments of the present specification can be devised for a method, a system, or a computer program product. Accordingly, embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, embodiments of the present specification can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer-readable program code thereon for use by or in connection with an instruction execution system.

Claims

1. An image detection method characterized by, The method comprises: obtaining a target image to be detected; extracting a plurality of types of image features pre-set from the target image respectively, and performing image detection on the target image based on the plurality of types of image features extracted to obtain a first detection result; and extracting image features from a plurality of image regions included in the target image respectively, and performing image detection on the target image based on the image features corresponding to the plurality of image regions extracted to obtain a second detection result; in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, determining that the target image fails the image detection.

2. The method of claim 1, wherein, The target image is a certificate image containing a user certificate, and the image detection comprises anti-fake detection on the user certificate contained in the target image; the response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection, determining that the target image fails the image detection, comprises: in response to at least one of the first detection result and the second detection result indicating that the user certificate contained in the target image is a fake certificate, determining that the user certificate contained in the target image is a fake certificate.

3. The method of claim 2, wherein, The method comprises: inputting the target image into a first model pre-trained, extracting a plurality of types of image features pre-set from the target image by the first model respectively, and performing image detection on the target image based on the plurality of types of image features extracted.

4. The method of claim 3, wherein, The plurality of types of image features include differential features of the target image; the first model includes a center difference convolution network; the plurality of types of image features extracted from the target image respectively comprise: inputting the target image into the center difference convolution network included in the first model, and performing center difference convolution processing on the target image by the center difference convolution network to obtain the differential features of the target image.

5. The method of claim 3, wherein, The plurality of types of image features include frequency domain features of the target image; the first model includes a discrete cosine transform network; the plurality of types of image features extracted from the target image respectively comprise: inputting the target image into the discrete cosine transform network included in the first model, and performing discrete cosine transform processing on the target image by the discrete cosine transform network to obtain the frequency domain features of the target image.

6. The method of claim 3, wherein, The plurality of types of image features include enhanced features based on feature distribution mixing EFDMix, and the first model includes an EFDMix network; The plurality of types of image features extracted from the target image respectively comprise: The target image is input into the EFDMix network included in the first model, and the target image is subjected to EFDMix-based image enhancement processing by the EFDMix network to obtain enhanced features of the target image.

7. The method of claim 3, wherein, The first model includes a binary classification network; the image detection on the target image based on the extracted multiple types of image features includes: The multiple types of image features extracted are fused to obtain corresponding first fused features; The first fused features are input into the binary classification network included in the first model, and whether the user certificate contained in the target image is a fake certificate is determined by the binary classification network based on the first fused features.

8. The method of claim 7, wherein, The loss function used by the binary classification network in the training process includes a binary cross-entropy loss function and a one-center constraint OCCL loss function; wherein the OCCL loss function is used to aggregate the training samples in the training sample set of the binary classification network in the feature space.

9. The method of claim 2, wherein, The multiple image regions included in the target image are extracted from the target image, and the image detection on the target image based on the extracted image features corresponding to the multiple image regions includes: The target image is input into a second model trained in advance, and the second model extracts image features from multiple image regions included in the target image, and performs image detection on the target image based on the extracted image features corresponding to the multiple image regions.

10. The method of claim 9, wherein The multiple image regions include: a full image region of the target image, a certificate region corresponding to the user certificate contained in the target image, and a background region other than the certificate region in the full image region; The image features corresponding to the multiple image regions include: full image features, certificate features, and background features; wherein the full image features include an embedding vector for representing image information contained in the full image region, the certificate features include an embedding vector for representing image information contained in the certificate region, and the background features include an embedding vector for representing image information contained in the background region.

11. The method of claim 10, wherein, The second model includes a ViT model and a SAM model; the image features are extracted from the multiple image regions included in the target image, including: The target image is input into the ViT model, and the ViT model extracts features from the target image to obtain full image features of the target image; The full image features are input into the SAM model, and the SAM model segments the full image features to obtain certificate features corresponding to the certificate region and background features corresponding to the background region.

12. The method of claim 11, wherein, The second model also includes an object detector; the full image features are input into the SAM model, and the SAM model segments the full image features, including: input the target image and prompt text for prompting the detection object as the certificate region into the object detector, detect the certificate region from the target image according to the prompt text by the object detector, and output a bounding box corresponding to the certificate region; input the full-image feature and the bounding box into the SAM model, and segment the full-image feature based on the bounding box by the SAM model.

13. The method of claim 9, wherein, The image detection on the target image based on the extracted image features corresponding to the plurality of image regions comprises: performing feature fusion on the extracted image features corresponding to the plurality of image regions to obtain corresponding second fusion features; comparing the second fusion features with a plurality of features contained in a pre-constructed fake certificate feature library, and determining whether the user certificate contained in the target image is a fake certificate based on a comparison result; wherein each feature contained in the fake certificate feature library is a fusion feature obtained by performing feature fusion on image features corresponding to a plurality of image regions contained in a certificate image containing a fake certificate.

14. The method of claim 13, wherein, The determination of whether the user certificate contained in the target image is a fake certificate based on the comparison result comprises: if the similarity between the second fusion features and any one of the plurality of features contained in the fake certificate feature library is greater than a preset threshold, it is determined that the user certificate contained in the target image is a fake certificate.

15. The method of claim 14, wherein, The determination of whether the user certificate contained in the target image is a fake certificate based on the comparison result comprises: perform K-means clustering on the plurality of features contained in the fake certificate feature library to obtain K feature clusters; wherein the K feature clusters correspond to K center features, the center feature corresponding to each feature cluster is used to represent the average feature of all features contained in the feature cluster; K is an integer greater than 1; if the similarity between the second fusion features and any one of the K center features is greater than a preset threshold, it is determined that the user certificate contained in the target image is a fake certificate.

16. An image detection apparatus characterized by comprising: The apparatus comprises: an acquisition unit configured to acquire a target image to be detected; a multi-type feature extraction unit configured to extract a plurality of types of image features preset from the target image, and perform image detection on the target image based on the extracted plurality of types of image features to obtain a first detection result; a multi-region feature extraction unit configured to extract image features from a plurality of image regions contained in the target image, and perform image detection on the target image based on the extracted image features corresponding to the plurality of image regions to obtain a second detection result; an image detection unit configured to determine that the target image fails the image detection in response to at least one of the first detection result and the second detection result indicating that the target image fails the image detection.

17. An electronic device, comprising: comprise: a memory and a processor; the memory has stored thereon computer programs / instructions executable by the processor; The processor executes the computer program / instructions to perform the method of any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that, A computer program product having stored thereon computer program / instructions which, when executed by a processor, implement the method of any one of claims 1 to 15.

19. A computer program product, characterised in that, The computer program product comprises computer program / instructions which, when executed by a processor, implement the method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • Certificate identification method and apparatus, electronic device and storage medium

    CN108229499A

  • Face image detection method and device, model training method and device and storage medium

    CN114913565A

  • Image detection method and device, readable medium and electronic equipment

    CN115294662A

  • Image processing method and device, storage medium and electronic equipment

    CN117576388A

  • Image auditing method and system based on image segmentation and retrieval enhancement generation technology

    CN118210937A