Certificate detection method

By acquiring video of documents under varying acquisition conditions, the system identifies and detects multiple anti-counterfeiting features in the documents, solving the problem of difficulty in identifying counterfeit documents and achieving more accurate and secure document detection.

CN120932157APending Publication Date: 2025-11-11HANGZHOU ANT KUAI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511064974.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing eKYC identity verification technologies, counterfeit documents are difficult to effectively identify through dynamic or three-dimensional anti-counterfeiting features, resulting in insufficient accuracy in document detection.

Method used

By acquiring video of documents under varying acquisition conditions, various anti-counterfeiting features in the documents are identified, and comprehensive detection is performed using corresponding detection methods, including adjustments to lighting and angles to highlight the anti-counterfeiting features.

Benefits of technology

It improves the accuracy and security of document detection, effectively identifies dynamic changes in genuine documents, and enhances the comprehensiveness and reliability of document detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932157A_ABST
    Figure CN120932157A_ABST
Patent Text Reader

Abstract

The invention discloses a certificate detection method, and the method comprises the steps: obtaining a target video which is obtained by carrying out the video collection of a certificate in the process of changing a collection condition; based on the type of the certificate, determining a plurality of anti-counterfeiting features to be detected in the certificate; based on the detection method corresponding to each anti-counterfeiting feature, detecting the target video to determine whether the certificate comprises each anti-counterfeiting feature; and determining the authenticity of the certificate based on the detection result. The method is more accurate in detection result and higher in safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification belong to the field of image processing technology, and in particular relate to a document detection method. Background Technology

[0002] In the financial sector, user authentication is often achieved through document verification. For example, in eKYC (electronic Know Your Customer) scenarios, a key step is verifying a user's identity based on a single image of their identification document. However, some criminals, engaged in illicit activities, deceive verification systems by submitting forged document images to gain illegal profits. This includes printing out color copies of documents, using photographs of documents, or even forging the document itself to impersonate a user's genuine identification for authentication. Therefore, it is crucial to identify this type of document forgery during document detection.

[0003] Currently, in eKYC identity verification, the single image of the document only displays limited information, and some dynamic or three-dimensional anti-counterfeiting features (such as laser patterns, color-changing ink, embossing, etc.) cannot be clearly displayed or fully presented. This allows counterfeiters to imitate the distinctive appearance of the document to pass verification. Therefore, a more reliable document detection solution is needed to ensure comprehensive verification of document authenticity. Summary of the Invention

[0004] The purpose of this invention is to provide a document detection method to improve the accuracy of document detection.

[0005] This specification provides a document detection method, comprising: acquiring a target video, the target video being obtained by capturing video of a document under varying acquisition conditions; determining multiple anti-counterfeiting features to be detected in the document based on the type of the document; detecting the target video based on detection methods corresponding to each of the anti-counterfeiting features to determine whether the document includes each of the anti-counterfeiting features; and determining the authenticity of the document based on the detection results.

[0006] A second aspect of this specification provides a computing device including a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method described in the first aspect.

[0007] In the solutions provided by the above embodiments of this specification, by detecting the target video of the document during the process of changing the acquisition conditions, it is possible to detect the dynamically changing anti-counterfeiting features in the real document. The various types of anti-counterfeiting features of the document are comprehensively detected using their corresponding detection methods, making the detection results of the document more accurate and the security stronger. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a comparison diagram of a genuine document and a counterfeit document shown in one embodiment of this specification;

[0010] Figure 2 This is a flowchart illustrating a document detection method according to one embodiment of this specification;

[0011] Figure 3 This is a schematic diagram illustrating the location of anti-counterfeiting features in a document according to one embodiment of this specification;

[0012] Figure 4 This is a schematic diagram of the document inspection process in one embodiment of this specification;

[0013] Figure 5 This is a schematic diagram of the structure of a video segmentation model in one embodiment of this specification;

[0014] Figure 6 This is a schematic diagram of a document detection device shown in one embodiment of this specification. Detailed Implementation

[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0016] As mentioned earlier, in eKYC identity verification, the authenticity of the document relies on the detection of its anti-counterfeiting features. Currently, documents made of materials such as PVC (Polyvinyl Chloride) and PC (Polycarbonate) typically possess anti-counterfeiting features related to lighting conditions and angles, such as laser patterns, color-changing inks, embossing, and 3D watermarks. Different anti-counterfeiting features will change under different conditions. For example, the Great Wall logo on Type I identity documents is usually laser-engraved or optically variable, and is only visible under specific lighting conditions; it is invisible under other lighting conditions. Similarly, the logo on Type II identity documents, made with optically variable ink, is only visible under specific lighting and angle conditions, and is invisible or difficult to detect under other conditions.

[0017] Figure 1 This is a schematic diagram illustrating an implementation scenario of one embodiment disclosed in this specification. This implementation scenario involves document anti-counterfeiting detection. It is understood that document anti-counterfeiting detection typically involves material forgery; for example, a printed image or photograph of a genuine document is a counterfeit. Additionally, there are other counterfeit documents produced by black market activities. Due to the difficulty of forgery, counterfeit documents do not possess the aforementioned dynamically changing anti-counterfeiting features. This embodiment in this specification uses varying acquisition conditions to highlight these anti-counterfeiting features, thereby detecting whether a document is genuine or counterfeit. (Refer to...) Figure 1 The example shows a genuine ID card and its corresponding counterfeit version. Both documents share the same text and user image. However, the genuine ID card utilizes optical holographic anti-counterfeiting film technology, exhibiting a laser-like reflective and color-changing visual effect under different light intensities and holding angles (i.e., acquisition condition 1). It also employs optically variable ink to create a special mark, which displays the same reflective and color-changing visual effect under another different light intensity and holding angle (i.e., acquisition condition 2), thus forming a unique anti-counterfeiting feature. The counterfeit ID card lacks these visual effects and is difficult to perfectly replicate using both anti-counterfeiting features. Figure 1 The dotted circle around the holographic pattern and the special mark illustrates the difference in visual effect. It's understandable that the actual visual effect is much richer; for example, the shape of the holographic pattern or the special mark can be various, and the color can be a variety of colors. However, if ekYC captures an image of a document for inspection, it will be difficult to simultaneously detect both of these different anti-counterfeiting features in a genuine document, thus making it impossible to accurately distinguish between genuine and counterfeit documents.

[0018] To address the aforementioned technical problems, embodiments of this specification provide a document detection method. During document detection, a target video is first acquired, obtained by capturing video of the document under varying acquisition conditions. Then, based on the document type, various anti-counterfeiting features to be detected in the document are identified. The target video is then detected using the detection methods corresponding to each anti-counterfeiting feature to determine whether the document contains each anti-counterfeiting feature. Finally, based on the detection results, the authenticity of the document is determined.

[0019] In this way, by detecting the target video captured by the document under changing acquisition conditions, it is possible to detect the dynamically changing anti-counterfeiting features in the real document. The various types of anti-counterfeiting features of the document can be comprehensively detected using the corresponding detection methods, making the detection results of the document more accurate and the security stronger.

[0020] Figure 2 A flowchart of a document detection method according to one embodiment is shown, which can be based on Figure 1 The implementation scenario shown illustrates that this detection process can be performed by any device, platform, or cluster of devices with computing and processing capabilities. For example... Figure 2 As shown, the document detection method in this embodiment includes the following steps:

[0021] In step 202, the target video is acquired.

[0022] The target video is obtained by capturing video of the document under varying acquisition conditions. This embodiment does not limit the specific method of acquiring the target video. The device for acquiring the target video can be a user's mobile terminal, a service terminal, or other devices with video acquisition capabilities. Acquisition conditions may include different lighting conditions and / or different angle conditions. It is understood that the above lighting conditions may include light intensity, light source type, lighting time, and light direction. The light source type affects the type of light, such as visible light or invisible light; the light direction affects the angle between the light and the document to be detected. The above angle conditions may include the angle between the device acquiring the target video and the document, the focal length of the device's camera, etc. Different angle conditions can be achieved by adjusting the tilt angle of the device or the document, or by adjusting the focal length of the device's camera.

[0023] In practice, to acquire target videos that meet the requirements of changing acquisition conditions, at least one target operation to be performed can be determined based on the type of identification document before acquiring the target video: prompting the user to adjust the pose of the device acquiring the target video, for example, the user can adjust the angle and position of the handheld mobile terminal; prompting the user to adjust the placement posture of the identification document, for example, the user can move the identification document to adjust the placement angle, height, or tilt state, etc.; prompting the user to control the device's sensors, for example, the user can turn on the device's flash or other lights, or adjust the focal length of the device's camera; controlling the device's sensors, for example, the system directly controls the device's flash to turn on or flash, and directly sets a new focal length for the camera without user operation; the sensors include the flash and / or the camera; then, the target video acquired by the device during the execution of the target operation is acquired.

[0024] The aforementioned prompts to the user can be text prompts on the device's target video acquisition interface, or audio prompts via a speaker or other sound-emitting device. This embodiment does not limit the specific prompting method or content. In other embodiments, the user can be directly prompted to change the lighting conditions, allowing them to adjust the ambient light or change their environment. Different types of documents require different target operations. For example, when the anti-counterfeiting features included in the document need to be detected under different lighting conditions, the required target operation is related to the lighting conditions, such as prompting the user to control the device's flash, or directly controlling the device's flash. Similarly, when the anti-counterfeiting features included in the document need to be detected under different angle conditions, the required target operation is related to the angle conditions, such as prompting the user to adjust the angle of the device or document, or directly controlling the device's camera focus. By comprehensively considering the changes in anti-counterfeiting features under different acquisition conditions and interacting with the user in real time to guide the acquisition, ensuring that anti-counterfeiting features are acquired as much as possible, the comprehensiveness and accuracy of the detection can be improved.

[0025] In one implementation, the process of acquiring the target video described above can be performed simultaneously with subsequent step 206. This involves real-time acquisition of the target video and detection of anti-counterfeiting features, using a dynamic guidance mechanism to prompt and interact with the user, encouraging their cooperation in the acquisition process. For example, when a user uses their mobile phone to acquire the target video for identity verification, a corresponding verification interface, such as an SDK (Software Development Kit) interface, can be opened for document verification. The user can then position the document according to the SDK interface prompts and acquire a video sequence of the target video. Every N seconds, this N-second video is used for anti-counterfeiting feature detection. When the target feature to be detected requires a change in lighting conditions, the SDK can automatically turn on the flash. When an angle change is required, the SDK reminds the user to adjust the angle of the phone or document. In other implementations, step 206 can be performed after acquiring the complete target video under varying acquisition conditions.

[0026] In step 204, based on the type of document, various anti-counterfeiting features to be detected in the document are determined.

[0027] In practice, anti-counterfeiting features on identification documents can include at least one of the following: optically variable ink, laser patterns, embossing, watermarks, holograms, iridescent effects, colored ultraviolet patterns, and 3D laser images. Different types of identification documents contain different types of anti-counterfeiting features. For example, Type 1 identification documents include features such as the Great Wall logo created using optical variable ink and embossed lettering created using laser engraving. Type 2 identification documents include features such as optically variable ink, colored ultraviolet patterns, and 3D laser images. Type 3 identification documents include laser patterns covering the entire document. A mapping relationship between different document types and different anti-counterfeiting features can be pre-established, allowing for the determination of multiple anti-counterfeiting features to be detected based on the document type. For instance, based on a pre-built mapping table between different document types and anti-counterfeiting features, the document type can be queried in the table to determine multiple anti-counterfeiting features matching the document type.

[0028] It should be noted that the embodiments in this specification do not restrict the execution order of steps 202 and 204; they can be executed sequentially or in parallel. The following two examples illustrate two execution orders:

[0029] In Example 1, the user can be informed first that the target video is being collected during the process of changing the collection conditions. Then, the target video is identified to determine the type of document in the target video. Based on the type of document, various anti-counterfeiting features to be detected in the document are then determined.

[0030] In Example 2, the user-selected or pre-specified document type can be obtained first, and based on the document type, multiple anti-counterfeiting features to be detected in the document can be determined. Then, based on the document type, at least one target operation to be performed can be determined. During the acquisition of the target video, at least one target operation can be performed to acquire the target video while the acquisition conditions are changing.

[0031] In this step, when determining the various anti-counterfeiting features to be detected based on different document types, the corresponding detection methods can also be determined. For example, an anti-counterfeiting feature library can be pre-built based on the document type. This library contains the names of various anti-counterfeiting features in different types of documents, triggering conditions (i.e., different acquisition conditions, specifically light intensity thresholds and angle ranges), the location of the anti-counterfeiting features, and the detection methods. Then, the target operation to be performed is determined based on the triggering conditions, such as brightening the light or adjusting the angle. For example, a portion of the anti-counterfeiting feature library can be shown in Table 1 below:

[0032] Document Types Anti-counterfeiting feature name Data collection conditions Location Detection methods Type 1 identification documents Category 1 anti-counterfeiting features Lighting conditions coordinate Video segmentation detection Type 1 identification documents Second type of anti-counterfeiting features angle condition coordinate Target detection Second type of identification document Category 1 anti-counterfeiting features Lighting conditions coordinate Video segmentation detection Second type of identification document Third type of anti-counterfeiting features Lighting and angle conditions coordinate Video segmentation detection Third type of identity document Third type of anti-counterfeiting features Lighting and angle conditions All documents Video segmentation detection

[0033] Table 1

[0034] Once the user-selected document type is recognized, the corresponding anti-counterfeiting feature library is loaded, and the capture of the target video begins. Furthermore, the detection priority of different anti-counterfeiting features can be initialized. For example, difficult-to-counterfeit dynamic features, such as optically variable ink, can be detected first, followed by static features, such as name tags and portraits, which are simpler to check. This priority setting helps ensure that more difficult-to-counterfeit anti-counterfeiting features are fully checked, improving verification efficiency.

[0035] In step 206, the target video is detected based on the detection method corresponding to each anti-counterfeiting feature to determine whether the document includes each anti-counterfeiting feature.

[0036] Different detection methods can be used for different types of anti-counterfeiting features to achieve accurate detection. For example, as shown in Table 1, the first and third types of anti-counterfeiting features, such as the Great Wall logo (made with optically variable ink), optically variable ink, and laser patterns, can be called dynamic gradient features, which exhibit dynamic gradation or change effects with changes in light and / or angle. For instance, the color of the anti-counterfeiting feature may change between yellow, green, and orange with changes in angle, or the feature may gradually appear from nothing as the light intensity increases. Another example is the second type of anti-counterfeiting feature, such as embossed fonts, which are only visible under specific angles or lighting conditions and cannot be seen under other angles or lighting conditions; these can be called dynamic 0 / 1 features. Furthermore, specific identifiers inherent to the document, such as the document name and portrait, can be called static features.

[0037] For dynamic gradient features, the presence of a security feature in a document can be determined by detecting whether the feature undergoes a preset gradient in consecutive video frames during changes in acquisition conditions. Each dynamic gradient feature corresponds to a different preset gradient. One implementation method involves segmenting the target video in response to the security feature being a dynamic gradient feature, obtaining a segmented video corresponding to the security feature. Each frame of the segmented video contains an image corresponding to the region where the security feature is located. The segmented video is then input into a classifier to obtain the detection result of the security feature, which includes the confidence level that the security feature has undergone a preset gradient.

[0038] The target video may include identification documents, as well as other unrelated objects such as the user's hands and the desktop. The target video can be segmented, isolating the regions containing the dynamic gradient features to be detected in each frame. For example, [the following can be done:] Figure 1 The region containing the laser pattern is segmented to obtain a segmented video corresponding to the anti-counterfeiting feature. Each frame of this segmented video may contain only the image corresponding to the region containing the anti-counterfeiting feature. This embodiment does not limit the segmentation method used; for example, optical flow or spatiotemporal segmentation based on convolutional neural networks can be used.

[0039] Considering that the position of anti-counterfeiting features in documents is usually fixed, alignment can be used to achieve fast segmentation and reduce the amount of computation. Specifically, the target video can be aligned to obtain an aligned video, and then the aligned video can be segmented based on the position of the anti-counterfeiting features corresponding to the document type to obtain the segmented video corresponding to the anti-counterfeiting features.

[0040] The locations of most of the anti-counterfeiting features on each type of document are fixed, for example, Figure 3 A template for a reference document is shown, with dashed boxes marking the position of each anti-counterfeiting feature. The document in each frame is aligned to ensure that the spatial relationship between the document image in each frame and the reference document template is consistent, resulting in an aligned video. Then, based on the document classification results, the position corresponding to each anti-counterfeiting feature is determined, and the area containing that position is segmented to obtain a segmented video corresponding to each anti-counterfeiting feature. Alignment can be achieved through image processing techniques such as image registration, transformation, or feature matching; this embodiment does not impose any limitations on this.

[0041] This embodiment does not limit the classifier used. For example, networks such as MLP (Multilayer Perceptron) and LSTM (Long Short-Term Memory) can be used. The classifier can be pre-trained using sample videos containing preset gradients, enabling it to learn to judge whether the anti-counterfeiting feature has undergone a preset gradient. In this way, the classifier can analyze the feature change patterns by segmenting the temporal features corresponding to the anti-counterfeiting feature in the video to obtain the detection result of the anti-counterfeiting feature. The detection result includes the confidence level of the anti-counterfeiting feature undergoing a preset gradient. The higher the confidence level, the higher the probability that the current feature is a real change, and the higher the probability that the anti-counterfeiting feature exists in the document. A detection threshold can be set; when the confidence level is higher than the detection threshold, it is determined that the document contains the anti-counterfeiting feature. It should be noted that for segmented videos corresponding to multiple anti-counterfeiting features, they can be input into a single classifier for classification to obtain detection results for multiple anti-counterfeiting features, or they can be input into different classifiers for classification to obtain detection results for multiple anti-counterfeiting features separately.

[0042] For anti-counterfeiting features that are not dynamically changing, such as dynamic 0 / 1 features and static features, the presence of the feature in a particular frame within a series of video frames can be detected to determine whether the document contains that anti-counterfeiting feature. As one implementation, in response to the anti-counterfeiting feature not being dynamically changing, video frames from the target video are input into the target detection model to obtain the detection result of the anti-counterfeiting feature.

[0043] Specifically, the detection model can be configured to input individual video frames from the target video one by one to detect whether anti-spoofing features exist in the video frame. Input can stop once an anti-spoofing feature is detected. Alternatively, the target video can be input into the target detection model to detect whether anti-spoofing features exist. This embodiment does not limit the specific target detection model used. For example, it can use a target detection model based on YOLOv10 (You Only Look Once version 10) or R-CNN (Region Convolutional Neural Network) as the target detection model.

[0044] The detection process for each anti-counterfeiting feature described above can be carried out in parallel or sequentially. Once an anti-counterfeiting feature is detected, the document is considered to possess that anti-counterfeiting feature. Each anti-counterfeiting feature only needs to be detected once.

[0045] In one embodiment, such as Figure 4As shown, when acquiring target video in real time, in response to the detection result indicating the existence of undetected anti-counterfeiting features, at least one target operation is executed according to the detection conditions corresponding to the anti-counterfeiting features. The target video acquired by the device during the execution of the target operation can be acquired in real time, and step 206 is continued to detect the undetected anti-counterfeiting features.

[0046] For example, after the target video is identified, it is compared with the list of anti-counterfeiting features to be detected. If there are still undetected anti-counterfeiting features, the detection conditions corresponding to the anti-counterfeiting features can be prompted to the user, allowing the user to cooperate in collecting new target videos in the next time period. The above process can be repeated until the anti-counterfeiting feature is detected. If the collection exceeds a certain time and the anti-counterfeiting feature is still not detected, the timeout is reached and the collection ends.

[0047] It should be noted that when performing at least one target operation, the priority of these operations can vary. For example, if the detection conditions for the anti-counterfeiting feature are related to lighting and angle, the target operation that changes the angle can be prioritized, such as prompting the user to adjust the angle of the document or device. This is because lighting conditions generally change with the angle, thus reducing the steps required to adjust lighting conditions and improving the user experience. Similarly, sensors that can directly control the device can have a higher priority than those that prompt the user, reducing user actions and improving the user experience.

[0048] In step 208, the authenticity of the document is determined based on the detection results.

[0049] When a document contains multiple anti-counterfeiting features, its authenticity can be determined based on the detection results corresponding to each feature. Specifically, for example... Figure 4As shown, the system can determine if a document is genuine if the detection results of multiple anti-counterfeiting features indicate that the number of detected anti-counterfeiting features in the document has reached a preset number. This embodiment does not limit the setting of the preset number; for example, if a document may contain nine anti-counterfeiting features, the preset number can be set to nine. When all anti-counterfeiting features are detected, the document is considered genuine, and the data collection ends. Alternatively, considering that detecting all anti-counterfeiting features would be time-consuming and computationally resource-intensive, impacting user experience, the preset number could be set to seven. When more than seven anti-counterfeiting features are detected, the document is considered genuine, and the data collection ends. Alternatively, the system can determine if a document is not genuine if the detection results of multiple anti-counterfeiting features within a preset time do not reach the preset number. For example, the preset time could be set to one minute. If the detection results within one minute still do not detect the preset number, the timeout occurs, detection stops, and the user is informed of the detection failure, allowing them to choose whether to re-detect. If they do not continue detection, the document is considered counterfeit.

[0050] In the solutions provided by the above embodiments in this specification, by detecting the target video captured by the document during the process of changing acquisition conditions, it is possible to detect the dynamically changing anti-counterfeiting features in the real document. The various types of anti-counterfeiting features of the document are comprehensively detected using their corresponding detection methods. For the detection algorithm of dynamic gradient feature sampling spatiotemporal joint, the changes of anti-counterfeiting features are accurately detected by verifying the physical change law of anti-counterfeiting features on the time axis, making the detection results of the document more accurate and the security stronger.

[0051] In one embodiment, to improve the detection accuracy of dynamic gradient features under changing acquisition conditions, when segmenting the target video to obtain the segmented video corresponding to the anti-counterfeiting feature in step 204 of the above embodiment, the target video can be input into an image encoder to obtain an image feature sequence, and the identifier of the anti-counterfeiting feature can be input into a prompt word encoder to obtain a prompt word feature sequence. Then, based on the prompt word feature sequence and the image feature sequence, a mask decoder generates the segmented video corresponding to the anti-counterfeiting feature. The identifier of the anti-counterfeiting feature can be the name of the anti-counterfeiting feature, such as a laser pattern, a purple-gold flower, or optically variable ink.

[0052] This embodiment uses a promptable visual segmentation model to segment the target video. The specific network structure of the model is not limited in this embodiment; for example, the SAM2 model (Segment Anything Model 2, a visual segmentation model) can be used. SAM2 is a model based on the Transformer architecture, and its structure can be as follows... Figure 5As shown, when the target video is input into the image encoder, the image encoder can use a Vit (Vision Transformer) base model to obtain an image feature sequence, which can be an image token sequence. The anti-counterfeiting feature identifier is input into the prompt word encoder, and the resulting prompt word feature sequence can be a prompt word token sequence. To more accurately indicate the model's position, the anti-counterfeiting feature identifier and its position identifier can also be input into the prompt word encoder to obtain a prompt word feature sequence. The anti-counterfeiting feature's position identifier can be the position coordinates shown in Table 1. The attention mechanism module can perform attention operations on the image feature sequence using the memory of previous video frames stored in the memory bank. Then, it and the prompt word feature sequence are input into the mask decoder. The mask decoder generates a segmentation mask in the video frame corresponding to the anti-counterfeiting feature. The segmentation mask indicates the position of the anti-counterfeiting feature in each video frame, thus obtaining the segmented video corresponding to the anti-counterfeiting feature. The segmentation mask can also be input into the memory bank module for storage. The segmented video can then be input into the classifier to obtain the detection result of the anti-counterfeiting feature.

[0053] The feature extraction capability of an image encoder has a crucial impact on the accuracy of video segmentation. In one embodiment, self-supervised pre-training can be combined to further improve the feature extraction capability of the image encoder. The image encoder is pre-trained using a generative self-supervised approach that randomly masks some image patches in the video and then reconstructs them. The training method can be to first acquire a first sample video, which can be any type of real document. The first sample video is obtained based on the first sample document. During acquisition, the acquisition can be carried out while the acquisition conditions are changing. After masking the first sample video, it is input into the image encoder to obtain the output first feature sequence. The masking process is used to mask some image patches in the video frame of the first sample video. Then, the first feature sequence is input into the decoder to obtain the output reconstructed video. Finally, the network parameters of the image encoder are adjusted based on the difference between the first sample video and the reconstructed video.

[0054] During masking, several pixel blocks at spatially identical locations can be masked in each video frame of the first sample video. A tube masking strategy can be used here, where a "tube" refers to a pixel block at a spatially identical location. For these spatially identical pixel blocks, they are either completely masked or not masked in all video frames along the temporal dimension. This alleviates the information leakage problem caused by temporal correlation in the video and improves the training effect of the encoder. For example, more than 95% of the pixel blocks in the video sequence can be masked, and then the remaining few image blocks can be used to reconstruct the entire video sequence to obtain the reconstructed video. Then, based on the difference between the first sample video and the reconstructed video, a loss function is constructed, and the network parameters of the image encoder are adjusted with the goal of minimizing the loss function, making the reconstructed video increasingly closer to the first sample video. When the model reconstructs the content information of multiple frames of the document and some anti-counterfeiting features under lighting and angle conditions at a high masking rate, it indicates that the pre-trained image encoder has a certain ability to represent and extract the anti-counterfeiting information of the document. The pre-training framework here includes an image encoder and a decoder. After pre-training, the image encoder will have a strong ability to extract spatiotemporal features from videos.

[0055] After obtaining the pre-trained image encoder, the image encoder, mask decoder, and classifier can be fine-tuned end-to-end in the following way to make the model more adaptable to the detection task of dynamic gradient features: First, obtain the anti-counterfeiting feature labels corresponding to the second sample video and the second sample document. The second sample video is based on the second sample document collected under different conditions. The anti-counterfeiting feature labels are used to mark the first identifier of the first anti-counterfeiting feature in the second sample document and whether the first anti-counterfeiting feature has undergone a preset gradient. Then, input the second sample video into the image encoder to obtain the second feature sequence, and input the first identifier into the prompt word encoder to obtain the first prompt sequence. Next, based on the first prompt sequence and the second feature sequence, the mask decoder generates the first segmented video corresponding to the first anti-counterfeiting feature. Input the first segmented video into the classifier to obtain the first detection result of the first anti-counterfeiting feature. The first detection result includes the confidence level that the first anti-counterfeiting feature has undergone a preset gradient. Based on the difference between the first detection result and the anti-counterfeiting feature label, adjust the network parameters of the image encoder, mask decoder, and classifier.

[0056] The second sample document can include both genuine and counterfeit documents. The constructed second sample video includes genuine document videos and counterfeit document videos. The anti-counterfeiting features of the counterfeit document video will not change, while the genuine document, under certain acquisition conditions, will have dynamic gradient features and dynamic 0 / 1 features that change, but static features will not change. The anti-counterfeiting features in the second sample video can be specifically labeled to obtain anti-counterfeiting feature tags. The anti-counterfeiting feature tags can include the first identifier of the first anti-counterfeiting feature (which can be labeled as the name of the anti-counterfeiting feature) and whether the first anti-counterfeiting feature has undergone a preset gradient (which can be labeled as whether there is a change, with 1 representing a change and 0 representing no change). The anti-counterfeiting feature tags can also include the location of the anti-counterfeiting feature (which can be labeled using a mask). The first identifier and the location coordinates of the anti-counterfeiting feature are input together into the prompt word encoder to obtain the first prompt sequence. For example, the anti-counterfeiting feature tags for video data 1 can be: Video data 1, laser pattern, the mask corresponding to the laser pattern, no change; Video data 1, purple and gold flower, mask, changed; Video data 1, color-changing ink, mask, changed. In this example, the loss function can be calculated by comparing the difference between the 0 or 1 labels with or without changes in the anti-counterfeiting feature labels and the confidence level in the first detection result. Then, with the goal of minimizing the loss function, the model parameters of the image encoder, mask decoder, and classifier are adjusted to improve the prediction accuracy of the model through supervised learning.

[0057] This specification also provides a document detection device for performing the methods provided in this specification. Figure 6 A schematic block diagram of a document verification device according to one embodiment is shown. Figure 6 As shown, the device includes:

[0058] The video acquisition module 601 is used to acquire the target video, which is obtained by capturing video of the document during the process of changing acquisition conditions;

[0059] The feature determination module 602 is used to determine various anti-counterfeiting features to be detected in the document based on the type of the document;

[0060] The video detection module 603 is used to detect the target video based on the detection methods corresponding to each anti-counterfeiting feature in order to determine whether the document contains each anti-counterfeiting feature.

[0061] The authenticity determination module 604 is used to determine the authenticity of a document based on the detection results.

[0062] This specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed in a computer, causes the computer to perform the document detection method provided in the foregoing embodiments.

[0063] This specification also provides a computing device in the embodiments, including a memory and a processor. The memory stores computer programs / instructions, and when the processor executes the computer programs / instructions, it implements the document detection methods provided in the foregoing embodiments.

[0064] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0065] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0066] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0067] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.

[0068] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0069] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0072] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0073] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0074] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0075] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0076] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0077] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0078] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.

Claims

1. A method for detecting identification documents, the method comprising: The target video is obtained based on video capture of the document during the process of changing acquisition conditions; Based on the type of the document, identify multiple anti-counterfeiting features to be detected in the document; Based on the detection methods corresponding to each of the anti-counterfeiting features, the target video is detected to determine whether the document includes each of the anti-counterfeiting features. Based on the results of the detection, the authenticity of the document is determined.

2. The method according to claim 1, wherein, The detection method based on the anti-counterfeiting features detects the target video to obtain the detection result of the anti-counterfeiting features, including: In response to the fact that the anti-counterfeiting feature is a dynamic gradient feature, the target video is segmented to obtain a segmented video corresponding to the anti-counterfeiting feature. Each video frame of the segmented video contains an image corresponding to the region where the anti-counterfeiting feature is located. The segmented video is input into a classifier to obtain the detection result of the anti-counterfeiting feature, and the detection result includes the confidence level of the anti-counterfeiting feature undergoing a preset gradual change.

3. The method according to claim 2, wherein, The step of segmenting the target video to obtain the segmented video corresponding to the anti-counterfeiting feature includes: The target video is input into an image encoder to obtain an image feature sequence, and the identifier of the anti-counterfeiting feature is input into a prompt word encoder to obtain a prompt word feature sequence. Based on the prompt word feature sequence and the image feature sequence, the mask decoder generates the segmented video corresponding to the anti-counterfeiting feature.

4. The method according to claim 3, wherein, The image encoder was trained using the following method: Acquire a first sample video, which is obtained based on the collection of a first sample document; After the first sample video is masked, it is input into the image encoder to obtain the first feature sequence output. The masking process is used to mask some image blocks in the video frames of the first sample video. The first feature sequence is input into the decoder to obtain the output reconstructed video; The network parameters of the image encoder are adjusted based on the differences between the first sample video and the reconstructed video.

5. The method according to claim 3 or 4, wherein, The image encoder, the mask decoder, and the classifier are trained in the following manner: Obtain the anti-counterfeiting feature label corresponding to the second sample video and the second sample document. The second sample video is obtained based on the second sample document under different conditions. The anti-counterfeiting feature label is used to mark the first identifier of the first anti-counterfeiting feature in the second sample document and whether the first anti-counterfeiting feature undergoes a preset gradient. The second sample video is input into the image encoder to obtain the second feature sequence, and the first identifier is input into the prompt word encoder to obtain the first prompt sequence; Based on the first prompt sequence and the second feature sequence, the mask decoder generates a first segmented video corresponding to the first anti-counterfeiting feature; The first segmented video is input into the classifier to obtain the first detection result of the first anti-counterfeiting feature. The first detection result includes the confidence level of the first anti-counterfeiting feature undergoing a preset gradient. Based on the difference between the first detection result and the anti-counterfeiting feature label, the network parameters of the image encoder, the mask decoder, and the classifier are adjusted.

6. The method according to claim 2, wherein, The step of segmenting the target video to obtain the segmented video corresponding to the anti-counterfeiting feature includes: The target video is aligned to obtain an aligned video; Based on the location of the anti-counterfeiting feature corresponding to the type of the document, the aligned video is segmented to obtain the segmented video corresponding to the anti-counterfeiting feature.

7. The method according to claim 1, wherein, The detection method based on the anti-counterfeiting features detects the target video to obtain the detection result of the anti-counterfeiting features, including: In response to the fact that the anti-counterfeiting feature is not a dynamic gradient feature, the video frames in the target video are input into the target detection model to obtain the detection result of the anti-counterfeiting feature.

8. The method according to claim 1, wherein, The acquisition of the target video includes: Based on the type of the document, at least one target operation to be performed is determined: prompting the user to adjust the pose of the device capturing the target video; prompting the user to adjust the placement of the document; prompting the user to control the sensors of the device; controlling the sensors of the device; the sensors include a flash and / or a camera; Obtain the target video captured by the device during the execution of the target operation.

9. The method according to claim 1, wherein, The acquisition of the target video includes: In response to the detection result indicating the existence of an undetected anti-counterfeiting feature, at least one of the target operations is performed according to the detection conditions corresponding to the anti-counterfeiting feature; Obtain the target video captured by the device during the execution of the target operation.

10. The method according to claim 8 or 9, wherein, When performing at least one of the target operations, the priority of performing the target operations is different.

11. The method according to claim 1, wherein, The anti-counterfeiting features include at least one of the following: optical color-changing ink, laser pattern, embossing, watermark, hologram, iridescent effect element, colored ultraviolet pattern, and stereoscopic laser image. The acquisition conditions include at least one of the following: lighting conditions and angle conditions.

12. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Certificate identification method and device, storage medium and electronic equipment

    CN118097679A

  • Certificate anti-counterfeiting detection method and device

    CN119478788A