Certificate detection method, computing device, storage medium and computer program product
By using an image processing model in the document detection, the thickness analysis of document images from multiple perspectives is solved, and the problem of inaccurate document detection in the prior art is achieved, and higher security and reliability are achieved.
Patent Information
- Application Number
- CN202510224048.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
When using computers to process documents, the prior art is susceptible to problems such as falsehood and error, which leads to security risks. Especially in Internet finance and personal information identification scenarios, how to accurately detect documents has become an urgent problem.
A certificate detection method is provided, by determining target images from multiple different perspectives, using an image processing model to analyze the document thickness, determine the document thickness information, and perform document detection based on this information.
This method can accurately detect the certificate, avoid security risks caused by false and error of the certificate, and improve the accuracy and reliability of the certificate detection.
Smart Images

Figure CN120147727A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification relate to the technical field of risk detection, and particularly to a document detection method. One or more embodiments of this specification also relate to a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the continuous development of computer technology, in the process of using a computer to provide services for users, operations involving the processing of document information are often involved, such as the processing of documents in scenarios such as Internet finance and personal information identification.
[0003] Currently, in the process of using a computer to process documents, due to problems such as false or incorrect documents, there will be relatively large security risks; therefore, how to accurately detect documents has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a document detection method. One or more embodiments of this specification also relate to a document detection device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a document detection method is provided, including:
[0006] Determine a plurality of target images corresponding to a target document, where the plurality of target images are images of the target document collected from multiple different perspectives;
[0007] Use an image processing model to perform document thickness analysis on the target document based on the plurality of target images, determine the document thickness information of the target document, and determine the document detection result of the target document based on the document thickness information.
[0008] According to the second aspect of the embodiments of this specification, a document detection device is provided, including:
[0009] An image determination module configured to determine a plurality of target images corresponding to a target document, where the plurality of target images are images of the target document collected from multiple different perspectives;
[0010] A document detection module configured to use an image processing model to perform document thickness analysis on the target document based on the plurality of target images, determine the document thickness information of the target document, and determine the document detection result of the target document based on the document thickness information.
[0011] According to a third aspect of the embodiments of the present specification, a computing device is provided, including:
[0012] a memory and a processor;
[0013] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned document detection method are implemented.
[0014] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned document detection method are implemented.
[0015] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned document detection method are implemented.
[0016] One or more embodiments of the present specification provide a document detection method; during the process of document detection, target images of a target document collected from multiple different perspectives are obtained, and an image processing model is used to analyze the thickness of the document for the target images to determine the document thickness information of the target document. Based on this document thickness information, the target document can be accurately risk-detected, so as to obtain an accurate document detection result and avoid security risks caused by problems such as false or incorrect documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an application schematic diagram of a document detection method provided by an embodiment of the present specification;
[0018] Figure 2 is a flowchart of a document detection method provided by an embodiment of the present specification;
[0019] Figure 3 is a structural schematic diagram of a feature fusion model in a document detection method provided by an embodiment of the present specification;
[0020] Figure 4 is a process schematic diagram of feature fusion in a document detection method provided by an embodiment of the present specification;
[0021] Figure 5 is a process schematic diagram of the processing of the CA network in a document detection method provided by an embodiment of the present specification;
[0022] Figure 6 is a process flowchart of the processing of a document detection method provided by an embodiment of the present specification;
[0023] Figure 7 It is a schematic structural diagram of a document detection device provided by an embodiment of this specification;
[0024] Figure 8 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific embodiments
[0025] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0026] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0027] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0028] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0029] First, the noun terms involved in one or more embodiments of this specification are explained.
[0030] EKYC (Enhanced Know Your Customer): is an enhanced customer identification process. It usually combines electronic and automated technologies to verify personal identity, assess risks, and meet regulatory compliance requirements. EKYC can authenticate the identity of the user based on the user's personal documents; in this process, the authenticity of the user's documents needs to be guaranteed. Based on this, the document detection method provided in one or more embodiments of this specification is required to detect the target document.
[0031] AIGC (Artificial Intelligence Generated Content): AIGC is a technology that uses artificial intelligence to create multimedia content such as images and videos. In the financial field, users may use AIGC to generate false personal documents and use the false personal documents for personal identity verification in EKYC, thus posing serious security risks to financial services.
[0032] SDK (Software Development Kit): It can be called a software development kit. SDK is a set of tools and documents that allow developers to create applications for specific platforms or services. In the document detection method provided in this specification, a software program with the function of collecting document images can be developed as an SDK tool and deployed on the client; thereby allowing users to collect personal documents from multiple angles based on the client SDK, which facilitates subsequent document detection based on the multi-frame images collected from multiple angles.
[0033] ResNeSt (Residual Nesting Network) is a deep learning architecture; this method introduces ResNeSt in the encoder and decoder of the feature fusion model; the split attention block in the ResNeSt can be used to perform feature enhancement processing on image features of multiple frames of images; the ResNeSt allows the ResNeSt model to adaptively adjust the image feature map, focusing on key image channels and spatial positions to capture the complex interactions between different image features.
[0034] Feature fusion model: It is a model that can extract features from multiple frames of documents and fuse multiple image features through LSM fusion. The thickness information of the document can be obtained through this feature fusion model. Considering that in EKYC in financial scenarios, the thickness of real documents is 1 mm, while fake documents such as color-printed, AIGC, and photocopied documents are very thin and have no thickness. Therefore, using thickness information to assist anti-counterfeiting detection can effectively distinguish between real and fake documents.
[0035] LSM (Local Salient Module): It refers to the Local Salient Module. The LSM is used to extract multi-scale feature maps, and can perform local feature extraction on multiple frames of images and achieve efficient feature fusion.
[0036] MLP (Multilayer Perceptron): Also known as the multi-layer perceptron, the MLP is a feedforward artificial neural network model composed of multiple layers of neurons, including an input layer, one or more hidden layers, and an output layer. In this method, the MLP can classify the authenticity of a certificate based on image features, so as to identify whether the certificate is genuine or fake.
[0037] With the continuous development of computer technology, in the process of using a computer to provide services to users, operations involving the processing of certificate information are often involved. For example, in scenarios such as Internet finance and personal information identification, the processing of certificates is involved. Currently, in the process of using a computer to process certificates, due to problems such as false or incorrect certificates, there will be relatively large security risks. For example, in EKYC in the financial scenario, how to defend against highly realistic fake certificates is a challenge.
[0038] In response to the above problems, a solution provided in this specification is an anti-counterfeiting detection solution for single-frame pictures, which can detect whether a certificate is genuine or fake; however, this solution has relatively large defects. Since the recognition accuracy of the single-frame algorithm is relatively low, the accuracy rate for highly realistic fake certificates (such as fax attacks, etc.) is not high.
[0039] Based on this, in this specification, a certificate detection method is provided. One or more embodiments of this specification simultaneously relate to a certificate detection device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0040] See Figure 1 , Figure 1 shows an application schematic diagram of a certificate detection method provided according to an embodiment of this specification. Based on Figure 1It can be seen that the user can collect multiple frames of images of a document (i.e., the target document) from multiple perspectives through the client 102 and send the multiple frames of images of the document to the server 104 for anti-counterfeiting detection; after receiving the multiple frames of images of the document, the server 104 will input the multiple frames of images of the document into a multi-frame anti-counterfeiting framework for detection; the multi-frame anti-counterfeiting framework can perform thickness analysis on the multiple frames of images, thereby determining the thickness information of the document using the multiple frames of images and performing anti-counterfeiting detection on the document based on the thickness information, so as to accurately obtain the document detection result of whether the document is genuine / fake. After obtaining the document detection result, the server 104 can send the document detection result to the client 102, thereby presenting the document detection result to the user.
[0041] See Figure 2 , Figure 2 shows a flowchart of a document detection method provided according to an embodiment of the present specification, which specifically includes the following steps.
[0042] Step 202: Determine multiple target images corresponding to the target document, where the multiple target images are images of the target document collected from multiple different perspectives.
[0043] Among them, the target document can be a document existing in physical form, and the target document has a thickness; for example, the target document can be a document existing in the form of a card (such as an ID card, a bank card, a social security card, etc.), a document existing in the form of a booklet (such as a passport, a student ID card, etc.).
[0044] The target image can be understood as an image corresponding to the target document; the target image can be collected from multiple different perspectives. For example, the different perspectives can be a front perspective, multiple side perspectives.
[0045] In one or more embodiments provided in the present specification, the determining multiple target images corresponding to the target document includes:
[0046] Receiving the multiple target images corresponding to the target document sent by the client, where the multiple target images are collected by the image acquisition module of the client from multiple different perspectives;
[0047] Among them, the image acquisition module can be understood as a module deployed in the client for collecting target images. For example, the image acquisition module can be a hardware device such as a camera or a webcam of the client, or the image acquisition module can be a software device such as a client SDK or an application program.
[0048] Specifically, the document detection method provided in this specification can be applied to the server, and the client corresponding to the server can quickly capture the target image of the target document from multiple different perspectives through the image acquisition module under the user's operation; after obtaining multiple target images, the client will send the multiple target images to the server for document detection, thereby improving the efficiency of document detection.
[0049] Taking the application of the certificate detection method provided in this specification in a financial scenario as an example, the certificate detection method is explained, wherein the target certificate is a card-type certificate, and the image acquisition module can be an SDK; it should be noted that this method can use the "multi-angle certificate" or "collect active light" + "multi-angle certificate" method to collect certificate images (i.e., target images). Among them, the collected active light refers to: an artificial light source (i.e., active light) actively provided during the image acquisition process; for example, the light source emitted by the flash on a mobile phone (i.e., the client) is used to collect the certificate image when the light source is irradiated on the certificate; wherein, a multi-angle certificate refers to: during the image acquisition process, the certificate image is collected from multiple angles (i.e., viewing angles).
[0050] Based on this, this method takes into account that in EKYC in financial scenarios, the thickness of the real certificate (i.e., the target certificate) is 1 mm, while fake certificates such as color-printed, AIGC, and photocopied certificates are very thin and have no thickness; therefore, by identifying the thickness, the real certificate and the fake certificate can be effectively distinguished.
[0051] In order to effectively identify the thickness of the certificate, the client SDK needs to collect certificate images at multiple angles when collecting images, and use artificial light sources to collect certificates. The multiple angles include a front view and multiple side view angles (i.e., other view angles except the front view angle); the multiple side view angles can be four different angles of the certificate: the left view angle, the right view angle, the upper view angle, and the lower view angle. Therefore, the certificate image collected by this method can be a front frame image (i.e., a front view angle image), angle frame images at four different angles (i.e., at least two other view angle images); the multi-angle frame images at four different angles can be understood as oblique angle frames.
[0052] In addition, in order to accurately identify the detailed image features of the front frame image and ensure the comprehensiveness of the image features, the present method can also introduce the flash frame image corresponding to the front frame image for document detection; the flash frame image is a document image captured when the light source is irradiated on the document, and the acquisition angle of the flash frame image is the same as that of the front frame. Based on this, the document image captured by the present method can be a front frame image (i.e., a front perspective image), a flash frame image (i.e., a light source image), and angle frame images at four different angles (i.e., at least two other perspective images).
[0053] Based on the above steps, to obtain accurate document thickness information, this method uses an image acquisition module to capture target images of the target document from multiple different perspectives, so as to facilitate subsequent accurate analysis of the document thickness information based on multiple target images.
[0054] Step 204: Use an image processing model to analyze the document thickness of the target document based on the multiple target images, determine the document thickness information of the target document, and determine the document detection result of the target document based on the document thickness information.
[0055] Among them, this image processing model can be understood as a model capable of detecting the target document based on the target image. For example, this image processing model can be a deep learning model; or
[0056] This image processing model can be an image anti-counterfeiting detection framework, which can determine the document thickness information of the target document and determine the authenticity result (i.e., the document detection result) of the target document based on this document thickness information; for example, this image anti-counterfeiting detection framework can be a multi-frame anti-counterfeiting framework based on card thickness information, used for anti-counterfeiting detection of documents through multi-frame images (i.e., multiple target images).
[0057] The document thickness information can be information used to represent the thickness of the target document. For example, this document thickness information can be a document thickness value (such as 1 millimeter, 2 millimeters, etc.), a document thickness feature (such as an embedded vector feature, an image coding feature, an image feature matrix, etc.).
[0058] The document detection result can be understood as the detection result for the target document; this document detection result can be information characterizing the target document as true / false. For example, this document detection result can be a value (such as value 0 for false, value 1 for true), a probability (such as 95% for false, 5% for true), etc., used to characterize the target document as true / false. Or
[0059] This document detection result can be information characterizing the risk level of the target document. For example, this document detection result can be a value (such as value 1 for first-level risk, value 2 for second-level risk, value 3 for third-level risk, etc.), a label (such as a first-level risk label, a second-level risk label, a third-level risk label), etc., used to characterize the risk level of the target document.
[0060] Specifically, this method can input multiple target images into an image processing model. In the image processing model, the thickness of the target document is analyzed based on the multiple target images to obtain the document thickness information of the target document; it is determined whether the document thickness information is equal to a preset thickness threshold (for example, 1 mm), and the document detection result of the target document is determined based on the judgment result; the judgment result includes being equal to the preset thickness threshold or not being equal to the preset thickness threshold; the corresponding document detection results are genuine documents or fake documents respectively.
[0061] In one or more embodiments provided in this specification, the multiple target images include a front view image and at least two other view images;
[0062] Before using the image processing model to analyze the thickness of the target document based on the multiple target images to determine the document thickness information of the target document, it further includes:
[0063] Input the multiple target images into the image processing model, and use the image processing model to determine the target image positions in the front view image and the at least two other view images;
[0064] Perform position alignment processing on the target image positions in each of the other view images and the target image positions in the front view image to obtain the aligned front view image and the aligned at least two other view images.
[0065] Among them, the target image position can be the center point position of the target document in the front view image and the at least two other view images; or, the target image position can be the four corner point positions of the target document in the front view image and the at least two other view images.
[0066] Continuing with the above example, the image processing model is a multi-frame anti-counterfeiting framework based on card thickness information (hereinafter simply referred to as the multi-frame anti-counterfeiting framework), and the target image position is the four corner points of the target document; based on this, after the front frame image and the angle frame images at four different angles are collected through the client SDK in this method, the multi-frame images (that is, the front frame image, multiple angle frame images) can be input into the multi-frame anti-counterfeiting framework for preprocessing.
[0067] In the preprocessing part, the model input is multi-frame images, and the specific preprocessing method is as follows:
[0068] 1. Determine the front frame image from the multi-frame images, and use the front frame image as the central image to identify the four corner points of the document in the front frame image.
[0069] 2. Identify the 4 corner points of the document in the multiple angle frame images.
[0070] 3. Align the four corner points of the document in the multiple angular frame images with the four corner points of the document in the front frame image, so as to align the multiple angular frame images with the front frame image, and obtain the aligned front frame image (i.e., the aligned front - view image) and the aligned multiple - angular frame images (i.e., the aligned at least two other - view images).
[0071] Subsequently, the aligned front frame image and the aligned multiple - angular frame images can be used to analyze the document thickness of the target document, and the document detection result of the target document can be determined according to the document thickness information.
[0072] Based on the content of the above - mentioned embodiments, before using the image - processing model to analyze the document thickness of the target document based on the multiple target images and determine the document thickness information of the target document, the method can use the positions of the target images in the front - view image and the at least two other - view images to align the front - view image and the at least two other - view images, so as to obtain the aligned front - view image and the aligned at least two other - view images; thus, the multiple images are adjusted to a unified size, and the model input data with a unified format is obtained, which is convenient for the subsequent image - processing model to analyze the document thickness of the target document based on the multiple target images with a unified format, so as to accurately determine the document thickness information of the target document.
[0073] It should be noted that the operation method of "using the front - view image and at least two other - view images to determine the document thickness information of the target document" in this method belongs to the same technical concept as "using the light - source image, the front - view image and at least two other - view images to determine the document thickness information" in one or more embodiments of this specification; for the above - mentioned operation method of "using the front - view image and at least two other - view images to determine the document thickness information of the target document", the operation method of "using the light - source image, the front - view image and at least two other - view images to determine the document thickness information" in this specification can be referred to, and no more details will be given here.
[0074] In one or more embodiments provided in this specification, the multiple target images include a light - source image, a front - view image and at least two other - view images;
[0075] Before using the image - processing model to analyze the document thickness of the target document based on the multiple target images and determine the document thickness information of the target document, it further includes:
[0076] Input the multiple target images into the image - processing model, and use the image - processing model to determine the positions of the target images in the light - source image, the front - view image and the at least two other - view images;
[0077] Perform position alignment processing on the target image positions in the light source image and each other perspective image with the target image position in the front perspective image to obtain the aligned light source image, the aligned front perspective image, and at least two aligned other perspective images.
[0078] Among them, the target image position can be the center point position of the target certificate in the light source image, the front perspective image, and at least two other perspective images; or, the target image position can be the four corner point positions of the target certificate in the light source image, the front perspective image, and at least two other perspective images.
[0079] Continuing with the above example, the image processing model is a multi-frame anti-counterfeiting framework based on card thickness information (hereinafter simply referred to as the multi-frame anti-counterfeiting framework), and the target image position is the four corner points of the target certificate; based on this, after collecting the front frame image, the flash frame image, and the angle frame images at four different angles through the client SDK, this method can input the multi-frame images (i.e., the front frame image, the flash frame image, and multiple angle frame images) into the multi-frame anti-counterfeiting framework for preprocessing.
[0080] In the preprocessing part, the model input is multi-frame images, and the specific preprocessing method is as follows:
[0081] 1. Determine the front frame image from the multi-frame images, and use the front frame image as the central image to identify the four corner points of the certificate in the front frame image.
[0082] 2. Identify the four corner points of the certificate in the flash frame image and multiple angle frame images.
[0083] 3. Align the four corner points of the certificate in the flash frame image and multiple angle frame images with the four corner points of the certificate in the front frame image, so as to align the flash frame image, multiple angle frame images with the front frame image, and obtain the aligned front frame image (i.e., the aligned front perspective image), the aligned flash frame image (i.e., the aligned light source image), and the aligned multi-angle frame images (i.e., the aligned at least two other perspective images).
[0084] Based on the content of the above embodiments, before using the image processing model to analyze the document thickness of the target document based on the multiple target images and determine the document thickness information of the target document, the method can align the light source image, the front view image, and the target image positions in the at least two other view images, so as to obtain the aligned light source image, the aligned front view image, and the aligned at least two other view images; thereby realizing adjusting multiple images to a unified size and obtaining model input data with a unified format, which is convenient for the subsequent image processing model to analyze the document thickness of the target document based on the multiple target images with a unified format, so as to accurately determine the document thickness information of the target document.
[0085] In one or more embodiments provided in this specification, the document thickness information is a document thickness feature;
[0086] The method of using the image processing model to analyze the document thickness of the target document based on the multiple target images, determine the document thickness information of the target document, and determine the document detection result of the target document based on the document thickness information includes Step 1 to Step 2:
[0087] Step 1: Use the image processing model to perform document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document, where the document thickness feature is used to represent the thickness information of the target document;
[0088] Among them, the document thickness feature can be understood as a feature used to represent the thickness information of the target document. For example, the document thickness feature can be the encoding feature of the target image. The document image feature can be understood as a feature used to represent the overall image information of the target document.
[0089] Specifically, the method can input multiple target images into the image processing model, and perform document feature analysis on each target image in the image processing model to obtain the document thickness feature of the target document and the document image feature of the target document.
[0090] Specifically, the use of the image processing model to perform document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document includes:
[0091] Input the multiple target images into the image processing model, where the image processing model includes an image feature processing unit and a thickness analysis unit;
[0092] Using the thickness analysis unit, encode and process the multiple target images to obtain the target image encoding features corresponding to each target image, and perform feature fusion on the multiple target image encoding features to obtain the document thickness feature of the target document;
[0093] Using the image feature processing unit, extract features from the multiple target images to obtain the image features corresponding to each target image, and perform feature fusion on the multiple image features to obtain the document image feature of the target document.
[0094] Among them, the image feature processing unit can be understood as the unit in the image processing model for feature extraction and feature fusion of multiple target images; for example, this image feature processing unit can be a sub-model in the image processing model, or multiple network layers in the image processing model; for example, this image feature processing unit can be a sub-model composed of a CNN model and a feature fusion model in the image processing model.
[0095] The thickness analysis unit can be understood as the unit in the image processing model for encoding and feature fusion of target images. For example, this thickness analysis unit can be a sub-model in the image processing model, or multiple network layers in the image processing model; for example, this thickness analysis unit can be a sub-model composed of a transform model and a feature fusion model in the image processing model.
[0096] It should be noted that this image feature processing unit can be the multi-frame representation branch in the image processing model, and this multi-frame representation branch is used to extract the representation of multi-frame images (i.e., the document image feature); this thickness analysis unit can be the depth prediction branch in the image processing model, and this depth prediction branch is used to predict the document thickness based on the front frame image and multi-angle frame images to obtain the feature representing the thickness information of the target document (i.e., the document thickness feature). In addition, this depth prediction branch can also perform image segmentation and document detection on the target document based on the front frame image and multi-angle frame images, so as to achieve multi-task processing.
[0097] Continuing with the above example, the multi-frame anti-counterfeiting framework in this method includes a depth prediction branch; the input data of this depth prediction branch are the aligned front frame (i.e., the front frame image) and the aligned multi-angle frames (i.e., multi-angle images); after obtaining the front frame (aligned front frame) and multi-angle frames (aligned multi-angle frames), perform image encoding processing on the front frame and multi-angle frames to obtain the image encoding features of the front frame and multi-angle frames; then fuse the image encoding features of the front frame and multi-angle frames to obtain Feature B (i.e., the document thickness feature).
[0098] The multi-frame anti-counterfeiting framework in this method includes a multi-frame representation branch (i.e., an image feature processing unit); the input data of this multi-frame representation branch can be multiple frames of images such as the aligned front frame, the aligned multi-angle frames, etc.; after obtaining multiple frames of pictures, the image features of each image are extracted respectively; then the multiple image features are fused to obtain Feature A (i.e., the document image feature). Or, the input data of this multi-frame representation branch can be multiple frames of images such as the aligned front frame, the aligned flash frame (i.e., the flash frame image), the aligned multi-angle frames, etc.; after obtaining multiple frames of pictures, the image features of each image are extracted respectively; then the multiple image features are fused to obtain Feature A (i.e., the document image feature).
[0099] Based on the above embodiments, it can be seen that the image processing model in this method uses the image feature processing unit and the thickness analysis unit to process multiple target images specifically, thereby improving the efficiency of image processing; and, processes multiple target images from two aspects of image feature processing and thickness analysis to obtain rich image feature data, improving the accuracy of subsequent document detection.
[0100] In one or more embodiments provided in this specification, the multiple target images include a front-view image and at least two other-view images, where the front-view image is an image of the target document collected from the front view, and the other-view images are images of the target document collected from other views except the front view;
[0101] The use of the thickness analysis unit to perform encoding processing on the multiple target images, obtain the target image encoding features corresponding to each target image, and perform feature fusion on the multiple target image encoding features to obtain the document thickness feature of the target document, includes:
[0102] Using the feature extraction sub-unit in the thickness analysis unit to extract features from the front-view image and the at least two other-view images to obtain the front-view image feature of the front-view image and the other-view image features of each other-view image;
[0103] Using the encoding fusion sub-unit in the thickness analysis unit to perform encoding processing on the front-view image feature and the other-view image features to obtain the front-view image encoding feature and the other-view image encoding feature, and perform feature fusion on the front-view image encoding feature and the other-view image encoding feature to obtain the document thickness feature of the target document.
[0104] Among them, the feature extraction subunit can be understood as a subunit for extracting features from the front - view image and at least two other - view images. This feature extraction subunit can be a sub - model or network layer in the thickness analysis unit. For example, this feature extraction subunit can be a transform model.
[0105] The encoding and fusion subunit can be understood as a subunit for encoding and fusing the features of the front - view image and the features of at least two other - view images; this feature encoding subunit can be a sub - model or network layer in the thickness analysis unit. For example, this feature encoding subunit can be an encoder or an encoding network layer in the feature fusion model.
[0106] Continuing with the above example, the depth prediction branch in the multi - frame anti - counterfeiting framework of this method predicts the front - view frame and multi - angle frames to obtain Feature B. Specifically, the processing method of this depth prediction branch can be as follows:
[0107] 1. Input the front - view frame and multi - angle frames into the transform model, and perform feature extraction based on the transform model to obtain the image features of the front - view frame and multi - angle frames (i.e., the front - view image features and other - view image features);
[0108] 2. Use the encoder in the feature fusion model to perform encoding processing on the image features of the front - view frame and multi - angle frames to obtain the image encoding features of the front - view frame and multi - angle frames (i.e., the front - view image encoding features and other - view image encoding features);
[0109] 3. Use the LSM (i.e., the feature extraction model) in the feature fusion model to fuse the image encoding features of the front - view frame and multi - angle frames through the fusion method of the LSM to obtain Feature B.
[0110] Based on the content of the above - mentioned embodiments, the thickness analysis unit in this method processes multiple target images through the feature extraction subunit and the encoding and fusion subunit, thereby using information from multiple dimensions such as the front - view frame and the tilted - angle frame to assist in reconstructing and identifying the depth information of the card (i.e., the document thickness information), facilitating the subsequent effective distinction between fake and genuine documents and improving the accuracy of document detection.
[0111] In one or more embodiments provided in this specification, using the encoding and fusion subunit in the thickness analysis unit to perform encoding processing on the front - view image features and the other - view image features to obtain the front - view image encoding features and the other - view image encoding features, and performing feature fusion on the front - view image encoding features and the other - view image encoding features to obtain the document thickness features of the target document, includes:
[0112] Using the encoding fusion subunit in the thickness analysis unit, perform attention analysis on the frontal view image features and the at least two other view image features respectively, to obtain the first attention parameter of the frontal view image and the second attention parameters corresponding to the at least two other view images;
[0113] Based on the first attention parameter, perform encoding processing on the at least two other view image features to obtain the other view image encoding features of each other view image;
[0114] Based on the second attention parameter, perform encoding processing on the frontal view image features to obtain the frontal view image encoding features of the frontal view image;
[0115] Perform feature fusion on the frontal view image encoding features and the other view image encoding features to obtain the document thickness feature of the target document.
[0116] Continuing with the above example, the multi-frame anti-counterfeiting framework provided in this specification can use a feature fusion model to detect the card thickness, so as to obtain feature B representing the document thickness information; in addition, this feature fusion model also has driving factors (i.e., network layers) for segmentation and detection, so that through these driving factors for segmentation and detection, document thickness map segmentation and document detection can be realized.
[0117] Based on Figure 3 it can be seen that Figure 3 is a schematic structural diagram of a feature fusion model in a document detection method provided by an embodiment of this specification; the feature fusion model in this specification mainly includes three parts: an encoder, a feature fusion module, and a semantic segmentation and object detection network.
[0118] Among them, this encoder can be understood as an encoder (Base ResNeSt Encoder, abbreviated as BRE) that is a residual network with split attention (ResNeSt) and adopts an autoencoder (AE) module; this encoder can be composed of two parts: a residual network with split attention (ResNeSt) and a feature refinement block based on a convolutional layer.
[0119] The split attention block in this ResNeSt can be used to enhance feature representation and can perform enhancement processing on the image features of the frontal frame and multi-angle frames; this ResNeSt allows the model to adaptively adjust the feature map, focus on key image channels and spatial positions, so as to capture the complex interactions between different features. Therefore, by using ResNeSt as part of the encoder, its ability to promote complex cross-feature interactions between the image features of the frontal frame and the multi-angle frames is enhanced.
[0120] Based on this, the two encoders in this method can utilize the split attention blocks in ResNeSt to perform attention processing on the image features of the frontal frame and the image features of the multi-angle frame respectively, obtaining the attention parameters of the frontal frame (i.e., the first attention parameters) and the attention parameters of the multi-angle frame (i.e., the second attention parameters);
[0121] Then, in order to achieve the ability of complex cross-feature interaction, parameter sharing can be carried out between the two encoders; one encoder can utilize the attention parameters of the frontal frame to perform feature enhancement processing on the image features of the multi-angle frame, obtaining the enhanced image coding features (i.e., the image coding features of other perspectives). Another encoder can utilize the attention parameters of the multi-angle frame to perform feature enhancement processing on the image features of the frontal frame, obtaining the enhanced image coding features (i.e., the image coding features of the frontal perspective).
[0122] The feature refinement block in this encoder can further process the enhanced image coding features to improve the quality of the features. For example, the feature refinement block can perform noise suppression or useful information enhancement on the enhanced image coding features.
[0123] In summary, the encoder in this method realizes the extraction of interactive features, thereby performing feature extraction on the features, and the network based on AE shows good reconstruction performance.
[0124] Among them, the feature fusion module is a feature fusion module based on the local significant feature extraction model (LSM); this feature fusion module has a parallel dual-branch LSM for extracting important local information during the cross-modal feature extraction process. This LSM consists of a neighborhood attention Transformer (NAT) based on a sliding window and a detail saliency module (DSM).
[0125] The neighborhood attention (NA) of this NAT can localize the attention range of each pixel to its nearest neighbors, converging to self-attention as the range increases, maintaining translational invariance, and reducing the complexity and localization problems of the self-attention mechanism in visual tasks.
[0126] The purpose of the DSM is to improve the representation of the structural attributes of the feature fusion model. The DSM can process the output of the NAT through a convolutional layer; for the output of this convolutional layer, the DSM provides a feature extraction sub-module formed by combining two different feature extraction branches. Among them, one branch simultaneously performs average pooling and max pooling to synchronize detail extraction with global information. After dual pooling, global average pooling is used to further expand the features. Then, the weights of different channels are calculated by assigning two fully connected layers and a sigmoid layer (an activation layer), which enhances the importance of feature description. Summing the output of one branch with the output of the other branch can extract more significant information from the image-coded features of the enhanced multi-angle frames, and also extract structural information from the image-coded features of the enhanced frontal frames.
[0127] After processing the image-coded features of the frontal frame and the multi-angle frame using the parallel dual-branch LSM in this feature fusion module, the processed image-coded features can also be feature-fused; by inputting the processed image-coded features into the LSM for feature fusion, the fused image features (i.e., Feature B) can be obtained.
[0128] In addition, based on Figure 3 it is known that the feature fusion model also includes a decoder; the decoder can include: a feature refinement block based on a convolutional layer of a residual network with split attention (ResNeSt), and a spatial attention module (SA) based on meta-learning; the spatial attention module can be used to enhance the generalization and transfer learning capabilities of the model and promote the forward transmission of cross-modal information across various datasets.
[0129] Specifically, the decoding method of this decoder is as follows: First, the features extracted by the feature fusion model are used as the input of the decoder. Two spatial attention network layers in the spatial attention module perform spatial attention processing on the image-coded features of the frontal frame and the multi-angle frame respectively, and then perform weighted processing on the processed image features of the frontal frame and the multi-angle frame to obtain and output the weighted image features of the frontal frame and the multi-angle frame.
[0130] Second, use ResNeSt and the feature refinement block to enhance the image features of the weighted frontal frame and the multi-angle frame, and use the enhanced image features of the frontal frame and the multi-angle frame as the model output.
[0131] Among them, the semantic segmentation network is used to perform document segmentation based on the model output to obtain a segmentation map of the thickness information of the card document; the object detection network is used to perform document detection based on the model output to obtain the document detection results of multiple frames of images. This method introduces a semantic segmentation network and an object detection network into the feature fusion model, realizes the refinement processing of the fusion features (i.e., the model output), and improves the overall fusion performance. Moreover, according to the thickness information of the document image, a thick map (i.e., a segmentation map of the thickness information of the card document) can be automatically generated according to the genuine and fake document labels. Among them, the thickness of the genuine document is about 1 mm, and the thickness of the fake document is 0 mm.
[0132] In one or more embodiments provided in this specification, the multiple target images include a light source image, a front view image, and at least two other view images. Among them, the light source image is an image collected under the condition of irradiating the target document with an artificial light source;
[0133] Using the image feature processing unit to extract features from the multiple target images to obtain image features corresponding to each target image, and performing feature fusion on the multiple image features to obtain the document image features of the target document, including:
[0134] Using the feature extraction subunit in the image feature processing unit to extract features from the light source image, the front view image, and the at least two other view images respectively to obtain the image features corresponding to the light source image, the front view image, and the at least two other view images;
[0135] Using the feature fusion subunit in the image feature processing unit to perform feature fusion on the multiple image features to obtain the document image features of the target document.
[0136] Among them, the feature extraction subunit can be understood as a subunit used to extract features from the light source image, the front view image, and at least two other view images. This feature extraction subunit can be a sub-model or network layer in the image feature processing unit; for example, this feature extraction subunit can be a CNN model.
[0137] This feature fusion subunit can be understood as a subunit used to perform feature fusion on multiple image features; this feature fusion subunit can be a sub-model or network layer in the image feature processing unit; for example, the feature fusion subunit can be a feature fusion network.
[0138] Continuing with the above example, the multi-frame anti-counterfeiting framework in this method includes a multi-frame representation branch (i.e., the image feature processing unit); the input data of this multi-frame representation branch is: the front frame (i.e., the front frame image) and the aligned multi-angle frames (i.e., the multi-angle images), or the input data of this multi-frame representation branch is multiple frames of images such as the front frame (i.e., the front frame image), the aligned post-flash frame (i.e., the flash frame image), and the aligned multi-angle frames (i.e., the multi-angle images).
[0139] After obtaining multiple frames of pictures, use the CNN model to extract the image features of each image (i.e., the image features corresponding to the light source image, the front view image, and the at least two other view images); then the feature fusion network (CA network) uses the CA fusion method to fuse multiple image features to obtain FeatureA (i.e., the document image feature); specifically, the way the multi-frame representation branch performs feature processing is as follows:
[0140] 1. Use the CNN model to extract features from multiple frames of pictures such as the aligned front frame, the aligned post-flash frame, and the aligned multi-angle frames respectively, to obtain the image feature of the front frame (CNN branch 1), the image feature of the flash frame (CNN branch 2), and the image feature of the multi-angle frame (CNN branch 3).
[0141] Alternatively, this method can also use the CNN model to extract features from multiple frames of pictures such as the aligned front frame and the aligned multi-angle frames respectively, to obtain the image feature of the front frame (CNN branch 1) and the image feature of the multi-angle frame (CNN branch 3).
[0142] It should be noted that the method of "fusing CNN branch 1 and CNN branch 3 to obtain feature A" in this method is the same as the method of "fusing CNN branch 1, CNN branch 2, and CNN branch 3 to obtain feature A" described below. For the operation method of "fusing CNN branch 1 and CNN branch 3", the operation steps of "fusing CNN branch 1, CNN branch 2, and CNN branch 3" described below can be referred to, and will not be elaborated here.
[0143] It should be noted that the CNN model in this method uses the method of sharing parameters during the process of feature extraction.
[0144] 2. Perform fusion using the card_CA fusion method to fuse CNN branch 1, CNN branch 2, and CNN branch 3 to obtain feature A. The card_CA fusion method can be referred toFigure 4 ; Figure 4 It is a schematic diagram of the feature fusion process in a document detection method provided by an embodiment of this specification.
[0145] Based on Figure 4 it can be seen that, first, for the three image features of CNN branch 1, CNN branch 2, and CNN branch 3, the CA network is respectively used to enhance the features to obtain the enhanced image features.
[0146] Secondly, the enhanced image features are concatenated to obtain the concatenated image features;
[0147] Finally, the concatenated image features are input into the fusion network layer for feature fusion to obtain Feature A.
[0148] It should be noted that the processing process of the CA network can be referred to Figure 5 , Figure 5 which is a schematic diagram of the processing process of the CA network in a document detection method provided by an embodiment of this specification. Among them, Figure 5 f(c×h×w) in
[0149] Based on Figure 5 it can be seen that the processing process of the CA network can be:
[0150] 1. Three 3×3 convolution operations are respectively performed on the image feature f(c×h×w) to obtain three convolution image features obtained by the three convolution operations (i.e., Figure 5 P(f), Q(f), H(f) in
[0151] 2. Reshape operations are performed on P(f) and Q(f) to obtain the adjusted convolution image features (i.e., Figure 5 M in P , M Q ).
[0152] 3. M P and M Q are multiplied element-wise to obtain the image feature M;
[0153] 4. Normalize the image feature M (softmax), and perform element-wise multiplication on the normalized image feature and H(f) to obtain the image feature f h (c×h×w);
[0154] 5. Take f h (c×h×w) and f h (c×h×w) and perform element-wise addition to obtain the enhanced image feature f CA (c×h×w) output by the CA network.
[0155] Based on the content disclosed in the above embodiments, it can be seen that in the document feature representation branch of this method, the CA fusion algorithm is used to decouple multiple image features (CNN branch), improving the accuracy of representation.
[0156] Step 2: Based on the document thickness feature and the document image feature, determine the document detection result of the target document.
[0157] Specifically, the document detection result is the document risk type, and the image processing model is the document detection model;
[0158] Determining the document detection result of the target document based on the document thickness feature and the document image feature includes:
[0159] Using the classification prediction unit in the document detection model, perform document risk classification processing on the target document based on the document thickness feature and the document image feature to obtain the document risk type corresponding to the target document.
[0160] Among them, the document risk type can be understood as the type information indicating whether there is a risk for the target document, and this document risk type can be the real document type or the fake document type. The document detection model can be understood as a model used to detect the risk or authenticity of the target document, and this document detection model can be a document anti-counterfeiting model, a multi-frame anti-counterfeiting framework based on thickness information.
[0161] Among them, this classification prediction unit can be understood as a unit that performs risk classification on the target document based on the document thickness feature and the document image feature; this classification prediction unit is a sub-model or network layer in the image processing model; for example, this classification prediction unit can be an MLP model.
[0162] Continuing with the above example, after the above multi-frame representation branch and depth prediction branch are executed to obtain Feature A and Feature B, Feature A and Feature B can be used to perform multi-branch fusion processing in the multi-frame anti-counterfeiting framework, thereby detecting the certificate; the specific processing method is as follows:
[0163] 1. Concatenate (concat) Feature A extracted using multi-frame multi-modalities and Feature B obtained using the supervision information of card thickness detection to obtain a concatenated feature;
[0164] 2. Input the concatenated feature into the MLP model, and use the MLP to classify and identify the fused concatenated feature to accurately obtain the type information of genuine and fake certificates (i.e., the certificate risk type), thereby realizing multi-frame multi-modal anti-counterfeiting detection and accurately predicting whether the current certificate is genuine or fake.
[0165] Based on the content of the above embodiments, it can be seen that the certificate anti-counterfeiting algorithm framework (i.e., the multi-frame anti-counterfeiting framework) of the present method adopts the thickness information prediction branch and the certificate feature representation branch to be fused together for joint anti-counterfeiting detection. Moreover, the thickness of the certificate is obtained based on the card multi-modal thickness segmentation prediction branch, thereby assisting anti-counterfeiting detection and improving the accuracy of prediction.
[0166] In one or more embodiments provided in this specification, after using the image processing model to perform certificate feature analysis on the multiple target images to obtain the certificate thickness feature of the target certificate and the certificate image feature of the target certificate, it further includes:
[0167] Use the feature decoding unit in the image processing model to perform decoding processing on the certificate thickness feature to obtain the target image feature of the target image;
[0168] Based on the target image feature, perform a target task for the target certificate to obtain a task execution result corresponding to the target certificate.
[0169] Among them, the feature decoding unit can be understood as a unit for performing decoding processing on the certificate thickness feature. The feature decoding unit can be a sub-model or network layer in the image processing model. For example, the feature decoding unit can be the decoder in the above embodiments.
[0170] The target task can be understood as a certificate processing task that needs to be performed on the target certificate. For example, the target task can be a certificate segmentation task and a certificate detection task.
[0171] Among them, the target image feature can be understood as a feature for accurately representing the certificate information in the multiple target images including the target certificate.
[0172] Continuing with the above example, based on Figure 3It can be seen that after the fused features (i.e., document thickness features) are extracted using the feature fusion model, the fused features are input into the decoder for decoding to obtain the model output (i.e., target image features).
[0173] Then, based on the model output, a document segmentation task and / or a document detection task is performed on the document, thereby obtaining a segmentation map of the card document thickness information and the document detection result of the picture.
[0174] Based on the content of the above embodiments, it can be known that the present method performs a target task on the target document through the target image features, thereby solving the requirements of document segmentation and document detection in actual scenarios.
[0175] In one or more embodiments provided in this specification, the target task includes a document segmentation task and / or a document detection task;
[0176] Performing the target task on the target document based on the target image features and obtaining the task execution result corresponding to the target document includes:
[0177] Using the document segmentation unit in the image processing model to perform the document segmentation task on the target document based on the target image features to obtain a document thickness segmentation image corresponding to the target document; and / or
[0178] Using the document detection unit in the image processing model to perform the document detection task on the target document based on the target image features to obtain a document detection result corresponding to the target document.
[0179] Among them, the document segmentation task can be understood as the task of segmenting a document thickness segmentation image corresponding to the target document from multiple target images; this document thickness segmentation image can be understood as the segmentation map of the card document thickness information in the above embodiments.
[0180] The document segmentation unit can be understood as a semantic segmentation model for segmenting and obtaining a document thickness segmentation image.
[0181] This document detection task can be understood as the task of detecting the document type of the target document and the document information contained in the target document; this document detection result can be the document type, document information, etc. of the target document.
[0182] This document detection unit can be a target detection model for detecting the target document.
[0183] Continuing with the above example, in this method, a semantic segmentation network and an object detection network are introduced into the feature fusion model. The semantic segmentation network is used to segment the certificate based on the model output to obtain a segmentation map of the thickness information of the card certificate. The object detection network is used to detect the certificate based on the model output to obtain the certificate detection results of multiple frames of images. By introducing the semantic segmentation network and the object detection network into the feature fusion model, the refinement processing of the fusion features is realized, and the overall fusion performance is improved.
[0184] In one or more embodiments provided in this specification, after determining the certificate detection result of the target certificate based on the certificate thickness information, the following is further included:
[0185] Sending the certificate detection result to the client.
[0186] Continuing with the above example, this method can send the detection results of whether the target certificate is a genuine certificate or a forged certificate to the client, and then display it to the user through the client, meeting the requirement for accurate detection of the target certificate.
[0187] One or more embodiments of this specification provide a certificate detection method. During the process of certificate detection, target images of the target certificate collected from multiple different perspectives are obtained, and an image processing model is used to analyze the certificate thickness of the target images to determine the certificate thickness information of the target certificate. Based on this certificate thickness information, the risk of the target certificate can be accurately detected, thereby obtaining accurate certificate detection results and avoiding security risks caused by problems such as false or incorrect certificates.
[0188] The following combines the attached Figure 6 , taking the application of the certificate detection method provided in this specification in the certificate anti-counterfeiting scenario as an example, to further illustrate the certificate detection method. Among them, Figure 6 FIG. shows the processing flowchart of a certificate detection method provided in an embodiment of this specification, which specifically includes the following steps.
[0189] Step 602: Image preprocessing.
[0190] Specifically, the processing method of this image preprocessing is as follows:
[0191] 1. Use the method of "collecting active light" + "multi-angle certificate" to collect certificate images.
[0192] Among them, the collected active light refers to the artificial light source (i.e., active light) actively provided during the image collection process. For example, the light source emitted by the flashlight on the mobile phone (i.e., the client) is used to collect the certificate image when the light source shines on the certificate.
[0193] Among them, the multi-angle certificate refers to: during the image acquisition process, the certificate images are acquired from multiple angles (i.e., perspectives).
[0194] Based on this, in the certificate detection method provided in this specification, a software program with the function of acquiring certificate images can be developed into an SDK tool and deployed on the client; thus enabling users to perform multi-angle acquisition of personal certificates based on the client SDK; when the client SDK acquires images, it can acquire certificate pictures from multiple angles, and can acquire certificates from the front perspective under the condition of irradiating the certificate with an artificial light source (flashlight).
[0195] Among them, the multiple angles include the front perspective and multiple side perspectives (i.e., other perspectives except the front perspective); the multiple side perspectives can be 4 different angles: the left side perspective, the right side perspective, the upper side perspective, and the lower side perspective of the certificate.
[0196] Therefore, the certificate images collected by this method can be front frame images, flash frame images, and angle frame images at 4 different angles.
[0197] 2. Use the model to perform multi-frame alignment on multiple frames of images
[0198] In the preprocessing part, the input of the multi-frame anti-counterfeiting framework is multiple frames of images; the method of multi-frame image alignment is to perform alignment operations based on the front frame; the steps of image alignment in the multi-frame anti-counterfeiting framework are as follows:
[0199] 1. Determine the front frame image from multiple frames of images, and use the front frame image as the central image (i.e., the reference image) to identify the four corner points of the certificate in the front frame image.
[0200] 2. Identify the 4 corner points of the certificate in the flash frame image and multiple angle frame images.
[0201] 3. Align the 4 corner points of the certificate in the flash frame image and multiple angle frame images with the 4 corner points of the certificate in the front frame image, so as to align the flash frame image, multiple angle frame images with the front frame image, and obtain the aligned front frame image, aligned flash frame image, and aligned multi-angle frame image.
[0202] Step 604: Multi-frame representation branch.
[0203] The multi-frame anti-counterfeiting framework includes a multi-frame representation branch; the input data of the multi-frame representation branch are multiple frames of images such as the aligned front frame, aligned flash frame, and aligned multi-angle frame; after obtaining multiple frames of images, the image features of each image are extracted respectively; then the multiple image features are fused to obtain Feature A.
[0204] Specifically, the execution steps of the multi-frame representation branch are as follows:
[0205] 1. Use the CNN model to extract features from multiple frames of images such as the aligned frontal frame, the aligned flash frame, and the aligned multi-angle frames, respectively, to obtain the image features of the frontal frame (CNN branch 1), the image features of the flash frame (CNN branch 2), and the image features of the multi-angle frames (CNN branch 3).
[0206] Specifically, first, input multiple frames of images such as an aligned frontal frame, an aligned flash frame, and aligned multi-angle frames into the CNN model.
[0207] Second, in the CNN model, use multiple network layers such as convolutional layers, pooling layers, and fully connected layers to capture the spatial hierarchical structure and image content in each frame of the image, and obtain the image features corresponding to each frame of the image.
[0208] It should be noted that in the process of feature extraction by the CNN model in this method, feature extraction is performed in a way of sharing parameters.
[0209] The shared parameters refer to: for the convolutional operations at all positions in each frame of the image, the same convolutional kernel weights and bias terms are used. In the CNN model, each convolutional layer has one or more convolutional kernels (also called filters). These convolutional kernels slide over the entire input image (i.e., the "convolution" operation) to generate feature maps. Performing feature extraction in a way of sharing parameters means that the convolutional kernels use the same convolutional kernel weights and bias terms for the convolutional operations at all positions in each frame of the image, so as to effectively capture the spatial hierarchical structure of the image.
[0210] 2. Perform fusion using the CA fusion method to fuse one CNN branch 1, one CNN branch 2, and multiple CNN branch 3 to obtain Feature A.
[0211] Among them, the image feature (CNN branch) is a feature matrix with length, width, and height.
[0212] For the CA fusion method, reference can be made to Figure 4 ; Figure 4 It is a schematic diagram of the feature fusion process in a document detection method provided by an embodiment of this specification.
[0213] Based on Figure 4 it can be known that first, for the three types of image features of CNN branch 1, CNN branch 2, and CNN branch 3, they are respectively input into the corresponding CA network for feature enhancement to obtain multiple enhanced image features.
[0214] It should be noted that the CA networks corresponding to the three types of image features are the same CA network, and each CA network can perform the same feature enhancement processing on the three types of image features.
[0215] Secondly, splice multiple enhanced image features to obtain spliced image features;
[0216] The method of splicing multiple enhanced image features can adopt the method of feature concatenation; the feature concatenation means: attaching multiple feature matrices together to form a single, longer feature vector; for example, multiple enhanced image features are the feature-enhanced image features corresponding to a front frame, a flash frame, and 4 angle frames, and each enhanced image feature is a feature matrix with a size of 9×9×9; based on this, the method of feature concatenation for these 6 enhanced image features is:
[0217] Stack 6 feature matrices with a size of 9×9×9 together to obtain a feature matrix with a size of 54×9×9 (i.e., the spliced image feature).
[0218] Finally, input the spliced image features into the fusion network layer for feature fusion to obtain Feature A.
[0219] The fusion network layer is used to perform feature fusion processing on multiple enhanced image features, and the fusion network layer can select a suitable network layer according to the actual application scenario. For example, the fusion network layer can be: a convolutional network layer.
[0220] In addition, it should be noted that the processing process of the CA network can be referred to Figure 5 , Figure 5 is a schematic diagram of the processing process of the CA network in a document detection method provided by an embodiment of this specification. Among them, Figure 5 f(c×h×w) in is CNNbranch 1, CNN branch 2 or CNN branch 3; (c×h×w) is the size of the image feature (CNN branch); Conv3×3 is a 3×3 convolution operation; among them, reshape means to re-adjust the shape of the image data to the required dimension without changing its data content.
[0221] Based on Figure 5 it can be known that the processing process of the CA network can be:
[0222] 1. Perform three 3×3 convolution operations (conv3×3) on the image feature f(c×h×w) to obtain three convolution image features obtained by the three convolution operations (i.e., Figure 5 P(f), Q(f), H(f) in).
[0223] Among them, the input data of the three 3×3 convolution operations are different, and the parameters during the convolution operation are also different; for example, Figure 5 P and Q in
[0224] respectively represent the width and height of the image feature, while H is used to represent the number of channels (depth) of the image feature. Based on this, the three 3×3 convolution operations can respectively perform convolution processing on the image feature in a targeted manner from three aspects: width, height, and depth; P(f) input after passing through the convolution layer refers to the width of the image feature, Q(f) refers to the height of the image feature, and H(f) refers to the depth (number of channels) of the image feature. Figure 5 in P 、M Q )。
[0225] 3. Multiply M P and M Q element-wise to obtain the image feature M;
[0226] 4. Normalize the image feature M (softmax), and multiply the normalized image feature with H(f) element-wise to obtain the image feature f h (c×h×w);
[0227] 5. Add f h (c×h×w) and f h (c×h×w) element-wise to obtain the enhanced image feature f CA (c×h×w) output by the CA network.
[0228] Step 606: Depth prediction branch.
[0229] The multi-frame anti-counterfeiting framework includes a depth prediction branch; the input data of the depth prediction branch are the aligned front frame and 4 aligned multi-angle frames; after obtaining the front frame and aligned multi-angle frames, perform image encoding processing on the front frame and multi-angle frames to obtain the image encoding features of the front frame and multi-angle frames; then perform LSM fusion on the image encoding features of the front frame and multi-angle frames to obtain feature B.
[0230] Specifically, the processing method of the depth prediction branch can be:
[0231] 1. Input the frontal frame and multi - angle frames into the Transformer model, perform feature extraction based on the Transformer model, and obtain the image features of the frontal frame (i.e., the Transformer result 1 in Figure 6 ) and the image features of the multi - angle frames (i.e., the Transformer result 2 in Figure 6 );
[0232] Among them, the processing method of the Transformer model for the frontal frame and multi - angle frames is as follows:
[0233] First, divide the frontal frame and multi - angle frames into multiple non - overlapping image patches (patches) of a fixed size, such as regions of 16×16 or 32×32 pixels. And each image patch is regarded as a single "token", and these high - dimensional vectors are mapped into a lower - dimensional space through a linear transformation (usually a fully - connected layer) to form patch embeddings of the image patches.
[0234] Second, since the Transformer does not have the ability to understand the order of elements in a sequence, it is necessary to add position encoding to the patch embeddings of each image patch to retain the spatial relationship between the image patches.
[0235] Using the multi - head self - attention mechanism, the patch embeddings of each image patch are interacted with the patch embeddings of all other image patches, calculate the similarity scores between them, and perform weighted summation on the patch embeddings of the image patches based on these scores to obtain new embeddings. Through this multi - head self - attention mechanism, the Transformer model focuses on information from different representational sub - spaces.
[0236] Finally, the new embeddings corresponding to each image patch in the frontal frame are used as the image features of the frontal frame; the new embeddings corresponding to each image patch in each angle frame are used as the image features of each angle frame.
[0237] 2. Use the encoder in the feature fusion model to perform encoding processing on the image features of the frontal frame and multi - angle frames to obtain the image encoding features of the frontal frame and multi - angle frames (i.e., Feature A);
[0238] Among them, this encoder can be understood as an encoder (Base ResNeSt Encoder, abbreviated as BRE) that is a residual network with split attention (ResNeSt) and adopts an auto - encoder (AE) module; this encoder can be composed of two parts: a residual network with split attention (ResNeSt) and a feature refinement block based on convolutional layers.
[0239] The split attention block in the ResNeSt can be used to enhance feature representation and perform enhancement processing on the image features of frontal frames and multi-angle frames. The ResNeSt allows the model to adaptively adjust the feature map, focusing on key image channels and spatial positions to capture complex interactions between different features. Therefore, by adopting ResNeSt as part of the encoder, its ability to promote complex cross-feature interactions between the image features of frontal frames and multi-angle frames can be enhanced.
[0240] Based on this, the two encoders in this method can utilize the split attention blocks in the ResNeSt to perform attention processing on the image features of frontal frames and multi-angle frames respectively, obtaining the attention parameters of frontal frames and the attention parameters of multi-angle frames.
[0241] Then, in order to achieve the ability of complex cross-feature interaction, parameter sharing can be carried out between the two encoders. One encoder can use the attention parameters of frontal frames to perform feature enhancement processing on the image features of multi-angle frames, obtaining enhanced image encoded features. The other encoder can use the attention parameters of multi-angle frames to perform feature enhancement processing on the image features of frontal frames, obtaining enhanced image encoded features (i.e., Feature A).
[0242] The feature refinement block in the encoder can further process the enhanced image encoded features to improve the quality of the features. For example, the feature refinement block can perform noise suppression or useful information enhancement on the enhanced image encoded features.
[0243] In summary, the encoder in this method realizes interactive feature extraction, thereby performing feature extraction on the features, and the network based on AE shows good reconstruction performance.
[0244] 3. Use the feature fusion module in the feature fusion model to fuse the image encoded features of frontal frames and multi-angle frames through the fusion method of LSM to obtain Feature B.
[0245] The feature fusion module in the feature fusion model is a feature fusion module based on the Local Significant Feature Extraction Model (LSM). The feature fusion module has a parallel dual-branch LSM for extracting important local information during cross-modal feature extraction. The LSM consists of a Neighborhood Attention Transformer (NAT) based on a sliding window and a Detail Salience Module (DSM).
[0246] The neighborhood attention (NA) of this NAT can localize the attention range of each pixel to its nearest neighbors, converge to self-attention as the range increases, maintain translational invariance, and reduce the complexity and localization problems of the self-attention mechanism in visual tasks.
[0247] The purpose of the DSM is to improve the representation of the structural attributes of the feature fusion model. The DSM can process the output of the NAT through a convolutional layer; for the output of this convolutional layer, the DSM provides a feature extraction sub-module formed by combining two different feature extraction branches. Among them, one branch performs average pooling and max pooling simultaneously to synchronize detail extraction with global information. After double pooling, global average pooling is used to further expand the features. Then, the weights of different channels are calculated by assigning two fully connected layers and a sigmoid layer, which enhances the importance of feature description. Summing the output of one branch with the output of the other branch can extract more significant information from the image-coded features of the enhanced multi-angle frames, and can also extract structural information from the image-coded features of the enhanced front frames.
[0248] After processing the image-coded features of the front frame and the multi-angle frame using the parallel dual-branch LSM of this feature fusion module, the processed image-coded features can also be feature-fused; by inputting the processed image-coded features into the LSM for feature fusion, the fused image features (i.e., Feature B) are obtained.
[0249] Step 608: Multi-branch fusion.
[0250] After completing the execution of the above multi-frame representation branch and depth prediction branch to obtain Feature A and Feature B, Feature A and Feature B can be used to perform multi-branch fusion processing in the multi-frame anti-counterfeiting framework to detect the certificate; the specific processing method is as follows:
[0251] 1. Feature concatenation is performed on Feature A extracted using multi-frame multi-modal and Feature B obtained using the supervision information of card thickness detection to obtain concatenated features;
[0252] 2. The concatenated features are input into the MLP classification model, and the MLP classification model is used to classify and identify the fused concatenated features to accurately obtain the type information of genuine and fake certificates, thereby realizing multi-frame multi-modal anti-counterfeiting detection and accurately predicting whether the current certificate is genuine or fake.
[0253] Among them, the method of using the MLP classification model to classify and identify the fused concatenated features is as follows:
[0254] First, the MLP classification model includes an input layer, multiple hidden layers, and an output layer. The input layer is used to receive concatenated features and input the concatenated features into the subsequent hidden layers;
[0255] Second, the hidden layer can perform non-linear transformation on the concatenated features through an activation function, so as to extract higher-level abstract features. By performing transformation processing on the concatenated features layer by layer through multiple hidden layers, the MLP classification model can learn the complex patterns and relationships inside the data;
[0256] Finally, based on the features obtained by processing all previous network layers, generate the prediction result of whether the certificate is genuine / fake and output it.
[0257] In addition, it should be noted that the feature fusion model also includes a decoder; the decoding method of the decoder is as follows:
[0258] First, the features extracted by the feature fusion module are used as the input of the decoder. The two spatial attention network layers in the spatial attention module perform spatial attention processing on the image encoding features of the front frame and the image encoding features of the multi-angle frame respectively, and then perform weighted processing on the processed image features of the front frame and the image features of the multi-angle frame to obtain and output the weighted image features of the front frame and the image features of the multi-angle frame.
[0259] Second, use ResNeSt and the feature refinement block to enhance the image features of the front frame and the image features of the multi-angle frame after weighting, and use the enhanced image features of the front frame and the image features of the multi-angle frame as the model output.
[0260] After obtaining the model output, multi-task processing can be performed based on the model output. Among them, the multi-task can include an image segmentation task and an object detection task.
[0261] Specifically, in this method, a semantic segmentation network and an object detection network are introduced into the feature fusion model. The semantic segmentation network is used to perform certificate segmentation based on the model output to obtain a segmentation map of the thickness information of the card certificate; the object detection network is used to perform certificate detection based on the model output to obtain the certificate detection results of multiple frames of images;
[0262] It should be noted that during the process of performing multiple tasks using the multi-frame anti-counterfeiting framework, it is necessary to use training samples and sample labels to train the feature fusion model; the training samples can be multi-frame images as samples; taking the scenario of training the feature fusion model for image segmentation tasks as an example, the feature fusion model performs image segmentation based on the training samples to obtain a predicted thickness map; then the loss function is calculated using the sample label (i.e., the actual thickness map) and the predicted thickness map; the parameters of the feature fusion model are adjusted using this loss function until the model training stop condition is reached, and the trained feature fusion model is obtained.
[0263] Based on the above embodiments, the document detection method in one or more embodiments of this specification provides a multi-frame document anti-counterfeiting detection algorithm based on thickness information; this algorithm is a color printing detection algorithm based on thickness information. Through multi-angle acquisition, multi-frame angle information is captured, and these pictures are comprehensively reconstructed to obtain the thickness of the card as auxiliary information. Then these pictures are combined with deep learning and attention mechanism to accurately determine whether the current document is a genuine or fake document information; improving the security level of EKYC.
[0264] Moreover, the multi-frame multi-modal document anti-counterfeiting algorithm framework provided in this method adopts the fusion of the thickness information prediction branch and the document feature representation branch to jointly perform anti-counterfeiting detection; through the card multi-modal thickness segmentation prediction branch, the thickness of the document is obtained to assist in anti-counterfeiting detection, so as to accurately perform anti-counterfeiting detection.
[0265] Corresponding to the above method embodiments, this specification also provides embodiments of a document detection device. Figure 7 The structural schematic diagram of a document detection device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0266] An image determination module 702, configured to determine a plurality of target images corresponding to a target document, where the plurality of target images are images of the target document collected from multiple different perspectives;
[0267] A document detection module 704, configured to use an image processing model to perform document thickness analysis on the target document based on the plurality of target images, determine the document thickness information of the target document, and determine the document detection result of the target document based on the document thickness information.
[0268] Optionally, the document thickness information is document thickness features;
[0269] The document detection module 704 is further configured to:
[0270] Using the image processing model, perform document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document, where the document thickness feature is used to represent the thickness information of the target document;
[0271] Based on the document thickness feature and the document image feature, determine the document detection result of the target document.
[0272] Optionally, the document detection module 704 is further configured to:
[0273] Input the multiple target images into the image processing model, where the image processing model includes an image feature processing unit and a thickness analysis unit;
[0274] Use the thickness analysis unit to perform encoding processing on the multiple target images to obtain target image encoding features corresponding to the target images, and perform feature fusion on the multiple target image encoding features to obtain the document thickness feature of the target document;
[0275] Use the image feature processing unit to perform feature extraction on the multiple target images to obtain image features corresponding to the target images, and perform feature fusion on the multiple image features to obtain the document image feature of the target document.
[0276] Optionally, the multiple target images include a front view image and at least two other view images, where the front view image is an image of the target document collected from the front view, and the other view images are images of the target document collected from other views other than the front view;
[0277] The document detection module 704 is further configured to:
[0278] Use the feature extraction sub-unit in the thickness analysis unit to perform feature extraction on the front view image and the at least two other view images to obtain the front view image feature of the front view image and the other view image features of the other view images;
[0279] Use the encoding fusion sub-unit in the thickness analysis unit to perform encoding processing on the front view image feature and the other view image features to obtain a front view image encoding feature and other view image encoding features, and perform feature fusion on the front view image encoding feature and the other view image encoding features to obtain the document thickness feature of the target document.
[0280] Optionally, the document detection module 704 is further configured to:
[0281] Using the encoding fusion subunit in the thickness analysis unit, perform attention analysis on the frontal view image features and the at least two other view image features respectively, to obtain the first attention parameter of the frontal view image and the second attention parameters corresponding to the at least two other view images;
[0282] Based on the first attention parameter, perform encoding processing on the at least two other view image features to obtain the other view image encoding features of each other view image;
[0283] Based on the second attention parameter, perform encoding processing on the frontal view image features to obtain the frontal view image encoding features of the frontal view image;
[0284] Perform feature fusion on the frontal view image encoding features and the other view image encoding features to obtain the document thickness feature of the target document.
[0285] Optionally, the multiple target images include a light source image, a frontal view image, and at least two other view images, where the light source image is an image collected when the target document is irradiated with an artificial light source;
[0286] The document detection module 704 is further configured to:
[0287] Using the feature extraction subunit in the image feature processing unit, perform feature extraction on the light source image, the frontal view image, and the at least two other view images respectively, to obtain the image features corresponding to the light source image, the frontal view image, and the at least two other view images;
[0288] Using the feature fusion subunit in the image feature processing unit, perform feature fusion on the multiple image features to obtain the document image feature of the target document.
[0289] Optionally, the document detection result is a document risk type, and the image processing model is a document detection model;
[0290] The document detection module 704 is further configured to:
[0291] Using the classification prediction unit in the document detection model, perform document risk classification processing on the target document based on the document thickness feature and the document image feature to obtain the document risk type corresponding to the target document.
[0292] Optionally, the document detection device further includes a task execution module, configured to:
[0293] Decode the document thickness feature by using the feature decoding unit in the image processing model to obtain the target image feature of the target image;
[0294] Based on the target image feature, execute the target task for the target document to obtain the task execution result corresponding to the target document.
[0295] Optionally, the target task includes a document segmentation task and / or a document detection task;
[0296] The task execution module is further configured to:
[0297] Use the document segmentation unit in the image processing model to execute the document segmentation task for the target document based on the target image feature to obtain the document thickness segmentation image corresponding to the target document; and / or
[0298] Use the document detection unit in the image processing model to execute the document detection task for the target document based on the target image feature to obtain the document detection result corresponding to the target document.
[0299] Optionally, the multiple target images include a light source image, a front view image, and at least two other view images;
[0300] The document detection device further includes an image processing module, which is configured to:
[0301] Input the multiple target images into the image processing model, and use the image processing model to determine the target image positions in the light source image, the front view image, and the at least two other view images;
[0302] Perform position alignment processing on the target image positions in the light source image and each other view image with the target image position in the front view image to obtain the aligned light source image, the aligned front view image, and the aligned at least two other view images.
[0303] Optionally, the image determination module 702 is further configured to:
[0304] Receive the multiple target images corresponding to the target document sent by the client, where the multiple target images are collected by the image acquisition module of the client from multiple different perspectives;
[0305] The document detection device further includes a result sending module, which is configured to:
[0306] Send the document detection result to the client.
[0307] One or more embodiments of this specification provide a document detection device; during the process of document detection, target images of a target document collected from multiple different perspectives are obtained, and an image processing model is used to analyze the thickness of the document for the target images to determine the document thickness information of the target document. Based on this document thickness information, the target document can be accurately risk-detected, so as to obtain an accurate document detection result and avoid security risks caused by problems such as false or incorrect documents.
[0308] The above is a schematic solution of a document detection device according to this embodiment. It should be noted that the technical solution of this document detection device and the technical solution of the above document detection method belong to the same concept. For the details not described in detail in the technical solution of the document detection device, reference can be made to the description of the technical solution of the above document detection method.
[0309] Figure 8 A structural block diagram of a computing device 800 according to an embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 through a bus 830, and a database 850 is used to store data.
[0310] The computing device 800 further includes an access device 840, and the access device 840 enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interfaces (for example, a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0311] In one embodiment of the present specification, the above components of the computing device 800, as well as Figure 8 other components not shown in Figure 8 the computing device structure block diagram shown are also connected to each other, for example, through a bus. It should be understood that
[0312] the computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smart phones), wearable computing devices (e.g., smart watches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.
[0313] The processor 820 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned document detection method are implemented.
[0314] Each embodiment in the present specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the document detection method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the document detection method embodiment.
[0315] One embodiment of the present specification further provides a computer-readable storage medium, which stores computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned document detection method are implemented.
[0316] Each embodiment in the present specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the document detection method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the document detection method embodiment.
[0317] One embodiment of the present specification further provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned document detection method are implemented.
[0318] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above document detection method belong to the same concept. For the details not described in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above document detection method.
[0319] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0320] The computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0321] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0322] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0323] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is only limited by the claims and their full scope and equivalents.
Claims
1. A document detection method, comprising: Determine a plurality of target images corresponding to the target document, wherein the plurality of target images are images of the target document captured at a plurality of different viewing angles; An image processing model is used to perform document thickness analysis on the target document based on the multiple target images to determine document thickness information of the target document, and a document detection result of the target document is determined based on the document thickness information.
2. The document detection method according to claim 1, wherein the document thickness information is a document thickness feature; The using of the image processing model to perform document thickness analysis on the target document based on the multiple target images, determining document thickness information of the target document, and determining a document detection result of the target document based on the document thickness information, includes: Using the image processing model, performing document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document, wherein the document thickness feature is used to represent the thickness information of the target document; Based on the document thickness feature and the document image feature, a document detection result of the target document is determined.
3. The document detection method according to claim 2, wherein the using the image processing model to perform document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document comprises: Inputting the plurality of target images into the image processing model, wherein the image processing model comprises an image feature processing unit and a thickness analysis unit; Utilizing the thickness analysis unit, encoding the plurality of target images to obtain target image encoding features corresponding to each target image, and fusing the plurality of target image encoding features to obtain the document thickness feature of the target document; The image feature processing unit is used to extract features from the multiple target images to obtain image features corresponding to each target image, and multiple image features are fused to obtain document image features of the target document.
4. The document detection method according to claim 3, wherein the plurality of target images include a front view image and at least two other view images, wherein: The front view image is an image of the target document captured at a front view, and the other view images are images of the target document captured at other view angles except the front view; The method of using the thickness analysis unit to encode the plurality of target images to obtain target image encoding features corresponding to each target image, and fusing the plurality of target image encoding features to obtain the document thickness feature of the target document includes: Using the feature extraction subunit in the thickness analysis unit, extract features from the front-view image and the at least two other-view images to obtain a front-view image feature of the front-view image and other-view image features of each of the other-view images; The coding fusion subunit in the thickness analysis unit is used to encode the front perspective image features and the other perspective image features to obtain the front perspective image coding features and the other perspective image coding features, and the front perspective image coding features and the other perspective image coding features are feature fused to obtain the document thickness features of the target document.
5. The document detection method according to claim 4, wherein the encoding fusion subunit in the thickness analysis unit is used to encode the front view image features and the other view image features to obtain the front view image encoding features and the other view image encoding features, and the front view image encoding features and the other view image encoding features are feature fused to obtain the document thickness features of the target document, including: Using the encoding fusion subunit in the thickness analysis unit, respectively perform attention analysis on the front view image features and the at least two other view image features to obtain a first attention parameter of the front view image and a second attention parameter corresponding to the at least two other view images; Based on the first attention parameter, encoding the at least two other-view image features to obtain other-view image encoding features of each other-view image; Based on the second attention parameter, encoding the frontal perspective image feature to obtain the frontal perspective image encoding feature of the frontal perspective image; The front view image coding features and the other view image coding features are fused to obtain the document thickness features of the target document.
6. The document detection method according to claim 3, wherein the plurality of target images include a light source image, a front view image and at least two other view images, wherein: The light source image is an image acquired when the target document is illuminated by an artificial light source; The method of using the image feature processing unit to extract features from the multiple target images to obtain image features corresponding to each target image, and performing feature fusion on the multiple image features to obtain the document image features of the target document includes: Using the feature extraction subunit in the image feature processing unit, respectively extracting features from the light source image, the frontal perspective image, and the at least two other perspective images to obtain image features corresponding to the light source image, the frontal perspective image, and the at least two other perspective images; The feature fusion subunit in the image feature processing unit is used to perform feature fusion on the multiple image features to obtain the document image features of the target document.
7. The document detection method according to claim 2, wherein the document detection result is a document risk type, and the image processing model is a document detection model; The determining the document detection result of the target document based on the document thickness feature and the document image feature includes: The classification prediction unit in the document detection model is utilized to perform document risk classification processing on the target document based on the document thickness feature and the document image feature to obtain the document risk type corresponding to the target document.
8. The document detection method according to claim 2, after using the image processing model to perform document feature analysis on the multiple target images to obtain the document thickness feature of the target document and the document image feature of the target document, further comprising: Using the feature decoding unit in the image processing model, the thickness feature of the document is decoded to obtain the target image feature of the target image; Based on the target image features, a target task for the target certificate is executed to obtain a task execution result corresponding to the target certificate.
9. The document detection method according to claim 8, wherein the target task comprises a document segmentation task and / or a document detection task; The step of executing the target task for the target certificate based on the target image feature and obtaining the task execution result corresponding to the target certificate includes: Utilizing the document segmentation unit in the image processing model, based on the target image features, the document segmentation task for the target document is performed to obtain a document thickness segmentation image corresponding to the target document; and / or The document detection unit in the image processing model is utilized to perform the document detection task for the target document based on the target image features, and obtain a document detection result corresponding to the target document.
10. The document detection method according to any one of claims 1 to 8, wherein the multiple target images include a light source image, a front view image, and at least two other view images; Before the method utilizes the image processing model to perform document thickness analysis on the target document based on the multiple target images to determine the document thickness information of the target document, the method further includes: Inputting the multiple target images into the image processing model, and using the image processing model to determine the positions of the target images in the light source image, the front view image, and the at least two other view images; The target image positions in the light source image and each other viewing angle image are aligned with the target image position in the front viewing angle image to obtain an aligned light source image, an aligned front viewing angle image and at least two aligned other viewing angle images.
11. The document detection method according to any one of claims 1 to 8, wherein the step of determining a plurality of target images corresponding to the target document comprises: Receiving the plurality of target images corresponding to the target certificate sent by the client, wherein the plurality of target images are acquired at a plurality of different viewing angles by an image acquisition module of the client; After determining the document detection result of the target document based on the document thickness information, the method further includes: The certificate detection result is sent to the client.
12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 11 are implemented.
13. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Cited By
Identity document detection method, computing device, storage medium, and computer program product
WO2026179205A1