Method and apparatus for detecting forged receipt, storage medium, and electronic device
By locating key information areas in the credential image and performing text anomaly detection, the problems of low efficiency and unexplainable results in counterfeit credential detection are solved, and efficient and accurate counterfeit credential identification is achieved, protecting the rights and interests of merchants and consumers.
Patent Information
- Application Number
- PCT/CN2025/081126
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-06
- Publication Date
- 2025-09-25
AI Technical Summary
Existing technologies are unable to effectively detect counterfeit credentials, resulting in unfair treatment of merchants and platforms and damage to consumer rights. Traditional methods are inefficient and the results are not interpretable.
By locating the key information areas of the credential image, text anomaly detection is performed, including text alignment and font consistency detection. The neural network model and text recognition technology are used to determine whether the credential image is a forged credential.
The accuracy and efficiency of counterfeit voucher detection are improved, false detection interference is reduced, the results are interpretable, and the rights and interests of merchants and consumers are protected.
Smart Images

Figure CN2025081126_25092025_PF_FP_ABST
Abstract
Description
Method, device, storage medium and electronic device for detecting forged credentials
[0001] This application claims priority to the Chinese patent application filed on March 19, 2024, with application number 202410317746.9 and invention name “A method, device, storage medium and electronic device for detecting forged certificates”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a method, device, storage medium, and electronic device for detecting forged credentials. Background Art
[0003] With the growth of the black and gray industry and the increasing prevalence of "wool gangs," some criminals are using forged credentials to deliberately deceive merchants or platforms in an attempt to illegally gain financial benefits. This practice severely undermines the fairness and integrity of the e-commerce industry, causing significant distress to merchants, platforms, and consumers. If merchants or platforms are forced to pay large amounts of unnecessary compensation due to misjudgment, this can cause significant financial losses to their operations. Furthermore, merchants or platforms are treated unfairly, not only facing the need to pay compensation but also potentially losing their reputation and customers. This behavior also severely infringes on consumer rights and fosters distrust in merchants or platforms. Therefore, effectively detecting forged credentials has become a pressing technical challenge. Summary of the Invention
[0004] The present application provides a method, device, storage medium and electronic device for detecting forged credentials, which are used to effectively detect forged credentials.
[0005] In a first aspect, a method for detecting forged credentials is provided, the method comprising:
[0006] Get the credential image;
[0007] Locating key information areas of the credential image;
[0008] Performing text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on the result of the text anomaly detection.
[0009] According to an achievable method in an embodiment of the present application, locating the key information area of the credential image includes:
[0010] The credential image is detected using a region detection model to obtain key information regions of the credential image, wherein the region detection model is pre-trained based on a neural network model.
[0011] According to an achievable method in an embodiment of the present application, locating the key information area of the credential image includes:
[0012] The voucher image is recognized by using a text recognition technology to obtain the text content in the voucher image, and the area where the text content is located is used as the key information area.
[0013] According to an achievable method in an embodiment of the present application, performing text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on the result of the text anomaly detection includes:
[0014] Masking the key information area in the voucher image;
[0015] Reconstructing the masked key information area using the credential image obtained after masking and the text content recognized from the key information area;
[0016] Determine the difference between the reconstructed key information region and the key information region before masking;
[0017] If the difference satisfies a preset first forgery condition, the credential image is determined to be a forged credential.
[0018] According to an achievable method in an embodiment of the present application, performing text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on the result of the text anomaly detection includes:
[0019] Performing text alignment detection and / or font consistency detection on the key information area;
[0020] If the text misalignment or inconsistent fonts in the key information area meet a preset second forgery condition, the certificate image is determined to be a forged certificate.
[0021] According to an achievable method in an embodiment of the present application, the text alignment detection includes at least one of the following:
[0022] Using the center position of the reference area of the credential image and the left and right edges of the text area in the key information area, detecting whether the text in the key information area is centered;
[0023] Using the same side edge of the reference area of the credential image and the text area in the key information area, detecting whether the text area in the key information area is aligned on the same side, where the same side is the left side or the right side;
[0024] Using the text line spacing in the reference area of the credential image and the text line spacing in the key information area, detecting whether the text line spacing in the key information area is aligned;
[0025] Utilizing an inline alignment detection model to detect whether the text in the key information area is inline aligned, the inline alignment detection model being pre-trained based on a neural network model;
[0026] The reference area is obtained by detecting the credential image using a region detection model.
[0027] According to an achievable method in an embodiment of the present application, the font consistency detection includes at least one of the following:
[0028] Detecting whether the fonts of the characters in the key information area are consistent using a font detection model, wherein the font detection model is pre-trained based on a neural network model;
[0029] Performing font recognition on the text in the key information area using a font recognition tool, and determining whether the fonts of the text in the key information area are consistent based on the recognition result;
[0030] Features of each character in the key information area are extracted, clustering is performed based on the features of each character, and whether the fonts of the characters in the key information area are consistent is determined according to the clustering result.
[0031] In a second aspect, a method for detecting forged credentials is provided, which is executed by a cloud server and includes:
[0032] Obtain the credential image sent by the user's device;
[0033] Locating key information areas of the credential image;
[0034] Performing a text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on a result of the text anomaly detection;
[0035] If it is determined that the credential image is a forged credential, the service requested by the user device based on the credential image is denied.
[0036] According to an achievable embodiment of the present application, the method further includes:
[0037] A response rejecting the service is returned to the user device, the response including information indicating that the credential image is a forged credential.
[0038] In a third aspect, a device for detecting forged credentials is provided, the device comprising:
[0039] an image acquisition unit configured to acquire a credential image;
[0040] an area positioning unit, configured to locate a key information area of the credential image;
[0041] The anomaly detection unit is configured to perform text anomaly detection on the key information area and determine whether the certificate image is a forged certificate based on the result of the text anomaly detection.
[0042] In a fourth aspect, a device for detecting forged credentials is provided, the device comprising:
[0043] an image acquisition unit, configured to acquire a credential image sent by a user device;
[0044] an area positioning unit, configured to locate a key information area of the credential image;
[0045] an anomaly detection unit configured to perform text anomaly detection on the key information area and determine whether the credential image is a forged credential based on a result of the text anomaly detection;
[0046] The service processing unit is configured to deny the service requested by the user equipment based on the credential image if it is determined that the credential image is a forged credential.
[0047] In a fifth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first or second aspects above.
[0048] In a sixth aspect, the present application provides an electronic device, including:
[0049] one or more processors; and
[0050] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first aspect or the second aspect.
[0051] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0052] 1) In view of the fact that fraudsters usually edit or replace the text in the voucher image when forging vouchers, the present application determines whether the voucher image is a forged voucher by performing text anomaly detection on the key information area, thereby effectively realizing forged voucher detection, reducing false detection interference in non-key information areas, and improving detection accuracy and efficiency compared to manual recognition methods.
[0053] 2) The present application detects whether a credential image is a forged credential through font alignment detection and / or font consistency detection, which on the one hand improves the effect of forged credential detection, and on the other hand makes the result of determining whether a credential image is a forged credential better interpretable.
[0054] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0056] FIG1 is a diagram of a system architecture applicable to an embodiment of the present application;
[0057] FIG2 is a flow chart of a method for detecting forged credentials provided in an embodiment of the present application;
[0058] FIG3 is a flowchart of the work flow of the text anomaly detection module provided in an embodiment of the present application;
[0059] FIG4 a is a schematic diagram of a voucher image (payment interface) provided in an embodiment of the present application;
[0060] FIG4 b is a schematic diagram of a voucher image (payment result interface) provided in an embodiment of the present application;
[0061] FIG5 is a schematic diagram of a voucher image (mailing and express delivery interface) provided in an embodiment of the present application;
[0062] FIG6 is a flow chart of another method for detecting forged credentials provided in an embodiment of the present application;
[0063] FIG7 is a schematic structural diagram of a device for detecting forged credentials provided in an embodiment of the present application;
[0064] FIG8 is a schematic structural diagram of another device for detecting forged credentials provided in an embodiment of the present application;
[0065] FIG9 is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0067] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0068] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0069] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0070] Where the solutions described in this specification and in the examples involve the processing of personal information, such processing will be conducted with a legitimate basis (e.g., with the consent of the personal information subject or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of these basic functions.
[0071] Currently, the phenomenon of criminals using forged credentials to deliberately deceive merchants, platforms, or consumers in an attempt to illegally obtain economic benefits is becoming increasingly serious, causing great trouble to merchants, platforms, or consumers. Traditionally, there are two main methods for detecting forged credentials:
[0072] The first method is through manual review. This method relies heavily on human ability and experience. On the one hand, it is not very effective, and on the other hand, it is inefficient and has high labor costs.
[0073] The second method uses a trained neural network model to classify credential images and directly determine whether they are forged. However, in real-world service scenarios, this method suffers from poor recognition performance when directly inputting credential images into the classification model due to the diversity of credential screenshot styles and the lack of textures found in traditional photos. Furthermore, the output is not interpretable, making it difficult for users to trust it.
[0074] In view of this, the present application provides a method for detecting forged credentials. To facilitate understanding of the present application, the system architecture on which the present application is based is first described. FIG1 shows an exemplary system architecture 100 to which embodiments of the present application can be applied. As shown in FIG1 , the system architecture 100 may include: a user device and a forged credential detection device located on a server side.
[0075] The user can upload the credential image through the user device, and the user device sends the credential image to the forged credential detection device on the server side.
[0076] User devices may include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, and personal computers (PCs). Smart mobile devices may include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected cars. Smart home devices may include smart TVs and smart refrigerators. Wearable devices may include smart watches, smart glasses, virtual reality devices, augmented reality devices, and mixed reality devices (i.e., devices that support both virtual reality and augmented reality).
[0077] The forged credential detection device can use the method provided in the embodiments of the present application to detect the credential image and obtain a detection result of whether it is a forged credential. Among them, the detection process of the forged credential detection device involves the use of regional detection models, etc., and may also involve other models.
[0078] The forged credential detection device can be deployed on a standalone server, within a server cluster, or even on a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a host product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosts and virtual private servers (VPS). In addition to the architecture shown in Figure 1, the forged credential detection device can also be deployed on a computer terminal with significant computing power.
[0079] As one possible implementation, a user can upload a credential image via a user device to a server-side forged credential detection device. The forged credential detection device then detects the credential image and determines whether it is forged. Based on this detection result, a determination can be made as to whether the user device should be provided with a service requested based on the credential image, and / or the detection result can be returned to the user terminal. The architecture shown in Figure 1 illustrates this implementation.
[0080] As another possible implementation, the forged credential detection device can be installed in the user terminal, for example, in the form of an application, a plug-in within the application, or a software development kit. When the user enters a credential image through the user terminal, the forged credential detection device detects the credential image. If the credential image is detected as forged, the image is not uploaded to the server. Otherwise, the image is uploaded to the server to obtain the corresponding service.
[0081] It should be understood that the user equipment, forged credential detection apparatus and regional detection model in Figure 1 are merely illustrative and any number of user equipment, forged credential detection apparatus and regional detection model may be provided as required.
[0082] FIG2 is a flow chart of a method for detecting forged credentials provided in an embodiment of the present application. The method may be performed by the forged credential detection device in the architecture shown in FIG1 . As shown in FIG2 , the method may include the following steps:
[0083] Step 201: Obtain a credential image.
[0084] Step 203: Locate the key information area of the voucher image.
[0085] Step 205: Perform text anomaly detection on the key information area, and determine whether the voucher image is a forged voucher based on the result of the text anomaly detection.
[0086] It can be seen that when fraudsters forge vouchers, they usually edit or replace the text in the voucher image through Photoshop (an image editing software) and other image processing software, or even use the editing function of a smartphone to do so. However, due to the varying technical levels of these fraudsters, forged vouchers are often full of obvious defects. For example, the font style is not uniform, the text is not neatly arranged, and the image quality is inconsistent. These suspicious details and errors have become the key evidence to reveal the fraudulent behavior. Therefore, the present application determines whether the voucher image is a forged voucher by performing text anomaly detection on the key information area, thereby effectively realizing forged voucher detection, reducing the interference of false detection in non-key information areas, and improving the detection accuracy and efficiency compared to manual recognition methods.
[0087] The following describes in detail each step in the flowchart shown in Figure 2 and the effects that can be produced, in conjunction with embodiments. It should be noted that the terms "first" and "second" in this disclosure do not restrict the size, order, or quantity of the terms, but are merely used to distinguish between them. For example, "first forgery condition" and "second forgery condition" are used to distinguish between different forgery conditions.
[0088] First, the above step 201, namely “obtaining a voucher image”, is described in detail with reference to the embodiment.
[0089] The sources of credential images vary depending on the application scenario. For example, in e-commerce scenarios, these images may include payment receipts, postage screenshots, business licenses, and qualification certificates. In recruitment scenarios, these images may include educational credentials, graduation certificates, award certificates, and resignation certificates. In financial scenarios, these images may include proof of relationship, notarized documents, and bank statements.
[0090] Taking the payment interface screenshot shown in Figure 4a as an example, criminals can deceive merchants or platforms by modifying the payment amount (e.g., ¥110.00) to obtain illegal profits.
[0091] The above step 203, namely "locating the key information area of the voucher image", is described in detail below with reference to an embodiment.
[0092] In a possible implementation, a region detection model may be used to detect the credential image to obtain the key information region of the credential image.
[0093] The region detection model can be pre-trained based on a neural network model. For example, the key information regions of multiple voucher image samples can be pre-calibrated to obtain a training sample set, where each training sample includes a voucher image sample and its annotated key information region. The pre-calibrated training samples are then used to train the neural network model to obtain a region detection model. Key information regions refer to areas of a voucher that are often susceptible to tampering. The neural network model can employ models such as object detection models and image segmentation models, for example, the YOLO (You Only Look Once) family of object detection models. During training, voucher image samples from the training set are input into a neural network model such as the YOLO model to obtain predicted key information regions for the voucher image samples. The training objective can be to minimize the difference between the predicted key information regions and the corresponding calibrated key information regions. A loss function can be constructed based on this training objective. The value of the loss function is used in each iteration to update the model parameters using methods such as gradient descent until a preset training termination condition is met. The training termination condition can include, for example, the loss function value being less than or equal to a preset loss function threshold or the number of iterations reaching a preset threshold.
[0094] Furthermore, when training the aforementioned region detection model, the reference regions in the training samples can be further calibrated. Reference regions are regions that can be used as a reference to assist in determining the authenticity of key information regions. These are typically regions in the credential image that are difficult or unlikely to be tampered with. During the training process, the credential image samples in the training samples are input into a neural network model such as YOLO. In addition to obtaining predicted key information regions for the credential image samples, reference regions are also predicted. The aforementioned training objective can further include minimizing the difference between the predicted reference regions and the corresponding calibrated reference regions.
[0095] In another possible implementation, text recognition technology can be used to identify the credential image to obtain the text content in the credential image, and the area containing the text content is used as the key information area. For example, the text recognition technology can be OCR (Optical Character Recognition) technology or ICR (Intelligent Character Recognition) technology.
[0096] For example, after identifying text content using text recognition technologies such as OCR, the text content can be divided or clustered according to its position. The positions of the outermost points of the text in a cluster in the upper, lower, left, and right directions are used as the positions of the four sides to construct a rectangular area as the key information area.
[0097] In addition to the above methods, other methods can also be used to achieve this, such as using preset rules (such as the area with a preset length from the upper edge of the credential image to the lower edge) to locate the key information area of the credential image, etc., which are not listed here one by one.
[0098] The above step 205, namely "performing a text anomaly detection on the key information area and determining whether the voucher image is a forged voucher based on the result of the text anomaly detection", is described in detail below with reference to an embodiment.
[0099] In one possible implementation, the text content identified in the key information area can be obtained, and then the key information area can be reconstructed to detect forged documents. Specifically, the key information area in the document image can be masked; the masked key information area can be reconstructed using the document image obtained after the masking process and the text content corresponding to the key information area; the difference between the reconstructed key information area and the key information area before the masking process is determined; if the difference meets a preset first forgery condition, the document image is determined to be a forged document.
[0100] Among them, the masked key information areas can be reconstructed using diffusion models, generative models or applications.
[0101] In one example, diffusion models such as Stable-Diffusion (a text graph model based on the latent diffusion model) and Latent Diffusion Model (LDM) can be used to reconstruct masked key information regions. For example, the key information region in the voucher image is masked, random noise is superimposed on the masked region, and the random noise and the feature representation corresponding to the text content of the key information region are input into the diffusion model. The diffusion model then uses the feature representation corresponding to the text content to denoise the random noise over T time steps, resulting in the reconstructed key information region. The denoising process can be understood as predicting normally distributed noise and performing denoising at each time step to restore the true image content. This process simulates the reverse process of a Markov chain of length T. Where T is the total time step of the diffusion model. A longer T improves the denoising effect, but also increases the impact on computational performance. Therefore, a trade-off between the two is necessary. Empirical or experimental values can be used.
[0102] In another example, a generative model such as a text-based image model or a multimodal image generation model can be used to guide the prediction of the masked key information area using the text content as the input of the generative model, and the predicted content is used as the reconstructed key information area.
[0103] In another example, the text content can be used as an input parameter to call the interface of the application from which the credential image originated, and the area corresponding to the text content can be regenerated. For example, for a screenshot of a payment interface generated by a payment application, the payment application's interface can be called, and the text content in the key information area of the payment interface can be used as an input parameter to reconstruct the key information area.
[0104] Since the reconstructed key information area is obtained based on the text content, it can be considered to be closer to the original content and not tampered with. Therefore, by comparing it with the key information area before masking, it can be determined whether the credential image is a forged credential.
[0105] For example, the first forgery condition may be that the difference between the reconstructed key information region and the key information region before masking is greater than or equal to a first preset threshold. This difference can be measured, for example, by similarity. For example, if the similarity between the reconstructed key information region extracted and reconstructed using a neural network model and the key information region before masking is less than or equal to a preset similarity threshold, the credential image is determined to be a forged credential.
[0106] In another possible implementation, in this step, text alignment detection and / or font consistency detection can be performed on the key information area; if the text misalignment or font inconsistency in the key information area meets the preset second forgery condition, the certificate image is determined to be a forged certificate.
[0107] For example, the second counterfeit condition may be that the number of misaligned characters in the key information area is greater than a second preset threshold, or that the number of characters with inconsistent fonts is greater than a third preset threshold, etc.
[0108] For example, as shown in Figure 3, the key information area is checked for text alignment and font consistency. If text misalignment or font inconsistency exists and a second preset forgery condition is met, the voucher image is determined to be a forged voucher. For example, if the number of misaligned characters in the key information area exceeds a second preset threshold or the number of characters with inconsistent fonts exceeds a third preset threshold, the voucher image is determined to be a forged voucher.
[0109] In some embodiments, the text alignment detection may include at least one of the following methods:
[0110] Method 1 uses the center position of the reference area of the voucher image and the left and right edges of the text area in the key information area to detect whether the text in the key information area is centered. The reference area is obtained by detecting the voucher image using a region detection model. That is, in step 203, the voucher image is input into the region detection model, and the region detection model outputs the key information area and the reference area in the voucher image, where the reference area has a corresponding key information area.
[0111] In one example, as shown in Figure 4b, the payee icon is typically untampered with and serves as a reference area for the payment price in the voucher image. The payment price area is a key information area in the voucher image. Using the text area within the payment price area, namely the left and right edges of "-12.00," and the center of the payee icon, the center alignment of "-12.00" in the payment price area is checked. If "-12.00" in the payment price area is not centered, the voucher image may be forged.
[0112] Method 2 uses the same side edge of the reference area of the voucher image and the text area in the key information area to detect whether the text area in the key information area is aligned on the same side, where the same side is the left side or the right side.
[0113] In one example, as shown in FIG4b , the transfer time is the key information area of the voucher image. The transfer order number is the reference area of the transfer time. The left edge of the text area in the transfer order number and the transfer time, i.e., “July 10, 2023 11:56:12”, is used to detect whether the left edge of the text area in the transfer time is aligned with the left edge of the text area in the transfer order number. In an embodiment of the present application, if the left side of the text area in the transfer time, i.e., “July 10, 2023 11:56:12”, is not aligned with the left side of the transfer order number, i.e., “10000499012023071000996808910063”, it means that the voucher image may be a forged voucher.
[0114] Method three uses the text line spacing in the reference area of the voucher image and the text line spacing in the key information area to detect whether the text line spacing in the key information area is aligned.
[0115] In one example, if the difference between the text line spacing in the key information area of the voucher image and the text line spacing in the reference area is greater than a preset threshold, the text line spacing in the key information area is determined to be misaligned, and the voucher image may be a forged voucher. The text line spacing can be determined by the distance between the bottom edge of one line of text and the top edge of the next line of text.
[0116] Method 4 uses an inline alignment detection model to check whether the text in the key information area is aligned within the line. This inline alignment detection model is pre-trained based on a neural network model. It should be noted that the inline alignment detection model primarily detects whether the text in the key information area has been partially tampered with, that is, whether individual characters are aligned with other text within the line. If the detection result shows inline misalignment, the voucher image is a forged voucher.
[0117] The intra-row alignment detection model can be pre-trained based on a neural network model. For example, the intra-row misaligned regions within the key information regions of multiple voucher image samples can be pre-calibrated to obtain a training sample set, where each training sample includes the voucher image sample and the intra-row misaligned regions within its annotated key information regions. The pre-calibrated training samples are then used to train the neural network model to obtain the intra-row alignment detection model. The key information regions refer to regions of the voucher that are often susceptible to tampering. The neural network model can employ models such as object detection models or image segmentation models, for example, the YOLO (You Only Look Once) family of object detection models. During training, the voucher image samples from the training set are input into a neural network model such as YOLO to obtain the predicted intra-row misaligned regions within the key information regions for the voucher image samples. The training objective can be to minimize the difference between the predicted intra-row misaligned regions and the corresponding calibrated intra-row misaligned regions. A loss function can be constructed based on this training objective. The value of the loss function is used in each iteration to update the model parameters using methods such as gradient descent until a preset training end condition is met. The training end conditions may include, for example, the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset threshold, etc.
[0118] In some embodiments, the font consistency detection includes at least one of the following methods:
[0119] Method 1: Use a font detection model to detect whether the fonts of the text in the key information area are consistent. The font detection model is pre-trained based on a neural network model. If the detection result shows that the fonts are inconsistent, it means that the certificate image may be a forged certificate.
[0120] The font detection model can be pre-trained based on a neural network model. For example, text regions with inconsistent fonts within the key information regions of multiple voucher image samples can be pre-labeled to obtain a training sample set, where each training sample includes a voucher image sample and its labeled text regions with inconsistent fonts. The pre-labeled training samples are then used to train a neural network model to obtain a font detection model. The critical information regions refer to areas of the voucher that are often susceptible to tampering. The neural network model can employ models such as object detection models and image segmentation models, for example, the YOLO (You Only Look Once) family of object detection models. During training, the voucher image samples from the training sample are input into a neural network model such as YOLO to obtain predicted text regions with inconsistent fonts within the key information regions for the voucher image samples. The training objective can be to minimize the difference between the predicted text regions with inconsistent fonts within the key information regions and the corresponding labeled text regions with inconsistent fonts. A loss function can be constructed based on this training objective. The value of the loss function is used in each iteration to update the model parameters using methods such as gradient descent until a preset training end condition is met. The training end conditions may include, for example, the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset threshold, etc.
[0121] Method 2: Use a font recognition tool to identify the font of the text in the key information area and, based on the recognition results, determine whether the fonts of the text in the key information area are consistent. If the fonts are inconsistent, the voucher image may be a forged voucher. Exemplary font recognition tools may include WhatTheFont or Adobe Font Finder.
[0122] Method three extracts the features of each character in the key information area, clusters them based on the features of each character, and determines whether the fonts of the characters in the key information area are consistent based on the clustering results. If the fonts are inconsistent, it indicates that the voucher image may be a forged voucher. For example, the features of the characters may include: stroke width, edge features, etc. The clustering algorithm may include K-Means clustering algorithm, mean shift clustering algorithm and / or hierarchical clustering algorithm, etc. A key information area can be considered to have several fonts according to the number of clusters it is clustered into.
[0123] In one example, as shown in Figure 5, the top box contains the sender and recipient's address information. The box in the lower left corner contains the amount and weight information. The box in the lower right corner contains the courier information. If the voucher image is genuine, the font and style of the amount, weight, and other information in the lower left box should match the font and style of the courier information in the lower right box. Otherwise, the voucher image may be a forgery.
[0124] As can be seen, the embodiments of the present application use font alignment detection and / or font consistency detection to detect whether a credential image is forged. This not only improves the effectiveness of forged credential detection, but also makes the result of determining whether a credential image is forged more interpretable. Whether it is font alignment anomalies or font inconsistencies, manual secondary review can be performed.
[0125] FIG6 is another flow chart of a method for detecting forged credentials provided by an embodiment of the present application, which can be executed by a cloud server. As shown in FIG6 , the method can include the following steps:
[0126] Step 601: The cloud server obtains a credential image sent by a user device.
[0127] In an embodiment of the present application, the cloud server obtains a voucher image sent by the user device. For example, the voucher image can be a return postage voucher image uploaded by the consumer to the e-commerce platform via the user device.
[0128] Step 603: The cloud server locates the key information area of the credential image.
[0129] In the embodiment of the present application, the cloud server locates the key information area of the credential image.
[0130] It should be noted that the method by which the cloud server locates the key information area of the credential image in step 603 is similar to the method by which the user device locates the key information area of the credential image in step 203 , and will not be described in detail here.
[0131] In step 605 , the cloud server performs text anomaly detection on the key information area and determines whether the voucher image is a forged voucher based on the result of the text anomaly detection.
[0132] In an embodiment of the present application, the cloud server performs text anomaly detection on the key information area and determines whether the credential image is a forged credential based on the result of the text anomaly detection.
[0133] It should be noted that the method by which the cloud server performs text anomaly detection on the key information area in step 605 is similar to the method by which the user device performs text anomaly detection on the key information area in step 205, and will not be repeated here.
[0134] In step 607 , if it is determined that the credential image is a forged credential, the cloud server denies the service requested by the user device based on the credential image.
[0135] In an embodiment of the present application, if it is determined that the credential image is a forged credential, the cloud server rejects the service requested by the user device based on the credential image.
[0136] In one example, when returning goods, a customer needs to provide proof of payment for the return postage and / or proof of the mailing slip. If the proof is identified as a forged document, the refund service will be denied.
[0137] In another example, when a merchant registers a store on an e-commerce website, he or she needs to provide business licenses, permits and other credentials. If the credentials are identified as forged, the merchant will be denied the store registration service.
[0138] In a possible implementation, the cloud server returns a denial of service response to the user device, where the response includes information indicating that the credential image is a forged credential.
[0139] In addition, it should be noted that the above method provided in the embodiments of the present application can be applied to a variety of application scenarios, including but not limited to the following:
[0140] (1) Merchant attitude complaint scenario: Consumers complain about abusive behavior by submitting images of abusive and harassing credentials such as text messages and call logs to the e-commerce platform. The e-commerce platform detects the credential images and rejects the consumer's complaint if they detect that the credential images are forged; otherwise, they accept the consumer's complaint and initiate follow-up service matters.
[0141] (2) Taobao return postage voucher scenario: When returning goods, consumers pay the postage in advance and upload a postage voucher image, submitting multiple postage payment records and other voucher images. The e-commerce platform detects the voucher image and refuses to refund the postage if it detects that the voucher image is forged; otherwise, the postage is refunded to the consumer.
[0142] (3) Taobao return complaint scenario: When a consumer initiates a return, they upload a proof image showing poor quality or insufficient quantity, and submit a photo and screenshot of the proof image. The e-commerce platform will check the proof image and reject the return if it detects that the proof image is forged. Otherwise, the platform will accept the consumer's complaint and initiate follow-up service.
[0143] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0144] According to another embodiment, based on the same concept as the above method embodiment, the present application embodiment proposes a device for detecting forged credentials. The schematic diagram of the structure is shown in Figure 7. The device includes:
[0145] The image acquisition unit 710 is configured to acquire a credential image.
[0146] The region locating unit 720 is configured to locate the key information region of the credential image.
[0147] The anomaly detection unit 730 is configured to perform text anomaly detection on the key information area, and determine whether the voucher image is a forged voucher based on the result of the text anomaly detection.
[0148] As one of the possible implementations, the region positioning unit 720 is specifically configured to detect the credential image using a region detection model to obtain the key information region of the credential image, where the region detection model is pre-trained based on a neural network model.
[0149] As another achievable manner, the region positioning unit 720 is further specifically configured to use text recognition technology to recognize the voucher image, obtain the text content in the voucher image, and use the region where the text content is located as the key information region.
[0150] As one of the feasible methods, the anomaly detection unit 730 is specifically configured to mask the key information area in the credential image; reconstruct the masked key information area using the credential image obtained after masking and the text content corresponding to the key information area; determine the difference between the reconstructed key information area and the key information area before masking; if the difference meets the preset first forgery condition, determine that the credential image is a forged credential.
[0151] As one of the feasible ways, the anomaly detection unit 730 is further specifically configured to perform text alignment detection and / or font consistency detection on the key information area; if the text misalignment or font inconsistency in the key information area meets the preset second forgery condition, the voucher image is determined to be a forged voucher.
[0152] As one of the achievable methods, the anomaly detection unit 730 is further specifically configured to use the center position of the reference area of the credential image and the left and right edges of the text area in the key information area to detect whether the text in the key information area is centered; or, use the same side edges of the reference area of the credential image and the text area in the key information area to detect whether the text area in the key information area is aligned on the same side, where the same side is the left or right side; or, use the text line spacing in the reference area of the credential image and the text line spacing in the key information area to detect whether the text line spacing in the key information area is aligned; or, use an in-line alignment detection model to detect whether the text in the key information area is in-line aligned, and the in-line alignment detection model is pre-trained based on a neural network model; wherein the reference area is obtained by detecting the credential image using the area detection model.
[0153] As one of the feasible ways, the anomaly detection unit 730 is further specifically configured to use a font detection model to detect whether the fonts of the text in the key information area are consistent, and the font detection model is pre-trained based on a neural network model; or, use a font recognition tool to perform font recognition on the text in the key information area, and determine whether the fonts of the text in the key information area are consistent based on the recognition results; or, extract the features of each text in the key information area, cluster them based on the features of each text, and determine whether the fonts of the text in the key information area are consistent based on the clustering results.
[0154] Based on the same concept as the above-mentioned method embodiment, the present application embodiment proposes a device for detecting counterfeit credentials. Its structural diagram is shown in Figure 8. The device includes: an image acquisition unit 810, an area positioning unit 820, an anomaly detection unit 830, and a service processing unit 840. The main functions of each component unit are as follows:
[0155] The image acquisition unit 810 is configured to acquire a credential image sent by a user device.
[0156] The region locating unit 820 is configured to locate the key information region of the voucher image.
[0157] The anomaly detection unit 830 is configured to perform text anomaly detection on the key information area, and determine whether the voucher image is a forged voucher based on the result of the text anomaly detection.
[0158] The service processing unit 840 is configured to deny the service requested by the user device based on the credential image if it is determined that the credential image is a forged credential.
[0159] As one possible implementation, the forged credential detection apparatus further includes a response returning unit configured to return a denial of service response to the user device, the response including information indicating that the credential image is a forged credential.
[0160] It should be noted that the image acquisition unit 810 , the region positioning unit 820 and the abnormality detection unit 830 are respectively the same as the image acquisition unit 710 , the region positioning unit 720 and the abnormality detection unit 730 .
[0161] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0163] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0164] And an electronic device comprising:
[0165] one or more processors; and
[0166] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0167] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.
[0168] 9 exemplarily illustrates the architecture of an electronic device, which may include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920 may be communicatively connected via a communication bus 930.
[0169] Among them, the processor 910 can be implemented by a general CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0170] The memory 920 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device 900 and a basic input and output system (BIOS) 922 for controlling the low-level operations of the electronic device 900. In addition, a web browser 923, a data storage management system 924, and a forged credential detection device 925, etc. can also be stored. The above-mentioned forged credential detection device 925 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory X20 and is called and executed by the processor X10.
[0171] The input / output interface 913 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0172] The network interface 914 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0173] The bus 930 comprises a pathway for transmitting information between the various components of the device (eg, the processor 910 , the video display adapter 911 , the disk drive 912 , the input / output interface 913 , the network interface 914 , and the memory 920 ).
[0174] It should be noted that although the above device only shows the processor 910, video display adapter 911, disk drive 912, input / output interface 913, network interface 914, memory 920, bus 930, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0175] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0176] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.
Claims
1. A method for detecting forged credentials, characterized in that: The method comprises: Get the credential image; Locating key information areas of the credential image; Performing text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on the result of the text anomaly detection.
2. The method according to claim 1, characterized in that The key information areas for locating the credential image include: The credential image is detected using a region detection model to obtain key information regions of the credential image, wherein the region detection model is pre-trained based on a neural network model.
3. The method according to claim 1, characterized in that The key information areas for locating the credential image include: The voucher image is recognized by using a text recognition technology to obtain the text content in the voucher image, and the area where the text content is located is used as the key information area.
4. The method according to claim 2 or 3, characterized in that Performing text anomaly detection on the key information area and determining whether the voucher image is a forged voucher based on the result of the text anomaly detection includes: Masking the key information area in the voucher image; Reconstructing the masked key information area using the credential image obtained after masking and the text content recognized from the key information area; Determine the difference between the reconstructed key information region and the key information region before masking; If the difference satisfies a preset first forgery condition, the credential image is determined to be a forged credential.
5. The method according to claim 2 or 3, characterized in that Performing text anomaly detection on the key information area and determining whether the voucher image is a forged voucher based on the result of the text anomaly detection includes: Performing text alignment detection and / or font consistency detection on the key information area; If the text misalignment or inconsistent fonts in the key information area meet a preset second forgery condition, the certificate image is determined to be a forged certificate.
6. The method according to claim 5, characterized in that The text alignment detection includes at least one of the following: Using the center position of the reference area of the credential image and the left and right edges of the text area in the key information area, detecting whether the text in the key information area is centered; Using the same side edge of the reference area of the credential image and the text area in the key information area, detecting whether the text area in the key information area is aligned on the same side, where the same side is the left side or the right side; Using the text line spacing in the reference area of the credential image and the text line spacing in the key information area, detecting whether the text line spacing in the key information area is aligned; Utilizing an inline alignment detection model to detect whether the text in the key information area is inline aligned, the inline alignment detection model being pre-trained based on a neural network model; The reference area is obtained by detecting the credential image using a region detection model.
7. The method according to claim 5, characterized in that The font consistency detection includes at least one of the following: Detecting whether the fonts of the characters in the key information area are consistent using a font detection model, wherein the font detection model is pre-trained based on a neural network model; Performing font recognition on the text in the key information area using a font recognition tool, and determining whether the fonts of the text in the key information area are consistent based on the recognition result; Features of each character in the key information area are extracted, clustering is performed based on the features of each character, and whether the fonts of the characters in the key information area are consistent is determined according to the clustering result.
8. A method for detecting forged credentials, executed by a cloud server, characterized in that: The method comprises: Obtain the credential image sent by the user's device; Locating key information areas of the credential image; Performing a text anomaly detection on the key information area, and determining whether the credential image is a forged credential based on a result of the text anomaly detection; If it is determined that the credential image is a forged credential, the service requested by the user device based on the credential image is denied.
9. The method according to claim 8, characterized in that The method further comprises: A response rejecting the service is returned to the user device, the response including information indicating that the credential image is a forged credential.
10. A device for detecting forged credentials, characterized in that: The device comprises: an image acquisition unit configured to acquire a credential image; an area positioning unit, configured to locate a key information area of the credential image; The anomaly detection unit is configured to perform text anomaly detection on the key information area and determine whether the certificate image is a forged certificate based on the result of the text anomaly detection.
11. A device for detecting forged documents, characterized in that: The device comprises: an image acquisition unit, configured to acquire a credential image sent by a user device; an area positioning unit, configured to locate a key information area of the credential image; an anomaly detection unit configured to perform text anomaly detection on the key information area and determine whether the credential image is a forged credential based on a result of the text anomaly detection; The service processing unit is configured to deny the service requested by the user equipment based on the credential image if it is determined that the credential image is a forged credential.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being configured to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method according to any one of claims 1 to 9.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Certificate information identification method and device and computer equipment
CN111259894A
Method for training image detection model, image detection method and corresponding device
CN117593570A
Detection method and device for forged certificate, storage medium and electronic equipment
CN118298440A
Information processing device, information processing method, program, and information processing system
JP2024032435A
Cited By
Method, system and equipment for identifying validity of voucher file and medium
CN120877295A