Credit material repeated detection method and device based on image fingerprints

By introducing joint modeling of perceptual hash fingerprints and deep semantic vectors, combined with user behavior and time series, the problem of image duplication detection in credit systems is solved, achieving efficient and accurate image identification and risk differentiation, and improving the anti-fraud capabilities of credit systems.

CN120876079AInactive Publication Date: 2025-10-31JIANGSU SUNING BANK CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511384006.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing credit systems lack the ability to understand the hierarchical structure of image content in image material duplication detection, making it difficult to identify the reuse of deformed materials and non-perfectly matched images. Furthermore, they lack a stable image fingerprinting mechanism, which cannot support efficient similarity comparison and contextual dimension judgment for large-scale historical material databases, leading to frequent credit fraud.

Method used

A method for detecting duplicate credit materials based on image fingerprints is adopted. By jointly modeling perceptual hash fingerprints and deep semantic vectors, and combining user identity and time series, structured detection results are generated, enabling multi-dimensional recognition and risk differentiation of images.

Benefits of technology

It improves the accuracy and sensitivity of image duplication detection, can identify slight image deformation, supports second-level response for a million-level credit image library, generates structured risk output, and enhances the accuracy and traceability of anti-fraud strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876079A_ABST
    Figure CN120876079A_ABST
Patent Text Reader

Abstract

The invention discloses a credit material repeated detection method and equipment based on image fingerprints. The method comprises the following steps: receiving an original image of a credit material uploaded by a user, and preprocessing the original image; extracting image fingerprint data from the preprocessed image, wherein the image fingerprint data comprises a perceptual hash fingerprint and a deep semantic vector; performing fingerprint matching on the original image and a historical image based on the image fingerprint data of the original image and the fingerprint data of the historical image in a historical database, and if the fingerprint matching is successful, determining that the original image of the credit material uploaded by the user is a repeated image; generating a repeated image judgment result in combination with the user identity and the time sequence; and outputting a structured detection result based on the structured interpretable output mechanism. According to the method, through joint modeling of the perceptual hash fingerprints and the deep semantic vectors, the problem that a traditional image deduplication method can only identify identical files is effectively solved, and the accuracy and sensitivity of identifying forged materials are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial risk control technology, specifically to a method and device for detecting duplicate credit materials based on image fingerprints. Background Technology

[0002] In modern digital lending processes, customers are increasingly shifting from offline paper submissions to online electronic uploads when submitting loan application materials to financial institutions. This is particularly true for micro and small enterprise loans, inclusive personal loans, and contactless approval processes. Customers typically submit various materials in image file format via mobile apps, lending platforms, or third-party interfaces. These materials include, but are not limited to: legal representative's ID card, business license, bank statements, property ownership certificate, lease agreement, mortgage registration certificate, and purchase and sale contract.

[0003] Traditional document review processes primarily rely on manual methods to view, identify, verify, and compare images submitted by clients to determine their completeness, clarity, compliance, and authenticity. However, with the rapid expansion of credit business and the increasing frequency of applications, manual review faces multiple challenges in terms of human resources, fraud prevention capabilities, and review efficiency. Therefore, more and more credit systems are attempting to introduce automated methods such as image recognition, OCR, and template comparison for preliminary identification and verification of materials.

[0004] However, existing technologies still have significant shortcomings in determining the authenticity of image materials, especially in determining whether image materials are reused, historical materials are copied, or images of others are maliciously reused. These shortcomings are mainly reflected in the following aspects: 1. Lack of systematic ability to identify image "reuse". Currently, most credit systems in the industry use simple methods such as filename comparison, image MD5 hashing, and pixel-level image comparison for "duplicate image recognition." These methods can only find completely identical image files, but cannot identify the following "variant duplicate materials" commonly seen in real-world business: (1) Re-screenshot, crop, and compress previously submitted materials before uploading them; (2) Malicious applicants steal other people's materials and evade detection by blurring the edges or retaking the photos; (3) The same material is used repeatedly in multiple user accounts; (4) After the image is tampered with and then saved again, the file level is different, but the content is similar.

[0005] Therefore, current systems generally lack a "repetition recognition engine" with the ability to understand image content at a hierarchical level, making it difficult to intercept the reuse of deformed materials and non-perfectly matched images, leading to frequent credit fraud incidents.

[0006] 2. Lack of a stable and robust image fingerprinting mechanism Existing image hashing methods, such as the traditional MD5 and SHA-1, are highly sensitive to minor changes in the image. For example, even a simple compression, resizing, or cropping of an image can cause a significant change in the hash value, making it impossible to determine its similarity to the original image.

[0007] Although some technologies attempt to introduce perceptual hashing (such as pHash, aHash, dHash) for approximate image matching, their low feature dimension (usually 64 bits) leads to problems such as high collision rate, high misjudgment rate, and weak robustness in large-scale image libraries, making them difficult to apply to scenarios in financial lending where risk is extremely sensitive.

[0008] Meanwhile, perceptual hashing-based methods are difficult to extend to deep image similarity modeling, meaning they cannot identify images with "semantic similarity in image content but significant changes in pixel structure".

[0009] 3. Unable to support efficient similarity comparison for large-scale historical material databases. In real-world financial scenarios, image databases are often massive, reaching millions of records, especially when full comparisons of historical data spanning multiple years are required. However, most existing systems lack indexable, scalable, and low-latency image fingerprint databases, and have not introduced high-performance approximate vector indexing structures, such as image vector matching engines based on FAISS, Annoy, and Milvus.

[0010] This results in the system only being able to perform local comparisons and not global searches when detecting image repetition, which greatly reduces the detection coverage and accuracy, and also fails to support the second-level recognition requirements under high concurrency.

[0011] 4. Lack of contextual judgment and risk level differentiation mechanism In real-world business scenarios, even if two images are similar, it doesn't necessarily indicate malicious intent. For example, the same customer might upload similar contract copies in two different transactions, or repeatedly use the same photo of collateral within the past six months. Therefore, determining image duplication cannot be based solely on similarity scores; it also requires intelligent judgment by considering contextual information such as user identity, time span, and business scenario.

[0012] Currently, most systems lack the ability to jointly model "image fingerprint + user behavior + time dimension", making it difficult to accurately distinguish between them. (1) Repeated use by the same user (reasonable behavior) (2) Repeated behavior across users (high-risk behavior) (3) Images with highly similar content but different generation paths (disguise and evasion) This "one-size-fits-all" identification method is prone to problems such as false positives, missed detections, or inability to trace the source.

[0013] Currently, the ability of most commercial banks, consumer finance platforms, and internet lending companies to perform duplicate material testing is generally limited to the following dimensions: (1) It can only find completely identical images, but cannot determine spoofed duplicate images. (2) Lack of cross-user image comparison logic makes it impossible to determine the behavior of uploading the same image from multiple accounts. (3) No structured index was built for historical materials, and there is a lack of an efficient image plagiarism detection database. (4) Lack of a risk control judgment system that links "image fingerprint + user behavior + submission time" Due to the aforementioned technical shortcomings, some users are able to "bypass detection" and "easily use the kits," directly leading to credit losses, soaring fraud rates, increased manual review costs, and decreased customer trust. Summary of the Invention

[0014] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and device for detecting duplicate credit materials based on image fingerprints.

[0015] To achieve the above objectives, in a first aspect, the present invention provides a method for duplicate detection of credit materials based on image fingerprints, comprising: Receive the original image of the credit materials uploaded by the user, and preprocess the original image; Image fingerprint data is extracted from the preprocessed image, and the image fingerprint data includes perceptual hash fingerprint and deep semantic vector; Based on the image fingerprint data of the original image and the fingerprint data of historical images in the historical database, the original image and the historical image are matched. If the fingerprint of the original image and the historical image are successfully matched, it is determined that the original image of the credit material uploaded by the user is a duplicate image. The system combines user identity and time series data to generate duplicate image detection results. The structured detection results are output based on a structured and interpretable output mechanism.

[0016] Furthermore, the preprocessing method is as follows: Convert the user-uploaded raw image into a grayscale image; Perform edge cropping to obtain the main body region of the original image; The cropped images are uniformly scaled to a fixed size.

[0017] Furthermore, the extraction method of the perceptual hash fingerprint is as follows: The grayscale matrix of the preprocessed image is then subjected to a two-dimensional discrete cosine transform, as follows: ; in, This is the result after the discrete cosine transform. The image after preprocessing. This is the discrete cosine transform. Extract the 8×8 low-frequency submatrix from the top left corner, as follows: ; in, To extract the 8×8 low-frequency submatrix from the top left corner; The mean of the low-frequency submatrix is ​​calculated as follows: ; in, The mean of the calculated low-frequency submatrix, Low-frequency sub-matrix The Middle One element; The generated perceptual hash fingerprint is as follows: ; in, The generated perceptual hash fingerprint.

[0018] Furthermore, the deep semantic vector is extracted in the following way: Deep feature extraction is performed on the preprocessed image based on a deep neural network model, as follows: ; in, The deep feature vector is extracted from the preprocessed image. This represents the forward inference of a deep neural network model. It is a real number; For the deep feature vector Normalization is performed to obtain the deep semantic vector, as follows: ; in, For the extracted deep semantic vector, Deep feature vectors The L2 norm.

[0019] Furthermore, the specific method for fingerprint matching between the original image and historical images is as follows: Hamming distance matching is performed based on the perceptual hash fingerprints of the original image and the historical images, as follows: ; in, Perceptual hash fingerprint of the calculated original image Perceptual hash fingerprints of historical images Hamming distance, These are the k-th perceptual hash fingerprints in the original image and the historical image, respectively; If the perceptual hash fingerprint of the original image Perceptual hash fingerprints of historical images Hamming distance This indicates that the structural fingerprint match was successful. The set Hamming distance threshold; The cosine similarity between the original image and historical images is calculated based on deep semantic vectors, as follows: ; in, To calculate the cosine similarity between the original image and the historical images, For the deep semantic vector of historical images, The dot product symbol for vectors; If the cosine similarity between the original image and the historical image This indicates that the semantic fingerprint match was successful. The set similarity threshold; When both the structural fingerprint and semantic fingerprint are successfully matched, it means that the fingerprints of the original image and the historical image have been successfully matched.

[0020] Furthermore, the method for generating duplicate image judgment results by combining user identity and time series is as follows: Extract the unique identifier of the current user, the unique identifier of the user to which the historical image belongs, the upload timestamp of the original image, and the upload timestamp of the historical image. If the unique identifier of the current user is inconsistent with the unique identifier of the user to which the historical image belongs, it is determined to be cross-user image reuse; otherwise, it is determined to be repeated use of historical materials of the same user. If the cosine similarity between the original image and the historical image gradually decreases over time, an image fingerprint evolution path is constructed and determined to be an image tampering chain evolution behavior.

[0021] Furthermore, the Hamming distance threshold The similarity threshold is 5. It is 0.9.

[0022] Furthermore, the structured detection results are based on a JSON structure and include image comparison conclusions, duplication types, similar image information, risk levels, and recommendations.

[0023] In a second aspect, the present invention provides a credit material duplication detection device based on image fingerprints, comprising a storage medium and a processor, wherein the storage medium stores a computer program, which, when executed by the processor, is used to implement the above-described method.

[0024] Beneficial effects: 1. This invention introduces a multi-fingerprint fusion strategy for the first time in the identification of the authenticity of credit materials. By jointly modeling perceptual hash fingerprints and deep semantic vectors, it achieves highly robust matching of material images under slight deformation such as compression, cropping, and screenshotting. This mechanism effectively solves the problem that traditional image deduplication methods can only identify completely identical files, thus improving the accuracy and sensitivity of identifying counterfeit materials. 2. This invention can distinguish different risk scenarios such as "duplication by the same user", "reuse across users", and "image tampering chain", and generate structured duplication type tags. It fills the gap in traditional image comparison systems that do not make sufficient use of behavioral information such as uploader upload time, and significantly enhances the accuracy and traceability of anti-fraud strategies. 3. This invention can support the approximate deduplication requirements of a credit image database of millions of images, with a response time of seconds. The system supports structured risk output and can be seamlessly integrated into the approval engine and risk control platform to achieve automated and real-time judgment of material authenticity. Attached Figure Description

[0025] Figure 1 This is a flowchart of a credit material duplication detection method based on image fingerprints according to an embodiment of the present invention. Detailed Implementation

[0026] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solutions of the present invention, and it should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0027] like Figure 1 As shown, this embodiment of the invention provides a method for duplicate detection of credit materials based on image fingerprints, including: The system receives raw images of credit materials uploaded by users and preprocesses these images. The preprocessing method is as follows: First, the user-uploaded raw image (Iraw) is converted to a grayscale image (Igray). This process removes color interference by weighted fusion of RGB three-channel brightness information, preserving image structure and texture information, and simplifying data dimensions.

[0028] Next, the crop_margin(Igray) operation is performed to obtain the main image area. This step is used to remove redundant areas around the image, such as the shooting background, frame edges, watermarks, or shadow interference, in order to focus on the core material content and improve fingerprint stability.

[0029] Finally, the cropped images are uniformly scaled to a fixed size, such as 128×128, which is the preprocessed image. The value range is 0-255. Size normalization ensures that subsequent fingerprint extraction is consistent in the input dimension, which is beneficial to the consistency and stability of fingerprint comparison.

[0030] Image fingerprint data is extracted from the preprocessed image. This fingerprint data includes perceptual hash fingerprints (pHash) and deep semantic vectors (semantic feature fingerprints) to balance accuracy, robustness, and computational efficiency. Perceptual hash fingerprinting aims to extract low-frequency information from the image and construct a 64-bit fingerprint consistent with the semantic content of the image. The specific extraction method is as follows: The grayscale matrix of the preprocessed image is then subjected to a two-dimensional discrete cosine transform, as follows: ; in, The result of the discrete cosine transform represents the frequency components of the image, which are closely related to the energy distribution of the image content. The image after preprocessing. Discrete cosine transform (DCT) represents the frequency domain transformation of an image. It is one of the key steps in the perceptual hash algorithm to capture the structural contour features of the image, thereby achieving more robust image similarity recognition.

[0031] Extract the 8×8 low-frequency submatrix from the top left corner (rows 0 to 7 and columns 0 to 7), as follows: ; in, Extracting the 8×8 low-frequency submatrix (64 values) in the top left corner is crucial for understanding the image structure and overall contour. This step extracts the "low-frequency core region" with the strongest structure and least variation from the image's frequency space, forming the basis for constructing a perceptual hash fingerprint (pHash).

[0032] The mean of the low-frequency submatrix is ​​calculated as follows: ; in, The mean of the calculated low-frequency submatrix, Low-frequency sub-matrix The Middle This step is crucial in the entire pHash perceptual hashing algorithm. Its function is to generate an average reference value for the low-frequency energy of the image, which is subsequently used to construct the image structural fingerprint, thereby enabling image similarity comparison.

[0033] The generated perceptual hash fingerprint is as follows: ; in, This generates a perceptual hash fingerprint. The structural information of the image is compressed into a 64-bit binary signature for subsequent tasks such as image duplication detection and image fingerprint comparison. This application extracts the frequency information of the image through discrete cosine transform and uses the representativeness of the image structure by the frequency intensity to create the fingerprint. The final 64-bit binary vector is the image's "identity signature."

[0034] To improve the robustness of image duplication detection, this application introduces a deep neural network model (ResNet-18) for deep feature extraction. Specifically, the extraction method of the aforementioned deep semantic vectors is as follows: Deep feature extraction is performed on the preprocessed image based on a deep neural network model, as follows: ; in, The deep feature vector extracted from the preprocessed image has a dimension of 512 and represents the deep semantic representation of the image. This represents the forward inference of a deep neural network model. real number For deep feature vectors Normalization is performed to obtain the deep semantic vector, as follows: ; in, The extracted deep semantic vectors are converted into unit vectors to facilitate subsequent cosine similarity calculations. Deep feature vectors The L2 norm (Euclidean Norm).

[0035] The fingerprint data of the original image is matched with the fingerprint data of historical images in the historical database. If the fingerprints of the original image and the historical images match successfully, the original image of the credit materials uploaded by the user is determined to be a duplicate image.

[0036] The specific method for matching fingerprints between the original image and historical images is as follows: Hamming distance matching is performed based on the perceptual hash fingerprints of the original image and historical images, as follows: ; in, Perceptual hash fingerprint of the calculated original image Perceptual hash fingerprints of historical images Hamming distance, These are the k-th perceptual hash fingerprints in the original image and the historical image, respectively, which count how many bits of the two image fingerprints are different in 64 bits.

[0037] If the perceptual hash fingerprint of the original image Perceptual hash fingerprints of historical images Hamming distance This indicates that the structural fingerprint match was successful. For the set Hamming distance threshold, it is preferable to use If the two image fingerprints differ by a maximum of 5 bits, they are considered to be "structurally identical", meaning the structural fingerprints match successfully.

[0038] The cosine similarity between the original image and historical images is calculated based on deep semantic vectors, as follows: ; in, To calculate the cosine similarity between the original image and the historical images, For the deep semantic vector of historical images, The dot product symbol for vectors If the cosine similarity between the original image and the historical image This indicates that the semantic fingerprint match was successful. This is the set similarity threshold. The closer the cosine similarity between the original image and the historical images is to 1, the more "semantically consistent" the two images are. The preferred value is 0.90. When the cosine similarity is greater than 90%, it is considered a duplicate image. This is suitable for situations where the content structure is consistent but the pixel variations are large, such as compression, color correction, filters, and font replacement. When both the structural fingerprint and semantic fingerprint match successfully, it means that the original image and historical image fingerprints match successfully, indicating that the user-uploaded original image is a highly similar duplicate image. When both the structural fingerprint and semantic fingerprint match fail, the user-uploaded original image can be directly judged as a non-duplicate image. When only the structural fingerprint or only the semantic fingerprint matches successfully, but the other fingerprint fails to pass the threshold, the system judges the currently uploaded original image as a low-confidence duplicate image. For such low-confidence results, different strategies can be adopted according to business needs, such as marking it as "awaiting manual review," delaying approval, or combining it with other risk characteristics (such as user behavior and time series) for comprehensive judgment.

[0039] To further enhance the risk control and behavioral understanding capabilities of the image duplication detection system, this invention also introduces user-dimensional and time-dimensional information to construct a behavior-aware duplication judgment mechanism. This mechanism generates duplicate image judgment results by combining user identity and time series, thereby subdividing image duplication situations into different risk types, thus supporting more accurate system response strategies. Based on obtaining the matching results of the currently uploaded original image and historical images, the following behavioral context variables are extracted: user_query: A unique identifier for the user who uploaded the image; user_match: The user to whom the historical image was matched; t_query: The timestamp of the current image upload; t_match: Timestamps of historical image uploads.

[0040] Based on the above variables, the following three-level decision logic is constructed: (1) Judgment of cross-user image reuse If the current uploading user's `user_query` does not match the historical image's owner's `user_match`, and the image fingerprint matches successfully (both the structural and semantic fingerprints exceed the corresponding thresholds), then it is determined to be cross-user image reuse. This behavior is highly risky and is commonly seen in forgery behaviors such as some users sharing images or using templates.

[0041] (2) Judgment of duplicate user historical materials Provided that the image fingerprint matching is successful (i.e., both the structural fingerprint and semantic fingerprint exceed the corresponding thresholds), if user_query=user_match, meaning the current image and a historical image belong to the same user, and the time difference between the two uploads exceeds a set threshold (e.g., 30 days), then it can be determined that the same user's historical materials are being reused. This type generally does not pose a risk, but can serve as a basis for optimizing material reuse and for business reminders.

[0042] (3) Image evolution chain (tampering path) recognition Assuming successful image fingerprint matching, if the current image and multiple historical images simultaneously possess the following characteristics: 1. Decreasing similarity trend (e.g., 0.98→0.94→0.91); 2. The increasing trend of upload time (e.g., January 1st → January 5th → January 9th); The system can then construct an image fingerprint evolution path and identify it as an image tampering chain evolution behavior. This situation is common when some users or black market groups generate variant images in batches to evade detection, and it has high identification and tracking value.

[0043] The above mechanism supports structured reuse type label output (e.g., "repeated type": "cross-user" / "historical reuse" / "tampering chain"), and can be combined with risk level scoring for subsequent strategy actions such as automatic approval, manual review and scheduling, and account locking.

[0044] It can also output structured detection results based on a structured and interpretable output mechanism. To facilitate the automated access and processing of image duplication detection results in business systems such as credit approval engines and risk strategy platforms, this invention designs a standardized structured result output format, which returns image comparison conclusions, duplication types, similar image information, risk levels, and system recommendations based on a JSON structure. This structure has good readability, parsability, and interface integration capabilities.

[0045] The structured output format is as follows: { "is_duplicate": true, "match_type": "Duplicate across users", "similar_images": [ { "img_id": "img_124", "user_id": "U2034", "similarity_score": 0.92, "match_time": "2024-12-03T10:01:00" } ], "risk_level": "high", "suggestion": "Suspected image reuse; please conduct a manual review." } The meanings of each field are as follows: • "is_duplicate": Boolean, indicating whether the current image is recognized as a duplicate image by the system; • "match_type": String that identifies the type of image repetition. Possible values ​​include "cross-user repetition", "history reuse", "image evolution chain", etc. • "similar_images": An array containing a list of historical images that match the current image. Each element includes: o "img_id": A unique identifier for the matched image; o "user_id": The user to whom the image was uploaded; o "similarity_score": Image similarity score (e.g., cosine similarity); o "match_time": The timestamp of the first appearance of the image, used to track the image's source and reuse time; • "risk_level": The system's risk level assessment for the repeated behavior of this image, which can be set to "low", "medium", or "high"; • "suggestion": The system provides automated processing suggestions, such as "automatically return the item", "mark for manual review", "lock account", etc.

[0046] This structured output can serve as a standard communication protocol between the image risk control judgment module and the business system, adapting to API interfaces, MQ message streams, or database calls, enabling the structured, standardized, and automated processing of image fingerprint recognition results.

[0047] Based on the above embodiments, those skilled in the art can easily understand that the present invention also provides a credit material duplication detection device based on image fingerprints, which includes a storage medium and a processor. The storage medium stores a computer program, which is executed by the processor to implement the above-described method.

[0048] Example 1: A Credit Document Duplicate Detection System Based on Image Fingerprint (1) Application Background In credit risk control scenarios, customers often submit loan application materials (such as business licenses, ID cards, bank statements, etc.) through online channels. However, because these materials are easily forged, copied, or submitted repeatedly, some users often use methods such as image tampering, screenshot splicing, and cross-user sharing to launch fraud, causing huge credit losses.

[0049] Traditional document verification systems primarily rely on OCR extraction and rule-based validation, which struggles to identify semantically duplicated or non-pixel-level tampered images. Therefore, this embodiment, based on the image fingerprint fusion detection method proposed in this invention, constructs an image duplication detection system for credit scenarios. This system can identify image reuse risks, trace evolution chains, and support automated approval systems in determining the authenticity of documents.

[0050] (2) System input The system receives the following input: Customers can upload image materials (supports JPG, PNG, PDF screenshots); User identification information (user_id) and upload time; Historical image fingerprint database (including past credit image materials); Business context information (business type, application amount, business channel, etc.).

[0051] (3) System Flow Step 1: Perform the following processing on the uploaded image: Grayscale conversion; Trim edges; All sizes are uniformly scaled to the standard 128×128 size.

[0052] Output standard image This is used for subsequent feature extraction.

[0053] Step 2: Generate image fingerprints from two pathways: Pathway 1: Perceptual hash (pHash) → Obtain a 64-bit binary perceptual hash fingerprint ; Pathway 2: Deep Feature Extraction → Outputting 512D Deep Feature Vectors using ResNet-18 And normalize it.

[0054] Step 3: Retrieve the historical image that is closest to the input image from the database: judge ; judge .

[0055] Step 4: Combine user behavior, time, and other contextual dimensions to make judgments. See Table 1 for details: Table 1: Logic Table for Identifying Repeated Images ; Step 5: Structured Output The system will output the recognition results in a structured format with the following fields: { "Repeated conclusion": "Cross-user image reuse", Similar images:["img_3872","img_3105"], Image similarity: [0.94, 0.91], "Duplicate type": "Reuse across users", Risk Level: "High" System suggestion: "Automatically revert and mark as a black market risk". "Repeating chain path":["img_3105"->"img_3872"->"img_query"], Audit Tag: "Image fingerprints match, behavioral paths indicate high risk of reuse." }

[0056] (4) Deployment environment description The deployment environment required to implement this method can be found in Table 2: Table 2: Deployment Environment Table ; (5) Performance evaluation The performance of this method can be seen in Table 3: Table 3: Performance Evaluation Table ; The above description is merely a preferred embodiment of the present invention. It should be noted that for those skilled in the art, other parts not specifically described are existing technology or common knowledge. Several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting duplicate credit materials based on image fingerprints, characterized in that, include: Receive the original image of the credit materials uploaded by the user, and preprocess the original image; Image fingerprint data is extracted from the preprocessed image, the image fingerprint data including perceptual hash fingerprint and deep semantic vector; Based on the image fingerprint data of the original image and the fingerprint data of historical images in the historical database, the original image and the historical image are matched. If the fingerprint of the original image and the historical image are successfully matched, it is determined that the original image of the credit material uploaded by the user is a duplicate image. The system combines user identity and time series data to generate duplicate image detection results. The structured detection results are output based on a structured and interpretable output mechanism.

2. The method for detecting duplicate credit materials based on image fingerprints according to claim 1, characterized in that, The preprocessing method is as follows: Convert the user-uploaded raw image into a grayscale image; Perform edge cropping to obtain the main body region of the original image; The cropped images are uniformly scaled to a fixed size.

3. The method for detecting duplicate credit materials based on image fingerprints according to claim 2, characterized in that, The method for extracting the perceptual hash fingerprint is as follows: The grayscale matrix of the preprocessed image is then subjected to a two-dimensional discrete cosine transform, as follows: ; in, This is the result after the discrete cosine transform. The image after preprocessing. This is the discrete cosine transform; Extract the 8×8 low-frequency submatrix from the top left corner, as follows: ; in, To extract the 8×8 low-frequency submatrix from the top left corner; The mean of the low-frequency submatrix is ​​calculated as follows: ; in, The mean of the calculated low-frequency submatrix, Low-frequency sub-matrix The Middle One element; The generated perceptual hash fingerprint is as follows: ; in, The generated perceptual hash fingerprint.

4. The method for detecting duplicate credit materials based on image fingerprints according to claim 3, characterized in that, The deep semantic vector is extracted in the following way: Deep feature extraction is performed on the preprocessed image based on a deep neural network model, as follows: ; in, The deep feature vector is extracted from the preprocessed image. This represents the forward inference of a deep neural network model. It is a real number; For the deep feature vector Normalization is performed to obtain the deep semantic vector, as follows: ; in, For the extracted deep semantic vector, Deep feature vectors The L2 norm.

5. The method for detecting duplicate credit materials based on image fingerprints according to claim 4, characterized in that, The specific method for matching fingerprints between the original image and historical images is as follows: Hamming distance matching is performed based on the perceptual hash fingerprints of the original image and the historical images, as follows: ; in, Perceptual hash fingerprint of the calculated original image Perceptual hash fingerprints of historical images Hamming distance, These are the k-th perceptual hash fingerprints in the original image and the historical image, respectively; If the perceptual hash fingerprint of the original image Perceptual hash fingerprints of historical images Hamming distance This indicates that the structural fingerprint match was successful. The set Hamming distance threshold; The cosine similarity between the original image and historical images is calculated based on deep semantic vectors, as follows: ; in, To calculate the cosine similarity between the original image and the historical images, For the deep semantic vector of historical images, The dot product symbol for vectors; If the cosine similarity between the original image and the historical image This indicates that the semantic fingerprint match was successful. The set similarity threshold; When both the structural fingerprint and semantic fingerprint are successfully matched, it means that the fingerprints of the original image and the historical image have been successfully matched.

6. The method for duplicate detection of credit materials based on image fingerprints according to claim 5, characterized in that, The specific method for generating duplicate image judgment results by combining user identity and time series is as follows: Extract the unique identifier of the current user, the unique identifier of the user to which the historical image belongs, the upload timestamp of the original image, and the upload timestamp of the historical image. If the unique identifier of the current user is inconsistent with the unique identifier of the user to which the historical image belongs, it is determined to be cross-user image reuse; otherwise, it is determined to be repeated use of historical materials of the same user. If the cosine similarity between the original image and the historical image gradually decreases over time, an image fingerprint evolution path is constructed and determined to be an image tampering chain evolution behavior.

7. The method for detecting duplicate credit materials based on image fingerprints according to claim 5, characterized in that, The Hamming distance threshold The similarity threshold is 5. It is 0.

9.

8. The method for duplicate detection of credit materials based on image fingerprints according to claim 1, characterized in that, The structured detection results are based on JSON and include image comparison conclusions, duplication types, similar image information, risk levels, and recommendations.

9. A device for detecting duplicate credit materials based on image fingerprints, comprising a storage medium and a processor, wherein the storage medium stores a computer program, characterized in that, When the computer program is executed by a processor, it is used to implement the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Image-based credit approval method and device, electronic equipment and medium

    CN113706293A

  • Credit fraud behavior identification method and device, equipment and storage medium

    CN114707420A

  • Vehicle insurance claim settlement fraud identification method, device and equipment and storage medium

    CN116993517A

  • Anti-fraud system and anti-fraud method based on image technology

    CN117523369A

  • Insurance intelligent claim settlement operation auxiliary system and method

    CN117808603A