Inspection drawing behavior detection method and device, electronic equipment and storage medium
By combining geolocation verification, basic feature comparison, deep learning models, and text recognition technology, the problem of low efficiency in manual review of image collection detection has been solved, and multi-dimensional accurate identification and real-time response to image collection behavior have been achieved.
Patent Information
- Application Number
- CN202511094668.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, image set detection mainly relies on manual review, which is inefficient and has limited coverage, making it difficult to meet the regulatory needs of real-time requirements and massive data scenarios.
By combining geolocation verification, basic feature comparison, deep learning models, and text recognition technology, the system automatically identifies image copying behavior. Specific steps include: geolocation verification based on merchant information, extracting MD5 values for preliminary duplicate detection, extracting feature vectors using deep learning models and text recognition technology, and performing similarity comparison.
It enables multi-dimensional and accurate identification of image manipulation, expands the detection coverage, improves detection efficiency, reduces the risk of regulatory blind spots and delayed anomaly detection, and lowers the difficulty of risk management.
Smart Images

Figure CN120932079A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, electronic device and storage medium for detecting image overlay behavior during inspections. Background Technology
[0002] With the rapid development of digital technology, images, as an important carrier of information, are widely used in business processes across various industries. Verifying the authenticity and uniqueness of images has become a crucial step in ensuring business compliance and preventing fraud risks. Especially in scenarios where business processes need to be recorded through images (such as on-site inspections, qualification reviews, and scene verification), quickly and accurately identifying "image misuse" (i.e., the repeated or highly reused use of the same image for recording different objects or scenes) has become an important requirement for improving business supervision efficiency and ensuring information authenticity. Image misuse not only leads to distorted business data but may also cause process compliance issues, increasing operational risks and management costs for enterprises. Therefore, the demand for efficient image misuse detection technology is increasingly urgent.
[0003] In existing technologies, the detection of image set manipulation primarily relies on manual review. Manual review depends on professionals comparing each image individually, visually judging whether there are duplicate or highly similar image sets. However, in scenarios with a large number of images, manual review has significant limitations: firstly, the sampling rate is limited, making it difficult to cover all data and easily creating regulatory blind spots, resulting in some image set manipulation going undetected; secondly, manual judgment is inefficient and difficult to adapt to business scenarios with high real-time requirements, potentially leading to delays in the detection of anomalies and increasing the difficulty of risk management. Summary of the Invention
[0004] In view of this, this application provides a method, device, electronic device and storage medium for detecting image hijacking behavior, which can accurately identify potential image hijacking behavior through comprehensive analysis of geographical location and multimodal features.
[0005] According to a first aspect of this application, a method for detecting image overlay behavior during inspections is provided, comprising:
[0006] Based on the merchant information associated with the inspection images to be detected, the geographical location of the images to be detected is verified.
[0007] If it is determined that the inspection image to be detected passes the geographical location verification, then the basic features of the inspection image to be detected are extracted, and preliminary repeat detection is performed based on the basic features;
[0008] If it is determined that the inspection image to be detected passes the preliminary repeated detection, then the image feature vector of the inspection image to be detected is extracted based on the deep learning model, and the text feature vector in the inspection image to be detected is extracted by combining text recognition technology.
[0009] The image feature vector and the text feature vector are compared with the feature vectors of historical inspection images stored in the vector database. Based on the similarity comparison results, it is determined whether the inspection image to be detected involves image overlay.
[0010] According to a second aspect of this application, a device for detecting image overlay behavior during inspections is provided, comprising:
[0011] The verification module is used to verify the geographical location of the inspection image to be detected based on the merchant information associated with the image.
[0012] The detection module is used to extract the basic features of the inspection image to be detected if it is determined that the inspection image to be detected passes the geographical location verification, and to perform preliminary repeat detection based on the basic features;
[0013] The extraction module is used to extract the image feature vector of the inspection image to be detected based on a deep learning model and extract the text feature vector of the inspection image to be detected by combining text recognition technology if it is determined that the inspection image to be detected has passed the preliminary repeated detection.
[0014] The judgment module is used to compare the similarity of the image feature vector and the text feature vector with the feature vectors of historical inspection images stored in the vector database, and to determine whether the inspection image to be detected has any image overlay behavior based on the similarity comparison result.
[0015] According to a third aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described method for detecting overlapping images during inspection.
[0016] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for detecting overlapping images during inspection.
[0017] By employing the aforementioned technical solutions, the inspection image matching detection method, device, electronic equipment, and storage medium provided in this application, through a multi-layered automated detection process, can effectively address the limitations of manual review in image matching detection. Specifically, this is reflected in the following aspects: First, based on the geographical location verification of merchant information, the location of inspection images is automatically verified, quickly eliminating image matching behaviors that clearly do not conform to geographical location logic without manual intervention, thus expanding the detection coverage. Second, by extracting basic features such as MD5 values for preliminary duplicate detection, completely duplicate image matching can be automatically screened, replacing the inefficient method of manual comparison one by one, thereby improving detection efficiency. Third, by combining image feature vectors extracted by deep learning models with text feature vectors extracted by text recognition technology, multi-dimensional similarity comparison is performed on images that pass the preliminary detection. Utilizing a vector database, efficient feature matching is achieved from massive historical data, which not only avoids the limitations of manual sampling to cover all data but also meets real-time detection requirements through an automated process. This reduces regulatory blind spots and delays in anomaly detection, thereby lowering the difficulty of risk management.
[0018] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0019] Figure 1 A flowchart illustrating an inspection image overlay behavior detection method provided in an embodiment of this application is shown.
[0020] Figure 2 A flowchart illustrating a method for detecting inspection overlay behavior according to another embodiment of this application is shown;
[0021] Figure 3 A schematic diagram of the structure of an inspection mapping behavior detection device provided in an embodiment of this application is shown. Detailed Implementation
[0022] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0023] In existing technologies, the detection of image set manipulation primarily relies on manual review. Manual review depends on professionals comparing each image individually, visually judging whether there are duplicate or highly similar image sets. However, in scenarios with a large number of images, manual review has significant limitations: firstly, the sampling rate is limited, making it difficult to cover all data and easily creating regulatory blind spots, resulting in some image set manipulation going undetected; secondly, manual judgment is inefficient and difficult to adapt to business scenarios with high real-time requirements, potentially leading to delays in the detection of anomalies and increasing the difficulty of risk management.
[0024] To address the aforementioned technical problems, embodiments of the present invention provide a method for detecting image overlay behavior during inspections, such as... Figure 1 As shown, the method includes:
[0025] Step 110: Based on the merchant information associated with the inspection image to be inspected, perform geolocation verification on the location where the inspection image to be inspected was taken.
[0026] The inspection images to be tested refer to photos taken by inspectors during their inspections of merchants, which are used to verify whether there is any misuse of images. These include storefront photos, interior photos, business license photos, transaction receipts, etc., used to record the actual operating status of the merchants. The associated merchant information refers to the relevant information registered in the system for the merchants corresponding to the inspection images to be tested. The core geographic information includes the geographic location information of the merchant's registered operating area (such as latitude and longitude range, geographic boundaries corresponding to the specific address, etc.), used to define the legal operating space of the merchant. The shooting location refers to the location data (usually latitude and longitude coordinates) recorded by the shooting device (such as a mobile phone or camera) through GPS or other positioning functions when the inspection images to be tested were taken, reflecting the actual shooting location of the image. Geographic location verification refers to the process of comparing the shooting location information of the inspection images to be tested with the geographic location information of the merchant's registered operating area through technical means (such as electronic fence technology) to determine whether the shooting location is within the merchant's registered operating area.
[0027] In this embodiment of the disclosure, the location of the image being photographed can be checked against the merchant information associated with the image to be inspected to determine whether the location meets the merchant's geographical requirements. Specifically, the information registered by the merchant in the system (such as the geographical location of the registered business area) is compared with the location information recorded when the image was photographed to verify whether the location of the image is within the reasonable geographical range corresponding to the merchant.
[0028] Automated geolocation verification eliminates the need for manual checking, quickly identifying images whose shooting locations do not match the registered operating area of the merchant, thus providing an initial spatial basis for identifying image misuse. This process reduces the need for manual review, avoids regulatory blind spots caused by the limited coverage of manual spot checks, and significantly improves detection efficiency. Verification can be completed instantly after image upload, promptly detecting image misuse with abnormal geolocations, laying the foundation for more accurate subsequent detection steps, and effectively reducing the risk of distorted inspection records due to image misuse.
[0029] Step 120: If the inspection image to be detected is determined to pass the geographical location verification, then extract the basic features of the inspection image to be detected, and perform preliminary repeat detection based on the basic features.
[0030] Among them, the basic features are feature values used to identify the uniqueness of an image, such as the MD5 value of the image. The MD5 value is a unique feature value calculated from image data using a hash algorithm, which can be used to quickly identify the uniqueness of an image (identical images will generate the same MD5 value); the preliminary duplicate detection refers to the process of comparing the basic features (such as the MD5 value) of the image to be inspected with the basic features of historical inspected images to determine whether the image is a set of images that are completely reused. It is the second verification of the image set detection.
[0031] In this embodiment of the disclosure, once the location of the inspection image to be detected has been verified and confirmed to be within the registered business area of the merchant, the basic features of the image (such as feature values used to identify the uniqueness of the image) can be further extracted. By comparing these basic features with the basic features of historical inspection images, it can be preliminarily determined whether the image is a duplicate set of images. This is the next step in the detection process after the geographical location verification is passed, used to quickly screen for completely duplicate images.
[0032] Based on successful geolocation verification, preliminary duplicate detection is performed by extracting basic features, enabling tiered filtering of image duplication. On one hand, completely duplicate images can be quickly identified through basic feature comparison, eliminating the need for more complex feature extraction steps and significantly reducing subsequent processing workload, thus improving overall detection efficiency. On the other hand, this preliminary duplicate detection requires no manual intervention and provides comprehensive coverage of all images that have passed geolocation verification, avoiding potential omissions in manual sampling. This reduces the risk of undetected image duplication at the fundamental level and provides pre-filtering support for more accurate deep detection.
[0033] Step 130: If the inspection image to be detected passes the initial repeated detection, then the image feature vector of the inspection image to be detected is extracted based on the deep learning model, and the text feature vector in the inspection image to be detected is extracted by combining text recognition technology.
[0034] Among them, image feature vector refers to a set of values obtained after processing the image to be inspected through a deep learning model. It can quantify the visual features (such as color, texture, shape, scene layout, etc.) of the image to be inspected and is used to measure the similarity between images. Text feature vector refers to a set of values obtained after processing the text information in the image to be inspected through text recognition technology. It can quantify the semantic and content features of the text and is used to supplement the shortcomings of image feature vector in text image comparison.
[0035] In this embodiment of the disclosure, after the image to be inspected passes the initial duplication detection (i.e., no complete duplication with historical images is found), a deep learning model can be further used to analyze the visual content of the image and extract vectors that can characterize the visual features of the image. Simultaneously, text recognition technology is used to identify the text information contained in the image and convert this text information into text feature vectors that can be used for comparison. This step involves a deeper level of feature extraction from the image to identify image sets that are not completely duplicated but are highly similar.
[0036] Building upon initial duplicate detection, deeper learning models and text recognition techniques are used to extract more refined feature vectors. This effectively compensates for the limitations of relying solely on basic features (such as MD5 values) to identify image-swapping behavior that has undergone minor manipulations (such as scaling, rotation, and text modification). On one hand, image feature vectors can capture both global and local visual features of images, making them suitable for similarity assessment of scene-based images. On the other hand, text feature vectors can extract content features from images containing text (such as business licenses and transaction vouchers), addressing the issue of misjudgment of images with similar layouts but different text. The combination of these two approaches enables multi-dimensional and accurate identification of image-swapping behavior, reducing missed and false positives. Furthermore, the entire process is automated, requiring no manual intervention, which improves the accuracy and efficiency of image-swapping detection in scenarios with massive amounts of images, further strengthening the ability to monitor the authenticity of inspection records.
[0037] Step 140: Compare the image feature vector and text feature vector with the feature vectors of historical inspection images stored in the vector database, and determine whether the inspection images to be detected contain image overlays based on the similarity comparison results.
[0038] Among them, the feature vectors of historical inspection images refer to the image feature vectors and text feature vectors obtained after feature extraction from images taken during past inspections, which are stored in the vector database as benchmark data for comparison; similarity comparison refers to the process of calculating the similarity between the feature vector of the inspection image to be detected and the feature vector of historical inspection images. Commonly used methods include cosine similarity and Euclidean distance, which are used to determine whether two images are highly similar; image overlay refers to the behavior in merchant inspection work where some inspectors, in order to complete their tasks, reuse all or part of the inspection images of a certain merchant in the inspection records of other merchants under their responsibility, without carrying out actual inspection work.
[0039] In this embodiment of the disclosure, the image feature vector of the inspection image to be detected extracted by the deep learning model and the text feature vector extracted by the text recognition technology can be compared with the image feature vector and text feature vector of the historical inspection images pre-stored in the vector database to calculate the degree of similarity between them; and then, based on the result of the degree of similarity, it can be determined whether the inspection image to be detected belongs to a set of reused images.
[0040] By comparing the multimodal feature vectors of the images to be detected with historical data, accurate identification of image-swapping behavior can be achieved. On the one hand, the vector database can efficiently store and retrieve massive feature vectors, solving the problem of low efficiency of manual comparison in scenarios with massive data, and ensuring that the detection process can respond in real time. On the other hand, by combining the dual comparison of image and text features, it can not only identify visually highly similar image-swapping (such as images of the same scene taken from different angles), but also distinguish images with similar layouts but different text content (such as business licenses of different merchants), effectively balancing the accuracy and recall of detection, reducing subjective errors and regulatory blind spots in manual review, and strengthening the control over the authenticity of inspection records.
[0041] In summary, the inspection image matching detection method provided by this invention, through a multi-level automated detection process, effectively addresses the limitations of manual review in image matching detection. Specifically: First, based on the geographical location verification of merchant information, the location of inspection images is automatically verified, quickly eliminating image matching behaviors that clearly do not conform to geographical location logic without manual intervention, thus expanding the detection coverage. Second, by extracting basic features such as MD5 values for preliminary duplicate detection, completely duplicate image matching can be automatically screened, replacing the inefficient method of manual comparison and improving detection efficiency. Third, by combining image feature vectors extracted by deep learning models with text feature vectors extracted by text recognition technology, multi-dimensional similarity comparison is performed on images that pass the preliminary detection. Utilizing a vector database, efficient feature matching is achieved from massive historical data, avoiding the limitations of manual sampling to cover all data, and meeting real-time detection requirements through an automated process. This reduces regulatory blind spots and delays in anomaly detection, thereby lowering the difficulty of risk management.
[0042] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, and to fully illustrate the implementation of this embodiment, this embodiment also provides another method for detecting image overlay behavior during inspections, such as... Figure 2 As shown, the method includes:
[0043] Step 210: Based on the merchant information associated with the inspection image to be inspected, perform geolocation verification on the shooting location of the inspection image to be inspected.
[0044] The associated merchant information includes the geographical location of the merchant's registered business area, and the inspection images to be tested also include the location where they were taken.
[0045] In this embodiment of the disclosure, the geographical location information of the merchant's registered business area associated with the inspection image to be detected (such as the latitude and longitude range corresponding to the business address registered by the merchant in the system) can be combined with the shooting location information of the image itself (such as GPS coordinates obtained through the shooting device) to delineate the registered business area of the merchant through electronic fence technology, and verify whether the actual shooting location of the image is within the area. If the shooting location is within the area, the verification is passed; if not, it can be directly determined as image overlay behavior, and the subsequent judgment process will not continue. Alternatively, a prompt message indicating that the inspection image to be detected has image overlay behavior can be output, and the user instruction on whether to continue to execute the subsequent steps can be received.
[0046] Accordingly, step 210 of the embodiment may specifically include the following steps:
[0047] Step 210-1: Based on geographic location information and shooting location information, verify whether the shooting location of the inspection image to be detected is within the registered business area of the merchant through electronic fence technology.
[0048] Among them, electronic fence technology is a spatial delimitation technology based on geographic information. By setting virtual geographic boundaries (such as a specific radius range centered on the merchant's registered address), it automatically determines whether the target location (such as the location where the picture was taken) is within the boundary, thereby realizing automated verification of the spatial range.
[0049] Step 210-2: If yes, then determine that the inspection image to be detected has passed the geographical location verification.
[0050] Step 210-3: If not, output a prompt message indicating that the inspection images to be detected contain image overlay behavior.
[0051] The alert message refers to the warning content generated for inspection images that are determined to be images used in a duplicate order. It is used to notify relevant personnel to conduct subsequent verification and ensure the authenticity of the inspection records.
[0052] When verification using geofencing technology reveals that the location of the image to be inspected is outside the registered operating area of a merchant, an automatic alert can be generated indicating that the image has been used in a duplicate order. This step, through automated judgment and immediate alerts, replaces the subjective judgment of images with abnormal geographical locations during manual review. It can immediately identify duplicate image behavior that does not conform to spatial logic after the image is uploaded. This not only avoids regulatory blind spots that may be missed by manual spot checks, but also improves the efficiency of detecting anomalies, reduces the risk of distorted inspection records due to duplicate image behavior, and provides timely evidence for subsequent verification and rectification.
[0053] Step 220: If the inspection image to be detected is determined to pass the geographical location verification, then extract the basic features of the inspection image to be detected, and perform preliminary repeat detection based on the basic features.
[0054] Among them, the basic features include at least the MD5 value.
[0055] In this embodiment of the disclosure, if it is determined that the inspection image to be detected has passed the geolocation verification, its basic features can be extracted and a preliminary duplicate detection can be performed based on the basic features. This step further screens for image overlay behavior through technical means under the premise of geolocation compliance: Specifically, the MD5 value of the inspection image to be detected (as a basic feature) can be calculated and compared with the MD5 value of historical inspection images stored in the preset database. If there is no identical MD5 value, it is determined that the preliminary duplicate detection has passed and the subsequent deep detection stage is entered; if there is an identical MD5 value, it is directly determined to be image overlay behavior, and the judgment process of the subsequent steps is no longer executed, or a prompt message indicating that the inspection image to be detected has image overlay behavior can be output, waiting to receive user instructions on whether to continue to execute the subsequent steps.
[0056] This process uses automated MD5 value comparison to quickly identify completely duplicate image sets. It avoids the inefficiency caused by the large number of images in manual review, and also makes up for the limited coverage of sampling methods. It can comprehensively screen all images verified by geolocation, reducing the risk of image set behavior going undetected. At the same time, it reduces the load for subsequent more refined feature extraction and detection, improving the overall efficiency and accuracy of image set detection.
[0057] Accordingly, step 220 of the embodiment may specifically include the following steps:
[0058] Step 220-1: Calculate the MD5 value of the inspection image to be tested.
[0059] In this embodiment of the disclosure, the binary data of the image to be inspected can be calculated using the MD5 hash algorithm to generate a unique 32-bit string (i.e., MD5 value). Since identical images have identical binary data, their corresponding MD5 values are also identical. However, images that have been modified (even slightly) will generate different MD5 values. Therefore, the MD5 value can be used as a "digital fingerprint" of an image to quickly identify completely duplicate images.
[0060] Step 220-2: Compare the MD5 value with the MD5 value of the historical inspection images stored in the preset database.
[0061] Step 220-3: If no identical MD5 value is found, the inspection image to be tested is deemed to have passed the preliminary duplicate detection.
[0062] Step 220-4: If the same MD5 value exists, output a prompt message indicating that the images to be inspected have been overlaid.
[0063] Step 230: If it is determined that the inspection image to be detected has passed the preliminary repeated detection, then the inspection image to be detected is scaled to at least two different scales, and combined with the inspection image to be detected at the original image scale to form a multi-scale image set.
[0064] In this embodiment of the disclosure, if it is determined that the inspection image to be detected has passed the initial repeated detection, it can be scaled up to at least two different scales (such as the original inspection image to be detected). The images are then multi-scaled (1x and 1 / 2x) and combined with the original images to be inspected to form a multi-scale image set. This step is to extract image features more comprehensively. By generating images at different resolutions, the impact of image scale differences caused by shooting distance and angle on feature extraction can be mitigated, allowing subsequent deep learning models to capture richer visual information (such as global scene and local details).
[0065] This multi-scale processing approach can improve the model's adaptability to images of different scales, thereby more robustly identifying similar images that have undergone scaling and other processing in image set detection, reducing missed detections caused by scale changes, and providing a more reliable foundation for subsequent feature fusion and comparison.
[0066] Step 240: Use the optimized ResNet101 model to extract features from each image in the multi-scale image set, and fuse multiple feature vectors corresponding to the multi-scale image set to obtain the image feature vector of the image to be inspected.
[0067] The optimized ResNet101 model removes the fully connected layers and adopts the GeM pooling method.
[0068] In this embodiment of the disclosure, for a multi-scale image set consisting of images at the original scale and at least two scaled dimensions, a modified ResNet101 model can be used to extract deep visual features from each image, resulting in multiple feature vectors. These feature vectors are then fused using GeM pooling to ultimately form the image feature vector of the image to be inspected. This process, by preserving the model's backbone network and optimizing the pooling method, enhances the ability to capture information from key regions of the image. Furthermore, by combining feature fusion from multiple scale images, it improves adaptability to images of different resolutions.
[0069] The optimized ResNet101 model, by removing fully connected layers and employing GeM pooling, avoids the limitations imposed by fully connected layers on feature extraction, better preserving global and local features of the image and improving the representational power of the feature vectors. Furthermore, feature fusion from multi-scale image sets effectively mitigates the impact of image scale differences caused by shooting distance and angle on detection, enabling the model to more robustly extract features from images of different scales and reducing feature loss due to image scaling and resolution changes. Image feature vectors obtained in this way, when compared with historical image feature vectors, can more accurately identify image set detection that has undergone scaling or other processing, thus improving the accuracy of image set detection.
[0070] Step 250: Extract text feature vectors from the inspection images to be detected using text recognition technology.
[0071] In this embodiment of the disclosure, textual information (such as the merchant name on a business license, the amount on a transaction voucher, etc.) contained in the inspection image to be detected can first be identified using OCR (Optical Character Recognition) technology. This textual information is then input into a natural language processing model (such as a BERT model). The natural language processing model performs semantic analysis and feature extraction on the textual information in the inspection image to be detected, ultimately generating a text feature vector that can quantitatively represent the textual content. This process can overcome the shortcomings of simply relying on image features to detect images containing text (such as form-type images). By characterizing the textual content, richer judgment criteria can be provided for image set detection.
[0072] Accordingly, step 250 of the embodiment may specifically include the following steps:
[0073] Step 250-1: Use OCR technology to identify the text information in the inspection image to be inspected.
[0074] In this embodiment of the disclosure, optical character recognition (OCR) technology can be used to automatically recognize printed text (such as the merchant name and registration number on a business license, or the amount and date on a transaction voucher) contained in the inspection image to be inspected, converting the text content in the inspection image into an editable text format. This process can extract key text information from the inspection image containing text, providing basic data for subsequent generation of text feature vectors through a natural language processing model.
[0075] Step 250-2: Process the text information using the BERT model and extract the text feature vector.
[0076] In this embodiment of the disclosure, textual information (such as the merchant name on a business license, information on a transaction voucher, etc.) in the image to be inspected, identified by OCR technology, can be input into the BERT model. Utilizing the BERT model's ability to understand the semantic context of text, the textual information is transformed into vectors that can quantify its semantic and content features. This process can deeply uncover the underlying meaning of the textual information, providing accurate feature basis for subsequent similarity comparison with the text feature vectors of historical images. It is particularly suitable for distinguishing form-type images with similar layouts but different textual content.
[0077] Step 260: Compare the image feature vector and text feature vector with the feature vectors of historical inspection images stored in the vector database, and determine whether the inspection images to be detected contain image overlays based on the similarity comparison results.
[0078] For the embodiments of this disclosure, step 260 may specifically include the following steps:
[0079] Step 260-1: Based on image feature vectors, retrieve historical inspection images from the vector database whose image similarity to the inspection image to be detected is greater than a first preset threshold, and construct a candidate image set.
[0080] The vector database is a distributed vector database, Milvus, used to store the image feature vectors and text feature vectors of each historical inspection image, as well as the corresponding retrieval index. The retrieval index is used to retrieve candidate image sets whose image similarity to the inspection image to be detected is greater than a first preset threshold, and to retrieve the text feature vector of each candidate image in the candidate image set. The first preset threshold refers to a pre-set image similarity threshold. When the image similarity between the inspection image and the historical image exceeds this value, it is determined that the two are highly similar in visual features and are included in the candidate image set.
[0081] In this embodiment of the disclosure, based on the image feature vector of the inspection image to be detected, the retrieval index (such as LSH, ANN and other indexing technologies) in the distributed vector database Milvus can be used to quickly filter out historical images in the stored historical inspection images that have an image similarity to the inspection image to be detected that exceeds a first preset threshold, thus forming a candidate image set.
[0082] By optimizing similarity calculation through retrieval indexing techniques (such as LSH and ANN), candidate images with similar image features can be located from a large amount of historical data in a short time, significantly improving retrieval speed and meeting the needs of real-time image set detection. At the same time, the centralized construction of candidate image sets can avoid redundant calculations that compare all historical images one by one, reduce system load, and lay the foundation for accurate judgment by combining text features, thus achieving a balance between efficiency and accuracy in image set detection.
[0083] Step 260-2: Based on the text feature vector, calculate the text similarity between the image to be detected and each candidate image in the candidate image set using cosine similarity or Euclidean distance.
[0084] In this embodiment of the disclosure, the text feature vectors of the image to be inspected and each image in the candidate image set can be used to quantify the similarity between the two in terms of text content using algorithms such as cosine similarity or Euclidean distance. This process further compares the candidate images selected by image feature filtering from the dimension of text content, which can effectively distinguish images with similar visual features but different text content (such as business licenses of different merchants), providing a more accurate basis for the final determination of image matching behavior.
[0085] Step 260-3: If at least one candidate image in the candidate image set has a text similarity greater than the second preset threshold, then it is determined that the image to be detected has image matching behavior.
[0086] The second preset threshold refers to a pre-set text similarity threshold. When the text similarity between the image to be detected and a candidate image exceeds this threshold, the two are determined to be highly similar in text content.
[0087] Step 260-4: If it is determined that there are no candidate images in the candidate image set with a text similarity greater than the second preset threshold, then it is determined that the images to be inspected do not have image overlay behavior.
[0088] In summary, the technical solution of this application, by constructing a multi-level automated detection process, can specifically address the limitations of manual review in existing technologies. Specifically: First, by leveraging electronic fence technology, based on the geographical location information of the merchant's registered operating area and the shooting location information of the inspection images, it automatically verifies whether the shooting location is within a compliant area. This allows for rapid screening of image duplication due to inconsistent geographical locations, achieving preliminary filtering without manual intervention, thus expanding the detection coverage and avoiding regulatory blind spots inherent in manual sampling. Second, by calculating the MD5 value and comparing it with historical data, it automatically identifies completely duplicate images. This method significantly improves detection efficiency and solves the problem that manual review is difficult to adapt to massive data scenarios. Furthermore, through multi-scale image processing, image feature extraction using an optimized ResNet101 model, and text feature extraction using OCR and BERT models, combined with the efficient retrieval capabilities of the Milvus vector database, it can accurately identify complex image set behaviors (such as similar images that have been scaled, rotated, etc.). The entire process requires no manual intervention, can respond to detection needs in real time, avoids the problem of delayed detection of anomalies, effectively reduces the difficulty of risk management, and comprehensively improves the accuracy and timeliness of image set detection.
[0089] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a device for detecting image overlay behavior during inspections, such as... Figure 3 As shown, the device includes: a verification module 31, a detection module 32, an extraction module 33, and a judgment module 34.
[0090] Verification module 31 is used to verify the geographic location of the inspection image based on the merchant information associated with the inspection image to be inspected;
[0091] The detection module 32 is used to extract the basic features of the inspection image to be detected if it is determined that the inspection image to be detected has passed the geographical location verification, and to perform preliminary repeat detection based on the basic features.
[0092] The extraction module 33 is used to extract the image feature vector of the inspection image based on the deep learning model and extract the text feature vector in the inspection image based on the text recognition technology if it is determined that the inspection image to be detected has passed the preliminary repeated detection.
[0093] The judgment module 34 is used to compare the similarity of the image feature vector and text feature vector with the feature vector of the historical inspection images stored in the vector database, and to determine whether the inspection image to be detected has any image matching behavior based on the similarity comparison result.
[0094] In some embodiments of this application, the associated merchant information includes the geographical location information of the merchant's registered business area, and the inspection image to be detected is also associated with the shooting location information. The verification module 31 can be used to verify, based on the geographical location information and the shooting location information, whether the shooting location of the inspection image to be detected is located within the merchant's registered business area using electronic fence technology; if yes, it is determined that the inspection image to be detected has passed the geographical location verification; if no, a prompt message indicating that the inspection image to be detected has image overlay behavior is output.
[0095] In some embodiments of this application, the basic features include at least an MD5 value. The detection module 32 is specifically used to calculate the MD5 value of the inspection image to be detected; compare the MD5 value with the MD5 values of historical inspection images stored in a preset database; if there is no identical MD5 value, it is determined that the inspection image to be detected has passed the preliminary duplicate detection; if there is an identical MD5 value, a prompt message indicating that the inspection image to be detected has image overlay behavior is output.
[0096] In some embodiments of this application, when extracting the image feature vector of the inspection image to be detected based on a deep learning model, the extraction module 33 can be specifically used to scale the inspection image to be detected to at least two different scales, forming a multi-scale image set with the inspection image to be detected at the original image scale; the optimized ResNet101 model is used to extract features from each image in the multi-scale image set, the optimized ResNet101 model removes the fully connected layer and uses the GeM pooling method; multiple feature vectors corresponding to the multi-scale image set are fused to obtain the image feature vector of the inspection image to be detected.
[0097] In some embodiments of this application, when extracting text feature vectors from the inspection image to be detected using text recognition technology, the extraction module 33 can be specifically used to recognize text information in the inspection image to be detected using OCR technology; and to process the text information using the BERT model to extract text feature vectors.
[0098] In some embodiments of this application, the judgment module 34 can be specifically used to retrieve historical inspection images from the vector database based on image feature vectors, whose image similarity to the inspection image to be detected is greater than a first preset threshold, and construct a candidate image set; based on text feature vectors, calculate the text similarity between the inspection image to be detected and each candidate image in the candidate image set using cosine similarity or Euclidean distance calculation methods; if at least one candidate image in the candidate image set has a text similarity greater than a second preset threshold, it is determined that the inspection image to be detected has image overlay behavior; if it is determined that there is no candidate image in the candidate image set with a text similarity greater than the second preset threshold, it is determined that the inspection image to be detected does not have image overlay behavior.
[0099] In some embodiments of this application, the vector database is a distributed vector database Milvus, which is used to store the image feature vectors and text feature vectors of each historical inspection image and the corresponding retrieval index. The retrieval index is used to retrieve, in the vector database, a set of candidate images whose image similarity to the inspection image to be detected is greater than a first preset threshold, and to retrieve the text feature vector of each candidate image in the candidate image set.
[0100] It should be noted that other corresponding descriptions of the functional units involved in the inspection map overlay behavior detection device provided in this embodiment can be found in the following references. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.
[0101] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method for detecting inspection map overlay behavior is shown.
[0102] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0103] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method for detecting inspection map overlay behavior is shown.
[0104] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0105] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0106] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0108] This invention addresses the limitations of manual review in existing technologies by constructing a multi-level automated detection process. Specifically: First, leveraging electronic fence technology, it automatically verifies whether the shooting location is within a compliant area based on the geographical location information of the merchant's registered operating area and the shooting location information of the images to be inspected. This allows for rapid screening of image mismatches based on geographical location, achieving initial filtering without manual intervention, expanding the detection coverage and avoiding regulatory blind spots in manual sampling. Second, by calculating MD5 values and comparing them with historical data, it automatically identifies completely duplicate images, significantly improving detection efficiency and solving the problem of manual review being unable to adapt to massive data scenarios. Third, through multi-scale image processing, image feature extraction using an optimized ResNet101 model, text feature extraction using OCR and BERT models, and the efficient retrieval capabilities of the Milvus vector database, it can accurately identify complex image mismatch behaviors (such as similar images after scaling and rotation). The entire process requires no manual intervention, responding to detection needs in real time, avoiding delays in anomaly detection, effectively reducing risk management difficulty, and comprehensively improving the accuracy and timeliness of image mismatch detection.
[0109] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0110] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for detecting inspection map overlay behavior, characterized in that, include: Based on the merchant information associated with the inspection images to be detected, the geographical location of the images to be detected is verified. If it is determined that the inspection image to be detected passes the geographical location verification, then the basic features of the inspection image to be detected are extracted, and preliminary repeat detection is performed based on the basic features; If it is determined that the inspection image to be detected passes the preliminary repeated detection, then the image feature vector of the inspection image to be detected is extracted based on the deep learning model, and the text feature vector in the inspection image to be detected is extracted by combining text recognition technology. The image feature vector and the text feature vector are compared with the feature vectors of historical inspection images stored in the vector database. Based on the similarity comparison results, it is determined whether the inspection image to be detected involves image overlay.
2. The method according to claim 1, characterized in that, The associated merchant information includes the geographical location information of the merchant's registered operating area. The inspection image to be tested is also associated with the shooting location information. Based on the merchant information associated with the inspection image to be tested, the geographical location verification of the shooting location of the inspection image to be tested is performed, including: Based on the geographic location information and the shooting location information, the electronic fence technology is used to verify whether the shooting location of the inspection image to be detected is within the registered business area of the merchant; If so, then it is determined that the inspection image to be detected has passed the geographical location verification; If not, output a message indicating that the image to be inspected contains overlayed images.
3. The method according to claim 1, characterized in that, The basic features include at least the MD5 value. The basic features of the inspection image to be detected are extracted, and preliminary duplicate detection is performed based on these basic features, including: Calculate the MD5 value of the inspection image to be detected; The MD5 value is compared with the MD5 value of historical inspection images stored in the preset database; If no identical MD5 values are found, the inspection image to be detected is determined to have passed the preliminary duplicate detection. If the same MD5 value exists, a prompt message will be output indicating that the images to be inspected have been overlaid.
4. The method according to claim 1, characterized in that, The step of extracting the image feature vector of the inspection image to be detected based on the deep learning model includes: The inspection image to be detected is scaled up to at least two different scales and combined with the inspection image to be detected at the original image scale to form a multi-scale image set. The optimized ResNet101 model is used to extract features from each image in the multi-scale image set. The optimized ResNet101 model removes the fully connected layers and uses the GeM pooling method. By fusing multiple feature vectors corresponding to the multi-scale image set, the image feature vector of the inspection image to be detected is obtained.
5. The method according to claim 1, characterized in that, Text feature vectors are extracted from the inspection images to be detected using text recognition technology, including: The text information in the inspection image to be detected is identified using OCR technology; The text information is processed using the BERT model to extract text feature vectors.
6. The method according to claim 1, characterized in that, The image feature vector and the text feature vector are compared with the feature vectors of historical inspection images stored in the vector database. Based on the similarity comparison results, it is determined whether the inspection image to be detected exhibits image overlay behavior, including: Based on the image feature vector, historical inspection images with an image similarity greater than a first preset threshold to the inspection image to be detected are retrieved from the vector database to construct a candidate image set. Based on the text feature vector, the text similarity between the image to be detected and each candidate image in the candidate image set is calculated using cosine similarity or Euclidean distance. If at least one candidate image in the candidate image set has a text similarity greater than the second preset threshold, then it is determined that the image to be detected has image matching behavior. If it is determined that there is no candidate image in the candidate image set whose text similarity is greater than the second preset threshold, then it is determined that the image to be detected does not have image overlay behavior.
7. The method according to claim 6, characterized in that, The vector database is a distributed vector database, Milvus, used to store image feature vectors and text feature vectors of each historical inspection image, as well as corresponding search indexes. The search indexes are used to retrieve candidate image sets whose image similarity to the inspection image to be detected is greater than a first preset threshold, and to retrieve the text feature vector of each candidate image in the candidate image set.
8. A device for detecting inspection map overlay behavior, characterized in that, include: The verification module is used to verify the geographical location of the inspection image to be detected based on the merchant information associated with the image. The detection module is used to extract the basic features of the inspection image to be detected if it is determined that the inspection image to be detected passes the geographical location verification, and to perform preliminary repeat detection based on the basic features; The extraction module is used to extract the image feature vector of the inspection image to be detected based on a deep learning model and extract the text feature vector of the inspection image to be detected by combining text recognition technology if it is determined that the inspection image to be detected has passed the preliminary repeated detection. The judgment module is used to compare the similarity of the image feature vector and the text feature vector with the feature vectors of historical inspection images stored in the vector database, and to determine whether the inspection image to be detected has any image overlay behavior based on the similarity comparison result.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.