Repeated image recognition method and device based on locality sensitive hashing algorithm, and medium

Through a duplicate image recognition method based on the locally sensitive hashing algorithm, a deep learning model is used to extract image features and map them to hash values, which solves the problem of inefficiency in identifying duplicate images in the fields of financial technology and medical health and elderly care, and achieves fast and accurate image recognition.

CN120747556APending Publication Date: 2025-10-03CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510826517.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the fields of financial technology and healthcare and elderly care, existing technologies are unable to efficiently and accurately identify duplicate images in images, resulting in inefficient and error-prone manual review, huge consumption of computing resources, and inability to meet actual application needs.

Method used

A repeated image recognition method based on the local sensitive hashing algorithm is adopted. The image visual features are extracted through a deep learning model and mapped into feature hash values. The hash value similarity is calculated to quickly identify similar images.

Benefits of technology

It improves the efficiency of the image recognition system in identifying duplicate images, simplifies the amount of data processing, can quickly identify similar images in the hash space, and improves recognition efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747556A_ABST
    Figure CN120747556A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image detection, and discloses a repeated image recognition method and device based on a locality sensitive hash algorithm, equipment and a medium, and the method comprises the steps: mapping a to-be-recognized visual feature into a to-be-recognized feature hash value through the locality sensitive hash algorithm; calculating the similarity between the preset feature hash value and the to-be-recognized feature hash value, and determining a similar image; and generating a repeated image recognition result through an image recognition algorithm, the similar image and the to-be-recognized image. Through the above mode, the key information of the image is accurately acquired by using the deep learning model to serve as the visual feature of the to-be-identified image, the visual feature is mapped into the feature hash value by using the preset locality sensitive hash algorithm, and the similarity between different images is determined through different feature hash values. And the similar images can be quickly identified in the hash space. The method can be applied to the fields of financial science and technology, medical health, old-age care and the like, and the efficiency of recognizing repeated images by an image recognition system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image detection technology, and in particular to a method, device, equipment and medium for repeated image recognition based on a local sensitive hashing algorithm. Background Art

[0002] With the rapid development of artificial intelligence (AI), it has found widespread application in image detection. Fields such as fintech, healthcare, and senior care all face the challenge of accurately identifying similar or duplicate images from massive amounts of image data. Traditional methods have numerous limitations when handling these tasks, making them difficult to meet practical application needs.

[0003] In the insurance claims process within the fintech sector, customer-provided photos of the accident scene are crucial for determining the extent of loss. However, due to variations in shooting angle, lighting conditions, and other factors, the appearance of the same object in different photos can vary significantly, making it extremely difficult to identify duplicate photos. Currently, claims adjusters rely heavily on manual visual review, which, faced with the sheer volume of photos, is inefficient and prone to errors. Once a customer uses photos of the same damaged object to file claims in multiple cases, manual review is difficult to detect, leading to claims leakage. While existing feature-matching-based methods can identify duplicate photos to a certain extent, they consume significant computational resources and, when faced with large datasets, are unable to meet the timeliness requirements of actual claims processing.

[0004] In the field of medical imaging in healthcare, health care and senior care, doctors often need to compare imaging data from multiple patient examinations or identify images of similar cases in medical research. However, images taken by different devices have different parameters, and images taken from different angles of the same slice also look very different. Traditional manual comparison is time-consuming and prone to misjudgment, which can easily delay the diagnosis of the disease and the progress of research. At the same time, with the surge in the amount of medical imaging data, storage space pressure is high. If duplicate images can be accurately identified, storage management can be optimized, but existing technologies lack efficient and accurate duplicate image recognition methods. Therefore, in the fields of financial technology, healthcare and senior care, how to improve the efficiency of image recognition systems in identifying duplicate images has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The present application provides a method, apparatus, device and medium for duplicate image recognition based on a local sensitive hashing algorithm to improve the efficiency of image recognition systems in identifying duplicate images in the fields of financial technology, medical health and elderly care.

[0006] In a first aspect, the present application provides a method for identifying repeated images based on a locality sensitive hashing algorithm, the method comprising:

[0007] Extracting at least one visual feature to be identified from the image to be identified by a preset deep learning model, and mapping each visual feature to be identified into a hash value of the feature to be identified by a preset local sensitive hashing algorithm;

[0008] Calculating similarities between each preset feature hash value and the feature hash value to be identified, and determining at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library;

[0009] A duplicate image recognition result is generated by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

[0010] In a second aspect, the present application further provides a repeated image recognition device based on a local sensitive hashing algorithm, the device comprising:

[0011] A module for determining hash values ​​of features to be identified, configured to extract at least one visual feature to be identified of an image to be identified by using a preset deep learning model, and map each of the visual features to be identified into a hash value of the feature to be identified by using a preset local sensitive hashing algorithm;

[0012] a similar image determination module, configured to calculate similarities between each preset feature hash value and the feature hash value to be identified, and determine at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library;

[0013] The repeated image recognition result generating module is used to generate repeated image recognition results by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

[0014] In a third aspect, the present application also provides a computer device, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the above-mentioned repeated image recognition method based on the local sensitive hashing algorithm when executing the computer program.

[0015] In a fourth aspect, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the repeated image recognition method based on the local sensitive hashing algorithm as described above.

[0016] The present application discloses a method, apparatus, device and medium for duplicate image recognition based on a local sensitive hashing algorithm. The duplicate image recognition method based on the local sensitive hashing algorithm includes extracting at least one visual feature to be recognized of an image to be recognized through a preset deep learning model, and mapping each of the visual features to be recognized into a feature hash value to be recognized through a preset local sensitive hashing algorithm; calculating the similarity between each preset feature hash value and the feature hash value to be recognized, and determining at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library; generating a duplicate image recognition result through a preset image recognition algorithm, each similar image and the image to be recognized. Through the above-mentioned method, the present application uses a deep learning model to extract the visual features of the image to be recognized, quickly and accurately obtains the key information of the image, and uses a preset local sensitive hashing algorithm to map the visual features into feature hash values, thereby simplifying the data processing amount, being able to quickly calculate the similarity between the feature hash values ​​of different images, so that similar images can be quickly recognized in the hash space, and improving the efficiency of the image recognition system in recognizing duplicate images in the fields of financial technology, medical health and elderly care. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 is a schematic flow chart of a repeated image recognition method based on a local sensitive hashing algorithm provided in the first embodiment of the present application;

[0019] Figure 2 is a schematic flow chart of a repeated image recognition method based on a local sensitive hashing algorithm provided in the second embodiment of the present application;

[0020] Figure 3 A schematic block diagram of a repeated image recognition device based on a locality sensitive hashing algorithm provided in an embodiment of the present application;

[0021] Figure 4 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0024] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0025] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] The embodiments of the present application provide a method, apparatus, device and medium for duplicate image recognition based on a local sensitive hashing algorithm. The duplicate image recognition method based on the local sensitive hashing algorithm can be applied to an image recognition system, using a deep learning model to extract the visual features of the image to be recognized, quickly and accurately obtaining the key information of the image, and using a preset local sensitive hashing algorithm to map the visual features into feature hash values, thereby simplifying the amount of data processing and being able to quickly calculate the similarity between different image feature hash values, so that similar images can be quickly recognized in the hash space, thereby improving the efficiency of the image recognition system in recognizing duplicate images in the fields of financial technology, medical health and elderly care. The server can be an independent server or a server cluster.

[0027] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0028] See also Figure 1 , Figure 1This is a schematic flow chart of a method for identifying duplicate images based on a local sensitive hashing algorithm, as provided in the first embodiment of this application. This method can be applied to image recognition systems to improve the efficiency of image recognition systems in identifying duplicate images in fields such as financial technology, healthcare, and elderly care.

[0029] like Figure 1 As shown, the repeated image recognition method based on the local sensitive hashing algorithm specifically includes steps S10 to S30.

[0030] Step S10: extracting at least one visual feature to be identified of the image to be identified by using a preset deep learning model, and mapping each visual feature to be identified into a hash value of the feature to be identified by using a preset local sensitive hashing algorithm;

[0031] Specifically, in the fintech business, the images to be identified can be photos of claims to be identified. These images are collected and aggregated to create an insurance claims image database. Standardized image preprocessing, including scaling and cropping, is performed to align the photos with similar angles, sizes, and pixels, eliminating size and angle differences between photos.

[0032] Using a pre-configured deep learning model (such as a convolutional neural network (CNN)), we extract features from pre-processed claim photos. The model automatically learns and extracts key visual features from the image, such as text, numbers (amounts, serial numbers, etc.), and layout structure (text layout, table lines, etc.) in the claim photos.

[0033] The extracted visual feature vectors are converted using a preset locality-sensitive hashing algorithm to generate corresponding feature hash values. Different feature vectors are mapped to different positions in the hash space, and similar feature vectors have closer hash values.

[0034] Step S20: Calculate the similarity between each preset feature hash value and the feature hash value to be identified, and determine at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library;

[0035] Specifically, in the field of medical care, health care and elderly care, the feature hash values ​​of all stored preset images are obtained from the medical imaging database. These preset images contain a large number of diagnosed and labeled case images, covering various disease types and imaging manifestations.

[0036] Calculate the Hamming distance or other similarity metrics between the hash value of the image feature to be identified and the hash value of the preset feature to obtain the similarity between the two. Similarity can be achieved using the Hamming distance or other similar methods. A smaller Hamming distance indicates more similar image features, which may correspond to similar pathological conditions or disease types.

[0037] Based on the set similarity threshold, if the similarity between the image to be identified and a preset image feature hash value is high enough, it is considered that a similar image has been found in the preset image library, which may indicate that the patient to be identified has similar disease conditions or imaging manifestations, which helps doctors make reference comparisons and preliminary diagnoses.

[0038] Step S30: Generate a duplicate image recognition result by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

[0039] In a specific embodiment, in the healthcare and elderly care fields, similar medical images are input into a preset image recognition algorithm (e.g., a deep learning-based medical imaging diagnostic model) along with the image to be identified. The algorithm then performs a comprehensive analysis based on the diagnostic information of the similar images (e.g., disease type, lesion severity, and other annotated data) and the characteristics of the image to be identified.

[0040] Based on the diagnostic results of similar images and the features to be identified, the final duplicate image recognition results are generated. This may include a preliminary diagnosis of the disease, an assessment of the likelihood of lesions, and recommended further examination directions. This provides auxiliary diagnostic support for doctors, improving diagnostic efficiency and accuracy. This can especially help doctors broaden their diagnostic thinking and reference range when dealing with complex and difficult cases or rare diseases.

[0041] In a specific embodiment, in the field of financial technology, similar bill images are input into a preset image recognition algorithm model (such as a similarity comparison network based on deep learning) along with the image to be identified. The algorithm comprehensively considers the overall similarity between the similar images and the image to be identified in terms of visual features, layout structure, etc., and combines financial bill business rules (such as the uniqueness of the same bill number) to ultimately generate duplicate image recognition results. The results can determine whether the bill is submitted as a duplicate and the authenticity of the bill, thereby helping financial institutions effectively prevent fraud and ensure the security and compliance of bill business.

[0042] The present embodiment discloses a method, apparatus, device and medium for duplicate image recognition based on a local sensitive hashing algorithm. The duplicate image recognition method based on the local sensitive hashing algorithm includes extracting at least one visual feature to be recognized of an image to be recognized through a preset deep learning model, and mapping each of the visual features to be recognized into a feature hash value to be recognized through a preset local sensitive hashing algorithm; calculating the similarity between each preset feature hash value and the feature hash value to be recognized, and determining at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library; generating a duplicate image recognition result through a preset image recognition algorithm, each similar image and the image to be recognized. Through the above-mentioned method, the present application uses a deep learning model to extract the visual features of the image to be recognized, quickly and accurately obtains the key information of the image, and uses a preset local sensitive hashing algorithm to map the visual features into feature hash values, thereby simplifying the data processing amount, being able to quickly calculate the similarity between the feature hash values ​​of different images, so that similar images can be quickly recognized in the hash space, and improving the efficiency of the image recognition system in recognizing duplicate images in the fields of financial technology, medical health and elderly care.

[0043] See also Figure 2 , Figure 2 This is a schematic flow chart of a repeated image recognition method based on a locally sensitive hashing algorithm provided by the second embodiment of the present application. The repeated image recognition method based on the locally sensitive hashing algorithm can be applied to an image recognition system, and is used to perform hash mapping on image features in a targeted manner by obtaining a preset locally sensitive hash function group independently generated based on the probability distribution of each preset visual feature in a preset image database. When processing large-scale image data, the time complexity of the hash operation is low, which greatly improves the efficiency of image feature processing. Each visual feature to be identified is spliced ​​according to the dimension of the preset locally sensitive hash function group to generate a low-dimensional hash feature vector, thereby achieving effective dimensionality reduction of the original high-dimensional visual features and improving the efficiency of the image recognition system in identifying repeated images in the fields of financial technology, medical health and elderly care.

[0044] based on Figure 1 The embodiment shown, this embodiment Figure 2 As shown, step S10 includes steps S101 to S103.

[0045] Step S101: obtaining a preset locality-sensitive hash function group, wherein each hash function in the preset locality-sensitive hash function group is independently generated based on the probability distribution of each preset visual feature in the preset image database;

[0046] Specifically, images in a pre-set image database are analyzed to extract visual features, including color histograms, texture features, and edge features. Different application areas prioritize different features. For example, in the financial sector, features such as text patterns and lines on bills may be of particular interest; in healthcare and elderly care, features such as the image's grayscale distribution and the shape and boundaries of lesions may be of particular interest. The probability distribution of these visual features is calculated to understand the frequency and statistical patterns of each feature in the image database, enabling the hash function to better adapt to the characteristics of the image data.

[0047] According to the probability distribution of visual features, multiple local sensitive hash functions are independently generated to form a preset local sensitive hash function group. Each hash function has specific parameters, such as a random projection vector, a bias term, etc.

[0048] For example, a random projection method is used to generate hash functions. For each hash function, a projection vector with the same dimension as the visual feature vector and a bias value are randomly generated. When calculating the hash value, a dot product operation is performed on the visual feature vector and the projection vector. After adding the bias value, a binary bit of the hash value is obtained through operations such as rounding. Different hash functions correspond to different random projection vectors and bias values, thus hashing the visual features from different perspectives.

[0049] Step S102: performing concatenation processing on each of the visual features to be identified according to the dimensions of the preset local sensitive hash function group to generate a low-dimensional hash feature vector;

[0050] In a specific embodiment, the image to be identified is processed to extract the same visual feature types as those in the preset image database, ensuring that the image to be identified and the preset image are comparable at the feature level. For example, if the images in the preset image database are analyzed based on color histograms and texture features, then color histograms and texture features should also be extracted for the image to be identified.

[0051] Determine the dimension of the preset local sensitive hash function group, that is, the number of hash functions in the function group. Assuming that the hash function group has k hash functions, then each visual feature will obtain k hash values ​​(binary bits) after being mapped by the hash function group.

[0052] Each visual feature to be identified is mapped using a preset set of locally sensitive hash functions to obtain a corresponding hash value sequence. The hash value sequences are concatenated in a specific order to form a low-dimensional hash feature vector. For example, if there are m visual features to be identified, and each feature is mapped to k binary bits using k hash functions, the dimension of the concatenated low-dimensional hash feature vector is m*k. This vector significantly reduces the dimensionality of the original visual feature vector while retaining the key information of the original feature.

[0053] Step S103: Generate the feature hash value to be identified according to each of the low-dimensional hash feature vectors.

[0054] Specifically, based on actual needs and application scenarios, determine how to convert the low-dimensional hash feature vector into the feature hash value to be identified, and combine the binary bits in the low-dimensional hash feature vector into a hash value string. For example, each binary bit is arranged in sequence to form a binary number, or other encoding methods are used to convert it into a compact hash value representation.

[0055] The concatenated low-dimensional hash feature vectors are processed according to defined rules to obtain a final hash value for the feature to be identified. This hash value can be quickly compared with other hash values ​​to determine the similarity between images. For example, in the scenario of duplicate financial bill detection, each bill image is processed through the above steps to obtain a unique feature hash value. By comparing the feature hash value of the newly submitted bill with the hash value of the stored bill, it is possible to quickly determine whether there are duplicate bills. In the retrieval of similar medical image cases, the feature hash value can be used to quickly find case images similar to the image to be identified from a large number of images, providing a reference for doctors' diagnosis.

[0056] This embodiment discloses a method, apparatus, device, and medium for identifying duplicate images based on a local sensitive hashing algorithm. The method comprises obtaining a preset local sensitive hash function group, wherein each hash function in the preset local sensitive hash function group is independently generated based on the probability distribution of each preset visual feature in the preset image database; concatenating each of the visual features to be identified according to the dimension of the preset local sensitive hash function group to generate a low-dimensional hash feature vector; and generating a hash value of the feature to be identified based on each of the low-dimensional hash feature vectors. Through the above-mentioned method, the present application can perform targeted hash mapping on image features by obtaining a preset local sensitive hash function group independently generated based on the probability distribution of each preset visual feature in the preset image database. When processing large-scale image data, the time complexity of the hash operation is low, which greatly improves the efficiency of image feature processing. The visual features to be identified are concatenated according to the dimension of the preset local sensitive hash function group to generate a low-dimensional hash feature vector, achieving effective dimensionality reduction of the original high-dimensional visual features, thereby improving the efficiency of image recognition systems in identifying duplicate images in the fields of financial technology, medical health, and elderly care.

[0057] based on Figure 2 In the embodiment shown, in this embodiment, step S103 includes:

[0058] The concatenated low-dimensional hash feature vectors are converted into binary codes using a preset binary coding rule to generate a binary hash string of fixed length;

[0059] The binary hash string is determined as the hash value of the feature to be identified.

[0060] Specifically, a suitable binarization threshold is determined according to actual needs and data characteristics, for example, it can be set to 0.5, or a threshold that can maximize the distinction of image features is selected by analyzing data distribution.

[0061] Define a binary encoding rule. Compare each element in the concatenated low-dimensional hash feature vector with a threshold. Elements greater than or equal to the threshold are encoded as 1, and elements less than the threshold are encoded as 0. The low-dimensional hash feature vector, mapped and concatenated using a preset group of locally sensitive hash functions, is used as input data. Each element in the low-dimensional hash feature vector is converted according to the preset binary encoding rule. Starting with the first element of the vector, the relationship between the value of each element and the set threshold is determined, and the element is encoded as 0 or 1 based on the comparison result.

[0062] The binary encoding results of all elements are concatenated to form a fixed-length binary string, which is the binary hash string. The length of this binary hash string is equal to the dimension of the low-dimensional hash feature vector, and the value at each position represents the binarization result of the corresponding element.

[0063] The generated binary hash string is verified, for example, to check whether it meets the expected length requirements and whether it complies with the binary encoding format requirements. The verified and corrected binary hash string is finally determined as the feature hash value to be identified. In the financial field, it is used to detect duplicate bill images, and in the medical field, it is used to retrieve similar case images.

[0064] In a specific embodiment, the concatenated low-dimensional hash feature vectors are converted into binary codes using a preset binary coding rule to generate a fixed-length binary hash string, including:

[0065] Converting the low-dimensional hash feature vector that is greater than or equal to a preset encoding threshold into a first code;

[0066] Converting the low-dimensional hash feature vector smaller than the preset encoding threshold into a second encoding;

[0067] The binary hash string is generated according to each of the first codes, each of the second codes and the preset binarization coding rule.

[0068] Specifically, based on the actual application scenario and data characteristics, a suitable preset coding threshold (for example, 0.5) is set as the standard for judging the size of the low-dimensional hash feature vector elements. It is stipulated that the low-dimensional hash feature vector elements greater than or equal to the preset coding threshold are converted to the first code (for example, 1), while the elements less than the preset coding threshold are converted to the second code (for example, 0). Starting from the first element of the low-dimensional hash feature vector, each element is judged and converted in turn. For each element, it is compared with the preset coding threshold. If it is greater than or equal to the threshold, it is converted to the first code; if it is less than the threshold, it is converted to the second code.

[0069] During the traversal process, the encoding value (0 or 1) obtained from each conversion is stored in a new string in turn. When all elements are judged and converted, the resulting string is a binary hash string. The obtained binary hash string is subjected to necessary verification to ensure that it accurately reflects the binary features of the original low-dimensional hash feature vector. The binary hash string that passes the verification is finally determined as the feature hash value to be identified.

[0070] based on Figure 1 In the embodiment shown, in this embodiment, step S30 includes:

[0071] Extracting key point features of the image to be identified and similar key point features of the similar images;

[0072] Specifically, keypoint detection is performed on both the image to be identified and similar images. A detection algorithm is used to identify keypoints in the image. These keypoints are typically points with distinct features in the image, such as corners, edges, and points of interest, which can reflect the local characteristics of the image. A descriptor is generated for each detected keypoint. The method for generating descriptors depends on the keypoint detection algorithm used.

[0073] Iteratively calculating the affine transformation matrix between the local key point feature to be identified and the similar local key point feature through the preset image recognition algorithm;

[0074] Specifically, a preset image recognition algorithm (such as a nearest neighbor matching algorithm based on a descriptor) is used to perform preliminary matching on the extracted key point features of the image to be recognized and the key point features of similar images to find possible corresponding point pairs.

[0075] In order to improve the accuracy of matching, a two-way matching verification method can be adopted, that is, for each key point of the image to be identified, the most similar key point is found in the similar image; at the same time, for each key point in the similar image, the most similar key point is also found in the image to be identified. Only point pairs with consistent two-way matching are considered valid matching point pairs.

[0076] If the key point features to be identified and the similar key point features exist in the affine transformation matrix and the feature point distance error is less than the preset distance error threshold, the image to be identified is determined to be a repeated image, and the image to be identified is determined to be the repeated image as the repeated image recognition result.

[0077] Specifically, according to the final affine transformation matrix, all key points of the image to be identified are mapped to the similar image coordinate system, the distance error between each mapped key point and the corresponding key point of the similar image is calculated, and each distance error is compared with a preset distance error threshold.

[0078] If the distance error between the feature point pairs is less than a preset threshold, it is considered that there is sufficient similarity between the image to be identified and the similar image, and the image to be identified is determined to be a duplicate image.

[0079] based on Figure 1 In the embodiment shown, in this embodiment, step S20 includes:

[0080] Calculating the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determining each of the similarities according to the Hamming distance;

[0081] The preset images in the preset image library whose similarity is greater than a preset similarity threshold are determined as the similar images.

[0082] Specifically, the similarity calculation formula is defined as: Similarity = 1-(Hamming distance) / (hash value length). For example, for a 64-bit hash value, Similarity = 1-(Hamming distance) / 64. According to the formula, the calculated Hamming distance is substituted into the formula to calculate the similarity between the feature hash value to be identified and each preset feature hash value. For example, if the Hamming distance is 8, then Similarity = 1-8 / 64 = 0.875.

[0083] Set an appropriate similarity threshold based on the actual application scenario and requirements. For example, set the threshold to 0.85. For each preset image, compare the calculated similarity between it and the image to be identified with the preset similarity threshold. For example, the calculated similarity of 0.875 is greater than the threshold of 0.85.

[0084] If the similarity is greater than or equal to the threshold, the preset image is determined as a similar image.

[0085] In a specific embodiment, after calculating the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determining each of the similarities according to the Hamming distance, the method includes:

[0086] If the preset image with the similarity greater than or equal to the preset similarity threshold does not exist in the preset image library, arranging the preset images in descending order according to the similarities to generate an image sorting list;

[0087] The first preset number of preset images in the image sorting list are determined as the similar images.

[0088] Specifically, for each preset image in the preset image library, check whether its similarity with the image to be identified is greater than or equal to the preset similarity threshold. If there is a preset image with a similarity greater than or equal to the threshold, directly proceed to the step of determining similar images; if there is no preset image with a similarity greater than or equal to the threshold, prepare to generate an image sorting list.

[0089] Based on the collected similarity data, the preset images are sorted from high to low according to the similarity to generate an image sorting list. According to actual needs, a preset number is set to determine how many preset images are selected from the sorting list as similar images. From the image sorting list, the preset number of preset images are selected as similar images.

[0090] See also Figure 3 , Figure 3The embodiment of the present application provides a schematic block diagram of a repeated image recognition device based on a local sensitive hashing algorithm, which is used to perform the aforementioned repeated image recognition method based on a local sensitive hashing algorithm. The repeated image recognition device based on a local sensitive hashing algorithm can be configured on a server.

[0091] like Figure 3 As shown, the repeated image recognition device 400 based on the local sensitive hashing algorithm includes:

[0092] A module 410 for determining hash values ​​of features to be identified is configured to extract at least one visual feature to be identified from an image to be identified using a preset deep learning model, and map each visual feature to be identified into a hash value of the feature to be identified using a preset local sensitive hashing algorithm.

[0093] A similar image determination module 420 is configured to calculate similarities between each preset feature hash value and the feature hash value to be identified, and determine at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library;

[0094] The duplicate image recognition result generating module 430 is configured to generate a duplicate image recognition result by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

[0095] Furthermore, the to-be-identified feature hash value determination module 410 includes:

[0096] a locality-sensitive hash function group acquisition submodule, configured to acquire a preset locality-sensitive hash function group, wherein each hash function in the preset locality-sensitive hash function group is independently generated based on a probability distribution of each preset visual feature in the preset image database;

[0097] A low-dimensional hash feature vector generation submodule, configured to perform concatenation processing on each of the visual features to be identified according to the dimensions of the preset local sensitive hash function group to generate a low-dimensional hash feature vector;

[0098] The feature hash value generating submodule to be identified is used to generate the feature hash value to be identified according to each of the low-dimensional hash feature vectors.

[0099] Furthermore, the submodule for generating hash values ​​of features to be identified includes:

[0100] A binary hash string generation unit is used to perform binary encoding conversion on the spliced ​​low-dimensional hash feature vectors according to a preset binary encoding rule to generate a binary hash string of fixed length;

[0101] The feature hash value determination unit to be identified is configured to determine the binary hash string as the feature hash value to be identified.

[0102] Furthermore, the binary hash string generation unit includes:

[0103] A first code conversion subunit, configured to convert the low-dimensional hash feature vector that is greater than or equal to a preset code threshold into a first code;

[0104] A second code conversion subunit, configured to convert the low-dimensional hash feature vector smaller than the preset code threshold into a second code;

[0105] The binary hash string generation subunit is configured to generate the binary hash string according to each of the first codes, each of the second codes and the preset binarization coding rule.

[0106] Furthermore, the repeated image recognition result generation module 430 includes:

[0107] A key point feature extraction submodule, configured to extract key point features of the image to be identified and similar key point features of similar images;

[0108] An affine transformation matrix calculation submodule, configured to iteratively calculate the affine transformation matrix between the local key point feature to be identified and the similar local key point feature using the preset image recognition algorithm;

[0109] The repeated image recognition result determination submodule is used to determine that the image to be identified is a repeated image if there are the key point features to be identified and the similar key point features in the affine transformation matrix whose feature point distance error is less than a preset distance error threshold, and determine that the image to be identified is the repeated image as the repeated image recognition result.

[0110] Furthermore, the similar image determination module 420 includes:

[0111] A similarity determination submodule, configured to calculate a Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determine each of the similarities according to the Hamming distance;

[0112] The similar image determination submodule is configured to determine each of the preset images in the preset image library whose similarity is greater than a preset similarity threshold as the similar image.

[0113] Furthermore, the similar image determination module 420 includes:

[0114] an image sorting list generating submodule, configured to, if the preset image having the similarity greater than or equal to the preset similarity threshold does not exist in the preset image library, sort the preset images in descending order according to the similarities to generate an image sorting list;

[0115] The similar image determination submodule is configured to determine the first preset number of preset images in the image sorting list as the similar images.

[0116] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0117] The above-mentioned device can be realized in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer device shown.

[0118] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0119] See Figure 4 The computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause a processor to perform any repeated image recognition method based on a locality-sensitive hashing algorithm.

[0121] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0122] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any repeated image recognition method based on the local sensitive hashing algorithm.

[0123] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0124] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0125] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0126] Extracting at least one visual feature to be identified from the image to be identified by a preset deep learning model, and mapping each visual feature to be identified into a hash value of the feature to be identified by a preset local sensitive hashing algorithm;

[0127] Calculating similarities between each preset feature hash value and the feature hash value to be identified, and determining at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library;

[0128] A duplicate image recognition result is generated by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

[0129] In one embodiment, each of the visual features to be identified is mapped to a hash value of the feature to be identified by a preset local sensitive hashing algorithm, so as to achieve:

[0130] Obtaining a preset locality-sensitive hash function group, wherein each hash function in the preset locality-sensitive hash function group is independently generated based on a probability distribution of each preset visual feature in the preset image database;

[0131] Concatenate the visual features to be identified according to the dimensions of the preset local sensitive hash function group to generate a low-dimensional hash feature vector;

[0132] The feature hash value to be identified is generated according to each of the low-dimensional hash feature vectors.

[0133] In one embodiment, the feature hash value to be identified is generated according to each of the low-dimensional hash feature vectors to achieve:

[0134] The concatenated low-dimensional hash feature vectors are converted into binary codes using a preset binary coding rule to generate a binary hash string of fixed length;

[0135] The binary hash string is determined as the hash value of the feature to be identified.

[0136] In one embodiment, the concatenated low-dimensional hash feature vectors are converted into binary codes using a preset binary coding rule to generate a fixed-length binary hash string for achieving:

[0137] Converting the low-dimensional hash feature vector that is greater than or equal to a preset encoding threshold into a first code;

[0138] Converting the low-dimensional hash feature vector smaller than the preset encoding threshold into a second encoding;

[0139] The binary hash string is generated according to each of the first codes, each of the second codes and the preset binarization coding rule.

[0140] In one embodiment, a preset image recognition algorithm, each of the similar images, and the image to be recognized are used to generate a duplicate image recognition result, so as to achieve:

[0141] Extracting key point features of the image to be identified and similar key point features of the similar images;

[0142] Iteratively calculating the affine transformation matrix between the local key point feature to be identified and the similar local key point feature through the preset image recognition algorithm;

[0143] If the key point features to be identified and the similar key point features exist in the affine transformation matrix and the feature point distance error is less than the preset distance error threshold, the image to be identified is determined to be a repeated image, and the image to be identified is determined to be the repeated image as the repeated image recognition result.

[0144] In one embodiment, similarities between each preset feature hash value and the hash value of the feature to be identified are calculated, and at least one similar image is determined from a preset image library based on each similarity, to achieve:

[0145] Calculating the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determining each of the similarities according to the Hamming distance;

[0146] The preset images in the preset image library whose similarity is greater than a preset similarity threshold are determined as the similar images.

[0147] In one embodiment, the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified is calculated, and each similarity is determined according to the Hamming distance, so as to achieve:

[0148] If the preset image with the similarity greater than or equal to the preset similarity threshold does not exist in the preset image library, arranging the preset images in descending order according to the similarities to generate an image sorting list;

[0149] The first preset number of preset images in the image sorting list are determined as the similar images.

[0150] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement any one of the repeated image recognition methods based on the local sensitive hashing algorithm provided in the embodiments of the present application.

[0151] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.

[0152] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A repeated image recognition method based on local sensitive hashing algorithm, characterized in that: include: Extracting at least one visual feature to be identified from the image to be identified by a preset deep learning model, and mapping each visual feature to be identified into a hash value of the feature to be identified by a preset local sensitive hashing algorithm; Calculating similarities between each preset feature hash value and the feature hash value to be identified, and determining at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library; A duplicate image recognition result is generated by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

2. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 1 is characterized in that: Mapping each of the visual features to be identified into a hash value of the feature to be identified by using a preset local sensitive hashing algorithm includes: Obtaining a preset locality-sensitive hash function group, wherein each hash function in the preset locality-sensitive hash function group is independently generated based on a probability distribution of each preset visual feature in the preset image database; Concatenate the visual features to be identified according to the dimensions of the preset local sensitive hash function group to generate a low-dimensional hash feature vector; The feature hash value to be identified is generated according to each of the low-dimensional hash feature vectors.

3. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 2 is characterized in that: Generating the feature hash value to be identified according to each of the low-dimensional hash feature vectors includes: The concatenated low-dimensional hash feature vectors are converted into binary codes using a preset binary coding rule to generate a binary hash string of fixed length; The binary hash string is determined as the hash value of the feature to be identified.

4. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 3 is characterized in that: The binary encoding conversion of the spliced ​​low-dimensional hash feature vectors by a preset binary encoding rule to generate a binary hash string of fixed length includes: Converting the low-dimensional hash feature vector that is greater than or equal to a preset encoding threshold into a first code; Converting the low-dimensional hash feature vector smaller than the preset encoding threshold into a second encoding; The binary hash string is generated according to each of the first codes, each of the second codes and the preset binarization coding rule.

5. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 1 is characterized in that: The generating of a duplicate image recognition result by using a preset image recognition algorithm, each of the similar images and the image to be recognized includes: Extracting key point features of the image to be identified and similar key point features of the similar images; Iteratively calculating the affine transformation matrix between the local key point feature to be identified and the similar local key point feature through the preset image recognition algorithm; If the key point features to be identified and the similar key point features exist in the affine transformation matrix and the feature point distance error is less than the preset distance error threshold, the image to be identified is determined to be a repeated image, and the image to be identified is determined to be the repeated image as the repeated image recognition result.

6. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 1 is characterized in that: The calculating the similarity between each preset feature hash value and the hash value of the feature to be identified, and determining at least one similar image from a preset image library according to each similarity, includes: Calculating the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determining each of the similarities according to the Hamming distance; The preset images in the preset image library whose similarity is greater than a preset similarity threshold are determined as the similar images.

7. The repeated image recognition method based on the local sensitive hashing algorithm according to claim 6 is characterized in that: After calculating the Hamming distance between each of the preset feature hash values ​​and the feature hash value to be identified, and determining each of the similarities according to the Hamming distance, the method further includes: If the preset image with the similarity greater than or equal to the preset similarity threshold does not exist in the preset image library, arranging the preset images in descending order according to the similarities to generate an image sorting list; The first preset number of preset images in the image sorting list are determined as the similar images.

8. A repeated image recognition device based on a local sensitive hashing algorithm, characterized in that: include: A module for determining hash values ​​of features to be identified, configured to extract at least one visual feature to be identified from an image to be identified using a preset deep learning model, and map each of the visual features to be identified into a hash value of the feature to be identified using a preset local sensitive hashing algorithm; a similar image determination module, configured to calculate similarities between each preset feature hash value and the feature hash value to be identified, and determine at least one similar image from a preset image library based on each similarity, wherein each preset feature hash value is a feature hash value corresponding to each preset image in the preset image library; The repeated image recognition result generating module is used to generate repeated image recognition results by using a preset image recognition algorithm, each of the similar images and the image to be recognized.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the repeated image recognition method based on the local sensitive hashing algorithm as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the repeated image recognition method based on the local sensitive hashing algorithm according to any one of claims 1 to 7.