Image deduplication method and device, equipment, storage medium and program product

By splitting images into local feature categories and performing deduplication based on mapping relationships, the problem of insufficient image deduplication accuracy in existing technologies is solved, realizing a more efficient image deduplication method that preserves important visual information and optimizes storage space utilization.

CN121808091APending Publication Date: 2026-04-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing image deduplication methods, especially those based on visual information and vector similarity, suffer from poor deduplication accuracy, resulting in the loss of necessary visual information from the reference image library.

Method used

Each image is split into one or more feature categories corresponding to local images, and deduplication is performed based on the mapping relationship between images and feature categories to generate a second image set. This ensures that the second image set covers all feature categories and reduces the number of images.

Benefits of technology

While reducing the storage space of the image set, it retains the local image features of all categories, improves the accuracy and stability of deduplication, and balances the integrity of local image feature coverage and the utilization of storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808091A_ABST
    Figure CN121808091A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an image deduplication method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a first image set comprising a plurality of images; on the basis of a mapping relation between each image and a feature category, performing duplicate removal processing on each image in the first image set to generate a second image set; wherein the feature category indicates a category corresponding to a local image contained in the image; the number of images in the second image set is less than that in the first image set, and the second image set completely covers each feature category. Therefore, while the storage space of the image set is reduced, local image features of all categories are reserved, and the integrity of local image feature coverage and the utilization rate of the image storage space are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and particularly relates to an image deduplication method and device, equipment, a storage medium and a program product. BACKGROUND

[0002] There are some services in the related art that perform image retrieval or image matching based on a specific type of object contained in an image, and these services contain a reference image library. The source of the reference image library is diverse, and there may be a large amount of image redundancy in the reference image library, which needs to be deduplicated. For example, taking the above-mentioned specific type of object as a geographical location, there are multiple photographed images for the same geographical location in the reference image library, so image deduplication is needed.

[0003] The current image deduplication method is mainly a deduplication algorithm based on the vector similarity of visual information, that is, the feature vectors of each image in the reference image library are clustered, and only a certain number of images are retained in each cluster. However, this deduplication method only considers the overall visual similarity of the image, so the removed images will cause some necessary visual information to be lost in the reference image library, resulting in the problem of poor deduplication accuracy. SUMMARY

[0004] To solve the above technical problems, the embodiments of the present disclosure provide an image deduplication method, device, equipment, storage medium and program product.

[0005] In a first aspect, the embodiments of the present disclosure provide an image deduplication method, which comprises: obtaining a first image set containing a plurality of images; performing deduplication processing on each image in the first image set based on a mapping relationship between each image and a feature category, to generate a second image set; wherein the feature category indicates a category corresponding to a local image contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category.

[0006] In a second aspect, the embodiments of the present disclosure also provide an image deduplication device, which comprises: a first image set obtaining module configured to obtain a first image set containing a plurality of images; The second image set generation module is configured to perform deduplication processing on each image in the first image set based on a mapping relationship between each image and a feature category, to generate a second image set, wherein the feature category indicates a category corresponding to a local image contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category.

[0007] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises: a processor; a memory configured to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the image deduplication method described in any of the embodiments of the present disclosure.

[0008] In a fourth aspect, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the image deduplication method described in any of the embodiments of the present disclosure.

[0009] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product configured to execute the image deduplication method described in any of the embodiments of the present disclosure.

[0010] The image deduplication method, device, equipment, storage medium and program product provided by the embodiments of the present disclosure can obtain a first image set containing a plurality of images; perform deduplication processing on each image in the first image set based on a mapping relationship between each image and a feature category, to generate a second image set; wherein the feature category indicates a category corresponding to a local image contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category; each image is discretized into one or more feature categories corresponding to local images contained in the image, and the deduplication problem of the first image set is converted into an image screening or filtering problem of selecting partial images to cover all feature categories, to complete the deduplication of the first image set, so that the storage space of the image set is reduced while retaining all local image features, and the integrity of the local image feature coverage and the utilization rate of the image storage space are taken into account.

[0011] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 A schematic flowchart of an image deduplication method provided in an embodiment of this disclosure; Figure 2 A schematic diagram illustrating the principle of an image deduplication method provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of the structure of an image deduplication device provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0016] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0018] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0019] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0020] Some related technologies involve image retrieval or matching based on specific categories of objects contained within images. These services utilize reference image libraries. These reference image libraries come from diverse sources and may contain a significant amount of image redundancy, necessitating image deduplication.

[0021] Taking the aforementioned specific category of objects as geographical locations as an example, the reference image library contains multiple images of the same geographical location, thus requiring image deduplication. Currently, there are two main methods for image deduplication related to geographical locations: one is to map each image in the reference image library to pre-divided regional grids based on geographical coordinates, with each regional grid retaining only a certain number of images; the other is a deduplication algorithm based on visual information vector similarity, which clusters the feature vectors of each image in the reference image library and restricts each cluster to retaining only a certain number of images.

[0022] However, the first deduplication method mentioned above can lead to image misclassification due to geographic coordinate positioning errors, resulting in the incorrect retention of images. The second deduplication method only considers the overall visual similarity of the images, ignoring the impact of local differences, which can lead to the loss of visual information about some geographic locations when removing images. In summary, existing image deduplication methods for associated geographic locations all suffer from poor deduplication accuracy.

[0023] Based on the above, this disclosure provides a technical solution for image deduplication. By splitting each image containing a specific category of objects into one or more feature categories corresponding to local images, and deduplicating the first image set according to the relationship between the images and feature categories, the local differences within the images are fully considered. This reduces the storage space of the image set while retaining the local image features of all categories, taking into account the integrity of the local image feature coverage and the utilization rate of image storage space, and also improving the accuracy and stability of image deduplication.

[0024] The image deduplication method provided in this disclosure is applicable to deduplication scenarios involving images containing at least one specific category of objects. This method can be executed by an image deduplication device, which can be implemented in software and / or hardware and integrated into an electronic device with image processing capabilities. This electronic device may include, but is not limited to, smartphones, personal digital assistants (PDAs), tablet computers (Tablet PCs), laptops, desktop computers, servers, or server clusters.

[0025] Figure 1 A flowchart illustrating an image deduplication method provided in an embodiment of this disclosure is shown. Figure 1 As shown, the image deduplication method may include the following steps: S110. Obtain the first image set containing multiple images.

[0026] Specifically, the electronic device first acquires a first image set containing redundant images. For example, the electronic device can access an image database related to the business scenario to acquire the first image set. This first image set contains multiple images that have not undergone deduplication. Each image contains at least one object of a specific category. This specific category of object can be an object defined by business requirements. For example, if the specific category of object is a geographical location, then each image in the first image set is an image containing at least one geographical location. Here, a geographical location is a location with a physical geographical scene, and the image can be understood as an image obtained by taking a picture of the physical geographical scene.

[0027] It should be noted that in the subsequent embodiments of this disclosure, specific categories of objects will be used as examples of geographical locations for illustration. Unless otherwise specified, each image to be deduplicated contains one or more geographical locations.

[0028] In some embodiments, for application scenarios where each image contains at least one geographic location, S110 includes: acquiring each image in the image database to be deduplicated; classifying each image into different geographic grids based on the geographic location information corresponding to the images, and generating a first image set corresponding to each geographic grid.

[0029] In this context, geographic location information refers to the three-dimensional geographic coordinates or latitude and longitude carried in the image's metadata, which can be obtained through the positioning sensor of the image-capturing device. The geographic grid is a pre-defined geographic area, the shape and size of which can be defined as needed. To mitigate the impact of positioning deviations in the geographic location information, the grid size of the geographic grid in this embodiment is larger than a preset distance value that is a multiple of the positioning deviation of the geographic location information. This preset multiple can be determined comprehensively based on factors such as device processing performance and processing speed required by the business. For example, the preset multiple can be set to a relatively large value, such as tens or hundreds. In this way, the grid size of the geographic grid is much larger than the positioning deviation of the geographic location information.

[0030] Specifically, considering that all the images in the image database may correspond to a relatively large geographical area, the amount of data processing may be too large. Therefore, in this embodiment, the geographical area covered by each image in the image database can be pre-divided into multiple geographical grids according to a set grid size. Then, the electronic device can perform grid positioning based on the geographical location information of the images, classify all the images into different geographical grids, and generate a first image set corresponding to each geographical grid, so as to reduce the amount of data for each deduplication process and thus improve the deduplication efficiency.

[0031] It should be noted that the first image set corresponding to each geographic grid can perform subsequent deduplication processing in parallel to further improve deduplication efficiency and shorten the time spent deduplication on the image database.

[0032] S120. Based on the mapping relationship between each image and the feature category, deduplication is performed on each image in the first image set to generate a second image set; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category.

[0033] The feature category indicates the category corresponding to a local image within the overall image. A local image is an image composed of pixels in a local region within the entire image; it is a part of the image that possesses static invariance and discriminability.

[0034] Specifically, the reason why deduplication schemes using image vector similarity in related technologies are ineffective is that an image may contain multiple geographical locations, while vector similarity only focuses on the overall similarity of the images. This leads to the easy loss of visual information (i.e., local image features) of geographical locations occupying local areas of the image. For example, image 1 captures locations A, B, and C, while image 2 captures locations B, C, and D. Most areas of these two images have the same content or high similarity, while location A in image 1 and location D in image 2 have relatively small proportions. Therefore, using vector similarity deduplication schemes, it is easy to extract image 1 or image 2, resulting in the loss of visual information of location A or location D.

[0035] Based on the above, the deduplication approach of this disclosure does not perform deduplication on the entire image, but rather divides the entire image into different local images according to specific categories of objects (such as geographical locations). Different local images correspond to different feature categories, for example, discretizing into one or more feature categories corresponding to local images with discriminative power. This disclosure does not limit the methods of division and discretization. Thus, there are relationships between images and feature categories, such as one-to-one or one-to-many correspondences, or more complex logical relationships (such as using some logical algorithms to express the relationship between images and feature categories, which are respectively dependent and independent variables). If there is redundancy among the images in the first image set, then there is repetition among the feature categories corresponding to the images. However, a feature category associated with an image can represent a specific category of objects in that image relatively fully / comprehensively. Therefore, in this disclosure, the deduplication between images can be driven by deduplication between feature categories, so as to retain all non-repeating feature categories using a relatively small number of images. In this way, while reducing the number of images in the first image set, all the necessary visual features in the first image set can be retained.

[0036] Therefore, after acquiring the first image set, the electronic device can obtain the relationship between each image and its feature category. For example, the feature categories contained in the images in the first image set are labeled when they are entered, so the correspondence or mapping relationship between the images and feature categories can be directly obtained. Furthermore, image processing such as feature extraction, object detection, and feature classification can be performed on each image in the first image set to establish the relationship between the images and feature categories. Then, using the relationship between the images and feature categories, a smaller number of images are selected to cover all feature categories as constraints for image filtering or selection. Deduplication is then performed on the images in the first image set to obtain the second image set. For example, if image 3 captures location A, image 4 captures locations B and C, and image 5 captures location B, only images 3 and 4 can be retained.

[0037] In some embodiments, S120 includes the following steps A and B.

[0038] Step A: Establish the mapping relationship between each image and the feature category identifier.

[0039] Feature category identifiers are identifying information about the feature category to which a local feature within an image belongs, and they are globally unique. Feature category identifiers can be digitally encoded values, codes, or symbols related to the feature category. For example, a feature category identifier can be understood as an identifier of a feature category related to a geographical location, such as location A, location B, etc. Local features are features of a local part of the image.

[0040] Specifically, given that manually establishing the relationship between images and feature categories introduces subjective errors and is time-consuming and labor-intensive, this embodiment can automatically establish the aforementioned relationship. Furthermore, considering that feature categories have actual category meanings, are not easily fully defined, and are not convenient for participation in calculations, this embodiment can digitize them as feature category identifiers. Similarly, for ease of calculation, local images can be digitized as corresponding features, that is, local images can be represented as local features.

[0041] Based on the above, the electronic device can perform feature extraction / object detection, feature classification, category identifier creation, and other related processing on each image in the first image set to determine the local features contained in each image and their corresponding feature category identifiers. Then, a mapping relationship can be established between each image and the feature category identifiers corresponding to each local feature contained within it. For example: Image - {Category Identifier 1, Category Identifier 2, ..., Category Identifier k}.

[0042] by Figure 2 For example, based on the mapping relationship above, we can obtain... Figure 2 The images shown are bipartite graphs between geographic locations: Image 1 - {Location A, Location B, Location C}, Image 2 - {Location B, Location C}, Image 3 - {Location C, Location D}, Image 4 - {Location D, Location E, Location F}, Image 5 - {Location C, Location E, Location F}.

[0043] Step B: Based on the mapping relationships, perform deduplication on each image in the first image set to generate the second image set.

[0044] Specifically, the electronic device can use the mapping relationship corresponding to each image as the input data for image deduplication, and adopt related algorithms in related technologies that can achieve the above-mentioned selection of a small number of images to cover all feature category identifiers, such as related algorithms for solving set coverage problems, filtering images one by one in order of covering feature category identifiers from most to least, and deduplicating them, to perform deduplication processing on each image in the first image set to obtain the second image set.

[0045] Continue with Figure 2 As shown in the example, after filtering, the electronic device can select images 1 and 4 as deduplication results, and only two images are needed to cover locations A to F.

[0046] In some embodiments, S130 can be implemented as follows: the deduplication problem based on the mapping relationship is modeled as a set covering problem. In this way, the aforementioned mapping relationship can be used as input data, and the constraint of minimizing the number of images in the second image set obtained after deduplication is used as the optimization solution. Using relevant algorithms for solving set covering problems (such as greedy algorithms and their related variants, heuristic algorithms, etc.), images containing uncovered feature category labels are continuously filtered from the first image set until the filtered images completely cover all feature category labels. Then the filtering process ends, and the filtered images constitute the second image set.

[0047] In other embodiments, S130 includes: selecting images from the first image set based on each mapping relationship to perform image combination, generating multiple image combinations that satisfy the first condition; and selecting an image combination that satisfies the second condition from each image combination as the second image set.

[0048] The first condition includes that the image combination completely covers the feature category identifiers, and that the images within the image combination are not duplicated. The second condition includes that the number of images in the image combination is less than the number of images in the first image set.

[0049] Specifically, the deduplication process based on mapping relationships can be as follows: First, the electronic device, based on the mapping relationship corresponding to each image and a first condition, performs different permutations and combinations of all images in the first image set to obtain multiple image combinations. Each such image combination satisfies the basic requirement of retaining all visual features. Then, the electronic device, based on a second condition, selects one image combination from all the image combinations to form the second image set.

[0050] In one example, if the second condition is met, an image combination can be randomly selected from the various image combinations as the second image set.

[0051] In another example, the second condition includes having the fewest images in the image combination. In this example, the image combination with the fewest images can be selected from all the image combinations as the second image set.

[0052] In another example, the second condition includes that the image quality of the image combination is higher than a preset quality threshold, and that the number of images in the image combination is less than the number of images in the first image set.

[0053] The preset quality threshold is a pre-set critical value used to filter relatively high-quality images or image combinations. If applied to the overall image quality of an image combination, the preset quality threshold can be implemented as a first quality threshold, which is a relatively small value, for example, slightly larger than the median of the image quality range, to ensure the overall quality of the image combination meets the standard. If applied to the quality of an individual image within the image combination, the preset quality threshold can be implemented as a second quality threshold, which is a value greater than the first quality threshold, for example, a value within a larger range of image quality values ​​(such as 70%~90%), to ensure the quality of the individual image meets the standard.

[0054] Specifically, since the reference image library serves as the baseline images for image retrieval or image matching, in addition to reducing redundancy, the images within it can also be of high quality. Based on this, this embodiment can perform image deduplication processing simultaneously by considering both image quality and the number of images.

[0055] For example, an electronic device can first perform a first image combination filtering based on the number of images, that is, filtering out image combinations with fewer images than the number of images in a first image set. Then, the electronic device performs a second image combination filtering based on image quality, that is, using relevant image quality assessment methods to calculate the overall image quality of the image combinations obtained from the first filtering, and filtering out image combinations whose overall image quality is greater than a first quality threshold. If there is only one image combination obtained from the second filtering, it is determined as the second image set. If there are multiple image combinations obtained from the second filtering, the image combination with the highest overall image quality or the fewest images is determined as the second image set.

[0056] For example, after the first screening, the electronic device can use image quality assessment methods to calculate the individual image quality of each image in each image combination obtained from the first screening, and count the number of images whose individual image quality is greater than a second quality threshold. Then, the image combinations obtained from the first screening with more than a preset number of images are used as the image combinations obtained from the second screening. If there is only one image combination obtained from the second screening, it is determined as the second image set. If there are multiple image combinations obtained from the second screening, the image combination with the largest number of images or the smallest number of images within the combination is determined as the second image set.

[0057] For example, an electronic device can, according to the foregoing description, first perform a first image combination filtering based on image quality, and then perform a second image combination filtering based on the number of images to generate a second image set.

[0058] In yet another example, the second condition includes that the image combination completely covers the feature category identifiers corresponding to the first and second features.

[0059] Among them, the first feature and the second feature are different local features among all the local features. They can be the local features that the business party pays more attention to or that are more important among all the local features.

[0060] Specifically, as described above, each image set contains multiple images that cover all local features without repeating feature category identifiers. Based on this, a more important first and second feature can be selected from all local features. Then, from each image set, a set that fully covers the feature category identifiers corresponding to the first and second features and has a relatively small number of images can be selected as the second image set. This ensures that the second image set contains fewer images that adequately represent the more important local features, achieving the characteristic of a small but high-quality second image set.

[0061] The image deduplication method provided in this disclosure can obtain a first image set containing multiple images; based on the mapping relationship between each image and a feature category, the images in the first image set are deduplicated to generate a second image set; wherein, the feature category indicates the category corresponding to the local images contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category; it achieves the discretization of each image into the feature category corresponding to one or more local images contained within it, and transforms the deduplication problem of the first image set into an image filtering or screening problem of selecting some images to cover all feature categories according to the relationship between the image and the feature category, so as to complete the deduplication of the first image set, thereby reducing the storage space of the image set while retaining the local image features of all categories, taking into account the integrity of the local image feature coverage and the utilization rate of image storage space.

[0062] In some embodiments, S120 includes: extracting local features from each image in the first image set to determine the local features contained in the image; performing cluster analysis on the local features contained in each image to determine the cluster to which the local features belong and generating a feature category identifier corresponding to the cluster; and establishing a mapping relationship between each image and at least one feature category identifier based on the relationship between the local features and the cluster.

[0063] Local features are numerical representations of pixels within a local region of an image (referred to as local images). They can be obtained through feature extraction of the local image or by directly vectorizing the local image. In the application example of geographic location, the corresponding local features can be features related to the geographic location of the image.

[0064] Specifically, the electronic device can employ local feature extraction algorithms from related technologies to extract local features from each image, obtaining local features contained in each image that are related to or have discriminative power for a specific category of object (such as a geographical location). Then, considering the continuous changes in geographical locations in practical applications and the impact of manual intervention on efficiency and accuracy, this embodiment can employ an unsupervised clustering algorithm to perform cluster analysis on all the local features obtained from each image, obtaining multiple clusters, each of which can serve as a feature category. Then, according to specific rules (such as predefined identification rules or identification generation rules provided by related technologies), a globally unique category identifier within the processing range of the first image set is generated for each cluster, serving as the feature category identifier for the corresponding cluster. Finally, based on the inclusion relationship between images and local features, the affiliation relationship between the aforementioned local features and clusters, and the correspondence between clusters and feature category identifiers, a mapping relationship is established between each image and at least one feature category identifier. This method of determining feature category identifiers through clustering can improve the uniqueness, comprehensiveness, and stability of the identified feature categories. This allows the deduplication scheme to better adapt to unknown categories and unknown scenarios, enhancing the scalability and scenario adaptability of the deduplication method, thereby further improving the efficiency and stability of image deduplication.

[0065] In some embodiments, performing local feature extraction on each image in the first image set and determining the local features contained in the image includes: performing object detection on each image using a first object detection model, determining the object detection box corresponding to a preset local scene detected in the image, and performing feature extraction or vectorization on the image within the object detection box to determine the local features contained in the image.

[0066] The first object detection model is a pre-trained machine learning model capable of detecting various preset local scenes from images. Preset local scenes are pre-defined, static local scenes with discriminative power and semantic (e.g., location semantics) representations of objects of specific categories. For example, preset local scenes are determined based on Points of Interest (POIs) in location-related business scenarios (e.g., local life services, travel, navigation, etc.). Here, POIs are geographic entities with clear geographic coordinates and category attributes, including but not limited to commercial facilities, natural landscapes, and public buildings. Given that local features may not cover all POIs, preset local scenes can be set as local features corresponding to the POIs, such as shop signs, streetlights, and traffic signs.

[0067] Specifically, in this embodiment, multiple preset local scenes can be predefined, and corresponding first object detection models can be trained. Then, each image is input into the first object detection model for object detection to detect at least one object detection box containing a preset local scene. Next, features are extracted or vectorized from the image within the object detection box to generate local features corresponding to that object detection box. This allows the acquisition of local features for each image. For example, if an image contains three shop signs and two streetlights, the electronic device can determine a total of five local features corresponding to these two preset local scenes. This allows the detection range of local features to be limited by preset local scenes, making the extraction target of local features clearer and eliminating interference from features unrelated to the location scene. This improves the accuracy and understandability of subsequent clustering analysis in determining the target category, thereby further improving the accuracy of image deduplication. Furthermore, preset local scenes can narrow the range of local features, reducing the data processing volume of subsequent clustering analysis and further improving the efficiency of image deduplication.

[0068] In other embodiments, local feature extraction is performed on each image in the first image set to determine the local features contained in the image, which includes: using a second object detection model or feature extraction model to extract local features from each image to determine the local features contained in the image.

[0069] The second object detection model is a pre-trained machine learning model capable of detecting discriminative local features from images. Compared to the first object detection model, it does not require a pre-defined local scene and is more versatile. The feature extraction model is a model or algorithm with local feature extraction capabilities. Examples include scale-invariant feature transformation algorithms and deep learning-based general local feature point detection and description algorithms for extracting scale-invariant and rotation-invariant local feature points from images.

[0070] Specifically, in this embodiment, the electronic device can directly utilize the second object detection model or feature extraction model to perform local feature extraction processing on each image, and determine the output features of the model as the local features corresponding to the image. This allows for the automatic extraction of significant local regions from images without relying on prior local scene rules, saving labor costs and further improving the automation level of local feature extraction and its adaptability to unknown local scenes. This further enhances the flexibility and scalability of local feature extraction, thereby further improving the flexibility, scene adaptability, and scalability of the deduplication scheme.

[0071] It should be noted that during the local feature extraction process described above, the number of local features output by the model can be adjusted by changing the confidence level of the first or second object detection model, thereby controlling the redundancy of the deduplication results. For example, increasing the confidence level can reduce the number of local features, thus reducing the number of target images retained to some extent and lowering the redundancy of the second image set.

[0072] It should also be noted that, according to the descriptions of the foregoing embodiments, for images that do not contain local features, it can be considered that they do not contain visual information that can identify objects of a specific category (such as geographical locations), and the electronic device can discard them as redundant images.

[0073] The following are embodiments of the image deduplication device provided in this disclosure. This device and the image deduplication methods in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the image deduplication device, please refer to the embodiments of the above image deduplication methods.

[0074] Figure 3 A schematic diagram of the structure of an image deduplication device provided in an embodiment of this disclosure is shown. Figure 3 As shown, the image deduplication device 300 may include: The first image set acquisition module 310 is used to acquire a first image set containing multiple images; The second image set generation module 320 is used to perform deduplication processing on each image in the first image set based on the mapping relationship between each image and the feature category to generate a second image set; wherein, the feature category indicates the category corresponding to the local image contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each feature category.

[0075] The image deduplication apparatus provided in this disclosure discretizes each image into a feature category corresponding to one or more local images contained within it. Based on the relationship between the images and the feature categories, the deduplication problem of the first image set is transformed into an image filtering or screening problem that selects a portion of the images to cover all feature categories, thereby completing the deduplication of the first image set. This reduces the storage space of the image set while retaining the local image features of all categories, thus balancing the integrity of local image feature coverage and the utilization rate of image storage space.

[0076] In some embodiments, the second image set generation module 320 includes: The mapping relationship establishment submodule is used to establish the mapping relationship between each image and the feature category identifier; where the feature category identifier indicates the category identifier corresponding to the local features contained in the image; the local features are the features of the local image; The image deduplication submodule is used to perform deduplication on each image in the first image set based on the mapping relationships, and generate a second image set.

[0077] In some embodiments, the image deduplication submodule is specifically used for: Based on the mapping relationships, images are selected from the first image set and combined to generate multiple image combinations that meet the first condition; wherein, the first condition includes that the image combination completely covers each feature category identifier and that the images in the image combination do not overlap with each other; Select an image combination from the various image combinations that meets the second condition, and use it as the second image set; wherein the second condition includes that the number of images in the image combination is less than the number of images in the first image set.

[0078] In some embodiments, the second condition includes at least one of the following: The image composition contains the fewest images; The image quality of the image combination is higher than a preset quality threshold, and the number of images in the image combination is less than the number of images in the first image set; The image combination completely covers the feature category identifiers corresponding to the first and second features; where the first and second features are different local features among the local features.

[0079] In some embodiments, the mapping relationship establishment submodule includes: The local feature determination unit is used to extract local features from each image in the first image set and determine the local features contained in the image. The feature category identifier generation unit is used to perform cluster analysis on the local features contained in each image, determine the cluster to which the local features belong, and generate the feature category identifier corresponding to the cluster. The mapping relationship establishment unit is used to establish a mapping relationship between each image and at least one feature category identifier based on the belonging relationship between local features and clusters.

[0080] In some embodiments, the local feature determination unit is specifically used for: The first object detection model is used to perform object detection on each image, determine the object detection box corresponding to the preset local scene detected in the image, and extract or vectorize the image within the object detection box to determine the local features contained in the image; wherein, the first object detection model is a pre-trained machine learning model that is capable of detecting each preset local scene in the image; Alternatively, a second object detection model or feature extraction model can be used to extract local features from each image to determine the local features contained in the image; wherein, the second object detection model is a pre-trained machine learning model capable of detecting discriminative local features from an image.

[0081] In some embodiments, each image in the first image set contains at least one geographic location; And / or, the preset local scene is determined based on points of interest in location-related business scenarios.

[0082] In some embodiments, the first image set acquisition module 310 is specifically used for: Retrieve each image from the image database to be deduplicated; Based on the geographic location information corresponding to the images, each image is classified into a different geographic grid, generating a first image set corresponding to each geographic grid; wherein, the grid size of the geographic grid is greater than a preset distance value that is a multiple of the set distance value, and the preset distance value is the positioning deviation of the geographic location information.

[0083] The image deduplication device provided in this disclosure can execute the image deduplication method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0084] It is worth noting that in the above-described embodiments of the image deduplication device, the various modules and sub-modules are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional module / sub-module are only for easy differentiation and are not used to limit the scope of protection of this disclosure.

[0085] This disclosure also provides an electronic device that may include a processor and a memory, the memory being used to store executable instructions. The processor can be used to read the executable instructions from the memory and execute them to implement the image deduplication method described above.

[0086] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0087] like Figure 4As shown, the electronic device 400 may include a processing unit 401 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output interface (I / O interface) 405 is also connected to the bus 404.

[0088] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data.

[0089] It should be noted that, Figure 4 The illustrated electronic device 400 is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein. That is, although... Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0090] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the image deduplication method of any embodiment of this disclosure.

[0091] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the image deduplication method in any embodiment of this disclosure.

[0092] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.

[0093] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0095] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform the image deduplication method described in any embodiment of this disclosure.

[0096] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0098] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0099] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0100] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0101] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image deduplication method, characterized in that, include: Get the first image set containing multiple images; Based on the mapping relationship between each image and the feature category, the images in the first image set are deduplicated to generate a second image set; wherein, the feature category indicates the category corresponding to the local image contained in the image; the number of images in the second image set is less than the number of images in the first image set, and the second image set completely covers each of the feature categories.

2. The method according to claim 1, characterized in that, The step of deduplicating each image in the first image set based on the mapping relationship between each image and a feature category to generate a second image set includes: Establish a mapping relationship between each image and a feature category identifier; wherein, the feature category identifier indicates the category identifier corresponding to a local feature contained within the image; the local feature is a feature of the local image; Based on the mapping relationships, the images in the first image set are deduplicated to generate the second image set.

3. The method according to claim 2, characterized in that, The step of performing deduplication processing on each of the images in the first image set based on the mapping relationships to generate the second image set includes: Based on the mapping relationships, images are selected from the first image set and combined to generate multiple image combinations that satisfy the first condition; wherein, the first condition includes that the image combination completely covers each of the feature category identifiers, and that the images in the image combination do not overlap with each other; Select one image combination from each of the image combinations that satisfies the second condition as the second image set; wherein the second condition includes that the number of images in the image combination is less than the number of images in the first image set.

4. The method according to claim 3, characterized in that, The second condition includes at least one of the following: The image combination contains the fewest images; The image quality corresponding to the image combination is higher than a preset quality threshold, and the number of images in the image combination is less than the number of images in the first image set; The image combination completely covers the feature category identifiers corresponding to the first feature and the second feature; wherein the first feature and the second feature are different local features among the local features.

5. The method according to claim 2, characterized in that, The step of establishing a mapping relationship between each image and the feature category identifier corresponding to the local features contained in the corresponding image includes: Local feature extraction is performed on each image in the first image set to determine the local features contained in the image; Cluster analysis is performed on the local features contained in each of the images to determine the cluster to which the local features belong, and the feature category identifier corresponding to the cluster is generated; Based on the relationship between the local features and the clusters, a mapping relationship is established between each image and at least one feature category identifier.

6. The method according to claim 5, characterized in that, The step of extracting local features from each of the images in the first image set to determine the local features contained in the image includes: The first object detection model is used to perform object detection on each of the images, determine the object detection box corresponding to the preset local scene detected in the image, and extract or vectorize the image within the object detection box to determine the local features contained in the image; wherein, the first object detection model is a pre-trained machine learning model capable of detecting each preset local scene in the image; Alternatively, a second object detection model or feature extraction model can be used to extract local features from each of the images to determine the local features contained in the images; wherein the second object detection model is a pre-trained machine learning model capable of detecting discriminative local features from images.

7. The method according to claim 6, characterized in that, Each of the images in the first image set contains at least one geographical location; And / or, the preset local scene is determined based on points of interest in location-related business scenarios.

8. The method according to claim 7, characterized in that, The step of obtaining a first image set containing multiple images includes: Obtain each of the images in the image database to be deduplicated; Based on the geographic location information corresponding to the images, each image is classified into a different geographic grid, generating a first image set corresponding to each geographic grid; wherein, the grid size of the geographic grid is greater than a preset distance value that is a multiple of a set, and the preset distance value is the positioning deviation of the geographic location information.

9. An electronic device, characterized in that, include: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the image deduplication method according to any one of claims 1-8.

10. A computer program product, characterized in that, The computer program product is used to implement the image deduplication method according to any one of claims 1-8.