Picture processing method and apparatus

CN122821174APending Publication Date: 2026-09-25SWEET POTATO TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611188278.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,由于链式结构的连通分量是基于两两图片的相似程度串联构建而成,因此,针对处于同一链式结构上互不相邻的图片之间存在并无直接相似关系的可能,进而导致对图片的误删

Benefits of technology

[0011]根据本说明书提供的图片处理方法,通过构建连通分量,将待处理图片集合中的各待处理图片划分为若干独立子图得到至少一个目标连通分量,使得各分量内的决策互不影响,降低计算复杂度。进一步的,在确定删除节点的过程中,在连通分量内部,且不在分量间传递,进而规避了沿路径传递删除的情况,从而使得在剔除视觉冗余的情况下,保留了图片集的视觉覆盖广度。更进一步的,在本说明书中根据图片熵在连通分量中确定目标保留节点以及删除节点,在提升确定保留节点准确性的同时,还提升了所保留图片的语义丰富程度。综上,根据本说明书提供的图片处理方法,实现了对图片的精准去重操作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821174A_ABST
    Figure CN122821174A_ABST
Patent Text Reader

Abstract

The specification provides a picture processing method and device, wherein the picture processing method comprises: obtaining a picture set to be processed; constructing at least one target connected component based on picture similarity between each picture to be processed in the picture set to be processed, wherein each node in the target connected component represents a picture to be processed, an edge between nodes represents picture similarity, and each node corresponds to picture entropy; in each target connected component, determining at least one target reserved node and at least one target deleted node corresponding to each target reserved node based on picture entropy corresponding to each node, obtaining a target reserved node set and a deleted node set corresponding to each target connected component; and determining at least one target deleted picture corresponding to the picture set to be processed based on the target reserved node set and the deleted node set corresponding to each target connected component. According to the picture processing method provided in the specification, picture deduplication accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to image processing methods. This specification also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the continuous development of computer technology, people are using more and more diverse methods to deduplicate images.

[0003] In scenarios involving deduplication of large image collections, a chain-like connected component structure is typically constructed between different images to facilitate deduplication within images belonging to the same chain. However, since the connected components of the chain structure are constructed based on the similarity of pairwise images, there is a possibility that non-adjacent images within the same chain structure may not have a direct similarity relationship, leading to accidental deletion of images. Furthermore, when multiple similar images exist for the same image, a deterministic retention decision is lacking.

[0004] Therefore, how to accurately deduplicate images has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of this specification provide an image processing method. This specification also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the aforementioned problems existing in the prior art.

[0006] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising: Get the collection of images to be processed; Based on the image similarity between each image in the set of images to be processed, at least one target connected component is constructed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy; In each target connected component, based on the image entropy corresponding to each node, at least one target retention node is determined and at least one target deletion node corresponding to each target retention node is determined, thereby obtaining the target retention node set and deletion node set corresponding to each target connected component. Based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component, at least one target deleted image is determined corresponding to the set of images to be processed.

[0007] According to a second aspect of the embodiments of this specification, an image processing apparatus is provided, comprising: The acquisition unit is configured to acquire a set of images to be processed. The construction unit is configured to construct at least one target connected component based on the image similarity between each image to be processed in the set of images to be processed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy; The processing unit is configured to determine at least one target retention node and at least one target deletion node corresponding to each target retention node in each target connected component based on the image entropy of each node, thereby obtaining a set of target retention nodes and a set of deletion nodes corresponding to each target connected component. The determining unit is configured to determine at least one target image to be deleted corresponding to the set of images to be processed, based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component.

[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described image processing method.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0011] According to the image processing method provided in this specification, by constructing connected components, each image in the image set to be processed is divided into several independent subgraphs to obtain at least one target connected component. This ensures that decisions within each component do not affect each other, reducing computational complexity. Furthermore, in determining the deletion node, deletion occurs within the connected component and is not propagated between components, thus avoiding deletion along the path. This preserves the visual coverage breadth of the image set while eliminating visual redundancy. Even further, this specification determines the target retention node and deletion node in the connected component based on image entropy, improving the accuracy of retaining node determination while also enhancing the semantic richness of the retained images. In summary, the image processing method provided in this specification achieves accurate image deduplication. Attached Figure Description

[0012] Figure 1A flowchart of an image processing method according to an embodiment of this specification is shown; Figure 2 A flowchart illustrating an image processing method for image deduplication provided in one embodiment of this specification is shown. Figure 3 A flowchart illustrating a local greedy iterative processing method provided in one embodiment of this specification is shown. Figure 4 This specification shows a schematic diagram of the structure of an image processing apparatus according to an embodiment of the present specification; Figure 5 This specification shows an architecture diagram of an image processing system provided in one embodiment; Figure 6 A structural block diagram of a computing device provided according to an embodiment of this specification is shown. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0017] The image processing methods provided in this manual are applicable to the field of image processing, such as image deduplication, including deduplication of face images. It should be noted that this manual does not limit the scenarios for image deduplication.

[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0019] Image entropy can be understood as an important metric in information theory, used to measure the complexity and information content of an image's grayscale distribution. A higher entropy value indicates richer image details and greater uncertainty; a low entropy value may indicate blurriness, overexposure, or underexposure. This specification uses image entropy to determine the richness of an image, thereby providing support for image processing.

[0020] A greedy algorithm can be understood as an iterative strategy that selects the current local optimum at each step. In this specification, the local greedy selection is performed round by round within each connected component.

[0021] A connected component (CC) can be understood as a set of mutually reachable nodes in a graph. It's important to understand that this specification only uses CC to partition the graph for candidate deletions and does not propagate deletion relationships.

[0022] The keeper node can be understood as the node selected to be retained in this round of iteration, and its direct similar neighbors will be marked as deletion candidates.

[0023] Duplicate nodes (dup) can be understood as nodes that are judged to be nearly duplicates and need to be deleted.

[0024] A canonical node can be understood as a representative, retained node corresponding to a duplicate node.

[0025] A direct similarity edge can be understood as an edge between two nodes whose similarity scores meet a threshold, indicating that the two nodes have a direct near-repetition relationship; deletion decisions are only generated at both ends of such edges.

[0026] Content Application Platform: This is an application platform that integrates various types of content. The content application platform provides users with a rich variety of multimedia content, including but not limited to live streaming, on-demand video, audio, text and image information, social interaction, shopping, etc.

[0027] Published Content: This refers to the content service provided to the content application platform. Published content consists of any content pre-published by the recommending account on the platform. For example, the published content could be a product recommendation note published by the recommending account on the platform. This product recommendation note can include product images, videos, descriptions, advantages and disadvantages, etc. The product recommendation note can be text and image content, text information, videos, etc. The content platform's server stores the published content from each recommending account and sends it to browsing accounts on the content platform through a content distribution algorithm deployed on the server.

[0028] "Responding to": "Responding to" generally refers to a reaction or response to a situation, problem, request, or event. In the technical field, it can refer to a system's reaction to input information or events. It can be understood as the system or software's perception and recognition of user input or operations, and the resulting corresponding actions or behaviors.

[0029] Triggered actions refer to actions such as clicking, swiping, or other interactive methods on content application platforms, published content, live streaming pages, and other display interfaces. These actions can trigger corresponding events and display the page or information associated with the triggered action. For example, clicking the recommendation information block on a product details page can display a list of recommended accounts.

[0030] With the development of computer technology, the demand for deduplication of massive images in the field of image processing is growing, and the corresponding deduplication methods are becoming increasingly diverse.

[0031] In certain application scenarios, such as user uploads or image scraping on content publishing platforms, the resulting image sets often contain a large number of nearly duplicate images. Nearly duplicate images can be understood as multiple derived images generated from the same original image after image editing operations such as cropping, compression, adding filters, and minor color adjustments. These derived images are visually similar but not entirely identical. The large presence of these nearly duplicate images in the image database significantly reduces the diversity of content within the database, thereby impacting the user's browsing experience and the content distribution quality of the recommendation system.

[0032] To address the need for processing nearly duplicate images, deduplication methods are typically based on hash algorithms or precise feature matching. However, these methods are usually sensitive to pixel-level changes in images and cannot effectively handle nearly duplicate images generated after slight transformations such as cropping, compression, and color adjustment. Their deduplication accuracy and recall rate are difficult to meet the needs of practical applications.

[0033] To further improve the detection and processing capabilities of near-duplicate images, graph computation-based deduplication schemes have gradually become the mainstream technical approach. Graph computation-based deduplication schemes can be understood as constructing a graph structure based on the similarity between images, with images as nodes and the similarity relationships between images as edges. Then, by calculating the connected components in the graph, the near-duplicate relationships between images are characterized, and corresponding deduplication decisions are made based on these connected components.

[0034] However, the aforementioned graph-based deduplication scheme still has significant drawbacks in practical applications. Specifically, because connected components in a graph structure may contain chain-like structures (i.e., nodes A and B are similar, nodes B and C are similar, but nodes A and C are not similar), nodes A and C, whose visual content is not actually repetitive, are grouped into the same connected component during the construction of connected components. When subsequent decisions are made to retain or delete images based on connected components, since non-adjacent images within the same connected component may not actually be similar, this method of processing images within the same connected component as a whole can easily lead to the accidental deletion of dissimilar images, thereby damaging the diversity and integrity of the image collection.

[0035] Furthermore, in the aforementioned graph structures, it is common for the same node to have similar relationships with multiple other nodes simultaneously. For example, the same image may have a high degree of similarity to several other images. In this scenario, the lack of a deterministic node selection mechanism makes it difficult to determine a unique and reasonable representative node from multiple candidate nodes similar to the given node. This leads to uncertainty in the deduplication decision, affecting the consistency and reproducibility of the deduplication results.

[0036] In summary, graph deduplication schemes based on connected components have significant shortcomings in terms of the accuracy of deletion decisions, the precision of controlling the deletion range, and the determinism of the decision results. There is an urgent need for an image processing method that can precisely control the deletion propagation range, provide a deterministic decision-making mechanism, and avoid erroneous deletions caused by chain propagation.

[0037] In view of this, this specification provides an image processing method to accurately identify and remove nearly duplicate images in a large-scale image library, while retaining the most information-rich representative images and effectively avoiding excessive deletion. This specification also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0038] Figure 1 A flowchart of an image processing method according to an embodiment of this specification is shown, specifically including the following steps 102-108: Step 102: Obtain the set of images to be processed.

[0039] The image set to be processed can be understood as an image database containing a large number of image files. It is important to understand that, in this specification, the image set to be processed contains nearly duplicate images. The images to be processed can be based on user-uploaded images or images determined by the content publishing platform.

[0040] In one specific embodiment provided in this specification, a set of images to be processed for image deduplication is obtained. This set of images can be obtained from the image database corresponding to the content publishing platform. This specification does not limit the method of image acquisition.

[0041] In one specific embodiment provided in this specification, after obtaining the image set, the following method is used to filter the nearly duplicate images in the image set to be processed, and to remove duplicates from the nearly duplicate images.

[0042] Step 104: Based on the image similarity between each image in the set of images to be processed, construct at least one target connected component, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy.

[0043] Image similarity can be understood as the degree of visual similarity between different images, calculated through visual feature extraction, and is usually quantified as a numerical value. It's important to understand that the image similarity in this specification is used to construct edges in the graph structure, and the method for determining image similarity is not limited. The graph structure in this specification can be a connected subgraph. However, it should be understood that the connected components in this specification are only used to partition the candidate deletion graph and do not propagate deletion relationships.

[0044] A target connected component can be understood as a connected subgraph displayed in a graph structure, determined by the image similarity between images, and composed of independent subgraphs connected by edges. It is important to clarify that there is no connected path between any two nodes in different connected components. In this specification, it can be understood as the physical boundary used to define subsequent retention or deletion decisions. The retention decision can be understood as the decision used to determine which nodes to retain. The deletion decision can be understood as the decision used to determine which nodes to delete.

[0045] A node can be understood as a vertex in a graph structure. In this specification, nodes are used to represent independent images to be processed. In addition, each node has corresponding node attribute information. Node attribute information is used to characterize the image features of the image to be processed, such as image entropy. It is also used to characterize the node features, such as node identifiers. It should be noted that the node identifiers in this specification can be understood as being determined according to the order in which the images were added to the database. For a node corresponding to an image added earlier, its node identifier is less than that of a node corresponding to an image added later.

[0046] An edge can be understood as an undirected line connecting two nodes. In this specification, it is used to represent a near-repetition relationship between two images. The existence of an edge between two nodes determines whether it is used to determine which node to retain and which to delete. It should be clarified that for two nodes sharing the same edge, one node to retain and one corresponding node to delete can be determined. However, this specification does not impose a mandatory limitation on this. In other words, for two nodes sharing the same edge, only one node to retain or only one node to delete can be determined. The specific retention or deletion strategy is determined based on subsequent decision-making methods.

[0047] Image entropy can be understood as image information entropy, an objective physical quantity (such as grayscale entropy or color entropy) used to measure the richness of information or visual complexity of an image. In this specification, it is used to measure whether an image has retention value, in order to determine retention nodes.

[0048] In one specific embodiment provided in this specification, the image similarity between images in a set of images to be processed is determined. At least one connected component is constructed based on the image similarity. In the connected component, nodes represent images, and edges represent the image similarity between images. Further, the connected components are used to divide the images in the set of images to be processed into independent partitions. It should be understood that the connected components do not interfere with each other. It should also be clarified that a connected component can be understood as a connected subgraph. In this connected subgraph, there are edges connecting the nodes.

[0049] In one specific implementation, image pairs with similar relationships are determined based on the image similarity of each image to be processed. Based on each image in the image pairs with similar relationships, the associated images corresponding to each image are determined, thereby constructing at least one connected component. It should be understood that in a single connected component, each node corresponds to at least one edge, while nodes belonging to different connected components do not have edges between them.

[0050] Furthermore, for ease of understanding, this specification uses the following method to construct target connected components for each image in the image set to be processed.

[0051] In one specific embodiment provided in this specification, at least one target connected component is constructed based on the image similarity between the images to be processed in the set of images to be processed, including: Based on the image similarity between each image in the set of images to be processed, image pairs with an image similarity greater than a preset similarity threshold are determined; Based on image pairs with an image similarity greater than a preset similarity threshold, at least one target connected component is constructed.

[0052] The preset similarity threshold can be understood as a numerical boundary condition pre-set and fixed in the configuration file. According to the preset similarity threshold, only when the visual similarity between two images reaches or exceeds this threshold is the two images considered to have a near-repetition relationship; otherwise, they are considered visually unrelated. It should be noted that the preset similarity threshold in this manual can be flexibly set according to actual application scenarios. A higher preset similarity threshold means a higher similarity between the selected two images.

[0053] Image pairs can be understood as unordered binary tuples formed by any two combinations of images in the set to be processed. In this specification, they can be understood as highly related binary tuples that have undergone similarity calculations and whose values ​​are greater than a preset similarity threshold. It should be noted that any image pair is directly determined based on common edges, and there are no skipped edges.

[0054] In one specific embodiment provided in this specification, the original similarity scores between all pairs of images are obtained, and the similarity of each original edge is numerically compared with a static preset similarity threshold. If the image similarity is greater than or equal to the preset similarity threshold, the image pair is retained for subsequent deduplication. Conversely, if the image similarity is less than the preset similarity threshold, the edge is discarded, does not participate in any subsequent calculations, and is marked as indicating that there is no near-duplicate relationship between the two images connected by the edge. It should be noted that images marked as having no near-duplicate relationship can be directly considered as retained nodes and continue to be stored in the image library to be processed until an image with a near-duplicate relationship with them is found in the library.

[0055] In one specific implementation, the image similarity scores between the images in the image set to be processed are compared with a preset similarity threshold to determine all image pairs with similarity scores greater than the preset similarity threshold. It should be understood that the preset similarity threshold is used to filter out image pairs with a high-confidence near-repetition relationship, eliminating weakly similar edges below or equal to the threshold, thus retaining only high-confidence edges with a clear visual near-repetition relationship. This avoids accidental deletion of images due to operations such as cropping.

[0056] Then, the image pairs selected above, whose similarity is greater than a preset similarity threshold, are used as a set of legal edges. At least one target connected component is constructed based on the connected component algorithm in graph theory. In this connected component, each node represents an image to be processed, and the edges between nodes represent the image similarity between the corresponding images of the two nodes. Since the construction of the connected components is based on high-confidence edges selected by the threshold, the nodes within each constructed target connected component are connected through high-confidence similarity relationships, while there are no similar edges between different connected components that satisfy the threshold condition, thus ensuring that each connected component is physically independent.

[0057] By employing the threshold filtering described above, and through the two-step processing based on connected components, a large number of weakly similar or noisy edges are excluded from the graph structure, reducing the data scale for subsequent iterative computations. Furthermore, dividing the entire set of images to be processed into multiple independent connected components ensures that processing decisions within each component do not cross component boundaries, laying the foundation for subsequent independent and parallel execution of multiple rounds of greedy iterations within each component.

[0058] According to a specific embodiment provided in this specification, a preset similarity threshold is set to prune weak connections. This results in several closely related components, thereby reducing long-tail latency caused by data skew. Furthermore, identifying high-confidence near-duplicate pairs based on the preset similarity threshold significantly improves the accuracy of deduplication results and prevents erroneous deletion of visually dissimilar but numerically slightly higher pairs. Moreover, by setting the preset similarity threshold, this specification further precisely controls the physical boundaries of connected components when duplicate images are identified, thus achieving distributed, transitive deletion.

[0059] Furthermore, given that at least one target connected component is obtained, this specification employs the following method to determine the retained nodes and deleted nodes corresponding to each connected component.

[0060] Step 106: In each target connected component, based on the image entropy corresponding to each node, determine at least one target retention node and at least one target deletion node corresponding to each target retention node, and obtain the target retention node set and deletion node set corresponding to each target connected component.

[0061] In this context, the target retention node can be understood as the node determined to be retained in the retention decision within the connected component. It is used to represent the image that is ultimately retained and will not be physically deleted.

[0062] The target deletion node can be understood as a node determined to need to be deleted in the deletion decision within a connected component. It is used to characterize images that are redundant compared to retained nodes and should be marked or cleaned up.

[0063] The target reserved node set can be understood as a summary list of all target reserved nodes belonging to the same connected component.

[0064] The set of deleted nodes can be understood as a summary list of all target deleted nodes belonging to the same connected component.

[0065] It should be clarified that different target connected components can be subject to the same retention or deletion decision, or they can be subject to different retention or deletion decisions. When different retention or deletion decisions are made, the decision can be determined based on the specific characteristics within the target connected component. This specification uses the same retention or deletion decision as the preferred strategy.

[0066] In one specific embodiment provided in this specification, local greedy decision-making is performed within each independent connected component. Optionally, for an image in any connected component, the image entropy corresponding to each image is determined, and the nodes to be retained and deleted in the connected component are determined based on the image entropy. It should be clarified that there is a correspondence between the nodes to be retained and the nodes to be deleted in this specification. Furthermore, each deleted node corresponds to a unique node to be retained.

[0067] For ease of understanding, this specification describes the method for determining the target retained node and the target deleted node for the same connected component in the following manner.

[0068] In a specific embodiment provided in this specification, in each target connected component, based on the image entropy corresponding to each node, at least one target retention node is determined and at least one target deletion node corresponding to each target retention node is determined, thereby obtaining a set of target retention nodes and a set of deletion nodes corresponding to each target connected component, including S1062-S1066: S1062. In each target connected component, based on image similarity, determine each node to be processed corresponding to the current iteration round and at least one associated node corresponding to each node to be processed.

[0069] The current iteration round can be understood as the loop counting state during execution. Each round represents a complete operation of determining the target node to retain and the target node to delete.

[0070] A node to be processed can be understood as any node in the set of nodes that are still active at the start of the current iteration (i.e., the remaining set) and have not yet been marked as either retained or deleted. These nodes are candidates for decision-making in the current iteration.

[0071] An associated node can be understood as a node in the target connected component corresponding to the current round that is connected to a node to be processed via a direct edge. It should be understood that the associated nodes in this specification are limited to one-hop direct neighbors, excluding two-hop, three-hop, or other indirect nodes within the same connected component; and the neighboring node must also be in the current set of nodes to be processed.

[0072] In one specific embodiment provided in this specification, in each round, the original edges are filtered based on the current set of nodes to be processed. Optionally, only edges where both ends of the node are alive are retained, forming the active edges for this round. Then, an adjacency list is built for each node to be processed based on the active edges for this round. It should be understood that in each Lenz, deleted nodes and their associated edges will not be considered as nodes to be processed in this round, so that the number of nodes to be processed in different iteration rounds gradually decreases with the number of iterations.

[0073] S1064. Among each node to be processed and at least one associated node corresponding to each node to be processed, based on the image entropy corresponding to each node to be processed, determine at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node.

[0074] The target node to be retained can be understood as the node (i.e., the Keeper) that is determined to be retained in the current iteration. Its characteristic is that it has the highest retention priority among all its associated nodes, where the retention priority is determined based on image entropy.

[0075] The target deletion node can be understood as an associated node that is directly overwritten by the target retention node in the current iteration. It is marked as redundant and will be removed from the pending set after this round.

[0076] In one specific implementation provided in this specification, within the neighborhood of each active node, the node with the highest image entropy is elected by comparing the image entropy values. It should be noted that if a node has no neighbors with higher entropy, it is designated as the target retention node (Keeper) for this round. Nodes directly adjacent to the Keeper are marked as the target deletion node (Dup) for this round. It should be understood that this specification allows multiple Keepers to be generated simultaneously within the same connected component, and the deletion decisions for each Keeper are executed in parallel without interference.

[0077] In a specific embodiment provided in this specification, among each node to be processed and at least one associated node corresponding to each node to be processed, based on the image entropy corresponding to each node to be processed, at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node are determined, including: Determine the node to be processed and at least one associated node to be processed corresponding to the node to be processed, wherein the node to be processed is any one of the nodes; If the image entropy corresponding to the node to be processed is greater than or equal to the image entropy corresponding to each associated node to be processed, the node to be processed is determined as the target retention node, and each associated node to be processed is determined as the target deletion node.

[0078] In one specific embodiment provided in this specification, at least one associated node is obtained from the one-hop radiation range centered on the node to be processed. It should be clarified that at least one edge scene includes isolated nodes (with zero associated nodes). In this case, the isolated node will directly become a retained node through subsequent determination. The image entropy of the node is compared one by one with the image entropy of all its directly active neighbor nodes. If the node's entropy is not lower than that of any of its neighbor nodes, it is identified as a Keeper. Similarly, the directly active neighbor nodes directly connected to the Keeper are unconditionally marked as Dup in this round.

[0079] According to a specific implementation provided in this specification, by employing a method of greater than or equal to, a node can still be identified as a target retention node when its image entropy is equal to the image entropy of all its associated nodes. Simultaneously, its associated nodes are identified as target deletion nodes, ensuring that at least one target retention node is generated in each iteration, driving the continuous shrinking of the set of nodes to be processed. Furthermore, during the determination process, once a node is selected as a target retention node, all its directly associated nodes are simultaneously identified as nodes to be deleted. For a star topology, one iteration is sufficient to classify all nodes within the structure, significantly reducing the total number of iterations and thus lowering the overall computational complexity and network communication overhead in a distributed environment.

[0080] In one specific embodiment provided in this specification, further, for the case where the same deleted node corresponds to at least two target retained nodes, this specification provides a deterministic node retention strategy. Optionally, the method further includes: In response to the existence of at least two initial target retention nodes corresponding to the same target deletion node, determine the node attribute information corresponding to each initial target retention node, wherein the initial target retention node is any one of the target retention nodes; Based on the node attribute information corresponding to each initial target retention node, and according to the node retention strategy, a single target retention node corresponding to the target deletion node is determined from at least two initial target retention nodes.

[0081] The initial target reserved nodes can be understood as indicating that these nodes are temporarily determined to be reserved nodes, but there may be conflicts with the same Dup later, which means that there is still a possibility that they may be rejected by other Keepers.

[0082] Node attribute information can be understood as a multi-dimensional decision-making basis for resolving conflicts. It can also be understood as providing input data for subsequent node retention strategies.

[0083] The node retention strategy can be understood as a predefined, deterministic decision rule used to uniquely select the final target Keeper from multiple conflicting Keepers.

[0084] In one specific embodiment provided in this specification, during multiple iterations, when multiple initial target retention nodes simultaneously point to the same target deletion node, it is necessary to determine the unique target retention node corresponding to this scenario. Specifically, once at least two initial target retention nodes are detected to correspond to the same target deletion node, the node attribute information corresponding to each conflicting initial target retention node is extracted. This information includes, but is not limited to, image entropy, direct image similarity with the target deletion node, and node identifier. Subsequently, according to a preset node retention strategy, a unique node is determined among these conflicting initial target retention nodes as the final mapping object of the target deletion node, i.e., the single target retention node.

[0085] In one implementation, the node retention strategy follows a preset priority order. For example, image entropy is used as the primary criterion, with higher image entropy being selected first. If multiple initial target retention nodes have the same image entropy, the direct image similarity between each node and the target deletion node is further compared, with the node having the higher similarity score being selected first. If both image entropy and image similarity are the same, the node identifier is used as the final deterministic decision criterion, selecting the node with the smallest node identifier as the unique retention node. Through the above conflict normalization process, it is ensured that each target deletion node ultimately corresponds to only one definite retention node, forming a globally unique and deterministic mapping relationship, avoiding decision conflicts caused by one-to-many mappings, and ensuring the reproducibility of computation results in a distributed environment.

[0086] For ease of understanding, this manual explains the method for determining the unique target reserved node in the following way.

[0087] In one specific embodiment provided in this specification, the node attribute information includes node identifier, image similarity between adjacent nodes, and image entropy.

[0088] The node identifier can be understood as a unique and deterministic code (such as a numeric ID or UUID) for each node. In this specification, it is used to determine the priority of its corresponding node.

[0089] Optionally, based on the node attribute information corresponding to each initial target retention node, and according to the node retention strategy, a single target retention node corresponding to the target deletion node is determined from at least two initial target retention nodes, including: Compare the image entropy corresponding to each initial target retention node, and determine the initial target retention node with the largest image entropy among all initial target retention nodes as the single target retention node corresponding to the target deletion node; and / or, Compare the image similarity between each initial target retention node and the target deletion node, and determine the initial target retention node with the highest image similarity between the initial target retention node and the target deletion node as the single target retention node corresponding to the target deletion node; and / or, Compare the node identifiers corresponding to each initial target retention node, and determine the initial target retention node with the smallest node identifier as the single target retention node corresponding to the target deletion node.

[0090] In one specific embodiment provided in this specification, during multiple iterations, when multiple initial target retention nodes correspond to the same target deletion node, a conflict normalization decision is executed according to a preset node retention strategy, so as to determine the only one from the multiple initial target retention nodes as the final mapping object of the target deletion node.

[0091] In one specific embodiment provided in this specification, the node retention strategy can be at least one of the following three methods: In one specific implementation, the image entropy corresponding to each initial target retention node is compared, and the initial target retention node with the largest image entropy value is determined as the single target retention node, thereby prioritizing the retention of nodes with richer image information.

[0092] In another specific implementation, the direct image similarity between each initial target retention node and the target deletion node is compared, and the initial target retention node with the highest similarity score is determined as the single target retention node, thereby prioritizing the retention of the node that is visually closest to the node to be deleted.

[0093] In another specific implementation, the node identifiers corresponding to each initial target retention node are compared, and the initial target retention node with the smallest node identifier is determined as the single target retention node, thereby using the uniqueness of the node identifier as the basis for the final determination.

[0094] By employing at least one of the above methods, a unique and definite target retention node mapping can be generated for each target deletion node in a conflict scenario, ensuring that each deletion node ultimately belongs to only one representative node.

[0095] Furthermore, in one example, image entropy can be used as the primary criterion, with higher image entropy being prioritized. If multiple initial target retention nodes have the same image entropy, the direct image similarity between each node and the target deletion node is further compared, with the node having the higher similarity score being prioritized. If both image entropy and image similarity are the same, the node identifier is used as the final deterministic decision criterion, selecting the node with the smallest node identifier as the unique retention node. Through the above conflict normalization process, it is ensured that each target deletion node ultimately corresponds to only one definite retention node, forming a globally unique and deterministic mapping relationship. This avoids decision conflicts caused by one-to-many mappings and ensures the reproducibility of computation results in a distributed environment.

[0096] S1066. Repeat the operation of the target retained nodes and target deleted nodes corresponding to each iteration round until the iteration stop condition is met, and obtain the target retained node set and the target deleted node set corresponding to each target connected component.

[0097] The iteration stopping condition can be understood as a preset termination signal used to end the loop. This specification includes, but is not limited to, the following: the set of nodes to be processed is empty, no nodes are retained in a certain round, or the maximum number of iteration rounds is reached.

[0098] In one specific implementation provided in this specification, the results determined in the current round are accumulated into the global set, while the Keeper and Dup of the current round are removed from the set to be processed. This process is repeated until convergence. This ensures that all nodes in all components belong to either the target set of retained nodes or the set of deleted nodes, thereby achieving full coverage.

[0099] In one specific implementation, within each target connected component, based on the current set of nodes to be processed, the nodes to be processed corresponding to this iteration round are determined according to image similarity, and at least one associated node directly connected to each node to be processed is identified. It should be noted that a node to be processed refers to an active node that has not yet been marked as retained or deleted at the start of the current iteration round, and an associated node refers to other active nodes in the current active subgraph that are directly connected to the node to be processed via a one-hop edge. The scope of associated nodes is strictly limited to directly adjacent one-hop neighbors, excluding indirect nodes with two or more hops.

[0100] Then, within the local neighborhood formed by the node to be processed and its corresponding associated nodes, the decision on which nodes to retain is made based on the image entropy of each node to be processed: if the image entropy of a node to be processed is greater than or equal to the image entropy of all its associated nodes, then the node to be processed is determined as the target node to retain in this iteration, and all its associated nodes are determined as target nodes to delete. The determination of each target node to retain is independent of each other, and multiple target nodes to retain can be generated simultaneously within the same connected component, achieving parallel election.

[0101] Then, all target retention nodes and their corresponding target deletion nodes generated in this round are removed from the set of nodes to be processed, resulting in an updated set of nodes to be processed. This updated set of nodes to be processed is used as the input data for the next iteration.

[0102] Repeat the above iterative operations: selecting active edges, electing retained nodes, marking deleted nodes, and updating the set of nodes to be processed, until the preset iteration stopping conditions are met. These stopping conditions include, but are not limited to: the set of nodes to be processed being empty, no retained nodes existing in a given iteration, or the number of iterations reaching the preset maximum iteration count. When the iteration terminates, all generated target retained nodes are aggregated to form the target retained node set for each target connected component, and all generated target deleted nodes are aggregated to form the deleted node set for each target connected component. These two sets together constitute a complete classification of all image nodes.

[0103] Through the aforementioned multi-round iterative mechanism, each round of iteration executes decisions on the continuously shrinking set of nodes to be processed until all nodes are classified as either retained or deleted, thus achieving layer-by-layer stripping of the graph structure and complete coverage of all nodes.

[0104] In one specific embodiment provided in this specification, the iteration stopping condition includes at least one of the following: The number of iteration rounds meets the iteration round threshold.

[0105] The iteration round threshold can be understood as a pre-set maximum allowed number of iterations. This threshold is a static preset value and does not change dynamically with the data distribution.

[0106] In one specific implementation provided in this specification, when the iteration round meets the iteration round threshold, all remaining nodes to be processed are regarded as keep nodes (Keeper) and no further stripping is performed to ensure processing efficiency.

[0107] The set of nodes to be processed is an empty set.

[0108] The fact that the set of nodes to be processed is empty can be understood as the fact that there are no unmarked nodes in the current set of remaining nodes (remaining). This state means that all nodes have completed the classification decision of "keeping" or "deleting", and there are no unresolved nodes.

[0109] In one specific embodiment provided in this specification, when the set of nodes to be processed is empty, it indicates that all nodes have undergone sufficient decision-making, either being marked as retained nodes or marked as deleted nodes and mapped to the corresponding retained nodes. Therefore, it can be concluded that the entire image has been completely deduplicated and classified, requiring no additional processing.

[0110] There are iteration rounds in which no node is retained.

[0111] The statement that there are no retained nodes in an iteration round can be understood as follows: in a certain iteration round, after comparing the neighborhoods of all nodes to be processed, no node can meet the conditions to become a Keeper.

[0112] In one specific implementation provided in this specification, when all remaining nodes are isolated nodes (without any active edges connecting them), in each iteration, since isolated nodes have no associated nodes, they are assumed to have no higher-priority neighbors and should be selected as Keepers. However, in some extreme cases (such as when all nodes in the current round have been "pre-marked" by other nodes in the previous round but have not yet been removed), there may be idle rounds without Keepers. It should be understood that in this case, there are no near-repeating relationships in the graph, and the remaining nodes are visually unique isolated images, which should all be retained.

[0113] According to a specific implementation method provided in this specification, by introducing an iteration round threshold, the maximum number of iterations of the algorithm is locked within a preset constant, so that the time complexity is strictly controlled at O(50 × E), that is, near-linear level.

[0114] In one specific embodiment provided in this specification, after constructing at least one target connected component, the method further includes: A set of nodes to be processed is determined, wherein the set of nodes to be processed is determined based on the nodes corresponding to each target connected component; In response to determining at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node, the set of nodes to be processed is updated to obtain the updated set of nodes to be processed. Based on the updated set of nodes to be processed, the operations of determining the target nodes to retain and the target nodes to delete are repeatedly executed until the iteration stopping condition is met.

[0115] The set of nodes to be processed can be understood as the complete set of nodes that are currently active and have not yet been marked as retained or deleted. This is the input scope for each round of iterative decision-making, used to determine which nodes participate in the election in this round.

[0116] In one specific embodiment provided in this specification, after constructing at least one target connected component, an initial set of nodes to be processed is first determined to support subsequent multi-round iterative decision-making processes. This set of nodes to be processed consists of all nodes contained in each target connected component; that is, in the initial state, all nodes are in a pending state and are eligible to participate in subsequent iterative decisions. In this case, an iterative loop process is entered. In each round of iteration, the target nodes to be retained and the target nodes to be deleted for the current round are determined based on the image entropy corresponding to each node. In response to this determination result, an update operation is performed on the current set of nodes to be processed.

[0117] Furthermore, the update operation can involve batch removing all target retention nodes and all target deletion nodes identified in this round from the current set of nodes to be processed. Nodes that are not removed are retained in the set of nodes to be processed, resulting in an updated set of nodes to be processed. This updated set of nodes to be processed serves as the input data for the next iteration. Based on the updated set of nodes to be processed, the operations of identifying target retention nodes and target deletion nodes, and updating the set of nodes to be processed, are repeated until a preset iteration stop condition is met, at which point the iteration terminates. Through this iterative loop mechanism, the size of the set of nodes to be processed is continuously reduced. Each iteration makes decisions on a decreasing subset of data until all nodes are marked as retention nodes or deletion nodes, or other preset stop conditions are triggered, thereby achieving complete classification processing of all image nodes.

[0118] According to a specific implementation method provided in this specification, each node to be processed in the current iteration round and at least one associated node corresponding to each node to be processed are determined, ensuring that the deletion effect is limited to a one-hop range. This avoids the propagation path of chain deletion, effectively preventing erroneous deletion due to indirect similarity, and maximizing the preservation of visual content diversity while removing redundancy. The target retention node is determined independently based on the local neighborhood of each node. This ensures that if two far apart high-entropy nodes are separated by a low-entropy chain in a connected component, they will be selected in parallel in the same round, further reducing the number of iterations. It should be noted that the range of associated nodes in each round in this specification is calculated in real-time based on the dynamic set of nodes to be processed after being stripped from the set of nodes to be processed in the previous round, and the deletion decision is strictly limited to directly adjacent nodes and does not propagate outwards along the path, thereby preventing high-entropy nodes from excessively deleting low-entropy but visually unique images.

[0119] Based on the above, it should be clarified that when a node has no higher entropy neighbors, the node will be automatically identified as a reserved node.

[0120] Step 108: Based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component, determine at least one target deleted image corresponding to the set of images to be processed.

[0121] In one specific embodiment provided in this specification, nodes on the graph are mapped back to physical images to obtain an image deletion list.

[0122] In one specific embodiment provided in this specification, after determining at least one target image to be deleted corresponding to the set of images to be processed, the method further includes: Add deletion markers to the images to be deleted for each target.

[0123] In this context, the deletion marker can be understood as a status label attached to image metadata, database fields, or storage systems.

[0124] In one specific embodiment provided in this specification, after determining at least one target image to be deleted corresponding to the set of images to be processed, the method further includes a tag writing operation in order to convert the determination result into identifiable and executable state information. Optionally, the mapping list of target images to be deleted is traversed, and for each target image to be deleted that is determined to be a redundant image, a deletion tag is added to the metadata record, database field, or storage system tag corresponding to the image. This deletion tag serves as image state identification information, used to characterize that the image has been determined to be a redundant image to be deleted. Through this tagging operation, the logical mapping relationship output by the algorithm is persisted as the physical state information of the image, enabling subsequent image processing flows to quickly identify and locate the image to be deleted based on the tag, without having to re-execute similarity calculation or graph iteration decision. At the same time, since the deletion tag is a reversible soft tagging operation rather than a physical deletion operation, the physical file of the tagged image is still retained in its original storage location, and it is only filtered at the business query layer or presentation layer through the state identifier. Thus, while implementing the deletion decision, the flexibility of data recovery and manual review is retained, providing a reliable trigger basis and operational prerequisite for further physical deletion or archiving operations.

[0125] In one specific embodiment provided in this specification, the method further includes: In response to receiving an image processing instruction; Based on the image processing instructions, images carrying deletion marks are identified in the set of images to be processed, and these images are deleted.

[0126] In this context, an image processing instruction can be understood as receiving an external signal that explicitly instructs the user to perform a deletion operation, through means such as API calls, message queues, scheduled task triggers, or user interface operations. This instruction may include parameters (such as specifying the deletion of a specific component or a marked image from a specific time period).

[0127] In one specific embodiment provided in this specification, after identifying at least one target image to be deleted corresponding to the set of images to be processed and adding deletion marks to each target image to be deleted, the method further includes deletion command response and execution steps to perform the final physical cleanup operation in a safe and controllable manner. Specifically, by continuously listening for or waiting for externally input image processing commands, when the image processing command is received, in response to the triggering of the command, a search operation is first performed in the set of images to be processed. By querying the marking status of the images, all images carrying deletion marks are filtered out to determine the target image list for this physical deletion operation. Subsequently, a deletion operation is performed on each image in the target image list. The deletion operation includes, but is not limited to: permanently erasing the image file from the stored images to be processed, migrating the image file to the recycle bin or archive storage area, or updating the image status to physically deleted and removing it from the business index. Through the above-described "command triggering - filtering and locating - physical execution" process, the images to be deleted identified in the soft marking stage are finally physically cleaned up. In this process, the physical deletion action is not automatically triggered when the marking is completed, but strictly depends on the receipt of external instructions. This creates a controllable operation interval between logical judgment and physical execution, providing an operation window for manual review, scheduled operation or batch maintenance. It also avoids irreversible data loss due to algorithm misjudgment or system anomalies, ensuring the safety, controllability and auditability of large-scale image cleaning operations.

[0128] According to a specific implementation provided in this specification, a large graph is divided into several independent subgraphs by constructing connected components, ensuring that decisions within each component do not affect each other, thus reducing computational complexity. Furthermore, in determining the deletion node, deletion occurs within the connected components and is not propagated between components, thereby avoiding deletion propagation along the path. This allows for the preservation of the visual coverage breadth of the image set while eliminating visual redundancy. Even further, this specification determines the target retention node and deletion node in the connected components based on image entropy, improving the accuracy of determining retention nodes while also enhancing the semantic richness of the retained images.

[0129] The following is in conjunction with the appendix Figure 2 Taking the image processing method provided in this specification as an example of its application in image deduplication, the image processing method will be further explained. Among other things, Figure 2 This specification shows a flowchart of an image processing method for image deduplication according to an embodiment, specifically including the following steps 202-212: Step 202: Construct candidate deletion edges.

[0130] In one specific embodiment provided in this specification, based on the input set of similar edges, similar edges with a similarity score greater than or equal to a preset similarity threshold (i.e., a direct greedy edge threshold) are selected, and the selected similar edges constitute a candidate deletion edge set. It is important to clarify that during the construction of candidate deletion edges, the deletion decision is limited to the two nodes directly connected by the candidate edge, and is not propagated outwards along the similar edges, thus avoiding the expansion of the deletion range. This prevents the accidental deletion of nodes within the same connected component.

[0131] Step 204: Calculate the connected components.

[0132] In one specific embodiment provided in this specification, based on the obtained set of candidate edges to be deleted, a connected component algorithm is run to divide each candidate edge into multiple target connected components. It should be noted that the target connected components are only used to limit the physical scope of subsequent iterations to achieve partitioned parallel processing; they themselves do not participate in passing on deletion decisions, and the final deletion relationship is not passed along the connected path.

[0133] Step 206: Construct nodes that retain priority.

[0134] In one specific implementation provided in this specification, a reservation priority is calculated for each node.

[0135] In one scenario, the priority rule could be set so that nodes with higher image entropy are retained more frequently. If image entropy information is missing, then all nodes have the same priority. If priorities are equal, the node identifier (ID) is used as the final criterion, with smaller IDs being retained more frequently.

[0136] Step 208: Normalize the direct edges within the connected components.

[0137] Candidate edges are normalized to undirected edge form, retaining only edges within the same connected component. It should be clarified that, for cases where multiple edges exist between the same pair of nodes, the node with the highest similarity score is taken as the direct similarity score for that pair.

[0138] Step 210: Multi-round local greedy iteration.

[0139] In one specific embodiment provided in this specification, for ease of understanding, the following references are made in this specification: Figure 3 The method shown is used to explain the local greedy iteration method for any round.

[0140] Figure 3 A flowchart illustrating a local greedy iterative processing method provided in one embodiment of this specification is shown.

[0141] like Figure 3As shown, the initialization process begins: all nodes are treated as the set of nodes to be processed (remaining), and the mapping accumulation set (mapping_acc) is initialized to an empty set. Then, the iteration loop begins: In one specific implementation, it is determined whether the current set of nodes to be processed is empty.

[0142] In one case, if the current set of nodes to be processed is empty, the iteration stops or jumps to step 212.

[0143] In another scenario, if the current set of nodes to be processed is not empty, then the active edges in this round are filtered.

[0144] In one specific implementation, only edges whose two endpoints both belong to the current set of nodes to be processed are retained as the set of active edges participating in the decision-making process in this round.

[0145] Furthermore, establish adjacency relationships.

[0146] Specifically, the set of active edges is expanded from undirected edges into a bidirectional adjacency list (oriented_edges) to facilitate subsequent neighbor traversal and comparison operations.

[0147] In one specific implementation, the node to be retained in this round (keeper) is selected from the constructed adjacency relationships.

[0148] Optionally, iterate through each active node in the current set of nodes to be processed, and check if the node has a direct neighbor with higher priority. Nodes without a direct neighbor with higher priority are identified as nodes to be retained in this round.

[0149] In one specific implementation, it is determined whether the set of nodes to be retained in this round is empty.

[0150] In one scenario, if the value is empty, the iteration stops, and the remaining nodes are considered reserved nodes and do not participate in the generation of subsequent output mappings.

[0151] In another case, if the set of retained nodes is not empty, then candidate nodes for deletion are generated.

[0152] It should be clarified that during the process of generating candidate nodes for deletion, only the direct similar neighbors (one-hop neighbors) of each retained node are marked as the target deletion node (dup). The scope of this marking is strictly limited to one hop and does not extend to two hops, three hops or other nodes within the same connected component.

[0153] Furthermore, if a target deletion node is adjacent to multiple retention nodes, a representative node mapping is determined. Specifically, if a target deletion node is adjacent to multiple retention nodes, a unique representative node is selected according to a three-level rule. This three-level rule can be understood as the node retention strategy following a preset priority order. For example, image entropy is used as the primary criterion, with higher image entropy being selected first. If multiple initial target retention nodes have the same image entropy, the direct image similarity between each node and the target deletion node is further compared, with higher similarity scores being selected first. If both image entropy and image similarity are the same, the node identifier is used as the final deterministic decision criterion, and the node with the smallest node identifier is selected as the unique retention node. Afterwards, a deterministic mapping from deletion node identifiers to representative node identifiers is generated, and this mapping is accumulated into a mapping accumulation set.

[0154] In one specific implementation, the active node set is updated. Specifically, the nodes to be retained in this round and the nodes to be deleted in this round are removed from the node set to be processed, and the node set after removal is used as the input for the next iteration.

[0155] During the iterative process of the above operations, it is determined whether the current iteration round meets the stopping condition, i.e., whether the preset maximum number of iterations has been reached. If it has been reached, the iteration stops, and the remaining unprocessed nodes are all considered as retained nodes. If it has not been reached, the active edges of this round are selected to facilitate the next iteration round.

[0156] Step 212: Output the results.

[0157] In one specific embodiment provided in this specification, the mapping relationship between all deleted nodes and their representative nodes is output. This mapping relationship includes the following fields: deleted node identifier (dup_id), representative node identifier (canonical_id), associated connected component identifier (component), and direct similarity score (direct_score). Nodes not appearing in the output are considered retained nodes.

[0158] To facilitate understanding, this specification provides an explanation of the above content in conjunction with a specific embodiment.

[0159] In one specific embodiment provided in this specification, the example of a set of images to be processed containing 10 images is used for explanation.

[0160] In one specific implementation, 10 images to be processed are obtained. The nodes corresponding to each image to be processed are determined to be nodes A to J.

[0161] In one specific implementation, determining the image entropy value and preset global priority of each node includes: The image entropy value of node A is 100. In the global priority ranking, it ranks first in priority and its connected component is component 1. The image entropy value of node F is 95. In the global priority ranking, it ranks 2nd in priority and belongs to component 2 of the connected components. The image entropy value of node G is 95, which is the same as the entropy value of node F. However, since the node identifier F is smaller than the node identifier G, according to the priority rule that the node with the smaller identifier takes precedence when the entropy values ​​are the same, node F ranks second, and node G ranks third in priority. The connected component to which node G belongs also belongs to component 2. The image entropy value of node B is 90. In the global priority ranking, it ranks 4th in priority and its connected component is component 1. The image entropy value of node C is 90, which is the same as the entropy value of node B. Since the node identifier B is less than that of C, according to the priority rule that the node with the smaller identifier takes precedence when the entropy values ​​are the same, node B ranks 4th, so the priority of node C ranks 5th. The connected component to which node C belongs belongs to component 1. The image entropy value of node H is 75. In the global priority ranking, it ranks 6th in priority and belongs to component 3 of the connected components. The image entropy value of node D is 60. In the global priority ranking, it ranks 7th in priority and belongs to component 1. The image entropy value of node I is 50. In the global priority ranking, it ranks 8th in priority and belongs to component 3. The image entropy value of node E is 40. In the global priority ranking, it ranks 9th in priority and belongs to component 2 of the connected components. If the image entropy value of node J is missing, its default value is set to 0 and it participates in the priority sorting. Its priority ranking is 10th and its connected component is component 3.

[0162] In a specific embodiment provided in this specification, the similarity edge relationships between the above-mentioned nodes are further as follows: In component 1, there are similar edges between nodes A and B with a similarity score of 0.95; between nodes A and C with a similarity score of 0.85; and between nodes B and D with a similarity score of 0.90.

[0163] Component 2 contains similar edges between nodes E and F with a similarity score of 0.95; and similar edges between nodes E and G with a similarity score of 0.90.

[0164] Component 3 contains similar edges between nodes H and I with a similarity score of 0.85; and similar edges between nodes I and J with a similarity score of 0.80.

[0165] In one specific implementation, when the preset similarity threshold is set to 0.75, it can be seen that the similarity scores of each similar edge corresponding to each connected component are greater than the preset similarity threshold. Therefore, the above 7 similar edges are all included in the candidate deletion edge set and participate in the subsequent target connected component construction and iterative decision-making process.

[0166] In a specific embodiment provided in this specification, based on the definitions of each node and the edges connecting the nodes obtained above, three mutually unconnected target connected components are obtained. Specifically, component 1 of the target connected components includes nodes A, B, C, and D; component 2 of the target connected components includes nodes E, F, and G; and component 3 of the target connected components includes nodes H, I, and J. As can be seen from the above, there are no similar edges connecting the connected components that satisfy the threshold condition, thus facilitating parallel computation of subsequent iterative processing on a per-component basis, and ensuring that deletion decisions within each component do not affect each other.

[0167] Based on the above, in a specific embodiment provided in this specification, candidate deletion edges are constructed.

[0168] In one specific embodiment, after obtaining the set of images to be processed, based on the image similarity between the images to be processed, similar edges with similarity scores greater than or equal to a preset similarity threshold are retained as candidate deletion edges. Specifically, in one specific embodiment provided in this specification, candidate deletion edges include: edges between nodes in component 1, component 2, and component 3.

[0169] In one specific embodiment provided in this specification, connected components are calculated.

[0170] In one specific implementation, based on the aforementioned set of candidate edges to be deleted, a connected component algorithm is run to obtain three unconnected target connected components. Component 1 contains nodes [A, B, C, D], component 2 contains nodes [E, F, G], and component 3 contains nodes [H, I, J]. There are no similar edges connecting the connected components. Subsequent iterative processing is performed in parallel on a per-component basis, and deletion decisions within each component do not affect each other.

[0171] In one specific embodiment provided in this specification, the construction node retains priority.

[0172] In one specific implementation, the retention priority is determined based on the image entropy value corresponding to each node, with higher entropy values ​​indicating higher priority; nodes J with missing entropy values ​​participate in the sorting with a default value of 0; for nodes with the same entropy value, the node with the smaller node identifier is given priority. Therefore, for the 10 images in the image set to be processed, the corresponding global priority order is: A>F>G>B>C>H>D>I>E>J.

[0173] In one specific embodiment provided in this specification, direct edges within a normalized connected component are defined.

[0174] Candidate edges within each connected component are normalized to undirected edges, and the highest similarity score between the same pair of nodes is taken as the direct similarity score of that pair of nodes.

[0175] Based on the above, this specification adopts the following method to perform multiple rounds of local greedy iteration.

[0176] In one specific implementation, during the first iteration, the set of nodes to be processed is initialized to all nodes [A, B, C, D, E, F, G, H, I, J], and the mapping accumulation set is empty.

[0177] In one specific embodiment provided in this specification, active edges are filtered based on the current set of nodes to be processed, retaining edges whose two endpoints are both in the set of nodes to be processed. Since all nodes are initially in the set of nodes to be processed, the edges connecting each node are considered active edges in this iteration round.

[0178] In one specific embodiment provided in this specification, after obtaining the active edges, the active edges are expanded into a bidirectional adjacency list. Each active node is traversed, and it is checked whether there is a directly adjacent node with higher priority. Specifically: As mentioned above, the directly associated nodes of node A are nodes B and C, and the priorities of nodes B and C are both lower than that of node A. Therefore, node A can be retained as the node corresponding to the current round. Similarly, the directly associated node of node F is node E, and as mentioned above, the priority of node E is lower than that of node F. Therefore, node F can be retained as the node corresponding to the current round. Similarly, the directly associated node of node G is node E, and the priority of node E is lower than that of node G. Therefore, node G can be retained as the node corresponding to the current iteration round. Similarly, the directly associated node of node H is node I, and the priority of node I is lower than that of node H. Therefore, node H can be retained as the node corresponding to the current round. It should be noted that the remaining nodes B, C, D, E, I, and J are not retained as nodes corresponding to the current iteration round because they have directly associated nodes with higher priority.

[0179] Based on the comparison process described above, nodes A, F, G, and H are selected as the corresponding keepers in this iteration. Then, the direct neighbors (one-hop neighbors) of each keeper are marked as the target deleters (Dup). That is, in the current iteration, nodes B and C of node A are marked as deleters; node E of node F is marked as a deleter; node E of node G is also marked as a deleter; and node I of node H is marked as a deleter.

[0180] Furthermore, based on the above, node E is deleted and simultaneously marked by both retained nodes F and G. A unique retained node is determined according to the preset node retention strategy. Specifically, the image entropy of nodes F and G is compared. Since their entropy values ​​are the same (both 95), a next-level comparison is performed. That is, the direct similarity between nodes F and G and node E is compared. The similarity between node F and node E is 0.95, and the similarity between node G and node E is 0.90. Therefore, node F has a higher similarity, and thus node F is determined as the unique representative node of node E. Based on the above, the mapping relationships generated in the current iteration are as follows: Node B maps to node A with a direct similarity score of 0.95. Node C maps to node A with a direct similarity score of 0.85. Node E maps to node F with a direct similarity score of 0.95. Node I maps to node H with a direct similarity score of 0.85.

[0181] In summary, the nodes to be retained in the current iteration [A, F, G, H] and the nodes to be deleted in the current iteration [B, C, E, I] are removed from the set of nodes to be processed, resulting in the updated set of nodes to be processed [D, J].

[0182] Furthermore, the iterative operation continues based on the updated set of nodes to be processed. Specifically, active edges are filtered based on the updated set of nodes to be processed [D, J]. However, since node B in node B of node D has already been marked as a deleted node and removed from the set of nodes to be processed, and since node I in node I of node J has already been marked as a deleted node and removed from the set of nodes to be processed, the condition that both ends of the node set are in the set of nodes to be processed is not met. Therefore, there are no active edges in this round.

[0183] Since neither node D nor node J currently has any active neighbors and neither has any directly adjacent nodes with higher priority, both nodes D and J are selected as retained nodes in this round, and no target deletion nodes are generated.

[0184] After removing nodes D and J from the set of nodes to be processed, the set of nodes to be processed becomes empty, satisfying the iteration stopping condition, and the iteration terminates.

[0185] Based on the above, the final output deletion mapping relationship is obtained.

[0186] In one specific implementation, the deleted nodes B and C in component 1 are both mapped to the representative node A. Specifically, the representative node of the deleted node B is A, and the direct similarity score between them is 0.95, which is the similarity score of the original similar edges between node B and node A. The representative node of the deleted node C is A, and the direct similarity score between them is 0.85, which is the similarity score of the original similar edges between node C and node A. In component 1, the representative node A itself and the other node D do not appear in the list of deleted nodes, indicating that nodes A and D are both determined to be retained nodes.

[0187] In one specific implementation, node E, which was deleted in component 2, is mapped to representative node F. The direct similarity score between them is 0.95, which is the similarity score of the original similar edges between node E and node F. Although representative node G has a similar edge with node E in component 2 (similarity 0.90), after filtering, node E is mapped to node F with a similarity of 0.95. Node G itself does not appear in the list of deleted nodes, therefore, it indicates that node G is determined to be a retained node.

[0188] In one specific implementation, the deleted node I in component 3 is mapped to the representative node H, and the direct similarity score between them is 0.85, which is the similarity score of the original similar edges between node I and node H. In component 3, the representative node H itself and the remaining node J do not appear in the list of deleted nodes, indicating that both node H and node J are determined to be retained nodes.

[0189] As stated above, the deleted nodes B, C, E, and I are mapped to their respective representative nodes A, F, and H, respectively, and this mapping relationship is globally unique and definite. Therefore, nodes A, D, F, G, H, and J are all considered as retained nodes, meaning the final retained nodes are [A, D, F, G, H, J], and the deleted nodes are [B, C, E, I]. This achieves a complete classification of all 10 nodes.

[0190] According to a specific implementation provided in this specification, although there is a similar edge (similarity 0.90) between nodes B and D in component 1, after node B is marked for deletion by node A and removed, node D, as an isolated node, is automatically selected as a retained node in the second round, unaffected by the deletion status of node B. This avoids the deletion decision not being propagated outward along the similar edge, effectively preventing false deletions caused by chain propagation. Simultaneously, in component 2, node E has similar relationships with both nodes F and G. Based on the node retention strategy, node F is uniquely determined as the representative node of node E, avoiding the uncertainty caused by multiple mappings. This effectively determines the validity, accuracy, and determinism in near-duplicate image deduplication scenarios. For example, in component 1, although there is a similar edge (similarity 0.90) between B and D, because B is marked for deletion by A and removed from the set of nodes to be processed in the first round, D loses its association with B in the second round. D is retained as an isolated node, unaffected by the deletion status of B, effectively preventing false deletions caused by chain deletions, thus ensuring that the deletion decision is not propagated outward along the similar edge.

[0191] Furthermore, as in component 1, after the first round of stripping (deleting nodes B and C), the second round strips the remaining isolated node D; similarly, after the first round of stripping (deleting node I), the second round strips the remaining isolated node J. This demonstrates the layer-by-layer shrinkage processing characteristic and ensures the global uniqueness of the mapping relationship, avoiding data conflicts caused by multiple mappings. Moreover, connected components are only used for parallel computing partitioning; deletion relationships are not propagated across components. For example, the deletion decision in component 1 is completely independent of components 2 and 3, allowing each component to be processed in parallel without interference.

[0192] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 4 A schematic diagram of the structure of an image processing apparatus according to an embodiment of this specification is shown. Figure 4 As shown, the device includes: Acquisition unit 402 is configured to acquire a set of images to be processed; The construction unit 404 is configured to construct at least one target connected component based on the image similarity between each image to be processed in the set of images to be processed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy; The processing unit 406 is configured to determine at least one target retention node and at least one target deletion node corresponding to each target retention node in each target connected component based on the image entropy corresponding to each node, thereby obtaining a set of target retention nodes and a set of deletion nodes corresponding to each target connected component. The determining unit 408 is configured to determine at least one target image to be deleted corresponding to the set of images to be processed based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component.

[0193] Furthermore, the processing unit 406 is further configured as follows: In each target connected component, based on image similarity, determine each node to be processed in the current iteration round and at least one associated node corresponding to each node to be processed; In each node to be processed and at least one associated node corresponding to each node to be processed, at least one target retention node and at least one target deletion node corresponding to the current iteration are determined based on the image entropy corresponding to each node to be processed. Repeat the operations of retaining the target nodes and deleting the target nodes for each iteration round until the iteration stop condition is met, and obtain the set of retaining the target nodes and the set of deleting the target nodes for each target connected component.

[0194] Furthermore, the processing unit 406 is further configured as follows: Determine the node to be processed and at least one associated node to be processed corresponding to the node to be processed, wherein the node to be processed is any one of the nodes; If the image entropy corresponding to the node to be processed is greater than or equal to the image entropy corresponding to each associated node to be processed, the node to be processed is determined as the target retention node, and each associated node to be processed is determined as the target deletion node.

[0195] Furthermore, the processing unit 406 is also configured as follows: In response to the existence of at least two initial target retention nodes corresponding to the same target deletion node, determine the node attribute information corresponding to each initial target retention node, wherein the initial target retention node is any one of the target retention nodes; Based on the node attribute information corresponding to each initial target retention node, and according to the node retention strategy, a single target retention node corresponding to the target deletion node is determined from at least two initial target retention nodes.

[0196] Furthermore, node attribute information includes node identifier, image similarity between adjacent nodes, and image entropy; Processing unit 406 is further configured as follows: Compare the image entropy corresponding to each initial target retention node, and determine the initial target retention node with the largest image entropy among all initial target retention nodes as the single target retention node corresponding to the target deletion node; and / or, Compare the image similarity between each initial target retention node and the target deletion node, and determine the initial target retention node with the highest image similarity between the initial target retention node and the target deletion node as the single target retention node corresponding to the target deletion node; and / or, Compare the node identifiers corresponding to each initial target retention node, and determine the initial target retention node with the smallest node identifier as the single target retention node corresponding to the target deletion node.

[0197] Furthermore, the processing unit 406 is also configured as follows: A set of nodes to be processed is determined, wherein the set of nodes to be processed is determined based on the nodes corresponding to each target connected component; In response to determining at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node, the set of nodes to be processed is updated to obtain the updated set of nodes to be processed. Based on the updated set of nodes to be processed, the operations of determining the target nodes to retain and the target nodes to delete are repeatedly executed until the iteration stopping condition is met.

[0198] Furthermore, building unit 404 is further configured as follows: Based on the image similarity between each image in the set of images to be processed, image pairs with an image similarity greater than a preset similarity threshold are determined; Based on image pairs with an image similarity greater than a preset similarity threshold, at least one target connected component is constructed.

[0199] Furthermore, the iteration stopping condition includes at least one of the following: The number of iteration rounds meets the iteration round threshold. The set of nodes to be processed is an empty set; There are iteration rounds in which no node is retained.

[0200] Furthermore, unit 408 is also configured as follows: Add deletion markers to the images to be deleted for each target.

[0201] Furthermore, unit 408 is also configured as follows: In response to receiving an image processing instruction; Based on the image processing instructions, images carrying deletion marks are identified in the set of images to be processed, and these images are deleted.

[0202] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0203] See Figure 5 , Figure 5 This specification illustrates an architecture diagram of an image processing system according to one embodiment of the present specification. The image processing system may include a client 100 and a server 200. Client 100 is used to send a set of images to be processed to server 200; Server 200 is used to obtain a set of images to be processed; based on the image similarity between the images to be processed in the set of images to be processed, at least one target connected component is constructed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node has an image entropy; in each target connected component, based on the image entropy corresponding to each node, at least one target retention node is determined and at least one target deletion node corresponding to each target retention node is determined, thereby obtaining a set of target retention nodes and a set of deletion nodes corresponding to each target connected component; based on the set of target retention nodes and the set of deletion nodes corresponding to each target connected component, at least one target deletion image corresponding to the set of images to be processed is determined; and at least one target deletion image is sent to client 100. Client 100 is also used to receive at least one target image to be deleted sent by server 200.

[0204] An image processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and the server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through the server 200. In an image processing scenario, the server 200 is used to provide image processing services between the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.

[0205] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In an image processing scenario, a user can publish a data stream to server 200 through client 100, and server 200 can generate at least one target image to be deleted based on the data stream, and push at least one target image to be deleted to other clients that have established communication.

[0206] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0207] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0208] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0209] It is worth noting that the image processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the image processing methods provided in the embodiments of this specification. In other embodiments, the image processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0210] Figure 6 A structural block diagram of a computing device according to an embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0211] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0212] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0213] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0214] The processor 620 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described image processing method.

[0215] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the image processing method described above.

[0216] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0217] The above is an illustrative embodiment of a computer-readable storage medium according to this invention. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the image processing method described above. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the image processing method described above.

[0218] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0219] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described image processing method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described image processing method.

[0220] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0221] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0222] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0223] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0224] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, characterized in that, include: Get the collection of images to be processed; Based on the image similarity between each image in the set of images to be processed, at least one target connected component is constructed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy; In each target connected component, based on the image entropy corresponding to each node, at least one target retention node is determined and at least one target deletion node corresponding to each target retention node is determined, thereby obtaining the target retention node set and deletion node set corresponding to each target connected component. Based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component, at least one target deleted image is determined corresponding to the set of images to be processed.

2. The method as described in claim 1, characterized in that, In each target connected component, based on the image entropy corresponding to each node, at least one target retention node is determined, and at least one target deletion node corresponding to each target retention node is determined, thus obtaining the set of target retention nodes and the set of deletion nodes corresponding to each target connected component, including: In each target connected component, based on image similarity, determine each node to be processed in the current iteration round and at least one associated node corresponding to each node to be processed; In each node to be processed and at least one associated node corresponding to each node to be processed, at least one target retention node and at least one target deletion node corresponding to the current iteration are determined based on the image entropy corresponding to each node to be processed. Repeat the operations of retaining the target nodes and deleting the target nodes for each iteration round until the iteration stop condition is met, and obtain the set of retaining the target nodes and the set of deleting the target nodes for each target connected component.

3. The method as described in claim 2, characterized in that, In each node to be processed and at least one associated node corresponding to each node to be processed, based on the image entropy corresponding to each node to be processed, at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node are determined, including: Determine the node to be processed and at least one associated node to be processed corresponding to the node to be processed, wherein the node to be processed is any one of the nodes; If the image entropy corresponding to the node to be processed is greater than or equal to the image entropy corresponding to each associated node to be processed, the node to be processed is determined as the target retention node, and each associated node to be processed is determined as the target deletion node.

4. The method as described in claim 3, characterized in that, The method further includes: In response to the existence of at least two initial target retention nodes corresponding to the same target deletion node, determine the node attribute information corresponding to each initial target retention node, wherein the initial target retention node is any one of the target retention nodes; Based on the node attribute information corresponding to each initial target retention node, and according to the node retention strategy, a single target retention node corresponding to the target deletion node is determined from at least two initial target retention nodes.

5. The method as described in claim 4, characterized in that, Node attribute information includes node identifier, image similarity between adjacent nodes, and image entropy; Based on the node attribute information corresponding to each initial target retention node, and according to the node retention strategy, a single target retention node corresponding to the target deletion node is determined from at least two initial target retention nodes, including: Compare the image entropy corresponding to each initial target retention node, and determine the initial target retention node with the largest image entropy as the single target retention node corresponding to the target deletion node; and / or, Compare the image similarity between each initial target retention node and the target deletion node, and determine the initial target retention node with the highest image similarity between the initial target retention node and the target deletion node as the single target retention node corresponding to the target deletion node; and / or, Compare the node identifiers corresponding to each initial target retention node, and determine the initial target retention node with the smallest node identifier as the single target retention node corresponding to the target deletion node.

6. The method as described in claim 2, characterized in that, After constructing at least one target connected component, the method further includes: A set of nodes to be processed is determined, wherein the set of nodes to be processed is determined based on the nodes corresponding to each target connected component; In response to determining at least one target retention node corresponding to the current iteration round and at least one target deletion node corresponding to each target retention node, the set of nodes to be processed is updated to obtain the updated set of nodes to be processed. Based on the updated set of nodes to be processed, the operations of determining the target nodes to retain and the target nodes to delete are repeatedly executed until the iteration stopping condition is met.

7. The method as described in claim 1, characterized in that, Based on the image similarity between the images in the set of images to be processed, at least one target connected component is constructed, including: Based on the image similarity between each image in the set of images to be processed, image pairs with an image similarity greater than a preset similarity threshold are determined; Based on image pairs with an image similarity greater than a preset similarity threshold, at least one target connected component is constructed.

8. The method as described in claim 2, characterized in that, The iteration stopping condition includes at least one of the following: The number of iteration rounds meets the iteration round threshold. The set of nodes to be processed is an empty set; There are iteration rounds in which no node is retained.

9. The method as described in claim 1, characterized in that, After determining at least one target image to be deleted corresponding to the set of images to be processed, the method further includes: Add deletion markers to the images to be deleted for each target.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: In response to receiving an image processing instruction; Based on the image processing instructions, images carrying deletion marks are identified in the set of images to be processed, and these images are deleted.

11. An image processing apparatus, characterized in that, include: The acquisition unit is configured to acquire a set of images to be processed. The construction unit is configured to construct at least one target connected component based on the image similarity between each image to be processed in the set of images to be processed, wherein each node in the target connected component represents an image to be processed, the edges between nodes represent image similarity, and each node corresponds to an image entropy; The processing unit is configured to determine at least one target retention node and at least one target deletion node corresponding to each target retention node in each target connected component based on the image entropy corresponding to each node, thereby obtaining a set of target retention nodes and a set of deletion nodes corresponding to each target connected component. The determining unit is configured to determine at least one target image to be deleted corresponding to the set of images to be processed, based on the set of target retained nodes and the set of deleted nodes corresponding to each target connected component.

12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.