Sample processing methods, annotation interface display methods, devices and electronic equipment

CN116758372BActive Publication Date: 2026-08-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

[0010]根据本公开的另一方面,提供了一种计算机程序产品,包括计算机程序,该计算机程序在被处理器执行时实现根据本公开实施例中任一的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758372B_ABST
    Figure CN116758372B_ABST
Patent Text Reader

Abstract

This disclosure provides a sample processing method, a labeling interface display method, a device, and an electronic device, relating to the field of artificial intelligence, specifically to the fields of face recognition and data labeling, and applicable to smart city scenarios. The specific implementation scheme is as follows: a first candidate image set is determined from a set of sample images, where the first candidate images in the first candidate image set are sample images in the sample image set whose similarity to a control image satisfies a first similarity condition; a second candidate image set is determined from the remaining sample image sets excluding the first candidate image set, where the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set; based on the first and second candidate image sets, the image set to be labeled corresponding to the control image is determined. The scheme provided by this disclosure can improve labeling quality and is beneficial for improving the accuracy of models trained using labeled sample image sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, specifically to the fields of facial recognition and data annotation technology, and can be applied in smart city scenarios. In particular, it relates to a sample processing method, an annotation interface display method, a device, and an electronic device. Background Technology

[0002] In facial recognition systems, deep neural network-based recognition algorithms are commonly used. However, training deep neural network models requires a large number of labeled facial images as training data. The quantity and quality of the labeled data largely determine the performance of the model. Summary of the Invention

[0003] This disclosure provides a sample processing method, a labeling interface display method, an apparatus, and an electronic device.

[0004] According to one aspect of this disclosure, a sample processing method is provided, comprising: determining a first candidate image set from a sample image set, wherein the first candidate images in the first candidate image set are sample images in the sample image set whose similarity to a control image satisfies a first similarity condition; determining a second candidate image set from the remaining sample image sets other than the first candidate image set, wherein the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set; and determining an image set to be labeled corresponding to the control image based on the first candidate image set and the second candidate image set, wherein the image to be labeled in the image set to be labeled comes from either the first candidate image set or the second candidate image set.

[0005] According to another aspect of this disclosure, a method for displaying an annotation interface is provided, comprising: displaying a reference image and auxiliary annotation information of the reference image in a reference area of ​​the annotation interface; displaying a first image to be annotated and auxiliary annotation information of the first image to be annotated in a first annotation area of ​​at least one annotation area of ​​the annotation interface; and, when an annotation operation on the first image to be annotated is obtained, displaying the annotation result of the first image to be annotated in the first annotation area; wherein the image set to be annotated is obtained according to the sample processing method described in any of the above embodiments.

[0006] According to another aspect of this disclosure, a sample processing apparatus is provided, comprising: a first determining unit, configured to determine a first candidate image set from a sample image set, wherein the first candidate images in the first candidate image set are sample images in the sample image set whose similarity to a control image satisfies a first similarity condition; a second determining unit, configured to determine a second candidate image set from the remaining sample image sets other than the first candidate image set, wherein the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set; and a third determining unit, configured to determine an image set to be labeled corresponding to a control image based on the first candidate image set and the second candidate image set, wherein the image to be labeled in the image set to be labeled comes from either the first candidate image set or the second candidate image set.

[0007] According to another aspect of this disclosure, a labeling interface display device is provided, comprising: a first display unit for displaying a reference image and auxiliary labeling information of the reference image in a reference area of ​​the labeling interface; a second display unit for displaying a first image to be labeled in a set of images to be labeled and auxiliary labeling information of the first image to be labeled in a first labeling area in at least one labeling area of ​​the labeling interface; and a third display unit for displaying the labeling result of the first image to be labeled in the first labeling area when a labeling operation on the first image to be labeled is obtained, wherein the set of images to be labeled is obtained according to the sample processing method described in any of the preceding embodiments of the claims.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the embodiments of this disclosure.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0011] The sample processing method, annotation interface display method, apparatus, and electronic device provided in this disclosure determine a first candidate image set from a sample image set, wherein the first candidate images in the first candidate image set are sample images in the sample image set whose similarity to a reference image satisfies a first similarity condition; determine a second candidate image set from the remaining sample image sets other than the first candidate image set, wherein the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set; and determine the image set to be annotated corresponding to the reference image based on the first and second candidate image sets, wherein the images to be annotated in the image set to be annotated come from either the first or second candidate image sets. When annotating the same face in the sample image set, the first candidate image set with a high similarity to the face can be selected from the sample image set based on the first similarity condition, and then a second candidate image set related to the first candidate image set can be selected from the remaining sample images based on the first preset spatiotemporal condition, and the image set to be annotated that can be used to annotate the face can be determined based on the first and second image sets. Since the second candidate image and the first candidate image have a first preset spatiotemporal relationship, some sample images that are of the same face but whose similarity does not meet the first similarity requirement can be added to the image set to be labeled, thereby improving the labeling quality and helping to improve the accuracy of the model trained using the labeled sample image set.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0014] Figure 1 This is a schematic diagram of the system structure for applying the sample processing method of this disclosure embodiment;

[0015] Figure 2 This is a flowchart of a sample processing method provided according to an embodiment of the present disclosure;

[0016] Figure 3 This is another flowchart of a sample processing method provided according to an embodiment of the present disclosure;

[0017] Figure 4 This is a flowchart of a method for displaying an annotation interface according to an embodiment of the present disclosure;

[0018] Figure 5A This is a schematic diagram of an annotated interface provided according to an embodiment of the present disclosure;

[0019] Figure 5B This is a schematic diagram of the annotation interface after the selection operation is obtained according to an embodiment of the present disclosure;

[0020] Figure 6 This is another flowchart of the sample processing method and annotation interface display method provided according to an embodiment of the present disclosure;

[0021] Figure 7 This is a schematic diagram of a sample processing apparatus provided according to an embodiment of the present disclosure;

[0022] Figure 8 This is a schematic diagram of a labeling interface display device provided according to an embodiment of the present disclosure;

[0023] Figure 9 This is a block diagram of an electronic device used to implement the sample processing method of the embodiments of this disclosure. Detailed Implementation

[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0025] This disclosure provides a sample processing method, apparatus, electronic device, and storage medium. Specifically, the sample processing method and annotation interface display method of this disclosure can be executed by an electronic device, which can be a terminal or a server. The terminal can be a smartphone, tablet, laptop, smart voice interaction device, smart home appliance, wearable smart device, aircraft, smart vehicle terminal, etc. The terminal can also include a client, which can be an audio client, video client, browser client, instant messaging client, or mini-program, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0026] In related technologies, training a deep facial recognition model requires sample data consisting of facial images labeled with an ID (Identity Document). That is, each person has a unique ID, and facial images belonging to the same person need to be labeled with the same ID.

[0027] The following are common data annotation schemes in related technologies:

[0028] 1. Obtain the face image to be labeled, input it into the pre-trained face recognition model, and obtain the face features.

[0029] 2. For each face image to be labeled, facial features are used to obtain faces similar to the current face through vector nearest neighbor search. These faces are then manually labeled to determine if they belong to the same person. The similarity score provided by the model can be referenced during labeling to reduce the probability of labeling errors.

[0030] Alternatively, facial feature clustering can be used to label data within a cluster.

[0031] The solutions mentioned above can assist in annotation through similarity and clustering, thereby improving annotation efficiency. However, annotation quality heavily relies on the accuracy of the pre-trained face recognition model, and it is difficult to correct for parts where the model performs poorly during annotation. For example:

[0032] 1. Two images, A and B, of the same person have very low similarity. During the annotation process, B does not appear in the list of similar images of A, so A and B cannot be labeled with the same ID.

[0033] 2. Two facial images, C and D, of different people, are judged by the model to have a high similarity. During annotation, the annotators are influenced by the similarity score given by the model and may very well label them with the same ID.

[0034] In summary, the poor quality of the labeled sample data in related technologies affects the accuracy of the trained deep face recognition models.

[0035] To address at least one of the aforementioned problems, embodiments of this disclosure provide a sample processing method, a labeling interface display method, an apparatus, and an electronic device. The method involves determining a first candidate image set from a set of sample images, where the first candidate images in the first candidate image set are sample images in the set whose similarity to a reference image satisfies a first similarity condition. A second candidate image set is determined from the remaining set of sample images excluding the first candidate image set, where the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set. Based on the first and second candidate image sets, a set of images to be labeled corresponding to the reference image is determined, where the images to be labeled in the set of images to be labeled originate from either the first or second candidate image set. When labeling the same face in the set of sample images, the method firstly relies on the first similarity condition to select a first candidate image set with high image similarity to the face from the set of sample images. Then, based on the first preset spatiotemporal condition, a second candidate image set related to the first candidate image set is selected from the remaining sample images. Finally, based on the first and second image sets, a set of images to be labeled that can be used to annotate the face is determined. Since the second candidate image and the first candidate image have a first preset spatiotemporal relationship, some sample images that are of the same face but whose similarity does not meet the first similarity requirement can be added to the image set to be labeled, thereby improving the labeling quality and helping to improve the accuracy of the model trained using the labeled sample image set.

[0036] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0037] Figure 1 This is a schematic diagram of the system structure for applying the sample processing method of this disclosure embodiment. Please refer to... Figure 1 The system includes a terminal 110 and a server 120, etc.; the terminal 110 and the server 120 are connected via a network, such as a wired or wireless network.

[0038] The terminal 10 can be used to display a graphical user interface. This terminal interacts with the user through the graphical user interface, for example, by downloading and installing a corresponding client, by calling and running a corresponding mini-program, or by logging into a website and presenting a corresponding graphical user interface. In this embodiment, the terminal 10 can be used to acquire a sample image set. The server 20 can process the sample image set, including determining a first candidate image set from the sample image set, where the first candidate images in the first candidate image set are sample images in the sample image set whose similarity to a control image satisfies a first similarity condition; determining a second candidate image set from the remaining sample image sets other than the first candidate image set, where the second candidate images in the second candidate image set satisfy a first preset spatiotemporal condition with any first candidate image in the first candidate image set; and determining an image set to be labeled corresponding to the control image based on the first and second candidate image sets, wherein the images to be labeled in the image set to be labeled come from either the first or second candidate image set.

[0039] It should be noted that the application can be an application installed on a desktop computer, an application installed on a mobile device, or a mini-program embedded in an application.

[0040] It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the embodiments of this disclosure are not limited in any way. On the contrary, the embodiments of this disclosure can be applied to any applicable scenario.

[0041] The following is a detailed description. It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.

[0042] Figure 2 This is a flowchart of a sample processing method provided according to an embodiment of this disclosure; please refer to... Figure 2 This disclosure provides a sample processing method 200, which can be applied to smart city and security scenarios. Method 200 includes the following steps S201 to S203.

[0043] Step S201: Determine a first candidate image set from the sample image set. The first candidate image in the first candidate image set is a sample image in the sample image set whose similarity to the control image satisfies the first similarity condition.

[0044] Step S202: Determine a second candidate image set from the remaining sample image sets other than the first candidate image set. The second candidate image in the second candidate image set satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set.

[0045] Step S203: Based on the first candidate image set and the second candidate image set, determine the image set to be labeled corresponding to the comparison image, and the image to be labeled in the image set to be labeled comes from the first candidate image set or the second candidate image set.

[0046] A sample image set can include multiple sample images. In smart city or security fields, this can be derived from videos or images captured by cameras. Taking road surveillance video as an example, the video can be first processed by extracting frames at a certain rate to obtain multiple images. Then, the facial regions from these images can be extracted to form the sample image set.

[0047] It can be understood that each sample image in the sample image set records a person's facial image. In addition, each face can correspond to an ID. Labeling the sample image set means labeling the face IDs in the sample images, thereby labeling all sample images belonging to the same face with the same ID, resulting in a labeled sample image set. This sample image set can be used for training a face recognition model.

[0048] The reference image can be a face image with a known ID or one that has already been labeled. Assuming the face ID of the reference image is 001, method 200 can find as many sample images as possible with an ID of 001 in the sample image set, i.e., the set of images to be labeled. Then, the images to be labeled in the set of images to be labeled can be labeled manually or automatically, i.e., the sample images with the same ID as the reference image are labeled as 001.

[0049] Specifically, step S201 can first select a first candidate image set from the sample image set based on the comparison image and the first similarity condition. It can be understood that the first candidate image set may include at least one first candidate image, which is a sample image from the sample image set and has a high similarity to the comparison image. Therefore, step S201 can obtain a first candidate image with a high similarity to the comparison image from the sample image set through one recall.

[0050] To prevent sample images showing the same face as the control image from being missed in step S201 due to factors such as image shooting angle and distortion, step S202 can continue to perform a secondary recall of sample images in the remaining sample image set. The remaining sample image set consists of all sample images remaining after removing all first candidate images from the first candidate image set.

[0051] The sample images recalled in the second round can be second candidate images, and these second candidate images can form a second candidate image set. Additionally, when recalling second candidate images, the first candidate images in the first sample image set can be used as a reference.

[0052] For example, for each first candidate image in the first candidate image set, all sample images in the remaining sample image set that can satisfy the first preset spatiotemporal conditions with the first candidate image can be selected through the first preset spatiotemporal conditions, and these sample images can be added to the second candidate image set. After traversing each first candidate image, the second candidate image set can be obtained.

[0053] The first preset spatiotemporal condition can be a time condition, a space condition, or both time and space conditions can be met simultaneously.

[0054] It is understandable that when acquiring each sample image in the sample image set, information such as the shooting time and location of that sample image can be acquired simultaneously. The time condition can include sample images from the remaining sample image set that satisfy a preset time interval between the shooting time of the corresponding first candidate image and the image's shooting time. For example, if the first candidate image was taken at 10:00 AM on May 3, 2022, and the preset time interval can be 30 minutes, then the first preset time condition is to select sample images taken between 9:30 AM and 10:30 AM on May 3, 2022 from the remaining sample image set, and use these as the second candidate image set.

[0055] Spatial conditions may include sample images from the remaining sample image set that satisfy a preset spatial positional relationship with the shooting location of the corresponding first candidate image. For example, if the first candidate image is taken at No. B on Road A, and the preset time interval can be 500m, then the first preset time condition is to select sample images from the remaining sample image set whose shooting locations are located within a range of 500m centered on No. B on Road A, and use them as the second candidate image set.

[0056] Of course, the first preset spatiotemporal conditions can include both time and space conditions. For example, sample images taken between 9:30 and 10:30 on May 3, 2022, and taken at locations within 500m of No. B on Road A, can be selected from the remaining sample image set and used as the second candidate image set.

[0057] Then, a set of images to be labeled that includes at least one image to be labeled can be selected from the first candidate image set and the second candidate image set. For example, the first candidate image set and the second candidate image set can be used as the entire set of images to be labeled.

[0058] It is understandable that, since the first candidate image set is selected based on the similarity between it and the object image, the faces in the images in the first candidate image set are likely to belong to the same ID as the faces in the comparison image.

[0059] Furthermore, since the activity range of the same person changes over time and space, but is limited in a short period of time or space, secondary recall can utilize the first preset spatiotemporal condition between the second candidate images in the second candidate image set and the first candidate images to further select sample images from the remaining image set that may belong to the same face ID. This reduces the number of sample images belonging to the same face ID but whose similarity does not meet the first preset similarity condition, i.e., those not selected in the first recall. Consequently, the image set to be labeled can comprehensively include the vast majority of sample images belonging to the same face ID, which is beneficial to improving the quality of sample labeling and further enhancing the accuracy of the model trained using the labeled sample image set.

[0060] In some embodiments, determining the second candidate image set from the remaining sample image set other than the first candidate image set in step S202 includes: determining at least one second candidate image from the remaining sample image set that satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set and whose similarity with the control image satisfies a second similarity condition. The second candidate image set includes at least one second candidate image, and the second similarity condition is lower than the first similarity condition.

[0061] In this embodiment, in addition to using the first preset spatiotemporal conditions to perform secondary recall on the remaining sample image set, a second similarity condition can also be added to the secondary recall.

[0062] It is understood that in the second candidate image set, the second candidate image not only satisfies the first preset spatiotemporal condition with a first candidate image, but also satisfies the second similarity condition between the second candidate image and the comparison image. The second similarity condition can be lower than the first similarity condition.

[0063] The first similarity condition can be a first similarity threshold, and the second similarity condition can be a second similarity threshold that is less than the first similarity threshold.

[0064] This embodiment enables secondary recall to be performed within a short time and space range with a lower similarity condition by setting a first preset spatiotemporal condition and a second similarity condition. This supplements a portion of sample images with lower similarity as images to be labeled in the image set to be labeled. Furthermore, by setting the second similarity condition, the number and quality of the second candidate images recalled can be controlled, reducing the difficulty of subsequent labeling.

[0065] Figure 3 This is another flowchart of a sample processing method provided according to an embodiment of the present disclosure; please refer to... Figure 3 This disclosure provides a sample processing method 300, which includes the following steps S301 to S304.

[0066] Step S301: Determine a first candidate image set from the sample image set. The first candidate image in the first candidate image set is a sample image in the sample image set whose similarity to the control image satisfies the first similarity condition.

[0067] Step S302: Determine at least one intermediate candidate image from the remaining sample image set that satisfies the first preset spatiotemporal condition with any first candidate image in the first candidate image set.

[0068] Step S303: Determine at least one second candidate image from at least one intermediate candidate image whose similarity to the control image satisfies the second similarity condition, wherein the second candidate image set includes at least one second candidate image.

[0069] Step S304: Based on the first candidate image set and the second candidate image set, determine the image set to be labeled corresponding to the comparison image, and the image to be labeled in the image set to be labeled comes from the first candidate image set or the second candidate image set.

[0070] The methods of steps S301 and S304 are the same as those of steps S201 and S203 described above. For details, please refer to the above embodiments, which will not be repeated here.

[0071] Steps S302 and S303 can be specific implementations of step S202. In this embodiment, intermediate candidate images that meet the first preset spatiotemporal conditions between the first sample image and the remaining sample image set can be selected first.

[0072] For example, for each first candidate image in the first candidate image set, all sample images that can satisfy the first preset spatiotemporal conditions with the first candidate image can be selected from the remaining sample image set by the first preset spatiotemporal conditions, and these sample images can be used as intermediate candidate images.

[0073] Then, the intermediate candidate images obtained in step S302 are selected, and at least one second candidate image whose similarity with the control image meets the second similarity requirement is selected. These second candidate images can constitute a second candidate image set.

[0074] In this embodiment, intermediate candidate images are first selected from the remaining sample image set based on the first preset spatiotemporal conditions. Then, second candidate images are selected from the intermediate candidate images based on the second similarity conditions and the comparison images to obtain the second candidate image set. This can achieve the recall of the remaining sample image set and supplement some sample images for the image set to be labeled.

[0075] Of course, in other embodiments, intermediate candidate images can be selected from the remaining sample image set first by using the second similarity condition and the comparison image, and then a second candidate image can be selected from the intermediate candidate images by using the first preset spatiotemporal condition to obtain the second candidate image set. The specific selection can be made according to actual needs.

[0076] In some embodiments, determining the set of images to be labeled corresponding to the comparison image based on the first candidate image set and the second candidate image set includes: filtering the first candidate image set and the second candidate image set based on the second preset spatiotemporal conditions to obtain the set of images to be labeled.

[0077] It is understood that in this embodiment, after obtaining the first candidate image set and the second candidate image set, the whole set can be filtered to obtain the image set to be labeled.

[0078] The first candidate image set and the second candidate image set can form the total candidate image set. The total candidate image set can include multiple sample images. These sample images are more likely to have the same ID as the control image. That is, through steps S201 and S202, we can obtain as many sample images as possible that have the same ID as the control image.

[0079] Then, the overall candidate image set can be filtered according to a second preset spatiotemporal condition, which can be a correlation between time and space. For example, if the shooting time interval between two sample images in the overall candidate image set is within a first time interval, but the shooting distance is greater than the first distance interval, that is, the shooting time interval between the two sample images is short, but the shooting distance is far. Since it is impossible for a person to teleport, it can be considered that the two people are not the same person. Therefore, some sample images in the overall candidate image set can be filtered out to obtain the image set to be labeled.

[0080] For example, for each current sample image in the overall candidate image set, a first group of sample images (let's say the number is 'a') that meet the first time interval requirement can be selected. Then, a second group of sample images (let's say the number is 'b') whose shooting distance does not meet the first distance interval requirement can be selected from this group. Then, the size of 'b' and 'ab' can be compared, i.e., comparing the number of sample images in the second group with the number of remaining sample images in the first group. If 'b' is much smaller than 'ab', for example, 'b' = 2 and 'ab' = 100, it means that the second group of sample images is likely to be images that do not meet the second preset spatiotemporal condition, and therefore, the second group of sample images can be removed from the overall candidate image set. Conversely, if 'b' is much larger than 'ab', it means that the current sample image is very likely not to have the same ID as other images in the first group, so the current sample image can be removed from the overall candidate image set. Alternatively, if 'b' and 'ab' are not significantly different, to avoid accidental deletion, the current sample image can be discarded. Then, the above selection method is performed on the next sample image.

[0081] This embodiment can filter out some sample images by using the second preset spatiotemporal condition, reducing the number of sample images in the image set that obviously do not belong to the same ID. This makes the images to be labeled in the image set largely sample images with the same ID as the control image, which is beneficial to improving the quality of the image set to be labeled. In addition, it is beneficial to reduce the difficulty and processing time of subsequent labeling work.

[0082] In some embodiments, determining the first candidate image set from the sample image set in step S201 may include: determining a first similarity between the sample face features of the first sample image in the sample image set and the control face features of the control image; if the first similarity is greater than or equal to a first similarity threshold, determining the first sample image corresponding to the first similarity as the first candidate image; and obtaining the first candidate image set based at least on the first candidate image.

[0083] In this embodiment, a first candidate image set can be recalled from the sample image set by similarity calculation. Similarity represents the degree of similarity between the faces in two images. High similarity indicates that the two images are more likely to have the same ID, while low similarity indicates that the two images are likely to have different IDs.

[0084] Specifically, we can first calculate the first similarity between each sample image and the control image. The first similarity condition can be a first similarity threshold. If the first similarity is greater than or equal to the first similarity threshold, it indicates that the sample image and the control image have a high similarity and may have the same ID. In this case, the sample image can be added to the first candidate image set as the first candidate image.

[0085] Similarity calculation can quickly determine the first candidate image set from the sample image set with high accuracy.

[0086] In some embodiments, method 200 may further include: inputting a comparison image into a first face recognition model to obtain comparison face features; and inputting a set of sample images into the first face recognition model to obtain sample face features of a first sample image in the set of sample images.

[0087] It can be understood that calculating the similarity between the control image and the sample image can be specifically implemented as calculating the similarity between the sample facial features of the sample image and the control facial features of the control image.

[0088] The first face recognition model can be a pre-trained face recognition model that can extract facial features from an image.

[0089] By inputting the control image and the sample image into the first face recognition model, the control face features and the sample face features can be obtained respectively. Then, by calculating the similarity between the control face features and the sample face features, the first similarity between the control image and the sample image can be calculated. This method is simple and easy to implement, and can also be used to iteratively update the first face recognition model.

[0090] In addition, it can be understood that after obtaining the set of images to be labeled according to method 200, the set of images to be labeled can be labeled, and then the entire sample image set can be labeled. The labeled sample image set can be used to train a new face recognition model, which can be an iterative version of the first face recognition model.

[0091] In some embodiments, determining the first candidate image set from the sample image set in step S201 includes: determining the image set to be clustered, the image set to be clustered including the sample image set and the control image; clustering the image set to be clustered to obtain a target image cluster containing the control image; and determining the remaining sample images in the target image cluster, excluding the control image, as the first candidate image set.

[0092] It's understandable that, besides obtaining the first candidate image set through similarity calculation, clustering can also be used. Clustering also uses a first similarity condition to group the image set containing the control image and all sample images; images with high similarity are grouped together. Then, from the resulting clusters, the cluster containing the control image (i.e., the target image cluster) is selected.

[0093] The remaining sample images in the target image cluster, excluding the control image, constitute the first candidate image set. It can be understood that the essence of clustering is also to group images based on similarity; that is, the first candidate images obtained must also satisfy the first similarity condition with the control images.

[0094] This embodiment uses clustering to recall the sample image set, thereby finding the first candidate image set with high similarity to the control image, and the accuracy is high.

[0095] Alternatively, it can be understood that during clustering, the sample face features of the sample image and the reference face features of the reference image can be clustered. That is, the face features corresponding to the image can be obtained through the first face recognition model mentioned above, and then clustering can be performed.

[0096] Figure 4 This is a flowchart of a method for displaying an annotation interface according to an embodiment of the present disclosure; Figure 5A This is a schematic diagram of an annotated interface provided according to an embodiment of this disclosure. Please refer to... Figure 4 and Figure 5A This embodiment provides a labeling interface display method 400, which includes the following steps S401 to S403.

[0097] Step S401: Display the reference image 511 and the auxiliary annotation information 512 of the reference image 511 in the reference area 510 of the annotation interface 500a.

[0098] Step S402: In the first annotation area 520 of at least one annotation area of ​​the annotation interface 500a, the first image 521 to be annotated in the image set to be annotated and the auxiliary annotation information 522 of the first image 21 to be annotated are displayed.

[0099] In step S403, upon obtaining the annotation operation on the first image to be annotated 521, the annotation result 524 of the first image to be annotated 521 is displayed in the first annotation region 520. The set of images to be annotated is obtained according to the sample processing method described in any of the above embodiments.

[0100] It is understandable that after obtaining the set of images to be labeled, the images can be labeled manually, for example, as shown in the labeling interface 500a.

[0101] The annotation interface 500a may include a reference area 510 and at least one annotation area (eight are shown in the figure). The reference area 510 may be displayed above, to the left, to the right, below, etc. of at least one annotation area.

[0102] Each annotation area can display an image to be annotated and its related information. Specifically, taking the first annotation area 520 as an example, it is one of these annotation areas. The first annotation area 520 can display an image to be annotated, namely the first image to be annotated 521, and the auxiliary annotation information 522 of the first image to be annotated 521. Of course, it can also display the first annotation control 523 of the first image to be annotated 521.

[0103] Users can annotate the first image to be annotated based on the reference image, the auxiliary annotation information of the reference image, the first image to be annotated, and the auxiliary annotation information of the first image to be annotated.

[0104] When the annotation result of the first image to be annotated cannot be confirmed through the face image, auxiliary annotation information can be used for annotation.

[0105] In some embodiments, the auxiliary annotation information for the comparison image includes at least one of the background image of the comparison image, the shooting time of the comparison image, and the shooting location of the comparison image; the auxiliary annotation information for the first image to be annotated includes at least one of the background image of the first image to be annotated, the shooting time of the first image to be annotated, and the shooting location of the first image to be annotated.

[0106] The background image can be a screenshot from the video corresponding to the first image to be labeled, and it can include the clothing and surrounding people and objects of the face image.

[0107] It is understandable that the control image and some sample images can also be extracted from the same video, and the control image is an already labeled image. When the labeling result of the first image to be labeled cannot be confirmed through the face image (control image and first image to be labeled), the shooting time, shooting location and background image can be used to make a judgment.

[0108] For example, if the clothing and appearance of the person in the comparison image and the first image to be labeled are the same, then they can be identified as having the same ID, given that the shooting time, location, and surrounding environment are the same.

[0109] Of course, supplementary annotation information may include at least one of the following: shooting time, shooting location, and background image.

[0110] Furthermore, multiple images to be labeled within multiple labeling areas can be sorted chronologically, allowing images captured around the same time to be located close together for easy reference during labeling. In addition to referencing the reference image and its auxiliary labeling information, users can also refer to other successfully labeled images with the same ID as the reference image, along with their auxiliary labeling information, for a comprehensive assessment. For example… Figure 5AThe diagram shows eight images to be labeled. After the first four images are labeled, the labels for the remaining four can be labeled by referring to the labels for the first four images, thereby further improving the quality of the labeling.

[0111] For example, users can use the following tagging clues to help determine whether it is the same ID:

[0112] 1. Assuming two images are taken on the same day, you can determine whether they belong to the same person by looking at their clothing.

[0113] 2. The images to be labeled on the annotation interface can be arranged in chronological order. If the images have changed from short hair to long hair to short hair again within a few days, they are likely not from the same person.

[0114] 3. Judge by the people around the baby. For example, it may be difficult to distinguish the baby's face, but if the adults around the baby in the two pictures are the same and the baby's face is also similar, it can be assumed that they are the same person.

[0115] It is understandable that a high similarity score given by the model may affect the user's annotation results for different people. However, in this embodiment, auxiliary annotation information can help users filter out images with high similarity but not belonging to the same ID, so that users can quickly identify images to be annotated that belong to the same ID as the comparison image. This has low dependence on facial features, thereby improving annotation efficiency and quality.

[0116] Additionally, after the user confirms the annotation results, the annotation results can be displayed on the annotation interface through annotation operations. For example... Figure 5A The first annotation control 523 may include three buttons: "same", "different" and "uncertain". The annotation operation can be the clicking operation of these three buttons.

[0117] "Same" indicates that the first image to be labeled and the reference image have the same ID. After the user confirms that the two are the same, they can click the "Same" button, and then the labeling result will be displayed as "Same" in the first labeling area.

[0118] Similarly, "different" indicates that the first image to be labeled and the reference image do not have the same ID. After the user confirms that the two are different, they can click the "different" button, and then the labeling result will be displayed as "same" in the first labeling area.

[0119] "Uncertain" indicates that it is uncertain whether the first image to be labeled and the reference image have the same ID, that is, the user's judgment result is uncertain. At this time, you can click the "Uncertain" button, and then the labeling result will be displayed as "Uncertain" in the first labeling area.

[0120] After the user selects the corresponding button, the corresponding annotation result can be displayed in the first annotation area 520, for example, in the upper right corner of the first annotation image.

[0121] Understandable. Figure 5A Each image to be labeled corresponds to a first labeling control. In other embodiments, the labeling interface can be set with a first labeling control. By first selecting the image to be labeled and then operating the first labeling control, the labeling operation of multiple images to be labeled can also be realized.

[0122] This embodiment displays auxiliary annotation information for both the comparison image and the first image to be annotated on the annotation interface, which allows users to quickly determine whether two images are of the same person without having to carefully compare facial features, thereby improving annotation efficiency and quality.

[0123] Figure 5B This is a schematic diagram of the annotation interface after obtaining the selection operation according to an embodiment of this disclosure; please refer to... Figure 5B In some embodiments, method 400 further includes: displaying the background image 530 of the first image to be annotated on the annotation interface when a selection operation on the first image to be annotated is obtained.

[0124] The selection operation can be a double-click, a single click, or other similar operation on the first image to be labeled. By selecting the first image to be labeled, its corresponding background image can be opened.

[0125] It is understandable that the auxiliary annotation information displayed in the annotation interface 500a may include the shooting time and shooting location. In order to display multiple images to be annotated in the annotation interface 500a, the background image may not be displayed in the annotation interface 500a. Instead, the annotation interface 500b may be opened by selection, and the background image 530 of the first image to be annotated may be displayed in the annotation interface 500b.

[0126] This embodiment can display multiple images to be annotated on the annotation interface, which makes it convenient for users to comprehensively refer to the reference image and its auxiliary annotation information, the annotated images to be annotated and their auxiliary annotation information, and the unannotated images to be annotated and their auxiliary annotation information, and to annotate the remaining unannotated images to be annotated. When needed, users can click to view the background image of the image to be annotated, which makes it convenient for users to obtain the background image in time to assist in the annotation, which helps the accuracy of the annotation results and further improves the quality of annotation.

[0127] It is understood that this embodiment takes manual annotation as an example. In other embodiments, for the current image to be annotated in the annotation set, the system can also automatically annotate the image set to be annotated. Automatic annotation can also annotate the current image to be annotated based on the reference image, the auxiliary annotation information of the reference image, the annotated image to be annotated, and the auxiliary annotation information of the annotated image to be annotated (if there is an annotated image to be annotated), the current image to be annotated, and the auxiliary annotation information of the current image to be annotated.

[0128] Furthermore, automatic annotation conditions can be set in the automatic annotation process. These conditions can be referenced from the annotation clue settings used in manual annotation.

[0129] Figure 6 This is another flowchart of the sample processing method and annotation interface display method provided according to an embodiment of this disclosure; please refer to... Figure 6 In one specific embodiment, method 600 can achieve higher quality face recognition data annotation at a faster speed. It can be applied to smart city and security scenarios. Method 600 includes the following steps S601 to S606.

[0130] Step S601: Obtain the face image to be labeled (sample image in the sample image set), its corresponding background image (background image), shooting time, and shooting location.

[0131] Step S602: Use a pre-trained face recognition model (first face recognition model) to obtain face features (including the reference face features of the reference image and the sample face features of the sample image).

[0132] Step S603: Based on the preset similarity thresholds (first similarity condition and second similarity condition), recall is performed within the entire dataset (sample image set) and a short spatiotemporal range (first preset spatiotemporal condition) (obtaining the first candidate image set and the second candidate image set).

[0133] Step S604: Filter the recall results according to the spatiotemporal conditions (second preset spatiotemporal conditions).

[0134] Step S605: The filtered results (the set of images to be labeled) are imported into the labeling system, and the background image, shooting time, and shooting location information are displayed in the labeling system (labeling interface).

[0135] Step S606: Manual labeling.

[0136] It is understood that the method 600 provided in this embodiment can use information such as time, space, and background image (auxiliary annotation information) to assist the annotator in annotation, and is an improvement on the similarity data annotation method in related technologies.

[0137] In the security field, besides the face image (sample image), the complete frame containing the face image can usually be obtained, becoming the background image. The background image contains information about the target person's body, companions, etc. Additionally, the time and location where the face image was captured are usually also obtainable.

[0138] The above-mentioned auxiliary annotation information can help us complete annotation better and faster during the data preparation and annotation execution stages.

[0139] 1. Data preparation stage:

[0140] In addition to recalling images across the entire dataset (determining the first candidate image set), since most pedestrians have limited activity range in a short period of time, recall can also be performed within a short spatiotemporal range with a lower similarity threshold to supplement a portion of candidate images with lower similarity (determining the second candidate image set).

[0141] Based on the spatiotemporal relationship (second preset spatiotemporal condition), assuming that the time interval between two images is short and the shooting distance is far apart, they can be considered not to be the same person. A batch of candidate images (a portion of sample images from the first and second candidate image sets) are filtered out using a preset threshold condition to obtain the set of images to be labeled.

[0142] 2. Annotation execution phase:

[0143] The annotation page (annotation interface) displays background image, shooting time, and shooting location information (auxiliary annotation information), and sorts the candidate image list (images to be annotated in the image set) by time. When the label of a face to be annotated is difficult to determine, annotators can use the following clues to determine whether two faces to be annotated are the same:

[0144] 1) Assuming that the two images are taken on the same day, we can determine whether they belong to the same person by looking at their clothing.

[0145] 2) Arranged in chronological order, if the images show changes from short hair to long hair and back to short hair within a few days, they are most likely not from the same person.

[0146] 3) Judge by the people around the baby. For example, it may be difficult to distinguish the baby's face, but if the adults around the baby in the two pictures are the same and the baby's face is also similar, it can be assumed that they are the same person.

[0147] Based on the above clues, annotators can quickly determine whether two images are of the same person without having to carefully compare facial features, thus improving annotation efficiency and quality.

[0148] This embodiment improves the quality and efficiency of face recognition data annotation by using auxiliary information. Although this embodiment also uses a pre-trained face recognition model, it reduces the quality problems of annotated data caused by the poor performance of the pre-trained model to a certain extent by using short-term spatiotemporal low similarity threshold recall and auxiliary information to help annotators make judgments.

[0149] Figure 7 This is a schematic diagram of a sample processing apparatus provided according to an embodiment of the present disclosure; please refer to... Figure 7 This disclosure also provides a sample processing apparatus 700, which includes:

[0150] The first determining unit 701 is used to determine a first candidate image set from the sample image set, wherein the first candidate image in the first candidate image set is a sample image in the sample image set whose similarity to the comparison image satisfies the first similarity condition;

[0151] The second determining unit 702 is used to determine a second candidate image set from the remaining sample image sets other than the first candidate image set, wherein the second candidate image in the second candidate image set satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set.

[0152] The third determining unit 703 is used to determine the set of images to be labeled corresponding to the comparison image based on the first candidate image set and the second candidate image set, wherein the images to be labeled in the set of images to be labeled come from the first candidate image set or the second candidate image set.

[0153] In some embodiments, the second determining unit 702 is further configured to: determine, from the remaining sample image set, at least one second candidate image that satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set and whose similarity with the control image satisfies a second similarity condition, wherein the second candidate dataset includes at least one second candidate image and the second similarity condition is lower than the first similarity condition.

[0154] In some embodiments, the second determining unit 702 is further configured to: determine at least one intermediate candidate image from the remaining sample image set that satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set; and determine at least one second candidate image from the at least one intermediate candidate image that satisfies a second similarity condition with a control image, wherein the second candidate image set includes at least one second candidate image.

[0155] In some embodiments, the third determining unit 703 is further configured to: filter the first candidate image set and the second candidate image set based on the second preset spatiotemporal conditions to obtain the image set to be labeled.

[0156] In some embodiments, the first determining unit 701 is further configured to: determine a first similarity between the sample face features of the first sample image in the sample image set and the control face features of the control image; if the first similarity is greater than or equal to a first similarity threshold, determine the first sample image corresponding to the first similarity as a first candidate image; and obtain a first candidate image set based at least on the first candidate image.

[0157] In some embodiments, the apparatus 700 further includes: a feature unit, configured to input a reference image into a first face recognition model to obtain reference face features; and input a set of sample images into the first face recognition model to obtain sample face features of a first sample image in the set of sample images.

[0158] In some embodiments, the first determining unit 701 is further configured to: determine an image set to be clustered, the image set to be clustered including a sample image set and a control image; cluster the image set to be clustered to obtain a target image cluster containing the control image; and determine the remaining sample images in the target image cluster, excluding the control image, as a first candidate image set.

[0159] Figure 8 This is a schematic diagram of a labeling interface display device according to an embodiment of this disclosure; please refer to... Figure 8 This embodiment also provides a labeling interface display device 800, which includes:

[0160] The first display unit 801 is used to display the comparison image and auxiliary annotation information of the comparison image in the comparison area of ​​the annotation interface.

[0161] The second display unit 802 is used to display a first image to be labeled and auxiliary labeling information of the first image to be labeled in a first labeling area of ​​at least one labeling area of ​​the labeling interface.

[0162] The third display unit 803 is used to display the annotation result of the first image to be annotated in the first annotation area when an annotation operation on the first image to be annotated is obtained; wherein the set of images to be annotated is obtained according to the sample processing method described in any of the above embodiments of the claims.

[0163] In some embodiments, the apparatus 800 further includes a background image display unit, configured to display a background image of the first image to be labeled on the labeling interface when a selection operation on the first image to be labeled is obtained.

[0164] In some embodiments, the auxiliary annotation information for the comparison image includes at least one of the background image of the comparison image, the shooting time of the comparison image, and the shooting location of the comparison image; the auxiliary annotation information for the first image to be annotated includes at least one of the background image of the first image to be annotated, the shooting time of the first image to be annotated, and the shooting location of the first image to be annotated.

[0165] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0166] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0167] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0168] This disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any of the above embodiments.

[0169] This disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method of any of the above embodiments.

[0170] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the above embodiments.

[0171] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0172] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded into random access memory (RAM) 903 from storage unit 908. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0173] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0174] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as sample processing methods and annotation interface display methods. For example, in some embodiments, the sample processing methods and annotation interface display methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the sample processing methods and annotation interface display methods described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the sample processing method and the annotation interface display method by any other suitable means (e.g., by means of firmware).

[0175] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0176] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0179] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0180] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0181] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0182] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A sample processing method, comprising: A first candidate image set is determined from the sample image set, wherein the first candidate image in the first candidate image set is a sample image in the sample image set whose similarity to the control image satisfies the first similarity condition; A second candidate image set is determined from the remaining sample image sets other than the first candidate image set. The second candidate image in the second candidate image set satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set, and the similarity with the comparison image satisfies a second similarity condition. The second candidate image set includes at least one second candidate image, and the second similarity condition is lower than the first similarity condition. The first preset spatiotemporal condition includes at least one of time condition and space condition. The first similarity condition is a first similarity threshold, and the second similarity condition is a second similarity threshold that is lower than the first similarity threshold. Based on the first candidate image set and the second candidate image set, a set of images to be labeled corresponding to the comparison image is determined, and the images to be labeled in the set of images to be labeled come from the first candidate image set or the second candidate image set.

2. The method according to claim 1, wherein, Determining a second candidate image set from the remaining sample image sets other than the first candidate image set further includes: Determine at least one intermediate candidate image from the remaining sample image set that satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set; At least one second candidate image is determined from the at least one intermediate candidate image that satisfies a second similarity condition with the control image, and the second candidate image set includes the at least one second candidate image.

3. The method according to claim 1, wherein, Based on the first candidate image set and the second candidate image set, the set of images to be labeled corresponding to the comparison image is determined, including: Based on the second preset spatiotemporal conditions, the first candidate image set and the second candidate image set are filtered to obtain the image set to be labeled.

4. The method according to any one of claims 1-3, wherein, The first candidate image set is determined from the sample image set, including: Determine the first similarity between the sample facial features of the first sample image in the sample image set and the control facial features of the control image; If the first similarity is greater than or equal to the first similarity threshold, the first sample image corresponding to the first similarity is determined as the first candidate image; The first candidate image set is obtained based at least on the first candidate image.

5. The method according to claim 4, further comprising: The comparison image is input into the first face recognition model to obtain the comparison face features; The sample image set is input into the first face recognition model to obtain the sample face features of the first sample image in the sample image set.

6. The method according to any one of claims 1-3, wherein, The first candidate image set is determined from the sample image set, including: A set of images to be clustered is determined, the set of images to be clustered including the sample image set and the control image; Cluster the set of images to be clustered to obtain a target image cluster that includes the reference image; The remaining sample images in the target image cluster, excluding the control image, are determined as the first candidate image set.

7. A method for displaying a labeling interface, comprising: The comparison area of ​​the annotation interface displays the comparison image and its auxiliary annotation information. In at least one annotation area of ​​the annotation interface, the first annotation area displays the first image to be annotated in the image set to be annotated and the auxiliary annotation information of the first image to be annotated; If a labeling operation is obtained on the first image to be labeled, the labeling result of the first image to be labeled is displayed in the first labeling area; The image set to be labeled is obtained by the sample processing method according to any one of claims 1-6.

8. The method according to claim 7, further comprising: If a selection operation is received on the first image to be annotated, the background image of the first image to be annotated is displayed on the annotation interface.

9. The method according to claim 7, wherein, The auxiliary annotation information of the comparison image includes at least one of the background image of the comparison image, the shooting time of the comparison image, and the shooting location of the comparison image; The auxiliary annotation information for the first image to be annotated includes at least one of the background image of the first image to be annotated, the shooting time of the first image to be annotated, and the shooting location of the first image to be annotated.

10. A sample processing apparatus, comprising: The first determining unit is configured to determine a first candidate image set from the sample image set, wherein the first candidate image in the first candidate image set is a sample image in the sample image set whose similarity to the comparison image satisfies a first similarity condition; The second determining unit is configured to determine a second candidate image set from the remaining sample image sets other than the first candidate image set. The second candidate image in the second candidate image set satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set, and the similarity with the comparison image satisfies a second similarity condition. The second candidate image set includes at least one second candidate image, and the second similarity condition is lower than the first similarity condition. The first preset spatiotemporal condition includes at least one of time condition and space condition. The first similarity condition is a first similarity threshold, and the second similarity condition is a second similarity threshold that is lower than the first similarity threshold. The third determining unit is used to determine the set of images to be labeled corresponding to the comparison image based on the first candidate image set and the second candidate image set, wherein the images to be labeled in the set of images to be labeled come from the first candidate image set or the second candidate image set.

11. The apparatus according to claim 10, wherein, The second determining unit is further configured to: Determine at least one intermediate candidate image from the remaining sample image set that satisfies a first preset spatiotemporal condition with any first candidate image in the first candidate image set; At least one second candidate image is determined from the at least one intermediate candidate image that satisfies a second similarity condition with the control image, and the second candidate image set includes the at least one second candidate image.

12. The apparatus according to claim 10, wherein, The third determining unit is also used for: Based on the second preset spatiotemporal conditions, the first candidate image set and the second candidate image set are filtered to obtain the image set to be labeled.

13. The apparatus according to any one of claims 10-12, wherein, The first determining unit is further configured to: Determine the first similarity between the sample facial features of the first sample image in the sample image set and the control facial features of the control image; If the first similarity is greater than or equal to the first similarity threshold, the first sample image corresponding to the first similarity is determined as the first candidate image; The first candidate image set is obtained based at least on the first candidate image.

14. The apparatus of claim 13, further comprising: The feature unit is used to input the comparison image into the first face recognition model to obtain the comparison face features; The sample image set is input into the first face recognition model to obtain the sample face features of the first sample image in the sample image set.

15. The apparatus according to any one of claims 10-12, wherein, The first determining unit is further configured to: A set of images to be clustered is determined, the set of images to be clustered including the sample image set and the control image; Cluster the set of images to be clustered to obtain a target image cluster that includes the reference image; The remaining sample images in the target image cluster, excluding the control image, are determined as the first candidate image set.

16. A labeling interface display device, comprising: The first display unit is used to display the comparison image and the auxiliary annotation information of the comparison image in the comparison area of ​​the annotation interface; The second display unit is used to display a first image to be labeled in the image set to be labeled and auxiliary labeling information of the first image to be labeled in a first labeling area of ​​at least one labeling area of ​​the labeling interface; The third display unit is used to display the annotation result of the first image to be annotated in the first annotation area when an annotation operation on the first image to be annotated is obtained; The image set to be labeled is obtained by the sample processing method according to any one of claims 1-6.

17. The apparatus of claim 16, further comprising: The background image display unit is used to display the background image of the first image to be labeled on the labeling interface when a selection operation on the first image to be labeled is obtained.

18. The apparatus according to claim 16, wherein, The auxiliary annotation information of the comparison image includes at least one of the background image of the comparison image, the shooting time of the comparison image, and the shooting location of the comparison image; The auxiliary annotation information for the first image to be annotated includes at least one of the background image of the first image to be annotated, the shooting time of the first image to be annotated, and the shooting location of the first image to be annotated.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN112948614A

  • Image annotation method, device, equipment and medium

    CN113569888A