Artificial intelligence-based image retrieval methods, devices, equipment, and storage media

By performing multiple similarity matching and expanded retrieval on the facial features of the search image, the problem of low accuracy in low-quality blurry image retrieval was solved, achieving higher recall and accuracy.

CN115481276BActive Publication Date: 2025-10-31BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211153215.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-10-31
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

In existing technologies, image retrieval methods based on urban road snapshots do not achieve high accuracy in retrieving low-quality, blurry images.

Method used

By extracting facial features from the image to be searched, target facial features are obtained. Based on the similarity between the target facial features and the facial features of each image in the facial feature database, a first-recall image set is retrieved from the facial feature database. Positive sample images are then determined based on the similarity between the image to be searched and the first-recall images to expand the search images. Images are then retrieved again using the facial features of the positive sample images to obtain a second-recall image set. Finally, the search result is determined from the two recall results.

Benefits of technology

It improved the recall and precision of the images to be searched, thus enhancing the image retrieval effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481276B_ABST
    Figure CN115481276B_ABST
Patent Text Reader

Abstract

This disclosure provides an image retrieval method, apparatus, device, and storage medium based on artificial intelligence, relating to the field of artificial intelligence, specifically image recognition and video analysis technologies, applicable to scenarios such as smart cities, urban governance, and emergency management. The specific implementation scheme is as follows: Facial features of the image to be searched are extracted to obtain target facial features; based on the target facial features, recall images corresponding to the image to be searched are retrieved from a facial feature database to obtain a primary recall image set; based on the similarity between the image to be searched and each recalled image in the primary recall image set, positive sample images are determined from the primary recall image set; based on the facial features of the positive sample images, recall images corresponding to the positive sample images are retrieved from the facial feature database to obtain a secondary recall image set; the retrieval result of the image to be searched is determined from the primary and secondary recall image sets, improving the recall rate and recall accuracy of the image to be searched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, specifically image recognition and video analysis technologies, which can be applied in scenarios such as smart cities, urban governance, and emergency management. Background Technology

[0002] In the construction of smart cities, urban road camera checkpoints are becoming increasingly common. Based on the images captured by various cameras on urban roads, image search and other methods are used to retrieve the images to be searched, which is conducive to the construction of smart cities. Summary of the Invention

[0003] This disclosure provides an image retrieval method, apparatus, device, and storage medium based on artificial intelligence.

[0004] According to one aspect of this disclosure, an artificial intelligence-based image retrieval method is provided, comprising:

[0005] The facial features of the image to be searched are extracted to obtain the target facial features;

[0006] Based on the similarity between the target facial features and the facial features of each image in the facial feature database, the corresponding recall image is retrieved from the facial feature database to obtain a recall image set; the facial feature database contains facial features of multiple images;

[0007] Based on the similarity between the image to be searched and each recalled image in the primary recall image set, positive sample images are determined from the primary recall image set;

[0008] Based on the facial features of the positive sample image and the similarity between the facial features of each image in the facial feature database, the recall image corresponding to the positive sample image is recalled from the facial feature database to obtain a secondary recall image set;

[0009] The retrieval results of the image to be searched are determined from the primary recall image set and the secondary recall image set.

[0010] According to another aspect of this disclosure, an artificial intelligence-based image retrieval device is provided, comprising:

[0011] The first feature extraction module is used to extract facial features from the image to be searched, thereby obtaining the target facial features;

[0012] The first image recall module is used to recall the image corresponding to the image to be searched from the face feature library based on the similarity between the target face feature and the face features of each image in the face feature library, thereby obtaining a recall image set; the face feature library contains the face features of multiple images;

[0013] The sample supplementation module is used to determine positive sample images from the primary recall image set based on the similarity between the image to be searched and each recalled image in the primary recall image set;

[0014] The second image recall module is used to recall the recall image corresponding to the positive sample image from the face feature library based on the similarity between the facial features of the positive sample image and the facial features of each image in the face feature library, so as to obtain a secondary recall image set.

[0015] The image retrieval module is used to determine the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set.

[0016] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the artificial intelligence-based image retrieval methods described in this disclosure.

[0020] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform any of the artificial intelligence-based image retrieval methods described in this disclosure.

[0021] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the artificial intelligence-based image retrieval method described in any one of this disclosure.

[0022] This disclosure improves the recall rate and recall precision of the images to be searched.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0024] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0025] Figure 1 This is a schematic diagram of an artificial intelligence-based image retrieval method according to the present disclosure;

[0026] Figure 2 This is another schematic diagram of the artificial intelligence-based image retrieval method according to this disclosure;

[0027] Figure 3 This is a schematic diagram illustrating a method for determining positive sample images according to this disclosure;

[0028] Figure 4 This is a schematic diagram illustrating an implementation method for determining search results from recalled images according to this disclosure;

[0029] Figure 5 This is a schematic diagram illustrating an implementation method for determining the recall score of recalled images according to this disclosure;

[0030] Figure 6 This is a schematic diagram of the framework of the AI-based image retrieval method according to this disclosure;

[0031] Figure 7 This is a schematic diagram of an artificial intelligence-based image retrieval device according to the present disclosure;

[0032] Figure 8 This is a block diagram of an electronic device used to implement the artificial intelligence-based image retrieval method according to the embodiments of this disclosure. Detailed Implementation

[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0034] With the development of technology, urban road surveillance cameras are becoming increasingly widespread, accumulating a large amount of captured image information in urban security networks. In the construction of smart cities, images captured by various cameras on urban roads can be searched using methods such as image search, which is beneficial to the construction of smart cities. For example, in public security scenarios, the image to be searched or the clue image can be searched in the image database composed of captured images to retrieve images with high confidence levels with the image to be searched, i.e., face search, so as to obtain more image clues, which is beneficial to the construction of smart cities and urban governance.

[0035] In public security scenarios, this technology extracts facial features from each captured image in the image database and stores these features in a face vector feature library. Then, when searching for an image, the facial features of the search image are extracted and compared one-to-N with the facial features of each captured image in the face vector feature library. The similarity between the facial features of the search image and those of each captured image in the library is calculated, and the top M images with similarity greater than a threshold are selected as the search results. Here, N represents the number of captured images in the face vector feature library, and M represents the number of search result images.

[0036] However, in practical applications, the initial images to be searched are usually blurry images with low image quality, which makes the image accuracy of the search results obtained by image search not high.

[0037] To address the aforementioned issues, this disclosure provides an AI-based image retrieval method. The method extracts facial features from the image to be searched, obtaining target facial features. Based on the similarity between the target facial features and the facial features of various images in a facial feature database, recall images corresponding to the image to be searched are retrieved from the database, resulting in a primary recall image set. Then, based on the similarity between the image to be searched and the recall images in the primary recall image set, positive sample images are determined from the primary recall image set to expand the search image pool. Further, based on the similarity between the facial features of the positive sample images and the facial features of various images in the facial feature database, recall images corresponding to the positive sample images are retrieved from the database, resulting in a secondary recall image set. Finally, the retrieval result for the image to be searched is determined from both the primary and secondary recall image sets. By selecting the retrieval result from the two recall results, the recall rate and accuracy of the image to be searched are improved.

[0038] The image retrieval method based on artificial intelligence provided in the embodiments of this disclosure will be described in detail below.

[0039] The AI-based image retrieval method provided in this disclosure can be applied to electronic devices, such as server devices, smart terminal devices, etc. The AI-based image retrieval method provided in this disclosure can also be applied in scenarios such as smart cities, urban governance, and emergency management.

[0040] See Figure 1 , Figure 1 A flowchart illustrating an artificial intelligence-based image retrieval method provided in this disclosure includes the following steps:

[0041] S101, extract the facial features of the image to be searched to obtain the target facial features.

[0042] When performing image search or face search, a face detection and recognition model is used to extract facial features from the search image to obtain the target face features. These facial features can be face vector features, and the face detection and recognition model can be trained using training sample images and the facial features of those images.

[0043] S102, based on the similarity between the target face features and the face features of each image in the face feature database, recall the corresponding recall image from the face feature database to obtain a recall image set.

[0044] In this embodiment of the disclosure, for images captured by cameras on urban roads, a face detection and recognition model can be used to extract facial features from each image, obtaining facial features for each image. The facial features of each image, along with the image identifiers of each image, are stored in a face feature database. This face feature database contains facial features of multiple images and image identifiers for each image.

[0045] In one example, a facial feature database can also be called a facial vector database. This database could be an index of Faiss (Facebook AI Similarity Search, a tool developed by Facebook's AI (Artificial Intelligence) team for large-scale similarity retrieval problems), or other self-built indexes, capable of quickly retrieving a predetermined number of target objects with similar vectors. There is a correspondence between the facial feature database and the captured image database; that is, the facial features and image identifiers of each image contained in the facial feature database correspond to the captured images in the captured image database, which consists of images captured by cameras on city roads.

[0046] The process involves calculating the target facial features of the image to be searched and comparing them with the facial features of each image in the facial feature database. This yields the similarity scores for each feature. Based on these similarities, a predetermined number of images with similarities greater than a set threshold and ranking highly in the facial feature database are quickly retrieved from the database, along with their corresponding image identifiers. Finally, based on these image identifiers, the corresponding recall image is determined from the captured image database, resulting in a single recall image set. The feature similarity can be calculated using cosine similarity or Euclidean distance between facial features, and the threshold and the number of similar features can be set according to specific requirements.

[0047] For example, by setting the threshold to 0.6 and the quantity to 100, the facial features and corresponding image identifiers of the top 100 images with a feature similarity greater than 0.6 can be quickly retrieved from the facial feature database. The 100 recall images corresponding to the image to be searched are determined from the captured image database to obtain a recall image set.

[0048] S103, Based on the similarity between the image to be searched and each recalled image in the primary recall image set, determine the positive sample image from the primary recall image set.

[0049] In one example, the cosine similarity or Euclidean distance between the image to be searched and each recalled image in a single recall image set is calculated to obtain the image similarity between the image to be searched and each recalled image. The P recalled images whose image similarity is greater than a preset threshold, or whose image similarity is greater than the preset threshold and whose ranking is the highest, are determined as positive sample images. The preset threshold and the value of P can be set according to actual needs, where P is an integer greater than or equal to 1.

[0050] To save computational resources, the similarity between the target facial features of the image to be searched and the facial features of each recalled image can be determined as the similarity between the image to be searched and each recalled image.

[0051] S104. Based on the facial features of the positive sample images and the similarity between them and the facial features of each image in the facial feature database, recall images corresponding to the positive sample images are retrieved from the facial feature database to obtain a secondary recall image set.

[0052] Positive sample images are determined from the recall images corresponding to the image to be searched. The facial features of the recall images are stored in the facial feature database, which in turn can determine the facial features of the positive sample images.

[0053] When there are multiple positive sample images, for each positive sample image, based on the similarity between the facial features of that positive sample image and the facial features of each image in the facial feature database, a secondary recall image corresponding to that positive sample image is retrieved from the facial feature database to obtain a secondary recall image set. For example, when there are multiple positive sample images, the recall images corresponding to each positive sample image can be stored in the same secondary recall image set, or they can be stored separately in different secondary recall image sets. Specifically, the process of retrieving the secondary recall image corresponding to that positive sample image from the facial feature database based on the similarity between the facial features of that positive sample image and the facial features of each image in the facial feature database to obtain a secondary recall image set can refer to the implementation process of step S102, and will not be elaborated further in this embodiment.

[0054] The positive sample image and the image to be searched describe the same person. The positive sample image is used to recall images again, which expands the recall results of the image to be searched and improves the recall rate of the image to be searched.

[0055] S105, determine the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set.

[0056] In one example, for the recalled images contained in the primary and secondary recall image sets, they can be sorted based on the similarity between each recalled image and the image that recalled it, calculated when the recalled image was recalled. Then, a set number of the top-ranked recalled images are determined as the retrieval results of the image to be searched.

[0057] Weights can be assigned to the recalled images in the primary and secondary recall image sets respectively. Then, based on the weight of each recalled image and the similarity between each recalled image and the image that recalled it, the score of each recalled image is calculated. The recalled images are then sorted according to the score, and the top-ranked number of recalled images are determined as the retrieval results of the image to be searched.

[0058] In this embodiment, positive sample images are determined from the primary recall image set of the image to be searched to augment the image sample. These augmented positive sample images are then used for further image recall, thereby improving the recall rate of the image to be searched. Based on the primary recall image set of the image to be searched and the secondary recall image set of the positive sample images, the retrieval result is determined from the two recall results, improving the recall accuracy of the image to be searched. In other words, both the recall rate and recall accuracy of the image to be searched are improved.

[0059] In one possible implementation, see Figure 2 , Figure 2 A flowchart illustrating another AI-based image retrieval method provided in this disclosure includes the following steps:

[0060] S201, Extract the facial features of the image to be searched to obtain the target facial features.

[0061] S202, based on the similarity between the target face features and the face features of each image in the face feature library, recall the corresponding recall image from the face feature library to obtain a recall image set; the face feature library contains the face features of multiple images.

[0062] The implementation process of steps S201-S202 can be referred to the implementation process of steps S101-S102 above, and will not be repeated here in this embodiment.

[0063] S203, extract the attribute features of the image to be searched to obtain the target attribute features.

[0064] An attribute extraction model is pre-trained based on training sample images and their attribute features. This pre-trained model is then used to extract attribute features from the search image, yielding the target attribute features.

[0065] In one possible implementation, the target attribute features may include at least one of the following: whether the face in the image is wearing a mask, whether the face is wearing glasses, the age of the face, the gender of the face, the brightness of the face, the blur of the face, the pitch angle of the face, and the rotation angle of the face. The brightness of the face can be the brightness of the image itself.

[0066] In this embodiment of the disclosure, attribute features such as whether the face in the image to be searched is wearing a mask, whether the face is wearing glasses, the age of the face, the gender of the face, the brightness of the face, the blur of the face, the pitch angle of the face, and the rotation angle of the face are extracted. These attribute features are used to more accurately determine the quality of the image to be searched and to provide a variety of positive sample images corresponding to the image to be searched.

[0067] S204. Based on the target attribute features, determine the image quality value of the image to be searched.

[0068] Based on the target attribute features, it is determined whether the image to be searched is a clear image with high image quality or a blurry image with low image quality. For example, if the target attribute features of the image to be searched include features such as a face wearing a mask or a large degree of blurriness, it indicates that the image to be searched is a blurry image with low image quality. Conversely, if the target attribute features of the image to be searched include features such as a face not wearing a mask or a small degree of blurriness, it indicates that the image to be searched is a clear image with high image quality.

[0069] In one possible implementation, determining the image quality value of the image to be searched based on target attribute features may include:

[0070] The target attribute features are encoded to obtain attribute feature codes; the attribute feature codes are input into a pre-trained face quality discrimination model to predict face quality and obtain the image quality value of the image to be searched; wherein, the pre-trained face quality discrimination model is trained based on the sample attribute feature codes and the image quality values ​​corresponding to the sample attribute feature codes.

[0071] The face quality discrimination model can score the quality of the search image. The higher the image quality value, the higher the image quality of the image being retrieved, that is, the clearer the image. For example, the image quality value of the search image can be a value between 0 and 1. The closer the image quality value is to 1, the higher the image quality of the image being retrieved. Of course, it can also be set to other ranges of values.

[0072] In this embodiment of the disclosure, a face quality discrimination model is pre-trained, and the face quality discrimination model is used to score the quality of the search image so as to accurately judge the quality of the search image based on the attribute features of the search image.

[0073] S205, acquires the shooting information of the image to be searched.

[0074] In one possible implementation, the shooting information may include at least one of the following: the type of camera used to capture the image, the shooting time, and license plate information in the image.

[0075] The type of image capturing equipment can be, for example, a vehicle-mounted camera or a person-mounted camera. A vehicle-mounted camera refers to a camera used to capture images of vehicles, while a person-mounted camera refers to a camera used to capture images of people. Capture information may also include: the camera's identification mark, and the latitude and longitude of the capturing equipment.

[0076] In this embodiment of the disclosure, the shooting information of the image to be searched is obtained in order to provide a variety of positive sample images corresponding to the image to be searched.

[0077] Steps S201-S202, S203-S204, and S205 can be executed synchronously or asynchronously.

[0078] S206. Based on image quality value, target attribute features, shooting information, and similarity between the image to be searched and each recalled image in the first recall image set, determine a preset number of positive sample images from the first recall image set.

[0079] In one example, priorities can be set for image quality value, target attribute features, shooting information, and the similarity between the image to be searched and each recalled image in the initial recall image set. Different priorities have different filtering conditions. Based on the priority order, a preset number of positive sample images are determined from the initial recall image set according to the different filtering conditions. Alternatively, a preset number of positive sample images can be determined randomly from the initial recall image set based on the image quality value, target attribute features, shooting information, and the filtering conditions corresponding to the similarity between the image to be searched and each recalled image in the initial recall image set. The preset number can be set according to actual needs.

[0080] S207. Based on the facial features of the positive sample images and the similarity between them and the facial features of each image in the facial feature database, recall images corresponding to the positive sample images are retrieved from the facial feature database to obtain a secondary recall image set.

[0081] S208, determine the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set.

[0082] The implementation process of steps S207-S208 can be referred to the implementation process of steps S104-S105 above, and will not be repeated here in this embodiment.

[0083] In this embodiment, positive sample images are determined from the primary recall image set of the image to be searched based on the image quality value, target attribute features, shooting information, and similarity between the image to be searched and each recalled image in the primary recall image set. This expands the diversity of positive samples for the image to be searched. The expanded positive sample images are then used for image recall again, thereby improving the recall rate of the image to be searched. Based on the primary recall image set of the image to be searched and the secondary recall image set of the positive sample images, the retrieval result is determined from the two recall results, improving the recall accuracy of the image to be searched. In other words, the recall rate of the image to be searched is improved while the recall accuracy is also improved.

[0084] In one possible implementation, such as Figure 3 As shown, the process of determining a preset number of positive sample images from a primary recall image set based on image quality values, target attribute features, shooting information, and the similarity between the image to be searched and each recalled image in the primary recall image set can include:

[0085] S301, obtain the similarity between the image to be searched and each recalled image in the first recall image set, and obtain the similarity of each image.

[0086] To save computational resources, the similarity between the image to be searched and each recalled image in a single recall image set can be the similarity between the target facial features of the image to be searched and the facial features of each recalled image.

[0087] S302, traverse each recalled image in the recalled image set once, and if the shooting information of the image to be searched includes the first specified information, add the recalled images with image similarity greater than the first threshold and shooting information including the second specified information to the positive sample image set.

[0088] In this embodiment, for images captured by cameras on urban roads or captured images in a captured image database, a face detection and recognition model is used to extract facial features from each captured image to obtain facial features of each captured image; an attribute extraction model is used to extract attribute features from each captured image to obtain attribute features of each captured image; and the attribute features of each captured image are encoded, and a face quality discrimination model is used to predict the quality of each captured image to obtain image quality values ​​of each captured image, thus obtaining the shooting information of each captured image. Further, the facial features, attribute features, image quality values, shooting information, and image identifiers of each captured image are stored in a face feature database. The captured image database, the face feature database, and the aforementioned face feature database have a corresponding relationship and are interconnected; the information of each captured image corresponds one-to-one, and they are used together to store the captured images and their corresponding feature information. Furthermore, when recalling the image corresponding to the image to be searched, it is possible to know the facial features, attribute features, image quality value, and shooting information of the recalled image.

[0089] The first specified information corresponds to the second specified information. For example, if the first specified information is that the image capturing device type is a vehicle-mounted camera, then the corresponding second specified information is that the image capturing device type is a person-mounted camera. The first threshold can be set according to actual needs.

[0090] For example, by traversing all recalled images in the recalled image set once, if the image to be searched includes vehicle / truck camera information, recalled images with an image similarity greater than a first threshold and whose image information includes human / truck camera information are added to the positive sample image set, thus achieving positive sample supplementation that complements vehicle / truck and human / truck images. Optionally, the image identifier of the recalled image can be added to the positive sample image set, and correspondingly, the positive sample image set contains the image identifier of the positive sample image.

[0091] During the process of traversing all recalled images in a single recall image set, if the number of positive sample images added to the positive sample image set reaches a preset number, the traversal stops; otherwise, all recalled images in the single recall image set are traversed, and recalled images that meet the conditions (the image to be searched includes the first specified information, the image similarity is greater than the first threshold, and the image to be searched includes the second specified information) are added to the positive sample image set. If, after traversing all recalled images in the single recall image set, the number of positive sample images added to the positive sample image set still does not reach the preset number, then step S303 is executed.

[0092] S303, when the number of positive sample images in the positive sample image set is less than the preset number, traverse each recalled image in the recalled image set once, and add the recalled images whose image similarity is greater than the second threshold, whose image quality value is opposite to the image quality value of the image to be searched, and which are not currently positive sample images to the positive sample image set.

[0093] For example, when the number of positive sample images in the positive sample image set is less than a preset number, the system iterates through all recalled images in the recall image set. If the image quality value of the image to be searched indicates a high-quality image, recalled images with an image similarity greater than a second threshold, an image quality value indicating a low-quality image, and not currently a positive sample image are added to the positive sample image set. Alternatively, when the number of positive sample images in the positive sample image set is less than a preset number, the system iterates through all recalled images in the recall image set. If the image quality value of the image to be searched indicates a low-quality image, recalled images with an image similarity greater than a second threshold, an image quality value indicating a high-quality image, and not currently a positive sample image are added to the positive sample image set, thus achieving positive sample supplementation that complements high-quality and low-quality images. The second threshold can be set according to actual needs.

[0094] During the process of traversing all recalled images in a single recall image set, if the number of positive sample images added to the positive sample image set reaches a preset number, the traversal stops; otherwise, all recalled images in the single recall image set are traversed, and recalled images that meet the conditions (image similarity greater than the second threshold, image quality value opposite to the image quality value of the image to be searched, and currently not a positive sample image) are added to the positive sample image set. If, after traversing all recalled images in the single recall image set, the number of positive sample images added to the positive sample image set still does not reach the preset number, then step S304 is executed.

[0095] S304, when the number of positive sample images in the positive sample image set is less than the preset number, traverse each recalled image in the recalled image set once, and if the target attribute feature of the image to be searched includes the first specified attribute, add the recalled image with image similarity greater than the third threshold, attribute feature including the second specified attribute, and currently not a positive sample image to the positive sample image set.

[0096] The first specified attribute corresponds to the second specified attribute. For example, if the first specified attribute is that the face in the image is wearing a mask, the corresponding second specified attribute is that the face in the image is not wearing a mask; or, if the first specified attribute is that the face in the image is wearing glasses, the corresponding second specified attribute is that the face in the image is not wearing glasses, and so on. The third threshold can be set according to actual needs.

[0097] For example, when the number of positive sample images in the positive sample image set is less than a preset number, the recall images in the recall image set are traversed once. If the target attribute feature of the image to be searched includes a face wearing a mask in the image, the recall images with an image similarity greater than the third threshold, attribute features including a face not wearing a mask in the image, and not currently positive sample images are added to the positive sample image set, so as to achieve positive sample supplementation with complementary attribute features, such as images with and without masks.

[0098] During the process of traversing all recalled images in a single recall image set, if the number of positive images added to the positive sample image set reaches a preset number, the traversal stops; otherwise, all recalled images in the single recall image set are traversed, and recalled images that meet the conditions (the target attribute features of the image to be searched include the first specified attribute, the image similarity is greater than the third threshold, the attribute features include the second specified attribute, and the image is not currently a positive sample image) are added to the positive sample image set. If, after traversing all recalled images in the single recall image set, the number of positive sample images added to the positive sample image set still does not reach the preset number, then step S305 is executed.

[0099] S305, when the number of positive sample images in the positive sample image set is less than the preset number, traverse each recalled image in the recalled image set once, and add the recalled images whose image similarity is between the fourth threshold and the fifth threshold and which are not currently positive sample images to the positive sample image set until the number of positive sample images in the positive sample image set reaches the preset number.

[0100] The fifth threshold varies with the similarity between the identified preceding positive sample images and the image to be searched. The fourth threshold can be set according to actual needs, and the initial value of the fifth threshold can also be set according to actual needs.

[0101] For example, the preset quantity is 3, the number of positive sample images in the positive sample image set is 1, the fourth threshold is 0.65, and the initial value of the fifth threshold is 0.95. After iterating through each recalled image in the recall image set once, if the image similarity between the current recalled image and the image to be searched is 0.9 (0.9 is between 0.65 and 0.95), and the current recalled image is not a positive sample image, then the current recalled image is added to the positive sample image set, and the number of positive sample images in the positive sample image set becomes 2. The fifth threshold is updated to the determined similarity of 0.9 between the previous positive sample image and the image to be searched. Continuing to iterate through each recalled image in the recall image set once more, if the image similarity between the current recalled image and the image to be searched is 0.85 (0.85 is between 0.65 and 0.9), and the current recalled image is not a positive sample image, then the current recalled image is added to the positive sample image set, and the number of positive sample images in the positive sample image set becomes 3. The iteration stops, thus achieving random supplementation of positive samples.

[0102] In this embodiment of the disclosure, based on the image to be searched and the recalled image, positive samples are sequentially supplemented from complementary scenarios of vehicle and card images, complementary scenarios of high-quality and low-quality images, complementary scenarios of attribute features such as wearing a mask and not wearing a mask, and random supplementation scenarios. This achieves diversified positive sample expansion, making the images recalled based on the image to be searched and the positive sample images more accurate.

[0103] In one possible implementation, such as Figure 4 As shown, the process of determining the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set in step S105 may include:

[0104] S401, deduplicatize each recalled image in the primary recall image set and the secondary recall image set to obtain a candidate image set.

[0105] The primary recall image set contains images recalled based on the search image, while the secondary recall image set contains images recalled based on positive sample images. Positive sample images are obtained by augmenting the search image. In other words, both the primary and secondary recall image sets contain images retrieved based on the search image. Therefore, the primary and secondary recall image sets may contain duplicate images. Consequently, the recalled images in both sets can be aggregated and deduplicated to obtain a candidate image set, which contains multiple candidate images.

[0106] S402, determine the recall count for each candidate image in the candidate image set.

[0107] Candidate images may be recalled by the image to be searched, or they may be recalled by different positive sample images. When summarizing and deduplicating the recalled images in the first and second recall image sets, the number of times each candidate image is recalled can be counted.

[0108] For example, if a candidate image is only recalled by the image to be searched, the recall count of the candidate image is 1. If another candidate image is recalled by both the image to be searched and 2 positive sample images, the recall count of the candidate image is 3.

[0109] S403, determine the recall score for each candidate image in the candidate image set.

[0110] In one example, for each candidate image, the recall score is determined based on the image similarity between the candidate image and the image that recalled the candidate image when the candidate image is recalled. For instance, the average of the image similarities corresponding to the candidate image when the candidate image is recalled can be used as the recall score for the candidate image.

[0111] S404, sort the candidate images based on the recall score and the number of times each candidate image has been recalled.

[0112] In one possible implementation, the recall score of each candidate image is used as the primary parameter, and the number of recalls for each candidate image is used as the secondary parameter to sort the candidate images.

[0113] For example, candidate image a has a recall score of 0.98 and a recall count of 3, candidate image b has a recall score of 0.98 and a recall count of 2, and candidate image c has a recall score of 0.96 and a recall count of 2. The ranking of the candidate images is then a, b, c.

[0114] In this embodiment of the disclosure, the recall score of the candidate image is used as the primary parameter and the number of recalls is used as the secondary parameter to sort the candidate images, so as to quickly and accurately determine the retrieval results of the image to be searched.

[0115] S405, the number of candidate images with recall scores greater than the sixth threshold and ranked at the top are determined as the retrieval results of the image to be searched.

[0116] The target quantity is set according to actual needs.

[0117] In this embodiment of the disclosure, duplicate images are deduplicated in the primary and secondary recall image sets to obtain a candidate image set containing multiple candidate images. Then, based on the recall count and recall score of each candidate image, the retrieval result of the image to be searched is determined, thereby improving the accuracy of the retrieval result.

[0118] In one possible implementation, such as Figure 5 As shown, the process of determining the recall score of each candidate image in the candidate image set in step S403 above may include:

[0119] S501, obtain the recall similarity of each candidate image in the candidate image set.

[0120] Here, recall similarity represents the similarity between the candidate image and the image that recalled the candidate image.

[0121] S502, for each candidate image, determine the preset dimensional attribute features corresponding to that candidate image.

[0122] In one possible implementation, the preset dimensional attribute features include: the spatial distance between the candidate image and the target image, whether the license plate information contained in the candidate image is the same as that contained in the target image, the shooting time difference between the candidate image and the target image, the image quality value of the candidate image, and the image quality value of the target image. The target image is the image from which the candidate images are recalled.

[0123] In this embodiment of the disclosure, a preset dimension attribute feature corresponding to each candidate image is determined so as to accurately predict the similarity weighting value of each candidate image.

[0124] S503, based on preset dimensional attribute features, predict the similarity weighting value of the candidate image to obtain the predicted weighting value.

[0125] In one possible implementation, preset dimensional attribute features are input into a pre-trained weighted adjustment model to obtain the predicted weighted value corresponding to the recall similarity of the candidate image; wherein, the pre-trained weighted adjustment model is trained based on the preset dimensional attribute features of the sample and the weighted value corresponding to the preset dimensional attribute features of the sample.

[0126] For example, for each candidate image, the preset dimensional attribute features corresponding to the candidate image are input into a pre-trained weighted model to predict the weighted value, thus obtaining the predicted weighted value corresponding to the recall similarity of the candidate image. In one example, the range of the predicted weighted value can be -0.1 to 0.1.

[0127] In this embodiment of the disclosure, a weighted model is trained using the preset dimension attribute features of the samples and the corresponding weighted values ​​of the preset dimension attribute features of the samples. The weighted model is then used to predict the predicted weighted values ​​corresponding to the recall similarity of the candidate images, so as to accurately correct the recall similarity of the candidate images.

[0128] S504. Using the predicted weighting value, the recall similarity of the candidate image is corrected to obtain the corrected similarity.

[0129] For example, candidate image a is recalled by image A to be retrieved, positive sample image B, and positive sample image C, with corresponding recall similarities of 0.9, 0.9, and 0.95, respectively. The prediction adjustment values ​​corresponding to the recall similarity of candidate image a are 0.02, -0.01, and 0.03, respectively. Then the corrected similarities of candidate image a are 0.92, 0.89, and 0.98, respectively.

[0130] S505, the corrected similarity mean of the candidate image is determined as the recall score of the candidate image.

[0131] For example, if the corrected similarity of candidate image a is 0.92, 0.89, and 0.98, then the recall score of the candidate image is (0.92+0.89+0.98) / 3=0.93.

[0132] In this embodiment of the disclosure, the similarity weighting value of the candidate image is predicted based on the preset dimensional attribute features corresponding to the candidate image, and the predicted weighting value is then used to correct the recall similarity of the candidate image, so as to improve the calculation accuracy of the recall score of the candidate image.

[0133] For example, such as Figure 6 As shown, for the image to be searched (or the query image of a face), the facial features of the image to be searched are extracted to obtain the target facial features, and the attribute features of the image to be searched are extracted to obtain the target attribute features. Based on the similarity between the target facial features and the facial features of each image in the facial feature database, the corresponding recall image is retrieved from the facial feature database to obtain a first recall image set. Figure 6 (Midvector recall).

[0134] Based on the similarity between the image to be searched and each recalled image in the primary recall image set, positive sample images are determined from the primary recall image set. Figure 6 (Supplementing positive samples), based on the facial features of the positive sample images and the similarity between them and the facial features of each image in the facial feature database, recall images corresponding to the positive sample images from the facial feature database, resulting in a secondary recall image set. Figure 6 (Second recall of samples from Zhongzheng).

[0135] The recalled images in the primary and secondary recall image sets are deduplicated to obtain a candidate image set. The recall count and recall score of each candidate image in the candidate image set are then determined. Figure 6 The system uses a comprehensive scoring method, taking the recall score of each candidate image as the primary parameter and the recall count of each candidate image as the secondary parameter. The candidate images are then ranked, and the top-ranked candidate images with recall scores greater than the sixth threshold are identified as the search results. Figure 6 (Result returned).

[0136] This disclosure also provides an image retrieval device based on artificial intelligence, see [link to relevant documentation]. Figure 7 The device includes:

[0137] The first feature extraction module 701 is used to extract facial features from the image to be searched, thereby obtaining the target facial features;

[0138] The first image recall module 702 is used to recall the image corresponding to the image to be searched from the face feature library based on the similarity between the target face features and the face features of each image in the face feature library, and obtain a recall image set; the face feature library contains the face features of multiple images;

[0139] The sample supplementation module 703 is used to determine positive sample images from the primary recall image set based on the similarity between the image to be searched and each recalled image in the primary recall image set;

[0140] The second image recall module 704 is used to recall the recall image corresponding to the positive sample image from the face feature library based on the similarity between the face features of the positive sample image and the face features of each image in the face feature library, and to obtain a secondary recall image set.

[0141] The image retrieval module 705 is used to determine the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set.

[0142] In this embodiment, positive sample images are determined from the primary recall image set of the image to be searched to augment the image sample. These augmented positive sample images are then used for further image recall, thereby improving the recall rate of the image to be searched. Based on the primary recall image set of the image to be searched and the secondary recall image set of the positive sample images, the retrieval result is determined from the two recall results, improving the recall accuracy of the image to be searched. In other words, both the recall rate and recall accuracy of the image to be searched are improved.

[0143] In one possible implementation, the above-described apparatus further includes:

[0144] The second feature extraction module is used to extract the attribute features of the image to be searched, and obtain the target attribute features;

[0145] The image quality determination module is used to determine the image quality value of the image to be searched based on the target attribute features;

[0146] The image information acquisition module is used to acquire the shooting information of the image to be searched;

[0147] The sample supplementation module is specifically used to determine a preset number of positive sample images from a primary recall image set based on image quality value, target attribute features, shooting information, and the similarity between the image to be searched and each recalled image in a primary recall image set.

[0148] In one possible implementation, the image quality determination module described above is specifically used for:

[0149] The target attribute features are encoded to obtain the attribute feature codes;

[0150] The attribute feature encoding is input into a pre-trained face quality discrimination model to predict face quality and obtain the image quality value of the image to be searched. The pre-trained face quality discrimination model is trained based on the sample attribute feature encoding and the image quality value corresponding to the sample attribute feature encoding.

[0151] In one possible implementation, the target attribute features include at least one of the following: whether the face in the image is wearing a mask, whether the face is wearing glasses, the age of the face, the gender of the face, the brightness of the face, the blur of the face, the pitch angle of the face, and the rotation angle of the face; the shooting information includes at least one of the following: the type of shooting device, the shooting time, and the license plate information in the image.

[0152] In one possible implementation, the sample supplementation module 703 includes:

[0153] The similarity acquisition submodule is used to obtain the similarity between the image to be searched and each recalled image in a single recall image set, thus obtaining the similarity of each image;

[0154] The first sample supplementation submodule is used to traverse each recalled image in the recalled image set once. If the shooting information of the image to be searched includes the first specified information, the recalled images with image similarity greater than the first threshold and shooting information including the second specified information are added to the positive sample image set.

[0155] The second sample supplementation submodule is used to traverse each recalled image in the recalled image set once when the number of positive sample images in the positive sample image set is less than a preset number, and add the recalled images whose image similarity is greater than the second threshold, whose image quality value is opposite to the image quality value of the image to be searched, and whose current image is not a positive sample image to the positive sample image set.

[0156] The third sample supplementation submodule is used to traverse each recalled image in the recalled image set once when the number of positive sample images in the positive sample image set is less than a preset number. If the target attribute feature of the image to be searched includes the first specified attribute, the recalled image with image similarity greater than the third threshold, attribute feature including the second specified attribute, and currently not a positive sample image is added to the positive sample image set.

[0157] The fourth sample supplementation submodule is used to traverse each recalled image in the recalled image set once when the number of positive sample images in the positive sample image set is less than a preset number. It adds recalled images whose image similarity is between the fourth threshold and the fifth threshold and which are not currently positive sample images to the positive sample image set until the number of positive sample images in the positive sample image set reaches the preset number. The fifth threshold changes with the determined similarity between the previous positive sample images and the image to be searched.

[0158] In one possible implementation, the image retrieval module 705 includes:

[0159] The image deduplication submodule is used to deduplicatize each recalled image in the primary recall image set and the secondary recall image set to obtain a candidate image set.

[0160] The recall count determination submodule is used to determine the recall count for each candidate image in the candidate image set;

[0161] The recall score determination submodule is used to determine the recall score value for each candidate image in the candidate image set;

[0162] The image ranking submodule is used to rank each candidate image based on its recall score and the number of times it has been recalled.

[0163] The result determination submodule is used to determine the number of candidate images with recall scores greater than the sixth threshold and ranked at the top as the retrieval results of the image to be searched.

[0164] In one possible implementation, the recall score determination submodule described above is specifically used for:

[0165] Obtain the recall similarity of each candidate image in the candidate image set. The recall similarity represents the similarity between the candidate image and the image that recalled the candidate image.

[0166] For each candidate image, determine the preset dimensional attribute features corresponding to that candidate image;

[0167] Based on the preset dimensional attribute features, the similarity weighting value of the candidate image is predicted to obtain the predicted weighting value.

[0168] By using the predicted weighting value, the recall similarity of the candidate image is corrected to obtain the corrected similarity.

[0169] The corrected mean similarity of the candidate image is determined as the recall score of the candidate image.

[0170] In one possible implementation, the above-mentioned prediction of the similarity weighting value of the candidate image based on preset dimensional attribute features to obtain the predicted weighting value includes:

[0171] The preset dimensional attribute features are input into the pre-trained weighted adjustment model to obtain the predicted weighted value corresponding to the recall similarity of the candidate image; wherein, the pre-trained weighted adjustment model is trained based on the preset dimensional attribute features of the sample and the weighted value corresponding to the preset dimensional attribute features of the sample.

[0172] In one possible implementation, the aforementioned preset dimension attribute features include: the spatial distance between the candidate image and the target image, whether the license plate information contained in the candidate image is the same as the license plate information contained in the target image, the shooting time difference between the candidate image and the target image, the image quality value of the candidate image and the image quality value of the target image, wherein the target image is the image of the recalled candidate image.

[0173] In one possible implementation, the above-mentioned image sorting submodule is specifically used for:

[0174] The recall score of each candidate image is used as the primary parameter, and the recall count of each candidate image is used as the secondary parameter to sort the candidate images.

[0175] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. It should be noted that the head model in this embodiment is not a head model specific to any particular user and does not reflect the personal information of any particular user. It should also be noted that the two-dimensional face images in this embodiment are from publicly available datasets.

[0176] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0177] This disclosure provides an electronic device, comprising:

[0178] At least one processor; and

[0179] A memory that is communicatively connected to at least one processor; wherein,

[0180] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform any of the methods of this disclosure.

[0181] This disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform any of the methods described in this disclosure.

[0182] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described in this disclosure.

[0183] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0184] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0185] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0186] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the AI-based image retrieval method described above. For example, in some embodiments, the AI-based image retrieval method described above can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the AI-based image retrieval method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the AI-based image retrieval method described above by any other suitable means (e.g., by means of firmware).

[0187] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0188] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0189] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0190] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0191] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0192] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0193] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0194] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An image retrieval method based on artificial intelligence, comprising: The facial features of the image to be searched are extracted to obtain the target facial features; Based on the similarity between the target facial features and the facial features of each image in the facial feature database, the corresponding recall image is retrieved from the facial feature database to obtain a recall image set; the facial feature database contains facial features of multiple images; The attribute features of the image to be searched are extracted to obtain the target attribute features; Based on the target attribute features, determine the image quality value of the image to be searched; Obtain the image capture information of the image to be searched; Based on the image quality value, the target attribute features, the shooting information, and the similarity between the image to be searched and each recalled image in the first recall image set, a preset number of positive sample images are determined from the first recall image set; Based on the facial features of the positive sample image and the similarity between the facial features of each image in the facial feature database, the recall image corresponding to the positive sample image is recalled from the facial feature database to obtain a secondary recall image set; The retrieval results of the image to be searched are determined from the primary recall image set and the secondary recall image set; The step of determining the image quality value of the image to be searched based on the target attribute features includes: The target attribute features are encoded to obtain attribute feature codes; The attribute feature encoding is input into a pre-trained face quality discrimination model to predict face quality, thereby obtaining the image quality value of the image to be searched; wherein, the pre-trained face quality discrimination model is trained based on the sample attribute feature encoding and the image quality value corresponding to the sample attribute feature encoding; The step of determining a preset number of positive sample images from the primary recall image set based on the image quality value, the target attribute features, the shooting information, and the similarity between the image to be searched and each recalled image in the primary recall image set includes: The similarity between the image to be searched and each recalled image in the first recall image set is obtained to obtain the image similarity. Traverse each recalled image in the first recalled image set. If the shooting information of the image to be searched includes the first specified information, add the recalled images with image similarity greater than the first threshold and shooting information including the second specified information to the positive sample image set. When the number of positive sample images in the positive sample image set is less than a preset number, each recalled image in the first recall image set is traversed, and the recalled images whose image similarity is greater than the second threshold, whose image quality value is opposite to the image quality value of the image to be searched, and whose current image is not a positive sample image are added to the positive sample image set. When the number of positive sample images in the positive sample image set is less than a preset number, each recalled image in the first recall image set is traversed. If the target attribute feature of the image to be searched includes a first specified attribute, the recalled image with an image similarity greater than a third threshold, an attribute feature including a second specified attribute, and not currently a positive sample image is added to the positive sample image set. When the number of positive sample images in the positive sample image set is less than a preset number, each recalled image in the first recall image set is traversed, and recalled images whose image similarity is between the fourth threshold and the fifth threshold and which are not currently positive sample images are added to the positive sample image set until the number of positive sample images in the positive sample image set reaches the preset number; wherein, the fifth threshold changes with the similarity between the determined previous positive sample image and the image to be searched.

2. The method according to claim 1, wherein, The target attribute features include at least one of the following: whether the face in the image is wearing a mask, whether the face is wearing glasses, the age of the face, the gender of the face, the brightness of the face, the blur of the face, the pitch angle of the face, and the rotation angle of the face; the shooting information includes at least one of the following: the type of shooting device, the shooting time, and the license plate information in the image.

3. The method according to claim 1, wherein, Determining the retrieval results of the image to be searched from the primary recall image set and the secondary recall image set includes: Deduplication is performed on each recalled image in the primary recall image set and the secondary recall image set to obtain a candidate image set; Determine the recall count for each candidate image in the candidate image set; Determine the recall score for each candidate image in the candidate image set; Based on the recall score and the number of times each candidate image was recalled, the candidate images were ranked. The number of candidate images whose recall scores are greater than the sixth threshold and whose ranking is among the top are determined as the retrieval results for the image to be searched.

4. The method according to claim 3, wherein, Determining the recall score for each candidate image in the candidate image set includes: Obtain the recall similarity of each candidate image in the candidate image set, where the recall similarity represents the similarity between the candidate image and the image that recalled the candidate image; For each candidate image, determine the preset dimensional attribute features corresponding to that candidate image; Based on the preset dimensional attribute features, the similarity weighting value of the candidate image is predicted to obtain the predicted weighting value. Using the predicted adjustment value, the recall similarity of the candidate image is corrected to obtain the corrected similarity. The corrected mean similarity of the candidate image is determined as the recall score of the candidate image.

5. The method according to claim 4, wherein, The step of predicting the similarity weighting value of the candidate image based on the preset dimensional attribute features to obtain the predicted weighting value includes: The preset dimension attribute features are input into a pre-trained weighting model to obtain the predicted weighting value corresponding to the recall similarity of the candidate image; wherein, the pre-trained weighting model is trained based on the preset dimension attribute features of the sample and the weighting value corresponding to the preset dimension attribute features of the sample.

6. The method according to claim 4, wherein, The preset dimension attribute features include: the spatial distance between the candidate image and the target image, whether the license plate information contained in the candidate image is the same as that contained in the target image, the shooting time difference between the candidate image and the target image, the image quality value of the candidate image and the image quality value of the target image, wherein the target image is the image of the recalled candidate image.

7. The method according to any one of claims 3-6, wherein, The process of ranking candidate images based on their recall score and the number of times each candidate image has been recalled includes: The recall score of each candidate image is used as the primary parameter, and the recall count of each candidate image is used as the secondary parameter to sort the candidate images.

8. An image retrieval device based on artificial intelligence, comprising: The first feature extraction module is used to extract facial features from the image to be searched, thereby obtaining the target facial features; The first image recall module is used to recall the image corresponding to the image to be searched from the face feature library based on the similarity between the target face feature and the face features of each image in the face feature library, thereby obtaining a recall image set; the face feature library contains the face features of multiple images; The second feature extraction module is used to extract the attribute features of the image to be searched to obtain the target attribute features; An image quality determination module is used to determine the image quality value of the image to be searched based on the target attribute features; An image information acquisition module is used to acquire the shooting information of the image to be searched; The sample supplementation module is used to determine a preset number of positive sample images from the primary recall image set based on the image quality value, the target attribute features, the shooting information, and the similarity between the image to be searched and each recalled image in the primary recall image set. The second image recall module is used to recall the recall image corresponding to the positive sample image from the face feature library based on the similarity between the facial features of the positive sample image and the facial features of each image in the face feature library, so as to obtain a secondary recall image set. An image retrieval module is used to determine the retrieval result of the image to be searched from the primary recall image set and the secondary recall image set; Specifically, the image quality determination module is used to encode the target attribute features to obtain attribute feature codes; input the attribute feature codes into a pre-trained face quality discrimination model to perform face quality prediction, and obtain the image quality value of the image to be searched; wherein, the pre-trained face quality discrimination model is trained based on the sample attribute feature codes and the image quality values ​​corresponding to the sample attribute feature codes; The sample replenishment module includes: The similarity acquisition submodule is used to acquire the similarity between the image to be searched and each recalled image in the first recall image set, and obtain the similarity of each image; The first sample supplementation submodule is used to traverse each recalled image in the first recalled image set, and when the shooting information of the image to be searched includes the first specified information, add the recalled images with image similarity greater than the first threshold and shooting information including the second specified information to the positive sample image set. The second sample supplementation submodule is used to, when the number of positive sample images in the positive sample image set is less than a preset number, traverse each recalled image in the first recalled image set and add recalled images whose image similarity is greater than a second threshold, whose image quality value is opposite to the image quality value of the image to be searched, and whose current image is not a positive sample image to the positive sample image set. The third sample supplementation submodule is used to traverse each recalled image in the first recall image set when the number of positive sample images in the positive sample image set is less than a preset number. If the target attribute feature of the image to be searched includes a first specified attribute, the recalled image with image similarity greater than a third threshold, attribute feature including a second specified attribute, and currently not a positive sample image is added to the positive sample image set. The fourth sample supplementation submodule is used to, when the number of positive sample images in the positive sample image set is less than a preset number, traverse each recalled image in the first recalled image set, and add recalled images whose image similarity is between a fourth threshold and a fifth threshold and which are not currently positive sample images to the positive sample image set, until the number of positive sample images in the positive sample image set reaches the preset number; wherein, the fifth threshold changes with the determined similarity between the previous positive sample image and the image to be searched.

9. The apparatus according to claim 8, wherein, The target attribute features include at least one of the following: whether the face in the image is wearing a mask, whether the face is wearing glasses, the age of the face, the gender of the face, the brightness of the face, the blur of the face, the pitch angle of the face, and the rotation angle of the face; the shooting information includes at least one of the following: the type of shooting device, the shooting time, and the license plate information in the image.

10. The apparatus according to claim 8, wherein the image retrieval module comprises: The image deduplication submodule is used to deduplicatize each recalled image in the primary recall image set and the secondary recall image set to obtain a candidate image set. The recall count determination submodule is used to determine the recall count of each candidate image in the candidate image set; The recall score determination submodule is used to determine the recall score value of each candidate image in the candidate image set; The image ranking submodule is used to rank each candidate image based on its recall score and the number of times it has been recalled. The result determination submodule is used to determine the number of candidate images with recall scores greater than the sixth threshold and ranked at the top as the retrieval results of the image to be searched.

11. The apparatus according to claim 10, wherein, The recall score determination submodule is specifically used for: Obtain the recall similarity of each candidate image in the candidate image set, where the recall similarity represents the similarity between the candidate image and the image that recalled the candidate image; For each candidate image, determine the preset dimensional attribute features corresponding to that candidate image; Based on the preset dimensional attribute features, the similarity weighting value of the candidate image is predicted to obtain the predicted weighting value. Using the predicted adjustment value, the recall similarity of the candidate image is corrected to obtain the corrected similarity. The corrected mean similarity of the candidate image is determined as the recall score of the candidate image.

12. The apparatus according to claim 11, wherein, The step of predicting the similarity weighting value of the candidate image based on the preset dimensional attribute features to obtain the predicted weighting value includes: The preset dimension attribute features are input into a pre-trained weighting model to obtain the predicted weighting value corresponding to the recall similarity of the candidate image; wherein, the pre-trained weighting model is trained based on the preset dimension attribute features of the sample and the weighting value corresponding to the preset dimension attribute features of the sample.

13. The apparatus according to claim 11, wherein, The preset dimension attribute features include: the spatial distance between the candidate image and the target image, whether the license plate information contained in the candidate image is the same as that contained in the target image, the shooting time difference between the candidate image and the target image, the image quality value of the candidate image and the image quality value of the target image, wherein the target image is the image of the recalled candidate image.

14. The apparatus according to any one of claims 10-13, wherein, The image sorting submodule is specifically used for: The recall score of each candidate image is used as the primary parameter, and the recall count of each candidate image is used as the secondary parameter to sort the candidate images.

15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image retrieval method and device, electronic equipment and storage medium

    CN113806582A