Multi-modal image retrieval method and device, computer equipment and storage medium
By searching and archive search in the image library and center of mass library, combined with topological recommendation technology, the problem of single modality and style of image search results is solved, and higher quality and diverse image search results are achieved.
Patent Information
- Application Number
- CN202510281373.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the modality and style of image search results are single, making it difficult to recall high-quality images of different modality and styles.
By conducting image search and archive search in the preset image library and center of mass library, combined with topological recommendation technology, the modality and style of the image to be retrieved can be expanded, and the recommended image is obtained, and a secondary image search is performed to obtain the image search results in a comprehensive way.
This increases the modality and style diversity of image search results, improves the recall rate of image search, and solves the problem of single modality and style of image search results.
Smart Images

Figure CN120216714A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a multi-modal image retrieval method, apparatus, computer device, and storage medium. Background Art
[0002] Person identity recognition has been widely applied in the fields of security, content analysis, auditing, interaction, etc. Currently, there are mainly methods such as identity recognition, image search by image, and file search by image.
[0003] Among them, image search by image is to extract features from face or body capture images to obtain corresponding modal information, and then compare and retrieve them one by one with the people in the existing database based on the extracted modal information. However, currently, most of the image search results are in single-modal image search mode and highly dependent on the style of the image to be retrieved (e.g., angle, color, quality, etc.). Usually, only images with similar styles and dimensions can be retrieved. Especially if the quality and angle of the image to be retrieved are not good, the style of the recall results is also similar, and it is difficult to recall high-quality results.
[0004] Regarding the problem of single modality and style in image search results in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] Based on this, it is necessary to provide a multi-modal image retrieval method, apparatus, computer device, and storage medium that can increase the modality and style of image search results for the above technical problems.
[0006] In the first aspect, in the present embodiment, a multi-modal image retrieval method is provided, including:
[0007] Performing image search in a preset image library corresponding to the modality of the image to be retrieved to obtain a first image search result; the image to be retrieved includes a face and / or a body;
[0008] Performing file search in a preset centroid library corresponding to the modality of the image to be retrieved, and determining the main file corresponding to the image to be retrieved from a preset file library according to the file search result;
[0009] Performing topological recommendation according to the main file and the first image search result to obtain a recommended image; and performing image search based on the recommended image to obtain a second image search result;
[0010] Obtaining an image retrieval result according to the predicted positive example images in the first image search result, the main file, and the second image search result.
[0011] In some of these embodiments, the method further includes:
[0012] Pre-construct a preset image library for corresponding modalities of human faces and human bodies;
[0013] Extract the face features and face attributes, as well as the human body features and human body attributes, from the preset image library.
[0014] In some embodiments, the performing an image search in the preset image library of the corresponding modality based on the image to be retrieved to obtain a first image search result includes:
[0015] Extract the face features and / or human body features from the image to be retrieved;
[0016] Perform an image search based on the similarity between the face features and / or human body features in the image to be retrieved and the image features in the preset image library to obtain a first image search result.
[0017] In some embodiments, the method further includes:
[0018] Pre-construct a preset archive library for each retrieval target;
[0019] Based on the face features and human body features in the preset archive library, construct a preset centroid library for corresponding modalities of human faces and human bodies.
[0020] In some embodiments, the performing an archive search in the preset centroid library of the corresponding modality based on the image to be retrieved and determining the main file corresponding to the image to be retrieved from the preset archive library according to the archive search result includes:
[0021] Extract the face features and / or human body features from the image to be retrieved;
[0022] Perform an archive search based on the similarity between the face features and / or human body features in the image to be retrieved and the centroid features in the preset centroid library to obtain an archive search result;
[0023] Use the preset archive library that meets the similarity requirement in the archive search result as the main file.
[0024] In some embodiments, the performing a topology recommendation based on the main file and the first image search result to obtain a recommended image includes:
[0025] Set the value intervals of each face attribute and each human body attribute respectively;
[0026] Based on the value interval of the face attribute, select a face capture recommendation image from the face result sets in the main file and the first image search result;
[0027] Based on the value interval of the human body attribute, select a human body capture recommendation image from the human body result sets in the main file and the first image search result;
[0028] Based on the face capture recommendation map and the body capture recommendation map, a recommended image is obtained.
[0029] In some of the embodiments, the method further includes:
[0030] Determine the predicted features in the second image search result;
[0031] Based on a preset classifier, perform classification prediction on the predicted features to obtain the predicted scores of each image in the second image search result;
[0032] According to the predicted scores and a preset threshold, determine the positive example images in the second image search result.
[0033] In a second aspect, a multi-modal image retrieval device is provided in this embodiment, including:
[0034] A first image search module, configured to perform image search in a preset image library corresponding to a modality based on an image to be retrieved, to obtain a first image search result; the image to be retrieved includes a face and / or a body;
[0035] An archive search module, configured to perform archive search in a preset centroid library corresponding to the modality based on the image to be retrieved, and determine a master file corresponding to the image to be retrieved from a preset archive library according to the archive search result;
[0036] A second image search module, configured to perform topology recommendation according to the master file and the first image search result to obtain a recommended image; and perform image search based on the recommended image to obtain a second image search result;
[0037] An image retrieval module, configured to obtain an image retrieval result according to the first image search result, the master file, and the predicted positive example images in the second image search result.
[0038] In a third aspect, a computer device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the multi-modal image retrieval method described in the first aspect above is implemented.
[0039] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, the multi-modal image retrieval method described in the first aspect above is implemented.
[0040] Compared with related technologies, the multimodal image retrieval method, device, computer device, and storage medium provided in this embodiment perform image search in a preset image library corresponding to the modality of the image to be retrieved to obtain a first image search result; the image to be retrieved includes a face and / or a human body; perform a file search in a preset centroid library corresponding to the modality of the image to be retrieved, and determine the main file corresponding to the image to be retrieved from a preset file library according to the file search result; perform topological recommendation based on the main file and the first image search result to obtain a recommended image; and perform image search based on the recommended image to obtain a second image search result; obtain an image retrieval result according to the predicted positive example images in the first image search result, the main file, and the second image search result. Through this embodiment, topological recommendation can be performed through file search and image search, the modality and style of the image to be retrieved are increased to obtain a recommended image, and then the recommended image is used for secondary image search to comprehensively obtain an image retrieval result, solving the problem of the single modality and style of the image search result.
[0041] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0043] Figure 1 is a hardware structure block diagram of a terminal for the multimodal image retrieval method in one embodiment;
[0044] Figure 2 is a flowchart of the multimodal image retrieval method in one embodiment;
[0045] Figure 3 is a schematic flowchart of multimodal image retrieval in one embodiment;
[0046] Figure 4 is a schematic flowchart of topological recommendation in one embodiment;
[0047] Figure 5 is a flowchart of the multimodal image retrieval method in another embodiment;
[0048] Figure 6 is a structure block diagram of the multimodal image retrieval device in one embodiment.
[0049] In the figure: 102, a processor; 104, a memory; 106, a transmission device; 108, an input / output device; 10, a first image search module; 20, an archive search module; 30, a second image search module; 40, an image retrieval module. Detailed implementation manners
[0050] To understand the purpose, technical solution and advantages of the present application more clearly, the present application will be described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0051] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application only distinguish similar objects and do not represent a specific order for the objects.
[0052] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 is the hardware structure block diagram of the terminal of the multi-modal image retrieval method of this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in Figure 1The structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 or different configurations from those shown in Figure 1 .
[0053] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the multi-modal image retrieval method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0054] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0055] In this embodiment, a multi-modal image retrieval method is provided. Figure 2 is a flowchart of the multi-modal image retrieval method in this embodiment, as shown in Figure 2 . The method includes the following steps:
[0056] Step S201, perform an image search in a preset image library corresponding to the modality of the image to be retrieved to obtain a first image search result; the image to be retrieved includes a face and / or a human body.
[0057] Specifically, the acquisition method of the image to be retrieved includes but is not limited to cameras set in public areas, etc. Specifically, it can be portrait capture, which includes a face and / or a human body. The corresponding modality of the feature to be retrieved is extracted using the face and / or the human body, and the face feature to be retrieved and / or the human body feature to be retrieved are respectively used for image retrieval in the preset image library corresponding to the modality, so as to retrieve and recall a face result set and / or a human body result set that meet the similarity threshold in the preset image library as the first image search result.
[0058] Among them, a bottom database pre-loaded with portrait captures (including face information and body information) is used as a preset image library for the corresponding modalities of faces and bodies.
[0059] Step S202: Conduct an archive search for the image to be retrieved in the preset centroid library of the corresponding modality, and determine the main file corresponding to the image to be retrieved from the preset archive library according to the archive search result.
[0060] Specifically, perform a clustering operation on the above-mentioned face information and body information, cluster portrait captures of the same retrieval target with different styles into the same category to obtain archive information, where the archive information includes the preset archive library for each retrieval target, and the multi-modal preset centroid libraries further refined based on each preset archive library, such as the body preset centroid library and the face preset centroid library, etc.
[0061] Extract the features to be retrieved of the corresponding modality using the face and / or body, conduct an archive search for the face features to be retrieved and / or the body features to be retrieved in the preset centroid library of the corresponding modality respectively, and obtain the archive search results for the face and the body respectively by calculating the feature similarity. The archive search results include the similarity results of the features in the preset centroid library and the preset archive library corresponding to the preset centroid library. Extract the preset archive library that meets the similarity requirement from the archive search results as the main file. When the image to be retrieved includes a face and a body, each corresponds to a preset archive library that meets the similarity requirement, and then further determine the unique main file.
[0062] Step S203: Perform topological recommendation based on the main file and the first image search result to obtain recommended images; and conduct an image search based on the recommended images to obtain the second image search result.
[0063] Specifically, use the face attributes (such as clarity, quality score, alignment score, whether wearing a mask, yaw angle, pitch angle, roll angle, age) and body attributes (such as whether associated with a face, image width, feature type, GPS) extracted by feature extraction to select recommended captures for topological recommendation through multi-attribute linkage from the main file and the first image search result. For the face result set and / or the body result set in the main file and the first image search result, perform topological recommendation using face attributes and body attributes respectively. Specifically, it is possible to comprehensively consider the face captures in the main file and the face result set, select captures using face attributes, comprehensively consider the body captures in the main file and the body result set, select captures using body attributes, and finally output face captures and body captures with multi-modal and multi-style characteristics as recommended images.
[0064] Perform image search on the recommended images based on topological recommendation, and perform image retrieval on the face features and / or the human body features to be retrieved in the corresponding modality's preset image library respectively, so as to retrieve and recall a face result set and / or a human body result set that meet the similarity threshold in the preset image library as the second image search result.
[0065] Step S204: Obtain an image retrieval result according to the predicted positive example images in the first image search result, the master file, and the second image search result.
[0066] Specifically, introduce a topological decision algorithm to perform classification prediction on the predicted features of the images in the second image search result. Among them, for the face result set and the human body result set in the second image search result, extract the corresponding predicted features respectively, and use a trained decision tree, GBDT (Gradient Boosting Decision Tree), etc. classifier to perform classification prediction on the predicted features, so as to screen out the positive example images that meet the similarity threshold in the second image search result and improve the accuracy of image search.
[0067] Integrate the master file, the images in the first image search result, and the predicted positive example images in the second image search result as the image retrieval result of the image to be retrieved.
[0068] Figure 3 is a schematic flowchart of multimodal image retrieval in this embodiment. As Figure 3 shown, in the preparation stage, a base library of portrait captures (including face information and human body information) is pre-loaded as the preset image library for the corresponding modalities of faces and human bodies, and clustering operations are performed on the above-mentioned face information and human body information to cluster portrait captures of the same retrieval target with different styles into the same class, obtaining a preset archive library for each retrieval target, and a multimodal preset centroid library further refined based on each preset archive library, such as a human body preset centroid library and a face preset centroid library, etc.
[0069] Input the image to be retrieved, perform image search and archive search in the preset image library and the preset centroid library respectively, perform topological recommendation according to the determined master file and the first image search result to obtain recommended images. Use the recommended images for secondary image search, and finally perform topological decision on the second image search result. According to the first image search result, the master file, and the predicted positive example images in the second image search result, output the image retrieval result of multimodal image retrieval.
[0070] It should be noted that the user information (including but not limited to user device information, user personal information, portrait captures, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by relevant departments of all parties.
[0071] Through the above steps, it is possible to perform topological recommendation by using the results of one-time image search and archive search, expand the modality and style diversity of the images to be retrieved, obtain recommended images, perform secondary image search using the recommended images, and comprehensively obtain the image retrieval results of the images to be retrieved through two image searches and archive search. Compared with the single-modal image search mode in related technologies, it is possible to use the clustered archive information as a supplement to the preset image library, and increase the style diversity of the retrieved images through topological recommendation, and be able to retrieve and recall images of different modalities and styles, solving the problem of single modality and style of image search results.
[0072] In some of these embodiments, the above method further includes the following steps:
[0073] Pre-construct a preset image library for the corresponding modalities of faces and human bodies; extract the face features and face attributes, as well as the human body features and human body attributes in the preset image library.
[0074] Specifically, pre-load the base library of portrait captures (including face information and human body information) as the preset image library for the corresponding modalities of faces and human bodies. In addition, extract face features and attributes, human body features and attributes through a feature extraction algorithm, and at the same time store the spatio-temporal information where the portrait capture image is located, and load the face information (including face features, face attributes and spatio-temporal information) and human body information (including human body features, human body attributes and spatio-temporal information) into the server memory and save them for a long time in turn. Among them, face attributes include but are not limited to clarity, quality score, alignment score, whether wearing a mask, yaw angle, pitch angle, roll angle, age, etc., and human body attributes include but are not limited to whether associated with a face, image width, feature type, GPS, etc.
[0075] By constructing a multi-modal preset image library in this embodiment, it is used for image search and further for subsequent archiving operations to construct the archive information of the retrieval target.
[0076] In some of these embodiments, the above method further includes the following steps:
[0077] Pre-construct a preset archive library for each retrieval target; based on the face features and human body features in the preset archive library, construct a preset centroid library for the corresponding modalities of faces and human bodies.
[0078] Specifically, perform a clustering operation on the above face information and human body information, cluster portrait captures of the same retrieval target with different styles into the same class to obtain archive information, and the archive information includes the preset archive library for each retrieval target, and the multi-modal preset centroid library further refined based on each preset archive library, such as the human body preset centroid library and the face preset centroid library, etc.
[0079] Each preset archive contains multiple different captures of the same retrieval target, with diverse capture modalities (including faces and human bodies) and capture styles. The corresponding preset centroid library for each preset archive includes multiple centroid features corresponding to each preset archive. The centroid feature is a virtual representation of the face / human body features, which can better summarize the key features of the face / human body. Exemplarily, the centroid feature can be obtained by averaging the face features / human body features of multiple images in the preset archive. The centroid feature can not only be generated based on actual capture images, but may also be the feature corresponding to a virtual generated image.
[0080] Through the portrait filing in this embodiment, different-style captures of the same retrieval target are grouped into archive information, providing a basis for subsequent multi-style topology recommendations. As a supplement to the multi-modal and multi-style retrieval graph, it can achieve multi-modal and multi-style recall during image retrieval.
[0081] In some of these embodiments, the above step S201 of performing an image search in the preset image library corresponding to the modality based on the image to be retrieved to obtain the first image search result includes the following steps:
[0082] Extract the face features and / or human body features in the image to be retrieved; perform an image search based on the similarity between the face features and / or human body features in the image to be retrieved and the image features in the preset image library to obtain the first image search result.
[0083] Specifically, use a feature extraction algorithm to extract the face features and / or human body features corresponding to the modality of the face and / or human body in the image to be retrieved, and retrieve the face features and human body features separately in the preset image library. Specifically, calculate the similarity between the face features / human body features to be retrieved and the image features in the preset image library, and obtain the face result set and / or human body result set whose similarity meets the similarity threshold as the first image search result. Exemplarily, the similarity threshold is 92 points.
[0084] By performing an image search in this embodiment and obtaining the first image search result from the preset image library based on the image to be retrieved, it can provide a basis for subsequent topology recommendations.
[0085] In some of these embodiments, the above step S202 of performing an archive search in the preset centroid library corresponding to the modality based on the image to be retrieved and determining the main file corresponding to the image to be retrieved from the preset archive includes the following steps:
[0086] Extract the face features and / or human body features in the image to be retrieved; perform an archive search based on the similarity between the face features and / or human body features in the image to be retrieved and the centroid features in the preset centroid library to obtain the archive search result; use the preset archive that meets the similarity requirement in the archive search result as the main file.
[0087] Specifically, the face and / or the human body are used to extract the to-be-retrieved features of the corresponding modality, and the to-be-retrieved face features and / or the to-be-retrieved human body features are respectively searched for files in the preset centroid libraries of the corresponding modalities. Specifically, the feature similarity between the face features and / or the human body features in the to-be-retrieved image and the centroid features in the preset centroid libraries of the corresponding modalities can be calculated to respectively obtain the file search results of the face and the human body. The file search results include the similarity results of the centroid features in the preset centroid libraries and the preset file libraries corresponding to the preset centroid libraries. When the to-be-retrieved image only contains a face or a human body, the preset file library corresponding to the similarity result with the highest score that meets the similarity threshold is used as the main file. Exemplarily, the similarity threshold is 92 points. When the to-be-retrieved image contains a face and a human body, there is respectively a preset file library that meets the similarity threshold and has the highest similarity result, and the preset file library with more face captures is selected as the only main file.
[0088] By performing file search based on the to-be-retrieved image and determining the main file in this embodiment, a file library containing faces and human bodies of different styles corresponding to the to-be-retrieved image can be obtained, providing a basis for subsequent topology recommendation.
[0089] In some of these embodiments, the above step S203 of performing topology recommendation based on the main file and the first image search result to obtain a recommended image includes the following steps:
[0090] The value intervals of each face attribute and each human body attribute are respectively set; based on the value intervals of the face attributes, face capture recommended images are selected from the face result sets in the main file and the first image search result; based on the value intervals of the human body attributes, human body capture recommended images are selected from the human body result sets in the main file and the first image search result; and the recommended image is obtained according to the face capture recommended images and the human body capture recommended images.
[0091] Specifically, the face attributes include but are not limited to clarity, quality score, alignment score, whether wearing a mask, yaw angle, pitch angle, roll angle, age, and the human body attributes include but are not limited to whether associated with a face, image width, feature type (including pedestrian and rider), GPS. The numerical values of the face attributes of the main file and the face result sets, and the numerical values of the human body attributes of the main file and the human body result sets are determined through feature extraction. The value intervals of each face attribute and each human body attribute are respectively set, and according to the numerical values of the face attributes of the face result sets and the main file, the face capture recommended images are selected by comprehensively considering the value intervals of each face attribute. According to the numerical values of the human body attributes of the human body result sets and the main file, the human body capture recommended images are selected by comprehensively considering the value intervals of each human body attribute. Thus, the face capture recommended images and the human body capture recommended images are used as the recommended image.
[0092] The face images in the comprehensive main file and face result set are selected according to the face attributes at a certain interval, and the recommended snapshots of all face attributes except whether a mask is worn are taken as a union. Taking clarity as an example, its value range is 1 to 100, and the value interval is set to 20, thereby obtaining several clarity ranges, and the face images fall into different clarity ranges. In each clarity range, a certain number of recommended face snapshots are selected as needed. If all face images have masks (value 1) and non-masks (value 0), at least 1 recommended face snapshot is taken for each, and finally the recommended face snapshots of all face attributes are combined.
[0093] Based on the human images in the main file and the human result set, the recommended human snapshots are selected through screening conditions according to human attributes and value intervals. The specific screening conditions are as follows:
[0094] If the number of human images associated with faces is greater than or equal to 100, pedestrian snapshots and rider snapshots are obtained from these human images associated with faces in a certain ratio according to the GPS, image width and feature type, with a value interval. For example, the ratio can be set to 1:1;
[0095] If the number of human body images with associated faces is less than 100, all human body images with associated faces will be used as recommended human body snapshots. Then, according to the GPS, image width and feature type, at the value interval, pedestrian snapshots and rider snapshots will be obtained from pure human body images without associated faces as recommended human body snapshots in a certain proportion.
[0096] Figure 4 is a schematic diagram of the topology recommendation process in this embodiment, such as Figure 4 As shown, taking the case where the image to be retrieved contains a human body and a face as an example, the image to be retrieved is input for image search and archive search respectively, wherein the face and the body in the image to be retrieved are searched in the face library and the body library of the preset image library respectively, and the first image search result including the face result set and the body result set is obtained, and the face and the body in the image to be retrieved are searched in the face preset centroid library and the body preset centroid library of the corresponding modality respectively, and the main file is determined. Based on the main file, the face result set and the body result set, a multi-modal topological recommendation with multi-attribute linkage is performed, and a face snapshot recommendation image and a body snapshot recommendation image are obtained as recommended images respectively.
[0097] By performing topological recommendation based on multi-attribute linkage of the main file of the archive search and the first image search result of the image search in this embodiment, the original image to be retrieved can be expanded in style and modality to obtain recommended images, thereby improving style diversity and further potentially improving the recall rate of image search using recommended images.
[0098] In some of these embodiments, the above method further includes the following steps:
[0099] Determine the predicted features in the second image search results; based on a preset classifier, perform classification prediction on the predicted features to obtain the predicted scores of each image in the second image search results; according to the predicted scores and a preset threshold, determine the positive example images in the second image search results.
[0100] Specifically, the face result set / human body result set in the second image search results includes the top K recommended captures of each recommended image and the corresponding similarity scores. Exemplarily, assuming the number of recommended images is 100 and the number of top K recommended captures is 300, then there are 300 captures in the face result set / human body result set.
[0101] For the face result set in the second image search results, its face prediction features include but are not limited to age, mask, quality score, clarity, gender, average index ReID score (the average of the similarity scores between the current capture and the recommended image). Use a trained classifier such as a decision tree, GBDT, etc. to perform classification prediction on the predicted features to obtain the predicted scores, and the captures corresponding to the predicted features that meet the preset threshold are positive example images.
[0102] For the human body result set in the second image search results, its human body prediction features include but are not limited to spatio-temporal information, average index ReID score (the average of the similarity scores between the current capture and the recommended image). Considering that the human body may change clothes across days and the credibility is relatively low, calculate the average index ReID score per day as the predicted score, and the captures that meet the preset threshold are positive example images.
[0103] By using the classifier in this embodiment to determine the positive and negative examples of the second image search results in the secondary image search, the accuracy of image search is improved.
[0104] The following describes and illustrates this embodiment through preferred embodiments.
[0105] Figure 5 is the flowchart of the multi-modal image retrieval method of this embodiment, as Figure 5 shown, the method includes the following steps:
[0106] Step S501, pre-construct a preset image library for the corresponding modalities of faces and human bodies; extract the face features and face attributes, as well as the human body features and human body attributes in the preset image library.
[0107] Step S502, perform clustering operations on the face information and human body information in the preset image library to construct a preset archive library for each retrieval target; based on the face features and human body features in the preset archive library, construct a preset centroid library for the corresponding modalities of faces and human bodies.
[0108] Step S503: Perform an image search based on the similarity between the face features and / or human body features in the image to be retrieved and the image features in the preset image library, and obtain the first image search result.
[0109] Step S504: Perform an archive search based on the similarity between the face features and / or human body features in the image to be retrieved and the centroid features in the preset centroid library, and obtain the archive search result; use the preset archive library that meets the similarity requirement in the archive search result as the main file.
[0110] Step S505: Perform a topology recommendation based on the main file and the first image search result to obtain the recommended image.
[0111] Step S506: Perform an image search based on the recommended image to obtain the second image search result; determine the positive example images in the second image search result through a preset classifier.
[0112] Step S507: Obtain the image retrieval result based on the first image search result, the main file, and the predicted positive example images in the second image search result.
[0113] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0114] Through this embodiment, using the archive information of the preset archive library and the preset centroid library as a supplement to the multi-modal and multi-style retrieval graph, the recall rate of different modalities and different styles during image search is improved. And a topology recommendation method combining multiple attributes is provided to obtain the recommended image, which improves the style diversity of the retrieval graph and potentially improves the image search recall rate. Finally, the face result set and the human body result set in the secondary image search contain many positive and negative example images, and a preset classifier is used to determine these positive and negative example images to improve the accuracy of image search.
[0115] In this embodiment, a multi-modal image retrieval device is also provided. This device is used to implement the above embodiment and the preferred implementation manners, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0116] Figure 6 is the structural block diagram of the multi-modal image retrieval device in this embodiment. As Figure 6 shown, this device includes:
[0117] The first image search module 10 is configured to perform image search in a preset image library of the corresponding modality based on the image to be retrieved, so as to obtain a first image search result; the image to be retrieved includes a face and / or a human body;
[0118] The file search module 20 is configured to perform file search in a preset centroid library of the corresponding modality based on the image to be retrieved, and determine a master file corresponding to the image to be retrieved from a preset file library according to the file search result;
[0119] The second image search module 30 is configured to perform topology recommendation according to the master file and the first image search result to obtain a recommended image; and perform image search based on the recommended image to obtain a second image search result;
[0120] The image retrieval module 40 is configured to obtain an image retrieval result according to the first image search result, the master file, and the predicted positive example images in the second image search result.
[0121] Through the device provided in this embodiment, it is possible to perform topology recommendation by using the results of one image search and file search, expand the modality and style diversity of the image to be retrieved, obtain a recommended image, perform secondary image search by using the recommended image, and comprehensively obtain the image retrieval result of the image to be retrieved through two image searches and file searches. Compared with the single-modal image search mode in the related art, it is possible to use the clustered file information as a supplement to the preset image library, and increase the style diversity of the retrieved images through topology recommendation, and it is possible to retrieve and recall images of different modalities and styles, solving the problem of single modality and style of the image search result.
[0122] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.
[0123] In this embodiment, a computer device is further provided, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0124] Optionally, the above computer device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0125] It should be noted that the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated in this embodiment.
[0126] In addition, in combination with the multi-modal image retrieval method provided in the above embodiments, a storage medium may also be provided in this embodiment to implement the method. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the multi-modal image retrieval methods in the above embodiments is implemented.
[0127] It should be understood that the specific embodiments described herein are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0128] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during the development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.
[0129] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0130] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A multimodal image retrieval method, characterized in that: include: Performing an image search in a preset image library of a corresponding modality based on the image to be retrieved, and obtaining a first image search result; The image to be retrieved includes a face and / or a body; Performing an archive search based on the image to be retrieved in a preset centroid library of the corresponding modality, and determining a master file corresponding to the image to be retrieved from the preset archive library according to the archive search results; Performing topological recommendation based on the main file and the first image search results to obtain a recommended image; and performing an image search based on the recommended image to obtain a second image search result; An image retrieval result is obtained based on the first image search result, the main file, and the predicted positive example images in the second image search result.
2. The multimodal image retrieval method according to claim 1, characterized in that: The method further comprises: Pre-build a preset image library of face and body corresponding modalities; The facial features and facial attributes, as well as the body features and body attributes in the preset image library are extracted.
3. The multimodal image retrieval method according to claim 2, characterized in that: The step of performing an image search in a preset image library of a corresponding modality based on the image to be retrieved to obtain a first image search result includes: Extracting facial features and / or body features in the image to be retrieved; An image search is performed based on the similarity between the facial features and / or body features in the image to be retrieved and the image features in the preset image library to obtain a first image search result.
4. The multimodal image retrieval method according to claim 1, characterized in that: The method further comprises: Pre-build preset archives for each search target; Based on the facial features and body features in the preset archive, a preset centroid library of corresponding modes of face and body is constructed.
5. The multimodal image retrieval method according to any one of claim 1 or claim 4, characterized in that: The method of performing an archive search in a preset centroid library of a corresponding modality based on the image to be retrieved, and determining a master file corresponding to the image to be retrieved from the preset archive library according to the archive search results, comprises: Extracting facial features and / or body features in the image to be retrieved; Performing archive search based on the similarity between the facial features and / or body features in the image to be retrieved and the centroid features in the preset centroid library to obtain archive search results; The preset archive library that meets the similarity requirement in the archive search results is used as the main archive.
6. The multimodal image retrieval method according to claim 1, characterized in that: The performing topological recommendation according to the main file and the first image search result to obtain a recommended image includes: Set the value intervals of each face attribute and each body attribute respectively; Based on the value interval of the facial attribute, selecting a facial snapshot recommendation image from the facial result set in the main file and the first image search result; Based on the value interval of the human attribute, selecting a human body snapshot recommendation image from the human body result set in the main file and the first image search result; A recommended image is obtained according to the recommended face snapshot image and the recommended body snapshot image.
7. The multimodal image retrieval method according to claim 1, characterized in that: The method further comprises: determining a predicted feature in the second image search result; Based on a preset classifier, classify and predict the prediction features to obtain a prediction score for each image in the second image search result; According to the prediction score and a preset threshold, a positive image in the second image search result is determined.
8. A multimodal image retrieval device, characterized in that: include: A first image search module, configured to perform an image search in a preset image library of a corresponding modality based on an image to be retrieved, and obtain a first image search result; The image to be retrieved includes a face and / or a body; An archive search module, used to perform an archive search in a preset centroid library of a corresponding modality based on the image to be retrieved, and determine a master file corresponding to the image to be retrieved from the preset archive library according to the archive search results; A second image search module, configured to perform topology recommendation based on the main file and the first image search results to obtain a recommended image; and performing an image search based on the recommended image to obtain a second image search result; The image retrieval module is used to obtain image retrieval results based on the first image search results, the main file and the predicted positive example images in the second image search results.
9. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the multimodal image retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal image retrieval method according to any one of claims 1 to 7 are implemented.