Retrieval method and apparatus, electronic device, computer readable storage medium
By generating multiple sketch images and combining them with a set of feature words, the target record set is retrieved from a pre-set database, which solves the problem of inaccurate retrieval results in the prior art and achieves higher retrieval accuracy and coverage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2026-03-24
AI Technical Summary
Existing fuzzy retrieval methods based on search text have the problem of inaccurate search results when retrieving image or video materials.
By acquiring the feature vocabulary set of the target object, multiple first-modal images (such as sketch images) are generated. Combined with the feature vocabulary set, the target record set is retrieved from the preset database. The preset image generation model and feature extraction model are used to improve the retrieval accuracy and coverage.
Without deviating from the limitations of characteristic vocabulary, the accuracy and diversity of search results are improved, the coverage of the search is enhanced, and the problem of low accuracy caused by fuzzy search is solved.
Smart Images

Figure CN116304150B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a retrieval method and apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of internet technology, when creating data objects such as images or videos, in order to improve the speed and effect of production, people usually search for materials from a material library that stores a large amount of image and / or video data based on the keywords corresponding to the object to be created, so as to quickly create data objects based on the searched materials.
[0003] For example, when creating advertisements, a common retrieval method is for electronic devices to receive the search text entered by the user, and then use fuzzy search to retrieve images, videos, etc. that match the search text from the material library as matching materials and display them to the user for selection.
[0004] Because this type of retrieval method in related technologies only uses fuzzy search based on the search text to retrieve matching materials, the retrieval results may not be accurate enough. Summary of the Invention
[0005] This disclosure provides a retrieval method and apparatus, electronic device, and computer-readable storage medium to accurately determine the target identity of a target object.
[0006] Firstly, this disclosure provides a retrieval method, which includes:
[0007] Obtain a first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object;
[0008] Based on the first feature vocabulary set, multiple first images of the first modality are generated, wherein the multiple first images correspond to the target object;
[0009] Based on the first feature vocabulary set and the multiple first images, a target record set is retrieved from a preset database;
[0010] The preset database is used to store object records of a first object. Each object record contains a second image and a second feature vocabulary set corresponding to the first object in a second modality. The second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image. The target record set includes at least one target object record. Each target object record is an object record in the preset database that satisfies a first preset condition. The object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record, which match at least one of the first feature vocabulary set and the plurality of first images.
[0011] Secondly, this disclosure provides a retrieval device, which includes:
[0012] The feature vocabulary acquisition unit is used to acquire a first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object;
[0013] An image acquisition unit is configured to generate multiple first images of a first modality based on the first feature vocabulary set, wherein the multiple first images correspond to the target object;
[0014] The retrieval unit is used to retrieve a target record set from a preset database based on the first feature vocabulary set and the multiple first images;
[0015] The preset database is used to store object records of a first object. Each object record contains a second image and a second feature vocabulary set corresponding to the first object in a second modality. The second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image. The target record set includes at least one target object record. Each target object record is an object record in the preset database that satisfies a first preset condition. The object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record, which match at least one of the first feature vocabulary set and the plurality of first images.
[0016] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described retrieval method.
[0017] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described retrieval method.
[0018] The embodiments provided in this disclosure obtain a first feature vocabulary set of the target object to be retrieved, and generate multiple first images corresponding to the target object with the first image modality based on the first feature vocabulary set. Then, a target record set corresponding to the target object can be retrieved from a preset database based on the multiple first images and the first feature vocabulary set.
[0019] In the embodiments provided in this disclosure, when retrieving a target record set corresponding to a target object from a preset database, it is not necessary to rely solely on the search text corresponding to the target object for retrieval. Instead, a first feature vocabulary set describing the object attributes of the target object is first obtained, and then multiple first images of a first modality corresponding to the target object, such as multiple sketch images, are automatically obtained based on the first feature vocabulary set. Subsequently, the target record set corresponding to the target object can be retrieved from the preset database based on the multiple first images of the first modality and the first feature vocabulary set. Compared to the text-based retrieval method in related technologies, since the multiple first images of the first modality automatically generated by the electronic device based on the first feature vocabulary set can complete the features not described in the feature vocabulary set without deviating from the limitations of the feature vocabulary in the first feature vocabulary set, the retrieval based on the multiple first images of the first modality and the first feature vocabulary set can increase the retrieval coverage while ensuring accuracy. Thus, the accuracy of the retrieved target record set can be improved without deviating from the features limited by the feature vocabulary in the first feature vocabulary set.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0022] Figure 1 A schematic diagram illustrating the implementation environment of the retrieval method provided in the embodiments of this disclosure;
[0023] Figure 2A flowchart of a retrieval method provided in an embodiment of this disclosure;
[0024] Figure 3 A flowchart for obtaining a first candidate record set provided in an embodiment of this disclosure;
[0025] Figure 4 A flowchart for obtaining first image feature data provided in an embodiment of this disclosure;
[0026] Figure 5 A flowchart for obtaining a target record set provided in this embodiment of the disclosure;
[0027] Figure 6 A block diagram of a retrieval device provided in an embodiment of this disclosure;
[0028] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0031] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0033] The use of user data in this technical solution complies with relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example, appropriate measures are taken to control access to personal information; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely identifying specific individuals.
[0034] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0035] In related technologies, when users create objects such as images or videos, they generally use the search text corresponding to the object to be created to retrieve matching materials from the material library. In this application scenario, after the electronic device receives the search text input by the user, the search method usually adopted is to extract the text features of the search text, and then, based on the text features, use fuzzy matching to match materials with image or video features that match the text features from the material library as matching materials.
[0036] However, since search terms can only describe some of the characteristics of an object, the number of matching materials obtained from fuzzy matching of search terms is usually large and often not accurate enough. After obtaining these matching materials, users often need to manually confirm them in order to obtain materials that meet their needs.
[0037] Please refer to Figure 1 This is a schematic diagram illustrating the implementation environment of the retrieval method provided in the embodiments of this disclosure. For example... Figure 1 As shown, the implementation environment may include server 101, terminal device 102, and network 103.
[0038] Server 101 can be a physical server, such as a blade server or a rack server, or it can be a virtual server, such as a server cluster deployed in the cloud. There are no restrictions on this.
[0039] In this embodiment of the disclosure, server 101 can be used to implement the retrieval method of any embodiment of the disclosure, so as to accurately retrieve the target record set corresponding to the target object based on the first feature vocabulary set of the target object to be retrieved.
[0040] Terminal device 102 may be a smartphone, laptop, desktop computer, tablet computer, etc. In this embodiment of the present disclosure, terminal device 102 may be used to provide server 101 with a first feature vocabulary set of the target object to be retrieved. The first feature vocabulary set may be used to describe the object attributes of the target object. The first feature vocabulary may be, for example, a key-value pair format of <attribute item: attribute value> to describe the object attributes of the target object from different aspects.
[0041] Network 103 can be a wireless network or a wired network, and it can be a local area network or a wide area network. Server 101 and terminal device 102 can communicate with each other through network 103.
[0042] In this embodiment of the disclosure, server 101 can be used to participate in implementing the retrieval method according to any embodiment of the disclosure. For example, it can be used to: obtain a first feature vocabulary set of a target object to be retrieved sent by terminal device 102, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object; generate multiple first images of a first modality based on the first feature vocabulary set, wherein the multiple first images correspond to the target object; retrieve a target record set from a preset database based on the first feature vocabulary set and the multiple first images; wherein the preset database is used to store object records of the first object, the object records include a second image of a second modality and a second feature vocabulary set corresponding to the first object, the second feature vocabulary in the second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image; the target record set includes at least one target object record, the target object record is an object record in the preset database that satisfies a first preset condition, the object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record, which matches at least one of the first feature vocabulary set and the multiple first images; and send the obtained target record set to terminal device 102 so that terminal device 102 can display the target record set for user viewing.
[0043] Understandable Figure 1 The implementation environment shown is merely illustrative and is by no means intended to limit this disclosure, its application, or its use. For example, although Figure 1 Only one server 101 and one terminal device 102 are shown, but this does not mean that the number of each is limited. The implementation environment may contain multiple servers 101 and multiple terminal devices 102.
[0044] To address the issue of low accuracy in retrieval processes in related technologies, this disclosure provides a retrieval method, please refer to the embodiments below. Figure 2This is a flowchart of the retrieval method provided in the embodiments of this disclosure. This method can be applied to electronic devices, such as... Figure 1 The server 101 shown; of course, in some embodiments, the electronic device can also be a terminal device, for example, it can also be directly a terminal device. Figure 1 The terminal device 102 shown can, upon obtaining a first feature vocabulary set of the target object to be retrieved, obtain a target record set corresponding to the target object retrieved by the server through interaction with the server. In this embodiment of the disclosure, unless otherwise specified, the method is described using an application to a server as an example.
[0045] like Figure 2 As shown, the retrieval method provided in this embodiment may include the following steps S201-S203, which will be described in detail below.
[0046] Step S201: Obtain the first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object.
[0047] The target object can be any virtual object. For example, the target object can be any material, such as advertising material, which can be a "human image". Of course, in actual implementation, the target object can also be materials such as "animals" or "vehicles" in advertising material, without special restrictions here.
[0048] The first feature vocabulary is used to describe the object attributes of an object from multiple dimensions using standardized feature words. That is, in related technologies, when retrieving advertising materials, the search is usually based on the descriptive information in the received requirements as keywords. However, this descriptive information is often colloquial and not standardized enough, and the matching materials retrieved based on this colloquial descriptive information are often inaccurate. Therefore, in this embodiment of the disclosure, in order to improve the accuracy of the search results, multiple standardized first feature vocabulary words can be used to describe the object attributes of the target object to be searched, thereby improving the accuracy of the matching materials retrieved by the electronic device.
[0049] In this embodiment of the disclosure, the first feature words in the first feature vocabulary set can be in the form of <attribute item: attribute value>. For example, when the target object is a "portrait" in advertising materials, the attribute items and their corresponding attribute value ranges can be as shown in Table 1:
[0050] Attributes Attribute value range gender men and women age Children, youth, middle-aged, and elderly face shape Square, triangle, oval, heart, circle Glasses Yes or no Eyebrow Thin, thick, long, short, narrow spacing, wide spacing nose shape Hump nose, straight nose, hooked nose, upturned nose hairstyle Short hair, straight hair, curly hair cheekbones Frontal sphenoid process, maxillary process, temporal process, orbital process scars, birthmarks, or moles Nothingness, Existence
[0051] Table 1
[0052] It should be noted that the contents of Table 1 are only used to illustrate the attribute items and their value ranges in the feature words. In actual implementation, the attribute items and their value ranges can also be set as needed, and no special restrictions are made here.
[0053] In addition, when constructing the first feature vocabulary set, the attribute value of a specific attribute item can be empty.
[0054] Step S202: Generate multiple first images of the first modality based on the first feature vocabulary set, wherein the multiple first images correspond to the target object.
[0055] In this embodiment of the disclosure, the first image of the first modality refers to an image with a relatively simple composition, such as a sketch, draft, or drawing; correspondingly, the second image of the second modality described below can be a color image with a relatively complex composition, such as an image with true colors acquired by an acquisition device.
[0056] That is, compared with the related technologies, when retrieving matching materials from the material library, there may be inaccuracies when only using the search text corresponding to the target object for retrieval. The embodiments of this disclosure are based on the first feature words in the standardized first feature word set for retrieval. On the other hand, in order to further improve the accuracy of the retrieval results, after obtaining the first feature word set, the electronic device automatically generates multiple first images of the first modality based on the first feature word set, so as to use the first images to further standardize the object attributes that the target object may have.
[0057] For example, in related technologies, when the advertising material to be searched is "portrait," the search text might be "a middle-aged man, without glasses, short hair, oval face, thick eyebrows." If a search is performed directly based on this description, the results are often inaccurate because the description may lack many features, and even the same feature may have different manifestations. However, in this embodiment, after obtaining the first feature vocabulary set of the target object, the electronic device can conveniently generate multiple sketch images corresponding to the first feature vocabulary set based on a pre-trained image generation model. Since the images can relatively accurately and comprehensively describe the features that the target object may possess, searching based on these sketch images can improve the accuracy of the retrieved target record set. At the same time, since different descriptions can be made for the same feature in the sketch images, searching based on these sketch images can ensure both accuracy and diversity of target object records in the target record set.
[0058] It should be noted that, in order to improve the retrieval speed, this embodiment is illustrated by the example of an electronic device automatically generating multiple first images of the first modality based on a first feature vocabulary set, that is, by the example of an electronic device automatically generating multiple sketch images. In actual implementation, it is also possible to directly generate an image with the same modality as the second image stored in the preset database based on the first feature vocabulary set. However, considering that it takes a relatively long time for an electronic device to generate images with rich colors and complex structures, and it will not have a significant effect on the accuracy of the retrieval results, this embodiment directly uses the example of generating multiple first images of the first modality.
[0059] Step S203: Based on the first feature vocabulary set and multiple first images, a target record set is retrieved from a preset database; wherein, the preset database is used to store object records of the first object, and the object record contains a second image of the second modality corresponding to the first object and a second feature vocabulary set, and the second feature vocabulary in the second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image; the target record set includes at least one target object record, and the target object record is an object record in the preset database that satisfies a first preset condition, and the object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record matches at least one of the first feature vocabulary set and multiple first images.
[0060] The preset database can be any type of database, such as MySQL (a relational database) or Redis (a key-value database); there are no special restrictions here.
[0061] In this embodiment of the disclosure, the preset database can be used to store object records of any first object. The first object can be any material obtained in advance, such as images, videos, etc., used as advertising materials. The object record can include the object identifier, name, one or more second images of the second modality, a second feature vocabulary set corresponding to each second image, and can also include first image feature data of the first modality corresponding to each second image.
[0062] It should be noted that, if the first object in the object record stored in the preset database is a video, the one or more second images of the second modality corresponding to the video can be video frames used to describe the key content of the video, such as key frames in the video.
[0063] In addition, in this embodiment of the present disclosure, it is not necessary to be limited to searching from a preset database based solely on the generated first image. Instead, if at least one of the first feature vocabulary set and / or multiple first images matches an object record in the preset database, the object record can be identified as the target object record.
[0064] As can be seen, the retrieval method provided in this embodiment of the present disclosure, when retrieving a target record set corresponding to a target object from a preset database, does not necessarily rely solely on the search text corresponding to the target object. Instead, it first obtains a first feature vocabulary set describing the object attributes of the target object, and then automatically obtains multiple first images of a first modality corresponding to the target object based on the first feature vocabulary set, such as multiple sketch images. Subsequently, the target record set corresponding to the target object can be retrieved from the preset database based on the multiple first images of the first modality and the first feature vocabulary set. Compared with the text-based retrieval method in related technologies, since the multiple first images of the first modality automatically generated by the electronic device based on the first feature vocabulary set can, without deviating from the limitation of the feature vocabulary in the first feature vocabulary set, complete the object features not described in these feature vocabulary sets as much as possible. This allows the retrieval of the target record set based on the multiple first images of the first modality and the first feature vocabulary set to increase the retrieval coverage while ensuring accuracy. Thus, the accuracy of the retrieved target record set can be improved without deviating from the features limited by the feature vocabulary in the first feature vocabulary set.
[0065] In some embodiments, the step S202 above, generating multiple first images of the first modality based on the first feature vocabulary set, includes: generating a third descriptive sentence for describing the target object based on the first feature vocabulary set; and inputting the third descriptive sentence multiple times into a preset image generation model for image generation processing to obtain multiple first images.
[0066] The third descriptive sentence is a statement generated based on the first feature vocabulary in the first feature vocabulary set according to a preset sentence pattern, used to describe the target object.
[0067] In this embodiment of the disclosure, when the target object to be searched is a "portrait" in the advertising material, the preset sentence can be in the form of "a {age feature}{gender}, {glasses feature}, {hair feature}, {face shape feature}, {eyebrow and eye feature}, {nose feature}, {mouth feature}, {other feature 1}, {other feature 2}". Of course, this is only an example for illustration. In actual implementation, the preset sentence can also be set according to the type of target object to be searched. No special limitation is made here.
[0068] For example, based on the feature vocabulary set {{'age':'middle-aged'}, {'gender':'male'}, {'face shape':'oval'}, {'glasses':'none'}, {'eyebrows':'thick'}, {'nose shape':"}, {'hairstyle':'short hair'}, {'cheekbones':"}, ...}, we can obtain the descriptive sentence "A middle-aged man, without glasses, short hair, oval face, thick eyebrows".
[0069] It should be noted that, in actual processing, in order to generate the third descriptive sentence quickly and accurately, a text generation model for generating descriptive sentences can be pre-trained. Thus, by inputting the first feature vocabulary set into the text generation model, a descriptive sentence with a preset sentence structure can be automatically generated. That is, in this embodiment of the disclosure, the step of generating a third descriptive sentence for describing the target object based on the first feature vocabulary set includes: inputting the first feature vocabulary set into the preset text generation model to obtain the third descriptive sentence.
[0070] The preset image generation model can be a pre-trained image generation model used to generate images of the first modality. For example, the preset image generation model can be the DALLE2 model.
[0071] After obtaining the third descriptive sentence, considering that even for the same content, the image generation model can have multiple representations of certain features when generating a sketch image, and also different representations of the same feature, the third descriptive sentence can be input multiple times into the pre-trained preset image generation model for image generation processing to obtain multiple first images.
[0072] It should be noted that in this embodiment, the example given is that the preset image generation model generates one first image each time based on the third descriptive sentence. In actual implementation, an image generation model that can generate multiple different images simultaneously based on the descriptive sentence can also be trained to reduce operations during image generation processing. No special limitation is made here.
[0073] As can be seen, in this embodiment of the present disclosure, in order to ensure the accuracy of the search results while also improving the diversity of the search results, after obtaining the first feature vocabulary set of the target object, a third descriptive sentence is generated based on the first feature vocabulary set, and the third descriptive sentence is input into a pre-trained preset image generation model, such as the DALLE2 model, for image generation processing. This can obtain multiple first images of the first modality that satisfy the first feature vocabulary set and also have diversity. Searching based on the first image can ensure the accuracy of the search results while also improving the diversity of the search results.
[0074] It should also be noted that, in some embodiments, in order to avoid the multiple first images automatically generated by the electronic device having too large a deviation, the multiple first images can also be displayed for user review after they are generated, so as to remove the first images that obviously do not meet the requirements.
[0075] In some embodiments, step S103 above, which retrieves a target record set from a preset database based on a first feature vocabulary set and multiple first images, may include: retrieving a first candidate record set from a preset database based on multiple first images, wherein the first candidate record set includes at least one candidate object record, and the candidate object record is an object record in which the second image matches any of the first images; and obtaining the target record set based on the first feature vocabulary set and the first candidate record set.
[0076] It should be noted that, in the process of realizing this disclosure, the inventors discovered that although there are retrieval methods in related technologies that use fuzzy retrieval based on images to obtain matching records, these methods often rely on images of the same modality. That is, they are often based on real-world images, requiring high standards for the retrieved images. Furthermore, the retrieval is typically based solely on image similarity matching. This image similarity matching often uses local visual features such as feature points and contours. On the one hand, this requires a high degree of matching between the retrieved image and the target object; otherwise, the matching records obtained based on the retrieved image may be inaccurate. On the other hand, the materials stored in the database, such as images or videos, may have been collected in different environments or by objects under different health conditions. Even when the object is a human figure, features such as facial features, body shape, and hair length can change. Therefore, the image similarity matching method based on local feature points in related technologies may also suffer from low accuracy.
[0077] In this embodiment of the disclosure, object records in the object records that match any one of the multiple first images in the first modality are firstly retrieved from the preset database based on the generated first images of the first modality. This is called "coarse screening". Then, the first candidate record set is "fine screening" based on the first feature vocabulary set, which can improve the accuracy of the search results.
[0078] Please refer to Figure 3 This is a flowchart of obtaining a first candidate record set provided in an embodiment of this disclosure. It addresses the problem in related technologies where image retrieval may only be based on matching local feature points, leading to inaccurate retrieval results. Figure 3As shown, in some embodiments, each object record stored in the preset database may also contain first image feature data of the first modality corresponding to the second image of the second modality; in this embodiment, retrieving a first candidate record set from the preset database based on multiple first images may include the following steps S301-S302:
[0079] Step S301: Input any first image into a preset image feature extraction model for image feature extraction processing to obtain second image feature data. The preset image feature extraction model includes a model for simultaneously extracting local detail features and global features of the image.
[0080] The preset image feature extraction model can be a pre-trained DOLG model. Since the DOLG model can extract feature data that integrates local detail features and global features of an image, using this preset image feature extraction model for image feature extraction processing can solve the problem of low accuracy that may be caused by similarity matching based only on local feature points in image features in related technologies.
[0081] Step S302: Perform similarity matching processing between the second image feature data and the first image feature data in each object record in the preset database, and select object records whose similarity meets the second preset condition to construct a first candidate record set.
[0082] In this embodiment of the disclosure, the similarity between different image feature data can be cosine similarity.
[0083] The second preset condition can be the object record corresponding to the similarity value located in the TopK or Top2*K position in the calculated similarity. Here, K can be 20 or can be set as needed, without special limitation.
[0084] As can be seen, in this embodiment of the present disclosure, by simultaneously extracting local detail features and global features of an image based on a preset image feature extraction model, the accuracy of retrieval results can be improved when performing image similarity matching. Furthermore, although the object records in the preset database store images of the second modality, the stored first image feature data is of the first modality, thereby enabling similarity matching based on image features of the same modality when performing image similarity matching, so as to further improve the accuracy of retrieval results.
[0085] Please refer to Figure 4 This is a flowchart of obtaining first image feature data provided in an embodiment of this disclosure. Figure 4As shown, in some embodiments, the first image feature data corresponding to each second image in the preset database can be obtained through the following steps S401-S402:
[0086] Step S401: Perform image mode conversion processing on the second image of the second mode to obtain the third image of the first mode, and the third image corresponds to the second image.
[0087] Since the second modality image contains far more information than the first modality image, in this embodiment of the disclosure, the second image can be pre-processed with image modality conversion to obtain the third image of the first modality, that is, the real material image stored in the preset database is converted into a sketch image.
[0088] Step S402: Input the third image into the preset image feature extraction model for image feature extraction processing to obtain the first image feature data.
[0089] After the third image is obtained through step S401, it can be input into a preset image feature extraction model, such as the DOLG model, for image feature extraction processing to obtain the first image feature data.
[0090] As can be seen, in this embodiment of the present disclosure, by storing the first image feature data of the first modality corresponding to the first image of the second modality in the object record of the preset database, matching can be performed based on the image feature data of the same modality during image matching processing, thereby improving the accuracy of the results.
[0091] Please refer to Figure 5 This is a flowchart of obtaining a target record set provided in an embodiment of this disclosure. Figure 5 As shown, in some embodiments, each feature word in the first feature vocabulary set and the second feature vocabulary set can be composed of attribute items and attribute values; in this embodiment, obtaining the target record set based on the first feature vocabulary set and the first candidate record set may include the following steps S501-S504:
[0092] Step S501: Based on the first feature vocabulary set and the second feature vocabulary set in the first candidate object record, obtain the third feature vocabulary set and the fourth feature vocabulary set, wherein the first candidate object record is any candidate object record in the first candidate record set, and the attribute values in the feature vocabulary contained in the third feature vocabulary set and the fourth feature vocabulary set are all non-empty and correspond to the same set of attribute items.
[0093] Specifically, since the object attributes of the target object, such as the face shape of a user's portrait, may change at different stages, the description of the target object in the first feature vocabulary set may not be accurate. Therefore, when "refining" the first candidate record set based on the first feature vocabulary set, it is not possible to simply match based on "keywords". In order to improve accuracy, in this embodiment of the disclosure, the attribute items of the first feature vocabulary set and the second feature vocabulary set in each candidate object record are first processed uniformly. That is, the first feature vocabulary set and the second feature vocabulary set are first filtered to obtain the third feature vocabulary set and the fourth feature vocabulary set with the same attribute item set and the attribute values in the attribute items are not empty.
[0094] For example, for feature vocabulary set 1: {<'attribute 1': 'x1'>, <'attribute 2': 'y1'>, <'attribute 3': ">, <'attribute 4': 'z1'>}; and feature vocabulary set 2: {<'attribute 1': ">, <'attribute 2': 'y2'>, <'attribute 3': 'm2'>, <'attribute 4': 'z2'>}, the set of attribute items that are contained in both sets and whose corresponding attribute values are not empty is {'attribute 2', 'attribute 4'}. Therefore, we can obtain feature vocabulary set 3 corresponding to feature vocabulary set 1: {<'attribute 2': 'y1'>, <'attribute 4': 'z1'>}, and feature vocabulary set 4 corresponding to feature vocabulary set 2: {<'attribute 2': 'y2'>, <'attribute 4': 'z2'>}.
[0095] Step S502: Generate a first descriptive sentence to describe the target object based on the third feature vocabulary set, and generate a second descriptive sentence to describe the first object corresponding to the second image in the first candidate object record based on the fourth feature vocabulary set.
[0096] The generation and processing of the first and second descriptive sentences are explained in the section on generating the third descriptive sentence above, and will not be repeated here.
[0097] Step S503: Obtain the first text feature data of the first descriptive sentence and the second text feature data of the second descriptive sentence.
[0098] Step S504: If the similarity between the first text feature data and the second text feature data meets the third preset condition, construct a target record set based on the first candidate object record.
[0099] After obtaining the first and second descriptive sentences, the matching target record set can be accurately selected from the first candidate record set based on the first feature vocabulary set by obtaining the text features of the first and second descriptive sentences respectively and calculating their similarity.
[0100] It should be noted that in the process of obtaining the first and second text feature data in step S503 above, the feature data can be extracted based on any text feature extraction model. However, in order to improve the accuracy of the final retrieval results, considering that in this embodiment of the present disclosure, when generating the first image, the third descriptive sentence generated based on the first feature vocabulary set is input into a preset image generation model, such as the DALLE2 model, where the text encoding layer first extracts text features, then the feature transformation layer maps the text features into image features, and finally the image features are input into the decoding layer for decoding to obtain the first image, it is necessary to consider that different text feature extraction models may have different attention scales to text semantics during training, which may lead to different extracted features. To perform similarity matching of text features at the same scale, in this embodiment of the disclosure, the steps S503 above, namely obtaining the first text feature data of the first descriptive sentence and obtaining the second text feature data of the second descriptive sentence, may include: inputting the first descriptive sentence into the text encoding layer of a preset image generation model for text encoding processing to obtain the first text feature data, and inputting the second descriptive sentence into the text encoding layer of the preset image generation model for text encoding processing to obtain the second text feature data. In this way, by using the text encoding layer in the preset image generation model used to generate the first image to extract the text feature data of the descriptive sentence, the feature extraction scale can be kept consistent, thereby improving the accuracy of the final retrieval results.
[0101] In some embodiments, obtaining a third feature vocabulary set and a fourth feature vocabulary set based on a first feature vocabulary set and a second feature vocabulary set in a first candidate object record may include: constructing a first attribute item set by obtaining attribute items with non-empty attribute values from the first feature vocabulary set, and constructing a second attribute item set by obtaining attribute items with non-empty attribute values from the second feature vocabulary set; obtaining the intersection of the first attribute item set and the second attribute item set; and filtering the third feature vocabulary set and the fourth feature vocabulary set from the first feature vocabulary set and the second feature vocabulary set based on the intersection.
[0102] That is, in this embodiment of the disclosure, when obtaining the third feature vocabulary set and the fourth feature vocabulary set, the attribute item sets whose attribute values are not empty can be obtained by first extracting the attribute item sets whose attribute values are not empty from the first and second feature vocabulary sets respectively, and obtaining the intersection of the two attribute item sets, so as to obtain the attribute item sets whose attribute values are not empty in common, and then filtering the first and second feature vocabulary sets based on the attribute item sets to obtain the third and fourth feature vocabulary sets respectively.
[0103] It should be noted that in the above embodiments, the target record set is obtained by first coarsely screening multiple first images from a preset database to obtain a first candidate record set, and then finely screening the first candidate record set to obtain a target record set based on a first feature vocabulary set. In some embodiments, the target record set can also be obtained by any one of the following A1-A3: A1: Retrieving the target record set from the preset database based on the first feature vocabulary set, wherein the target object record in the target record set is the object record in the preset database that matches the second feature vocabulary set with the first feature vocabulary set; A2: Retrieving the target record set from the preset database based on multiple first images, wherein the target object record in the target record set is the object record in the preset database that matches any of the second images with the first image; A3: Retrieving the second candidate record set from the preset database based on the first feature vocabulary set, and obtaining the target record set based on multiple first images and the second candidate record set, wherein the candidate object record in the second candidate object record set is the object record in the preset database that matches the second feature vocabulary set with the first feature vocabulary set.
[0104] Additionally, it should be noted that the above embodiments illustrate the application of this method to material retrieval, such as in the scenario of searching for advertising materials. In actual implementation, this method can also be applied to other application scenarios, such as in the scenario of user identification.
[0105] For example, when this retrieval method is applied to a user identity determination scenario, the first feature words in the first feature word set mentioned in step S201 can be feature words used to describe the user attributes of the target user whose identity is to be determined. The multiple first images can be multiple sketch images corresponding to the target user generated by an electronic device based on the first feature word set. In addition, the object record in the preset database can be the user record of any first user, where the first user is any user whose identity has been determined. The second image in the object record is the image of the second modality of the first user corresponding to the object record, that is, the real image of the user collected by the acquisition device. The second feature words in the second feature word set in the object record can be used to describe the user attributes possessed by the first user corresponding to the second image. Then, when determining the target identity of the target user, a target record set can be constructed by retrieving user records that meet the first preset conditions from the preset database based on the first feature word set and the multiple first images, and the identity of the first user corresponding to the user record with the highest matching degree in the target record set can be determined as the target identity of the target user.
[0106] Compared to related technologies that often rely solely on simple descriptions of the target user to manually create one or more sketches and then search a pre-defined database based on those sketches, which can lead to inaccuracies in user identification, the retrieval method provided in this embodiment addresses this issue. Since multiple first images of the first modality can be automatically generated by an electronic device, the method overcomes the time-consuming and laborious nature of manually creating sketches, which often results in sketches covering only a limited range of features. Automatic generation of first images, without deviating from the constraints of the first feature vocabulary set, allows for the rapid acquisition of multiple first images that satisfy feature diversity. Retrieving these images and the first feature vocabulary set enables a more accurate determination of the target user's identity based on the retrieved target record set.
[0107] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0108] In addition, this disclosure also provides a retrieval device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the retrieval methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding records in the method section and will not be repeated here.
[0109] Figure 6 This is a block diagram of a retrieval device provided in an embodiment of the present disclosure.
[0110] Reference Figure 6 This disclosure provides a retrieval device 600, which includes a feature word acquisition unit 601, an image acquisition unit 602, and a retrieval unit 603.
[0111] The feature vocabulary acquisition unit 601 is used to acquire a first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object.
[0112] The image acquisition unit 602 is used to generate multiple first images of a first modality based on a first feature vocabulary set, wherein the multiple first images correspond to a target object.
[0113] The retrieval unit 603 is used to retrieve a target record set from a preset database based on a first feature vocabulary set and multiple first images; wherein, the preset database is used to store object records of a first object, and the object record contains a second image of a second modality corresponding to the first object and a second feature vocabulary set, the second feature vocabulary in the second feature vocabulary set being used to describe the object attributes possessed by the first object corresponding to the second image; the target record set includes at least one target object record, the target object record being an object record in the preset database that satisfies a first preset condition, the object record satisfying the first preset condition including: the second image and / or the second feature vocabulary set contained in the object record matching at least one of the first feature vocabulary set and multiple first images.
[0114] In some embodiments, when the retrieval unit 603 retrieves a target record set from a preset database based on a first feature vocabulary set and multiple first images, it can be used to: retrieve a first candidate record set from the preset database based on multiple first images, wherein the first candidate record set includes at least one candidate object record, and the candidate object record is an object record containing a second image that matches any of the first images; and obtain the target record set based on the first feature vocabulary set and the first candidate record set.
[0115] In some embodiments, the object record further includes first image feature data of a first modality, which corresponds to a second image of a second modality. When the retrieval unit 603 retrieves a first candidate record set from a preset database based on multiple first images, it can be used to: input any first image into a preset image feature extraction model for image feature extraction processing to obtain second image feature data, wherein the preset image feature extraction model includes a model for simultaneously extracting local detail features and global features of the image; perform similarity matching processing between the second image feature data and the first image feature data in each object record in the preset database, and select object records whose similarity meets the second preset condition to construct the first candidate record set.
[0116] In some embodiments, the retrieval device further includes an image feature data extraction unit, which can be used to: perform image mode conversion processing on the second image of the second mode to obtain a third image of the first mode, wherein the third image corresponds to the second image; and input the third image into a preset image feature extraction model for image feature extraction processing to obtain first image feature data.
[0117] In some embodiments, each feature word in the first feature word set and the second feature word set consists of an attribute item and an attribute value. In this embodiment, when the retrieval unit 603 obtains the target record set based on the first feature word set and the first candidate record set, it can be used to: obtain a third feature word set and a fourth feature word set based on the first feature word set and the second feature word set in the first candidate object record, wherein the first candidate object record is any candidate object record in the first candidate record set, and the attribute values in the feature words contained in the third feature word set and the fourth feature word set are all non-empty and correspond to the same set of attribute items; generate a first descriptive sentence for describing the target object based on the third feature word set, and generate a second descriptive sentence for describing the first object corresponding to the second image in the first candidate object record based on the fourth feature word set; obtain the first text feature data of the first descriptive sentence, and obtain the second text feature data of the second descriptive sentence; and construct the target record set based on the first candidate object record when the similarity between the first text feature data and the second text feature data meets a third preset condition.
[0118] In some embodiments, when the retrieval unit 603 obtains a third feature vocabulary set and a fourth feature vocabulary set based on a first feature vocabulary set and a second feature vocabulary set in a first candidate object record, it can be used to: construct a first attribute item set by obtaining attribute items with non-empty attribute values from the first feature vocabulary set, and construct a second attribute item set by obtaining attribute items with non-empty attribute values from the second feature vocabulary set; obtain the intersection of the first attribute item set and the second attribute item set; and filter the third feature vocabulary set and the fourth feature vocabulary set from the first feature vocabulary set and the second feature vocabulary set based on the intersection.
[0119] In some embodiments, when the image acquisition unit 602 generates multiple first images of a first modality based on a first feature vocabulary set, it can be used to: generate a third descriptive sentence for describing a target object based on the first feature vocabulary set; input the third descriptive sentence multiple times into a preset image generation model for image generation processing to obtain multiple first images.
[0120] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0121] Reference Figure 7This disclosure provides an electronic device 700, which includes: at least one processor 701; at least one memory 702; and one or more I / O interfaces 703 connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs that can be executed by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 to enable the at least one processor 701 to perform the above-described retrieval method.
[0122] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the aforementioned retrieval method. The computer-readable storage medium may be volatile or non-volatile.
[0123] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described retrieval method.
[0124] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0125] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store preset (desired) information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0126] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0127] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0128] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0129] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0130] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0131] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0133] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A retrieval method, characterized in that, include: Obtain a first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object; Based on the first feature vocabulary set, multiple first images of the first modality are generated, wherein the multiple first images correspond to the target object; Based on the first feature vocabulary set and the multiple first images, a target record set is retrieved from a preset database; The preset database is used to store object records of a first object. Each object record contains a second image and a second feature vocabulary set corresponding to the first object in a second modality. The second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image. The target record set includes at least one target object record. Each target object record is an object record in the preset database that satisfies a first preset condition. The object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record, which match at least one of the first feature vocabulary set and the plurality of first images.
2. The method according to claim 1, characterized in that, The step of retrieving a target record set from a preset database based on the first feature vocabulary set and the multiple first images includes: Based on the plurality of first images, a first candidate record set is retrieved from the preset database, wherein the first candidate record set includes at least one candidate object record, and the candidate object record is an object record containing a second image that matches any of the first images; The target record set is obtained based on the first feature vocabulary set and the first candidate record set.
3. The method according to claim 2, characterized in that, The object record also includes first image feature data of the first modality, and the first image feature data corresponds to the second image of the second modality; The step of retrieving a first candidate record set from the preset database based on the plurality of first images includes: Any of the first images is input into a preset image feature extraction model for image feature extraction processing to obtain second image feature data, wherein the preset image feature extraction model includes a model for simultaneously extracting local detail features and global features of the image; The second image feature data is matched with the first image feature data in each object record in the preset database to perform similarity matching, and the object records whose similarity meets the second preset condition are selected to construct the first candidate record set.
4. The method according to claim 3, characterized in that, The first image feature data in the object record is obtained through the following steps: The second image of the second mode is subjected to image mode conversion processing to obtain the third image of the first mode, and the third image corresponds to the second image; The third image is input into the preset image feature extraction model for image feature extraction processing to obtain the first image feature data.
5. The method according to claim 2, characterized in that, Each feature word in the first feature vocabulary set and the second feature vocabulary set consists of an attribute item and an attribute value; The step of obtaining the target record set based on the first feature vocabulary set and the first candidate record set includes: Based on the first feature vocabulary set and the second feature vocabulary set in the first candidate object record, a third feature vocabulary set and a fourth feature vocabulary set are obtained, wherein the first candidate object record is any candidate object record in the first candidate record set, and the attribute values in the feature vocabulary contained in the third feature vocabulary set and the fourth feature vocabulary set are all non-empty and correspond to the same set of attribute items. Based on the third feature vocabulary set, a first descriptive sentence for describing the target object is generated; and based on the fourth feature vocabulary set, a second descriptive sentence for describing the first object corresponding to the second image in the first candidate object record is generated. Obtain the first text feature data of the first description sentence, and obtain the second text feature data of the second description sentence; If the similarity between the first text feature data and the second text feature data meets a third preset condition, the target record set is constructed based on the first candidate object record.
6. The method according to claim 5, characterized in that, The process of obtaining the third and fourth feature vocabulary sets based on the first feature vocabulary set and the second feature vocabulary set in the first candidate object record includes: Obtain attribute items whose attribute values are not empty from the first feature vocabulary set to construct a first attribute item set; and obtain attribute items whose attribute values are not empty from the second feature vocabulary set to construct a second attribute item set. Obtain the intersection of the first set of attribute items and the second set of attribute items; Based on the intersection, the third and fourth feature vocabulary sets are obtained by filtering from the first and second feature vocabulary sets.
7. The method according to claim 1, characterized in that, The step of generating multiple first images of the first modality based on the first feature vocabulary set includes: Based on the first feature vocabulary set, a third descriptive sentence is generated to describe the target object; The third descriptive sentence is input multiple times into a preset image generation model for image generation processing to obtain the multiple first images.
8. The method according to claim 2, characterized in that, The step of retrieving the target record set from a preset database based on the first feature vocabulary set and the multiple first images further includes any one of the following: Based on the first feature vocabulary set, the target record set is retrieved from the preset database, wherein the target object record in the target record set is an object record in the preset database that matches the second feature vocabulary set and the first feature vocabulary set. Based on the plurality of first images, a target record set is retrieved from the preset database, wherein the target object record in the target record is an object record in the preset database whose second image matches any of the first images; Based on the first feature vocabulary set, a second candidate record set is retrieved from the preset database, and based on the multiple first images and the second candidate record set, the target record set is obtained, wherein the candidate object records in the second candidate record set are object records in the preset database whose second feature vocabulary set matches the first feature vocabulary set.
9. A retrieval device, characterized in that, include: The feature vocabulary acquisition unit is used to acquire a first feature vocabulary set of the target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe the object attributes possessed by the target object; An image acquisition unit is configured to generate multiple first images of a first modality based on the first feature vocabulary set, wherein the multiple first images correspond to the target object; The retrieval unit is used to retrieve a target record set from a preset database based on the first feature vocabulary set and the multiple first images; The preset database is used to store object records of a first object. Each object record contains a second image and a second feature vocabulary set corresponding to the first object in a second modality. The second feature vocabulary set is used to describe the object attributes possessed by the first object corresponding to the second image. The target record set includes at least one target object record. Each target object record is an object record in the preset database that satisfies a first preset condition. The object record that satisfies the first preset condition includes: the second image and / or the second feature vocabulary set contained in the object record, which match at least one of the first feature vocabulary set and the plurality of first images.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the retrieval method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the retrieval method as described in any one of claims 1-8.
Citation Information
Patent Citations
Image-text data processing method and processor
CN115563334A