Image retrieval method and device, electronic equipment, storage medium and product

Through fine-tuning training of image encoder and text encoder, the feature vector of fisheye images is generated, which solves the problem that deep learning models cannot effectively extract fisheye image characteristics and improves the accuracy of image retrieval.

CN120407837APending Publication Date: 2025-08-01HUIZHOU DESAY SV AUTOMOTIVE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510530831.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing deep learning models cannot effectively extract image characteristics in fisheye images, resulting in reduced image retrieval accuracy.

Method used

By fine-tuning training of the image encoder and text encoder, the image feature vectors and text feature vectors of the fisheye image are generated, and the matching function is used to improve the matching accuracy of text description and image features.

Benefits of technology

The accuracy of image retrieval is improved, especially when processing fisheye images, the image features can be more accurately recognized, and the effect of image retrieval is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407837A_ABST
    Figure CN120407837A_ABST
Patent Text Reader

Abstract

The invention discloses an image retrieval method and device, electronic equipment, a storage medium and a product. The method comprises the following steps: constructing text description according to a text category; according to the text description, performing retrieval in a database by using a text encoder after fine tuning training to obtain a first target image; wherein the database comprises a plurality of fisheye images and image feature vectors of the fisheye images, the image feature vectors of the fisheye images are obtained through the image encoder after fine-tuning training, and the image encoder after fine-tuning training is obtained by performing fine-tuning training on the image encoder based on the fisheye images. According to the method, the image encoder is subjected to fine tuning training based on the fisheye image, so that the image encoder subjected to fine tuning training has the capability of accurately identifying the image feature vector in the fisheye image, the accuracy of text description and image matching can be improved, and the accuracy of image retrieval can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of image processing, and in particular, to an image retrieval method, apparatus, electronic device, storage medium, and product. Background Art

[0002] With the development of deep learning technology, deep learning models have been applied to various downstream tasks in the transportation field and have shown higher potential compared to traditional models.

[0003] In intelligent transportation applications, most of the images captured by in-vehicle panoramic cameras are fisheye images. A fisheye image is a special type of image with a very large viewing angle, but it introduces strong barrel distortion, causing straight lines to become curved and edges to stretch, resulting in a spherical or hemispherical visual effect.

[0004] However, existing deep learning models cannot effectively extract the image characteristics in fisheye images, thus reducing the accuracy of image retrieval. Summary of the Invention

[0005] The present invention provides an image retrieval method, apparatus, electronic device, storage medium, and product to solve the problem that existing deep learning models cannot effectively extract the image characteristics in fisheye images, thus reducing the accuracy of image retrieval.

[0006] According to one aspect of the present invention, there is provided an image retrieval method, including:

[0007] Constructing a text description according to the text category;

[0008] Retrieving a first target image in a database according to the text description by using a fine-tuned text encoder;

[0009] Wherein, the database includes a plurality of fisheye images and the image feature vectors of the fisheye images, the image feature vectors of the fisheye images are obtained by the fine-tuned image encoder, and the fine-tuned image encoder is obtained by fine-tuning the image encoder based on fisheye images; the fine-tuned text encoder has the function of matching text feature vectors with image feature vectors.

[0010] According to another aspect of the present invention, there is provided an image retrieval apparatus, including:

[0011] A construction module for constructing a text description according to the text category;

[0012] A retrieval module for retrieving a first target image in a database according to the text description by using a fine-tuned text encoder;

[0013] Among them, the database includes multiple fisheye images and the image feature vectors of the fisheye images. The image feature vectors of the fisheye images are obtained by the image encoder after fine-tuning training, and the image encoder after fine-tuning training is obtained by fine-tuning and training the image encoder based on the fisheye images; the text encoder after fine-tuning training has the function of matching the text feature vector with the image feature vector.

[0014] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0015] At least one processor;

[0016] And a memory communicatively connected to the at least one processor;

[0017] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image retrieval method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the image retrieval method according to any embodiment of the present invention when executed by a processor.

[0019] According to another aspect of the present invention, there is provided a computer program product including a computer program that implements the image retrieval method according to any embodiment of the present invention when executed by a processor.

[0020] The technical solution of the embodiment of the present invention fine-tunes and trains the image encoder based on the fisheye images, so that the image encoder after fine-tuning training has the ability to accurately identify the image feature vectors in the fisheye images, thereby improving the accuracy of matching between the text description and the image, and solving the problem that the existing deep learning model cannot effectively extract the image characteristics in the fisheye images, thereby reducing the accuracy of image retrieval, and achieving the beneficial effect of improving the accuracy of image retrieval.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic flow chart of an image retrieval method provided in Embodiment 1 of the present invention;

[0024] Figure 2 It is a schematic flow chart of an image retrieval method provided in Embodiment 2 of the present invention;

[0025] Figure 3 It is a schematic flow chart of an image retrieval method provided in Embodiment 3 of the present invention;

[0026] Figure 4 It is a schematic structural diagram of an image retrieval device provided in Embodiment 4 of the present invention;

[0027] Figure 5 It is a schematic structural diagram of an electronic device for an image retrieval method provided in Embodiment 5 of the present invention. Detailed implementation manners

[0028] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. It should be understood that the various steps recorded in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0029] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0030] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and do not limit the scope of these messages or information.

[0033] Embodiment 1

[0034] Figure 1 It is a schematic flowchart of an image retrieval method provided for Embodiment 1 of the present invention. This method is applicable to the situation where a user constructs an image set with specific category requirements, such as a fisheye image set of the traffic light category. This method can be executed by an image retrieval device, where the device can be implemented by software and / or hardware and is generally integrated on an electronic device. In this embodiment, the electronic device includes but is not limited to: computer devices.

[0035] As Figure 1 shown, an image retrieval method provided for Embodiment 1 of the present invention includes the following steps:

[0036] S110. Construct a text description according to the text category.

[0037] Among them, the text category can be provided by the user. If the user wants to retrieve images of a specific category, the text category can be provided. The text category represents the category of the object in the image and can also represent the category of the scene in the image. For example, if the user wants to retrieve images of the traffic light category, the provided text category can be "traffic light"; if the user wants to retrieve images at night, the provided text category can be "night"; if the user wants to retrieve images with traffic lights in a night scene, the provided text categories can be "night" and "traffic light".

[0038] In this embodiment, a corresponding text description can be constructed according to the text category provided by the user. There are various ways to construct the text description: Method 1: Construct the text description based on a template; Method 2: Construct the text description based on rules; Method 3: Generate the text description based on a model.

[0039] Exemplarily, when the text category is "traffic light", the constructed text description can be "an image with a traffic light"; when the text category is "night", the constructed text description can be "an image of a night scene"; when the text category is "traffic light" and "night", the constructed text description can be "an image including a traffic light in a night scene".

[0040] S120: According to the text description, use the fine-tuned text encoder to retrieve the first target image in the database.

[0041] Among them, the database includes multiple fish-eye images and the image feature vectors of the fish-eye images. The image feature vectors of the fish-eye images are obtained by the fine-tuned image encoder, and the fine-tuned image encoder is obtained by fine-tuning the image encoder based on the fish-eye images; the fine-tuned text encoder has the function of matching the text feature vector with the image feature vector.

[0042] Among them, the fine-tuned text encoder has the function of extracting the text feature vector by performing feature extraction on the text description; the fine-tuned image encoder has the function of extracting the image feature vector by performing feature extraction on the fish-eye image; the fine-tuned text encoder and the fine-tuned image encoder have the function of matching the text feature vector with the image feature vector.

[0043] In this embodiment, the image feature vectors included in the database are obtained by the fine-tuned image encoder performing image feature encoding on the fish-eye images; the text feature vectors included in the database are obtained by the fine-tuned text encoder performing text feature encoding on the text descriptions.

[0044] It should be noted that in the intelligent transportation application scenario, most of the images captured by the panoramic camera are fish-eye images. A fish-eye image is a special image, and its characteristic is that the viewing angle is extremely large, but it will introduce strong barrel distortion, resulting in straight lines becoming curved and the edges being stretched, forming a spherical or hemispherical visual effect.

[0045] In this embodiment, the text description is tokenized and then input into the trained text encoder for encoding to output a text feature vector with a fixed dimension. The text feature vector is matched with the image feature vectors of all the fish-eye images in the database, and a fish-eye image is determined as the first target image according to the matching result.

[0046] An image retrieval method provided by Embodiment 1 of the present invention first constructs a text description according to a text category; then, according to the text description, uses a fine-tuned text encoder to perform retrieval in a database to obtain a first target image; wherein, the database includes a plurality of fisheye images and image feature vectors of the fisheye images, the image feature vectors of the fisheye images are obtained by the fine-tuned image encoder, and the fine-tuned image encoder is obtained by fine-tuning the image encoder based on the fisheye images; the fine-tuned text encoder has the function of matching text feature vectors with image feature vectors. Using the above method to fine-tune the image encoder based on the fisheye images enables the fine-tuned image encoder to have the ability to accurately identify the image feature vectors in the fisheye images, thereby improving the accuracy of the matching between the text description and the image and enhancing the accuracy of image retrieval.

[0047] Based on the above embodiment, a variant embodiment of the above embodiment is proposed. It should be noted here that, for the sake of brevity of description, only the differences from the above embodiment are described in the variant embodiment.

[0048] In one embodiment, the method further includes: obtaining the fine-tuned text encoder and the fine-tuned image encoder after fine-tuning training;

[0049] The obtaining the fine-tuned text encoder and the fine-tuned image encoder after fine-tuning training includes: obtaining a plurality of first image feature vectors based on a plurality of first fisheye images and an image encoder; constructing a plurality of text descriptions of a plurality of first object categories according to the objects in the plurality of first fisheye images and / or constructing a plurality of text descriptions of a plurality of first scene categories according to the scenes in the plurality of first fisheye images; performing word segmentation on the plurality of text descriptions of the plurality of first object categories and / or the plurality of text descriptions of the plurality of first scene categories and then inputting them into the text encoder for feature vector encoding to obtain a plurality of first text feature vectors; for each first image feature vector, calculating the similarity between the first image feature vector and the plurality of first text feature vectors, calculating a contrast loss according to the similarity, and fine-tuning the image encoder and the text encoder according to the contrast loss until the similarity of the positive samples is higher than that of the negative samples, to obtain the fine-tuned text encoder and the fine-tuned image encoder; wherein, the positive sample is an image text vector in which the first image feature vector matches the first text feature vector, and the negative sample is an image text vector other than the matching image text vector among the plurality of first text feature vectors and the plurality of first image feature vectors.

[0050] Among them, a tokenizer can be used to tokenize the text description to obtain a sequence of tokens. Inputting the sequence of tokens into a RoBERTa encoder for feature encoding can obtain a sequence of text features. Taking the output of the first token of the last network layer of the RoBERTa encoder as the feature vector of the entire text description, i.e., the first text feature vector, the corresponding formula is as follows:

[0051] Z = RoBERTa(x TEXT )

[0052] X TEXT = Z[0]

[0053] In the above formula, x TEXT represents the sequence of tokens, Z represents the sequence of text features, Z[0] represents the output of the first token, and X TEXT represents the first text feature vector.

[0054] Among them, a contrastive loss is used to fine-tune and train the image encoder and the text encoder, and the corresponding formula is as follows:

[0055]

[0056] In the above formula, L cl represents the contrastive loss, I represents an image feature vector, T represents a text feature vector, sim(I, T) represents the similarity calculation between the image feature vector I and the text feature vector T, and τ represents the temperature coefficient to control the attention to positive and negative samples.

[0057] Furthermore, the first image feature vector includes scene information and an object feature vector. Correspondingly, obtaining multiple first image feature vectors based on multiple first fisheye images and an image encoder includes: for each first fisheye image, obtaining the scene information in the image; for each first fisheye image, after dividing the first fisheye image into multiple image blocks, performing dimensional transformation on each image block and performing feature extraction to obtain an image vector; performing linear transformation on the image vector to obtain a QKV vector; inputting the QKV vector into the image encoder for feature extraction to obtain an object feature vector.

[0058] Among them, the scene information may include time, weather, location, etc. The scene information needs to be extracted from the entire image, and the scene information in the image can be obtained through various methods, such as methods based on traditional computer vision and methods based on deep learning.

[0059] Among them, for the key object information in the image, such as pedestrians, motor vehicles, traffic lights, zebra crossings, etc., image feature extraction can be performed only based on the object information itself.

[0060] In this embodiment, given a fisheye image of H×W×C with a height of H, a width of W, and a channel number of C, the image is divided into multiple image patches of P×P; each image patch can be stretched into an image vector of dimension P 2 ×C; the stretched image patch is passed through a fully connected layer for feature extraction to obtain an image vector of a fixed size, and the corresponding formula is as follows:

[0061] X0 = W·Flatten(x IMAGE ) + b

[0062] In the above formula, x IMAGE represents the image patch, Flatten(x IMAGE ) represents stretching the image patch, W represents the weight matrix, b represents the bias term, and X0 represents the image vector.

[0063] The image vector is multiplied by the Query weight, Key weight, and Value weight respectively to obtain the Q (Query), K (Key), and V (Value) vectors, that is, the QKV vectors; the QKV vectors are input into the image encoder for feature extraction to obtain the object feature vector.

[0064] Further, inputting the QKV vectors into the image encoder for feature extraction to obtain the object feature vector includes: inputting the QKV vectors into the image encoder, and calculating the scaled dot product attention through the attention network layer of the image encoder to obtain the multi-head attention vector; the multi-head attention vector is integrated through the two-layer feed-forward network layer of the image encoder and then the object feature vector is output.

[0065] The corresponding formula is as follows:

[0066]

[0067] X IMAGE = max(0, XW1 + b1)W2 + b2

[0068] In the above formula, Attention(Q, K, V) represents calculating the scaled dot product attention, is used for scaling to prevent the dot product from being too large and causing the gradient to disappear, X represents the multi-head attention vector, and X IMAGE represents the object feature vector.

[0069] Embodiment 2

[0070] Figure 2 It is a schematic flowchart of an image retrieval method provided by the second embodiment of the present invention. The second embodiment is optimized on the basis of the above embodiments. For the content not detailed in this embodiment, please refer to Embodiment 1.

[0071] AsFigure 2 As shown in Figure 2 , an image retrieval method provided by the second embodiment of the present invention includes the following steps:

[0072] S210. Build a database.

[0073] Specifically, building the database includes: obtaining a plurality of second image feature vectors based on a plurality of second fisheye images and the image encoder after fine-tuning training; constructing text descriptions of object categories for the objects in the plurality of second fisheye images to obtain a plurality of second text descriptions, segmenting the plurality of second text descriptions and inputting them into the text encoder after fine-tuning training for feature extraction to obtain a plurality of second text feature vectors; storing the plurality of second fisheye images, the plurality of second text descriptions, the plurality of second image feature vectors, and the plurality of second text feature vectors in the database.

[0074] Among them, the process of obtaining a plurality of second image feature vectors based on a plurality of second fisheye images and the image encoder after fine-tuning training can refer to the process of obtaining a plurality of first image feature vectors based on a plurality of second fisheye images and the image encoder after fine-tuning training, which will not be elaborated here.

[0075] S220. Construct a text description according to the text category.

[0076] S230. Use the text encoder after fine-tuning training to perform text feature encoding on the text description to obtain a target text feature vector.

[0077] Among them, the text description is segmented and input into the text encoder after fine-tuning training for feature vector encoding to obtain a target text feature vector.

[0078] S240. Use the text encoder after fine-tuning training to calculate the similarity between the target text feature vector and a plurality of second image feature vectors in the database to obtain a plurality of first similarity scores.

[0079] The formula for similarity calculation is as follows:

[0080]

[0081] In the above formula, s(text,image) represents the similarity between the target text feature vector and the second image feature vector.

[0082] S250. Determine the first target image according to the plurality of first similarity scores.

[0083] Among them, the first target image can be determined in the following several ways:

[0084] Method 1: Use the fisheye image in the database with the highest corresponding similarity score as the first target image;

[0085] Method 2: Sort the fisheye images in the database in descending order of the first similarity score, and use the fisheye images with the top several rankings as the first target images;

[0086] Method 3: Use the fisheye images with corresponding similarity scores greater than a preset value as the first target images.

[0087] S260. According to the target fisheye image, use the fine-tuned trained image encoder to retrieve a second target image in the database.

[0088] Specifically, according to the target fisheye image, using the fine-tuned trained image encoder to retrieve a second target image in the database includes: obtaining a target image feature vector based on the target fisheye image and the fine-tuned trained image encoder; using the fine-tuned trained image encoder to calculate the similarity between the target image feature vector and multiple second image feature vectors in the database to obtain multiple second similarity scores; and determining the second target image according to the multiple second similarity scores.

[0089] Among them, the process of obtaining the target image feature vector based on the target fisheye image and the fine-tuned trained image encoder can refer to the process of obtaining multiple first image feature vectors based on multiple second fisheye images and the fine-tuned trained image encoder, which will not be elaborated here.

[0090] An image retrieval method provided in Embodiment 2 of the present invention uses a fine-tuned trained image encoder to perform image feature encoding on fisheye images to obtain second image feature vectors, uses a fine-tuned trained text encoder to perform text feature encoding on text descriptions to obtain second text feature vectors, constructs a database based on the first image feature vectors and the second text feature vectors, and uses the fine-tuned trained text encoder to perform text feature encoding on text descriptions constructed according to text categories to obtain target text feature vectors. Furthermore, the matching function of the fine-tuned trained text encoder is used to accurately match the target text feature vectors with multiple second image feature vectors in the database, which can improve the accuracy of image retrieval.

[0091] Embodiment 3

[0092] Figure 3 As shown in the figure, which is a schematic flowchart of an image retrieval method provided in Embodiment 3 of the present invention. Embodiment 3 is optimized on the basis of the above embodiments. For the content not elaborated in this embodiment, please refer to the above embodiments.

[0093] As Figure 3 shown, an image retrieval method provided in Embodiment 2 of the present invention includes the following steps:

[0094] S310. Build a database.

[0095] S320. Build a text description according to the text category.

[0096] S330. According to the text description, use the fine-tuned text encoder to retrieve the first target image in the database.

[0097] S340. According to the target fisheye image, use the fine-tuned image encoder to retrieve the second target image in the database.

[0098] S350. Perform image detection on the fisheye image to obtain the position bounding box of the object and the category information of the object.

[0099] Specifically, performing image detection on the fisheye image includes: extracting features of the fisheye image through the fine-tuned image encoder to obtain a feature map; using a detector to extract the position bounding box of the object and the category information of the object from the feature map.

[0100] Among them, the detector includes multiple convolutional layers and regression layers, and through the collaborative work of multiple network layers, the position boundary of the object and the category information of the object can be extracted from the feature map.

[0101] S360. Augment the data in the database.

[0102] Specifically, augmenting the data in the database includes: cropping the object image from the fisheye image according to the position bounding box and the category information of the object; inputting the object image into the fine-tuned image encoder for image feature encoding to obtain an object feature vector; storing the object feature vector, the category information of the object, and the scene information obtained from the fisheye image into the database.

[0103] Among them, the object image can be obtained from the fisheye image according to the detected position bounding box and the category information of the object. Exemplarily, if the category information is "pedestrian", find the pedestrian within the position bounding box in the fisheye image, and crop the corresponding image part to obtain the object image.

[0104] Among them, if the text description corresponding to the category information of the object is stored in the database, the category information of the object can be directly stored in the database; if there is no text description corresponding to the category information of the object in the data, it is necessary to build the corresponding text description according to the category information of the object, and store the category information of the object and the built text description into the database together.

[0105] In this embodiment, the scene information of the fisheye image can also be stored in the database. The corresponding text description, category information, fisheye image, scene information, text feature vector, and image feature vector are numbered the same and stored in a folder in the database. Thus, when retrieving the fisheye image from the database, other information with the same number can be obtained from the corresponding folder.

[0106] An image retrieval method provided in Embodiment 3 of the present invention. After performing image retrieval, the method can also perform image detection. The fine-tuned trained image encoder is used to effectively extract features from the fisheye image to obtain a feature map, and image detection is performed according to the feature map. Integrating image retrieval and image detection into one can reduce the overlapping steps in the two processes, improve the processing efficiency, and achieve the detection of fisheye images. In addition, by expanding the data in the database, the method can achieve efficient and automated management of newly acquired data, enrich the types of data in the database, and further improve the accuracy of image retrieval.

[0107] Embodiment 4

[0108] Figure 4 It is a schematic structural diagram of an image retrieval device provided in Embodiment 4 of the present invention. The device is applicable to the situation where a user constructs an image set with specific category requirements, such as a fisheye image set of the traffic light category. The device can be implemented by software and / or hardware and is generally integrated on an electronic device.

[0109] As Figure 4 shown, the device includes: a construction module 110 and a first retrieval module 120.

[0110] The construction module 110 is used to construct a text description according to the text category;

[0111] The first retrieval module 120 is used to retrieve the first target image in the database according to the text description by using the fine-tuned trained text encoder;

[0112] Among them, the database includes multiple fisheye images and the image feature vectors of the fisheye images. The image feature vectors of the fisheye images are obtained by the fine-tuned trained image encoder, and the fine-tuned trained image encoder is obtained by fine-tuning the image encoder based on the fisheye image; the fine-tuned trained text encoder has the function of matching the text feature vector with the image feature vector.

[0113] In this embodiment, the device first constructs a text description according to the text category through the construction module 110; then, through the first retrieval module 120, according to the text description, using the fine-tuned text encoder, retrieves the first target image in the database; wherein, the database includes a plurality of fish-eye images and the image feature vectors of the fish-eye images, and the image feature vectors of the fish-eye images are obtained through the fine-tuned image encoder, and the fine-tuned image encoder is obtained by fine-tuning the image encoder based on the fish-eye images; the fine-tuned text encoder has the function of matching text feature vectors with image feature vectors.

[0114] This embodiment provides an image retrieval device, which can improve the accuracy of image retrieval.

[0115] Furthermore, the device further includes: a training module, which is used to obtain the fine-tuned text encoder and the fine-tuned image encoder after fine-tuning training.

[0116] On the basis of the above optimization, the training module includes:

[0117] The first generation sub-module is used to obtain a plurality of first image feature vectors based on a plurality of first fish-eye images and an image encoder;

[0118] The construction sub-module is used to construct text descriptions of a plurality of first object categories according to the objects in the plurality of first fish-eye images and / or construct text descriptions of a plurality of first scene categories according to the scenes in the plurality of first fish-eye images;

[0119] The second generation sub-module is used to tokenize the text descriptions of the plurality of first object categories and / or the text descriptions of the plurality of first scene categories and then input them into the text encoder for feature vector encoding to obtain a plurality of first text feature vectors;

[0120] The fine-tuning sub-module is used to calculate the similarity between the first image feature vector and the plurality of first text feature vectors for each first image feature vector, calculate the contrast loss according to the similarity, and fine-tune the image encoder and the text encoder according to the contrast loss until the similarity of the positive samples is higher than that of the negative samples, so as to obtain the fine-tuned text encoder and the fine-tuned image encoder; wherein, the positive sample is an image text vector in which the first image feature vector matches the first text feature vector, and the negative sample is an image text vector other than the matching image text vector among the plurality of first text feature vectors and the plurality of first image feature vectors.

[0121] Based on the above technical solution, the first image feature vector includes scene information and an object feature vector, and the first generation sub-module includes:

[0122] An acquisition unit, configured to acquire scene information in the image for each first fisheye image;

[0123] A first transformation unit, configured to, for each first fisheye image, after dividing the first fisheye image into a plurality of image blocks, perform dimensional transformation on each image block and perform feature extraction to obtain an image vector;

[0124] A second transformation unit, configured to perform a linear transformation on the image vector to obtain a QKV vector;

[0125] An input unit, configured to input the QKV vector into the image encoder to perform feature extraction to obtain an object feature vector.

[0126] Based on the above technical solution, the input unit includes:

[0127] A calculation sub-unit, configured to input the QKV vector into the image encoder, and calculate scaled dot-product attention through the attention network layer of the image encoder to obtain a multi-head attention vector;

[0128] An integration sub-unit, configured to perform information integration on the multi-head attention vector through two feed-forward network layers of the image encoder and then output an object feature vector.

[0129] Further, the device further includes a construction module, and the construction module is configured to construct the database.

[0130] Based on the above optimization, the construction module includes:

[0131] A third generation sub-module, configured to obtain a plurality of second image feature vectors based on a plurality of second fisheye images and the image encoder after fine-tuning training;

[0132] A fourth generation sub-module, configured to construct a text description of the object category for the objects in the plurality of second fisheye images to obtain a plurality of second text descriptions, tokenize the plurality of second text descriptions, and input them into the text encoder after fine-tuning training for feature extraction to obtain a plurality of second text feature vectors;

[0133] A storage sub-module, configured to store the plurality of second fisheye images, the plurality of second text descriptions, the plurality of second image feature vectors, and the plurality of second text feature vectors in the database.

[0134] Further, the first retrieval module 120 includes:

[0135] An encoding sub-module, configured to perform text feature encoding on the text description using the text encoder after fine-tuning training to obtain a target text feature vector;

[0136] The first similarity calculation sub-module is used to calculate the similarity between the target text feature vector and multiple second image feature vectors in the database using the text encoder after fine-tuning training, and obtain multiple first similarity scores;

[0137] The first determination sub-module is used to determine the first target image according to the multiple first similarity scores.

[0138] Furthermore, the device further includes a second retrieval module, which is used to retrieve a second target image in the database according to the target fisheye image using the image encoder after fine-tuning training.

[0139] On the basis of the above optimization, the second retrieval module includes:

[0140] The fifth generation sub-module is used to obtain a target image feature vector based on the target fisheye image and the image encoder after fine-tuning training;

[0141] The second similarity calculation sub-module is used to calculate the similarity between the target image feature vector and multiple second image feature vectors in the database using the image encoder after fine-tuning training, and obtain multiple second similarity scores;

[0142] The second determination sub-module is used to determine the second target image according to the multiple second similarity scores.

[0143] Furthermore, the device further includes a detection module, which is used to perform image detection on the fisheye image.

[0144] On the basis of the above optimization, the detection module includes:

[0145] The feature extraction sub-module is used to extract features of the fisheye image through the image encoder after fine-tuning training to obtain a feature map;

[0146] The extraction sub-module is used to extract the position bounding box and category information of the object from the feature map using a detector.

[0147] Furthermore, the device further includes an expansion module, which is used to perform data expansion on the database.

[0148] On the basis of the above optimization, the expansion module includes:

[0149] The cropping sub-module is used to crop the object image from the fisheye image according to the position bounding box and the category information of the object;

[0150] The feature encoding sub-module is used to input the object image into the image encoder after fine-tuning training for image feature encoding to obtain an object feature vector;

[0151] A storage sub-module, configured to store the object feature vector, the category information of the object, and the scene information obtained from the fish-eye image into the database.

[0152] The above image retrieval device can execute the image retrieval method provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.

[0153] Embodiment 5

[0154] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0155] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0156] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0157] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the image retrieval method.

[0158] In some embodiments, the image retrieval method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image retrieval method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the image retrieval method by any other suitable means (e.g., by means of firmware).

[0159] The various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] In some embodiments, the image retrieval method may be implemented as a computer program that is invisibly included in a computer program product. When the computer program is executed by a processor, it implements the image retrieval method of the present invention. The computer program product can be understood as a software product that mainly implements its solution through a computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0161] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0163] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0164] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0165] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0166] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image retrieval method, characterized in that, The method includes: Constructing a text description according to the text category; Retrieving a first target image in the database using the text encoder after fine-tuning training according to the text description; Wherein, the database includes a plurality of fisheye images and image feature vectors of the fisheye images, the image feature vectors of the fisheye images are obtained by the image encoder after fine-tuning training, and the image encoder after fine-tuning training is obtained by fine-tuning training the image encoder based on the fisheye images; the text encoder after fine-tuning training has the function of matching text feature vectors and image feature vectors.

2. The method according to claim 1, wherein The method further includes: obtaining the text encoder after fine-tuning training and the image encoder after fine-tuning training; The obtaining the text encoder after fine-tuning training and the image encoder after fine-tuning training includes: Obtaining a plurality of first image feature vectors based on a plurality of first fisheye images and an image encoder; Constructing text descriptions of a plurality of first object categories according to the objects in the plurality of first fisheye images and / or constructing text descriptions of a plurality of first scene categories according to the scenes in the plurality of first fisheye images; Performing word segmentation on the text descriptions of the plurality of first object categories and / or the text descriptions of the plurality of first scene categories and then inputting them into a text encoder for feature vector encoding to obtain a plurality of first text feature vectors; For each first image feature vector, calculating the similarity between the first image feature vector and the plurality of first text feature vectors, calculating a contrast loss according to the similarity, and fine-tuning and training the image encoder and the text encoder according to the contrast loss until the similarity of the positive samples is higher than that of the negative samples, to obtain the text encoder after fine-tuning training and the image encoder after fine-tuning training; Wherein, the positive sample is an image-text vector in which the first image feature vector matches the first text feature vector, and the negative sample is an image-text vector other than the matching image-text vector among the plurality of first text feature vectors and the plurality of first image feature vectors.

3. The method according to claim 2, wherein The first image feature vector includes scene information and an object feature vector. Correspondingly, the obtaining a plurality of first image feature vectors based on a plurality of second fisheye images and the image encoder after fine-tuning training includes: For each first fisheye image, obtaining the scene information in the image; For each first fisheye image, after dividing the first fisheye image into a plurality of image blocks, performing dimensional transformation on each image block and then performing feature extraction to obtain an image vector; Performing a linear transformation on the image vector to obtain a QKV vector; Inputting the QKV vector into the image encoder for feature extraction to obtain an object feature vector.

4. The method according to claim 3, wherein The inputting the QKV vector into the image encoder for feature extraction to obtain an object feature vector includes: Inputting the QKV vector into the image encoder, and calculating scaled dot-product attention through the attention network layer of the image encoder to obtain a multi-head attention vector; Integrate the multi - head attention vector through two - layer feed - forward network layers of the image encoder and then output the object feature vector.

5. The method according to claim 1 or 2, characterized in that, The method further includes: constructing the database; The constructing the database includes: Obtain multiple second image feature vectors based on multiple second fisheye images and the fine - tuned image encoder; Construct text descriptions of object categories for objects in the multiple second fisheye images to obtain multiple second text descriptions. After tokenizing the multiple second text descriptions, input them into the fine - tuned text encoder for feature extraction to obtain multiple second text feature vectors; Store the multiple second fisheye images, the multiple second text descriptions, the multiple second image feature vectors, and the multiple second text feature vectors in the database.

6. The method according to claim 1, characterized in that, According to the text description, use the fine - tuned text encoder to retrieve the first target image in the database, including: Use the fine - tuned text encoder to perform text feature encoding on the text description to obtain the target text feature vector; Use the fine - tuned text encoder to calculate the similarity between the target text feature vector and multiple second image feature vectors in the database to obtain multiple first similarity scores; Determine the first target image according to the multiple first similarity scores.

7. The method according to claim 1, characterized in that The method further includes: according to the target fisheye image, use the fine - tuned image encoder to retrieve the second target image in the database.

8. The method according to claim 7, wherein The retrieving the second target image according to the target fisheye image and using the fine - tuned image encoder in the database includes: Obtain the target image feature vector based on the target fisheye image and the fine - tuned image encoder; Use the fine - tuned image encoder to calculate the similarity between the target image feature vector and multiple second image feature vectors in the database to obtain multiple second similarity scores; Determine the second target image according to the multiple second similarity scores.

9. The method according to claim 1, characterized in that, The method further includes: performing image detection on the fisheye image; The performing image detection on the fisheye image includes: Extract features of the fisheye image through the fine - tuned image encoder to obtain a feature map; Use the detector to extract the position bounding box of the object and the category information of the object from the feature map.

10. The method according to claim 9, characterized in that The method further includes: performing data augmentation on the database; The performing data augmentation on the database includes: Crop the object image from the fisheye image according to the position bounding box and the category information of the object; Input the object image into the fine - tuned image encoder for image feature encoding to obtain the object feature vector; Store the object feature vector, the category information of the object, and the scene information obtained from the fisheye image in the database.

11. An image retrieval device, characterized in that, The device includes: A construction module, configured to construct text descriptions according to text categories; A retrieval module, configured to retrieve the first target image in the database according to the text description by using the fine - tuned text encoder; Among them, the database includes a plurality of fish-eye images and the image feature vectors of the fish-eye images. The image feature vectors of the fish-eye images are obtained through the image encoder after fine-tuning training, and the image encoder after fine-tuning training is obtained by fine-tuning the image encoder based on the fish-eye images; the text encoder after fine-tuning training has the function of matching the text feature vectors with the image feature vectors.

12. An electronic device, characterized in that, The electronic device includes: At least one processor; And a memory communicatively connected to the at least one processor; Among them, the memory stores a computer program executable by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor can execute the image retrieval method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the image retrieval method according to any one of claims 1-10 when executed by a processor.

14. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program implements the image retrieval method according to any one of claims 1-10 when executed by a processor.

Citation Information

Cited By

  • Image retrieval method and device, electronic equipment and storage medium

    CN121561128A