Image retrieval method, device and equipment based on endogenous relevance of image database
By semantic encoding of the images in the image database and the images to be retrieved, and using the intrinsic correlation between images to be retrieved, the problem of insufficient accuracy and correlation of image search results in the prior art is solved, and a more efficient and accurate image search effect is achieved.
Patent Information
- Application Number
- CN202510212682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the accuracy and correlation of image retrieval results are low, especially when dealing with complex scenarios, the feature extraction capability of deep learning models in the semantic dimension is limited, and the contextual relationships and scene similarities between images are not fully utilized.
By using an image search method based on the endogenous correlation of the image database, by semantic encoding of each image in the image database and the image to be retrieved, the semantic expression optimization of the image to be retrieved is performed to enhance the part related to the target in its semantic coding feature representation.
It significantly improves the accuracy and efficiency of image retrieval, can more accurately identify and match images semantically similar to the image to be retrieved, and improves the correlation of search results.
Smart Images

Figure CN120196782A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and in particular to an image retrieval method, device and equipment based on the endogenous relevance of an image database. Background Art
[0002] With the rapid development of the Internet and the popularization of digital media, the quantity of image data has increased explosively. How to quickly and accurately find the information required by users from the vast amount of image data has become an important research topic. Early image retrieval systems usually required users to manually add descriptive tags to each image and then perform retrieval through text matching. This method relies on manual annotation, which not only has a huge workload but is also easily affected by subjective factors, making it difficult to guarantee the accuracy of retrieval results.
[0003] In recent years, deep learning technology has been widely applied to the field of image retrieval due to its powerful representation learning ability. However, as the depth of the deep neural network increases, the features will gradually become abstract. If only deep features are considered, some important details in the image may be ignored, resulting in limited feature extraction ability of traditional deep learning models in the semantic dimension and unsatisfactory image retrieval effects in complex scenarios. Moreover, when most image retrieval systems extract image features, they usually regard each image as an independent individual and directly calculate the similarity, ignoring the correlation relationships such as the context relationship and scene similarity between images, resulting in room for improvement in the accuracy and relevance of retrieval results. Summary of the Invention
[0004] The present invention provides an image retrieval method, device and equipment based on the endogenous relevance of an image database to solve the technical problem of low accuracy and relevance of image retrieval results in the prior art.
[0005] The present invention provides an image retrieval method based on the endogenous relevance of an image database, including: obtaining a to-be-retrieved image input by a user; performing image semantic encoding on the to-be-retrieved image to obtain a to-be-retrieved image semantic encoding feature vector; optimizing the semantic query response of the to-be-retrieved image semantic encoding feature vector based on a set of library image semantic encoding feature vectors obtained from the image database to obtain an optimized semantic encoding vector of the to-be-retrieved image; generating a retrieval result based on the semantic relevance between each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the to-be-retrieved image.
[0006] According to an image retrieval method based on the endogenous relevance of an image database provided by the present invention, before obtaining the to-be-retrieved image input by a user, it further includes: obtaining an image database; performing image semantic encoding on each image in the image database to obtain a set of library image semantic encoding feature vectors.
[0007] An image retrieval method based on the endogenous relevance of an image database, which optimizes the semantic query response for the semantic encoding feature vector of the image to be retrieved based on the set of semantic encoding feature vectors of the library images obtained from the image database, and obtains the optimized semantic encoding vector of the image to be retrieved, including: extracting the significant features of the set of semantic encoding feature vectors of the library images as a prompt template; based on the prompt template, performing cross-domain optimization query encoding based on the attention mechanism on the semantic encoding feature vector of the image to be retrieved and the set of semantic encoding feature vectors of the library images, so as to obtain the optimized semantic encoding vector of the image to be retrieved.
[0008] An image retrieval method based on the endogenous relevance of an image database, which extracts the significant features of the set of semantic encoding feature vectors of the library images as a prompt template, including: constructing a key matrix based on the set of semantic encoding feature vectors of the library images; extracting the maximum value of each key vector in the key matrix to obtain the key matrix significant feature vector as the prompt template.
[0009] An image retrieval method based on the endogenous relevance of an image database, which constructs a key matrix based on the set of semantic encoding feature vectors of the library images, including: linearly embedding and encoding each semantic encoding feature vector in the set of semantic encoding feature vectors of the library images using a key embedding matrix to obtain a set of linearly transformed semantic encoding feature vectors of the library images; using the linearly transformed semantic encoding feature vectors of the library images as key vectors, and arranging the set of linearly transformed semantic encoding feature vectors of the library images in a matrix to obtain the key matrix.
[0010] An image retrieval method based on the endogenous relevance of an image database, which performs cross-domain optimization query encoding based on the attention mechanism on the semantic encoding feature vector of the image to be retrieved and the set of semantic encoding feature vectors of the library images based on the prompt template, so as to obtain the optimized semantic encoding vector of the image to be retrieved, including: linearly embedding and encoding the semantic encoding feature vector of the image to be retrieved using a query embedding matrix and a value embedding matrix respectively to obtain a value vector and a query vector; inputting the query vector, the value vector, each key vector in the key matrix, and the prompt template into a heterogeneous transformer structure optimized based on template prompts respectively to obtain a set of cross-domain optimized query encoding feature vectors of the image to be retrieved; calculating the position-wise mean vector of the set of cross-domain optimized query encoding feature vectors of the image to be retrieved to obtain the optimized semantic encoding vector of the image to be retrieved.
[0011] The present invention also provides an image retrieval device based on the endogenous relevance of an image database, including: an image acquisition module for acquiring a to-be-retrieved image input by a user; an image semantic encoding and processing module for performing image semantic encoding on the to-be-retrieved image to obtain a to-be-retrieved image semantic encoding feature vector; a semantic query response optimization processing module for performing semantic query response optimization on the to-be-retrieved image semantic encoding feature vector based on a set of library image semantic encoding feature vectors obtained from the image database to obtain an optimized semantic encoding vector of the to-be-retrieved image; and a retrieval result generation module for generating a retrieval result based on the semantic relevance between each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the to-be-retrieved image.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the image retrieval method based on the endogenous relevance of an image database as described in any one of the above.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image retrieval method based on the endogenous relevance of an image database as described in any one of the above.
[0014] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the image retrieval method based on the endogenous relevance of an image database as described in any one of the above.
[0015] The image retrieval method, device, and equipment based on the endogenous relevance of an image database provided by the present invention respectively perform image semantic encoding on each image in the image database and the to-be-retrieved image input by the user by using an image processing technology based on deep learning to capture the semantic encoding feature representations of each image, and utilize the internal relevance between each image in the image database and the to-be-retrieved image to optimize the semantic expression of the to-be-retrieved image and enhance the part related to the target in the semantic encoding feature representation of the to-be-retrieved image; and then perform semantic matching calculations between the optimized semantic encoding features of the to-be-retrieved image and each image in the image database respectively, so as to intelligently return the image most relevant to the to-be-retrieved image as the retrieval result. In this way, the present invention can make full use of the internal correlation structure information in the image database and effectively improve the accuracy and efficiency of image retrieval. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0017] Figure 1 It is one of the schematic flowcharts of the image retrieval method based on the endogenous relevance of the image database provided by the embodiments of the present invention.
[0018] Figure 2 It is another schematic flowchart of the image retrieval method based on the endogenous relevance of the image database provided by the embodiments of the present invention.
[0019] Figure 3 It is the schematic structural diagram of the image retrieval device based on the endogenous relevance of the image database provided by the embodiments of the present invention.
[0020] Figure 4 It is the schematic physical structure diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0021] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.
[0022] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0023] Deep learning models have taken a dominant position in image retrieval applications due to their excellent feature learning capabilities. However, as the depth of neural networks increases, the extracted features become increasingly abstract. Relying solely on deep features may miss some subtle but important elements in images, making traditional deep learning-based models ineffective in processing semantic-level feature extraction, especially performing poorly in complex scenarios.
[0024] In addition, most existing image retrieval solutions tend to view each image in isolation when processing image feature extraction, only comparing them based on their similarity, and failing to effectively utilize the inherent relationships between images, such as context relationships and scene consistency. This approach ignores the mutual relevance within the image collection, thus limiting the quality and relevance of retrieval results and still leaving room for improvement.
[0025] Based on this, the present invention provides an image retrieval method based on the endogenous relevance of an image database, which can improve the accuracy and efficiency of image retrieval.
[0026] Please refer to Figure 1 , Figure 1 which is one of the schematic flowcharts of the image retrieval method based on the endogenous relevance of the image database provided by the embodiments of the present invention. In this embodiment, the image retrieval method based on the endogenous relevance of the image database includes steps S110 to S140, and the specific steps are as follows: S110: Obtain the image to be retrieved input by the user.
[0027] S120: Perform image semantic encoding on the image to be retrieved to obtain the semantic encoding feature vector of the image to be retrieved.
[0028] S130: Optimize the semantic query response for the semantic encoding feature vector of the image to be retrieved based on the set of semantic encoding feature vectors of the library images obtained from the image database, to obtain the optimized semantic encoding vector of the image to be retrieved.
[0029] S140: Generate retrieval results based on the semantic relevance between each semantic encoding feature vector in the set of semantic encoding feature vectors of the library images and the optimized semantic encoding vector of the image to be retrieved.
[0030] In this embodiment, first, the image to be retrieved input by the user is obtained. Considering that directly comparing the pixel values of images not only has a large computational amount but is also easily affected by factors such as illumination, perspective, and occlusion. Therefore, to improve the efficiency and accuracy of image retrieval, in this embodiment, image semantic encoding is also performed on each image in the image database and the image to be retrieved, so as to perform query matching based on the semantic content of the images to achieve more accurate image retrieval.
[0031] The image database contains a collection of images for searching, and the image to be retrieved entered by the user is the image that the user hopes to find similar or relevant to. Only when these two types of images are obtained can the subsequent image retrieval process begin to intelligently return the most relevant image to the image to be retrieved as the retrieval result, thereby improving the accuracy and efficiency of image retrieval.
[0032] Optionally, in an embodiment of the present application, a graph data image library containing specific themes or categories can be constructed, such as landscape photos, portraits, etc.; it can also be a general image database covering various types of images. These images can come from public data sets or private databases collected and organized by enterprises or individuals. The process of obtaining the image database generally includes tasks such as image acquisition, storage, and index establishment to ensure that the images can be accessed and processed efficiently.
[0033] Optionally, in an embodiment of the present application, for the image to be retrieved entered by the user, it can be obtained in the following two ways: 1) The user directly submits it by uploading or inputting an image file.
[0034] 2) The user takes a real-time image through the camera, and the system automatically receives this image as the retrieval object.
[0035] In some embodiments, the steps before obtaining the image to be retrieved entered by the user may specifically further include: obtaining the image database; performing image semantic encoding on each image in the image database to obtain a set of library image semantic encoding feature vectors.
[0036] In this embodiment, it also includes image semantic encoding performed on each image in the image database and the image to be retrieved.
[0037] In a specific example of the present application, an image semantic encoder based on a CNN-RNN hybrid network model can be used to process each image in the image database and the image to be retrieved respectively.
[0038] Among them, the image semantic encoder based on the CNN-RNN hybrid network model is a deep learning model that combines Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), and is specifically used for the image semantic encoding task.
[0039] The CNN model extracts the basic visual features of an image by performing sliding convolutional processing on the image. The image features output by the CNN model are input into the RNN model for processing, so as to utilize the sequential information processing ability of the RNN model, fully understand the global context semantic information of the image, and transform the visual content of the image into a high-level semantic feature representation, so as to obtain a set of library image semantic encoding feature vectors and the semantic encoding feature vectors of the image to be retrieved, thereby providing a data basis for subsequent semantic matching calculations.
[0040] Specifically, based on the CNN-RNN hybrid network model, the input image is first processed by the CNN. The basic visual features of the image are extracted through the convolutional layer. These features are usually local patterns in the image, such as edges, textures, etc. Through multi-level convolutional operations and pooling operations, the CNN can automatically identify important patterns in the image and form a preliminary representation of the image content.
[0041] The feature representation output by the CNN model is input into the RNN model. The RNN is good at processing sequential information and can understand the sequential relationship and context information between different parts of the image. In the context of image semantic encoding, the RNN can understand the relationship between regions in the image, thereby constructing a higher-level semantic description.
[0042] In this way, the visual content of the image is transformed into a high-level semantic feature representation, that is, a set of library image semantic encoding feature vectors and the semantic encoding feature vectors of the image to be retrieved. These feature vectors can be used for semantic matching calculations between images to achieve more accurate image retrieval.
[0043] Above, the image semantic encoder based on the CNN-RNN hybrid network model can effectively understand and encode the semantic information of the image by combining the powerful image feature extraction ability of the CNN and the sequence modeling ability of the RNN, providing a solid data basis for subsequent image retrieval.
[0044] Optionally, in an embodiment of the present application, each image in the image database and the image to be retrieved are respectively input into the image semantic encoder based on the CNN-RNN hybrid network model to obtain a set of library image semantic encoding feature vectors and the semantic encoding feature vectors of the image to be retrieved.
[0045] In this embodiment, based on the set of library image semantic encoding feature vectors, the semantic query response of the semantic encoding feature vectors of the image to be retrieved can be optimized.
[0046] It should be understood that, considering that in an image database, there may be certain semantic correlations between different images. For example, images from different perspectives of the same scene usually have similar semantic content. And this kind of correlation can provide additional information for image retrieval, thereby improving the accuracy and efficiency of retrieval.
[0047] Based on this, in an embodiment of the present invention, an image retrieval method based on the endogenous correlation in an image database is proposed. By utilizing the inherent correlation between images, the semantic expression of the to-be-retrieved image input by the user is optimized to enhance the part related to the target in its semantic encoding feature representation, thereby improving the accuracy and relevance of retrieval.
[0048] In some embodiments, the step of optimizing the semantic query response for the to-be-retrieved image semantic encoding feature vector based on the set of library image semantic encoding feature vectors obtained from the image database to obtain the optimized semantic encoding vector of the to-be-retrieved image may specifically include: Extracting the significant features of the set of library image semantic encoding feature vectors as a prompt template; based on the prompt template, performing cross-domain optimized query encoding based on the attention mechanism on the to-be-retrieved image semantic encoding feature vector and the set of library image semantic encoding feature vectors to obtain the optimized semantic encoding vector of the to-be-retrieved image.
[0049] In some embodiments, the step of extracting the significant features of the set of library image semantic encoding feature vectors as a prompt template may specifically include: Constructing a key matrix based on the set of library image semantic encoding feature vectors; extracting the maximum value of each key vector in the key matrix to obtain the key matrix significant feature vector as the prompt template.
[0050] In some embodiments, the step of constructing a key matrix based on the set of library image semantic encoding feature vectors may specifically include: Performing linear embedding encoding on each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors using a key embedding matrix to obtain a set of linearly transformed library image semantic encoding feature vectors; using the linearly transformed library image semantic encoding feature vectors as key vectors, arranging the set of linearly transformed library image semantic encoding feature vectors in a matrix to obtain the key matrix.
[0051] In some embodiments, the step of performing cross-domain optimized query encoding based on the attention mechanism on the to-be-retrieved image semantic encoding feature vector and the set of library image semantic encoding feature vectors based on the prompt template to obtain the optimized semantic encoding vector of the to-be-retrieved image may specifically include: The semantic encoding feature vectors of the image to be retrieved are linearly embedded and encoded using a query embedding matrix and a value embedding matrix respectively to obtain value vectors and query vectors; the query vectors, value vectors, each key vector in the key matrix, and the prompt template are respectively input into a heterogeneous transformer structure optimized based on template prompts to obtain a set of cross-domain optimized query encoding feature vectors of the image to be retrieved; the position-wise mean vector of the set of cross-domain optimized query encoding feature vectors of the image to be retrieved is calculated to obtain the optimized semantic encoding vector of the image to be retrieved.
[0052] In some embodiments, the step of respectively inputting the query vectors, value vectors, each key vector in the key matrix, and the prompt template into a heterogeneous transformer structure optimized based on template prompts to obtain a set of cross-domain optimized query encoding feature vectors of the image to be retrieved may specifically include: Multiplying the query vector by the transposed vector of the key vector and then dividing by the two-norm of the prompt template to obtain an attention score matrix; passing the attention score matrix through the softmax function and then multiplying by the prompt template to obtain a template prompt optimized attention weight vector; calculating the position-wise dot product between the template prompt optimized attention weight vector and the value vector to obtain the cross-domain optimized query encoding feature vector of the image to be retrieved.
[0053] In summary, in the embodiments of the present invention, the following cross-domain optimized query formula is used to process the set of semantic encoding feature vectors of the library images and the semantic encoding feature vectors of the image to be retrieved to obtain the optimized semantic encoding vector of the image to be retrieved, where the cross-domain optimized query formula is: ; ; ; ; ; ; ; ; Wherein, represents the set of semantic encoding feature vectors of the library images, , , and are respectively the first, second, th, and th semantic encoding feature vectors in the set of semantic encoding feature vectors of the library images, has a value of the number of semantic encoding feature vectors of the library images, , and respectively represent the key embedding matrix, the query embedding matrix, and the value embedding matrix, , and respectively represent different bias terms, represents the key matrix, , , and are respectively the first, second, th, and th key vectors in the key matrix, that is, the semantic encoding feature vectors of the library images after linear transformation, represents the max function, represents the prompt template, represents the semantic encoding feature vector of the image to be retrieved, and respectively represent the query vector and the value vector, represents the transpose of a vector, represents the L2 norm of a vector, is the softmax function, represents matrix multiplication operation, represents element-wise multiplication, represents the th cross-domain optimized query encoding feature vector of the image to be retrieved in the set of cross-domain optimized query encoding feature vectors of the image to be retrieved, represents the optimized semantic encoding vector of the image to be retrieved.
[0054] That is, to process the set of semantic encoding feature vectors of the library images and the semantic encoding feature vector of the image to be retrieved to obtain the optimized semantic encoding vector of the image to be retrieved. First, use the key embedding matrix to perform a linear transformation on each semantic encoding feature vector in the set of semantic encoding feature vectors of the library images to map it to a new latent feature space, making it suitable for subsequent attention mechanism calculations.
[0055] Next, take each linearly transformed semantic encoding feature vector of the library images as a key vector and arrange them into a key matrix for subsequent parallel processing. Then, select the maximum values of each key vector in the key matrix to form a key matrix significant feature vector to capture important image semantic information and use it as a prompt template, thereby guiding the subsequent semantic optimization process of the image to be retrieved to focus more on the key semantic parts of each image in the image database and reduce noise interference.
[0056] Furthermore, the query embedding matrix and the value embedding matrix are respectively applied to the semantic encoding feature vector of the image to be retrieved to generate corresponding query vectors and value vectors. Through a specially designed heterogeneous transformer structure, the query vectors, value vectors, each key vector in the key matrix, and the prompt template are interactively encoded. By introducing an additional prompt template in the traditional attention mechanism, the weight allocation strategy is automatically adjusted, so as to optimize the query response to the semantic features of the image to be retrieved, so as to generate a more accurate and relevant output of the semantic features of the image to be retrieved.
[0057] Finally, the cross-domain optimized query results of the image to be retrieved and each image in the image database are fused by position mean to integrate all the query information of the library images, and an optimized semantic encoding vector of the image to be retrieved is generated, so as to effectively emphasize the semantic features related to the target in the image to be retrieved, which helps to more accurately identify and match the images semantically similar to the image to be retrieved in the image database.
[0058] In some embodiments, the step of generating the retrieval result based on the semantic correlation between each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the image to be retrieved may specifically include: Calculate the semantic matching degree between the optimized semantic encoding vector of the image to be retrieved and each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors to obtain a sequence of semantic matching degrees; return the image corresponding to the maximum value in the sequence of semantic matching degrees as the retrieval result.
[0059] In this embodiment, it is necessary to further calculate the semantic matching degree between the optimized semantic encoding vector of the image to be retrieved and each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors to quantitatively represent the similarity between the image to be retrieved and each image in the database.
[0060] Exemplarily, the step of calculating the semantic matching degree between the optimized semantic encoding vector of the image to be retrieved and each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors to obtain a sequence of semantic matching degrees may specifically include: Calculate the cosine similarity between the optimized semantic encoding vector of the image to be retrieved and each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors as the semantic matching degree to obtain a sequence of semantic matching degrees.
[0061] Cosine Similarity is a method for measuring the angle between two non-zero vectors, used to evaluate the similarity of the directions of these two vectors. It has extensive applications in multiple fields such as information retrieval, natural language processing, and recommendation systems. The calculation of cosine similarity is based on the vector space model, which measures the similarity between two vectors by calculating the cosine value of the included angle between them. For two vectors A and B, their cosine similarity is defined as the dot product of the two vectors divided by the product of the lengths (i.e., norms) of the two vectors.
[0062] In this embodiment, cosine similarity is used to calculate the similarity between the optimized semantic encoding vector of the image to be retrieved and the semantic encoding feature vectors of each image in the image database. Specifically, it measures the semantic matching degree between them by calculating the cosine similarity between the optimized semantic encoding vector of the image to be retrieved and each semantic encoding feature vector in the set of semantic encoding feature vectors of the library images.
[0063] Among them, the obtained cosine similarity value ranges from -1 to 1. The closer the value is to 1, the more similar the two vectors are; close to 0 indicates that the two vectors are almost orthogonal; close to -1 indicates that the directions of the two vectors are almost completely opposite. Through the calculated cosine similarity sequence, the image most relevant to the image to be retrieved can be found as the retrieval result, thus achieving the purpose of intelligent retrieval.
[0064] In this embodiment, cosine similarity is used as a metric to measure the semantic matching degree between the image to be retrieved and each image in the database, and a sequence of semantic matching degrees is generated by calculation. Furthermore, by comparing and judging each semantic matching degree, the image corresponding to the maximum value in the sequence of semantic matching degrees is returned as the retrieval result.
[0065] Above, the image retrieval method based on the endogenous relevance of the image database proposed in the embodiment of the present invention adopts deep learning technology, especially combines the advantages of CNN and RNN, can better capture the semantic information of images, and greatly improves the intelligent level of image retrieval. And it can effectively overcome the problems existing in current image retrieval, such as the limitations of deep learning models in processing image semantic features and the problem of not fully utilizing the inherent connections between images. By using the CNN-RNN hybrid network model for image semantic encoding and optimizing semantic expression using the endogenous relevance in the image database, this technology can significantly improve the accuracy and efficiency of image retrieval.
[0066] In addition, the embodiment of the present invention enhances the relevance of the image retrieval results and improves the user interaction experience by introducing the cross-domain optimization query encoding technology based on the attention mechanism.
[0067] It should be added that the embodiments of the present invention can be applied to application scenarios that require quickly and accurately retrieving relevant information from a large number of images, such as image search in social media, image recognition in smart homes, visual inspection in industrial automation, intelligent album management, online shopping product recommendation, etc. By improving the accuracy of image retrieval, it helps to improve user satisfaction.
[0068] Please refer to Figure 2 , Figure 2 which is the second flowchart of the image retrieval method based on the endogenous relevance of the image database provided by the embodiments of the present invention.
[0069] In this embodiment, an image database and a to-be-retrieved image input by a user are obtained; semantic encoding is respectively performed on each image in the image database and the to-be-retrieved image to obtain a set of library image semantic encoding feature vectors and the to-be-retrieved image semantic encoding feature vector; based on the set of library image semantic encoding feature vectors, semantic query response optimization is performed on the to-be-retrieved image semantic encoding feature vector to obtain the optimized semantic encoding vector of the to-be-retrieved image; based on the semantic relevance between each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the to-be-retrieved image, a fused image can be generated as the retrieval result.
[0070] Some image retrieval technologies in the related art mainly focus on feature extraction of a single image without considering the relevance between images. Although such technologies can provide quick retrieval results in some cases, due to the lack of understanding of the overall structure of the image set, the relevance of the retrieval results may be relatively low.
[0071] The embodiments of the present invention further optimize the semantic expression of the to-be-retrieved image and enhance the accuracy of the retrieval results by utilizing the internal relevance between each image in the image database and the to-be-retrieved image. At the same time, the embodiments of the present invention also provide multiple implementation paths, such as using different combinations of neural network architectures or adjusting parameter settings in the algorithm. These flexible implementation methods can facilitate adaptation to different application scenarios and technical environments.
[0072] The present invention also provides an image retrieval device based on the endogenous relevance of the image database. The image retrieval device based on the endogenous relevance of the image database provided by the present invention will be described below. The image retrieval device based on the endogenous relevance of the image database described below can be mutually corresponding and referred to with the image retrieval method based on the endogenous relevance of the image database described above.
[0073] Please refer to Figure 3 , Figure 3It is a schematic structural diagram of an image retrieval device based on the endogenous relevance of an image database provided by an embodiment of the present invention. In this embodiment, the image retrieval device based on the endogenous relevance of the image database may include an image acquisition module 310, an image semantic encoding and processing module 320, a semantic query response optimization and processing module 330, and a retrieval result generation module 340.
[0074] The image acquisition module 310 is configured to acquire the to-be-retrieved image input by the user.
[0075] The image semantic encoding and processing module 320 is configured to perform image semantic encoding on the to-be-retrieved image to obtain a to-be-retrieved image semantic encoding feature vector.
[0076] The semantic query response optimization and processing module 330 is configured to perform semantic query response optimization on the to-be-retrieved image semantic encoding feature vector based on the set of library image semantic encoding feature vectors obtained from the image database to obtain an optimized semantic encoding vector of the to-be-retrieved image.
[0077] The retrieval result generation module 340 is configured to generate a retrieval result based on the semantic relevance between each library image semantic encoding feature vector in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the to-be-retrieved image.
[0078] In some embodiments, the image acquisition module 310 may also be configured to acquire the image database; the image semantic encoding and processing module 320 may also be configured to perform image semantic encoding on each image in the image database to obtain a set of library image semantic encoding feature vectors.
[0079] In some embodiments, the semantic query response optimization and processing module 330 may specifically be configured to: extract the significant features of the set of library image semantic encoding feature vectors as a prompt template; based on the prompt template, perform cross-domain optimization query encoding based on the attention mechanism on the to-be-retrieved image semantic encoding feature vector and the set of library image semantic encoding feature vectors to obtain an optimized semantic encoding vector of the to-be-retrieved image.
[0080] In some embodiments, the semantic query response optimization and processing module 330 may specifically be configured to: construct a key matrix based on the set of library image semantic encoding feature vectors; extract the maximum value of each key vector in the key matrix to obtain a key matrix significant feature vector as a prompt template.
[0081] In some embodiments, the semantic query response optimization processing module 330 may specifically be configured to: perform linear embedding encoding on each of the library image semantic encoding feature vectors in the set of library image semantic encoding feature vectors using a key embedding matrix to obtain a set of linearly transformed library image semantic encoding feature vectors; use the linearly transformed library image semantic encoding feature vectors as key vectors, and arrange the set of linearly transformed library image semantic encoding feature vectors in a matrix to obtain a key matrix.
[0082] In some embodiments, the semantic query response optimization processing module 330 may specifically be configured to: perform linear embedding encoding on the to-be-retrieved image semantic encoding feature vector using a query embedding matrix and a value embedding matrix respectively to obtain a value vector and a query vector; input the query vector, the value vector, each key vector in the key matrix, and a prompt template into a heterogeneous transformer structure optimized based on template prompts to obtain a set of cross-domain optimized query encoding feature vectors of the to-be-retrieved image; calculate the position-wise mean vector of the set of cross-domain optimized query encoding feature vectors of the to-be-retrieved image to obtain an optimized semantic encoding vector of the to-be-retrieved image.
[0083] On the other hand, an embodiment of the present invention further provides an electronic device. Please refer to Figure 4 , Figure 4 which is a schematic physical structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 4 shown, the electronic device may include a memory 420, a processor 410, and a computer program stored in the memory 420 and executable on the processor 410. When the processor 410 executes the program, it implements the image retrieval method based on the endogenous relevance of the image database provided by the above-mentioned various methods.
[0084] Optionally, the electronic device may further include a communication bus 430 and a communication interface 440. Among them, the processor 410, the communication interface 440, and the memory 420 complete mutual communication through the communication bus 430. The processor 410 may call the computer program in the memory 420 to execute the image retrieval method based on the endogenous relevance of the image database. The method may include: Obtaining a to-be-retrieved image input by a user; performing image semantic encoding on the to-be-retrieved image to obtain a to-be-retrieved image semantic encoding feature vector; performing semantic query response optimization on the to-be-retrieved image semantic encoding feature vector based on a set of library image semantic encoding feature vectors obtained from an image database to obtain an optimized semantic encoding vector of the to-be-retrieved image; generating a retrieval result based on the semantic relevance between each of the library image semantic encoding feature vectors in the set of library image semantic encoding feature vectors and the optimized semantic encoding vector of the to-be-retrieved image.
[0085] In addition, when the logical instructions in the above-mentioned memory 420 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0086] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image retrieval method based on the endogenous relevance of the image database provided by the above-mentioned various methods. The steps and principles have been introduced in detail in the above methods and will not be repeated here.
[0087] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the image retrieval method based on the endogenous relevance of the image database provided by the above-mentioned various methods. The steps and principles have been introduced in detail in the above methods and will not be repeated here.
[0088] The non-transitory computer-readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROM, EPROM, EEPROM, non-volatile memories (NANDFLASH), solid-state drives (SSD)).
[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image retrieval method based on the intrinsic correlation of an image database, characterized in that: include: Obtaining the image to be retrieved input by the user; Performing image semantic coding on the image to be retrieved to obtain a semantic coding feature vector of the image to be retrieved; Based on the set of library image semantic coding feature vectors obtained from the image database, the semantic query response optimization is performed on the semantic coding feature vectors of the image to be retrieved to obtain the optimized semantic coding vectors of the image to be retrieved; A search result is generated based on the semantic correlation between each library image semantic coding feature vector in the set of library image semantic coding feature vectors and the optimized semantic coding vector of the image to be searched.
2. The image retrieval method based on the intrinsic correlation of the image database according to claim 1 is characterized in that: Before obtaining the image to be retrieved input by the user, the method further includes: Acquire the image database; Perform image semantic coding on each image in the image database to obtain a set of semantic coding feature vectors of the library images.
3. The image retrieval method based on the intrinsic correlation of the image database according to claim 1 is characterized in that: The method of optimizing the semantic query response of the semantic coding feature vector of the image to be retrieved based on the set of library image semantic coding feature vectors obtained from the image database to obtain the optimized semantic coding vector of the image to be retrieved includes: Extracting salient features of the set of semantically encoded feature vectors of the library image as a prompt template; Based on the prompt template, a set of the semantic coding feature vector of the image to be retrieved and the semantic coding feature vector of the library image is subjected to cross-domain optimized query encoding based on an attention mechanism to obtain the optimized semantic coding vector of the image to be retrieved.
4. The image retrieval method based on the intrinsic correlation of the image database according to claim 3 is characterized in that: The extracting the significant features of the set of semantic encoding feature vectors of the library image as a prompt template includes: constructing a key matrix based on a set of semantically encoded feature vectors of the library images; The maximum value of each key vector in the key matrix is extracted to obtain a key matrix significant feature vector as the prompt template.
5. The image retrieval method based on the intrinsic correlation of the image database according to claim 4 is characterized in that: The constructing a key matrix based on a set of library image semantic encoding feature vectors includes: Using a key embedding matrix, linear embedding coding is performed on each of the set of library image semantic coding feature vectors to obtain a set of library image semantic coding feature vectors after linear transformation; The semantic coding feature vector of the library image after the linear transformation is used as the key vector, and the set of the semantic coding feature vectors of the library image after the linear transformation is arranged in a matrix to obtain the key matrix.
6. The image retrieval method based on the intrinsic correlation of the image database according to claim 4 is characterized in that: The step of performing cross-domain optimized query encoding based on an attention mechanism on a set of the semantic encoding feature vector of the image to be retrieved and the semantic encoding feature vector of the library image based on the prompt template to obtain the optimized semantic encoding vector of the image to be retrieved includes: Using the query embedding matrix and the value embedding matrix, linear embedding coding is performed on the semantic coding feature vector of the image to be retrieved, so as to obtain a value vector and a query vector; Inputting the query vector, the value vector, each key vector in the key matrix and the prompt template into a heterogeneous transformer structure based on template prompt optimization respectively, so as to obtain a set of cross-domain optimized query encoding feature vectors of the image to be retrieved; The position-based mean vector of the set of cross-domain optimized query encoding feature vectors of the image to be retrieved is calculated to obtain the optimized semantic encoding vector of the image to be retrieved.
7. An image retrieval device based on the intrinsic correlation of an image database, characterized in that: include: An image acquisition module is used to acquire an image to be retrieved input by a user; An image semantic coding processing module is used to perform image semantic coding on the image to be retrieved to obtain a semantic coding feature vector of the image to be retrieved; A semantic query response optimization processing module is used to optimize the semantic query response of the semantic coding feature vector of the image to be retrieved based on the set of library image semantic coding feature vectors obtained from the image database to obtain an optimized semantic coding vector of the image to be retrieved; The retrieval result generating module is used to generate a retrieval result based on the semantic correlation between each library image semantic coding feature vector in the set of library image semantic coding feature vectors and the optimized semantic coding vector of the image to be retrieved.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the image retrieval method based on the intrinsic correlation of the image database as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image retrieval method based on the intrinsic correlation of the image database as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the image retrieval method based on the intrinsic correlation of the image database as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Plastic injection molding and recycling integrated equipment and method thereof
CN120481176A