Picture retrieval method and device and electronic equipment

The pre-trained style feature extraction model generates style feature vectors, which solves the problems of inaccurate and high cost of image search in the prior art, and achieves efficient retrieval of similar styles.

CN120296190APending Publication Date: 2025-07-11NETEASE (SHANGHAI) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166871.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing image search methods based on labels or semantic embeddings are prone to inaccurate search results and increase labor and maintenance costs, especially when the number of defined tags is too small or the description information is not detailed.

Method used

Through pre-trained style feature extraction model, the style information of the search information is extracted, the style feature vector is generated, and similar pictures are searched from the vector database to reduce dependence on image tags and text information.

Benefits of technology

Improve the accuracy and effectiveness of image retrieval and reduce labor and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296190A_ABST
    Figure CN120296190A_ABST
Patent Text Reader

Abstract

The invention provides a picture retrieval method and device and electronic equipment, and the method comprises the steps: inputting retrieval information into a pre-trained feature extraction model, and obtaining an initial feature vector corresponding to the retrieval information; inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; and determining a target vector from a vector database according to the style feature vector corresponding to the retrieval information, and determining a picture corresponding to the target vector as a picture corresponding to the retrieval information. In the mode, the picture style information in the retrieval information is extracted through the pre-trained style feature extraction model, the style feature vector of the retrieval information is obtained, the picture similar to the style feature vector of the retrieval information is searched from the vector database, excessive picture label information and text information are not needed, and the retrieval efficiency is improved. Therefore, the picture with the similar painting style can be accurately retrieved, the retrieval precision and the retrieval effect are improved, and the labor cost and the maintenance cost are further reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to an image retrieval method, apparatus, and electronic device. Background Art

[0002] Currently, image search is usually based on tag or semantic embedding retrieval methods. In this retrieval method, tags and semantic feature vectors need to be defined to retrieve corresponding images. However, if the number of defined tags is too small or the description information of the images is not detailed enough, problems such as inaccurate search results and too few supported painting styles are likely to occur, affecting the retrieval effect; if the number of defined tags is too large or the description information of the images is too detailed, it will greatly increase the labor cost and maintenance cost. Summary of the Invention

[0003] In view of this, the purpose of the present disclosure is to provide an image retrieval method, apparatus, and electronic device. Through a pre-trained style feature extraction model, the painting style information in the retrieval information can be extracted to obtain the style feature vector of the retrieval information, and then images similar to the style feature vector of the retrieval information can be searched from the vector database. Without excessive image tag information and text information, images with similar painting styles can be accurately retrieved to improve the retrieval accuracy and effect, and thus reduce the labor cost and maintenance cost.

[0004] In a first aspect, an embodiment of the present disclosure provides an image retrieval method, which includes: inputting retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; determining a target vector from the vector database according to the style feature vector corresponding to the retrieval information, and determining the image corresponding to the target vector as the image corresponding to the retrieval information; the vector database includes multiple style feature vectors and the images corresponding to the style feature vectors.

[0005] In a second aspect, an embodiment of the present disclosure provides an image retrieval apparatus, which includes: a feature extraction module for inputting retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; a style feature extraction sub-model for inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; an image retrieval module for determining a target vector from the vector database according to the style feature vector corresponding to the retrieval information, and determining the image corresponding to the target vector as the image corresponding to the retrieval information; the vector database includes multiple style feature vectors and the images corresponding to the style feature vectors.

[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the picture retrieval method according to any one of the first aspect.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the picture retrieval method according to any one of the first aspect.

[0008] The embodiments of the present disclosure bring the following beneficial effects:

[0009] The present disclosure provides a picture retrieval method, apparatus and electronic device. The retrieval information is input into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; the initial feature vector corresponding to the retrieval information is input into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; the target vector is determined from the vector database according to the style feature vector corresponding to the retrieval information, and the picture corresponding to the target vector is determined as the picture corresponding to the retrieval information; the vector database includes a plurality of style feature vectors and the pictures corresponding to the style feature vectors. In this way, the pre-trained style feature extraction model can extract the painting style information in the retrieval information to obtain the style feature vector of the retrieval information, and then search for pictures similar to the style feature vector of the retrieval information from the vector database. Without excessive picture label information and text information, pictures with similar painting styles can be accurately retrieved, improving the retrieval accuracy and retrieval effect, and thus reducing the labor cost and maintenance cost.

[0010] Other features and advantages of the present disclosure will be described in the following specification, and will, in part, be obvious from the specification, or be learned by practicing the present disclosure. The objectives and other advantages of the present disclosure are realized and obtained by the structures particularly pointed out in the specification, claims and drawings.

[0011] To make the above objectives, features and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 A flowchart of a picture retrieval method provided by an embodiment of the present disclosure;

[0014] Figure 2 A schematic flow diagram of a picture retrieval method provided by an embodiment of the present disclosure;

[0015] Figure 3 A schematic structural diagram of a picture retrieval device provided by an embodiment of the present disclosure;

[0016] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specific embodiments

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions of the present disclosure with reference to the drawings. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.

[0018] Picture retrieval refers to a method of analyzing or identifying pictures and searching for pictures according to certain specific conditions. Common picture retrievals include retrieving pictures by text, by tags, by pictures, etc. Generally, there are two ways to implement picture retrieval. One is to attach meta-information such as tags and descriptions to pictures and then perform picture retrieval through this meta-information; the other is to convert pictures and query information into vectors of a fixed length through a certain algorithm and then retrieve pictures according to the similarity of the vectors. When the number of pictures involved in picture retrieval is large, tools such as databases need to be used to provide functions such as adding, modifying, deleting, and querying picture data to improve the retrieval efficiency.

[0019] The existing technical solutions are mainly divided into two types: the tag-based retrieval method and the semantic embedding-based retrieval method. Specifically, for the tag-based retrieval method, a series of semantic tags (such as black and white, comics, Chinese style, etc.) are first determined, and then each picture is added with its corresponding tag description. To simplify the problem, it can be considered that different tags are independent of each other. That is to say, the process of adding picture tags can be regarded as a multi-label classification problem. During retrieval, the query condition is converted into corresponding tags, and then these tags are used to retrieve the corresponding pictures.

[0020] Specifically, for the semantic embedding-based retrieval method, the picture and the retrieval condition are converted into fixed-length semantic feature vectors through a specific embedding algorithm. This kind of semantic feature vector contains the semantic information of the input data. For different types of data (such as text, image, tag, etc.), specific algorithms can be used to embed them into the same semantic space, so that the similarity of the obtained semantic feature vectors can represent the semantic similarity between the data. Currently, the mainstream semantic embedding method is to use deep learning algorithms, and different embedding models are used for different types of data and embedded into a unified semantic space.

[0021] However, if the number of defined tags is too small or the description information of the pictures is not detailed enough, problems such as inaccurate search results and too few supported painting styles are likely to occur, affecting the retrieval effect. If the number of defined tags is too large or the description information of the pictures is too detailed, it will greatly increase the labor cost and maintenance cost. Based on this, an image retrieval method, device, and electronic device provided by the embodiments of the present disclosure can be applied to electronic devices such as mobile phones, computers, notebooks, and servers, and especially can be applied to devices configured with an application program having an image retrieval function.

[0022] To facilitate the understanding of this embodiment, first, a detailed introduction to an image retrieval method disclosed by the embodiments of the present disclosure is provided. As Figure 1 shown, the method includes the following steps:

[0023] Step S102, input the retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information;

[0024] The above retrieval information is usually the information submitted by the user or the information selected by the user. The above retrieval information can be image information and text information, or it can be image information only, or text information only; the image information among them can be an image including tags or an image without tags, and the text information is usually a written description of the image. For example, the retrieval information includes image information and text information. The image information can be an animal in anime style, and the text information can be "a cat in anime style"; for example, the retrieval information includes text information, and the text information can be "a sketch of a game character"; for example, the retrieval information includes image information, and the image information can be a sketch of a game character's home, or an anime-style picture, etc.

[0025] The above feature extraction model usually includes an image feature extraction sub-model and a text feature extraction sub-model. The image feature extraction sub-model and the text feature extraction sub-model among them can be open-source pre-trained models, such as deep learning image models and deep learning text models. After training with specified training data, the above image feature extraction sub-model and text feature extraction sub-model can be obtained.

[0026] Optionally, if the retrieval information is image information and text information, input the image information into the pre-trained image feature extraction sub-model to obtain the image feature vector corresponding to the image information, input the text information into the pre-trained text feature extraction sub-model to obtain the text feature vector corresponding to the text information, and use the image feature vector and the text feature vector as the initial feature vectors.

[0027] Optionally, if the retrieval information is image information and text information, input the image information into the pre-trained image feature extraction sub-model to obtain the image feature vector corresponding to the image information, input the text information into the pre-trained text feature extraction sub-model to obtain the text feature vector corresponding to the text information, determine the image feature vector similar to the text feature vector to obtain the image feature vector corresponding to the text information, and use the image feature vector corresponding to the image information and the image feature vector corresponding to the text information as the initial feature vectors.

[0028] Optionally, if the retrieval information is text information, input the text information into the pre-trained text feature extraction sub-model to obtain the text feature vector corresponding to the text information, and determine the image feature vector similar to the text feature vector to obtain the image feature vector corresponding to the text information.

[0029] Optionally, if the retrieval information is image information, input the image information into the pre-trained text feature extraction sub-model to obtain the image feature vector corresponding to the image information.

[0030] The main function of the above-mentioned feature extraction module is to convert the retrieval information (image information and / or text information) input by the user into a feature vector of a fixed length. For image information, the obtained image feature vector after conversion removes data redundancy and can more efficiently represent the main feature information contained in the image information; for text information, the converted text feature vector can be aligned with the image feature vector in the same feature space, and the semantics of the text and the semantics of the image are correlated with each other through the similarity of the feature vectors.

[0031] Step S104: Input the initial feature vector corresponding to the retrieval information into the pre-trained style feature extraction model to obtain the style feature vector corresponding to the retrieval information.

[0032] Optionally, input the image feature vector corresponding to the image information into the pre-trained style feature extraction model to obtain the style feature vector corresponding to the image information; input the text feature vector corresponding to the text information into the pre-trained style feature extraction model to obtain the style feature vector corresponding to the text information.

[0033] Optionally, the above-mentioned style feature extraction model takes the feature vector of text or picture as input, and calculates the style feature vector through methods such as feature mapping, channel similarity calculation, feature dimensionality reduction, and deep learning model, etc., for calculating the style similarity between pictures and between pictures and texts.

[0034] The above-mentioned style feature extraction model is used to extract the style feature vector related to the painting style of the image from the initial feature vector, with the aim of searching for pictures with a painting style similar to the retrieval information.

[0035] Step S106: Determine the target vector from the vector database according to the style feature vector corresponding to the retrieval information, and determine the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes multiple style feature vectors and the pictures corresponding to the style feature vectors.

[0036] The style feature vectors and the pictures corresponding to the style feature vectors in the above-mentioned vector database are pre-generated according to the sample images. The sample images can be an open-source image sample set or the image information input by the user, etc. Usually, the style feature vectors are used as indexes in the vector database, and the pictures corresponding to the style feature vectors are the index values.

[0037] Optionally, compare the style feature vector corresponding to the retrieval information with the style feature vectors in the vector database, determine the target vector similar to the style feature vector corresponding to the retrieval information, and then determine the picture corresponding to the target vector as the picture corresponding to the retrieval information.

[0038] This embodiment provides an image retrieval method. The retrieval information is input into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information. The retrieval information is image information and / or text information. The initial feature vector corresponding to the retrieval information is input into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information. The target vector is determined from the vector database according to the style feature vector corresponding to the retrieval information, and the image corresponding to the target vector is determined as the image corresponding to the retrieval information. The vector database includes multiple style feature vectors and the images corresponding to the style feature vectors. In this way, the pre-trained style feature extraction model can extract the painting style information in the retrieval information to obtain the style feature vector of the retrieval information, and then search for images similar to the style feature vector of the retrieval information from the vector database. Without too much image label information and text information, images with similar painting styles can be accurately retrieved, improving the retrieval accuracy and retrieval effect, and thus reducing the labor cost and maintenance cost.

[0039] For the step of inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information, a possible implementation: The initial feature vector corresponding to the retrieval information is input into a style feature extraction sub-model to obtain an original style feature vector corresponding to the retrieval information. The original style feature vector corresponding to the retrieval information is input into a dimensionality reduction sub-model to perform dimensionality reduction processing on the original style feature vector to obtain a style feature vector corresponding to the retrieval information.

[0040] Generally, the original style feature vector may have a large amount of data, which is difficult to be directly used as an index for retrieval in the vector database. At the same time, due to the large amount of data, the computational time cost is relatively high, and it is difficult to calculate the similarity between vectors. Therefore, in this embodiment, the original style feature vector is input into the dimensionality reduction sub-model to select appropriate features from the original style feature vector for dimensionality reduction to obtain the final style feature vector. The above-mentioned dimensionality reduction sub-model is used to select appropriate features from the style feature vector for dimensionality reduction to improve the retrieval efficiency of images.

[0041] In the above method, the dimensionality reduction sub-model can reduce the amount of data of the original style feature vector, eliminate most of the data irrelevant to the painting style, reduce the amount of calculation, and at the same time increase the proportion of painting style information in the style feature vector, improving the data processing speed of the model and the retrieval efficiency of images.

[0042] For the step of determining the target vector from the vector database according to the style feature vector corresponding to the retrieval information, a possible implementation manner: For each style feature vector in the vector database, calculate the similarity value between the style feature vector and the style feature vector corresponding to the retrieval information, and determine the style feature vector whose similarity value meets the preset threshold as the target vector.

[0043] The above-mentioned preset threshold is usually set in advance according to needs. Optionally, similarity values can be calculated by methods such as the cosine similarity algorithm, Manhattan distance, and Jaccard similarity coefficient. For example, through the pre-similarity algorithm, the cosine value of the included angle between two vectors is calculated to measure the similarity between vectors. The closer the cosine value is to 1, the higher the similarity between the two vectors; conversely, the closer the cosine value is to -1, the lower the similarity; when the cosine value is 0, it means the two vectors are orthogonal and irrelevant. In this embodiment, this method can quickly find the vector with a style similar to the retrieval information from the style feature vectors because it does not depend on the length of the vectors and only focuses on the direction, which is more effective for measuring the similarity of such feature directions as image styles.

[0044] Another example is through the Manhattan distance method, which calculates the sum of the absolute values of the differences between two vectors in each dimension. In this embodiment, combined with the deep learning model, the Manhattan distance can measure the vector differences in multiple dimensions (such as line thickness, color, etc.) in detail, and combine the distances in multiple dimensions to more accurately find the picture with a style similar to the retrieval information.

[0045] The above-mentioned feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; for the step of inputting the retrieval information into the pre-trained feature extraction model to obtain the initial feature vector corresponding to the retrieval information, a possible implementation manner:

[0046] If the retrieval information includes image information, input the image information into the image feature extraction sub-model to obtain the initial feature vector corresponding to the image information;

[0047] If the retrieval information includes text information, input the text information into the text feature extraction sub-model to obtain the text feature vector corresponding to the text information, and determine the image feature vector similar to the text feature vector to obtain the initial feature vector corresponding to the text information.

[0048] That is to say, if the retrieved information includes image information and text information, two image feature vectors will finally be obtained. Since the information included in the image feature vector usually is more than that included in the text feature vector, generally, the image feature vector includes the content of the image, such as image color information, person information, landscape information, style information, etc. However, the text feature vector does not include the image content. Therefore, determining the image feature vector similar to the text feature vector in this embodiment can further improve the accuracy and reliability of image retrieval.

[0049] After the step of inputting the image information into the image feature extraction sub-model to obtain the initial feature vector corresponding to the image information if the retrieved information includes image information, the above method further includes: converting the image feature vector corresponding to the image information into a feature vector with a specified length and embedding it into a specified semantic feature space.

[0050] The purpose of this embodiment is to convert pictures or texts with unfixed sizes into feature vectors with generalized semantics.

[0051] For the above step of determining the image feature vector similar to the text feature vector, a possible implementation manner: converting the text feature vector corresponding to the text information into a feature vector with a specified length and embedding it into a specified semantic feature space; determining the image feature vector similar to the text feature vector corresponding to the text information through the position information of the vectors in the semantic feature space.

[0052] The above feature extraction model can learn the correlation features of images and texts, embed the image feature vector and the text feature vector into a specified semantic feature space to facilitate the association of images and texts with the same meaning. Furthermore, it can determine the image feature vector similar to the text feature vector corresponding to the text information. The user can input the text information to determine the similar image feature vector, and then retrieve more accurate and reliable pictures.

[0053] The above method further includes: if the retrieved information includes image information, saving the image information and the style feature vector corresponding to the image information to the vector database.

[0054] To improve the richness of the style feature vectors and the pictures corresponding to the style feature vectors in the vector database, after the user inputs the retrieved information, if the retrieved information includes image information, the image information and the style feature vector corresponding to the image information can be directly saved to the vector database. As the model is used for a longer time, the vectors and pictures in the vector database will become richer, and thus the retrieval effect, accuracy, usability, and reliability of the pictures will be higher.

[0055] The above method further includes: obtaining a sample image, inputting the sample image into a pre-trained feature extraction model, and obtaining an image feature vector of the sample image through an image feature extraction sub-model in the feature extraction model; inputting the image feature vector into a pre-trained style feature extraction model to obtain a style feature vector of the sample image, and saving the sample image and the style feature vector of the sample image into a vector database.

[0056] The above sample image can be an image related to the project, or an open-source sample set, or various images collected by developers, etc. The above sample image can be a labeled image or an unlabeled image.

[0057] In actual implementation, after developers have trained all the models, obtain the retrieval images with labels of the project, and convert the images into style feature vectors through the image feature extraction sub-model and the style feature extraction model, and store them in the vector database.

[0058] The training process of the model is described below. The above feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the above method further includes: obtaining first training data; where the first training data includes first picture data and first text data; training the image feature extraction sub-model through the first picture data to obtain a trained image feature extraction sub-model; training the text feature extraction sub-model through the first text data to obtain a trained text feature extraction sub-model.

[0059] The above first training data can be an open-source data set or the project data of the project. The project refers to the project to which the model is to be applied. For example, if the model is mainly applied to the field of image retrieval of art paintings, the project data therein is images and texts related to art paintings. Another example is that if the model is mainly applied to the field of security monitoring image retrieval, the project data therein is images and texts related to security monitoring. Another example is that if the model is mainly applied to the field of game character painting retrieval, the project data therein is images and texts related to game character painting.

[0060] Optionally, for each first training data, input the first picture data into the image feature extraction sub-model, obtain an image feature vector of the first picture data through the forward propagation algorithm, calculate the loss value according to the loss function, and then calculate the gradient through the backward propagation algorithm to update the model parameters. Similarly, input the first text data into the text feature extraction sub-model, obtain a text feature vector of the first text data through the forward propagation algorithm, calculate the loss value according to the loss function, and then calculate the gradient through the backward propagation algorithm to update the model parameters. Until the termination condition is reached, a trained image feature extraction sub-model and a trained text feature extraction sub-model are obtained. The termination condition can be reaching a specified number of iterations, and the loss value satisfies a preset loss value, etc.

[0061] Optionally, use the masked image modeling method to train the image feature extraction sub-model with the first picture data. If the first picture data has no labels, divide the image into several image patches and randomly mask a certain proportion of the image patches; transmit all the image patches to the image feature extraction sub-model to be trained and add a restoration module; use the restoration of the masked image patches as a template to train the image feature extraction sub-model and the restoration module simultaneously, and finally obtain the trained image feature extraction sub-model.

[0062] Optionally, use methods such as contrastive learning to train the feature extraction model with the first picture data and the first text data. Calculate the similarity / distance between the extracted image feature vectors and text feature vectors to make the features corresponding to relevant pictures and texts closer and the features corresponding to irrelevant pictures and texts farther away; the picture and text information generated by the project may be irrelevant, and it is necessary to use CLIP (Contrastive Language-Image Pretraining) to calculate similarity for screening, or use manual verification methods for screening.

[0063] The above first text data is the data label of the first picture data; the steps of training the image feature extraction sub-model with the first picture data to obtain the trained image feature extraction sub-model and training the text feature extraction sub-model with the first text data to obtain the trained text feature extraction sub-model have another possible implementation:

[0064] For each first training data, input the first picture data into the image feature extraction sub-module to obtain the image feature vector corresponding to the first picture data, and embed the image feature vector into the specified semantic feature space;

[0065] Input the first text data into the text feature extraction sub-module to obtain the text feature vector corresponding to the first text data, and embed the text feature vector into the specified semantic feature space. Determine the semantic similarity value between the first picture data and the first text data according to the positional relationship between the vectors in the specified semantic feature space; update the model parameters of the image feature extraction sub-module and the text feature extraction sub-module according to the semantic similarity value until the preset termination condition is reached, and obtain the trained image feature extraction sub-model and the trained text feature extraction sub-model.

[0066] Generally, the closer the positions of the vectors are, the more similar the two vectors are. The above preset termination condition can be that the semantic similarity values are all greater than the preset threshold, or that all the semantic similarity values remain unchanged, or that a specified number of iterations are reached, etc.

[0067] The above-mentioned style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the above-mentioned method further includes: obtaining second training data; wherein, the second training data includes second picture data and / or second text data, inputting the second training data into a pre-trained feature extraction model to obtain a feature vector corresponding to the second training data; training the style feature extraction sub-model with the feature vector corresponding to the second training data to obtain a trained style feature extraction sub-model.

[0068] The above-mentioned second training data can be an open-source data set or data related to the project. The second training data can be data with annotations or data without annotations, such as pictures annotated by the author, etc.

[0069] Specifically, input the feature vector corresponding to the second training data into the style feature extraction sub-model, calculate the prediction result of the feature vector through forward propagation, calculate the loss value according to the loss function, and then calculate the gradient through the backpropagation algorithm to update the model parameters. Until the termination condition is reached, a trained style feature extraction sub-model is obtained. The termination condition can be reaching a specified number of iterations, the loss value satisfying a preset loss value, etc.

[0070] The above-mentioned method further includes: inputting the feature vector corresponding to the second training data into the trained style feature extraction sub-model to obtain the original style feature vector corresponding to the second training data; performing data analysis on the original style feature vector corresponding to the second training data to determine the model parameters of the dimensionality reduction sub-model to obtain a trained dimensionality reduction sub-model.

[0071] Optionally, perform data analysis on the original style feature vector corresponding to the second training data through a specified method to determine the parameters used by the specified method, and determine the specified method and the parameters as the dimensionality reduction sub-model. The specified method can be a variance comparison method or a principal component analysis method.

[0072] For example, in the variance comparison method, after extracting the original style feature vector of the second training data, according to the similarity degree of this batch of data, select the largest or smallest batch of feature components, thereby reducing the data volume of the style feature vector. Another example is the principal component analysis method. Extract the original style feature vector of the second training data, and then use the principal component analysis algorithm on this batch of original style feature vectors for feature dimensionality reduction, thereby reducing the data volume of the style feature vector.

[0073] The above method further includes: determining third training data from the retrieved information and the picture corresponding to the retrieved information; the third training data includes third picture data and / or third text data; training the pre-trained image feature extraction sub-model with the third picture data to update the model parameters of the image feature extraction sub-model; training the pre-trained text feature extraction sub-model with the third text data to update the model parameters of the text feature extraction sub-model.

[0074] After the model is launched for a period of time, the model will determine a large amount of project data. Therefore, the third training data can be determined from the retrieved information and the picture corresponding to the retrieved information. If the third training data is the third picture data, the third picture data is used to train the image feature extraction sub-model, and then the model parameters of the image feature extraction sub-model are updated iteratively. If the third training data is the third text data, the third text data is used to train the text feature extraction sub-model, and then the model parameters of the text feature extraction sub-model are updated iteratively.

[0075] In the above case, if the third training data includes the third picture data and the third text data, the third text data is the data label of the third picture data. The third picture data and the third text data are used to train the image feature extraction sub-model and the text feature extraction sub-model at the same time, and then the model parameters of the image feature extraction sub-model and the text feature extraction sub-model are updated iteratively.

[0076] In the above manner, by using the data related to the project to train the feature extraction model, the ability of the model to represent the feature information of the project data is improved.

[0077] The above method further includes: determining fourth training data from the retrieved information and the picture corresponding to the retrieved information; the fourth training data includes fourth picture data and fourth text data, and the fourth text data is the data label of the fourth picture data; inputting the fourth picture data into the pre-trained image feature extraction sub-model to obtain the image feature vector corresponding to the fourth picture data; training the pre-trained style feature extraction sub-model with the image feature vector corresponding to the fourth picture data and the fourth text data to update the model parameters of the style feature extraction sub-model, and obtaining the original style feature vector corresponding to the fourth picture data; training the pre-trained dimensionality reduction sub-model with the original style feature vector corresponding to the fourth picture data to update the model parameters of the dimensionality reduction sub-model.

[0078] In the initial stage of the project, the amount of data is small. The style feature vector obtained by the style feature extraction model using the second training data may not accurately represent the painting style or may have a poor effect on newly emerging painting styles.

[0079] After the project has been running for a certain period of time, the data generated in the project (i.e., the above-mentioned fourth training data) can be used to optimize and iterate the style feature extraction model. For example, the annotation information such as author - picture can be used, and all the pictures uploaded by the same author are regarded as the same painting style to train the style feature extraction sub - model.

[0080] In the above - mentioned method, using the fourth training data can more accurately extract and screen out the feature vectors representing painting style information.

[0081] It should be noted that the above - mentioned fourth training data usually requires manual inspection and label verification to make the finally determined fourth training data more accurate.

[0082] See Figure 2 the flow schematic diagram of the picture retrieval method shown. Input the picture data into the image feature representation model (i.e., the image feature extraction sub - model) to obtain the image feature vector, input the text data into the text feature representation model (i.e., the text feature extraction sub - model) to obtain the text feature vector, then determine the image feature vector similar to the text feature vector in the text feature vector, and finally input the image feature vector into the style feature representation module (i.e., the style feature extraction model) for style feature extraction and feature dimension reduction, and finally obtain the style feature vector. Compare the style feature vector with the vectors in the vector database, and finally obtain the target vector, and output the picture corresponding to the target vector.

[0083] Corresponding to the above - mentioned method embodiments, the present disclosure provides a picture retrieval device, as Figure 3 shown. The device includes:

[0084] A feature extraction module 301, configured to input the retrieval information into a pre - trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information;

[0085] A style feature extraction module 302, configured to input the initial feature vector corresponding to the retrieval information into a pre - trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information;

[0086] A picture retrieval module 303, configured to determine a target vector from the vector database according to the style feature vector corresponding to the retrieval information, and determine the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes a plurality of style feature vectors and the pictures corresponding to the style feature vectors.

[0087] An embodiment of the present disclosure provides an image retrieval device, which inputs retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; the initial feature vector corresponding to the retrieval information is input into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; a target vector is determined from a vector database according to the style feature vector corresponding to the retrieval information, and the image corresponding to the target vector is determined as the image corresponding to the retrieval information; the vector database includes a plurality of style feature vectors and the images corresponding to the style feature vectors. In this way, the pre-trained style feature extraction model can extract the painting style information in the retrieval information to obtain the style feature vector of the retrieval information, and then search for images similar to the style feature vector of the retrieval information from the vector database. Without too much image label information and text information, images with similar painting styles can be accurately retrieved, improving the retrieval accuracy and retrieval effect, and thus reducing the labor cost and maintenance cost.

[0088] The above-mentioned pre-trained style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the above-mentioned style feature extraction module is further configured to: input the initial feature vector corresponding to the retrieval information into the style feature extraction sub-model to obtain the original style feature vector corresponding to the retrieval information; input the original style feature vector corresponding to the retrieval information into the dimensionality reduction sub-model to perform dimensionality reduction processing on the original style feature vector to obtain the style feature vector corresponding to the retrieval information.

[0089] The above-mentioned feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the above-mentioned feature extraction module is further configured to: if the retrieval information includes image information, input the image information into the image feature extraction sub-model to obtain the initial feature vector corresponding to the image information; if the retrieval information includes text information, input the text information into the text feature extraction sub-model to obtain the text feature vector corresponding to the text information, and determine the image feature vector similar to the text feature vector to obtain the initial feature vector corresponding to the text information.

[0090] The above-mentioned feature extraction module is further configured to: convert the image feature vector corresponding to the image information into a feature vector with a specified length and embed it into a specified semantic feature space.

[0091] The above-mentioned feature extraction module is further configured to: convert the text feature vector corresponding to the text information into a feature vector with a specified length and embed it into a specified semantic feature space; determine the image feature vector similar to the text feature vector corresponding to the text information through the position information of the vectors in the semantic feature space.

[0092] The above-mentioned image retrieval module is further configured to: for each style feature vector in the vector database, calculate the similarity value between the style feature vector and the style feature vector corresponding to the retrieval information, and determine the style feature vector whose similarity value meets the preset threshold as the target vector.

[0093] The above-mentioned device further includes a database update module, configured to: if the retrieval information includes image information, save the image information and the style feature vector corresponding to the image information to the vector database.

[0094] The above-mentioned device further includes a database determination module, configured to: obtain a sample image, input the sample image into a pre-trained feature extraction model, and obtain the image feature vector of the sample image through the image feature extraction sub-model in the feature extraction model; input the image feature vector into a pre-trained style feature extraction model to obtain the style feature vector of the sample image, and save the sample image and the style feature vector of the sample image to the vector database.

[0095] The above-mentioned feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the above-mentioned device further includes a feature extraction model training module, configured to: obtain first training data; wherein, the first training data includes first picture data and first text data; train the image feature extraction sub-model through the first picture data to obtain a trained image feature extraction sub-model; train the text feature extraction sub-model through the first text data to obtain a trained text feature extraction sub-model.

[0096] The above-mentioned first text data is the data label of the first picture data; the above-mentioned feature extraction model training module is further configured to: for each first training data, input the first picture data into the image feature extraction sub-module to obtain the image feature vector corresponding to the first picture data, and embed the image feature vector into a specified semantic feature space; input the first text data into the text feature extraction sub-module to obtain the text feature vector corresponding to the first text data, and embed the text feature vector into the specified semantic feature space, and determine the semantic similarity value between the first picture data and the first text data according to the positional relationship between the vectors in the specified semantic feature space; update the model parameters of the image feature extraction sub-module and the text feature extraction sub-module according to the semantic similarity value until a preset termination condition is reached, to obtain a trained image feature extraction sub-model and a trained text feature extraction sub-model.

[0097] The above-mentioned style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the above-mentioned device further includes a style feature extraction model training module, which is used for: obtaining second training data; wherein, the second training data includes second picture data or second text data, inputting the second training data into a pre-trained feature extraction model to obtain a feature vector corresponding to the second training data; training the style feature extraction sub-model through the feature vector corresponding to the second training data to obtain a trained style feature extraction sub-model.

[0098] The above-mentioned device further includes a dimensionality reduction sub-model training module: inputting the feature vector corresponding to the second training data into the trained style feature extraction sub-model to obtain an original style feature vector corresponding to the second training data; performing data analysis on the original style feature vector corresponding to the second training data to determine the model parameters of the dimensionality reduction sub-model, and obtaining a trained dimensionality reduction sub-model.

[0099] The above-mentioned device further includes a first update module, which is further used for: determining third training data from the retrieval information and the pictures corresponding to the retrieval information; the third training data includes third picture data and / or third text data; iteratively training the pre-trained image feature extraction sub-model through the third picture data to update the model parameters of the image feature extraction sub-model; training the pre-trained text feature extraction sub-model through the third text data to update the model parameters of the text feature extraction sub-model.

[0100] If the above-mentioned third training data includes third picture data and third text data, the third text data is the data label of the third picture data.

[0101] The above-mentioned device further includes a second update module, which is further used for: determining fourth training data from the retrieval information and the pictures corresponding to the retrieval information; the fourth training data includes fourth picture data and fourth text data, and the fourth text data is the data label of the fourth picture data; inputting the fourth picture data into the pre-trained image feature extraction sub-model to obtain an image feature vector corresponding to the fourth picture data; training the pre-trained style feature extraction sub-model through the image feature vector corresponding to the fourth picture data and the fourth text data to update the model parameters of the style feature extraction sub-model, and obtaining an original style feature vector corresponding to the fourth picture data; training the pre-trained dimensionality reduction sub-model through the original style feature vector corresponding to the fourth picture data to update the model parameters of the dimensionality reduction sub-model.

[0102] The picture retrieval device provided by the embodiments of the present disclosure has the same technical features as the picture retrieval method provided by the above-mentioned embodiments, so it can also solve the same technical problems and achieve the same technical effects.

[0103] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned picture retrieval method. The electronic device can be a server or a terminal device.

[0104] As shown in Figure 4 , the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above-mentioned picture retrieval method.

[0105] Furthermore, Figure 4 the electronic device shown in also includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0106] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only a bidirectional arrow is used in, but it does not mean that there is only one bus or one type of bus.

[0107] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step and logic block diagram disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0108] The processor in the above electronic device can implement the following operations in the above picture retrieval method by executing machine-executable instructions:

[0109] Input the retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; input the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; determine a target vector from the vector database according to the style feature vector corresponding to the retrieval information, and determine the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes a plurality of style feature vectors and the pictures corresponding to the style feature vectors.

[0110] The above-mentioned pre-trained style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the step of inputting the initial feature vector corresponding to the retrieval information into the pre-trained style feature extraction model to obtain the style feature vector corresponding to the retrieval information includes: inputting the initial feature vector corresponding to the retrieval information into the style feature extraction sub-model to obtain the original style feature vector corresponding to the retrieval information; inputting the original style feature vector corresponding to the retrieval information into the dimensionality reduction sub-model to perform dimensionality reduction processing on the original style feature vector to obtain the style feature vector corresponding to the retrieval information.

[0111] The above-mentioned feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the step of inputting the retrieval information into the pre-trained feature extraction model to obtain the initial feature vector corresponding to the retrieval information includes: if the retrieval information includes image information, inputting the image information into the image feature extraction sub-model to obtain the initial feature vector corresponding to the image information; if the retrieval information includes text information, inputting the text information into the text feature extraction sub-model to obtain the text feature vector corresponding to the text information, and determining the image feature vector similar to the text feature vector to obtain the initial feature vector corresponding to the text information.

[0112] After the step of, if the retrieval information includes image information, inputting the image information into the image feature extraction sub-model to obtain the initial feature vector corresponding to the image information, the method further includes: converting the image feature vector corresponding to the image information into a feature vector of a specified length and embedding it into a specified semantic feature space.

[0113] The above-mentioned step of determining the image feature vector similar to the text feature vector includes: converting the text feature vector corresponding to the text information into a feature vector of a specified length and embedding it into a specified semantic feature space; determining the image feature vector similar to the text feature vector corresponding to the text information through the position information of the vectors in the semantic feature space.

[0114] The above-mentioned step of determining the target vector from the vector database according to the style feature vector corresponding to the retrieval information includes: for each style feature vector in the vector database, calculating the similarity value between the style feature vector and the style feature vector corresponding to the retrieval information, and determining the style feature vector whose similarity value meets the preset threshold as the target vector.

[0115] The above-mentioned method further includes: if the retrieval information includes image information, saving the image information and the style feature vector corresponding to the image information to the vector database.

[0116] The above method further includes: obtaining a sample image, inputting the sample image into a pre-trained feature extraction model, and obtaining an image feature vector of the sample image through an image feature extraction sub-model in the feature extraction model; inputting the image feature vector into a pre-trained style feature extraction model to obtain a style feature vector of the sample image, and saving the sample image and the style feature vector of the sample image into a vector database.

[0117] The above feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the method further includes: obtaining first training data; wherein, the first training data includes first picture data and first text data; training the image feature extraction sub-model with the first picture data to obtain a trained image feature extraction sub-model; training the text feature extraction sub-model with the first text data to obtain a trained text feature extraction sub-model.

[0118] The above first text data is a data label of the first picture data; the steps of training the image feature extraction sub-model with the first picture data to obtain a trained image feature extraction sub-model and training the text feature extraction sub-model with the first text data to obtain a trained text feature extraction sub-model include: for each piece of first training data, inputting the first picture data into the image feature extraction sub-module to obtain an image feature vector corresponding to the first picture data, and embedding the image feature vector into a specified semantic feature space; inputting the first text data into the text feature extraction sub-module to obtain a text feature vector corresponding to the first text data, and embedding the text feature vector into the specified semantic feature space, and determining a semantic similarity value between the first picture data and the first text data according to the positional relationship between the vectors in the specified semantic feature space; updating the model parameters of the image feature extraction sub-module and the text feature extraction sub-module according to the semantic similarity value until a preset termination condition is reached, to obtain a trained image feature extraction sub-model and a trained text feature extraction sub-model.

[0119] The above style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the method further includes: obtaining second training data; wherein, the second training data includes second picture data or second text data, inputting the second training data into a pre-trained feature extraction model to obtain a feature vector corresponding to the second training data; training the style feature extraction sub-model with the feature vector corresponding to the second training data to obtain a trained style feature extraction sub-model.

[0120] The above method further includes: inputting the feature vector corresponding to the second training data into the trained style feature extraction sub-model to obtain an original style feature vector corresponding to the second training data; performing data analysis on the original style feature vector corresponding to the second training data to determine the model parameters of the dimensionality reduction model, to obtain a trained dimensionality reduction model.

[0121] The above method further includes: determining third training data from the retrieved information and the pictures corresponding to the retrieved information; the third training data includes third picture data and / or third text data; iteratively training the pre-trained image feature extraction sub-model with the third picture data to update the model parameters of the image feature extraction sub-model; training the pre-trained text feature extraction sub-model with the third text data to update the model parameters of the text feature extraction sub-model.

[0122] In the above case, if the third training data includes third picture data and third text data, the third text data is the data label of the third picture data.

[0123] The above method further includes: determining fourth training data from the retrieved information and the pictures corresponding to the retrieved information; the fourth training data includes fourth picture data and fourth text data, and the fourth text data is the data label of the fourth picture data; inputting the fourth picture data into the pre-trained image feature extraction sub-model to obtain the image feature vector corresponding to the fourth picture data; training the pre-trained style feature extraction sub-model with the image feature vector corresponding to the fourth picture data and the fourth text data to update the model parameters of the style feature extraction sub-model, and obtaining the original style feature vector corresponding to the fourth picture data; training the pre-trained dimensionality reduction sub-model with the original style feature vector corresponding to the fourth picture data to update the model parameters of the dimensionality reduction sub-model.

[0124] In this way, the pre-trained style feature extraction model can extract the painting style information in the retrieved information to obtain the style feature vector of the retrieved information, and then search for pictures similar to the style feature vector of the retrieved information from the vector database. Without excessive picture label information and text information, pictures with similar painting styles can be accurately retrieved, improving the retrieval accuracy and retrieval effect, and thus reducing the labor cost and maintenance cost.

[0125] This embodiment further provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above picture retrieval method.

[0126] The machine-executable instructions stored in the above machine-readable storage medium can, by executing the machine-executable instructions, implement the following operations in the above picture retrieval method:

[0127] Input the retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; input the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; determine a target vector from a vector database according to the style feature vector corresponding to the retrieval information, and determine the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes multiple style feature vectors and the pictures corresponding to the style feature vectors.

[0128] The above-mentioned pre-trained style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the step of inputting the initial feature vector corresponding to the retrieval information into the pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information includes: inputting the initial feature vector corresponding to the retrieval information into the style feature extraction sub-model to obtain an original style feature vector corresponding to the retrieval information; inputting the original style feature vector corresponding to the retrieval information into the dimensionality reduction sub-model to perform dimensionality reduction processing on the original style feature vector to obtain a style feature vector corresponding to the retrieval information.

[0129] The above-mentioned feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the step of inputting the retrieval information into the pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information includes: if the retrieval information includes image information, input the image information into the image feature extraction sub-model to obtain an initial feature vector corresponding to the image information; if the retrieval information includes text information, input the text information into the text feature extraction sub-model to obtain a text feature vector corresponding to the text information, and determine an image feature vector similar to the text feature vector to obtain an initial feature vector corresponding to the text information.

[0130] After the step of, if the retrieval information includes image information, inputting the image information into the image feature extraction sub-model to obtain an initial feature vector corresponding to the image information, the method further includes: converting the image feature vector corresponding to the image information into a feature vector of a specified length and embedding it into a specified semantic feature space.

[0131] The above-mentioned step of determining an image feature vector similar to the text feature vector includes: converting the text feature vector corresponding to the text information into a feature vector of a specified length and embedding it into a specified semantic feature space; determining an image feature vector similar to the text feature vector corresponding to the text information through the position information of the vectors in the semantic feature space.

[0132] The step of determining the target vector from the vector database according to the style feature vector corresponding to the retrieval information includes: for each style feature vector in the vector database, calculating the similarity value between the style feature vector and the style feature vector corresponding to the retrieval information, and determining the style feature vector whose similarity value meets the preset threshold as the target vector.

[0133] The above method further includes: if the retrieval information includes image information, saving the image information and the style feature vector corresponding to the image information to the vector database.

[0134] The above method further includes: obtaining a sample image, inputting the sample image into a pre-trained feature extraction model, and obtaining the image feature vector of the sample image through the image feature extraction sub-model in the feature extraction model; inputting the image feature vector into a pre-trained style feature extraction model to obtain the style feature vector of the sample image, and saving the sample image and the style feature vector of the sample image to the vector database.

[0135] The above feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; the method further includes: obtaining first training data; wherein, the first training data includes first picture data and first text data; training the image feature extraction sub-model through the first picture data to obtain a trained image feature extraction sub-model; training the text feature extraction sub-model through the first text data to obtain a trained text feature extraction sub-model.

[0136] The above first text data is the data label of the first picture data; the steps of training the image feature extraction sub-model through the first picture data to obtain a trained image feature extraction sub-model and training the text feature extraction sub-model through the first text data to obtain a trained text feature extraction sub-model include: for each first training data, inputting the first picture data into the image feature extraction sub-module to obtain the image feature vector corresponding to the first picture data, and embedding the image feature vector into a specified semantic feature space; inputting the first text data into the text feature extraction sub-module to obtain the text feature vector corresponding to the first text data, and embedding the text feature vector into the specified semantic feature space, and determining the semantic similarity value between the first picture data and the first text data according to the positional relationship between the vectors in the specified semantic feature space; updating the model parameters of the image feature extraction sub-module and the text feature extraction sub-module according to the semantic similarity value until the preset termination condition is reached, to obtain a trained image feature extraction sub-model and a trained text feature extraction sub-model.

[0137] The above-mentioned style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; the method further includes: obtaining second training data; wherein, the second training data includes second picture data or second text data, inputting the second training data into a pre-trained feature extraction model to obtain a feature vector corresponding to the second training data; training the style feature extraction sub-model with the feature vector corresponding to the second training data to obtain a trained style feature extraction sub-model.

[0138] The above-mentioned method further includes: inputting the feature vector corresponding to the second training data into the trained style feature extraction sub-model to obtain an original style feature vector corresponding to the second training data; performing data analysis on the original style feature vector corresponding to the second training data to determine the model parameters of the dimensionality reduction model, thereby obtaining a trained dimensionality reduction model.

[0139] The above-mentioned method further includes: determining third training data from the retrieval information and the pictures corresponding to the retrieval information; the third training data includes third picture data and / or third text data; iteratively training the pre-trained image feature extraction sub-model with the third picture data to update the model parameters of the image feature extraction sub-model; training the pre-trained text feature extraction sub-model with the third text data to update the model parameters of the text feature extraction sub-model.

[0140] In the above case, if the third training data includes third picture data and third text data, the third text data is the data label of the third picture data.

[0141] The above-mentioned method further includes: determining fourth training data from the retrieval information and the pictures corresponding to the retrieval information; the fourth training data includes fourth picture data and fourth text data, and the fourth text data is the data label of the fourth picture data; inputting the fourth picture data into the pre-trained image feature extraction sub-model to obtain an image feature vector corresponding to the fourth picture data; training the pre-trained style feature extraction sub-model with the image feature vector corresponding to the fourth picture data and the fourth text data to update the model parameters of the style feature extraction sub-model and obtain an original style feature vector corresponding to the fourth picture data; training the pre-trained dimensionality reduction sub-model with the original style feature vector corresponding to the fourth picture data to update the model parameters of the dimensionality reduction sub-model.

[0142] In this way, the pre-trained style feature extraction model can extract the painting style information in the retrieval information to obtain the style feature vector of the retrieval information, and then search for pictures similar to the style feature vector of the retrieval information in the vector database. Without excessive picture label information and text information, pictures with similar painting styles can be accurately retrieved, improving the retrieval accuracy and retrieval effect, and thus reducing the labor cost and maintenance cost.

[0143] A computer program product of a picture retrieval method, apparatus, and system provided by an embodiment of the present disclosure includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein.

[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system and apparatus can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0145] In addition, in the description of the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.

[0146] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program code.

[0147] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present disclosure. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0148] Finally, it should be noted that the above embodiments are only specific implementation manners of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for image retrieval, characterized in that, The method includes: Inputting the retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; Inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; Determining a target vector from a vector database according to the style feature vector corresponding to the retrieval information, and determining the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes a plurality of style feature vectors and the pictures corresponding to the style feature vectors.

2. The method according to claim 1, characterized in that, The pre-trained style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; The step of inputting the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information includes: Inputting the initial feature vector corresponding to the retrieval information into the style feature extraction sub-model to obtain an original style feature vector corresponding to the retrieval information; Inputting the original style feature vector corresponding to the retrieval information into the dimensionality reduction sub-model, performing dimensionality reduction processing on the original style feature vector to obtain a style feature vector corresponding to the retrieval information.

3. The method according to claim 1, characterized in that, The feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; The step of inputting the retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information includes: If the retrieval information includes image information, inputting the image information into the image feature extraction sub-model to obtain an initial feature vector corresponding to the image information; If the retrieval information includes text information, inputting the text information into the text feature extraction sub-model to obtain a text feature vector corresponding to the text information, and determining an image feature vector similar to the text feature vector to obtain an initial feature vector corresponding to the text information.

4. The method according to claim 3, wherein After the step of inputting the image information into the image feature extraction sub-model to obtain an initial feature vector corresponding to the image information if the retrieval information includes image information, the method further includes: Converting the image feature vector corresponding to the image information into a feature vector with a specified length and embedding it into a specified semantic feature space.

5. The method according to claim 4, wherein The step of determining an image feature vector similar to the text feature vector includes: Converting the text feature vector corresponding to the text information into a feature vector with a specified length and embedding it into the specified semantic feature space; Determining an image feature vector similar to the text feature vector corresponding to the text information through the position information of the vectors in the semantic feature space.

6. The method according to claim 1, wherein The step of determining a target vector from a vector database according to the style feature vector corresponding to the retrieval information includes: For each style feature vector in the vector database, calculating a similarity value between the style feature vector and the style feature vector corresponding to the retrieval information, and determining the style feature vector whose similarity value meets a preset threshold as the target vector.

7. The method according to claim 1, characterized in that, The method further includes: If the retrieved information includes image information, save the image information and the style feature vector corresponding to the image information to the vector database.

8. The method according to claim 1, wherein The method further includes: Obtain a sample image, input the sample image into a pre-trained feature extraction model, and obtain the image feature vector of the sample image through the image feature extraction sub-model in the feature extraction model; Input the image feature vector into a pre-trained style feature extraction model to obtain the style feature vector of the sample image, and save the sample image and the style feature vector of the sample image to the vector database.

9. The method according to claim 1, characterized in that The feature extraction model includes an image feature extraction sub-model and a text feature extraction sub-model; The method further includes: Obtain first training data; wherein, the first training data includes first picture data and first text data; Train the image feature extraction sub-model with the first picture data to obtain a trained image feature extraction sub-model, and train the text feature extraction sub-model with the first text data to obtain a trained text feature extraction sub-model.

10. The method according to claim 9, wherein The first text data is the data label of the first picture data; The steps of training the image feature extraction sub-model with the first picture data to obtain a trained image feature extraction sub-model and training the text feature extraction sub-model with the first text data to obtain a trained text feature extraction sub-model include: For each piece of the first training data, input the first picture data into the image feature extraction sub-module to obtain the image feature vector corresponding to the first picture data, and embed the image feature vector into a specified semantic feature space; Input the first text data into the text feature extraction sub-module to obtain the text feature vector corresponding to the first text data, and embed the text feature vector into the specified semantic feature space. According to the positional relationship between the vectors in the specified semantic feature space, determine the semantic similarity value between the first picture data and the first text data; According to the semantic similarity value, update the model parameters of the image feature extraction sub-module and the text feature extraction sub-module until a preset termination condition is reached, to obtain a trained image feature extraction sub-model and a trained text feature extraction sub-model.

11. The method according to claim 1, characterized in that, The style feature extraction model includes a style feature extraction sub-model and a dimensionality reduction sub-model; The method further includes: Obtain second training data; wherein, the second training data includes second picture data or second text data, input the second training data into a pre-trained feature extraction model to obtain the feature vector corresponding to the second training data; Train the style feature extraction sub-model with the feature vector corresponding to the second training data to obtain a trained style feature extraction sub-model.

12. The method according to claim 11, wherein The method further includes: Input the feature vector corresponding to the second training data into the trained style feature extraction sub-model to obtain the original style feature vector corresponding to the second training data; Perform data analysis on the original style feature vectors corresponding to the second training data to determine the model parameters of the dimensionality reduction model, and obtain a trained dimensionality reduction model.

13. The method according to claim 1, characterized in that The method further includes: Determine third training data from the retrieval information and the pictures corresponding to the retrieval information; the third training data includes third picture data and / or third text data; Iteratively train the pre-trained image feature extraction sub-model with the third picture data to update the model parameters of the image feature extraction sub-model; train the pre-trained text feature extraction sub-model with the third text data to update the model parameters of the text feature extraction sub-model.

14. The method according to claim 13, wherein If the third training data includes the third picture data and the third text data, the third text data is the data label of the third picture data.

15. The method according to claim 1, characterized in that, The method further includes: Determine fourth training data from the retrieval information and the pictures corresponding to the retrieval information; the fourth training data includes fourth picture data and fourth text data, and the fourth text data is the data label of the fourth picture data; Input the fourth picture data into the pre-trained image feature extraction sub-model to obtain the image feature vector corresponding to the fourth picture data; Train the pre-trained style feature extraction sub-model with the image feature vector corresponding to the fourth picture data and the fourth text data to update the model parameters of the style feature extraction sub-model, and obtain the original style feature vector corresponding to the fourth picture data; Train the pre-trained dimensionality reduction sub-model with the original style feature vector corresponding to the fourth picture data to update the model parameters of the dimensionality reduction sub-model.

16. An image retrieval device, characterized in that, The device includes: A feature extraction module, configured to input retrieval information into a pre-trained feature extraction model to obtain an initial feature vector corresponding to the retrieval information; the retrieval information is image information and / or text information; A style feature extraction sub-model, configured to input the initial feature vector corresponding to the retrieval information into a pre-trained style feature extraction model to obtain a style feature vector corresponding to the retrieval information; A picture retrieval module, configured to determine a target vector from a vector database according to the style feature vector corresponding to the retrieval information, and determine the picture corresponding to the target vector as the picture corresponding to the retrieval information; the vector database includes multiple style feature vectors and the pictures corresponding to the style feature vectors.

17. An electronic device, characterized in that, Comprising a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the picture retrieval method according to any one of claims 1-15.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the picture retrieval method according to any one of claims 1-15.