Text and image retrieval method and device, and computer readable storage medium
By using the image encoder, text encoder, and preset cyclic affine transformation module in the machine learning model, the consistency problem between text and image retrieval was solved, and accurate target recognition in vehicle retrieval was achieved.
Patent Information
- Application Number
- CN202211550479.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-12-05
AI Technical Summary
In the current technology, text-based vehicle retrieval has not yet been able to achieve accurate consistency between text and image retrieval in the field of video surveillance, especially in large vehicle image databases where it is difficult to accurately identify target vehicles.
A machine learning model is employed, including an image encoder, a text encoder, a preset cyclic affine transformation module, and a post-processing module. By acquiring the target vehicle image and text description, the preset cyclic affine transformation module performs multi-angle deep feature fusion, and the post-processing module establishes the correlation between image features and text features.
It achieves accurate consistency between text and image retrieval in vehicle search, improving the accuracy and efficiency of target vehicle identification.
Smart Images

Figure CN115934992B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video monitoring or computer technology, in particular to a text and image retrieval method and device and a computer readable storage medium. BACKGROUND
[0002] In the prior art, under the premise of natural language, text-based vehicle retrieval aims to identify the image of a target vehicle from a large vehicle image database. Since in most real application scenarios, text description is more easily accessible than other types of queries, text-based vehicle retrieval is of great significance in the field of video monitoring and has attracted more and more attention. However, most of the existing methods at present mainly focus on image-based vehicle retrieval, while text-based vehicle retrieval research is still in its infancy. Therefore, how to accurately realize the consistency of text and image retrieval in vehicle retrieval is an urgent problem to be solved. SUMMARY
[0003] The embodiments of the present application provide a text and image retrieval method, device and computer readable storage medium, which can accurately realize the consistency of text and image retrieval in vehicle retrieval.
[0004] In a first aspect, the embodiments of the present application provide a text and image retrieval method applied to an electronic device, wherein the electronic device is configured with a machine learning model, the machine learning model comprises an image encoder, a text encoder, at least one preset cyclic affine transformation module and a post-processing module, and the method comprises:
[0005] obtaining a target vehicle image and a text description of the target vehicle image;
[0006] inputting the target vehicle image into the image encoder to obtain a first image feature;
[0007] inputting the text description into the text encoder to obtain a first text feature;
[0008] inputting the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain a second image feature and a second text feature;
[0009] inputting the second image feature into the post-processing module for post-processing to obtain a processing result, and determining a target retrieval result according to the processing result and the second text feature.
[0010] Secondly, embodiments of this application provide a text and image retrieval device applied to an electronic device. The electronic device is equipped with a machine learning model, which includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The device includes: an acquisition unit, an extraction unit, a transformation unit, and a processing unit.
[0011] The acquisition unit is used to acquire a target vehicle image and a text description of the target vehicle image;
[0012] The extraction unit is used to input the target vehicle image into the image encoder to obtain a first image feature; and to input the text description into the text encoder to obtain a first text feature;
[0013] The transformation unit is used to input the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature.
[0014] The processing unit is used to input the second image features into the post-processing module for post-processing to obtain a processing result, and to determine the target retrieval result based on the processing result and the second text features.
[0015] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing the steps in the first aspect of embodiments of this application.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of embodiments of this application.
[0017] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of embodiments of this application. The computer program product may be a software installation package.
[0018] Implementing the embodiments of this application has the following beneficial effects:
[0019] As can be seen, the text and image retrieval method, apparatus, and computer-readable storage medium described in the embodiments of this application are applied to electronic devices. The electronic devices are equipped with machine learning models, which include: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The model acquires a target vehicle image and a text description of the target vehicle image. The target vehicle image is input to the image encoder to obtain a first image feature; the text description is input to the text encoder to obtain a first text feature; the first image feature and the first text feature are input to at least one preset cyclic affine transformation module to obtain a second image feature and a second text feature; the second image feature is input to the post-processing module for post-processing to obtain a processing result; and the target retrieval result is determined based on the processing result and the second text feature. On the one hand, vehicle images can be input to the text encoder and image encoder respectively to obtain corresponding text features and image features. On the other hand, the text features and image features can be fused from multiple angles in at least one preset cyclic affine transformation module to obtain fused text features and image features. The image features are then post-processed and mapped to the text features, thereby establishing a deep correlation between image features and text features. This enables accurate consistency between text and image retrieval in vehicle retrieval. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1A This is a schematic flowchart of a text and image retrieval method provided in an embodiment of this application;
[0022] Figure 1B This is a schematic diagram of the structure of a machine learning model provided in an embodiment of this application;
[0023] Figure 1C This is a schematic diagram of the structure of another machine learning model provided in an embodiment of this application;
[0024] Figure 1D This is a schematic diagram of the structure of the DSE module in a machine learning model provided in an embodiment of this application;
[0025] Figure 2 This is a flowchart illustrating another text and image retrieval method provided in an embodiment of this application;
[0026] Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0027] Figure 4 This is a block diagram of the functional units of a text and image retrieval device provided in an embodiment of this application. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0029] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0030] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0031] The electronic devices described in this application embodiments may include smartphones (such as Android phones, iOS phones, Windows Phones, etc.), tablet computers, PDAs, dashcams, servers, laptops, mobile internet devices (MIDs) or wearable devices (such as smartwatches, Bluetooth headsets), etc. The above are merely examples and not exhaustive, and include but are not limited to the above-mentioned electronic devices.
[0032] The embodiments of this application will be described in detail below.
[0033] Please see Figure 1A , Figure 1AThis is a flowchart illustrating a text and image retrieval method provided in an embodiment of this application. As shown in the figure, it is applied to an electronic device. The electronic device is equipped with a machine learning model, which includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The text and image retrieval method includes:
[0034] 101. Obtain the target vehicle image and a text description of the target vehicle image.
[0035] This system can acquire images of the target vehicle via a camera, recognize the images to obtain corresponding text descriptions, or input the target vehicle image into a preset neural network model to obtain a text description of the target vehicle image. The preset neural network model can include at least one of the following: convolutional neural network model, fully connected neural network model, recurrent neural network model, etc., without limitation. The text description is used to describe the target vehicle image using strings, text, arrays, etc.
[0036] In the embodiments of this application, such as Figure 1B As shown, the machine learning model may include: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. Any one of these modules can be preset or set by system default. Each preset cyclic affine transformation module can specify the number of iterations or a loop condition; if the number of iterations or the loop condition is met, the next stage is performed.
[0037] In this embodiment of the application, when a target vehicle is detected, a camera can be used to take a picture of the target vehicle to obtain an image of the target vehicle.
[0038] 102. Input the target vehicle image into the image encoder to obtain the first image feature.
[0039] In this embodiment of the application, the target vehicle image can be input into the image encoder to obtain the first image feature, or the target vehicle image can be preprocessed and then the preprocessed target vehicle image can be input into the image encoder to obtain the first image feature. The preprocessing can include at least one of the following: image denoising, image enhancement, target extraction, etc., which are not limited here. The first image feature can include at least one of the following: feature points, feature textures, feature vectors, feature values, etc., which are not limited here.
[0040] 103. Input the text description into the text encoder to obtain the first text feature.
[0041] In this embodiment of the application, the text description of the target vehicle image can be input into a text encoder to obtain the first text feature. Alternatively, the target vehicle image can be preprocessed, and then the text description of the preprocessed target vehicle image can be extracted and input into a text encoder to obtain the first text feature. The preprocessing can include at least one of the following: image denoising, image enhancement, target extraction, character recognition, etc., which are not limited here. The first text feature can include at least one of the following: string, sentence, feature vector, feature value, etc., which are not limited here.
[0042] 104. Input the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature.
[0043] In this embodiment, the first image features and the first text features can be input into at least one preset cyclic affine transformation module to obtain the second image features and the second text features.
[0044] 105. Input the second image features into the post-processing module for post-processing to obtain the processing result, and determine the target retrieval result based on the processing result and the second text features.
[0045] In this embodiment of the application, the second image features can be input into the post-processing module for post-processing to obtain the processing result, namely the processed image features, and the retrieval consistency can be achieved based on the image features and the second text features.
[0046] Optionally, the post-processing module includes a second convolution module, an activation function module, and a Gaussian interpolation module. Therefore, step 105, which involves inputting the second image features into the post-processing module for post-processing to obtain a processing result, and determining the target retrieval result based on the processing result and the second text features, may include the following steps:
[0047] 51. The second image features are sequentially input into the second convolution module, the activation function module, and the Gaussian interpolation module for processing to obtain the third image features;
[0048] 52. Determine the target retrieval result based on the second text feature and the third image feature.
[0049] In this embodiment, the post-processing module may include a second convolution module, an activation function module, and a Gaussian interpolation module. Specifically, the second image features can be sequentially input into the second convolution module, the activation function module, and the Gaussian interpolation module for processing to obtain the third image features. The third image features are the processing result, and the target retrieval result is determined based on the second text features and the third image features.
[0050] In this embodiment of the application, the preset cyclic affine transformation module includes: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module;
[0051] The first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features; the downsampling module is used to output the image features.
[0052] The second affine transformation module is used to perform an affine transformation on the first image features based on the text features output by the dynamic semantic combination module, so as to add spatial attention of text features to the image features;
[0053] The first convolutional module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolutional module. The dynamic semantic combination module is also used to output text features.
[0054] In this embodiment, a cross-modal affine transformation method is proposed to achieve the fusion of image perception and text semantic information. A cascaded cyclic affine transformation module connects the fusion blocks at each stage to the image-text encoder, and a spatial attention module is added to the image encoder to improve the semantic consistency between the text and the original image, thereby supervising the image encoder to extract more image content that conforms to the text description.
[0055] In the embodiments of this application, such as Figure 1C As shown, the preset cyclic affine transformation module may include: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module. Here, TF represents text features, which can be specifically represented as: W1, W2, ..., Wn. DSE represents the dynamic semantic combination module. The specific structure of DSE is as follows... Figure 1D As shown, W1, W2, ..., Wn represent different word vector subspaces, where n represents the number of subspaces and the dimension (i.e., subspace granularity) of each subspace. DSE first divides word features into subspaces of multiple granularities to construct a complete semantic space. Then, it configures a dynamic subspace router to generate a stage-aware path, resulting in more accurate and diverse semantic recombination results. Typically, given a list of subspace numbers and their corresponding word segmentation features, an attention mechanism is used to calculate the recombination of semantics represented by different subspaces.
[0056] Where W1, W2, ..., Wn represent the word vector subspace of the previous stage, and I1, I2, I3, ..., In represent the weights obtained by the image encoder, which are applied to the word vector subspace to obtain a new word vector space combination that is more suitable for the image.
[0057] In this embodiment, the dynamic semantic combination module allows text features at each stage (i.e., text and images from historical stages) to adaptively recombine according to the state of the historical stage, providing diverse and accurate semantic guidance. The dynamic semantic combination module can dynamically select the words to be recombinated at each stage to provide diverse and accurate semantic guidance. Because image encoders are a process from local to global, the text should also evolve synchronously during this process, providing semantic guidance from coarse-grained to fine-grained (e.g., from "vehicle" to "car" to "SUV"), to better guide the extraction of image features at each stage. Furthermore, by dynamically combining text features at different stages, previously used semantic information can be suppressed during the extraction process, and new, consistent semantic information can be activated, preventing the repeated generation of the same semantics and alleviating the problem of repetitive rendering.
[0058] In the embodiments of this application, such as Figure 1C As shown, the first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module performs an affine transformation on the first image features based on the first text features to incorporate spatial attention of text features into the image features; the downsampling module is used to output the image features. The second affine transformation module performs an affine transformation on the first image features based on the text features output by the dynamic semantic combination module to incorporate spatial attention of text features into the image features. The first convolution module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolution module. The dynamic semantic combination module is also used to output text features, enabling global allocation of text information during image encoding, thereby achieving the fusion of image perception and text semantic information.
[0059] In this embodiment, the training data can be a dataset composed of text-image pairs. Each pair contains a vehicle image captured by a specified surveillance camera, which is input into a text encoder and an image encoder to train a text-image model. The network structure is as follows: Figure 1CAs shown, the text encoder can be extracted by BERT, ATT represents the spatial attention module, and (I1, I2, I3, ..., In) in the DSE module represent the weight values of each word corresponding to the image region given by ATT. The image regions extracted by different feature layers have different focuses. Each dynamically combining semantics (DSE) module can recombine word features according to historical conditions.
[0060] In practice, it can start with the text description of the desired image and the initial image (random embedding, scene description in splines or pixels, or any image created in a distinguishable way), then run a cyclic affine transformation, and add multiple cascaded DSE modules to obtain more accurate word information to improve stability. The resulting image region embedding obtained by each DSE module is used to form the final image through a weighted hierarchical image integration, so that each cyclic affine transformation module only focuses on adding corresponding details based on the feature information of the previous stage, rather than the complete image.
[0061] Optionally, the first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features, including:
[0062] S1. Determine the aggregation weight based on the first text feature;
[0063] S2. Spatial attention embedding is performed on the first image features to obtain the image feature embedding after spatial attention;
[0064] S3. Determine the text information to be added to the image weights based on the image feature embedding after spatial attention and the aggregation weights;
[0065] S4. Perform cross-modal processing on the text information with added image weights using a specified activation function to obtain cross-modal features;
[0066] S5. Perform an affine transformation on the cross-modal features.
[0067] In this embodiment of the application, the specified activation function can be preset or defaulted to by the system. The specified activation function may include the tanh activation function.
[0068] In this embodiment, the aggregation weight can be determined based on the first text feature, and then the first image feature can be spatially attention-embedded to obtain the spatially attention-embedded image feature. The text information to be added to the image weight can be determined based on the spatially attention-embedded image feature and the aggregation weight. The text information to be added to the image weight can be cross-modal processed using a specified activation function to obtain cross-modal features. The cross-modal features can be affine transformed.
[0069] In practice, learnable full-time projective image features can be used to perform spatial attention on the image, resulting in a small number of N image vectors, which are then used to obtain aggregation weights. The aggregation process is as follows:
[0070]
[0071] in, W represents the image feature embedding after spatial attention. i-1 W′ represents the textual feature information of the current stage. i-1 This indicates the text information that is added to the image weights.
[0072] Then the recombined W′ i-1 W″ was obtained after cross-modal processing. i-1 The cross-modal processing procedure, which applies to the image features in the next stage, is as follows:
[0073] W″ i-1 =tanh(W′) i-1 )
[0074] In this embodiment, the tanh activation function is used instead of the softmax function because softmax maximizes the probability and suppresses other probabilities from approaching zero. Extremely small probabilities hinder backpropagation of gradients, exacerbating the instability of image encoder training. In contrast, the tanh function prevents attention probabilities from approaching zero, improving the efficiency of backpropagation and thus forcing the generator to synthesize more relevant information.
[0075] Optionally, before step 101, the following steps may also be included:
[0076] The machine learning model is trained using a preset loss function to obtain a machine learning model that meets preset requirements;
[0077] The preset loss function consists of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function.
[0078] The local loss function is obtained based on the principle of local similarity, specifically by averaging the matching scores of a sentence in the text with the most relevant object in the image.
[0079] The global loss function is obtained based on global similarity, specifically: it is obtained based on the degree of matching between text vectors and image vectors;
[0080] The cyclic affine transformation loss function is obtained based on the matching degree between the image and the text description.
[0081] In this embodiment, the preset loss function can be preset or defaulted to by the system. The preset loss function can be the loss function of at least one of the following modules: image encoder, text encoder, at least one preset cyclic affine transformation module, and post-processing module.
[0082] In a specific implementation, the machine learning model can be trained using a preset loss function (LOSS) to obtain a machine learning model that meets the preset requirements. The preset requirements can be set in advance or are system defaults. The preset requirements may include at least one of the following: reaching a set number of training iterations, satisfying preset convergence conditions, etc., which are not limited here. The set number of training iterations and the preset convergence conditions can both be set in advance or are system defaults.
[0083] The preset loss function can be composed of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function. The local loss function is based on the principle of local similarity, specifically: it is obtained by averaging the matching scores of a sentence in the text with the most relevant object in the image. The global loss function can be obtained based on global similarity, specifically: it is obtained based on the degree of matching between the text vector and the image vector. The cyclic affine transformation loss function can be obtained by calculating the matching degree between the image and the text description.
[0084] In this embodiment of the application, during the encoder process, image features can be extracted based on text information through multiple cascaded cyclic affine transformations. By calculating the similarity loss function between the text description and the vehicle image, and performing gradient updates and backpropagation to the image encoder, the image encoder is forced to extract more relevant image features.
[0085] In this embodiment, the preset loss function can be composed of a local loss function, a global loss function, and a cyclic affine transformation loss function. It can introduce external unstructured parameters to generate image-text pairs. Through contrastive learning, specifically utilizing contrastive losses between image-to-text, image region-to-word, and text-to-image pairs, alignment between images and text is generated. First, local similarity and global similarity are calculated. Local similarity is the average matching score of the most relevant objects between a sentence and an image, as shown below:
[0086]
[0087] Where, N w w represents the number of words in the text.T I represents text, and I represents an image.
[0088] Furthermore, global similarity uses the traditional cosine distance to measure the text vector w. v and image vector I v The degree of matching between them is represented as follows:
[0089]
[0090] Next, the cyclic affine transformation loss function can be used to calculate the matching degree between the image and the text description. It can be defined as the Kullback-Leibler divergence between the standard Gaussian distribution and the Gaussian branch of the training text, calculated as follows:
[0091] L CA =D KL (N(μ(s),∑(s)||N(0,I)))
[0092] Where s represents the matched sentence, The sentence represents a mismatch, and x represents the corresponding real image of s. This represents the generated pseudo-image corresponding to s.
[0093] Yes, the loss function (LOSS) consists of the three components mentioned above:
[0094]
[0095] In this embodiment of the application, in order to achieve better word-image region alignment, external unstructured knowledge is introduced to reconstruct the text, and synonym enhancement is performed on some words. Some object instances in the image are explicitly modeled to smoothly connect the text and the image. The mutual information between correspondences is maximized through contrastive learning. During the learning process, the contrastive loss between the image and the sentence and between the image region and the word is learned to strengthen the alignment between the generated image and the text, thereby enhancing the robust alignment between the image and the text.
[0096] As can be seen, the text and image retrieval method described in this application embodiment is applied to an electronic device. The electronic device is equipped with a machine learning model, which includes an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The method acquires a target vehicle image and a text description of the target vehicle image. The target vehicle image is input to the image encoder to obtain a first image feature; the text description is input to the text encoder to obtain a first text feature; the first image feature and the first text feature are input to at least one preset cyclic affine transformation module to obtain a second image feature and a second text feature; the second image feature is input to the post-processing module for post-processing to obtain a processing result; and the target retrieval result is determined based on the processing result and the second text feature. On the one hand, the vehicle image can be input to the text encoder and the image encoder respectively to obtain corresponding text features and image features. On the other hand, the text features and image features can be fused from multiple angles in at least one preset cyclic affine transformation module to obtain fused text features and image features. The image features are then post-processed and mapped to the text features, thereby establishing a deep correlation between image features and text features. This enables accurate consistency between text and image retrieval in vehicle retrieval.
[0097] With the above Figure 1A The illustrated embodiments are consistent with those shown in the example. Please refer to [link / reference]. Figure 2 , Figure 2 This is a flowchart illustrating another text and image retrieval method provided in this application embodiment, applied to an electronic device. The electronic device is equipped with a machine learning model, which includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. As shown in the figure, this text and image retrieval method includes:
[0098] 201. The machine learning model is trained using a preset loss function to obtain a machine learning model that meets preset requirements; wherein the preset loss function consists of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function.
[0099] 202. Obtain the target vehicle image and a text description of the target vehicle image.
[0100] 203. Input the target vehicle image into the image encoder to obtain the first image feature.
[0101] 204. Input the text description into the text encoder to obtain the first text feature.
[0102] 205. Input the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature.
[0103] 206. Input the second image features into the post-processing module for post-processing to obtain the processing result, and determine the target retrieval result based on the processing result and the second text features.
[0104] The specific descriptions of steps 201-206 above can be found in the above descriptions. Figure 1A The corresponding steps of the described text and image retrieval methods will not be repeated here.
[0105] As can be seen, the text and image retrieval method described in this application embodiment has the following characteristics: First, vehicle images can be input into a text encoder and an image encoder respectively to obtain corresponding text features and image features. Second, the text features and image features can be fused into multi-angle deep features in at least one preset cyclic affine transformation module to obtain fused text features and image features. Then, the image features are post-processed and mapped to the text features. Third, its loss function can introduce external unstructured parameters to generate image-text. Through contrastive learning, specifically using the contrastive loss between image to text, image region to word, and text to image, the alignment between image and text is generated. Thus, the correlation between image features and text features is deeply established, and the consistency of text and image retrieval can be accurately achieved in vehicle retrieval.
[0106] Consistent with the above embodiments, please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs. The electronic device is configured with a machine learning model, which includes an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The one or more programs are stored in the memory and configured to be executed by the processor. In this embodiment, the programs include instructions for performing the following steps:
[0107] Obtain an image of the target vehicle and a text description of the target vehicle image;
[0108] The target vehicle image is input into the image encoder to obtain the first image feature;
[0109] The text description is input into the text encoder to obtain the first text feature;
[0110] The first image feature and the first text feature are input into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature;
[0111] The second image features are input into the post-processing module for post-processing to obtain the processing result. The target retrieval result is determined based on the processing result and the second text features.
[0112] Optionally, the preset cyclic affine transformation module includes: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module;
[0113] The first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features; the downsampling module is used to output the image features.
[0114] The second affine transformation module is used to perform an affine transformation on the first image features based on the text features output by the dynamic semantic combination module, so as to add spatial attention of text features to the image features;
[0115] The first convolutional module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolutional module. The dynamic semantic combination module is also used to output text features.
[0116] Optionally, the first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features, including:
[0117] The aggregation weight is determined based on the first text feature;
[0118] The first image features are spatially attention-embedded to obtain spatially attention-embedded image features.
[0119] The text information to be added to the image weights is determined based on the image feature embedding after spatial attention and the aggregation weights;
[0120] The text information with added image weights is processed cross-modally using a specified activation function to obtain cross-modal features;
[0121] Perform an affine transformation on the cross-modal features.
[0122] Optionally, the post-processing module includes: a second convolution module, an activation function module, and a Gaussian interpolation module. Regarding the step of inputting the second image features into the post-processing module for post-processing to obtain a processing result, and determining the target retrieval result based on the processing result and the second text features, the above program includes instructions for performing the following steps:
[0123] The second image features are sequentially input into the second convolution module, the activation function module, and the Gaussian interpolation module for processing to obtain the third image features;
[0124] The target retrieval result is determined based on the second text feature and the third image feature.
[0125] Optionally, the above procedure may also include instructions for performing the following steps:
[0126] The machine learning model is trained using a preset loss function to obtain a machine learning model that meets preset requirements;
[0127] The preset loss function consists of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function.
[0128] The local loss function is obtained based on the principle of local similarity, specifically by averaging the matching scores of a sentence in the text with the most relevant object in the image.
[0129] The global loss function is obtained based on global similarity, specifically: it is obtained based on the degree of matching between text vectors and image vectors;
[0130] The cyclic affine transformation loss function is obtained based on the matching degree between the image and the text description.
[0131] As can be seen, the electronic device described in this application embodiment is equipped with a machine learning model, which includes an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The model acquires a target vehicle image and a text description of the target vehicle image. The target vehicle image is input to the image encoder to obtain a first image feature; the text description is input to the text encoder to obtain a first text feature; the first image feature and the first text feature are input to at least one preset cyclic affine transformation module to obtain a second image feature and a second text feature; the second image feature is input to the post-processing module for post-processing to obtain a processing result; and the target retrieval result is determined based on the processing result and the second text feature. On the one hand, the vehicle image can be input to the text encoder and the image encoder respectively to obtain corresponding text features and image features. On the other hand, the text features and image features can be fused from multiple angles in at least one preset cyclic affine transformation module to obtain fused text features and image features. The image features are then post-processed and correlated with the text features, thereby establishing a deep correlation between image features and text features. This enables accurate consistency between text and image retrieval in vehicle retrieval.
[0132] Figure 4 This is a functional unit block diagram of a text and image retrieval device 400 according to an embodiment of this application. The text and image retrieval device 400 is applied to an electronic device equipped with a machine learning model. The machine learning model includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The device includes: an acquisition unit 401, an extraction unit 402, a transformation unit 403, and a processing unit 404.
[0133] The acquisition unit 401 is used to acquire a target vehicle image and a text description of the target vehicle image;
[0134] The extraction unit 402 is used to input the target vehicle image into the image encoder to obtain a first image feature; and to input the text description into the text encoder to obtain a first text feature;
[0135] The transformation unit 403 is used to input the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature.
[0136] The processing unit 404 is used to input the second image features into the post-processing module for post-processing to obtain a processing result, and determine the target retrieval result based on the processing result and the second text features.
[0137] Optionally, the preset cyclic affine transformation module includes: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module;
[0138] The first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features; the downsampling module is used to output the image features.
[0139] The second affine transformation module is used to perform an affine transformation on the first image features based on the text features output by the dynamic semantic combination module, so as to add spatial attention of text features to the image features;
[0140] The first convolutional module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolutional module. The dynamic semantic combination module is also used to output text features.
[0141] Optionally, the first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features, including:
[0142] The aggregation weight is determined based on the first text feature;
[0143] The first image features are spatially attention-embedded to obtain spatially attention-embedded image features.
[0144] The text information to be added to the image weights is determined based on the image feature embedding after spatial attention and the aggregation weights;
[0145] The text information with added image weights is processed cross-modally using a specified activation function to obtain cross-modal features;
[0146] Perform an affine transformation on the cross-modal features.
[0147] Optionally, the post-processing module includes: a second convolution module, an activation function module, and a Gaussian interpolation module. In the process of inputting the second image features into the post-processing module for post-processing to obtain a processing result, and determining the target retrieval result based on the processing result and the second text features, the processing unit 404 is specifically used for:
[0148] The second image features are sequentially input into the second convolution module, the activation function module, and the Gaussian interpolation module for processing to obtain the third image features;
[0149] The target retrieval result is determined based on the second text feature and the third image feature.
[0150] Optionally, the device 400 is further specifically used for:
[0151] The machine learning model is trained using a preset loss function to obtain a machine learning model that meets preset requirements;
[0152] The preset loss function consists of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function.
[0153] The local loss function is obtained based on the principle of local similarity, specifically by averaging the matching scores of a sentence in the text with the most relevant object in the image.
[0154] The global loss function is obtained based on global similarity, specifically: it is obtained based on the degree of matching between text vectors and image vectors;
[0155] The cyclic affine transformation loss function is obtained based on the matching degree between the image and the text description.
[0156] As can be seen, the text and image retrieval device described in this application embodiment is applied to an electronic device. The electronic device is equipped with a machine learning model, which includes an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The device acquires a target vehicle image and a text description of the target vehicle image. The target vehicle image is input to the image encoder to obtain a first image feature; the text description is input to the text encoder to obtain a first text feature; the first image feature and the first text feature are input to at least one preset cyclic affine transformation module to obtain a second image feature and a second text feature; the second image feature is input to the post-processing module for post-processing to obtain a processing result; and the target retrieval result is determined based on the processing result and the second text feature. On the one hand, vehicle images can be input to the text encoder and the image encoder respectively to obtain corresponding text features and image features. On the other hand, the text features and image features can be fused from multiple angles in at least one preset cyclic affine transformation module to obtain fused text features and image features. The image features are then post-processed and mapped to the text features, thereby establishing a deep correlation between image features and text features. This enables accurate consistency between text and image retrieval in vehicle retrieval.
[0157] It is understood that the functions of each program module of the text and image retrieval device in this embodiment can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above method embodiments, and will not be repeated here.
[0158] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.
[0159] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.
[0160] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0161] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0162] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0163] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0164] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0165] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0166] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0167] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A text and image retrieval method, characterized in that, The method, applied to an electronic device equipped with a machine learning model, includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. Obtain an image of the target vehicle and a text description of the target vehicle image; The target vehicle image is input into the image encoder to obtain the first image feature; The text description is input into the text encoder to obtain the first text feature; The first image feature and the first text feature are input into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature; The second image features are input into the post-processing module for post-processing to obtain the processing result. The target retrieval result is determined based on the processing result and the second text features. The preset cyclic affine transformation module includes: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module; The first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features; the downsampling module is used to output the image features. The second affine transformation module is used to perform an affine transformation on the first image features based on the text features output by the dynamic semantic combination module, so as to add spatial attention of text features to the image features; The first convolutional module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolutional module. The dynamic semantic combination module is also used to output text features.
2. The method according to claim 1, characterized in that, The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features, including: The aggregation weight is determined based on the first text feature; The first image features are spatially attention-embedded to obtain spatially attention-embedded image features. The text information to be added to the image weights is determined based on the image feature embedding after spatial attention and the aggregation weights; The text information with added image weights is processed cross-modally using a specified activation function to obtain cross-modal features; Perform an affine transformation on the cross-modal features.
3. The method according to claim 1 or 2, characterized in that, The post-processing module includes a second convolution module, an activation function module, and a Gaussian interpolation module. The step of inputting the second image features into the post-processing module for post-processing to obtain a processing result, and determining the target retrieval result based on the processing result and the second text features, includes: The second image features are sequentially input into the second convolution module, the activation function module, and the Gaussian interpolation module for processing to obtain the third image features; The target retrieval result is determined based on the second text feature and the third image feature.
4. The method according to claim 1 or 2, characterized in that, The method further includes: The machine learning model is trained using a preset loss function to obtain a machine learning model that meets preset requirements; The preset loss function consists of three parts: a local loss function, a global loss function, and a cyclic affine transformation loss function. The local loss function is obtained based on the principle of local similarity, specifically by averaging the matching scores of a sentence in the text with the most relevant object in the image. The global loss function is obtained based on global similarity, specifically: it is obtained based on the degree of matching between text vectors and image vectors; The cyclic affine transformation loss function is obtained based on the matching degree between the image and the text description.
5. A text and image retrieval device, characterized in that, The device is applied to electronic devices, which are equipped with a machine learning model. The machine learning model includes: an image encoder, a text encoder, at least one preset cyclic affine transformation module, and a post-processing module. The device includes: an acquisition unit, an extraction unit, a transformation unit, and a processing unit. The acquisition unit is used to acquire a target vehicle image and a text description of the target vehicle image; The extraction unit is used to input the target vehicle image into the image encoder to obtain a first image feature; and to input the text description into the text encoder to obtain a first text feature; The transformation unit is used to input the first image feature and the first text feature into the at least one preset cyclic affine transformation module to obtain the second image feature and the second text feature. The processing unit is used to input the second image features into the post-processing module for post-processing to obtain a processing result, and to determine the target retrieval result based on the processing result and the second text features. The preset cyclic affine transformation module includes: a first affine transformation module, a first convolution module, a second affine transformation module, a downsampling module, and a dynamic semantic combination module; The first affine transformation module is connected to the first convolution module, the first convolution module is connected to the second affine transformation module, and the second affine transformation module is connected to the downsampling module. The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features; the downsampling module is used to output the image features. The second affine transformation module is used to perform an affine transformation on the first image features based on the text features output by the dynamic semantic combination module, so as to add spatial attention of text features to the image features; The first convolutional module is connected to the dynamic semantic combination module, which edits the first text features based on the image features of the first convolutional module. The dynamic semantic combination module is also used to output text features.
6. The apparatus according to claim 5, characterized in that, The first affine transformation module is used to perform an affine transformation on the first image features based on the first text features, so as to add spatial attention of text features to the image features, including: The aggregation weight is determined based on the first text feature; The first image features are spatially attention-embedded to obtain spatially attention-embedded image features. The text information to be added to the image weights is determined based on the image feature embedding after spatial attention and the aggregation weights; The text information with added image weights is processed cross-modally using a specified activation function to obtain cross-modal features; Perform an affine transformation on the cross-modal features.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory being used to store one or more programs and configured to be executed by the processor, the programs including instructions for performing the steps of the method as described in any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Text domain image retrieval
US20200175061A1