A cross-modal retrieval method and system
By constructing a visual-language search model and optimizing the loss function, the inaccurate search problem caused by noise interference in cross-modal search is solved, and more accurate search results and completion of incomplete data are achieved.
Patent Information
- Application Number
- CN202211435114.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The prior art is susceptible to noise interference in cross-modal retrieval, resulting in inaccurate search results, especially inaccurate search of incomplete text data.
By building an initial vision-language search model, including visual encoder, text encoder and cross-modal decoder, and setting image reconstruction loss function, noise adaptive comparison loss function and image description loss function, optimize the model to generate reconstructed text data and improve the accuracy of the search results.
It effectively avoids the search model overfitting the data set of pictures and texts containing noise, improves the accuracy of the search results, and can complete incomplete text data.
Smart Images

Figure CN115718815B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and more specifically, to a cross-modal retrieval method and system. Background Art
[0002] In recent years, the field of deep learning has developed vigorously, and computer vision and natural language processing have developed most rapidly. Visual language pre-training technology connects the two fields of computer vision and natural language processing for joint training, projects the visual modality and the text modality into a unified representation space, and aligns the visual modality and the text modality. Visual language models greatly reduce the requirements for manually labeled data. It can learn the weak correlation between the visual modality and the text modality from a large number of image-text pairs crawled from the network. Eventually, its zero-shot classification performance exceeds that of supervised models. Visual language pre-training models are affected by noise interference, and the training data set used for model training requires a refined screening process to ensure the quality of the data set. For a large number of image-text data sets crawled from the network, the manually designed filtering strategy cannot ensure that the model is not affected by noise; the noise in the image-text data set mainly comes from inaccurate and incomplete descriptions of text for pictures. The visual language model trained using the noisy image-text data set will greatly reduce the model performance due to being trapped in the noise, and the retrieval results are inaccurate; when the text information or image information is incomplete, the retrieval results cannot even be obtained.
[0003] The prior art discloses a cross-modal retrieval method, device, storage medium, and terminal based on semantic enhancement. The method includes constructing a cross-modal retrieval model, and training the cross-modal retrieval model based on an image-text retrieval data set to obtain a trained cross-modal retrieval model; determining target query data and a target modality data set, and obtaining the overall semantic similarity between the target query data and each target modality data based on the trained cross-modal retrieval model; selecting a preset number of target modality data corresponding to the overall semantic similarity in descending order of the overall semantic similarity in the target modality data set, and determining the retrieval result. This application has high requirements for manually labeled data, requires a large amount of complete image-text data for training, and is easily affected by noise; and for incomplete text data, the image data cannot be accurately retrieved. Summary of the Invention
[0004] In order to overcome the defect of inaccurate retrieval results in the above-mentioned prior art cross-modal retrieval, the present invention provides a cross-modal retrieval method and system, which can obtain accurate cross-modal retrieval results, and realize the complementation and filling of incomplete text pair data.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] The present invention provides a cross-modal retrieval method, including:
[0007] S1: Obtain an image-text pair dataset, including corresponding image data and text data;
[0008] S2: Construct an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder;
[0009] S3: Randomly cover pixel blocks on the image data to obtain a masked block image; randomly mask the text data to obtain masked text data;
[0010] S4: Input the masked block image and the image data into the vision encoder to obtain a masked block image encoding and an image data encoding, and set an image reconstruction loss function according to the masked block image encoding and the image data;
[0011] S5: Input the image data into a preset vision concept vocabulary library to obtain vision concept words; and input the vision concept words and the masked text data into the text encoder to obtain a vision concept-enhanced text encoding;
[0012] S6: Set an image description loss function according to the text data, the vision concept-enhanced text encoding, and the image data encoding;
[0013] S7: Input the image data, the text data, and the vision concept-enhanced text encoding into the cross-modal decoder, generate a pure text data encoding according to the text data and the vision concept-enhanced text encoding, and generate a reconstructed text data according to the image data and the vision concept-enhanced text encoding;
[0014] S8: Calculate the image-text pair noise probability according to the image data encoding and the pure text data encoding, and set a noise adaptive contrast loss function;
[0015] S9: Use the noise probability as the replacement probability, and replace the corresponding text data with the reconstructed text data according to the replacement probability to obtain reconstructed image-text pair data;
[0016] S10: Construct a total loss function according to the image reconstruction loss function, the noise adaptive contrast loss function, and the image description loss function, and optimize the total loss function by using the reconstructed image-text pair data to obtain an optimized vision-language retrieval model;
[0017] S11: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain a retrieval result.
[0018] The present invention obtains an image-text pair dataset, where there are defects such as noise influence or incomplete text between the mutually corresponding image data and text data; constructs an initial vision-language retrieval model, including a vision encoder for maintaining high-quality visual feature representation while reducing computational costs; a text encoder for encoding text data and auxiliary visual concept words; a cross-modal decoder for synthesizing semantically consistent reconstructed text data; constructs a total loss function according to an image reconstruction loss function, a noise adaptive contrast loss function, and an image caption loss function, optimizes the initial vision-language retrieval model to obtain an optimized vision-language retrieval model, uses the image-text pair noise probability for noise adaptive regularization to avoid severe deviation from noise, enhances the robustness of the retrieval model, can effectively avoid overfitting of the retrieval model to the image-text pair dataset containing noise, the generated reconstructed text data contains rich and detailed image descriptions, improves the accuracy of retrieval results, and can also complete the image-text pair dataset with incomplete text data.
[0019] Preferably, in the step S4, the image reconstruction loss function is:
[0020]
[0021] In the formula, L IR is the image reconstruction loss value, N represents the number of image-text pair data, x i represents the i-th image data, x' i represents the i-th masked block image, V e (x' i ) represents the encoding of the i-th masked block image, and ‖*‖ represents the calculation of the second norm.
[0022] Preferably, in the step S8, the specific method for calculating the image-text pair noise probability according to the image data encoding and the pure text data encoding is:
[0023] S8.1: For each group of image-text pairs, calculate the similarity of the image data relative to the text data and the similarity of the text data relative to the image data;
[0024] S8.2: Use the similarity of the image data relative to the text data and the similarity of the text data relative to the image data of all image-text pairs to calculate the total image-text contrast learning loss;
[0025] S8.3: Use a two-component Gaussian mixture model to calculate the image-text pair noise probability according to the image-text contrast loss.
[0026] Preferably, in the step S8.1, the specific method for calculating the similarity of the image data relative to the text data and the similarity of the text data relative to the image data is:
[0027]
[0028] In the formula, represents the similarity of the i-th image data relative to the j-th text data, represents the similarity of the j-th text data relative to the i-th image data.
[0029] Preferably, in the step S8.2, the specific method for calculating the total image-text contrast learning loss function is:
[0030]
[0031]
[0032]
[0033]
[0034] In the formula, B represents the number of image pairs input in this batch, represents the similarity of the i-th text data relative to the i-th image data, represents the similarity of the i-th image data relative to the i-th text data; L ITC (x i ,y i ) represents the i-th image-text contrast loss, and L ITC represents the total image-text contrast learning loss; represents the image-to-text contrast learning loss, represents the text-to-image contrast learning loss.
[0035] Preferably, in the step S8.3, the specific method for calculating the image-text pair noise probability according to the image-text contrast loss is:
[0036]
[0037] ∈ i = p(μ h )p(L ITC (x i ,y i )|μ h ) / p(L ITC (x i ,y i ))
[0038] In the formula, p(*) represents obtaining the probability distribution, θ represents the parameters of the two-component Gaussian mixture model, m represents the number of components of the two-component Gaussian mixture model, γ m is the mixing covariance of the m-th component of the two-component Gaussian mixture model, φ(*|m) represents obtaining the probability density of the m-th component of the two-component Gaussian mixture model, and μ hRepresents the high mean component, ∈ i Represents the noise probability of the i-th text-image pair.
[0039] To align different modalities, existing vision-language pre-training models use text-image contrastive loss to ensure that the positive text-image pairs {x i , y i} i=y are aligned in the same feature space, while the negative text-image pairs {x i , y i} i≠j are the opposite; however, existing text-image contrastive loss forces the vision-language pre-training model to align the features of each image-text pair without considering the cases where there is noise, reducing the model performance; the present invention coordinates cross-modal alignment to different degrees according to the text-image pair noise probability, and the text-image pair noise probability indicates the degree of semantic mismatch between the text-image pair; the present invention preferentially fits clean text-image pairs, followed by text-image pairs containing noise, that is, uses a two-component Gaussian mixture model to fit the text-image contrastive loss of each text-image pair; according to the text-image pair noise probability, regularize the true alignment labels to different degrees, and use lower regularization for text-image pair data with low text-image pair noise probability to learn alignment, and use higher regularization for text-image pair data with high text-image pair noise probability to learn alignment, which can effectively avoid overfitting noise.
[0040] Preferably, in the step S8, the specific method for setting the noise adaptive contrastive loss function according to the text-image pair noise probability is:
[0041] Set the text-image pair smoothing rate w i :
[0042] w i = λ ∈ i
[0043] where w i represents the smoothing rate of the i-th text-image pair, λ represents a hyperparameter, and ∈ i represents the noise probability of the i-th text-image pair;
[0044] Calculate according to the text-image pair smoothing rate and
[0045]
[0046]
[0047] In the formula, B represents the number of image pairs; represents the noise adaptive image-to-text contrastive learning loss, represents the noise adaptive text-to-image contrastive learning;
[0048] Calculate the noise adaptive contrast loss function:
[0049]
[0050] In the formula, L NITC represents the noise adaptive contrast loss value.
[0051] Preferably, in step S5, the preset visual concept vocabulary library is constructed by relying on various noun concepts in the corpus collected by the parsing network. Denote the preset visual concept vocabulary library as Q. Input the image data into the preset visual concept vocabulary library Q, and obtain the visual concept word q by calculating the similarity with the nouns. The similarity between the image data and the nouns is calculated by the following formula:
[0052] sim(x i , Q) = cos(V e (x), T e (Q))
[0053] Sort all similarities from largest to smallest in value, and select the top k nouns with the largest similarities as the visual concept word q.
[0054] Preferably, in step S6, the image description loss function is specifically:
[0055] L LM = -E (x,y)~D logp(y_t|C d (y τ<t , [V e (x), T e ([q, y m )]))
[0056] In the formula, L LM represents the image description loss value, C d represents the cross-modal decoder, y t represents the t-th word in the text data, q represents the visual concept word, y m is the masked text data, and T e is the text encoder.
[0057] Preferably, in step S10, the total loss function is specifically:
[0058] L = L IR + α·L LM + β·L NITC
[0059] In the formula, L represents the total loss function value, α represents the first weight coefficient, and β represents the second weight coefficient.
[0060] The present invention also provides a cross-modal retrieval system. Based on the above cross-modal retrieval method, the system includes:
[0061] A data acquisition module, configured to acquire a graph-text pair dataset, including corresponding image data and text data;
[0062] A model construction module, configured to construct an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder;
[0063] A masking module, configured to randomly cover pixel blocks on the image data to obtain a masked block image; and randomly mask the text data to obtain masked text data;
[0064] An image encoding module, configured to input the masked block image and the image data into the vision encoder to obtain image data encoding and masked block image encoding, and set an image reconstruction loss function according to the masked block image encoding and the image data;
[0065] A visual concept enhancement module, configured to input the image data into a preset visual concept vocabulary library to obtain visual concept words; and input the visual concept words and the masked text data into the text encoder to obtain visually concept-enhanced text encoding;
[0066] An image description loss function setting module, configured to set an image description loss function according to the text data, the visually concept-enhanced text encoding, and the image data encoding;
[0067] A pure text and reconstructed text generation module, configured to input the image data, the text data, and the visually concept-enhanced text encoding into the cross-modal decoder to obtain pure text data encoding and reconstructed text data;
[0068] A noise adaptation module, configured to calculate the graph-text pair noise probability according to the image data encoding and the pure text data encoding, and set a noise adaptation contrast loss function;
[0069] A graph-text pair data reconstruction module, configured to use the noise probability as the replacement probability, and replace the corresponding text data with the reconstructed text data according to the replacement probability to obtain reconstructed graph-text pair data;
[0070] A model optimization module, configured to construct a total loss function according to the image reconstruction loss function, the noise adaptation contrast loss function, and the image description loss function, and optimize the total loss function using the reconstructed graph-text pair data to obtain an optimized vision-language retrieval model;
[0071] A cross-modal retrieval module, configured to input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
[0072] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0073] After obtaining the graphic-text pair dataset, the present invention constructs an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder; respectively sets an image reconstruction loss function, a noise adaptive contrast loss function, and an image description loss function, and constructs a total loss function to optimize the initial vision-language retrieval model to obtain an optimized vision-language retrieval model; uses the graphic-text pair noise probability for noise adaptive regularization to avoid serious deviation from noise and enhance the robustness of the retrieval model, which can effectively avoid overfitting of the retrieval model to the graphic-text pair dataset containing noise; by retrieving visual concept words and image data in a preset visual concept vocabulary library and inputting them into the cross-modal decoder to generate reconstructed text data, the relevance between graphic-text pairs is improved; the reconstructed text data contains rich and detailed image descriptions, improving the accuracy of retrieval results, and can also complete the graphic-text pair dataset with incomplete text data. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a flowchart of a cross-modal retrieval method described in Embodiment 1.
[0075] Figure 2 It is a schematic structural diagram of the initial vision-language retrieval model described in Embodiment 2.
[0076] Figure 3 It is a schematic diagram of the generated reconstructed text data described in Embodiment 2.
[0077] Figure 4 It is a comparison schematic diagram of the graphic-text pair sample loss distribution and the two-component Gaussian mixture model predicted noise probability described in Embodiment 2.
[0078] Figure 5 It is a schematic structural diagram of a cross-modal retrieval system described in Embodiment 3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0080] To better illustrate this embodiment, some components in the drawings are omitted, enlarged, or reduced, and do not represent the dimensions of the actual product;
[0081] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0082] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0083] Embodiment 1
[0084] This embodiment provides a cross-modal retrieval method, as Figure 1 shown, including:
[0085] S1: Obtain a dataset of image-text pairs, including corresponding image data and text data;
[0086] S2: Construct an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder;
[0087] S3: Randomly cover pixel blocks on the image data to obtain a masked block image; randomly mask the text data to obtain masked text data;
[0088] S4: Input the masked block image and the image data into the vision encoder to obtain the masked block image encoding and the image data encoding, and set an image reconstruction loss function according to the masked block image encoding and the image data;
[0089] S5: Input the image data into a preset vision concept vocabulary library to obtain vision concept words; and input the vision concept words and the masked text data into the text encoder to obtain text encoding enhanced by vision concepts;
[0090] S6: Set an image description loss function according to the text data, the text encoding enhanced by vision concepts, and the image data encoding;
[0091] S7: Input the image data, the text data, and the text encoding enhanced by vision concepts into the cross-modal decoder, generate pure text data encoding according to the text data and the text encoding enhanced by vision concepts, and generate reconstructed text data according to the image data and the text encoding enhanced by vision concepts;
[0092] S8: Calculate the image-text pair noise probability according to the image data encoding and the pure text data encoding, and set a noise adaptive contrast loss function;
[0093] S9: Use the noise probability as the replacement probability, and use the reconstructed text data to replace the corresponding text data according to the replacement probability to obtain reconstructed image-text pair data;
[0094] S10: Construct a total loss function according to the image reconstruction loss function, the noise adaptive contrast loss function, and the image description loss function, and optimize the total loss function using the reconstructed image-text pair data to obtain an optimized vision-language retrieval model;
[0095] S11: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
[0096] In the specific implementation process, this embodiment first obtains a dataset of image-text pairs, where there are defects such as noise interference or incomplete text between the corresponding image data and text data; then constructs an initial vision-language retrieval model, including a vision encoder for maintaining high-quality visual feature representations while reducing computational costs; a text encoder for encoding text data and auxiliary visual concept words; a cross-modal decoder for synthesizing semantically consistent reconstructed text data; constructs a total loss function according to the image reconstruction loss function, the noise adaptive contrast loss function, and the image caption loss function, optimizes the initial vision-language retrieval model, obtains the optimized vision-language retrieval model, uses the image-text pair noise probability for noise adaptive regularization to avoid serious deviations from noise, enhances the robustness of the retrieval model, can effectively avoid overfitting of the retrieval model to the dataset of image-text pairs containing noise, the generated reconstructed text data contains rich and detailed image descriptions, improves the accuracy of the retrieval results, and can also complete the dataset of image-text pairs with incomplete text data.
[0097] Embodiment 2
[0098] This embodiment provides a cross-modal retrieval method, including:
[0099] S1: Obtain a dataset of image-text pairs, including corresponding image data and text data;
[0100] S2: Construct an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder;
[0101] As Figure 2 shown, the vision encoder is used to maintain high-quality visual feature representations while reducing computational costs; the text encoder is used to encode text data and auxiliary visual concept words; the cross-modal decoder is used to synthesize semantically consistent reconstructed text data;
[0102] S3: Randomly cover pixel blocks on the image data to obtain a masked block image; randomly mask the text data to obtain masked text data; specifically:
[0103] Embed an additional [CLS] token into the linearly projected image data, output the [CLS] token to represent the global image feature, and randomly mask image blocks and skip masked tokens during implementation to obtain a masked block image to reduce computational loss.
[0104] S4: Input the masked block image and the image data into the vision encoder to obtain the masked block image encoding and the image data encoding, and set the image reconstruction loss function according to the masked block image encoding and the image data;
[0105] The image reconstruction loss function is:
[0106]
[0107] Wherein, L IR is the image reconstruction loss value, N represents the number of image-text pair data, x i represents the i-th image data, x' i represents the i-th masked block image, V e (x' i ) represents the encoding of the i-th masked block image, and ‖*‖ represents the calculation of the second norm.
[0108] S5: Input the image data into a preset visual concept vocabulary library to obtain visual concept words; and input the visual concept words and the masked text data into a text encoder to obtain text encoding enhanced by visual concepts;
[0109] S6: Set an image description loss function according to the text data, the text encoding enhanced by visual concepts, and the image data encoding; specifically:
[0110] L LM = -E (x,y)~D logp(y t |C d (y τ<t , [V e (x), T e ([q, y m )]))
[0111] Wherein, L LM represents the image description loss value, C d represents the cross-modal decoder, y t represents the t-th word in the text data, q represents the visual concept word, y m is the masked text data, T e is the text encoder
[0112] S7: Input the image data, the text data, and the text encoding enhanced by visual concepts into the cross-modal decoder, generate pure text data encoding according to the text data and the text encoding enhanced by visual concepts, and generate reconstructed text data according to the image data and the text encoding enhanced by visual concepts;
[0113] S8: Calculate the image-text pair noise probability according to the image data encoding and the pure text data encoding, and set a noise adaptive contrast loss function;
[0114] The specific method for calculating the image-text pair noise probability is:
[0115] S8.1: For each group of image-text pairs, calculate the similarity of the image data relative to the text data and the similarity of the text data relative to the image data;
[0116]
[0117] In the formula, represents the similarity of the i-th image data relative to the j-th text data, represents the similarity of the j-th text data relative to the i-th image data.
[0118] S8.2: Using the similarity of the image data relative to the text data and the similarity of the text data relative to the image data of all image-text pairs, calculate the total image-text contrastive learning loss;
[0119]
[0120]
[0121]
[0122]
[0123] In the formula, B represents the number of image pairs input in this batch, represents the similarity of the i-th text data relative to the i-th image data, represents the similarity of the i-th image data relative to the i-th text data; L ITC (x i , y i ) represents the i-th image-text contrastive loss, L ITC represents the total image-text contrastive learning loss; represents the image-to-text contrastive learning loss, represents the text-to-image contrastive learning loss.
[0124] S8.3: Using the two-component Gaussian mixture model, calculate the image-text pair noise probability according to the image-text contrastive loss;
[0125]
[0126] ∈ i = p(μ h )p(L ITC (x i , y i )|μ h ) / p(L ITC (x i , y i ))
[0127] In the formula, p(*) represents obtaining the probability distribution, θ represents the two-component Gaussian mixture model parameter, m represents the number of components of the two-component Gaussian mixture model, γ mThe mixing covariance of the m-th two-component Gaussian mixture model component, where φ(*|m) represents the probability density of obtaining the m-th two-component Gaussian mixture model component, and μ h represents the high-mean component, and ∈ i represents the noise probability of the i-th image-text pair.
[0128] To align different modalities, existing vision-language pre-training models use image-text contrastive loss to ensure that the positive image-text pairs {x i , y i} i=y are aligned in the same feature space, while the negative image-text pairs {x i , y i} i≠j are the opposite; however, the existing image-text contrastive loss forces the vision-language pre-training model to align the features of each image-text pair without considering the case where there is noise, reducing the model performance; the present invention coordinates cross-modal alignment to different degrees according to the image-text pair noise probability, and the image-text pair noise probability indicates the degree of semantic mismatch between the image-text pair; the present invention preferentially fits clean image-text pairs, followed by image-text pairs containing noise, that is, uses a two-component Gaussian mixture model to fit the image-text contrastive loss of each image-text pair.
[0129] The specific method for setting the noise-adaptive contrastive loss function is as follows:
[0130] Set the image-text pair smoothing rate w i :
[0131] w i = λ∈ i
[0132] where w i represents the smoothing rate of the i-th image-text pair, λ represents a hyperparameter, and ∈ i represents the noise probability of the i-th image-text pair;
[0133] Calculate and
[0134]
[0135]
[0136] In the formula, B represents the number of image pairs; represents the noise-adaptive image-to-text contrastive learning loss, represents the noise-adaptive text-to-image contrastive learning;
[0137] Calculate the noise-adaptive contrastive loss function:
[0138]
[0139] In the formula, L NITC represents the noise adaptive contrast loss value.
[0140] According to the noise probability of the image-text pair, the image-text pair smoothing rate is introduced to regularize the true alignment label to different degrees. For the image-text pair data with low noise probability, lower regularization is used to learn the alignment, and for the image-text pair data with high noise probability, higher regularization is used to learn the alignment, which can effectively avoid overfitting to noise.
[0141] S9: Use the noise probability as the replacement probability, and replace the corresponding text data with the reconstructed text data according to the replacement probability to obtain the reconstructed image-text pair data;
[0142] S10: Construct the total loss function according to the image reconstruction loss function, the noise adaptive contrast loss function and the image description loss function, and optimize the total loss function with the reconstructed image-text pair data to obtain the optimized visual-language retrieval model;
[0143] The specific form of the total loss function is:
[0144] L = L IR + α·L LM + β·L NITC
[0145] In the formula, L represents the total loss function value, α represents the first weight coefficient, and β represents the second weight coefficient.
[0146] S11: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain the retrieval result.
[0147] In the specific implementation process, the model training and optimization process is divided into three stages: noise-aware pre-training, text generation, and concept-enhanced pre-training. In the noise-aware pre-training stage, first, the model is preheated through E IR 、L LM and L ITC supervision for E e iterative rounds. Then, based on the noise probability ∈ ITC of the i-th image-text pair estimated by L i , in the next E t iterative rounds, L ITC is replaced with L NITC to implement noise adaptive regularization. In the text generation stage, in order to obtain better generation ability, the visual concept words are retrieved to supplement the description of the text for the image, and the reconstructed text data is generated for each image-text pair, and the noise probability ∈ iAs the replacement probability, the original text data is replaced with the reconstructed text data to form a modified image-text pair dataset; in the concept enhancement pre-training stage, within E f iterations, the modified image-text pair dataset is used to optimize the vision-language retrieval model to obtain an optimized vision-language retrieval model. As shown in Figure 3 , it is a schematic diagram of generating the reconstructed text data in this embodiment; in the figure, x represents the image data, y represents the text data, q represents the visual concept word, and y' represents the reconstructed text data. It can be seen from the figure that compared with the original text data, the reconstructed text data shows a more specific meaning and contains more semantic information; at the same time, integrating the visual concept word into the cross-modal decoder helps to enrich the synthetic data and increase the description information of the existing objects. As shown in Figure 4 , it is a comparison schematic diagram of the sample loss distribution of the image-text pair and the predicted noise probability of the two-component Gaussian mixture model; among them, Figure 4 a shows three examples with different distribution positions and predicted noise probabilities, Figure 4 b-4d are three image-text pairs at the corresponding positions; it can be seen from the figure that as the sample loss of the image-text pair increases, the noise probability of the image-text pair also increases, and the correlation also increases;
[0148] Embodiment 3
[0149] This embodiment provides a cross-modal retrieval system. Based on the cross-modal retrieval method described in Embodiment 1 or 2, as shown in Figure 5 , it includes:
[0150] A data acquisition module for acquiring an image-text pair dataset, including corresponding image data and text data;
[0151] A model construction module for constructing an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder;
[0152] A masking module for randomly covering pixel blocks on the image data to obtain a masked block image; and randomly masking the text data to obtain masked text data;
[0153] An image encoding module for inputting the masked block image and the image data into the vision encoder to obtain the image data encoding and the masked block image encoding, and setting an image reconstruction loss function according to the masked block image encoding and the image data;
[0154] A visual concept enhancement module for inputting the image data into a preset visual concept vocabulary library to obtain visual concept words; and inputting the visual concept words and the masked text data into the text encoder to obtain text encoding with enhanced visual concepts;
[0155] An image description loss function setting module, configured to set an image description loss function according to text data, text encoding enhanced by visual concepts, and image data encoding;
[0156] A pure text and reconstructed text generation module, configured to input image data, text data, and text encoding enhanced by visual concepts into a cross-modal decoder to obtain pure text data encoding and reconstructed text data;
[0157] A noise adaptation module, configured to calculate the image-text pair noise probability according to image data encoding and pure text data encoding, and set a noise adaptation contrast loss function;
[0158] An image-text pair data reconstruction module, configured to use the noise probability as a replacement probability, and replace the corresponding text data with the reconstructed text data according to the replacement probability to obtain reconstructed image-text pair data;
[0159] A model optimization module, configured to construct a total loss function according to the image reconstruction loss function, the noise adaptation contrast loss function, and the image description loss function, and optimize the total loss function by using the reconstructed image-text pair data to obtain an optimized visual-language retrieval model;
[0160] A cross-modal retrieval module, configured to input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
[0161] Identical or similar reference numerals correspond to identical or similar components;
[0162] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0163] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A cross-modal retrieval method, characterized in that, Including: S1: Obtain an image-text pair dataset, including corresponding image data and text data; S2: Construct an initial vision-language retrieval model, including a vision encoder, a text encoder, and a cross-modal decoder; S3: Randomly cover pixel blocks on the image data to obtain a masked block image; Randomly mask the text data to obtain masked text data; S4: Input the masked block image and the image data into the vision encoder to obtain the masked block image encoding and the image data encoding, and set an image reconstruction loss function according to the masked block image encoding and the image data; S5: Input the image data into a preset vision concept vocabulary library to obtain vision concept words; and input the vision concept words and the masked text data into the text encoder to obtain text encoding enhanced with vision concepts; S6: Set an image captioning loss function according to the text data, the text encoding enhanced with vision concepts, and the image data encoding; S7: Input the image data, the text data, and the text encoding enhanced with vision concepts into the cross-modal decoder, generate pure text data encoding according to the text data and the text encoding enhanced with vision concepts, and generate reconstructed text data according to the image data and the text encoding enhanced with vision concepts; S8: Calculate the image-text pair noise probability according to the image data encoding and the pure text data encoding, and set a noise adaptive contrast loss function; S9: Use the noise probability as the replacement probability, and use the reconstructed text data to replace the corresponding text data according to the replacement probability to obtain reconstructed image-text pair data; S10: Construct a total loss function according to the image reconstruction loss function, the noise adaptive contrast loss function, and the image captioning loss function, and optimize the total loss function using the reconstructed image-text pair data to obtain an optimized vision-language retrieval model; S11: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
2. The cross-modal retrieval method according to claim 1, characterized in that, In step S4, the image reconstruction loss function is: where L IR is the image reconstruction loss value, N represents the number of text-image pair data, x i represents the i-th image data, x′ i represents the i-th masked block image, V e (x′ i ) represents the encoding of the i-th masked block image, and ‖*‖ represents the calculation of the second norm.
3. The cross-modal retrieval method according to claim 1, characterized in that, In step S8, the specific method for calculating the image-text pair noise probability according to the image data encoding and the pure text data encoding is: S8.1: For each group of image-text pairs, calculate the similarity of the image data relative to the text data and the similarity of the text data relative to the image data; S8.2: Use the similarity of the image data relative to the text data and the similarity of the text data relative to the image data of all image-text pairs to calculate the total image-text contrastive learning loss; S8.3: Use a two-component Gaussian mixture model to calculate the image-text pair noise probability according to the image-text contrastive loss.
4. The cross-modal retrieval method according to claim 3, characterized in that, In step S8.1, the specific method for calculating the similarity of the image data relative to the text data and the similarity of the text data relative to the image data is: In the formula, represents the similarity of the i-th image data to the j-th text data, represents the similarity of the j-th text data to the i-th image data.
5. The cross-modal retrieval method according to claim 4, characterized in that, In step S8.2, the specific method for calculating the total image-text contrastive learning loss is: Where B represents the number of image pairs input in this batch, represents the similarity between the i-th text data and the i-th image data, represents the similarity between the i-th image data and the i-th text data; L ITC (x i ,y i ) represents the i-th text-image contrast loss, L ITC represents the total text-image contrast learning loss; represents the image-to-text contrast learning loss, represents the text-to-image contrast learning loss.
6. The cross-modal retrieval method according to claim 5, characterized in that, In step S8.3, the specific method for calculating the image-text pair noise probability according to the image-text contrastive loss is: ∈ i = p(μ h ) p(L ITC (x i , y i ) | μ h ) / p(L ITC (x i , y i )) where p(*) represents obtaining the probability distribution, θ represents the parameters of the two-component Gaussian mixture model, m represents the number of components of the two-component Gaussian mixture model, and γ m is the mixing covariance of the m-th component of the two-component Gaussian mixture model, φ(*|m) represents obtaining the probability density of the m-th component of the two-component Gaussian mixture model, and μ h represents the high-mean component, and ∈ i represents the noise probability of the i-th text-image pair.
7. The cross-modal retrieval method according to claim 6, wherein In step S8, the specific method for setting the noise adaptive contrast loss function according to the image-text pair noise probability is: Set the graphic-text pair smoothing rate w for the noise probability according to the graphic and text i : w i = λ ∈ i where, w i represents the smoothing rate of the i-th text-image pair, λ represents a hyperparameter, and ∈ i represents the noise probability of the i-th text-image pair; Smoothing rate calculation based on graphics and text and where B represents the number of image pairs; represents the noise - adaptive image - to - text contrastive learning loss, represents the noise - adaptive text - to - image contrastive learning; Calculate the noise adaptive contrast loss function: where L NITC represents the noise adaptive contrast loss value.
8. The cross-modal retrieval method according to claim 7, wherein In step S6, the image captioning loss function is specifically: L LM = -E (x,y)~D log p(y t | C d (y τ<t , [V e (x), T e ([q, y m )])) Where L LM represents the image description loss value, C d represents the cross-modal decoder, y t represents the t-th word in the text data, q represents the visual concept word, and y m is the masked text data, and T e is the text encoder.
9. The cross-modal retrieval method according to claim 7, wherein In step S10, the total loss function is specifically: L = L IR + α·L LM + β·L NITC In the formula, L represents the total loss function value, α represents the first weight coefficient, and β represents the second weight coefficient.
10. A cross-modal retrieval system, based on the cross-modal retrieval method according to claims 1-9, wherein The system includes: A data acquisition module for acquiring a graphic-text pair dataset, including corresponding image data and text data; A model construction module for constructing an initial visual-language retrieval model, including a visual encoder, a text encoder, and a cross-modal decoder; A masking module for randomly covering pixel blocks on the image data to obtain a masked block image; and randomly masking the text data to obtain masked text data; An image encoding module for inputting the masked block image and the image data into the visual encoder to obtain image data encoding and masked block image encoding, and setting an image reconstruction loss function according to the masked block image encoding and the image data; A visual concept enhancement module for inputting the image data into a preset visual concept vocabulary library to obtain visual concept words; and inputting the visual concept words and the masked text data into the text encoder to obtain visually concept-enhanced text encoding; An image description loss function setting module for setting an image description loss function according to the text data, the visually concept-enhanced text encoding, and the image data encoding; A pure text and reconstructed text generation module for inputting the image data, the text data, and the visually concept-enhanced text encoding into the cross-modal decoder to obtain pure text data encoding and reconstructed text data; A noise adaptation module for calculating the graphic-text pair noise probability according to the image data encoding and the pure text data encoding, and setting a noise adaptation contrast loss function; A graphic-text pair data reconstruction module for using the noise probability as a replacement probability, and replacing the corresponding text data with the reconstructed text data according to the replacement probability to obtain reconstructed graphic-text pair data; A model optimization module for constructing a total loss function according to the image reconstruction loss function, the noise adaptation contrast loss function, and the image description loss function, and optimizing the total loss function with the reconstructed graphic-text pair data to obtain an optimized visual-language retrieval model; A cross-modal retrieval module for inputting the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
Citation Information
Patent Citations
Object perception image fusion method for multi-modal target tracking
CN112862860A
Audio-visual speech enhancement
US20210134312A1