Training Method of Graphic and Text Retrieval Model Based on OSCAR and Method for Implementing Graphic and Text Retrieval
Through the OSCAR model training method of the graphic search model, negative samples of different difficulties are generated and new loss functions are designed, which solves the problems of generalization ability and matching accuracy of the existing models, and achieves more efficient and accurate graphic search results.
Patent Information
- Application Number
- CN202311395517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-10-26
AI Technical Summary
The existing graphic and text retrieval models have problems such as insufficient generalization ability, heterogeneity differences and semantic differences in image-text matching, resulting in inaccurate matching.
A new graphic and text search model is built based on the OSCAR model. By obtaining image-text sample pairs in the training set, feature extraction and negative sample generation are carried out, and a new loss function is designed to fine-tune the model to improve the generalization ability of the model and the accuracy of graphic and text search.
It improves the generalization ability of the graphic and text retrieval model, improves the accuracy and efficiency of the model in image-text matching, and solves the shortcomings of existing models in heterogeneity and semantic differences.
Smart Images

Figure CN117390213B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular, to a training method for a graphic and text retrieval model based on OSCAR and a method for implementing graphic and text retrieval. Background Art
[0002] The purpose of graphic and text retrieval is to associate a given picture with its corresponding text description, so as to achieve the matching between images and texts. Graphic and text retrieval plays a key role in many important cross-modal tasks, such as semantic image retrieval, image description, visual quality assurance, etc. However, graphic and text matching faces some important challenges, mainly including heterogeneity differences and semantic differences. Heterogeneity differences refer to the inconsistent feature representations of image and text data from different modalities, while semantic differences refer to the misalignment problems that occur when capturing the cross-modal correspondence between images and texts.
[0003] Currently, many studies use pre-training modules such as convolutional neural networks and recurrent neural networks to extract image and text features to bridge the heterogeneity differences. However, the feature extractors in these pre-training modules have not been specially trained or networked with image-text pair data, so they cannot achieve good image or text embedding effects. Another common graphic and text matching method is to use triplet loss to encourage the model to make the similarity score of positive image-text pairs higher than that of negative image-text pairs. However, the existing cost functions do not fully consider the difficulty of negative samples, which is one of the main reasons for the inaccurate graphic and text matching of the model. Some studies have shown that increasing the batch size to obtain more negative samples will lead to a sharp increase in computational complexity, and at the same time, the return on performance improvement will gradually decrease.
[0004] Currently, in the context of visual language tasks, the OSCAR model has very powerful performance. It has been pre-trained on millions of pairs of image-text pairs, jointly processes images and texts to obtain meaningful feature representations, can capture the intricate associations between texts and images, and learn more discriminative image-text embeddings. The OSCAR model has good learning and understanding abilities for the feature representations of images and texts, but the generalization ability of the model is still weak.
[0005] Therefore, the present application constructs a new graphic and text retrieval model based on the OSCAR model to improve the generalization ability of the model and the accuracy and efficiency of graphic and text retrieval of the model. Summary of the Invention
[0006] The purpose of the present invention is to provide a training method for a graphic and text retrieval model based on OSCAR and a method for implementing graphic and text retrieval, which can improve the generalization ability of the model and the accuracy and efficiency of graphic and text retrieval of the model.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] In a first aspect, the present invention provides a training method for an image-text retrieval model based on OSCAR, the method comprising:
[0009] Obtain a training set, the training set including a plurality of image-text sample pairs;
[0010] Input the plurality of image-text sample pairs in the training set into the pre-trained model OSCAR for visual language tasks to perform feature extraction to generate image feature representations and text feature representations;
[0011] Take each sample in the training set as an anchor sample, and based on the image feature representation and the text feature representation, generate a plurality of negative samples with different difficulties corresponding to the anchor sample; the generated negative samples and the anchor sample form generated negative sample pairs;
[0012] Calculate the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair;
[0013] Calculate a loss function based on the positive similarity and the negative similarity, and fine-tune the pre-trained model OSCAR through the loss function to obtain a completed OSCAR image-text retrieval model.
[0014] Further, the step of taking each sample in the training set as an anchor sample and generating a plurality of negative samples with different difficulties corresponding to the anchor sample based on the image feature representation and the text feature representation includes:
[0015] Select a sample as the anchor sample q, the sample being an image sample or a text sample;
[0016] Based on the anchor sample q, perform global semantic clustering on each sample in the training set to obtain a negative sample clustering set G = {g1, g2,..., g M}, where g i = {x i1 , x i2 ,..., x iN} represents a negative sample set of N negative samples with similar semantics, and x ij represents the j-th negative sample in the negative sample set g i , i takes any integer from 1 to M, and j takes any integer from 1 to N;
[0017] Calculate the similarity and the corresponding weight between each negative sample and the anchor sample q based on a kernel function, and perform weighted averaging to obtain a plurality of negative samples with different difficulties.
[0018] Further, calculating the similarity and corresponding weights between each negative sample and the anchor sample q based on a kernel function, and performing weighted averaging to obtain multiple negative samples of different difficulties, including:
[0019] Calculating the similarity between each negative sample and the anchor sample based on the Gaussian radial basis function:
[0020]
[0021] where k represents the similarity between the anchor sample q and the negative sample x jn and ‖·‖ represents the norm distance, and σ is the width parameter;
[0022] Calculating the weight W corresponding to the similarity between each negative sample and the anchor sample according to the following formula n :
[0023] J(W) = min||X - W n ||
[0024] where J(W) is the cost function representing the error in the least squares method, W is the weight matrix to be optimized; X represents the input data matrix, each row represents a negative sample, and each column represents a feature; W n is the weight value of the weight matrix W; ||·|| represents calculating the error;
[0025] Calculating the generated negative sample through weighted averaging:
[0026]
[0027] where represents the generated negative sample corresponding to the anchor sample.
[0028] Further, the loss function can be expressed as:
[0029]
[0030] where v represents the image feature representation, c represents the text feature representation; s vc+ represents the positive similarity when the anchor sample is an image sample, s cv+ represents the positive similarity when the anchor sample is a text sample; S vc represents the set of positive and negative similarities when the anchor sample is an image sample, S cv represents the set of positive and negative similarities when the anchor sample is a text sample; and represent the penalty terms; τ is a hyperparameter; ||·|| represents the set size.
[0031] Further, input multiple image-text sample pairs in the training set into the pre-trained OSCAR image-text retrieval model to perform feature extraction to generate an image feature representation and a text feature representation, including:
[0032] Obtain the image samples in the training set, extract the regional visual features and regional position features of the image samples, and perform a linear combination of the regional visual features and the regional position features to obtain an image embedding; the image samples contain n object regions;
[0033] Obtain the text samples in the training set, use the word segmentation technology to divide the text samples into multiple tokens, and obtain the text embedding corresponding to each token based on the OSCAR-base model;
[0034] Based on the image embedding and the text embedding, use the attention mechanism to generate a joint feature representation, and generate the image feature representation and the text feature representation through average pooling.
[0035] In a second aspect, the present invention also provides a method for implementing image-text retrieval using the OSCAR image-text retrieval model. The OSCAR image-text retrieval model is trained by the above training method. The method includes:
[0036] Obtain the target text and target image to be retrieved;
[0037] Based on the text encoder in the image-text retrieval model, perform feature extraction on the target text to obtain a text feature representation;
[0038] Based on the image encoder in the image-text retrieval model, perform feature extraction on the target image to obtain an image feature representation;
[0039] Based on the text feature representation and the image feature representation, determine the image retrieval result of the target text in the target image, and / or determine the text retrieval result of the target image in the target text.
[0040] In a third aspect, the present invention also provides an OSCAR-based image-text retrieval model training device. The device includes:
[0041] A data acquisition module for acquiring a training set, where the training set includes multiple image-text sample pairs;
[0042] A feature extraction module for inputting multiple image-text sample pairs in the training set into the pre-trained model OSCAR for visual language tasks to perform feature extraction to generate an image feature representation and a text feature representation;
[0043] A negative sample synthesis module, configured to use each sample in the training set as an anchor sample, and based on the image feature representation and the text feature representation, generate multiple negative samples with different difficulties corresponding to the anchor sample; the generated negative samples and the anchor sample form generated negative sample pairs;
[0044] A similarity calculation module, configured to calculate the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair;
[0045] A contrastive loss calculation module, configured to calculate a loss function based on the positive similarity and the negative similarity, and fine-tune the pre-trained model OSCAR through the loss function to obtain a trained OSCAR image-text retrieval model.
[0046] In a fourth aspect, the present invention further provides a computer device, which includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the method as described in any one of the above.
[0047] In a fifth aspect, the present invention further provides a computer-readable storage medium, which stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the method as described in any one of the above.
[0048] The beneficial effects of the present invention are as follows: A training method for an image-text retrieval model based on OSCAR provided by an embodiment of the present invention uses the vision-language pre-trained model OSCAR to extract features from image samples and text samples, generates non-challenging negative samples through a negative sample synthesis module, increases the difficulty between the image and the text, designs a loss function using the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair, and trains the target OSCAR model based on the new loss function, improving the generalization ability of the image-text retrieval model, and further improving the efficiency and accuracy of the model for image-text retrieval.
[0049] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and combines the accompanying drawings to describe in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic flowchart of a training method for an image-text retrieval model based on OSCAR provided by an embodiment of the present invention;
[0051] Figure 2 is a schematic flowchart of a method for implementing image-text retrieval provided by an embodiment of the present invention;
[0052] Figure 3 The structural block diagram of a training device for an image - text retrieval model based on OSCAR provided by an embodiment of the present invention;
[0053] Figure 4 The structural schematic diagram of a computer device provided by an embodiment of the present invention. Specific embodiments
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] In addition, the term "and / or" herein is only used to describe an association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0056] An embodiment of the present application provides an image - text retrieval model training method based on OSCAR. The execution subject of this training method includes, but is not limited to, one of electronic devices such as a server, a terminal, etc. that can be configured to execute this method provided by the embodiments of the present application.
[0057] Please refer to Figure 1 As shown, it is a flowchart of an image - text retrieval model training method based on OSCAR provided by an embodiment of the present invention. In this embodiment, this training method includes:
[0058] Step S101: Obtain a training set, where the training set includes multiple image - text sample pairs.
[0059] In the embodiments of the present invention, the data set can be obtained from a specified open - source natural language learning model corpus, or a large number of image - text pairs can be obtained from a specified website using a Python script with data scraping capabilities. The embodiments of the present invention do not specifically limit the acquisition method of the training set.
[0060] Step S102: Input the multiple image - text sample pairs in the training set into the pre - trained model OSCAR pre - oriented to visual - language tasks to perform feature extraction to obtain image feature representations and text feature representations.
[0061] It is understandable that the pre-trained OSCAR vision-language model has been pre-trained on millions of image-text pairs, capable of jointly processing images and text to obtain meaningful feature representations, capturing the intricate associations between text and images, and learning more discriminative image-text embeddings. That is to say, the OSCAR model has good learning and understanding capabilities for the feature representations of images and text, and can extract richer feature information from images and text.
[0062] Specifically, the steps for generating image feature representations and text feature representations based on the pre-trained OSCAR model include:
[0063] 1) Obtain image samples in the training set, extract the regional visual features and regional position features of the image samples, and perform a linear combination of the regional visual features and regional position features to obtain image embeddings.
[0064] Among them, the image samples are divided into n object regions.
[0065] In one example, use the Faster R-CNN model pre-trained on the Visual Genome dataset to extract the regional visual features and regional position features of the image, and perform a linear combination of the regional visual features and regional position features through linear projection to obtain the image embedding corresponding to each image sample.
[0066] 2) Obtain text samples in the training set, use the tokenization technique to divide the text samples into multiple tokens, and obtain the text embedding corresponding to each token based on the OSCAR-base model.
[0067] In the embodiments of the present invention, for a given text sample c, first use the tokenization technique to divide it into z tokens, that is, c = {o1, o2,..., o z}, and then obtain the text embedding E tok corresponding to each token through the OSCAR-base model, then there is:
[0068]
[0069] Among them, represents the i-th token of the text sample.
[0070] Thus, the text embedding corresponding to the text sample is expressed as:
[0071] 3) Based on the image embedding and the text embedding, use the attention mechanism to generate a joint feature representation, and generate the image feature representation and the text feature representation through average pooling.
[0072] In this embodiment, the obtained image embedding and text embedding are input into a single Transformer model in the OSCAR vision-language model to obtain the joint feature representation of the image and text. Then, the local features of the image and the local features of the text are mapped into global features with a smaller dimension through average pooling, and the average information of the features is retained to generate the image feature representation and the text feature representation. Among them, the Transformer model is based on the attention mechanism, captures the complex relationships between image and text elements, and obtains the joint feature representation of the image-text pair based on the mutual relationship between the image and the text.
[0073] Step S103: Take each sample in the training set as an anchor sample, and generate multiple negative samples with different difficulties corresponding to the anchor sample based on the image feature representation and the text feature representation. Among them, the generated negative samples and the anchor sample form the generated negative samples.
[0074] In the embodiment of the present invention, considering that during the training process of the image-text retrieval model, the diversity of samples will affect the retrieval effect of the model. By designing a negative sample synthesis module to generate negative samples with different difficulties and training the model with challenging negative samples, the generalization ability of the model can be improved.
[0075] Among them, the step of generating multiple negative samples with different difficulties corresponding to the anchor sample includes:
[0076] 1) Select a sample as the anchor sample q, and the sample is an image sample or a text sample.
[0077] In the following embodiments of this article, taking the anchor sample q as an image sample as an example for specific illustration.
[0078] 2) Based on the anchor sample q, perform global semantic clustering on each sample in the training set to obtain the negative sample clustering set G = {g1, g2,..., g M}, where g i = {x i1 , x i2 ,..., x iN} represents a negative sample set of N negative samples with similar semantics, and x ij represents the jth negative sample in the negative sample set g i , i takes any integer from 1 to M, and j takes any integer from 1 to N.
[0079] Specifically, select negative samples that do not match the anchor sample in the mini-batch of the training set, perform the k-means algorithm on these negative samples, and divide these negative samples into multiple different negative sample sets according to semantics. These negative sample sets form the final negative sample clustering set G = {g1, g2,..., g M}, where each element in the clustering set G represents a group of semantically similar negative samples. Among them, the number of the negative sample set is determined by the parameter k, which is usually specified before the algorithm is executed.
[0080] 3) Calculate the similarity between each negative sample and the anchor sample q based on the kernel function and the corresponding weights, and perform weighted averaging to obtain multiple negative samples with different difficulties.
[0081] In the embodiment of the present invention, the adopted kernel function is the Gaussian radial basis function. Specifically, the steps of calculating the similarity between each negative sample and the anchor sample q based on the kernel function and the corresponding weights, and performing weighted averaging to obtain multiple negative samples with different difficulties include:
[0082] 1) Calculate the similarity between each negative sample and the anchor sample based on the Gaussian radial basis function:
[0083]
[0084] where k represents the similarity between the anchor sample q and the negative sample x in , ||·|| represents the norm distance, and σ is the width parameter;
[0085] 2) Calculate the weight w corresponding to the similarity between each negative sample and the anchor sample according to the following formula n :
[0086] J(W) = min||X - W||
[0087] where J(w) represents the cost function representing the error in the least squares method, w represents the weight matrix to be optimized; X represents the input data matrix, each row represents a sample, and each column represents a feature; ||·|| represents the norm distance;
[0088] In the embodiment of the present application, the weight matrix is optimized by the least squares method. The goal of the least squares method is to adjust the parameter matrix
[0089] 3) Calculate the generated negative sample through weighted averaging:
[0090]
[0091] where represents the generated negative sample corresponding to the anchor sample.
[0092] It can be understood that if the anchor sample is an image sample v, the generated negative sample is a text negative sample If the anchor sample is a text sample c, the generated negative sample is an image sample
[0093] Step S104: Calculate the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair.
[0094] The samples that match the anchor sample are positive samples. A positive sample and the anchor sample form a positive sample pair, and the positive similarity between the image and the text in the positive sample pair is calculated. The samples that do not match the anchor are negative samples. A negative sample and the anchor sample form a negative sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated sample pair is calculated. The positive similarity and the negative similarity are combined to form the third similarity.
[0095] Step S105: Calculate the loss function based on the positive similarity and the negative similarity, and fine-tune the pre-trained model OSCAR through the loss function to obtain the completed and trained OSCAR image-text retrieval model.
[0096] The embodiment of the present invention proposes a brand-new loss function based on InfoCMR for comparing positive and negative samples from different sources. The loss function can be expressed by the formula:
[0097]
[0098] where, v represents the image feature representation, and c represents the text feature representation; s vc+ represents the positive similarity when the anchor sample is an image sample, and s cv+ represents the positive similarity when the anchor sample is a text sample; S vc represents the set of positive similarity and negative similarity when the anchor sample is an image sample, and S cv represents the set of positive similarity and negative similarity when the anchor sample is a text sample; and represent penalty terms; τ is a hyperparameter; ||·|| represents the set size.
[0099] Among them, in order to reduce the risk of model overfitting, an additional penalty term is introduced. Z Gaussian noise vectors are randomly sampled from the Gaussian distribution. Each vector has the same dimension as the anchor vector corresponding to the anchor sample in the embedding space. These Gaussian noise vectors form high-confidence negative sample pairs with each sample in the batch, which helps to smooth the representation space. It should be noted that these Gaussian noise vectors do not participate in the formation of positive sample pairs.
[0100] The loss function designed by the present invention integrates the information among positive samples, negative samples, and generated negative samples, and further reduces the heterogeneity difference in image-text matching. At the same time, an additional penalty term is added to the loss function. By randomly sampling Gaussian noise vectors, high-confidence negative sample pairs are formed, reducing the risk of overfitting, helping to smooth the representation control, and improving the generalization ability of the model.
[0101] A training method for an image-text retrieval model based on OSCAR provided by an embodiment of the present invention uses the vision-language pre-training model OSCAR to extract features from image samples and text samples, generates unchallenging negative samples through a negative sample synthesis module, increases the difficulty between images and texts, designs a loss function using the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair, and trains the target OSCAR model based on the new loss function, improving the generalization ability of the image-text retrieval model, and further improving the efficiency and accuracy of the model for image-text retrieval.
[0102] Please refer to Figure 2 , which is a schematic flowchart of an image-text retrieval method implemented by an image-text retrieval model trained by the method as described above provided by an embodiment of the present invention. The method includes:
[0103] Obtain a target text and a target image to be retrieved;
[0104] Extract features from the target text based on the text encoder in the image-text retrieval model to obtain a text feature representation;
[0105] Extract features from the target image based on the image encoder in the image-text retrieval model to obtain an image feature representation;
[0106] Based on the text feature representation and the image feature representation, determine the image retrieval result of the target text in the target image, and / or determine the text retrieval result of the target image in the target text.
[0107] Using the image-text retrieval method provided by an embodiment of the present invention can improve the efficiency and accuracy of image-text retrieval.
[0108] Please refer to Figure 3 , which is a structural block diagram of a training device for an image-text retrieval model based on OSCAR provided by an embodiment of the present invention. The device includes:
[0109] A data acquisition module 310, configured to acquire a training set, where the training set includes a plurality of image-text sample pairs;
[0110] A feature extraction module 320, configured to input the plurality of image-text sample pairs in the training set into the pre-training model OSCAR for visual language tasks to perform feature extraction to generate an image feature representation and a text feature representation;
[0111] A negative sample synthesis module 330, configured to use each sample in the training set as an anchor sample, and based on the image feature representation and the text feature representation, generate a plurality of negative samples with different difficulties corresponding to the anchor sample; the generated negative samples and the anchor sample form a generated negative sample pair;
[0112] A similarity calculation module 340 is configured to calculate the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair.
[0113] A contrastive loss calculation module 350 is configured to calculate a loss function based on the positive similarity and the negative similarity, and fine-tune the pre-trained model OSCAR through the loss function to obtain a completed trained OSCAR image-text retrieval model.
[0114] Please refer to Figure 4 , which is a schematic structural diagram of a computer device provided by an embodiment of the present invention. The computer device may include a memory 20, a processor 21, and a bus, and may also include a computer program stored in the memory 20 and executable on the processor 21.
[0115] Among them, the memory 20 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 20 may be an internal storage unit of the computer device in some embodiments, such as the mobile hard disk of the computer device. The memory 20 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory 20 may include both an internal storage unit and an external storage device of the computer device. The memory 20 can be used not only to store application software and various types of data installed in the computer device, but also to temporarily store data that has been output or will be output.
[0116] The processor 21 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 21 is the control core (Control Unit) of the computer device, connecting various components of the entire electronic device through various interfaces and lines, and executing various functions of the computer device and processing data by running or executing programs or modules stored in the memory 20 and calling data stored in the memory 20.
[0117] The bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to implement connection communication between the memory 20 and at least one processor 21, etc.
[0118] Figure 4 Only a computer device with components is shown. Those skilled in the art can understand that Figure 4 the shown structure does not constitute a limitation on the computer device, and it can include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0119] For example, although not shown, the computer device can also include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor 21 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source can also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The computer device can also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0120] Furthermore, the computer device can also include a network interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device and other computer devices.
[0121] Optionally, the computer device can also include a user interface. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device and to display a visual user interface.
[0122] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0123] If the modules / units integrated in a computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).
[0124] The present invention also provides a computer-readable storage medium storing a computer program, which is executed by a processor of an electronic device.
[0125] In addition, in various embodiments of the present invention, each functional module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0126] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0127] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.
[0128] Embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0129] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as second are used to represent names and do not represent any specific order.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A training method for a graphic and text retrieval model based on OSCAR, characterized in that, The method includes: Obtaining a training set, where the training set includes multiple image-text sample pairs; Inputting the multiple image-text sample pairs in the training set into the pre-trained model OSCAR for visual language tasks to perform feature extraction to obtain image feature representations and text feature representations; Taking each sample in the training set as an anchor sample, and based on the image feature representation and the text feature representation, generating multiple negative samples with different difficulties corresponding to the anchor sample; the generated negative samples and the anchor sample form generated negative sample pairs; Calculating the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair; Calculating a loss function based on the positive similarity and the negative similarity, and fine-tuning the visual language pre-trained model OSCAR through the loss function to obtain a completed OSCAR image-text retrieval model; The step of taking each sample in the training set as an anchor sample and generating multiple negative samples with different difficulties corresponding to the anchor sample based on the image feature representation and the text feature representation includes: Selecting a sample as the anchor sample q, where the sample is an image sample or a text sample; Based on the anchor sample q, perform global semantic clustering on each sample in the training set to obtain a negative sample clustering set G = {g1, g2, …, g M}, where g i = {x i1 , x i2 , …, x iN} represents a negative sample set of N negative samples with similar semantics, x ij represents the j-th negative sample in the negative sample set g i , i takes any integer from 1 to M, and j takes any integer from 1 to N; Calculating the similarity and the corresponding weight between each negative sample and the anchor sample q based on a kernel function, and performing weighted averaging to obtain multiple negative samples with different difficulties; The step of calculating the similarity and the corresponding weight between each negative sample and the anchor sample q based on a kernel function and performing weighted averaging to obtain multiple negative samples with different difficulties includes: Calculating the similarity between each negative sample and the anchor sample based on a Gaussian radial basis function: where k represents the similarity between the anchor sample q and the negative sample x jn ||·|| represents the norm distance, and σ is the width parameter; Calculate the weight W corresponding to the similarity between each negative sample and the anchor sample according to the following formula n : J(W) = min ||X - W n || Among them, J(W) is the cost function representing the error in the least squares method, and W is the weight matrix to be optimized; X represents the input data matrix, where each row represents a negative sample and each column represents a feature; W n is the weight value of the weight matrix W; ||·|| represents the calculation of the error; Calculating the generated negative samples through weighted averaging: Among them, represents the generated negative sample corresponding to the anchor sample; If the anchor sample is the image sample v, the generated negative sample is the text negative sample If the anchor sample is the text sample c, the generated negative sample is the image sample 2. The training method according to claim 1, characterized in that, The loss function is expressed as: Among them, v represents the image feature representation, and c represents the text feature representation; s vc+ represents the positive similarity when the anchor sample is an image sample, and s cv+ represents the positive similarity when the anchor sample is a text sample; S vc represents the set of positive and negative similarities when the anchor sample is an image sample, and S cv represents the set of positive and negative similarities when the anchor sample is a text sample; and represent the penalty term; τ is a hyperparameter; ||·|| represents the set size.
3. The training method according to claim 1, wherein The step of inputting the multiple image-text sample pairs in the training set into the pre-trained OSCAR image-text retrieval model to perform feature extraction to generate image feature representations and text feature representations includes: Obtaining the image samples in the training set, extracting the regional visual features and regional position features of the image samples, and performing a linear combination of the regional visual features and the regional position features to obtain an image embedding; the image sample contains n object regions; Obtaining the text samples in the training set, dividing the text samples into multiple tokens by using a tokenization technique, and obtaining the text embedding corresponding to each token based on the OSCAR-base model; Based on the image embedding and the text embedding, using an attention mechanism to generate a joint feature representation, and generating the image feature representation and the text feature representation through average pooling.
4. A method for implementing image-text retrieval using the OSCAR image-text retrieval model, where the OSCAR image-text retrieval model is trained by the training method described in any one of claims 1 to 3, and is characterized in that, The method includes: Obtaining a target text and a target image to be retrieved; Performing feature extraction on the target text based on the text encoder in the image-text retrieval model to obtain a text feature representation; Performing feature extraction on the target image based on the image encoder in the image-text retrieval model to obtain an image feature representation; Based on the text feature representation and the image feature representation, determining the image retrieval result of the target text in the target image, and / or determining the text retrieval result of the target image in the target text.
5. A training device using the OSCAR-based graphic and text retrieval model training method according to any one of claims 1-3, characterized in that, The device includes: A data acquisition module, configured to acquire a training set, where the training set includes a plurality of image-text sample pairs; A feature extraction module, configured to input the plurality of image-text sample pairs in the training set into a pre-trained model OSCAR for visual language tasks to perform feature extraction to obtain an image feature representation and a text feature representation; A negative sample synthesis module, configured to use each sample in the training set as an anchor sample, and based on the image feature representation and the text feature representation, generate a plurality of negative samples with different difficulties corresponding to the anchor sample; the generated negative samples and the anchor sample form generated negative sample pairs; A similarity calculation module, configured to calculate the positive similarity between the image and the text in the positive sample pair, and the negative similarity between the image and the text in the negative sample pair and the generated negative sample pair; A contrastive loss calculation module, configured to calculate a loss function based on the positive similarity and the negative similarity, and fine-tune the pre-trained model OSCAR through the loss function to obtain a completed OSCAR image-text retrieval model.
6. A computer device, characterized in that, The computer device includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Image sample self-adaptive depth metric learning method
CN111144466A
Method and system for rapid retrieval of target images based on artificial intelligence
US20230134531A1