Image-text retrieval model acquisition method and device and storage medium
By generating mask diagrams and segmenting image blocks, the problem of high consumption of training resources of existing graphic and text retrieval models is solved, and more efficient model training and more accurate retrieval effects are achieved.
Patent Information
- Application Number
- CN202510069306.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The training of existing graphic and text retrieval models requires a large amount of computing resources, resulting in high training costs, long time and low training efficiency.
By generating a mask map of the target image, segmenting the image blocks and sorting them, removing the background image blocks, and training the FLIP model with the remaining image blocks to obtain the graphic and text retrieval model.
It effectively reduces the scale of the model, reduces the computing resources required for training, and improves the accuracy and training efficiency of the model.
Smart Images

Figure CN120067378A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of graphic and text retrieval. Specifically, the present application relates to a method, device, and storage medium for obtaining a graphic and text retrieval model. Background Art
[0002] Graphic and text retrieval is to retrieve a corresponding target image (such as a pedestrian image) by inputting a text description. Compared with traditional image retrieval or combined with a fixed keyword retrieval method, graphic and text retrieval is more open and more in line with human abstract thinking. Graphic and text retrieval can be widely used in fields such as security monitoring and intelligent traffic management to quickly find the captured pictures of specific objects (such as pedestrians and specific vehicles) through natural language description, and obtain the time and location of their appearance.
[0003] Since natural language and images belong to different modalities, and there are differences in their representation forms and semantic structures, most current graphic and text retrieval methods directly use a graphic and text comparison model with a very large number of parameters to fuse the language description and image of pedestrians, resulting in the need to use a large amount of computing resources during training, high training costs, long training time, and low training efficiency. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, and storage medium for obtaining a graphic and text retrieval model, which can solve the problems that existing model training requires a large amount of computing resources, high training costs, long time consumption, and low training efficiency.
[0005] To achieve this purpose, the embodiments of the present application provide the following several solutions.
[0006] According to one aspect of the embodiments of the present application, a method for obtaining a graphic and text retrieval model is provided, including:
[0007] Generating a mask image corresponding to the target image, segmenting the target image and the mask image to obtain a first image block corresponding to the target image and a second image block corresponding to the mask image, the first image block corresponding to the second image block, and the target image including the target object corresponding to the graphic and text retrieval model;
[0008] Sorting the second image blocks, and removing some second image blocks based on the sorting result and the number of background image blocks in the second image blocks;
[0009] Obtaining a target first image block corresponding to the remaining second image blocks, and training a FLIP model with the target first image blocks to obtain a graphic and text retrieval model.
[0010] In a possible implementation manner, the generating a mask image corresponding to the target image includes:
[0011] Scale the target image in the dataset to a preset size, where the preset size includes the length and width of the target image, and the length and width are integer multiples of a predetermined value;
[0012] Input the target image of the preset size into a preset segmentation model to obtain the mask image, and preprocess the mask image.
[0013] In a possible implementation, the number of the first images is the same as that of the second images, and the number of the first image blocks is determined according to the preset size and the predetermined value.
[0014] In a possible implementation, the sorting of the second image blocks includes:
[0015] Obtain the number of preset pixels in each second image block, and sort the second image blocks based on the number. The preset pixels are used to indicate the target object.
[0016] In a possible implementation, the removing of some second image blocks based on the sorting result and the number of background image blocks in the second image blocks includes:
[0017] Obtain the number of background image blocks in the second image blocks, and remove some second image blocks according to the comparison result between the number of background image blocks and a predetermined threshold, the sorting result, and a predetermined removing number.
[0018] In a possible implementation, the predetermined threshold is half of the number of the second image blocks, and the removing number of the second image blocks is equal to the predetermined threshold.
[0019] In a possible implementation, the training of the FLIP model using the target first image blocks includes:
[0020] Generate a training dataset using the target first image blocks, where each training batch in the training dataset includes a picture and text;
[0021] Train the FLIP model based on the training dataset, and during the training process, determine the labels output by the picture and the text according to a preset label assignment mechanism.
[0022] In a possible implementation, the determination of the label includes:
[0023] Calculate the label through formula (1), and the formula (1) is:
[0024] where P i,j is the label output by the picture with index i and the text with index j, and ε i,iIt is determined according to the ratio of the number of preset pixels in the second image block removed from the mask image corresponding to the image with index i to the total number of pixels in the mask image. It is determined based on the similarity between the text corresponding to index i and the text corresponding to index j.
[0025] According to one aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described above.
[0026] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the steps of the method described above are implemented.
[0027] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:
[0028] The method for obtaining a graphic and text retrieval model provided by the embodiments of the present application generates a mask image corresponding to a target image, divides the target image and the mask image to obtain a first image block corresponding to the target image and a second image block corresponding to the mask image. The first image block corresponds to the second image block. The target image includes a target object corresponding to the graphic and text retrieval model; sorts the second image blocks, removes some second image blocks based on the sorting result and the number of background image blocks in the second image blocks; obtains the target first image blocks corresponding to the remaining second image blocks, and uses the target first image blocks to train the FLIP model to obtain the graphic and text retrieval model. The embodiments of the present application use FLIP for model training, effectively reducing the model scale, effectively reducing the information coding amount of the images used in the training of the graphic and text retrieval model, thereby effectively reducing the computing resources required for training, and being able to extract key second image blocks through the mask image to participate in model training, reducing the influence of background factors on the model retrieval effect, while reducing the training cost, improving the accuracy of the model and reducing the training time-consuming, and improving the model training efficiency. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.
[0030] Figure 1 It is a flowchart of the method for obtaining a graphic and text retrieval model provided by the embodiments of the present application;
[0031] Figure 2 It is a schematic diagram of mask image processing provided by the embodiments of the present application;
[0032] Figure 3 It is a schematic diagram of the removal and mapping of the second image blocks provided by the embodiments of the present application;
[0033] Figure 4 This is a structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0034] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the implementation manners described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0035] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here may include a wireless connection or a wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" indicates being implemented as "A", or being implemented as "A", or being implemented as "A and B".
[0036] To make the purpose, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0037] The technical solutions of the embodiments of the present invention and the technical effects produced by the technical solutions of the present invention will be described below through the description of several exemplary implementation manners. It should be noted that the following implementation manners can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different implementation manners, they will not be described repeatedly.
[0038] A method, device, and storage medium for obtaining a graphic and text retrieval model provided by the present application aim to solve at least one technical problem existing in the prior art.
[0039] In the embodiments of the present application, a method for obtaining a graphic and text retrieval model is provided. As Figure 1 shown, the method for obtaining the graphic and text retrieval model includes:
[0040] S101: Generate a mask graph corresponding to the target image, and segment the target image and the mask graph to obtain a first image block corresponding to the target image and a second image block corresponding to the mask graph.
[0041] Optionally, the first image block corresponds to the second image block, and the target image includes a target object corresponding to the text-image retrieval model. The target object may be a pedestrian, an animal, a vehicle, or other objects that require text-image retrieval. Hereinafter, the target object being a pedestrian will be taken as an example for illustration.
[0042] In one embodiment, the number of the first image blocks is the same as that of the second image blocks, and they correspond to each other one by one.
[0043] Optionally, generating a mask map corresponding to the target image includes: scaling the target image in the dataset to a preset size, where the preset size includes the length and width of the target image, and the length and width are integer multiples of a predetermined value; inputting the target image of the preset size into a preset segmentation model to obtain a mask map, and preprocessing the mask map.
[0044] Optionally, the dataset contains target images for model training. When scaling, all the target images in the dataset are scaled to the preset size.
[0045] Optionally, the size of the predetermined value can be determined according to the input requirements for the model or the number of image partitions. Specifically, it can be 16, 15, 18, or other values.
[0046] In one embodiment, the predetermined value can be 16, and the target images in the dataset are uniformly scaled to a pixel size of H×W, where H is the height of the scaled target image, W is the width of the scaled target image, and both H and W are integer multiples of 16.
[0047] Optionally, the target object is a pedestrian, and the preset segmentation model can be Mask RCNN or other models that can recognize the contour of the pedestrian in the target image and generate a mask map according to the recognition result. The generated mask map can be a human-body-background mask map.
[0048] Optionally, the mask map can be a binary map, in which the pixel values of the pixels indicating the target object are different from those of the pixels indicating the background.
[0049] In one embodiment, when the target object is a pedestrian, the segmentation model Mask RCNN can be used to recognize the target image and generate a human-body-background mask map. The pixel values of the pixels indicating the human body part in the mask map are 255, and the pixel values of the pixels indicating the background part are 0.
[0050] Optionally, the preprocessing of the mask map includes removing noise from the mask map and smoothing the mask map.
[0051] In one embodiment, the target object is a pedestrian, such as Figure 2As shown, the background of the target image can be segmented by the segmentation model Mask RCNN to obtain a mask image. A morphological opening operation of erosion followed by dilation is performed on the mask image using a kernel of size 5×5 to remove small noises in the mask image and smooth the mask image. Among them, Figure 2 The left image in the middle is the target image, the middle image is the mask image obtained by background segmentation through the segmentation model, and the right image is the mask image after the opening operation.
[0052] Optionally, the number of the first images is the same as that of the second images, and the number of the first image blocks is determined according to a preset size and a predetermined value.
[0053] Optionally, the calculation formula for the number of the first image blocks can be:
[0054] , where P is the number of the first image blocks, H is the length of the first image after scaling, W is the width of the first image after scaling, and 16 represents the predetermined value. Based on the number of the first image blocks, the mask image and the scaled target image are divided into multiple image blocks, where the length and width of the first image blocks and the second image blocks are the same.
[0055] Optionally, when performing image division to obtain the first image blocks and the second image blocks, record the positions of the first image blocks in the mask image and the positions of the second image blocks in the target image.
[0056] S102: Sort the second image blocks, and remove some of the second image blocks based on the sorting result and the number of background image blocks in the second image blocks.
[0057] Optionally, sorting the second image blocks includes: obtaining the number of preset pixels in each second image block, and sorting the second image blocks based on the number, where the preset pixels are used to indicate the target object.
[0058] Optionally, the second image blocks can be sorted in ascending order of the number of the preset pixels.
[0059] In one embodiment, the preset pixels are the pixels with a pixel value of 255. After obtaining the second image blocks, the second image blocks are sorted in ascending order of the number of preset pixels in each second image block.
[0060] Optionally, removing some of the second image blocks based on the sorting result and the number of background image blocks in the second image blocks includes: obtaining the number of background image blocks in the second image blocks, and removing some of the second image blocks according to the comparison result between the number of background image blocks and a predetermined threshold, the sorting result, and a predetermined removal quantity.
[0061] Optionally, the predetermined threshold can be determined according to the number of second image blocks. Specifically, the predetermined threshold can be half of the number of second image blocks, and the number of second image blocks to be removed can be equal to the predetermined threshold.
[0062] Optionally, the number of second image blocks to be removed can be determined according to the image block removal information of the FLIP model to be trained. By removing the second image blocks first, it is possible to avoid the problem that some important image blocks are removed when randomly removing image blocks during the training of the FLIP model.
[0063] Optionally, the background image blocks can be second image blocks with all pixel values being 0 (i.e., there is no preset pixel with a pixel value of 255). After obtaining the second image blocks, the background image blocks are identified according to the pixel values of the pixels in each second image block, the number of the background image blocks is counted, and the second image blocks to be removed are determined according to the comparison result of the number with the number of second image blocks.
[0064] In one embodiment, as Figure 3 shown, Figure 3 In the left figure in the middle, it is the division result of the mask map. The gray part in the middle figure represents the removed second image blocks, and the right figure is the mapping result of the non-removed second image blocks in the scaled target image. The predetermined threshold can be half of the number of second image blocks, and the number of second image blocks is P. After sorting the second image blocks according to the number of preset pixels, the number of background image blocks is obtained. If the number of background image blocks is greater than then randomly remove background image blocks from the background image blocks; if the number of background image blocks is less than or equal to then according to the sorting result, in the order of increasing number of preset pixels, remove the first second image blocks from the sorting queue of the second image blocks. After removing image blocks, the number of obtained image blocks is half of the original, thereby effectively reducing the number of computing resources required for model training. Among them, after removing image blocks, the number of remaining second image blocks is consistent with the number of image blocks required as input for the FLIP model.
[0065] S103: Obtain the target first image blocks corresponding to the remaining second image blocks, and use the target first image blocks to train the FLIP model to obtain a text-image retrieval model.
[0066] Optionally, when dividing the first image blocks and the second image blocks, the corresponding relationship between the first image blocks and the second image blocks is pre-stored. After removing the second image blocks, the first image blocks corresponding to the remaining second image blocks are determined according to the corresponding relationship, and the first image blocks are determined as the target first image blocks.
[0067] Optionally, training the FLIP model using the target first image patch includes: generating a training dataset using the target first image patch, where each training batch in the training dataset includes an image and text, training the FLIP model based on the training dataset, and during the training process, determining the labels output for the image and text according to a preset label assignment mechanism. Among them, the training method can be self-supervised training.
[0068] Optionally, the training dataset may include multiple training batches. In each training batch, the image and text have a one-to-one correspondence. The image is generated based on the target first image patch, the text is generated based on the text annotation of the target image, and each image and text is assigned a corresponding index, which indicates the correspondence between the image and text. Among them, if the indices of the image and text are the same, it means they correspond; if the indices are different, it means they do not correspond. Determine the labels output for the image and sample according to this correspondence and the preset label assignment mechanism.
[0069] Optionally, the preset label assignment mechanism is:
[0070] Calculate the label through formula (1), and formula (1) is:
[0071] where P i,j is the label output for the image with index i and the text with index j, and ε i,i is determined according to the ratio of the number of preset pixels in the second image patch removed from the mask image corresponding to the image with index i to the total number of pixels in the mask image, and is determined based on the similarity between the text corresponding to index i and the text corresponding to index j.
[0072] Optionally, when calculating ε i,i , the mask image corresponding to the image with index i and the text with index j can be obtained, the second image patch removed from the mask image can be obtained, the number of pixels with a pixel value of 255 (i.e., the pixels indicating the target object) in the removed second image patch can be calculated, and the ratio of this number to the total number of pixels in the mask image can be obtained, and the value of this ratio can be determined as ε i,i , and the value range of ε i,i can be 0 - 0.5.
[0073] Optionally, when calculating , the cosine similarity between the text corresponding to index i and the text corresponding to index j can be calculated first, and this cosine similarity can be normalized to between 0 and 1 to obtain which describes the similarity degree between the two texts.
[0074] In one embodiment, during the model training process, when the image corresponding to index i exactly corresponds to the text with index i, that is, i = j, the output label is P i,i = 1 - ε i,i , ε i,i is the ratio of the number of pixels with a pixel value of 255 in the removed second image patch to the total number of pixels in the mask image, and ε i,i ranges from 0 to 0.5. ε i,i can make the output label negatively correlated with the ratio of the removed preset pixels. In a training batch, each image corresponds to only one text, and the image and the corresponding text form a positive sample, and the rest are negative samples. However, in the retrieval of target images, there is generally a situation where the images are different but the text descriptions are similar. If directly determined as negative samples, the correlation information between samples cannot be fully utilized to improve the model accuracy. Therefore, when the image and the text do not correspond, that is, i ≠ j, the cosine similarity s between the text at index i and the text at index j can be calculated using the publicly available Sentence - BERT model i,j , and then s i,j is normalized to between 0 and 1 to obtain The finally output label is , and its output label is the similarity degree of the two text descriptions
[0075] The method for obtaining the image - text retrieval model provided by the embodiments of the present application improves the target image patch pattern input during the training process. The instance segmentation model is used to divide the target object and the background in the training set to generate a mask image that can describe the branch part of the target object. The mask image is processed using the opening operation of erosion followed by dilation in morphological image processing, and then the mask image is evenly divided into P second image patches. By calculating the size of the mask area on each second image patch (indicating the number of pixels of the target object), sorting is performed, and the first M second image patches with fewer mask areas are obtained. According to the size of M, finally, a selective random removal number of second image patches, and the remaining second image patches are input into the image encoding end of FLIP, and the label assignment mechanism is improved for training. On the premise of retaining the memory - saving characteristics of FLIP, the image encoding part tries to retain the important branch part of the target object as much as possible, while reducing the influence of interference information such as the background, and improving the retrieval accuracy of the target object of the model. Therefore, the present application has the following advantages
[0076] 1. Different from existing image-text retrieval technologies, based on the FLIP model, this application changes the original random masking input method. By using an image segmentation model, a foreground-background mask map of the target object is generated. Combining with image opening operation processing, half of the key image patches of the target object's branches are selected from the original target image and input into the model for training, reducing the influence of background factors on the model retrieval effect and making the model pay more attention to the target object's main body.
[0077] 2. Improve the label assignment mechanism. During the training process, the output label of the real sample is no longer a fixed value of "0" or "1", but is related to the similarity between the input image data and the descriptive text, making the model more generalizable.
[0078] Based on the same inventive concept, an embodiment of this application provides an electronic device, as Figure 4 shown. Figure 4 The electronic device 2000 shown in the figure includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are communicatively connected, such as connected through a bus 2002.
[0079] The processor 2001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field-Programmable Gate Array, field-programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of this application. The processor 2001 can also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0080] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 2002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0081] The memory 2003 can be a ROM (Read-Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (random access memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read-Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0082] Optionally, the electronic device 2000 may further include a communication unit 2004. The communication unit 2004 can be used for receiving and sending signals. The communication unit 2004 can allow the electronic device 2000 to communicate with other devices wirelessly or wiredly to exchange data. It should be noted that in practical applications, the communication unit 2004 is not limited to one.
[0083] Optionally, the electronic device 2000 may further include an input unit 2005. The input unit 2005 can be used for receiving input digital, character, image, and / or sound information, or generating key signal inputs related to the user settings and function controls of the electronic device 2000. The input unit 2005 can include, but is not limited to, one or more of a touch screen, a physical keyboard, function keys (such as volume control buttons, power on / off buttons, etc.), a trackball, a mouse, a joystick, a shooting device, a pickup, etc.
[0084] Optionally, the electronic device 2000 may further include an output unit 2006. The output unit 2006 can be used for outputting or displaying the information processed by the processor 2001. The output unit 2006 can include, but is not limited to, one or more of a display device, a speaker, a vibration device, etc.
[0085] Although the figure shows an electronic device 2000 with various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0086] Optionally, the memory 2003 is used to store a computer program for executing the solution of this application, and is controlled by the processor 2001 to execute. The processor 2001 is used to execute the computer program stored in the memory 2003 to implement the steps of any method provided by the embodiments of this application.
[0087] Based on the same inventive concept, the embodiments of this application provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided by this application / implements the steps of various optional embodiments of the method provided by this application.
[0088] Terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than the illustrated or textually described order.
[0089] It should be understood that although the flowchart of the embodiments of this application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of this application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of this application do not limit this.
[0090] The above are only optional implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.
Claims
1. A method for obtaining a graph-text retrieval model, characterized in that: include: Generate a mask image corresponding to a target image, segment the target image and the mask image to obtain a first image block corresponding to the target image and a second image block corresponding to the mask image, wherein the first image block corresponds to the second image block, and the target image includes a target object corresponding to the image-text retrieval model; sorting the second image blocks, and removing some of the second image blocks based on the sorting result and the number of background image blocks in the second image blocks; A target first image block corresponding to the remaining second image block is obtained, and a FLIP model is trained using the target first image block to obtain an image-text retrieval model.
2. The method for acquiring a graph-text retrieval model according to claim 1, characterized in that: The generating of a mask image corresponding to the target image comprises: Scaling the target image in the data set to a preset size, where the preset size includes the length and width of the target image, and the length and width are integer multiples of a predetermined value; The target image of a preset size is input into a preset segmentation model to obtain the mask image, and the mask image is preprocessed.
3. The method for acquiring a graph-text retrieval model according to claim 2, characterized in that: The number of the first image and the number of the second image are the same, and the number of the first image blocks is determined according to the preset size and the predetermined value.
4. The method for acquiring a graph-text retrieval model according to claim 1, characterized in that: The step of sorting the second image blocks comprises: The number of preset pixels in each second image block is obtained, and the second image blocks are sorted based on the number, wherein the preset pixels are used to indicate the target object.
5. The method for acquiring a graph-text retrieval model according to claim 1, characterized in that: The removing part of the second image blocks based on the sorting result and the number of background image blocks in the second image block comprises: The number of background image blocks in the second image block is obtained, and part of the second image blocks are removed according to a comparison result between the number of background image blocks and a predetermined threshold, the sorting result, and a predetermined removal number.
6. The method for acquiring a graph-text retrieval model according to claim 5, characterized in that: The predetermined threshold is half of the number of the second image blocks, and the number of the second image blocks removed is equal to the predetermined threshold.
7. The method for acquiring a graph-text retrieval model according to claim 4, characterized in that: The step of training the FLIP model using the target first image block includes: Generate a training data set using the target first image block, wherein each training batch in the training data set includes images and texts; The FLIP model is trained based on the training data set, and during the training process, labels of the image and the text output are determined according to a preset label assignment mechanism.
8. The method for acquiring a graph-text retrieval model according to claim 7, characterized in that: The preset label allocation mechanism is: The label is calculated by formula (1), which is: Among them, P i,j The label output for the image with index i and the text with index j, ε i,i is determined according to the ratio of the number of preset pixels in the second image block shifted out of the mask image corresponding to the picture with index i to the total number of pixels in the mask image, It is determined based on the similarity between the text corresponding to index i and the text corresponding to index j.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method according to any one of claims 1 to 8 are implemented.