Method for constructing retrieval data set based on video source data set and electronic equipment
By constructing a retrieval dataset based on the video source dataset and using the coherence in the video to construct consistent image pairs, the inconsistency problem in the zero-sample combined image retrieval dataset is solved, and the retrieval performance and accuracy of the model are improved.
Patent Information
- Application Number
- CN202511134324.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-13
AI Technical Summary
In existing zero-shot combinatorial image retrieval datasets, there is a lack of consistency between image pairs, which leads to interference during model training and reduces the accuracy and reliability of the retrieval model.
By constructing a retrieval dataset method based on a video source dataset, a reference image and its corresponding target image are determined from the same video data, and highly consistent image pairs are constructed using the coherence in the video. The target image corresponding to the reference image is constructed as a positive sample set, and other images are constructed as a negative sample set to provide examples of incorrect queries.
It effectively reduces the problem of visual and semantic inconsistency in traditional retrieval datasets, improves the retrieval performance of the model, provides a real and accurate retrieval model foundation, and improves the accuracy of image retrieval tasks.
Smart Images

Figure CN120726422A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of multimedia technology, and more particularly, to a method and electronic device for constructing a retrieval dataset based on a video source dataset. Background Art
[0002] In the field of image retrieval, the task of composed image retrieval (CIR) based on text descriptions has made significant progress, especially in the zero-shot composed image retrieval (ZS-CIR) task. Zero-shot composed image retrieval aims to address an important challenge: retrieving relevant images (called target images) from an image library using a query consisting of a reference image and a corresponding descriptive text, in the absence of explicit training samples. Although a large number of studies in recent years have improved the performance of CIR, especially through the use of deep learning and cross-modal learning methods, existing zero-shot composed image retrieval datasets still have some serious problems that affect the performance and credibility of models in practical applications.
[0003] One prominent issue is that image pairs (i.e., reference and target images) in existing zero-shot combinatorial image retrieval datasets often lack consistency. Specifically, there may be visual or semantic mismatches between image pairs. For example, some images may be taken with different backgrounds, angles, or lighting conditions, resulting in significant visual differences. Semantically, objects or scenes in images may overlap but lack strict association. This makes it difficult for the model to capture effective visual-textual association information during training. Furthermore, the semantic associations between image pairs are loose and visually distinct, failing to represent the natural consistency between images in the real world. This inconsistency leads to interference during model training, reducing the accuracy and reliability of the retrieval model. Summary of the Invention
[0004] The present disclosure provides a method and electronic device for constructing a retrieval dataset based on a video source dataset, to solve at least one of the above problems.
[0005] According to a first aspect of an embodiment of the present disclosure, a method for constructing a retrieval dataset based on a video source dataset is provided, the method comprising: acquiring first video data and second video data; determining multiple reference images from multiple frames of first images of the first video data; for each reference image, performing the following operations: determining similarity information between the reference image and each frame of the first image, and determining the first image whose similarity information meets a preset similarity condition as a candidate target image of the reference image; generating relative description information between the reference image and the candidate target image as candidate relative description information, wherein the relative description information is used to describe the content association between the two images; constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information; and constructing a retrieval dataset based on the positive sample sets of each reference image and their relative description information, the negative sample set, and the multiple frames of second images of the second video data.
[0006] Optionally, determining multiple reference images from multiple frames of first images of the first video data includes: dividing the multiple frames of first images into at least one image subset according to a set number of frames; extracting multiple candidate reference images from the at least one image subset; determining the visual similarity and semantic similarity between the multiple candidate reference images; and determining images whose visual similarity is less than a first visual threshold and whose semantic similarity is less than a semantic threshold from the multiple candidate reference images to obtain the multiple reference images.
[0007] Optionally, the similarity information includes at least one of the following: visual similarity, semantic similarity, wherein the determining of the first image whose similarity information satisfies a preset similarity condition as a candidate target image of the reference image includes: determining the first image whose each similarity in the similarity information is within a preset value range corresponding to the similarity as the preliminary target image of the reference image; and determining the preliminary target image whose visual similarity with each other preliminary target image is less than or equal to a second visual threshold as a candidate target image of the reference image.
[0008] Optionally, constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information includes: determining the candidate relative description information that meets preset conditions among the candidate relative description information as the reference relative description information; determining the text similarity between each candidate relative description information and the reference relative description information as the candidate similarity of the candidate relative description information; and constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate similarity.
[0009] Optionally, the method of determining the candidate relative description information that meets the preset conditions among the candidate relative description information as the reference relative description information includes: for each candidate relative description information, determining the text similarity between the candidate relative description information and other candidate relative description information as the first similarity of the candidate relative description information; determining the maximum value among the various first similarities of the candidate relative description information, and determining the larger of the maximum value and a preset minimum similarity as the second similarity of the candidate relative description information; determining the maximum value among the second similarities of all candidate relative description information as the reference similarity, and determining the candidate relative description information corresponding to the reference similarity as the reference relative description information.
[0010] Optionally, constructing the positive sample set and the negative sample set of the reference image based on the candidate similarities and the candidate target image includes: determining the final relative description information based on the candidate relative description information corresponding to the candidate similarities greater than or equal to the reference similarity in each candidate similarity; selecting the candidate target image corresponding to the final relative description information to obtain the positive sample set of the reference image; and determining the set of images in the candidate target image that are not included in the positive sample set of the reference image as the negative sample set of the reference image.
[0011] Optionally, the release time of the first video data and the second video data are both later than a preset time, wherein the preset time is the latest time when the disclosed image processing model completes pre-training when executing the method of constructing a retrieval dataset based on a video source dataset.
[0012] According to a second aspect of an embodiment of the present disclosure, a device for constructing a retrieval dataset based on a video source dataset is provided, the device comprising: an acquisition unit configured to acquire first video data and second video data; a determination unit configured to determine multiple reference images from multiple frames of first images of the first video data; an execution unit configured to, for each reference image, perform the following operations: determine similarity information between the reference image and each frame of the first image, and determine the first image whose similarity information meets a preset similarity condition as a candidate target image of the reference image; generate relative description information between the reference image and the candidate target image as candidate relative description information, wherein the relative description information is used to describe the content association between the two images; construct a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information; a construction unit configured to construct a retrieval dataset based on the positive sample sets of each reference image and their relative description information, the negative sample set, and multiple frames of second images of the second video data.
[0013] Optionally, the determination unit is further configured to: divide the multiple frames of first images into at least one image subset according to a set number of frames; extract multiple candidate reference images from the at least one image subset; determine the visual similarity and semantic similarity between the multiple candidate reference images; determine images from the multiple candidate reference images whose visual similarity is less than a first visual threshold and whose semantic similarity is less than a semantic threshold, to obtain the multiple reference images.
[0014] Optionally, the similarity information includes at least one of the following: visual similarity, semantic similarity, and the execution unit is further configured to: determine the first image whose each similarity in the similarity information is within a preset value range corresponding to the similarity as the preliminary target image of the reference image; and determine the preliminary target image whose visual similarity with each other preliminary target image is less than or equal to a second visual threshold as the candidate target image of the reference image.
[0015] Optionally, the execution unit is further configured to: determine the candidate relative description information that meets preset conditions among the candidate relative description information as reference relative description information; determine the text similarity between each candidate relative description information and the reference relative description information as the candidate similarity of the candidate relative description information; and construct a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate similarity.
[0016] Optionally, the execution unit is further configured to: for each candidate relative description information, determine the text similarity between the candidate relative description information and other candidate relative description information as the first similarity of the candidate relative description information; determine the maximum value among the various first similarities of the candidate relative description information, and determine the larger of the maximum value and a preset minimum similarity as the second similarity of the candidate relative description information; determine the maximum value among the second similarities of all candidate relative description information as the reference similarity, and determine the candidate relative description information corresponding to the reference similarity as the reference relative description information.
[0017] Optionally, the execution unit is further configured to: determine the final relative description information based on the candidate relative description information corresponding to the candidate similarities greater than or equal to the reference similarity in each candidate similarity; select the candidate target image corresponding to the final relative description information to obtain the positive sample set of the reference image; and determine the set of images in the candidate target image that are not included in the positive sample set of the reference image as the negative sample set of the reference image.
[0018] Optionally, the release time of the first video data and the second video data are both later than a preset time, wherein the preset time is the latest time when the disclosed image processing model completes pre-training when executing the method of constructing a retrieval dataset based on a video source dataset.
[0019] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, prompt the at least one processor to execute a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0020] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to execute a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0021] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by at least one processor, prompt the at least one processor to execute a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0022] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: According to the method and electronic device for constructing a retrieval dataset based on a video source dataset disclosed in the present disclosure, a retrieval dataset is constructed based on the video source dataset. Since each frame in the video carries continuous image information, the scenes, characters and activities expressed therein are often more visually coherent. By determining a reference image and its corresponding target image from the same video data (i.e., the first video data), it is possible to utilize the strong visual and semantic consistency that the images in the same video data naturally possess, and effectively construct image pairs with high consistency, thereby reducing the common visual and semantic inconsistency problems in traditional retrieval datasets. In addition, by constructing the target image corresponding to the reference image as a positive sample set, and constructing the other images in the candidate target images as a negative sample set of the reference image, it is possible to distinguish it from the existing retrieval dataset, provide examples of erroneous queries, lay the foundation for building a more realistic and accurate retrieval model, and effectively improve the retrieval performance of the model.
[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0025] Figure 1 is a flowchart of a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0026] Figure 2 The present invention is a flowchart of constructing a positive sample set and a negative sample set of a reference image according to a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0027] Figure 3 It is a schematic diagram of a framework of a method for constructing a retrieval dataset based on a video source dataset according to a specific embodiment of the present disclosure.
[0028] Figure 4 is a schematic diagram of retrieving a data set according to a specific embodiment of the present disclosure.
[0029] Figure 5 4 is a block diagram of an apparatus for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure.
[0030] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0032] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0033] It should be noted that the phrase "at least one of the several items" in this disclosure includes three types of parallel situations: "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. For another example, "performing at least one of step 1 and step 2" means the following three parallel situations: (1) performing step 1; (2) performing step 2; and (3) performing both step 1 and step 2.
[0034] Hereinafter, a method and an electronic device for constructing a retrieval dataset based on a video source dataset according to exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0035] Figure 1 This is a flowchart of a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure. This method can be executed on an electronic device with sufficient computing power. The constructed retrieval dataset includes, but is not limited to, a combined image retrieval dataset, specifically, a zero-sample combined image retrieval dataset.
[0036] Reference Figure 1 In step S101, first video data and second video data are acquired.
[0037] It should be understood that the naming of the first and second video data is intended to indicate that they serve different purposes. Furthermore, as will be discussed later, the first video data serves as the source of reference images and their positive and negative sample sets, while the second video data serves as the source of other noise images in the complete search pool, simulating actual search scenarios. In other words, the search pool contains not only the query structure for the reference image constructed based on the first video data, but also images that are not part of the query structure, constructed based on the second video data.
[0038] As an example, a plurality of different types of video source data sets may be selected and divided into a plurality of categories and subcategories as needed, and then the first video data and the second video data may be selected as needed.
[0039] In step S102 , a plurality of reference images are determined from a plurality of first image frames of the first video data.
[0040] This step extracts multiple frames of images from the first video data to obtain a reference image as a basis. It should be noted that the first image is an image in the first video data that serves as a reference image source. This means that all images in the first video data can be used as first images, or only some images in the first video data can be used as first images. This is equivalent to first performing a preliminary screening of the images in the first video data, such as but not limited to extracting frames in the first video data as first images according to a fixed period (e.g., 1s), extracting key frames in the first video data as first images, and then further determining multiple reference images from the multiple frames of first images obtained from the preliminary screening. As an example, different reference images representing different contents in the first video data can be extracted to enrich the data.
[0041] In step S103 , for each reference image, a positive sample set and a negative sample set of the reference image are constructed based on multiple frames of first images of the first video data.
[0042] Figure 2 The execution flow of this step is shown.
[0043] Reference Figure 2 In step S201, similarity information between the reference image and the first image of each frame is determined, and the first image whose similarity information meets a preset similarity condition is determined as a candidate target image of the reference image.
[0044] By selecting the first image of each frame based on relative consistency, we ensure that each candidate target image has a high degree of visual and semantic relevance to the corresponding reference image. The resulting candidate target images serve as the final basis for determining the positive and negative sample sets.
[0045] In step S202 , relative description information between the reference image and the candidate target image is generated as candidate relative description information.
[0046] Relative description information is used to describe the content relationship between two images. Specifically, a multi-stage Large Language Model (LLM) is used to assist in the generation process, generating relative description information for each pair of reference image and candidate target image. Each generated relative description accurately reflects the relationship between the reference image and the candidate target image, ensuring semantic consistency and accuracy of the description content, eliminating ambiguity in image descriptions, and ensuring the adaptability of the generated data for image retrieval tasks (including but not limited to zero-shot composite image retrieval tasks).
[0047] In step S203 , a positive sample set and a negative sample set of a reference image are constructed based on the candidate target image according to the candidate relative description information.
[0048] For example, the Vision Transformer (ViT) model and the CLIP (Contrastive Language–Image Pre-training) model can be used to screen reference images and candidate target images, eliminating images that are irrelevant to the reference image's content or visually redundant. The remaining images ultimately constitute a set of positive samples, ensuring visual and semantic consistency between each pair of reference and target images. The excluded candidate target images can then be used to construct a set of negative samples, which are images that are very similar to the reference image but do not meet the retrieval requirements, thereby providing high-quality examples of incorrect queries.
[0049] Return to reference Figure 1 In step S104, a retrieval data set is constructed based on the positive sample set of each reference image and its relative description information, the negative sample set, and multiple frames of second images of the second video data.
[0050] This step combines the different video source data and the generated relative description information to construct a complete image retrieval dataset. The dataset can contain multiple queries and multiple images, and each query is accompanied by multiple positive target images (i.e., images in the positive sample set) and negative target images (i.e., images in the negative sample set). This ensures that the dataset can comprehensively evaluate the capabilities of retrieval models (including but not limited to zero-shot combined image retrieval models) and effectively support corresponding image retrieval tasks.
[0051] As an example, after the final dataset is constructed, it can be evaluated in multiple dimensions to ensure the quality of the dataset, avoid bias in the data source, and provide real and high-quality test data for subsequent image retrieval tasks.
[0052] According to the method for constructing a retrieval dataset based on a video source dataset of an exemplary embodiment of the present disclosure, the retrieval dataset is constructed based on the video source dataset. Since each frame in the video carries continuous image information, the scenes, characters and activities it expresses are often more visually coherent. By determining a reference image and its corresponding target image from the same video data (i.e., the first video data), it is possible to utilize the strong visual and semantic consistency that the images in the same video data naturally possess, and effectively construct image pairs with high consistency, thereby reducing the common visual and semantic inconsistency problems in traditional retrieval datasets. In addition, by constructing the target image corresponding to the reference image as a positive sample set, and constructing other images in the candidate target images as a negative sample set of the reference image, it is possible to distinguish it from the existing retrieval dataset, provide examples of erroneous queries, lay the foundation for building a more realistic and accurate retrieval model, and improve the accuracy of using such a retrieval model to perform image retrieval tasks.
[0053] Next, a method for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure is further introduced.
[0054] Regarding step S102, determining multiple reference images, this step optionally includes: dividing the multiple first image frames into at least one image subset according to a set number of frames; extracting multiple candidate reference images from the at least one image subset; determining visual and semantic similarities between the multiple candidate reference images; and determining images from the multiple candidate reference images whose visual similarity is less than a first visual threshold and whose semantic similarity is less than a semantic threshold, thereby obtaining multiple reference images. By subsetting the multiple first image frames, first images that are temporally close and have strong content relevance can be initially grouped into one image subset, facilitating the extraction of candidate reference images with different content. For example, a Large Vision Language Model (LVLM) combined with prompt words can be used to preliminarily screen these image subsets and select candidate reference images. The prompt words can be used, for example, to avoid selecting similar images when screening candidate reference images. Furthermore, by further calculating the visual and semantic similarities between the multiple candidate reference images and selecting images with less similarity, candidate reference images that are too visually and semantically identical can be removed, further ensuring the relative independence of each selected reference image.
[0055] As an example, when determining the visual similarity and semantic similarity between multiple candidate reference images, for example, each candidate reference image can be traversed, and the visual similarity and semantic similarity between the candidate reference image and each other candidate reference image can be calculated. Based on this, it is determined whether to retain the candidate reference image. For example, the average of the calculated visual similarities can be used as the visual similarity of the candidate reference image, and the average of the calculated semantic similarities can be used as the semantic similarity of the candidate reference image. Then, combined with the first visual threshold and the semantic threshold, a comparison is made to determine whether to retain the candidate reference image. For another example, it is possible to determine whether each calculated visual similarity is less than the first visual threshold. If so, the candidate reference image is retained; otherwise, it is excluded. The same applies to the determination of semantic similarity. With respect to the latter, in order to reduce the amount of calculation, when traversing subsequent candidate reference images, similarity calculations can be performed only for the currently remaining candidate reference images, and the candidate reference images that have been excluded are no longer calculated. Of course, other reasonable methods can also be used, and this disclosure does not limit this.
[0056] As an example, the first visual threshold is, for example, 0.4, 0.5, 0.6, etc., and the semantic threshold is, for example, 0.7, 0.8, 0.9, etc.
[0057] Regarding step S201 in step S103, specifically how to determine the first image whose similarity information meets the preset similarity condition as a candidate target image of the reference image, optionally, the similarity information includes at least one of the following: visual similarity, semantic similarity, and the above operation includes: determining the first image whose each similarity in the similarity information is within the preset value range corresponding to the similarity as the preliminary target image of the reference image; determining the preliminary target image whose visual similarity with each other preliminary target image is less than or equal to the second visual threshold as the candidate target image of the reference image.
[0058] By first screening the first images whose similarities are within a preset value range as preliminary target images based on the similarity information between the reference image and each first image, it is possible to obtain preliminary target images that have a certain degree of similarity but not identical visually and / or semantically with the reference image with the help of a moderate preset value range, which helps to improve the quality of the retrieval data set. For the case where the similarity information includes two or more similarities, such as visual similarity and semantic similarity, as an example, the complete similarity information of all first images can be determined at the same time and compared with the corresponding preset value ranges respectively; or one of the similarities of all first images can be determined first, and the first images whose similarities are within the corresponding preset value range can be screened, and then the next similarity of the screened first images can be calculated, and further screened in combination with the corresponding preset value range until all similarity calculations and corresponding screening are completed, which can reduce the amount of calculation for the similarity and value range comparison. Of course, other reasonable methods can also be used, and the present disclosure does not limit this.
[0059] On this basis, by continuing to calculate and screen the visual similarity of the multiple preliminary target images obtained by screening, the risk of each target image corresponding to each reference image containing visually redundant images can be reduced, thereby improving the quality of the obtained candidate target images.
[0060] Regarding step S203 in step S103, i.e., how to construct positive and negative sample sets, optionally, the operation of constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information in step S203 includes: determining the candidate relative description information that meets a preset condition among the candidate relative description information as the reference relative description information; determining the text similarity between each candidate relative description information and the reference relative description information as the candidate similarity of the candidate relative description information; and constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate similarity. The candidate relative description information between each candidate target image and the reference image obtained in step S202 can reflect the content association between these candidate target images and the reference image. By determining the information that meets the preset condition, specifically, the preset condition can be used to find the most representative content among the candidate relative description information, which can be used as a reference for evaluating other candidate relative description information. The candidate target images corresponding to the candidate relative description information with a higher text similarity are classified into the positive sample set, thereby obtaining a positive sample set that is closely associated with the reference image.
[0061] Further optionally, the above-mentioned operation of determining the candidate relative description information that meets the preset conditions in each candidate relative description information as the reference relative description information includes: for each candidate relative description information, determining the text similarity between the candidate relative description information and other candidate relative description information as the first similarity of the candidate relative description information; determining the maximum value among the first similarities of the candidate relative description information, and determining the larger of the maximum value and the preset minimum similarity as the second similarity of the candidate relative description information; determining the maximum value among the second similarities of all candidate relative description information as the reference similarity, and determining the candidate relative description information corresponding to the reference similarity as the reference relative description information. By calculating the first similarity and further obtaining the second similarity, it is possible to understand the content representativeness of each information in all candidate relative description information, that is, whether the information can represent other information. The maximum value in the second similarity can reflect the closest degree of similarity, and the candidate relative description information corresponding to the maximum value is the reference relative description information with the most content representativeness.
[0062] Optionally, the above-mentioned operation of constructing a positive sample set and a negative sample set of a reference image based on candidate target images according to candidate similarities includes: determining final relative description information based on candidate relative description information corresponding to candidate similarities greater than or equal to the reference similarity among each candidate similarity; selecting candidate target images corresponding to the final relative description information to obtain a positive sample set of the reference image; and determining a set of images among the candidate target images that are not included in the positive sample set of the reference image as a negative sample set of the reference image. By using the reference similarity as a reference basis for determining whether each candidate relative description information has a high textual similarity with the reference relative description information, the judgment standard can be improved, and final relative description information and corresponding positive target images that are highly matched with the reference image can be obtained, thereby improving the quality of the positive sample set.
[0063] In addition to the problems with the prior art described above, another issue with existing combined image retrieval datasets is that many existing zero-shot combined image retrieval datasets often contain a large number of datasets that have already been used for pre-training (for example, pre-training of deep learning models such as CLIP). Because these pre-training datasets contain a wide range of image data, the model in the zero-shot combined image retrieval task has already been trained on certain images, thus not being truly in a zero-shot environment. This "data leakage" phenomenon prevents the model from truly facing a zero-shot scenario during the testing phase, thus losing a true test of its zero-shot learning capabilities. Consequently, the retrieval task cannot truly test the model's generalization and zero-shot capabilities. For example, the candidate relative description information corresponding to candidate similarities greater than or equal to the reference similarity can be directly determined as the final relative description information. Alternatively, the candidate relative description information corresponding to candidate similarities greater than or equal to the reference similarity can be further quality-processed, such as using a large language model to transform it according to preset description requirements to obtain the final relative description information. This is not limited in this disclosure.
[0064] To solve this problem, optionally, the release time of the first video data and the second video data of the exemplary embodiment of the present disclosure are both later than a preset time, wherein the preset time is the latest time when the disclosed image processing model completes pre-training when executing the method of constructing a retrieval dataset based on a video source dataset.
[0065] By using video data released after the latest pre-training date for currently available image processing models, the retrieval dataset can be constructed using video data not previously included in the pre-training dataset. This ensures that the created retrieval dataset is independent of existing pre-trained models and ensures that the model is trained and tested in a true zero-shot environment. This means that the model relies on a direct match between relative descriptive information and image content, rather than relying on previously seen data. This feature avoids the effects of data leakage, making this method more suitable for real-world zero-shot learning scenarios, effectively testing the model's generalization capabilities, and promoting innovation in image retrieval technology across a variety of practical applications. For example, for pre-trained models such as CLIP, the preset deadline can be March 31, 2022. This ensures that the video content used is not previously included in the training data of pre-trained models such as CLIP. This prevents the video dataset from being trained on pre-trained models such as CLIP, thus reducing the risk of data bias.
[0066] Next, combine Figure 3 Introducing a method for constructing a retrieval dataset based on a video source dataset according to a specific embodiment of the present disclosure, the constructed retrieval dataset is named ZeroSight. Figure 3 White rectangles are used to schematically represent the processed images.
[0067] Reference Figure 3 , this specific embodiment includes the following steps.
[0068] Step 1: Video dataset screening.
[0069] This step corresponds to Figure 1 In step S101, eligible video data is screened from high-quality video datasets released after March 31, 2022, ensuring that the selected videos are of high quality and representativeness. To provide high-quality video sources, videos with clear image quality and rich scene changes can be selected to ensure that the subsequently extracted images have sufficient visual information and diversity. Multiple different types of video source datasets can be selected and divided into multiple categories and subcategories as needed.
[0070] Step 2: Video frame extraction.
[0071] This step extracts one frame per second from the filtered video data, and the extracted video frames are used as the first images. The set of these first images is recorded as F (Right now Figure 3 (see the lower row of images in step 2 of the previous section).
[0072] Step 3: Generate reference image.
[0073] This step corresponds to Figure 1 In step S102, based on the first image extracted in step 2, multiple different reference images are further extracted from the same video (i.e., the first video data). In order to ensure the diversity of ZeroSight, the reference images can be non-redundant and have significant differences. To this end, the first image set extracted can be first F Divide into equally spaced subsets , expressed by the following formula.
[0074]
[0075] In the above formula, Representing a collection F The number of elements in , that is, the number of the first image extracted in step 2. Represents an equidistant subset The set of each subset Contains at most 10 frames of images.
[0076] Then, we use existing LVLM (such as GPT-4, Generative Pre-trained Transformer 4) to preliminarily screen these subsets and select candidate reference images. , expressed by the following formula.
[0077]
[0078] The meaning of this formula is that if the first candidate reference image is selected, it is only necessary to use LVLM to select each subset in turn. If a new candidate reference image is selected after several candidate reference images have been selected, the previous candidate reference image that has been selected must be input into the LVLM to continue screening the subset, so that each selected candidate reference image is different from the previous one, thereby ensuring that similar images are avoided when initially selecting candidate reference images. It is a hint to use LVLM when selecting the first candidate reference image. It is a hint to use LVLM when continuing to select new candidate reference images when there are already several selected candidate reference images. Each subset determines at most one candidate reference image, and the candidate reference images selected in this process constitute the candidate reference image set. ,in, Representing a collection C The number of elements in the subset To further ensure the relative independence of each selected image, additional steps may be performed.
[0079] First, a visual transformer (ViT) model is used to remove candidate reference images that are too visually consistent, as expressed in the following formula.
[0080]
[0081] In the above formula, It represents a set of relatively independent candidate reference images at the visual level, where each element is a candidate reference image. is the upper limit of the visual similarity threshold for screening, that is, the first visual threshold. The visual similarity of each candidate reference image with other candidate reference images is checked in turn, and for each candidate reference image, all candidate reference images whose visual similarity exceeds the upper limit of the visual similarity threshold are deleted.
[0082] Then, the CLIP model is used to remove candidate reference images that are too semantically consistent. The final result is a set of reference images that are relatively independent in both visual and semantic terms. Each element in the set is a reference image, which is expressed as the following formula.
[0083]
[0084] In the above formula, Is the upper limit of the semantic similarity threshold for screening, that is, the semantic threshold. Check the semantic similarity of each candidate reference image with other candidate reference images in turn, and delete all candidate reference images whose semantic similarity exceeds the upper limit of the semantic similarity threshold for each candidate reference image. Figure 3 In step 3 of the above diagram, the multiple images on the left represent candidate reference images. After the first screening criteria (visual similarity < 50%, semantic similarity < 80%) are applied, the multiple reference images on the right are obtained. It should be understood that the ellipsis at the bottom of the three reference images on the right represent other reference images. Figure 3 The ellipsis marks in other parts also indicate omitted content, which will not be explained one by one in the following text.
[0085] Step 4: Generate candidate target images.
[0086] This step corresponds to Figure 2 In step S201 , multiple target images that are similar to but not visually identical to the reference image are selected to construct consistent image pairs, thereby ensuring visual and semantic coherence of the image pairs.
[0087] After generating a reference image set for a video, In the frame image (i.e., the first image of multiple frames), multiple target images are constructed for each reference image. The goal is to select an image that has a certain degree of similarity with the reference image but is not completely identical. To do this, you can first use Filter out reference images A set of images with a certain visual similarity , expressed by the following formula.
[0088]
[0089] In the above formula, and They are the lower and upper limits of the visual similarity threshold, respectively, and constitute the preset value range of visual similarity.
[0090] Next, to further ensure that the filtered similar images have semantic connections with the reference image, the semantic similarity between these images and the reference image can be considered, and CLIP can be used to further filter out the similar images. , that is, with A set of images with a certain semantic similarity is expressed by the following formula.
[0091]
[0092] In the above formula, and They are the lower and upper limits of the semantic similarity threshold, respectively, and constitute the preset value range of semantic similarity.
[0093] Finally, to ensure that the target image set corresponding to each reference image does not contain images that are too visually redundant, we can use To filter out The images that are too visually similar in are expressed as follows.
[0094]
[0095] This indicates the reference image The final candidate target image set, each element of which is a reference image Here, Is the upper limit of the visual similarity threshold, that is, the second visual threshold. Corresponding to Figure 3 In step 4 of the figure, each dotted box on the left represents an image set. After the second screening criterion (visual similarity < 85%), the candidate target image set represented by the corresponding dotted box on the right is obtained. .
[0096] Then, for the A video (ie, the second video data) is extracted every several frames (eg, including but not limited to 5 frames), and these images are Combined to form a complete retrieval pool (i.e. retrieval dataset) for model retrieval.
[0097] Step 5: Generate candidate relative descriptions.
[0098] Based on the image pairs generated in step 4, LVLM (such as GPT-4) is used to generate relative descriptions for each pair of images to obtain candidate relative description information, ensuring semantic consistency between each pair of images while avoiding unnecessary noise and redundant information, and enhancing the accuracy and relevance of the text description.
[0099] Step 6: Generate relative description information, positive sample set and negative sample set.
[0100] This step constructs and annotates the retrieval dataset. Based on the image pairs generated in Step 4 and the candidate relative descriptions generated in Step 5, a dataset containing multiple positive and negative sample images and relative descriptions is constructed. This ensures the diversity and complexity of the dataset, avoids dataset bias, and further improves the generalization and robustness of the model. To provide valid samples for dataset construction, each image and relative description pair must be rigorously screened and annotated, with a clear distinction between positive and negative samples, ensuring the high quality and comprehensiveness of the dataset.
[0101] Specifically, after generating the reference image set and the candidate target image set, in step 5, LVLMs (such as GPT-4) are first used to generate candidate relative description information for each reference image and each candidate target image in its corresponding candidate target image set. Then, , that is, reference images The candidate relative description set is expressed by the following formula.
[0102]
[0103] In the above formula, It is with reference images The corresponding candidate target image set, yes The candidate target images, Expressed as reference image and the The candidate relative description information generated by the candidate target image is then calculated using BERT. The text similarity between all candidate relative description pairs in is expressed by the following formula.
[0104]
[0105] This means that the calculation Middle The text similarity between a candidate relative description information and all other candidate relative description information in the same set is the first similarity. In order to obtain the final relative description, the following function is defined.
[0106]
[0107] In the above formula, Is a variable parameter, indicating the preset minimum similarity. Represents the second similarity. Then construct the set , expressed by the following formula.
[0108]
[0109] In the above formula, represents the reference similarity, Represents the reference relative description information, and filters out candidate relative description information whose text similarity (specifically BERT similarity) with the reference relative description information is greater than or equal to the reference similarity, which is equivalent to Filter out the proportion of candidate relative description information, and these candidate relative description information are similar to each other (text similarity is greater than or equal to ), and the obtained is used to generate the final relative description A subset of . Corresponding to Figure 3 In step 5, each dotted box on the left represents a candidate relative description set After filtering by the third screening criterion (i.e., the text similarity of candidate relative description information with a proportion of k>=k), the candidate relative description information marked with "×" is filtered out, and the set represented by the corresponding dotted box on the right is obtained. .
[0110] Step 6 specifically includes two parts: step 6.1 and step 6.2.
[0111] Step 6.1 is the generation of multiple target images based on Can be the first reference images Generate the final positive sample set , which is expressed by the following formula. The positive sample is the positive target image.
[0112]
[0113] No. reference images The negative sample set It can be expressed by the following formula: Positive samples are negative target images.
[0114]
[0115] The negative sample set is unique to the retrieval dataset of the present disclosure and is different from other datasets.
[0116] Step 6.2 is relative description generation, which can perform quality checks on the constructed dataset to remove potential errors and inconsistent samples, and optimize the dataset to ensure that each sample can provide an effective learning signal for model training. Specifically, LLMs (such as GPT-4) can be used to adjust ,right Make a semantic summary description, according to the corresponding Generate the reference images The unified final relative description of , expressed by the following formula.
[0117]
[0118] In the above formula, Is used to generate a summary Hint of the meaning of the text.
[0119] Step 7: Query generation.
[0120] Finally, we get a triple , as the query structure in the retrieval dataset. Figure 4 is a schematic diagram of a query structure in a retrieval data set according to a specific embodiment of the present disclosure, with reference to Figure 4 Each column represents a query, and the text at the top indicates the type of relative description information. Of course, the types here are only examples, and other reasonable types can also be included, which are not listed here one by one. The image in the first row is the reference image, the text in the second row is the relative description text, that is, the relative description information, the image in the third row is the positive target image, and the image in the fourth row is the negative target image.
[0121] The generated retrieval dataset can be used for model training and its performance can be evaluated on real zero-shot tasks. This ensures that the proposed method can effectively retrieve unseen image and text pairs and improves the model's zero-shot generalization ability. To provide a reliable evaluation basis, multi-dimensional evaluation metrics are used, including retrieval precision, recall rate, and F1 score, to comprehensively reflect the model's performance in real zero-shot tasks.
[0122] To evaluate this embodiment, training and testing were performed on the ZeroSight dataset, which contains 5,746 training queries and 1,437 test queries. The ZeroSight dataset is divided into two categories for evaluation: ZS-CIR and CIR. The ZS-CIR method retrieves target images using reference images and corresponding relative description text, while the CIR method is trained on the training set and only requires reference images and corresponding description text for testing.
[0123] Tables 1 and 2 below compare different methods using three evaluation metrics: mAP@k (mean average precision at k), PNR-mAP@k (positive-negative ratio mean average precision at k), and average. mAP@k measures retrieval accuracy, while PNR-mAP@k evaluates retrieval precision. In Table 1, we compare the following methods (where training-free means that the corresponding methods can be used directly without additional training): CIReVL (Compositional Image Retrieval through Vision-by-Language), published in the 2024 ICLR (International Conference on Learning Representations); LDRE (LLM-based Divergent Reasoning and Ensemble), published in the 2024 SIGIR (ACM Special Interest Group on Information Retrieval), the International Conference on Information Retrieval (ACM is the Association for Computing Machinery); SEIZE (Semantic Editing Increment for ZS-CIR), published in the 2024 MM (ACM International Conference on Multimedia); PALAVRA (Personalizing Language Vision Representations), published in the 2022 ECCV (European Conference on Computer Science). Vision, the European International Conference on Computer Vision);SEARLE-OTI (Zero-shot Composed Image Retrieval With Optimization-based Textual Inversion), published at the 2023 ICCV (International Conference on Computer Vision); SEARLE (Zero-shot Composed Image Retrieval With Textual Inversion), published at the 2023 ICCV; Captioning (a method that uses a pre-trained captioning model to generate captions for a reference image and extracts textual features of the captions through the CLIP text encoder for retrieval; Text-only (a method that uses only the relative description features extracted by the CLIP text encoder as retrieval features to calculate retrieval similarity); Image-only (an image-only method that uses only the features of the reference image extracted by the CLIP image encoder to calculate retrieval similarity); Pic2Word (a training-dependent method that uses a pre-trained text inversion network to capture pseudo-word tags through contrastive loss optimization and combines them with relative descriptions for retrieval), published at the 2023 CVPR (IEEE / CVF Conference on Computer Vision). Vision and Pattern Recognition, IEEE / CVF Conference on Computer Vision and Pattern Recognition); LinCIR (Language-only training for CIR), published in the 2024 CVPR conference. In Table 2, we compare the following methods: APTEMIS (Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity), published in the 2022 ICLR conference; AMC (Adaptive Multi-expert Collaborative Network), published in the 2023 TOMM journal (ACM Transactions on Multimedia Computing, Communications, and Applications).VAL (Visiolinguistic Attention Learning), published at the 2020 CVPR conference; TIRG (Text Image Residual Gating), published at the 2019 CVPR conference; MAAF (Modality-Agnostic Attention Fusion), published in 2020; DCNet (Dual Composition Network), published at the 2021 AAAI Conference on Artificial Intelligence; DWC (Dynamic Weighted Combiner), published at the 2024 AAAI conference; and CLIP4CIR (a two-stage approach that combines task-oriented fine-tuning with training a combination network that performs fine-grained merging of multimodal features), published in the 2023 TOMM journal. As shown in Tables 1 and 2, this specific embodiment significantly outperforms other methods on the ZeroSight dataset, demonstrating excellent performance.
[0124] Table 1
[0125] Table 2
[0126] This specific embodiment proposes a new method for constructing a zero-shot combined image retrieval dataset based on a video source dataset. This method effectively avoids the problem of image pair inconsistency in existing methods by extracting consistent image pairs from video data and generating accurate text descriptions in combination with a large language model. By filtering video data based on release time, the problems of pre-training data leakage and pre-training dataset dependence brought by existing datasets can be solved, and the retrieval task can be carried out in a real zero-shot environment, making the zero-shot combined image retrieval task more reliable and challenging, improving the performance and generalization ability of the retrieval model, providing a new solution for the field of zero-shot combined image retrieval, and also laying the foundation for building a more realistic and accurate retrieval model.
[0127] Figure 5 is a block diagram of an apparatus for constructing a retrieval dataset based on a video source dataset according to an exemplary embodiment of the present disclosure. Figure 5The device 500 for constructing a retrieval dataset based on a video source dataset includes an acquisition unit 501 , a determination unit 502 , an execution unit 503 , and a construction unit 504 .
[0128] The acquisition unit 501 may acquire first video data and second video data.
[0129] The determining unit 502 may determine a plurality of reference images from a plurality of first image frames of the first video data.
[0130] The execution unit 503 can perform the following operations for each reference image: determine the similarity information between the reference image and the first image of each frame, and determine the first image whose similarity information meets the preset similarity condition as the candidate target image of the reference image; generate relative description information between the reference image and the candidate target image as candidate relative description information, wherein the relative description information is used to describe the content association between the two images; and construct a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information.
[0131] The construction unit 504 may construct a retrieval data set based on the positive sample set of each reference image and its relative description information, the negative sample set, and the multiple frames of second images of the second video data.
[0132] Optionally, the determination unit 502 may also: divide the multiple frames of first images into at least one image subset according to a set number of frames; extract multiple candidate reference images from the at least one image subset; determine the visual similarity and semantic similarity between the multiple candidate reference images; determine the images whose visual similarity is less than the first visual threshold and whose semantic similarity is less than the semantic threshold from the multiple candidate reference images, to obtain multiple reference images.
[0133] Optionally, the similarity information includes at least one of the following: visual similarity, semantic similarity, and the execution unit 503 may also: determine the first image whose each similarity in the similarity information is within a preset value range corresponding to the similarity as the preliminary target image of the reference image; determine the preliminary target image whose visual similarity with each other preliminary target image is less than or equal to a second visual threshold as the candidate target image of the reference image.
[0134] Optionally, the execution unit 503 can also: determine the candidate relative description information that meets the preset conditions among the candidate relative description information as the reference relative description information; determine the text similarity between each candidate relative description information and the reference relative description information as the candidate similarity of the candidate relative description information; and construct a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate similarity.
[0135] Optionally, the execution unit 503 may also: determine, for each candidate relative description information, the text similarity between the candidate relative description information and other candidate relative description information as the first similarity of the candidate relative description information; determine the maximum value among the various first similarities of the candidate relative description information, and determine the larger of the maximum value and a preset minimum similarity as the second similarity of the candidate relative description information; determine the maximum value among the second similarities of all candidate relative description information as the reference similarity, and determine the candidate relative description information corresponding to the reference similarity as the reference relative description information.
[0136] Optionally, the execution unit 503 may also: determine the final relative description information based on the candidate relative description information corresponding to the candidate similarities greater than or equal to the reference similarity in each candidate similarity; select the candidate target image corresponding to the final relative description information to obtain a positive sample set of the reference image; and determine the set of images in the candidate target image that are not included in the positive sample set of the reference image as a negative sample set of the reference image.
[0137] Optionally, the release time of the first video data and the second video data are both later than a preset time, wherein the preset time is the latest time when the disclosed image processing model completes pre-training when executing the method of constructing a retrieval dataset based on a video source dataset.
[0138] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0139] Figure 6 FIG. 6 shows a structural block diagram of an electronic device 600 according to an exemplary embodiment of the present disclosure.
[0140] Reference Figure 6 The electronic device 600 includes: at least one memory 601 and at least one processor 602, wherein the at least one memory 601 stores computer executable instructions. When the computer executable instructions are executed by the at least one processor 602, the at least one processor is prompted to execute the method for constructing a retrieval dataset based on a video source dataset as described in the above exemplary embodiment.
[0141] As an example, electronic device 600 may be a PC, tablet device, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, electronic device 600 is not necessarily a single electronic device 600, but may also be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction set) individually or in combination. Electronic device 600 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device 600 that is interconnected locally or remotely (e.g., via wireless transmission) via an interface.
[0142] In electronic device 600, processor 602 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor 602 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0143] The processor 602 can execute instructions or codes stored in the memory 601, wherein the memory 601 can also store data. Instructions and data can also be sent and received over the network via the network interface device, wherein the network interface device can use any known transmission protocol.
[0144] The memory 601 may be integrated with the processor 602, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, the memory 601 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. The memory 601 and the processor 602 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that the processor 602 can access files stored in the memory.
[0145] In addition, the electronic device 600 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 600 may be connected to each other via a bus and / or a network.
[0146] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, the at least one processor is prompted to execute the method for constructing a retrieval dataset based on a video source dataset as described in the above exemplary embodiment. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as a multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be executed in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0147] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including computer instructions, which, when executed by at least one processor, execute the method for constructing a retrieval dataset based on a video source dataset as described in the above exemplary embodiments.
[0148] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0149] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for constructing a retrieval dataset based on a video source dataset, characterized in that: The method comprises: Acquire first video data and second video data; determining a plurality of reference images from a plurality of first image frames of the first video data; For each reference image, do the following: Determining similarity information between the reference image and the first image of each frame, and determining the first image whose similarity information meets a preset similarity condition as a candidate target image of the reference image; generating relative description information between the reference image and the candidate target image as candidate relative description information, wherein the relative description information is used to describe the content association between the two images; Constructing a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate relative description information; A retrieval data set is constructed based on a positive sample set of each reference image and its relative description information, a negative sample set, and multiple frames of second images of the second video data.
2. The method according to claim 1, wherein The determining of a plurality of reference images from a plurality of first image frames of the first video data comprises: Dividing the plurality of first image frames into at least one image subset according to a set number of frames; extracting a plurality of candidate reference images from the at least one subset of images; determining visual similarity and semantic similarity between the plurality of candidate reference images; Determine, from the multiple candidate reference images, images whose visual similarity is less than a first visual threshold and whose semantic similarity is less than a semantic threshold, to obtain the multiple reference images.
3. The method according to claim 1, wherein The similarity information includes at least one of the following: visual similarity and semantic similarity, wherein determining the first image whose similarity information satisfies a preset similarity condition as a candidate target image for the reference image includes: Determine the first image whose similarities in the similarity information are within a preset value range corresponding to the similarities as a preselected target image of the reference image; A preliminary selected target image whose visual similarity with other preliminary selected target images is less than or equal to a second visual threshold is determined as a candidate target image of the reference image.
4. The method according to any one of claims 1 to 3, characterized in that The constructing, based on the candidate target image and according to the candidate relative description information, a positive sample set and a negative sample set of the reference image includes: Determining candidate relative description information that meets a preset condition among the candidate relative description information as reference relative description information; determining a text similarity between each candidate relative description information and the reference relative description information as a candidate similarity of the candidate relative description information; According to the candidate similarities, a positive sample set and a negative sample set of the reference image are constructed based on the candidate target image.
5. The method according to claim 4, wherein The step of determining the candidate relative description information satisfying a preset condition among the candidate relative description information as the reference relative description information includes: For each candidate relative description information, determining a text similarity between the candidate relative description information and other candidate relative description information as a first similarity of the candidate relative description information; Determine a maximum value among the first similarities of the candidate relative description information, and determine the larger of the maximum value and a preset minimum similarity as the second similarity of the candidate relative description information; The maximum value among the second similarities of all candidate relative description information is determined as a reference similarity, and the candidate relative description information corresponding to the reference similarity is determined as the reference relative description information.
6. The method according to claim 5, wherein The constructing of a positive sample set and a negative sample set of the reference image based on the candidate target image according to the candidate similarity comprises: Determining final relative description information based on candidate relative description information corresponding to candidate similarities greater than or equal to the reference similarity among the candidate similarities; Selecting a candidate target image corresponding to the final relative description information to obtain a positive sample set of the reference image; A set of images in the candidate target images that are not included in the positive sample set of the reference image is determined as a negative sample set of the reference image.
7. The method according to any one of claims 1 to 3, characterized in that The release time of the first video data and the second video data is later than a preset time, wherein the preset time is the latest time when the disclosed image processing model completes pre-training when executing the method for constructing a retrieval dataset based on a video source dataset.
8. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer-executable instructions, When the computer executable instructions are executed by the at least one processor, the computer executable instructions prompt the at least one processor to execute the method for constructing a retrieval dataset based on a video source dataset according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by at least one processor, the instructions cause the at least one processor to perform the method for constructing a retrieval dataset based on a video source dataset according to any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the computer instructions cause the at least one processor to execute the method for constructing a retrieval dataset based on a video source dataset according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video image clustering method and system
CN101359368A
Model optimization method and device, equipment, storage medium and program product
CN115129908A
Domain-adaptive retrieval enhancement generation method and system
CN119669400A
Image retrieval method, device and equipment based on multi-modal semantics
CN120407825A
Method and system for generating candidate vocabulary
US20230334087A1