Image-text pair data set construction method, electronic equipment, storage medium and program product
By segmenting the images of the graphic and text pairs in instances, identifying and separating the target instances to match the text and text, the problem of low quality data sets in the prior art is solved, and high-quality graphic and text pairs are screened and data set construction is realized.
Patent Information
- Application Number
- CN202510526456.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When constructing a picture-text data set, it is difficult to accurately screen out picture-text pairs with high matching degree, resulting in low quality of the data set. Especially when the image content is complex and there are many elements, it is susceptible to interference from background and irrelevant elements.
By segmenting the images of the graphic and text pairs in an instance, identifying and separating the target instances that can represent the core content of the image, matching the text of the graphic and text pairs based on the instance segmentation information, high-quality target graphic and text pairs are selected, and finally building a high-quality graphic and text pair data set.
It effectively reduces interference from background and other irrelevant elements in the image, improves the accuracy of picture and text matching, and builds a high-quality picture and text data set, which improves the overall quality of the data set.
Smart Images

Figure CN120045954A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data set construction, and in particular to a method for constructing a text-image data set, an electronic device, a storage medium, and a program product. Background Art
[0002] With the rapid development of artificial intelligence technology, the industry's demand for high-quality image and text data sets is increasing.
[0003] A high-quality image-text pair dataset is inseparable from a high-quality image-text pair, and the quality of an image-text pair is positively correlated with the degree of matching between the image and the text in the image-text pair. Therefore, the traditional method of constructing an image-text pair dataset uses feature extraction technology and a vector similarity algorithm to calculate the vector similarity between the feature vector of the image in the image-text pair and the feature vector of the text as the degree of matching between the image and the text in the image-text pair, thereby screening out image-text pairs with a high degree of matching for the construction of the image-text pair dataset.
[0004] However, this method is difficult to accurately screen out image-text pairs with high matching degree when the content of the images in the image-text pairs is relatively complex and there are many elements, resulting in low quality of the constructed image-text pair dataset. Summary of the invention
[0005] The main purpose of the present application is to provide a method for constructing a picture-text pair data set, an electronic device, a storage medium and a program product, aiming to solve the technical problem of low quality of the picture-text pair data set constructed in the related art.
[0006] To achieve the above objectives, the present application provides a method for constructing a picture-text pair dataset, the method comprising: Performing instance segmentation on the image of the image-text pair to obtain an instance segmentation result of the image, wherein the instance segmentation result includes a target instance and instance segmentation information of the target instance; According to the instance segmentation information of the target instance, the target instance is matched with the text of the image-text pair to obtain an instance matching result of the image-text pair; If the instance matching result is a match, the image-text pair is determined as a target image-text pair, and an image-text pair data set is constructed based on at least one of the target image-text pairs.
[0007] In one example, the step of performing instance segmentation on the image of the image-text pair to obtain the instance segmentation result of the image includes: Performing instance segmentation on the image of the image-text pair to obtain at least one instance of the image and instance segmentation information of the at least one instance; Determining a target instance from the at least one instance according to instance segmentation information of the at least one instance; Use the target instance and the instance segmentation information of the target instance as the instance segmentation result of the image.
[0008] In one instance, the instance segmentation information includes category labels. The step of determining the target instance from the at least one instance according to the instance segmentation information of the at least one instance includes: Determine the target category label from the category labels of the at least one instance; Determine the instance with the category label being the target category label as the target instance from the at least one instance.
[0009] In one instance, the step of determining the target category label from the category labels of the at least one instance includes: According to the category labels of the at least one instance, count the repetition times of different category labels in the at least one instance; Determine the category label with the highest repetition times as the target category label.
[0010] In one instance, the instance segmentation information includes a bounding box mask. The step of determining the target category label from the category labels of the at least one instance includes: Determine the instance area of the at least one instance according to the bounding box mask of the at least one instance; According to the instance area and category label of the at least one instance, count the instance areas corresponding to different category labels in the at least one instance; Determine the category label with the largest corresponding instance area as the target category label.
[0011] In one instance, the instance segmentation information includes a bounding box mask. The step of determining the target instance from the at least one instance according to the instance segmentation information of the at least one instance includes: Determine the weight factor of the at least one instance according to the bounding box mask of the at least one instance, and calculate the weight of the at least one instance according to the weight factor of the at least one instance, where the weight factor includes at least one of instance area, instance position, and instance depth of field; Determine the instance with the weight meeting the preset weight condition as the target instance from the at least one instance.
[0012] In one instance, the instance segmentation information includes category labels. The step of matching the target instance with the text of the text-image pair according to the instance segmentation information of the target instance to obtain the instance matching result of the text-image pair includes: Perform text classification on the text of the text-image pair to obtain the text classification result of the text; If the class label of the target instance matches the text classification result, determine that the instance matching result of the text-image pair is a match.
[0013] In one example, the instance segmentation information includes a bounding box mask. The step of matching the target instance with the text of the text-image pair according to the instance segmentation information of the target instance to obtain the instance matching result of the text-image pair includes: Extract image features of the target instance to obtain an image feature vector of the target instance, and extract text features of the text of the text-image pair to obtain a text feature vector of the text. Based on a preset vector similarity algorithm, calculate the vector similarity between the image feature vector and the text feature vector. According to the weight of the target instance, perform weighted calculation on the vector similarity to obtain the instance-text similarity between the target instance and the text, where the weight of the target instance is determined based on the bounding box mask of the target instance. When the instance-text similarity is greater than a preset similarity threshold, determine that the instance matching result of the text-image pair is a match.
[0014] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the text-image pair dataset construction method as described above.
[0015] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the text-image pair dataset construction method as described above.
[0016] In addition, to achieve the above object, the present application further provides a program product, which is a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps of the text-image pair dataset construction method as described above.
[0017] The embodiments of the present application provide a method for constructing a text-image pair dataset, an electronic device, a storage medium, and a program product. The technical solution of the embodiments of the present application is to perform instance segmentation on the images of the text-image pairs to obtain the instance segmentation results of the images, where the instance segmentation results include target instances and the instance segmentation information of the target instances; according to the instance segmentation information of the target instances, match the target instances with the texts of the text-image pairs to obtain the instance matching results of the text-image pairs; if the instance matching results are matching, determine the text-image pairs as target text-image pairs, and construct a text-image pair dataset based on at least one target text-image pair, so that the embodiments of the present application can accurately identify and separate the target instances that can represent the core content of the image from the images with relatively complex content and more elements through instance segmentation technology, and then match the target instances with the texts of the text-image pairs according to the instance segmentation information of the target instances, effectively reducing the interference of the background and other irrelevant elements in the images, screening out high-quality target text-image pairs, improving the accuracy of text-image matching, and then constructing a high-quality text-image pair dataset based on the high-quality target text-image pairs, finally solving the technical problem of the low quality of the text-image pair dataset constructed in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart provided for the first embodiment of the method for constructing a text-image pair dataset of the present application; Figure 2 It is a schematic flowchart provided for the second embodiment of the method for constructing a text-image pair dataset of the present application; Figure 3 It is a schematic flowchart provided for the third embodiment of the method for constructing a text-image pair dataset of the present application; Figure 4 It is a schematic flowchart provided for the fourth embodiment of the method for constructing a text-image pair dataset of the present application; Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the method for constructing a text-image pair dataset in the embodiments of the present application.
[0021] The implementation of the purpose, functional features, and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0022] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0023] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0024] Traditional methods for constructing text-image pair datasets can accurately screen out text-image pairs with a high degree of matching for constructing text-image pair datasets when facing text-image pairs with relatively simple image content and few elements. However, when facing text-image pairs with relatively complex image content and many elements, it is difficult to extract effective image feature vectors from the complex images and is easily interfered by the complex and diverse element content in the images, thereby affecting the quality of the screened text-image pairs and ultimately resulting in a low-quality text-image pair dataset.
[0025] In response to this, the main solution of the embodiments of the present application is: performing instance segmentation on the image of the text-image pair to obtain the instance segmentation result of the image, where the instance segmentation result includes a target instance and instance segmentation information of the target instance; matching the target instance with the text of the text-image pair according to the instance segmentation information of the target instance to obtain the instance matching result of the text-image pair; if the instance matching result is a match, determining the text-image pair as a target text-image pair, and constructing a text-image pair dataset based on at least one of the target text-image pairs.
[0026] The embodiments of the present application use the instance segmentation technology to accurately identify and separate the target instance that can represent the core content of the image from the image with relatively complex content and many elements. Thus, according to the instance segmentation information of the target instance, the target instance is matched with the text of the text-image pair, effectively reducing the interference of the background and other irrelevant elements in the image, screening out high-quality target text-image pairs, improving the accuracy of text-image matching, and then constructing a high-quality text-image pair dataset based on the high-quality target text-image pairs, ultimately solving the technical problem of the low quality of the text-image pair dataset obtained in the related technology.
[0027] It should be noted that the execution subject of the embodiments of the present application is an electronic device, which may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions, tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs (Televisions) and desktop computers, or any electronic device capable of implementing the above functions. The embodiments of the present application do not make specific limitations in this regard. Taking the electronic device as the execution subject as an example, the following embodiments of the present application will be described.
[0028] To better understand the technical solution of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0029] The present application proposes a method for constructing a text-image pair dataset in the first embodiment.
[0030] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart provided for the first embodiment of the method for constructing a text-image pair dataset of the present application.
[0031] In this embodiment, the method for constructing a text-image pair dataset may include steps S100 to S300: Step S100, perform instance segmentation on the image of the text-image pair to obtain the instance segmentation result of the image, where the instance segmentation result includes target instances and instance segmentation information of the target instances; Those skilled in the art know that a text-image pair refers to paired data composed of an image and its corresponding text description (i.e., text), which is usually used to train multimodal models (such as text-image retrieval and image description generation). Instance segmentation is a technology that combines object detection and semantic segmentation, and can identify the precise boundaries (masks), class labels, and confidence levels of each independent object in the image.
[0032] It should be noted that a target instance refers to an instance that can represent the core content of the image, and the instance segmentation information generally includes the mask, class label, confidence level, and spatial location (bounding box coordinates or center point coordinates) of the instance. Among them, the mask and the bounding box coordinates can be collectively referred to as the border mask.
[0033] In this embodiment, the target instance that can represent the core content of the image can be directly segmented from the image through a specific instance segmentation algorithm (for example, the specific instance segmentation algorithm obtained by combining the instance segmentation technology and the saliency detection technology), or at least one instance of the image can be segmented first, and then the target instance that can represent the core content of the image can be determined from them.
[0034] Exemplarily, in a feasible implementation manner, step S100 may include steps S110 to S130: Step S110, perform instance segmentation on the image of the text-image pair to obtain at least one instance of the image and the instance segmentation information of the at least one instance; Step S120, determine the target instance from the at least one instance according to the instance segmentation information of the at least one instance; Step S130, use the target instance and the instance segmentation information of the target instance as the instance segmentation result of the image.
[0035] In this implementation manner, at least one instance of the image in the text-image pair and the instance segmentation information of the at least one instance are segmented through a common instance segmentation algorithm, and then the template instance that can represent the core content of the image is determined from the at least one instance through the instance segmentation information of the at least one instance, so that the target instance and its corresponding instance segmentation information are used as the instance segmentation result of the image for subsequent screening of high-quality text-image pairs.
[0036] It should be noted that when determining the target instance according to the instance segmentation information in this implementation manner, the instance with the highest number of repeated class labels can be determined as the target instance, or the instance with the center point coordinates closest to the center of the image can be determined as the target instance, or the instance with the largest area corresponding to the bounding box coordinates can be determined as the target instance. This implementation manner does not make specific restrictions on this, and users can flexibly set the selection criteria for the target instance according to actual needs, as long as it is ensured that the target instance can represent the core content of the image.
[0037] Step S200, match the target instance with the text of the text-image pair according to the instance segmentation information of the target instance to obtain the instance matching result of the text-image pair; It should be noted that the instance matching result refers to the matching evaluation result of the semantic association degree between the target instance and the text description of the text-image pair.
[0038] In this embodiment, when the instance matching result is a match, it indicates that there is a significant semantic association between the target instance and the text description, and when the instance matching result is a mismatch, it indicates that the semantic association between the target instance and the text description is insufficient.
[0039] Exemplarily, in a feasible implementation, the instance segmentation information includes class labels, and step S200 may include steps S210 to S220: Step S210, perform text classification on the text of the text-image pair to obtain the text classification result of the text; As known to those skilled in the art, text classification refers to the process of assigning a given text to one or more predefined categories through natural language processing techniques. It identifies and differentiates different topics or categories based on the content and semantic features of the text. The text classification result is one or a set of class labels determined according to the text content, and these labels represent the theme of the text or the main object described. For example, if the text mainly describes "cat", the text classification result may be "animal" or more specifically "cat".
[0040] Step S220, if the class label of the target instance matches the text classification result, determine that the instance matching result of the text-image pair is a match.
[0041] In this embodiment, by comparing the class label of the target instance and the text classification result, it is determined whether there is a significant semantic association between the target instance and the text. If the class label of the target instance matches the text classification result, it means that the core content in the image is consistent with the theme described in the text, indicating a high matching degree between the target instance and the text description. Therefore, it can be considered that there is a significant semantic association between them.
[0042] It should be noted that in this embodiment, it can be determined whether the class label of the target instance matches the text classification result by judging whether the class label of the target instance is the same as the text classification result. If the class label of the target instance is the same as the text classification result, it is determined that the class label of the target instance matches the text classification result. It can also be determined whether the class label of the target instance matches the text classification result by calculating the semantic vector similarity between the class label of the target instance and the text classification result. When the semantic vector similarity between the class label of the target instance and the text classification result is greater than a preset threshold, it is determined that the class label of the target instance matches the text classification result. This embodiment does not make specific limitations on this.
[0043] It is worth mentioning that when the text in the text-image pair is a text description of an image type, there is no need to perform text classification on the text, and it can be directly judged whether the class label of the target instance matches the text, so as to quickly determine the instance matching result of the text-image pair.
[0044] Step S300, if the instance matching result is a match, determine the text-image pair as the target text-image pair, and construct a text-image pair data set based on at least one target text-image pair.
[0045] It should be noted that the target image-text pair refers to a high-quality image-text pair with a relatively high matching degree between the image and the text.
[0046] In this embodiment, since the target instance can represent the core content of the image, when the instance matching result is a match, indicating that there is a significant semantic association between the target instance and the text description, the matching degree between the image and the text of this image-text pair is relatively high, and this image-text pair can be regarded as a high-quality image-text pair, that is, the target image-text pair, for constructing a high-quality image-text pair dataset. Correspondingly, when the instance matching result is a mismatch, indicating that the semantic association between the target instance and the text description is insufficient, the matching degree between the image and the text of this image-text pair is relatively low, and this image-text pair can be regarded as a low-quality image-text pair.
[0047] When dealing with complex images, traditional methods are easily interfered by the background or other irrelevant elements, and it is difficult to screen out high-quality image-text pairs, resulting in a low-quality image-text pair dataset. In this embodiment, through instance segmentation technology, the target instance that can represent the core content of the image is accurately identified and separated from the image with relatively complex content and many elements. Then, according to the instance segmentation information of the target instance, the target instance is matched with the text of the image-text pair, effectively reducing the interference of the background and other irrelevant elements in the image, screening out high-quality target image-text pairs, improving the accuracy of image-text matching, and then constructing a high-quality image-text pair dataset based on the high-quality target image-text pairs, finally solving the technical problem of the low quality of the image-text pair dataset constructed in the related technology.
[0048] It is worth mentioning that after constructing the image-text pair dataset based on at least one target image-text pair, this image-text pair dataset can be used to train or fine-tune the model, thereby improving the model's ability.
[0049] Exemplarily, the Chinese CLIP (Contrastive Language–Image Pretraining) model is a vision-language model that supports the understanding and processing of the Chinese language. Since the image-text pair dataset used in the training process of the Chinese CLIP model is not constructed from image-text pairs in the Chinese environment, there will inevitably be some defects when processing image-text pair data in the Chinese environment through the Chinese CLIP model. Therefore, through the method provided in this embodiment, high-quality target image-text pairs can be screened out from the image-text pairs in the Chinese environment, and a high-quality image-text pair dataset composed of target image-text pairs in the Chinese environment can be constructed. Then, the Chinese CLIP model can be fine-tuned (referring to the process of further training the model with a relatively small-scale dataset for a specific task or domain based on an existing pre-trained model) through this image-text pair dataset, so that the fine-tuned Chinese CLIP model performs better in the Chinese environment.
[0050] Based on the above first embodiment, a method for constructing a text-image pair dataset according to the second embodiment of the present application is proposed.
[0051] In the second embodiment of the present application, for the same or similar content as the above embodiment, reference can be made to the above introduction and will not be elaborated hereinafter.
[0052] Please refer to Figure 2 , Figure 2 which is a schematic flowchart provided for the second embodiment of the method for constructing a text-image pair dataset of the present application.
[0053] In this embodiment, step S120 determines a target instance from at least one instance according to the instance segmentation information of at least one instance, which may include steps S121 to S122: Step S121, determine a target category label from the category labels of at least one instance; It should be noted that the target category label refers to the category label that can represent the core content of the image.
[0054] In this embodiment, the image includes at least one instance, and each instance has its corresponding category label (for example, "cat", "dog", "car", etc.). However, not all category labels of the instances are related to the core content of the image. Therefore, it is necessary to screen out the target category label that can best reflect the core content of the image from these category labels.
[0055] In one example, step S121 may include steps A10 to A20: Step A10, count the repetition times of different category labels in at least one instance according to the category labels of at least one instance; Step A20, determine the category label with the highest repetition times as the target category label.
[0056] In this embodiment, by counting the repetition times of different category labels in all instances of the image and selecting the category label with the highest repetition times as the target category label, it cleverly focuses on the most common object type in the image, thus effectively highlighting the core content of the image and reducing the interference of the background and other non-critical elements. This screening mechanism based on frequency (i.e., repetition times) is not only simple to implement and has high computational efficiency, but also particularly applicable to image scenarios dominated by a certain type of object, ensuring that the target instance can accurately represent the core content of the image.
[0057] In another example, the instance segmentation information includes a border mask, and step S121 may further include steps B10 to B30: Step B10, determine the instance area of at least one instance according to the border mask of at least one instance; It should be noted that the instance area can be represented by the number of pixels. The instance area can be the number of pixels occupied by the bounding box of the instance or the number of pixels occupied by the mask of the instance. Among them, the number of pixels occupied by the bounding box can be calculated according to the bounding box coordinates in the bounding box mask, and the number of pixels occupied by the mask can be calculated according to the mask in the bounding box mask.
[0058] Step B20: According to the instance areas and class labels of at least one instance, count the instance areas corresponding to different class labels in at least one instance. Step B30: Determine the class label with the largest corresponding instance area as the target class label.
[0059] Based on the bounding box mask in the instance segmentation information, this embodiment quantifies the area occupied by each instance in the image (i.e., the instance area) by calculating the number of pixels covered by the bounding box or mask of the instance. Subsequently, it aggregates and statistically analyzes the total areas corresponding to different classes according to the class labels, and finally determines the class label with the largest total area as the target class label. Thus, the instance area is introduced as the core screening index, providing a more objective and stable decision basis for determining the target class label. Furthermore, the visual salience of the instance in the image is directly reflected through the physical dimension of the instance area, effectively capturing the core object that dominates the image and determining the target instance that best represents the core content of the image.
[0060] Through the above implementation, this embodiment can accurately identify the target class label that best represents the core content of the image from multiple class labels, providing a reliable basis for the subsequent determination of the target instance.
[0061] Step S122: From at least one instance, determine the instance with the class label being the target class label as the target instance.
[0062] After determining the target class label that can represent the core content of the image in this embodiment, it is then necessary to screen out the instances whose class labels are consistent with the target class label from all the instances in the image and determine them as the target instances, ensuring that the selected target instances can represent the core content of the image and excluding the interference of the background and other irrelevant elements in the image. Thus, even in the case of relatively complex image content and many elements, high-quality target text-image pairs can be accurately screened out, thereby improving the quality of the constructed text-image pair dataset.
[0063] It is worth mentioning that after filtering out the instances whose category labels are the same as the target category labels, these instances can be directly used as the final target instances, or these instances can be used as candidate target instances first, and then the confidence levels of the candidate target instances are further compared, and the instance with the highest confidence level is selected as the final target instance, so as to effectively exclude misdetected or low-quality instances and ensure that the target instances can represent the core content of the image.
[0064] It is not difficult to understand that in addition to determining the final target instance from multiple candidate target instances through the confidence level, the instance closer to the center of the image can be preferentially selected as the final target instance according to the center point coordinates of each candidate target instance, or the instance with a larger area can be preferentially selected as the final target instance according to the bounding box coordinates of each candidate target instance, or by integrating various instance segmentation information, the weights of each candidate target instance are calculated, and the instance with the largest weight is preferentially selected as the target instance.
[0065] Based on the above first embodiment, a method for constructing a text-image pair dataset according to the third embodiment of the present application is proposed.
[0066] In the third embodiment of the present application, the same or similar content as that in the above embodiment can be referred to the above introduction and will not be elaborated hereinafter.
[0067] Please refer to Figure 3 , Figure 3 which is a schematic flowchart provided for the third embodiment of the method for constructing a text-image pair dataset of the present application.
[0068] In this embodiment, step S120 for determining a target instance from at least one instance according to the instance segmentation information of at least one instance may further include steps S123 to S124: Step S123, determining a weight factor of at least one instance according to the bounding box mask of at least one instance, and calculating the weights of at least one instance according to the weight factors of at least one instance, where the weight factor includes at least one of instance area, instance position, and instance depth of field; It should be noted that the instance position refers to the center point coordinates of the instance, and the instance depth of field is used to represent whether the instance is in the foreground or background of the image.
[0069] It should also be noted that the weight is used to represent the importance or contribution degree of the instance in the image. The larger the weight, the higher the importance or contribution degree of the instance in the image, and the more it can represent the core content of the image.
[0070] It is worth mentioning that in addition to instance area, instance position, and instance depth of field, the weight factor may also include the number of times of category label repetition, etc.
[0071] It is not difficult to understand that the closer the instance position is to the center of the image (i.e., the smaller the distance between the instance position and the center of the image), the larger the instance area, and the higher the repetition times of the category label, the greater the corresponding calculated weight. For the instance depth of field, the weight of an instance in the foreground is greater than that of an instance in the background.
[0072] Step S124: From at least one instance, determine the instance whose weight meets the preset weight condition as the target instance.
[0073] It should be noted that the preset weight condition is a preset judgment criterion for screening target instances that can represent the core content of the image from all instances in the image. Exemplarily, the preset weight condition can be the largest weight, the weight being greater than a preset value, or the weights being ranked top three in descending order.
[0074] In this embodiment, by determining the weight factors of each instance and comprehensively calculating the weights of each instance respectively, the importance of the instance in the image is reflected by the weight. Furthermore, through the preset weight condition, the target instances that can represent the core content of the image are accurately screened, reducing the interference of other elements such as complex backgrounds in the image, and ensuring that the image-text pairs used to construct the image-text pair dataset are all high-quality target image-text pairs.
[0075] Through the comprehensive application of instance segmentation technology and weight factors, this embodiment can more precisely identify each instance in the image and its importance, reducing the interference of complex backgrounds and other elements. At the same time, the multi-dimensional weight calculation enables this embodiment to evaluate the importance of instances from multiple perspectives, avoiding the bias caused by single features, thereby better understanding the image content, providing strong support for subsequent image-text matching, and further obtaining a higher confidence level of image-text matching in complex scenarios, ensuring that the final image-text matching result is more accurate and reliable.
[0076] Based on the above embodiments, a method for constructing an image-text pair dataset according to the fourth embodiment of the present application is proposed.
[0077] In the fourth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be elaborated hereinafter.
[0078] Please refer to Figure 4 , Figure 4 which is a schematic flowchart provided for the fourth embodiment of the method for constructing an image-text pair dataset of the present application.
[0079] In this embodiment, the instance segmentation information includes a border mask. Step S200 matches the target instance with the text of the image-text pair according to the instance segmentation information of the target instance to obtain the instance matching result of the image-text pair, which may include steps S230 to S260: Step S230: Extract image features from the target instance to obtain the image feature vector of the target instance, and extract text features from the text of the text-image pair to obtain the text feature vector of the text. In this embodiment, a deep learning model can be used to extract an image feature vector that can represent its visual features from the target instance. At the same time, natural language processing technology is used to extract a text feature vector representing semantic information from the text description. These feature vectors can effectively capture the core content of the target instance and the text, providing a basis for subsequent similarity calculation.
[0080] Step S240: Calculate the vector similarity between the image feature vector and the text feature vector based on a preset vector similarity algorithm. As known to those skilled in the art, the vector similarity algorithm is an algorithm for calculating the similarity degree between vectors. Commonly used vector similarity algorithms include the cosine similarity algorithm, the Euclidean distance algorithm, etc.
[0081] In this embodiment, by calculating the vector similarity between the image feature vector and the text feature vector, the degree of semantic association between the target instance and the text description can be quantified, so as to evaluate whether they match.
[0082] Step S250: Perform weighted calculation on the vector similarity according to the weight of the target instance to obtain the instance-text similarity between the target instance and the text, where the weight of the target instance is determined based on the bounding box mask of the target instance. It should be noted that the instance-text similarity refers to the overall degree of semantic association between the target instance and the text description. Since there may be more than one target instance and the weights of different target instances in the image are different, it is necessary to perform weighted calculation on the vector similarity between the image feature vectors and the text feature vectors of each target instance through the weights of each target instance, so as to obtain the instance-text similarity that can accurately reflect the overall semantic association degree between the target instance and the text description. Furthermore, through this instance-text similarity, it is judged whether the target instance matches the text of the text-image pair to obtain the instance matching result of the text-image pair.
[0083] Step S260: When the instance-text similarity is greater than the preset similarity threshold, determine that the instance matching result of the text-image pair is a match.
[0084] It should be noted that the preset similarity threshold is a preset threshold for judging whether the instance matching result of the text-image pair is a match.
[0085] In this embodiment, when the similarity of the instance text is greater than the preset similarity threshold, it can be determined that the instance matching result of the image-text pair is a match, that is, it is determined that the matching degree between the image and the text in the image-text pair is relatively high, and the image-text pair is a high-quality image-text pair, which can be used as a target image-text pair to construct a high-quality image-text pair dataset.
[0086] This embodiment provides a more accurate image-text matching strategy by combining the similarity calculation of the image feature vector and the text feature vector, and introducing a weight factor of the target instance for weighted adjustment on this basis. This strategy not only considers the direct relevance between the target instance and the text content, but also fully considers the relative importance of the target instance in the image and its visual saliency. Therefore, compared with the traditional methods that only rely on category labels or simple frequency statistics, this embodiment can more effectively screen out high-quality image-text pairs, reduce the interference caused by complex backgrounds and other irrelevant elements, and further improve the overall quality and usability of the image-text pair dataset. In addition, by flexibly setting the preset similarity conditions, the requirements of different application scenarios can also be adapted to ensure that the selected target image-text pairs are all high-quality image-text pairs.
[0087] In addition, please refer to Figure 5 , Figure 5 which is a schematic diagram of the device structure of the hardware operating environment involved in the method for constructing an image-text pair dataset in an embodiment of the present application.
[0088] The present application also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method for constructing an image-text pair dataset in the above embodiment.
[0089] Next, refer to Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing an embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), vehicle-mounted terminals, etc., and fixed terminals such as desktop computers, or any electronic device capable of implementing the above functions. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0090] As Figure 5As shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display, a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0091] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0092] The electronic device provided by the present application adopts the method for constructing a text-image pair dataset in the above-mentioned embodiment, and can solve the technical problem of low quality of the constructed text-image pair dataset in the related art. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the method for constructing a text-image pair dataset provided by the above-mentioned embodiment, and other technical features in the electronic device are the same as those disclosed in the method of the above-mentioned embodiment, and will not be elaborated here.
[0093] It should be understood that the various parts disclosed in the present application may be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0094] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the above-mentioned claims.
[0095] In addition, the present application also provides a storage medium, which is a computer-readable storage medium and has computer-readable program instructions (i.e., computer programs) stored thereon. The computer-readable program instructions are used to execute the steps of the method for constructing a text-image pair dataset in the above-mentioned embodiments.
[0096] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical fibers, portable compact disk read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0097] The above computer-readable storage medium can be included in an electronic device; or it can exist separately without being assembled into the electronic device.
[0098] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device: performs instance segmentation on the image of the text-image pair to obtain the instance segmentation result of the image, where the instance segmentation result includes the target instance and the instance segmentation information of the target instance; matches the target instance with the text of the text-image pair according to the instance segmentation information of the target instance to obtain the instance matching result of the text-image pair; if the instance matching result is a match, determines the text-image pair as the target text-image pair, and constructs a text-image pair dataset based on at least one target text-image pair.
[0099] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network or a wide area network, or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0101] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0102] The computer-readable storage medium provided by this application stores computer-readable program instructions (i.e., computer programs) for performing the steps of the above-mentioned method for constructing a text-image pair dataset, and can solve the technical problem of the low quality of the text-image pair dataset constructed in the related art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for constructing a text-image pair dataset provided in the above embodiments, and will not be elaborated here.
[0103] In addition, an embodiment of the present application further provides a program product, which is a computer program product and includes a computer program. When the computer program is executed by a processor, the steps of the method for constructing a text-image pair data set in the above embodiment are implemented.
[0104] The computer program product provided by the present application can solve the technical problem that the quality of the constructed text-image pair data set is relatively low in the related art. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the method for constructing a text-image pair data set provided by the above embodiment, and will not be elaborated here.
[0105] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for constructing a picture-text pair dataset, characterized in that: The method comprises: Performing instance segmentation on the image of the image-text pair to obtain an instance segmentation result of the image, wherein the instance segmentation result includes a target instance and instance segmentation information of the target instance; According to the instance segmentation information of the target instance, the target instance is matched with the text of the image-text pair to obtain an instance matching result of the image-text pair; If the instance matching result is a match, the image-text pair is determined as a target image-text pair, and an image-text pair data set is constructed based on at least one of the target image-text pairs.
2. The method for constructing a picture-text pair dataset as claimed in claim 1, characterized in that: The step of performing instance segmentation on the image of the image-text pair to obtain the instance segmentation result of the image comprises: Performing instance segmentation on the image of the image-text pair to obtain at least one instance of the image and instance segmentation information of the at least one instance; Determining a target instance from the at least one instance according to instance segmentation information of the at least one instance; The target instance and the instance segmentation information of the target instance are used as the instance segmentation result of the image.
3. The method for constructing a picture-text pair data set as claimed in claim 2, characterized in that: The instance segmentation information includes a category label, and the step of determining a target instance from the at least one instance according to the instance segmentation information of the at least one instance includes: Determining a target category label from the category label of the at least one instance; From the at least one instance, an instance whose class label is a target class label is determined as a target instance.
4. The method for constructing a picture-text pair data set as claimed in claim 3, characterized in that: The step of determining a target category label from the category label of the at least one instance comprises: According to the category label of the at least one instance, obtaining by counting the number of repetitions of different category labels in the at least one instance; The category label with the highest number of repetitions is determined as the target category label.
5. The method for constructing a picture-text pair data set as claimed in claim 3, characterized in that: The instance segmentation information includes a bounding box mask, and the step of determining a target category label from the category label of the at least one instance includes: Determining an instance area of the at least one instance according to the bounding box mask of the at least one instance; Obtaining, according to the instance area and the category label of the at least one instance, instance areas corresponding to different category labels in the at least one instance; The category label with the largest corresponding instance area is determined as the target category label.
6. The method for constructing a picture-text pair data set as claimed in claim 2, characterized in that: The instance segmentation information includes a bounding box mask, and the step of determining a target instance from the at least one instance according to the instance segmentation information of the at least one instance includes: Determine a weight factor of the at least one instance according to the bounding box mask of the at least one instance, and calculate a weight of the at least one instance according to the weight factor of the at least one instance, wherein the weight factor includes at least one of an instance area, an instance position, and an instance depth of field; From the at least one instance, an instance whose weight satisfies a preset weight condition is determined as a target instance.
7. The method for constructing a picture-text pair dataset according to any one of claims 1 to 6, characterized in that: The instance segmentation information includes a category label. The step of matching the target instance with the text of the image-text pair according to the instance segmentation information of the target instance to obtain an instance matching result of the image-text pair includes: Performing text classification on the text of the image-text pair to obtain a text classification result of the text; If the category label of the target instance matches the text classification result, then the instance matching result of the image-text pair is determined to be a match.
8. The method for constructing a picture-text pair dataset according to any one of claims 1 to 6, characterized in that: The instance segmentation information includes a border mask. The step of matching the target instance with the text of the image-text pair according to the instance segmentation information of the target instance to obtain an instance matching result of the image-text pair includes: Performing image feature extraction on the target instance to obtain an image feature vector of the target instance, and performing text feature extraction on the text of the image-text pair to obtain a text feature vector of the text; Based on a preset vector similarity algorithm, the vector similarity between the image feature vector and the text feature vector is calculated; According to the weight of the target instance, weighted calculation is performed on the vector similarity to obtain the instance text similarity between the target instance and the text, wherein the weight of the target instance is determined based on the border mask of the target instance; When the instance text similarity is greater than a preset similarity threshold, it is determined that the instance matching result of the image-text pair is a match.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method for constructing a picture-text pair data set as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method for constructing a picture-text pair data set according to any one of claims 1 to 8 is implemented.
11. A program product, characterized in that The program product is a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the method for constructing a picture-text pair data set as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Image-text matching method and device, storage medium and equipment
CN110147457A
Cross-modal matching method and related device, electronic equipment and storage medium
CN115270754A
Image data set generation method, device and equipment and computer readable storage medium
CN119723240A
Image labeling method and device, equipment and storage medium
CN119851272A
Aligning unlabeled images to surrounding text
US20210209353A1