Text-based image retrieval method and device and storage medium
By segmenting and feature extraction of the image to be retrieved, combined with text feature matching, the problem of low image retrieval accuracy in the prior art is solved, and a more efficient image retrieval effect is achieved.
Patent Information
- Application Number
- CN202510072666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-20
AI Technical Summary
The existing text-based image retrieval methods cannot achieve the unity of fuzzy retrieval and fine retrieval, and cannot effectively mine multi-attribute information of images and text, resulting in poor retrieval accuracy.
By performing image segmentation processing on the image to be retrieved, global and local image features are extracted, and image enhancement features are generated through feature fusion and stitching processing; at the same time, text feature extraction is performed on the prompt text, and the feature matching results are used to determine whether the image to be retrieved is the target image.
The unity of blurred and fine retrieval of image retrieval is realized, the utilization of multi-attribute information of images and text is improved, and the accuracy of image retrieval is significantly improved.
Smart Images

Figure CN120179845A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data retrieval, and in particular to a text-based image retrieval method, device, and storage medium. Background Art
[0002] In a system platform involving video images, image retrieval is one of the very common functions. It is widely used in fields such as smart home and public security.
[0003] Currently, common image retrieval methods include retrieving a database image using a template image and retrieving a database image using a prompt text, etc. In practical applications, it is usually difficult to obtain a template image in advance. And text description provides a relatively comprehensive way to describe the attribute information of the object to be retrieved. Therefore, the retrieval method based on prompt text is more flexible and general, and has a wide range of application scenarios.
[0004] However, in the existing text-based image retrieval methods, usually the object to be retrieved in the image is first detected, and then the retrieval result is determined according to the matching result between the prompt text and the object to be retrieved. Such methods cannot achieve the unity of fuzzy retrieval and fine retrieval, and cannot mine the multi-attribute information of images and texts, resulting in poor accuracy of the retrieval method. Summary of the Invention
[0005] This application provides at least a text-based image retrieval method, device, equipment, and computer-readable storage medium.
[0006] In the first aspect of this application, a text-based image retrieval method is provided, including: performing image segmentation processing on the obtained image to be retrieved to obtain at least one sub-image to be retrieved; respectively performing image feature extraction processing on the image to be retrieved and each sub-image to be retrieved to obtain the global image feature of the image to be retrieved and the local image features of each sub-image to be retrieved; performing fusion and splicing processing according to the global image feature and each local image feature to obtain an image enhancement feature; performing text feature extraction processing on the obtained prompt text to obtain a prompt text feature; determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature.
[0007] In an embodiment, the performing image segmentation processing on the obtained image to be retrieved to obtain at least one sub-image to be retrieved includes: performing object detection processing on the image to be retrieved to obtain the target object in the image to be retrieved; performing segmentation processing on the image to be retrieved according to the component structure information of the target object to obtain at least one sub-image to be retrieved, and different sub-images to be retrieved include different object components of the target object.
[0008] In one embodiment, after segmenting the image to be retrieved according to the component structure information of the target object to obtain at least one sub-image to be retrieved, the method includes: performing feature extraction processing on the target object in the image to be retrieved to obtain the global image features; and performing feature extraction processing on the object components in each sub-image to be retrieved to obtain local image features for each sub-image.
[0009] In one embodiment, the step of performing fusion and stitching processing on the global image features and the local image features to obtain enhanced image features includes: performing feature fusion processing on the global image features and the local image features to obtain fused image features; and performing feature stitching processing on the fused image features and the local image features to obtain the enhanced image features.
[0010] In one embodiment, the step of determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the enhanced image features and the prompt text features includes: performing normalization processing on the feature matching result to obtain a feature matching probability; in response to the feature matching probability being greater than a preset matching threshold, determining the image to be retrieved as the target image; and in response to the feature matching probability being less than or equal to the preset matching threshold, performing filtering processing on the image to be retrieved.
[0011] In one embodiment, the feature matching result includes a feature similarity. The step of determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the enhanced image features and the prompt text features includes: in response to the feature similarity between the enhanced image features and the prompt text features being greater than a preset similarity threshold, determining the image to be retrieved as the target image; and in response to the feature similarity being less than or equal to the preset similarity threshold, performing filtering processing on the image to be retrieved.
[0012] In one embodiment, the step of performing text feature extraction processing on the obtained prompt text to obtain prompt text features includes: performing segmentation and recombination processing on the prompt text according to the text semantic information of the prompt text to obtain at least one local prompt text; and performing feature extraction processing on at least one local prompt text to obtain at least one prompt text feature.
[0013] In one embodiment, the step of performing segmentation and recombination processing on the prompt text according to the text semantic information of the prompt text to obtain at least one local prompt text includes: determining the component structure information of the target object in the prompt text according to the text semantic information; and performing segmentation and recombination processing on the prompt text according to the component structure information to obtain at least one local prompt text.
[0014] The second aspect of the present application provides a text-based image retrieval device, including: an image segmentation module for performing image segmentation processing on the retrieved image to be obtained, to obtain at least one sub-image to be retrieved; an image feature extraction module for performing image feature extraction processing on the retrieved image and each sub-image to be retrieved respectively, to obtain the global image feature of the retrieved image and the local image features of each sub-image to be retrieved; a feature enhancement module for performing fusion and stitching processing according to the global image feature and each local image feature, to obtain an enhanced image feature; a text feature extraction module for performing text feature extraction processing on the obtained prompt text, to obtain a prompt text feature; and an image retrieval module for determining whether the retrieved image is the target image corresponding to the prompt text according to the feature matching result between the enhanced image feature and the prompt text feature.
[0015] The third aspect of the present application provides an electronic device, including a memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the above text-based image retrieval method.
[0016] The fourth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above text-based image retrieval method is implemented.
[0017] In the above solution, by performing image segmentation processing on the retrieved image to be obtained, at least one sub-image to be retrieved is obtained, so as to perform multi-granularity analysis on the retrieved image; image feature extraction processing is respectively performed on the retrieved image and each sub-image to be retrieved, to obtain the global image feature of the retrieved image and the local image features of each sub-image to be retrieved; fusion and stitching processing is performed according to the global image feature and each local image feature, to obtain an enhanced image feature, and the fusion of multi-visual features enables the neural network model to obtain as much detailed image information as possible and improves the expression ability of visual features; text feature extraction processing is performed on the obtained prompt text, to obtain a prompt text feature, and thus image features similar in features can be found according to the prompt text feature; it is determined whether the retrieved image is the target image corresponding to the prompt text according to the feature matching result between the enhanced image feature and the prompt text feature, and thus text-based image retrieval can be realized and the image retrieval accuracy can be improved.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to illustrate the technical solutions of the present application.
[0020] Figure 1 is a schematic flowchart of an exemplary embodiment of the text-based image retrieval method of the present application;
[0021] Figure 2 is an exemplary image segmentation schematic diagram in the text-based image retrieval method of the present application;
[0022] Figure 3 is an exemplary feature fusion and stitching schematic diagram in the text-based image retrieval method of the present application;
[0023] Figure 4 is an exemplary overall flowchart of the text-based image retrieval method of the present application;
[0024] Figure 5 is an exemplary model training flowchart of the text-based image retrieval method of the present application;
[0025] Figure 6 is an exemplary text mask schematic diagram in the text-based image retrieval method of the present application;
[0026] Figure 7 is a block diagram of a text-based image retrieval device shown in an exemplary embodiment of the present application;
[0027] Figure 8 is a schematic structural diagram of an embodiment of an electronic device of the present application;
[0028] Figure 9 is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. Detailed Embodiments
[0029] The following will describe the solutions of the embodiments of the present application in detail with reference to the accompanying drawings of the specification.
[0030] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0031] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.
[0032] For ease of understanding, one of the applicable scenarios of the text-based image retrieval method of the present application is now exemplified. The text-based image retrieval method aims to retrieve the corresponding target image through a given query text (prompt text) (for example, searching among multiple base images in a database according to the prompt text, or searching among multiple objects to be retrieved in the images to be retrieved according to the prompt text). This method can be applied to a variety of fields such as smart homes and public security, and is not limited here.
[0033] Exemplarily, the method may include but is not limited to searching for images of animals such as people and pets, and / or searching for images of other objects with certain structural information. The following text mainly uses the retrieval of animal images as an example for explanation, which will not be repeated here. For example, the application may determine the target image from multiple images to be retrieved based on the prompt text, or determine the image of the area where the target object is located from an image to be retrieved based on the prompt text, which is not limited here.
[0034] In the existing text-based image retrieval methods, most of them first detect the position of the target object in the image to be retrieved through an external target detector, and then match the image of the target object based on the text information. However, this type of method cannot achieve the unification of fuzzy retrieval and fine retrieval of image retrieval, and cannot mine the multi-attribute information of the overall image and text, resulting in poor robustness and practicality of the text-image retrieval method.
[0035] See also Figure 1 , Figure 1 This is a flowchart of an exemplary embodiment of the text-based image retrieval method of the present application. Specifically, it may include the following steps:
[0036] Step S110, performing image segmentation processing on the acquired image to be retrieved to obtain at least one sub-image to be retrieved.
[0037] Among them, the image to be retrieved can be obtained in real time or from the base images in the database. There can be one or more images to be retrieved. In the actual application scenario of the image retrieval method of the present application, it can be to analyze each of the obtained multiple images to be retrieved one by one, or to select at least two images from the obtained multiple images to be retrieved for parallel analysis, thereby improving the image retrieval speed, which is not limited here. The specific execution method can be determined by the device specifications (computing power) and / or real-time operating load of the image retrieval device responsible for executing the image retrieval method, which will not be elaborated here.
[0038] Exemplarily, there can be one or more methods for performing image segmentation processing on the image to be retrieved. For example, there can be a preset image segmentation template, which can be determined in advance according to the structural information of the target object to be retrieved (such as body parts, etc.) or according to other segmentation criteria. Then, the image to be retrieved can be segmented according to the image segmentation template, or the target object (the image area where the target object is located) in the image to be retrieved can be segmented according to the image segmentation template, which is not limited here. For another example, it can be to determine the target object and the structural information of the target object in the image to be retrieved through a target detection algorithm or a target detection model, and then perform image segmentation processing on the image to be retrieved according to the structural information of the target object. Among them, it can be understood that the template-based segmentation method is fast and has low requirements for the operating environment, and the segmentation method based on the detection algorithm or the detection model has high accuracy and high requirements for the operating environment. The specific selection can be made adaptively according to the actual application scenario, which will not be elaborated here.
[0039] An example is given in combination with a specific application scenario. In a real-time detection scenario, video data of a target scenario can be collected by an image acquisition device and input into a target detection module for target detection, and then the detection result can be obtained. where N is the number of target objects detected in the current video frame. and respectively represent the upper left coordinate and the lower right coordinate of the i-th target box, f i and c i respectively represent the confidence level and the category of the i-th detected target.
[0040] Further, after performing image segmentation processing on the image to be retrieved, at least one sub-image to be retrieved can be obtained. Each sub-image to be retrieved may respectively include object components of the target object. For example, if the image segmentation template is determined based on the upper body and lower body of the target object, which is equivalent to the image segmentation boundary being the junction between the upper body and the lower body (such as the waistline position), then two sub-images to be retrieved can be obtained after segmentation, and the two sub-images to be retrieved respectively include the upper body image and the lower body image of the target object. It should also be noted that the method of performing image segmentation processing on the image to be retrieved can also be to copy an image to be retrieved for image segmentation, and then at least one sub-image to be retrieved and the original image to be retrieved can be obtained.
[0041] Step S120: Perform image feature extraction processing on the image to be retrieved and each sub-image to be retrieved respectively to obtain the global image features of the image to be retrieved and the local image features of each sub-image to be retrieved.
[0042] Illustrated in combination with the foregoing steps, after obtaining the sub-images to be retrieved of the image to be retrieved, image feature extraction processing can be performed on the image to be retrieved and each sub-image to be retrieved respectively to obtain the global image features of the image to be retrieved and the local image features of each sub-image to be retrieved. Among them, performing image feature extraction processing on the image to be retrieved can similarly be performing image feature extraction processing on the target object in the image to be retrieved, which will not be elaborated here.
[0043] Exemplarily, the method of image feature extraction processing can include but is not limited to being implemented using models such as an Image Encoder. For example, the image encoding module can be a ViT-B model (Vision Transformer-Base model) with a twelve-layer network structure, which may include but is not limited to having a SelfAttention module and a Feed Forward module, etc. The input can be the image to be retrieved (full-body image) and the sub-images to be retrieved (upper body image and lower body image) obtained after being processed through the foregoing steps. The image encoding module encodes the input images into Image tokens in sequence, and then inputs them into the model to extract image features and obtain multi-granularity image features (global image features and local image features).
[0044] Step S130: Perform fusion and stitching processing based on the global image features and each local image feature to obtain enhanced image features.
[0045] The above steps are referred to for illustration. The global image features can represent the overall features of the target object, while the local image features can represent the detailed features of the target object. Therefore, after obtaining the global image features and the local image features, the global image features and each of the local image features can be subjected to fusion and stitching processing to obtain image enhancement features. The image enhancement features include both the overall visual analysis of the target object and the detailed visual analysis of the target object, improving the expression ability of the visual features.
[0046] Exemplarily, in the specific implementation process, it may be to perform feature stitching on the global image features and each of the local image features to obtain image enhancement features. Or it may be to perform feature fusion on the global image features and each of the local image features to obtain image fusion features; then perform feature stitching on the image fusion features and each of the local image features to obtain image enhancement features.
[0047] Step S140: Perform text feature extraction processing on the obtained prompt text to obtain prompt text features.
[0048] Among them, the prompt text is the text used to retrieve images. The method for obtaining the prompt text may include, but is not limited to, generating the prompt text according to the received text information, or generating the prompt text through audio-text conversion technology based on the received audio information, which is not limited here.
[0049] Exemplarily, the description of the image feature extraction method in the foregoing steps can be referred to by analogy. In this application, a text encoding module (Text Encoder) can be used to perform text feature extraction processing on the prompt text to obtain prompt text features. Among them, the text encoding module can be a six-layer Transformer model, whose weight initialization comes from the BERT-Base model. The input data is the obtained prompt text, and then it is subjected to normalization processing. After that, byte pair encoding is used to encode the text information into text tokens, and then it is input into the Transformer model for text feature extraction to obtain prompt text features.
[0050] Step S150: Determine whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement features and the prompt text features.
[0051] The above steps are referred to for illustration. The prompt text features represent the features of the target object to be retrieved, while the image enhancement features represent the features included in the image to be retrieved. The prompt text features and the image enhancement features are matched through image-text matching technology to obtain a feature matching result, and then it can be determined whether the image to be retrieved is the target image corresponding to the prompt text (equivalent to the image that the prompt text needs to search for).
[0052] Among them, there may be one or more methods for matching image and text features, which are not limited here. For example, the image and text matching process may be to calculate the feature similarity between the image enhancement feature and the prompt text feature, compare the feature similarity with the preset similarity threshold, thereby obtaining a feature matching result, and then determining whether the image to be retrieved is the target image corresponding to the prompt text. For another example, the image and text matching process may be performed by a neural network, and the image enhancement feature and the prompt text feature are input into a pre-trained neural network for feature fusion and other processing to obtain an image and text fusion feature, and then the fully connected layer in the neural network may be used to predict whether the image and text pair of the image and text fusion feature matches. Specifically, the neural network may output a matching probability according to the input image and text fusion feature, and compare its matching probability with the preset matching threshold, thereby obtaining a feature matching result, and then determining whether the image to be retrieved is the target image corresponding to the prompt text.
[0053] It can be seen that the present application obtains at least one sub-image to be retrieved by performing image segmentation processing on the acquired image to be retrieved, so as to perform multi-granularity analysis on the image to be retrieved; performs image feature extraction processing on the image to be retrieved and each sub-image to be retrieved respectively, so as to obtain the global image feature of the image to be retrieved and the local image feature of each sub-image to be retrieved; performs fusion and splicing processing based on the global image feature and each local image feature to obtain image enhancement feature, and the fusion of multiple visual features enables the neural network model to obtain more detailed image information as much as possible, thereby improving the expression ability of visual features; performs text feature extraction processing on the acquired prompt text, and obtains prompt text feature, thereby searching for image features with similar features based on the prompt text feature; determines whether the image to be retrieved is the target image corresponding to the prompt text based on the feature matching result between the image enhancement feature and the prompt text feature, thereby realizing text-based image retrieval and improving image retrieval accuracy.
[0054] On the basis of the above embodiments, the embodiment of the present application exemplarily illustrates the steps of performing image segmentation processing on the acquired image to be retrieved to obtain at least one sub-image to be retrieved. Specifically, the method of this embodiment includes the following steps:
[0055] The image to be retrieved is subjected to object detection processing to obtain the target object in the image to be retrieved; the image to be retrieved is segmented according to the component structure information of the target object to obtain at least one sub-image to be retrieved, and different sub-images to be retrieved include different object components of the target object.
[0056] In combination with the description of the foregoing embodiments, the method for image segmentation processing in the method provided by this application may at least include: first, perform object detection processing on the image to be retrieved to obtain the target object in the image to be retrieved. Then, the component structure information of the target object can be identified or, according to a preset image segmentation template, the image to be retrieved can be segmented according to the component structure information of the target object to obtain at least one sub-image to be retrieved. Different sub-images to be retrieved include different object components of the target object (for example, it can be divided according to the segmentation standard of the upper body and the lower body, or according to the segmentation standard of the torso and the limbs, or other segmentation standards (such as a 50-50 segmentation ratio, etc.), which is not limited here).
[0057] Exemplarily, it can be referred to as Figure 2 shown in Figure 2 is an exemplary image segmentation schematic diagram in the text-based image retrieval method of this application. In the actual application process, the obtained image to be retrieved can be segmented to obtain a segmented image. Then, the segmented image can be resized according to a preset image size to obtain a sub-image to be retrieved. For example, the segmented upper body image is resized so that its image size is 256*256, and it is used as the sub-image to be retrieved. Similarly, the image to be retrieved can also be resized so that the image sizes of the image to be retrieved and the sub-image to be retrieved are aligned.
[0058] Based on the above embodiments, the embodiments of this application will describe the steps after segmenting the image to be retrieved according to the component structure information of the target object to obtain at least one sub-image to be retrieved. Specifically, the method of this embodiment includes the following steps:
[0059] Perform feature extraction processing on the target object in the image to be retrieved to obtain global image features; perform feature extraction processing on the object components in each sub-image to be retrieved to obtain local image features.
[0060] In combination with the foregoing steps, the image to be retrieved contains the target object, and the sub-image to be retrieved contains the object components of the target object. Therefore, after completing the image segmentation, the image to be retrieved and the sub-images to be retrieved can be input into the image encoding module. The image encoding module performs feature extraction processing on the target object in the image to be retrieved to obtain global image features; and performs feature extraction processing on the object components in each sub-image to be retrieved to obtain local image features.
[0061] Based on the above embodiments, the embodiments of this application will describe the steps of performing fusion and splicing processing on the global image features and each local image feature to obtain image enhancement features. Specifically, the method of this embodiment includes the following steps:
[0062] Perform feature fusion processing on the global image feature and each local image feature to obtain an image fusion feature; perform feature splicing processing on the image fusion feature and each local image feature to obtain an image enhancement feature.
[0063] In combination with the foregoing embodiments, in the present application, the method of combining the global image feature and the local image feature to enhance the visual feature detail information can be referred to as Figure 3 as shown Figure 3 is an exemplary feature fusion and splicing schematic diagram in the text-based image retrieval method of the present application. Performing feature fusion processing on the global image feature and each local image feature can specifically be adding the global image feature and each local image feature to obtain an image fusion feature. Then, perform feature splicing processing on the image fusion feature and each local image feature to obtain an image enhancement feature. The image enhancement feature enables the neural network model for subsequent data processing to obtain as much detailed image information as possible, enhancing the expression ability of the visual feature. Among them, the local image feature can be copied so that a part of the local image feature can be used for feature fusion, and another part of the local image feature can be used for feature splicing.
[0064] Based on the above embodiments, the embodiments of the present application describe the step of determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature. Specifically, the method of this embodiment includes the following steps:
[0065] Perform normalization processing on the feature matching result to obtain a feature matching probability; in response to the feature matching probability being greater than a preset matching threshold, determine the image to be retrieved as the target image; in response to the feature matching probability being less than or equal to the preset matching threshold, perform filtering processing on the image to be retrieved.
[0066] In combination with the foregoing embodiments, the method of determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature in the method of the present application can be to use a neural network to match the image enhancement feature and the prompt text feature.
[0067] Specifically, the image enhancement features and the prompt text features are input into the CrossEncoder for feature fusion processing to obtain the image-text fusion features. The image-text fusion features are input into the Image Text Match (IMT) for image-text matching processing, and then the feature matching result (image-text matching result) output by the image-text matching module can be obtained. The feature matching result is input into the Sigmoid function (Sigmoid can map the value to between 0 and 1 through a smooth step function, which is applicable to binary classification problems), that is, the image-text matching result can be normalized to between 0 and 1 to obtain the feature matching probability (image-text matching probability). Then, according to the comparison result between the preset matching threshold and the value of the feature matching probability, the image-text matching results with the feature matching probability less than the preset matching threshold are deleted, and the image-text matching results with the feature matching probability greater than or equal to the preset matching threshold are retained. Optionally, the image region where the successfully matched target object in the image to be retrieved can also be visually displayed (such as displaying the bounding box of the target object, etc.), and the final image retrieval result is returned.
[0068] In summary, the overall process of the method of this application in the application process can be referred to as Figure 4 shown in Figure 4 FIG. is an exemplary overall process schematic diagram of the text-based image retrieval method of this application, showing one of the implementable ways of the method of this application, and is also equivalent to the inference process when applying the image retrieval model of this application. The specific steps can be understood by referring to the description of the foregoing embodiments, and will not be elaborated here.
[0069] For another example, during actual image retrieval, the input prompt text can be copied three times, and then the image data to be retrieved is subjected to image preprocessing operations to obtain multi-granularity images (upper body image, lower body image of the human body, and overall image). Then the multi-granularity images are input into the image encoding module to extract image features, and multi-visual feature fusion is performed to obtain image enhancement features. The prompt text information is input into the text encoding module to extract prompt text features. Then, they are input into the cross-encoder together with the image enhancement features, and finally input into the image-text matching module to obtain the matching results of the prompt text and the multi-granularity images (upper body image, lower body image, and overall image). Each image-text matching result is input into the Sigmoid function to normalize the image-text matching result to between 0 and 1. Then, filtering is performed according to the preset matching threshold, the image-text matching results below the threshold are deleted, the image-text matching results above the threshold are retained, and the image region where the successfully matched target object is located is visualized, and the final image retrieval result is returned.
[0070] Based on the above embodiments, the embodiments of the present application illustrate the steps of determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature. Among them, the feature matching result includes the feature similarity. Specifically, the method of this embodiment includes the following steps:
[0071] In response to the feature similarity between the image enhancement feature and the prompt text feature being greater than the preset similarity threshold, determine the image to be retrieved as the target image; in response to the feature similarity being less than or equal to the preset similarity threshold, perform filtering processing on the image to be retrieved.
[0072] Combined with the foregoing embodiments for illustration, the method for determining the feature matching result between the image enhancement feature and the prompt text feature can also be to calculate the feature similarity between the image enhancement feature and the prompt text feature. The method for determining whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result can also be to judge whether the image to be retrieved is the target image corresponding to the prompt text according to the numerical comparison result between the feature similarity and the preset similarity threshold.
[0073] Exemplarily, if the feature similarity is greater than the preset similarity threshold, it indicates that the image to be retrieved and the prompt text are similar, so it can be determined that the image to be retrieved is the target image corresponding to the prompt text. If the feature similarity is less than or equal to the preset similarity threshold, it indicates that the image to be retrieved and the prompt text are not similar, so it can be determined that the image to be retrieved is not the target image corresponding to the prompt text, perform filtering processing on the image to be retrieved, and then other images to be retrieved can be continuously obtained for comparison with the prompt text.
[0074] Based on the above embodiments, the embodiments of the present application exemplarily illustrate the steps of performing text feature extraction processing on the obtained prompt text to obtain prompt text features. Specifically, the method of this embodiment includes the following steps:
[0075] Perform segmentation and recombination processing on the prompt text according to the text semantic information of the prompt text to obtain at least one local prompt text; respectively perform feature extraction processing on at least one local prompt text to obtain at least one prompt text feature.
[0076] Combined with the foregoing embodiments for illustration, for the obtained image to be retrieved, it can be divided into multiple sub-images to be retrieved, and then the local image features of the sub-images to be retrieved are fused with the global image features of the image to be retrieved to enhance the detail expression ability of the visual features.
[0077] Similarly, in this embodiment, for the received prompt text, the prompt text can be input into a text encoding model to extract its features as the prompt text features; alternatively, a pre-trained large language model can be used to segment and reorganize the prompt text according to the text semantic information of the prompt text to obtain local prompt texts. Then, the prompt text and the local prompt texts are input into the text encoding model for feature extraction to obtain the prompt text features. Among them, the text semantic information refers to the semantics represented by the prompt text.
[0078] Segmentation according to the text semantic information of the prompt text may result in one local prompt text or multiple local prompt texts. If the local prompt text is the same as the prompt text, the text features of the local prompt text or the prompt text are extracted to obtain the prompt text features. If the local prompt text is different from the prompt text, the text features of the prompt text and the text features of each local prompt text are respectively extracted to obtain the prompt text features.
[0079] Exemplarily, for the image retrieval application scenario, the prompt text is usually some attribute descriptions of the target object to be retrieved. By analyzing the semantic information of the prompt text, it can be known that these attribute descriptions in the prompt text are for one part or multiple parts of the target object, and thus the prompt text can be segmented into one local prompt text or multiple local prompt texts accordingly.
[0080] For example, taking Chinese and English scenarios as examples (it can actually be applied to more language scenarios, not limited here), an original prompt text is "The man has short black spiked hair. He is wearing a purple short sleeved polo, black shorts and white socks with black tennis shoes." or "This person has short black hair. He is wearing a purple short-sleeved polo shirt, black shorts and white socks and black tennis shoes.". The large language model can know from the text semantic information of the prompt text that the prompt text describes the upper body clothing attributes and lower body clothing attributes of "The man". Therefore, the large language model can divide the original prompt text into the upper body partial prompt text "The man has short black spiked hair. He is wearing a purple short sleeved polo." or "This person has short black hair. He is wearing a purple short sleeved polo shirt." and the lower body partial prompt text "He has on black shorts and white sockswith black tennis shoes." or "He is wearing black shorts and white socks and black tennis shoes". Then the upper body part prompt text, the lower body part prompt text and the whole body prompt text (original prompt text) are input into the text encoding module for text feature extraction, and the prompt text feature can be obtained. In addition, the prompt text can also include the attributes such as the posture, age, gender, color, etc. of the target object, which will not be described in detail here.
[0081] The text features and image features obtained according to the above method are used for image-text matching retrieval, which can not only analyze and search based on the text description and image description of the whole body, but also focus on the text description and image description of half body or more details, thereby improving the image retrieval accuracy.
[0082] Optionally, when determining the segmentation criteria for image segmentation, it can also be determined according to the text semantic information of the prompt text. For example, through the above method, the large language model can determine that the prompt text includes a text description of the upper body of the target object and a text description of the lower body of the target object according to the text semantic information of the prompt text. Therefore, the corresponding image segmentation criteria can be determined according to its text semantic information, so that the image to be retrieved is segmented into an upper body sub-image to be retrieved and a lower body sub-image to be retrieved. It should be noted that such optional methods are applicable to scenarios where the prompt text is obtained first and then image segmentation is performed. The methods in the foregoing embodiments can be applicable to scenarios where the prompt text is obtained first and then image segmentation is performed, and scenarios where image segmentation is performed first and then the prompt text is obtained (in some scenarios, the base images in the database can be segmented and stored in advance), which is not limited here.
[0083] Based on the above embodiments, the embodiments of the present application illustrate the step of performing segmentation and recombination processing on the prompt text according to the text semantic information of the prompt text to obtain at least one local prompt text. Specifically, the method of this embodiment includes the following steps:
[0084] Determine the component structure information of the target object in the prompt text according to the text semantic information; perform segmentation and recombination processing on the prompt text according to the component structure information to obtain at least one local prompt text.
[0085] It will be described in combination with the foregoing embodiments. Similarly, reference may be made to the description of the image segmentation method in the foregoing embodiments. In the prompt text, the component structure information of the target object in the prompt text can also be determined according to its text-to-speech information. For example, the prompt text in the foregoing embodiment "The man has short black spiked hair. He is wearing a purple short sleeved polo, black shorts and white socks with black tennis shoes." For the target object "The man", it can be known that the upper body text description of the target object is "The man has short black spiked hair. He is wearing a purple short sleeved polo", and the lower body text description is "black shorts and white socks with black tennis shoes." Therefore, the upper body partial prompt text "The man has short black spiked hair. He is wearing a purple short sleeved polo." and the lower body partial prompt text "He has on black shorts and white socks with black tennis shoes." can be generated accordingly.
[0086] It should also be noted that the segmentation standard for the text segmentation method can be similarly referred to the description of the image segmentation method in the foregoing embodiments, and its segmentation standard can also be adaptively adjusted according to the specific application scenario, not limited to the upper and lower body segmentation methods, which will not be elaborated here.
[0087] Based on the above embodiments, the embodiments of the present application can also exemplarily illustrate the neural network training process involved in the present application. Exemplarily, in the training process of the image retrieval model of the present application, an Image Text Contrast (IMC) module and a Masked Color Modeling (MCM) module can be introduced to assist in model training, and the above modules are discarded during the actual inference process of the image retrieval model.
[0088] Exemplarily, reference can be made to Figure 5 as shown Figure 5 is a schematic diagram of an exemplary model training process in the text-based image retrieval method of the present application. Some of the processes can be referred to the description of the application process in the foregoing embodiments, which will not be elaborated here.
[0089] Specifically, during the model training process, it is also necessary to obtain the images to be retrieved and the prompt texts as training data. Similarly, it is also necessary to preprocess the images to be retrieved and the prompt texts (such as the segmentation processing, size adjustment, text generation, etc. in the foregoing embodiments), to obtain the images to be retrieved, the sub-images to be retrieved, as well as the prompt texts and the local prompt texts.
[0090] Then, on the one hand, the images to be retrieved and the sub-images to be retrieved can be input into the image encoding module for feature extraction to obtain global image features and local image features.
[0091] On the other hand, the prompt texts and the local prompt texts can be input into the text encoding module for feature extraction to obtain prompt text features (for the convenience of description, hereinafter, the features corresponding to the prompt texts will be referred to as global text features, and the features corresponding to the local prompt texts will be referred to as local text features). The image features and text features corresponding to multiple granularities are input into the image-text contrast module for image-text contrast learning to align the image and text features. After that, the aligned image and text features are input into the multi-modal encoding module (image-text encoding module) for image and text feature fusion, and then input into the image-text matching module for image-text matching to obtain the image-text matching result.
[0092] Specifically, the image-text contrast module uses the method of contrast learning to align the image features and text features. Its input is the image features passing through the image encoding module and the text features obtained through the text encoding module. The image encoding module, the text encoding module, and the image-text contrast module can also be summarized as an image alignment module, which will not be elaborated here. Among them, the process of image-text alignment is briefly to introduce an image-text contrast learning loss function for the embedded features of the image and the embedded features of the text, and align the features of the image and the text in advance before the image and text feature fusion. First, the visual features obtained through the image encoding module and the text features obtained through the text encoding module are respectively passed through the linear projection layer to obtain 1×256-dimensional vectors, and then the image-text similarity s = g i (F l ) T g t (F T ) is calculated, where g i , g t are the linear projection layers. Then, for each image and text, the Softmax-normalized image-to-text and text-to-image similarities are calculated, and then the cross-entropy loss is calculated with the pre-set true label.
[0093] The local image-text features and the global image-text features are aligned respectively. After alignment, the image-text features alleviate the problem that the image features and the text features are difficult to interact in their respective spaces, improve the interaction of the image-text features, and make it easier for the multi-modal encoding module to perform cross-modal learning. In addition, the local and global image-text features are aligned respectively, which not only ensures the independence of a single granularity but also realizes the unity of coarse-grained and fine-grained. Then, the image-text matching module fuses the aligned image features, and then inputs them together with the aligned text features into the multi-modal encoding module for feature fusion. Finally, a fully connected layer can be used to predict whether the image features and the text features in the image-text pair match or not.
[0094] On the other hand, based on the above method, the prompt text and the local prompt text can also be input into the attribute masking module. The attribute masking module can be used to mask the attribute information of the target object in the text (the attribute can include but is not limited to one or more of the attributes such as the pose, age, gender, color, etc. of the target object). For the convenience of description, the following mainly takes the color of the target object (such as the clothing color) as an example for illustration. It is equivalent to that the attribute masking module can also be a color masking module. After the text is input into the masking module, the masked text can be obtained. Then, the masked text is input into the text encoding module for feature extraction, and the masked text features can be obtained. The corresponding multi-granularity image features and the masked text features are input into the multi-modal encoding module (image-text encoding module) for image-text feature fusion, and then input into the masked color modeling module MCM (if other attributes are masked, it can be similarly replaced with other masked attribute modeling modules) for masked color attribute prediction, thereby further deepening the image-text feature fusion effect. For the specific masked image modeling method, existing neural network pre-training learning methods can be referred to, which will not be elaborated here.
[0095] For example, it can be referred to as Figure 6 shown Figure 6This is an exemplary text mask schematic diagram in the text-based image retrieval method of this application. After inputting the prompt text "This person has short black hair. He is wearing a purple short-sleeved polo shirt, black shorts, white socks, and black tennis shoes." into the color mask module, the masked text "This person has short [MASK] hair. He is wearing a [MASK] short-sleeved polo shirt, [MASK] shorts, [MASK] socks, and [MASK] tennis shoes." can be obtained. Similarly, the local prompt text can also be masked in the same way, which will not be elaborated here. It should be noted that in specific application scenarios, the prompt text for image retrieval usually reflects the appearance information of the target object. Therefore, color attributes are even more crucial factors for appearance retrieval. Different target objects can be quickly identified and distinguished according to different colors. Thus, a mask color modeling module is proposed to enhance the perceptual correspondence ability of image-text information. During the color masking process, common colors can be counted and stored in the color space. Then, according to the colors in the color space, the text description can be color-masked.
[0096] It should be noted that the text features in mask color modeling are different from the text features in the image-text alignment module and the image-text matching module on the other hand. The input text information needs to undergo color attribute masking operations to mask the colors in the text information, and then the masked text features are extracted by inputting it into the text encoding module. After that, it is input into the multi-modal encoding module together with the image enhancement features after multi-visual feature fusion for color mask image-text feature fusion. Finally, it is input into the mask color attribute modeling MCM to predict the masked color attributes, which improves the association ability between the image and the text, realizes more refined semantic alignment, and improves the accuracy of image retrieval.
[0097] In summary, the text-based image retrieval method of this application can achieve multi-granularity retrieval. Through the association of the image to be retrieved, the sub-image to be retrieved, the global prompt text, and the local prompt text, the unification of coarse-grained fuzzy retrieval and fine-grained precise retrieval is completed, improving the practical application scope of image retrieval. And the local image features and the global image features are added and fused, and the fused features and the local image features are concatenated to obtain enhanced visual features. The multi-visual feature fusion enables the model to obtain as much detailed image information as possible, improving the accuracy of image retrieval. Moreover, during the model training process, taking the attribute information (such as color attributes) in the text description as the core, masking the color attributes in the prompt text according to the text color space, and predicting the color attributes in the masked text through mask color attribute modeling, improves the association ability between the image and the text, realizes more refined semantic alignment, and improves the accuracy of image retrieval.
[0098] Further, it should be noted that the execution subject of the text-based image retrieval method can be a text-based image retrieval device. For example, the text-based image retrieval method can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the text-based image retrieval method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0099] Figure 7 It is a block diagram of a text-based image retrieval device shown in an exemplary embodiment of the present application. As Figure 7 shown, the exemplary text-based image retrieval device 700 includes: an image segmentation module 710, an image feature extraction module 720, a feature enhancement module 730, a text feature extraction module 740, and an image retrieval module 750. Specifically:
[0100] The image segmentation module 710 is configured to perform image segmentation processing on the acquired image to be retrieved, so as to obtain at least one sub-image to be retrieved.
[0101] The image feature extraction module 720 is configured to perform image feature extraction processing on the image to be retrieved and each sub-image to be retrieved respectively, so as to obtain the global image feature of the image to be retrieved and the local image features of each sub-image to be retrieved.
[0102] The feature enhancement module 730 is configured to perform fusion splicing processing according to the global image feature and each local image feature, so as to obtain an enhanced image feature.
[0103] The text feature extraction module 740 is configured to perform text feature extraction processing on the acquired prompt text, so as to obtain a prompt text feature.
[0104] The image retrieval module 750 is configured to determine whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the enhanced image feature and the prompt text feature.
[0105] In the exemplary text-based image retrieval device, through image segmentation processing on the acquired image to be retrieved, at least one sub-image to be retrieved is obtained for multi-granularity analysis of the image to be retrieved; image feature extraction processing is respectively performed on the image to be retrieved and each sub-image to be retrieved to obtain the global image feature of the image to be retrieved and the local image features of each sub-image to be retrieved; fusion and stitching processing is performed according to the global image feature and each local image feature to obtain an image enhancement feature. The fusion of multiple visual features enables the neural network model to obtain as much detailed image information as possible and improves the expression ability of visual features; text feature extraction processing is performed on the acquired prompt text to obtain a prompt text feature, and thus image features with similar features can be searched according to the prompt text feature; it is determined whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature, thereby enabling text-based image retrieval and improving the accuracy of image retrieval.
[0106] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment and will not be elaborated here. In practical applications, the device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0107] Among them, the functions of each module can be referred to in the embodiment of the text-based image retrieval method and will not be elaborated here.
[0108] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of an electronic device of the present application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is used to execute program instructions stored in the memory 101 to implement the steps in any of the above embodiments of the text-based image retrieval method. In a specific implementation scenario, the electronic device 100 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 100 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited here.
[0109] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps in any of the above-described embodiments of the text-based image retrieval method. The processor 102 may also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 102 may be implemented jointly by integrated circuit chips.
[0110] In this exemplary electronic device, by performing image segmentation processing on the acquired image to be retrieved, at least one sub-image to be retrieved is obtained for multi-granularity analysis of the image to be retrieved; image feature extraction processing is respectively performed on the image to be retrieved and each sub-image to be retrieved to obtain the global image feature of the image to be retrieved and the local image features of each sub-image to be retrieved; fusion stitching processing is performed based on the global image feature and each local image feature to obtain an image enhancement feature. The fusion of multiple visual features enables the neural network model to obtain as much detailed image information as possible and improves the expression ability of visual features; text feature extraction processing is performed on the acquired prompt text to obtain prompt text features, whereby image features similar to the features can be searched according to the prompt text features; it is determined whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature, thereby enabling text-based image retrieval and improving the accuracy of image retrieval.
[0111] Please refer to Figure 9 , Figure 9 is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 110 stores program instructions 111 that can be run by a processor, and the program instructions 111 are used to implement the steps in any of the above-described embodiments of the text-based image retrieval method.
[0112] In the exemplary storage medium, by running the program instructions in the storage medium, image segmentation processing is performed on the acquired image to be retrieved, and at least one sub-image to be retrieved is obtained for multi-granularity analysis of the image to be retrieved; image feature extraction processing is respectively performed on the image to be retrieved and each sub-image to be retrieved to obtain the global image features of the image to be retrieved and the local image features of each sub-image to be retrieved; fusion stitching processing is performed according to the global image features and each local image feature to obtain image enhancement features. The fusion of multi-visual features enables the neural network model to obtain as much detailed image information as possible and improves the expression ability of visual features; text feature extraction processing is performed on the acquired prompt text to obtain prompt text features, and thus image features with similar features can be found according to the prompt text features; it is determined whether the image to be retrieved is the target image corresponding to the prompt text according to the feature matching result between the image enhancement features and the prompt text features, and thus image retrieval based on text can be realized and the image retrieval accuracy can be improved.
[0113] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0114] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated here.
[0115] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0116] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
Claims
1. A text-based image retrieval method, characterized in that: The method comprises: Performing image segmentation processing on the acquired image to be retrieved to obtain at least one sub-image to be retrieved; Performing image feature extraction processing on the image to be retrieved and each sub-image to be retrieved respectively, to obtain the global image feature of the image to be retrieved and the local image feature of each sub-image to be retrieved; Perform fusion and splicing processing according to the global image features and each local image feature to obtain an image enhancement feature; Performing text feature extraction processing on the acquired prompt text to obtain prompt text features; It is determined whether the image to be retrieved is a target image corresponding to the prompt text according to a feature matching result between the image enhancement feature and the prompt text feature.
2. The method according to claim 1, characterized in that The step of performing image segmentation processing on the acquired image to be retrieved to obtain at least one sub-image to be retrieved includes: Performing object detection processing on the image to be retrieved to obtain a target object in the image to be retrieved; The image to be retrieved is segmented according to the component structure information of the target object to obtain at least one sub-image to be retrieved, and different sub-images to be retrieved include different object components of the target object.
3. The method according to claim 2, characterized in that After the image to be retrieved is segmented according to the component structure information of the target object to obtain at least one sub-image to be retrieved, the method includes: Performing feature extraction processing on the target object in the image to be retrieved to obtain the global image feature; Feature extraction is performed on each object component in each sub-image to be retrieved to obtain each local image feature.
4. The method according to claim 1, characterized in that The fusing and splicing processing is performed according to the global image features and each local image feature to obtain the image enhancement feature, including: Performing feature fusion processing on the global image features and each local image feature to obtain an image fusion feature; The image fusion feature and each local image feature are subjected to feature splicing processing to obtain the image enhancement feature.
5. The method according to claim 1, characterized in that: The determining, based on the feature matching result between the image enhancement feature and the prompt text feature, whether the image to be retrieved is a target image corresponding to the prompt text comprises: Normalizing the feature matching results to obtain feature matching probability; In response to the feature matching probability being greater than a preset matching threshold, determining the image to be retrieved as the target image; In response to the feature matching probability being less than or equal to the preset matching threshold, filtering the image to be retrieved.
6. The method according to claim 1, characterized in that The feature matching result includes feature similarity, and the step of determining whether the image to be retrieved is a target image corresponding to the prompt text according to the feature matching result between the image enhancement feature and the prompt text feature includes: In response to the feature similarity between the image enhancement feature and the prompt text feature being greater than a preset similarity threshold, determining the image to be retrieved as the target image; In response to the feature similarity being less than or equal to the preset similarity threshold, filtering the image to be retrieved.
7. The method according to claim 1, characterized in that The acquired prompt text is subjected to text feature extraction processing to obtain prompt text features, including: Segmenting and reorganizing the prompt text according to the text semantic information of the prompt text to obtain at least one partial prompt text; Perform feature extraction processing on at least one local prompt text to obtain at least one prompt text feature.
8. The method according to claim 7, characterized in that The segmentation and reorganization of the prompt text according to the text semantic information of the prompt text to obtain at least one partial prompt text includes: Determining component structure information of a target object in the prompt text according to the text semantic information; The prompt text is segmented and reorganized according to the component structure information to obtain at least one local prompt text.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.