Entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompt

By adopting entity inference and distillation methods based on heterogeneous prompts of scene text in cross-modal retrieval, the problems of low retrieval performance and semantic ambiguity in the case of scene text in the prior art are solved, and a more efficient and robust multimodal entity understanding and retrieval effect is achieved.

CN120144804APending Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510419675.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the case of scene text, the cross-modal graphic and text retrieval performance is low, and there are problems of local noise and semantic ambiguity.

Method used

The entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous cues are adopted. Through the discriminant entity reasoning module and the perceptual entity distillation module, multimodal entity reasoning and robust representation learning are performed using visual and text heterogeneous cues strategies.

Benefits of technology

It improves the accuracy and robustness of cross-modal retrieval, can more effectively understand and extract valuable entity information in scene text, and reduces local noise and semantic ambiguity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144804A_ABST
    Figure CN120144804A_ABST
Patent Text Reader

Abstract

The invention discloses an entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompt. Firstly, a visual and text heterogeneous prompt reasoning strategy is determined, then feature extraction is performed on an input image-text pair through a feature extraction module, and valuable scene text information is extracted from OCR features extracted from an image encoder through a perception entity distillation module so as to learn better visual representation. Cross-modal retrieval is realized by calculating the similarity between visual representation and text features, and finally, the overall framework is trained by utilizing the triple loss sensed by OCR (Optical Character Recognition). On the basis of a heterogeneous prompt reasoning strategy, the scene texts in the images are aligned with the scene texts in the titles, so that the corresponding relation between the images and the titles is enhanced in a more detailed manner. Through the perception entity distillation module, beneficial information of the scene text at the image end is distilled and extracted, so that more robust representation is learned, and efficient scene text perception cross-modal retrieval is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer science and informatics, particularly information retrieval and text processing technologies, and specifically relates to an entity reasoning and distilled cross-modal retrieval method based on scene text heterogeneous cues. Background Art

[0002] The purpose of cross-modal retrieval is to retrieve relevant images in a database through text queries, or to retrieve relevant text in the database through images. Recently, researchers have made significant progress through methods such as globally aligning images and texts or locally aligning salient image regions and corresponding words using object detectors. However, these object detector-based methods cannot capture the text appearing in the images, resulting in poor performance in the presence of scene text (as shown in Figure 1 ). To address this issue, the industry has tried many solutions to explore cross-modal image-text retrieval under scene text perception to comprehensively understand visual semantics. Recent work has attempted to fully utilize scene text from both global alignment and local alignment directions, but there is still a problem with independent scene text lacking visual context. Specifically, local alignment methods align the text OCR from the caption with visual components at a fine-grained level. Although these methods are helpful for interpretation, overemphasis on local regions often leads to sensitivity to local noise. Correspondingly, the goal of global alignment methods is to globally align text queries with the visual appearance and scene text of the image. Researchers reason about the relationships between scene texts through a graph convolutional network (GCN) to obtain high-level semantic information. Although the relationships between scene texts are important, focusing only on them without other visual entities may lead to ambiguity problems. To overcome this limitation, another group of researchers regards scene text as a third modality outside of images and texts. Scene text is fused with other visual components through fusion tokens to aggregate key visual entities and scene text semantics, and then globally aligned with the caption. This method provides efficient global alignment, but due to its overly simple global matching, the accuracy is limited. Therefore, the above problems urgently need to be improved. Summary of the Invention

[0003] Aiming at the problems of local noise and semantic ambiguity caused in the above image-text retrieval technology in the presence of scene text, the present invention provides an entity reasoning and distilled cross-modal retrieval method based on scene text heterogeneous cues, and proposes a heterogeneous cue-guided entity reasoning and distilled network to reason discriminative multi-modal entities and learn an attribute-centered representation method.

[0004] In fact, humans are good at performing cross-modal retrieval of scene text perception by carefully comparing discriminative entities in images and text, especially scene text, that is, discovering whether there is shared scene text in both the image and the corresponding text. In addition, humans can comprehensively understand the scene text in the image based on the context information of the image and extract valuable entity information from it, which helps to alleviate semantic ambiguity. Inspired by this, the present invention believes that: (1) Aligning the discriminative scene text in images and text is beneficial for accurate fine-grained retrieval. (2) Utilizing multi-modal context to comprehensively understand scene text is beneficial for robust retrieval. In view of this, first, the present invention designs a Discriminative Entity Inference (DEI) module to predict discriminative scene text words in the title in the form of text prompts. At the same time, in terms of visual prompts, the present invention uses the bounding boxes detected by OCR to highlight the scene text in the image, enhance the context information, and alleviate local noise. Based on the heterogeneous prompt inference strategy, we align the scene text in these images with the scene text in the title to enhance their corresponding relationships in more detail. Secondly, in order to reduce semantic ambiguity, we design a Perceptual Entity Distillation (PED) module to distill and extract the beneficial information of the scene text at the image end, so as to learn more robust representations. Specifically, PED combines OCR features from different modalities (e.g., visual, semantic, and positional), and extracts valuable scene text entity representations centered on attributes through the slot attention mechanism.

[0005] An entity inference and distillation cross-modal retrieval method based on heterogeneous prompts of scene text, comprising the following steps:

[0006] Step 1: Determine the visual and text heterogeneous prompt inference strategy.

[0007] Step 2: Extract features from the input image-text pair through a feature extraction module, extract visual features through an image encoder, and extract text features through a text encoder.

[0008] Step 3: Extract valuable scene text information from the OCR features extracted by the image encoder through a Perceptual Entity Distillation (PED) module to learn better visual representations, and achieve cross-modal retrieval by calculating the similarity between the visual representations and the text features.

[0009] Step 4: Use the OCR-aware triplet loss to train the heterogeneous prompt-guided entity inference and distillation framework composed of the visual and text heterogeneous prompt inference strategy, the feature extraction module, and the Perceptual Entity Distillation (PED) module.

[0010] The beneficial effects of the present invention are as follows:

[0011] 1. The present invention proposes a new heterogeneous prompt entity inference and distillation framework for exploring discriminative multi-modal entities and obtaining attribute-centered representations, thereby enhancing the understanding of scene text.

[0012] 2. The present invention designs a visual and text heterogeneous prompt inference module. First, for text prompts, a Discriminative Entity inference (DEI) module is introduced to infer discriminative entity words in the text. Then, for visual prompts, we highlight the scene text in the image based on the bounding boxes detected by OCR to enhance the context information. We align the scene text at the text end and the image end, and enhance the corresponding relationship through a heterogeneous prompt strategy, so as to align the discriminative entities existing in the two modalities.

[0013] 3. The present invention designs a perceptual entity distillation module, which distills and extracts beneficial information of scene text at a fine-grained level to obtain a robust representation of scene text.

[0014] 4. A large number of experiments show that this method achieves efficient scene text-aware cross-modal retrieval on the CTC and TextCaps datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0016] Figure 1 It is a comparison diagram between the traditional cross-modal retrieval and the cross-modal retrieval method of the present invention.

[0017] Figure 2 It is a flowchart of the method of the embodiment of the present invention.

[0018] Figure 3 It is a performance comparison between the method of the present invention and the existing method on the CTC dataset.

[0019] Figure 4 It is a performance comparison between the method of the present invention and the existing method on the TextCaps dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0021] As Figure 2 shown, the entity inference and distillation cross-modal retrieval method based on scene text heterogeneous prompts includes the following steps:

[0022] Step 1: Determine the visual and text heterogeneous prompt inference strategy.

[0023] The described visual and text heterogeneous prompt inference strategy consists of two parts: visual prompts and text prompts. Visual prompts and text prompts can be regarded as steps for preprocessing images and text.

[0024] (1-1) Visual prompts. In visual prompts, to enhance the context information of scene text in the image and reduce local noise, we designed a visual prompt to retain the context information around the scene text through pixel-level prompts. Specifically, on the image side, for a given image containing scene text, use OCR (Optical Character Recognition) to detect the position of each scene text and obtain the bounding box of the scene text where n represents the maximum number of scene texts in the picture. Generally, n is less than 36, (x i , y i ), i = 1, 2, 3, 4 represent the four vertices of the bounding box, and the order is [top left, top right, bottom right, bottom left] in sequence. Then (in the embodiment, through the red solid line) connect the four vertices of the scene text bounding box in the image at the pixel level to obtain the picture processed by the visual prompt.

[0025] (1-2) Text prompts. In text prompts, first use the Discriminative Entity Inference (DEI) module to infer the scene text words in the text. For example, given the text "He is wearing a No. 63 jersey.", we can infer that the number "63" is likely to be a discriminative scene text word. Therefore, the scene text words in the text can be inferred through the DEI module. Technically, the described DEI module consists of a pre-trained BERT (Bidirectional Encoder Representations from Transformers), two layers of MLP layers, and a softmax layer. The DEI model is trained by calculating the cross-entropy loss between the output of the softmax layer of the DEI module and the true labels of the words (scene text words are represented by 1, and non-scene text words are represented by 0) on the TextCaps-OCR dataset. After the above separate training, the trained DEI module is used to infer the input text query to obtain the scene text words in the text. Then, combine the scene text words with the following prompt template to form the final text prompt:

[0026] “There are scene texts: [ST] in this image.”

[0027] where ST represents the scene text words inferred from the text by DEI.

[0028] After being processed by the heterogeneous prompting strategy (visual prompting and text prompting) in this step, the scene text on the image side and the scene text on the text side can be more finely aligned.

[0029] Step 2: Feature extraction based on the feature extraction module.

[0030] For a given image-text pair, visual features are extracted through the image encoder of the feature extraction module and text features are extracted through the text encoder of the feature extraction module based on the visual and text heterogeneous prompting inference strategy.

[0031] (2-1) Visual feature extraction through the image encoder:

[0032] In the image encoder, first, the visual prompting strategy in Step 2 is used to process the picture to facilitate the extraction of global features, and then local feature extraction and OCR feature extraction are performed on the original picture.

[0033] Global feature extraction: For the picture processed by visual prompting, a pre-trained CLIP image encoder with frozen parameters is used to extract global features, obtaining a 512-dimensional global feature G.

[0034] Local feature extraction: For the picture processed by visual prompting, pre-trained Faster-RCNN is used to extract local features, thereby obtaining local features where M is the maximum value of the number of significant objects in the picture.

[0035] OCR feature extraction: For the original picture, in order to extract OCR features from the input picture, we first use an OCR system (Paddle OCR is used in this embodiment) to obtain the set of scene text words and corresponding bounding boxes where N is the maximum number of OCRs detected in the picture. According to the set S, we extract features of the scene text in the picture from multiple dimensions to achieve a comprehensive understanding of the scene text. The specific operations are as follows:

[0036] (i) Use pre-trained FastText to extract features of the scene text words, obtaining 300-dimensional word embedding features (ii) Use pre-trained Faster-RCNN to extract visual information features of the objects within the OCR bounding boxes to obtain 2048-dimensional appearance features (iii) Use pre-trained PHOC to extract features of the scene text words to obtain 604-dimensional PHOC semantic features (iv) Perform tiling operation on the bounding boxes of the scene text and then input them as numbers to obtain 4-dimensional bounding box features Final OCR Features are defined as:

[0037]

[0038] where W 1 , W 2 , W 3 and W 4 are learnable projection matrices, LN(·) is the layer normalization operation, and σ is the LeakyReLU activation function.

[0039] (2-2) Extract text features through the text encoder:

[0040] In the text encoder, first use the DEI module in step 2 to infer the original text to obtain scene text words, thereby obtaining text prompts. Then, for the text prompts, use the CLIP pre-trained text encoder with frozen parameters to extract features, obtaining 512-dimensional text prompt features; for the original text, use GRU (Gated Recurrent Unit) to extract original text features. Finally, fuse the text prompt features and the original text features through an addition operation to obtain the final text feature T.

[0041] Step 3, extract valuable scene text features E from the OCR feature X ocr to learn better visual representations.

[0042] After obtaining the OCR feature X ocr containing various information, since the correlation between them is not established, they are unorganized, not compact, and will cause semantic ambiguity problems. To solve this problem, extract valuable scene text information from the OCR feature X ocr through the Perceptual Entity Distillation (PED) module, so that the OCR features will compete with each other through the slot attention mechanism, thereby obtaining more robust scene text features E.

[0043] Specifically, use a set of L learnable slots to distill L attribute-centered representations from the OCR feature X ocr through the slot attention mechanism. Specifically, the PED module first initializes L learnable attribute slots from the Gaussian distribution N(0,1) where D is the embedding dimension of each slot. Then, through multiple (4 layers are selected in this embodiment) weight-sharing cross-attention layers, through multiple iterations, we can distill more discriminative scene text features, and the formula is as follows:

[0044]

[0045] Among them, represents the cross-attention layer, and t represents the t-th iteration (in this embodiment, t = 1, 2, 3, 4). Specifically, the operation of cross-attention is as follows: First, perform layer normalization on L initialized attribute slots E t-1 and X ocr ; then, project the input OCR feature X ocr into the key and value in the attention mechanism; project E t-1 into the query D h represents the dimension of the hidden layer state; after that, the attention matrix between X ocr and E t-1 is calculated according to the following formula:

[0046]

[0047] where Softmax(·) is calculated along the direction of the slots, and this operation encourages competition between the slots to obtain information beneficial for retrieval. Based on the relationship matrix A t in the t-th iteration, the PED module updates the slots, and the formula is as follows:

[0048]

[0049] where W 0 represents the learnable projection matrix, and FFN represents the feed-forward neural network, which consists of a linear layer, an activation layer (we use GRLU here), and a layer normalization layer. Finally, a more robust scene text feature E is obtained:

[0050] E = Avg(E t )

[0051] where Avg(·) represents the average pooling operation.

[0052] Finally, the global feature G, the local feature and the scene text feature E are fused according to the following formula to obtain the final visual representation V of the image end:

[0053] I obj = GRU(GCN(X obj ))

[0054] V = (E⊙G + G)⊙I obj + I obj

[0055] where GCN represents the graph convolutional network, GRU represents the gated recurrent unit, and ⊙ represents the element-wise product.

[0056] Step 4: Train the heterogeneous prompt-guided entity reasoning and distillation framework composed of the visual and text heterogeneous prompt reasoning strategy, the feature extraction module, and the Perceptual Entity Distillation (PED) module using the OCR-aware triplet loss.

[0057] To align the visual feature V and the text feature T, use the OCR-aware triplet ranking loss as the training objective, which is defined as follows:

[0058]

[0059] where (V, T) represents an image-text positive sample pair in a batch, and represents an image-text negative sample pair in a batch, α > 0 is the gain parameter, [x] + = max(x, 0), and ε(·) represents the cosine similarity function.

[0060] During the experiment, we set the number of epochs to 30, the batch size to 300, the learning rate to 0.0002, and after every 10 epochs, the learning rate was reduced to 10% of the previous value. We selected Adam as the optimizer to update the gradients.

[0061] Through the above specific implementation manners, the present invention can enhance the model's understanding of scene text and achieve more accurate retrieval accuracy.

[0062] As Figure 3 shown, the method of the present invention (HOPID) releases the potential of heterogeneous scene text on the CTC-1K and CTC-5K test sets through the visual prompt and text prompt strategies, achieving a significant improvement of +58.5 in the RSUM metric compared to the single-modal CLIP model, and enabling a 9.3% improvement in the I2T-R@1 metric compared to ViSTA with the help of the Perceptual Entity Distillation (PED) module.

[0063] In addition, as Figure 4 shown, on the TextCaps dataset, the present invention captures diverse representations of scene text through the heterogeneous prompt strategy, refreshing the performance record with RSUM = 435.4, and improving the retrieval performance by 11% compared to the Dual Encoder model based on single-modal OCR modeling.

[0064] The above content is a further detailed description of the present invention in combination with specific / preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, they can make several substitutions or modifications to these described implementation manners, and these substitution or modification manners should all be regarded as belonging to the protection scope of the present invention.

[0065] The parts not described in detail in the present invention belong to the well-known technologies to those skilled in the art.

Claims

1. Entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompts, characterized by: The steps include: Step 1: Determine the reasoning strategy for heterogeneous visual and textual cues; Step 2: Extract features from the input image-text pair through the feature extraction module, extract visual features through the image encoder, and extract text features through the text encoder; Step 3: Extract valuable scene text information from the OCR features extracted by the image encoder through the Perceptual Entity Distillation (PED) module to learn better visual representations and achieve cross-modal retrieval by calculating the similarity between the visual representation and text features; Step 4: The heterogeneous cue-guided entity reasoning and distillation framework consisting of the visual and textual heterogeneous cue reasoning strategy, feature extraction module and perceptual entity distillation (PED) module is trained using the OCR-aware triplet loss.

2. The entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompts according to claim 1 is characterized in that: The described visual and textual heterogeneous cue reasoning strategy consists of two parts: visual cue and textual cue; (1-1) Visual cues; Specifically, on the image side, for a given image containing scene text, optical character recognition (OCR) is used to detect the location of each scene text and obtain the bounding box of the scene text. Where n represents the maximum number of scene texts contained in the image, (x i ,y i ), i = 1, 2, 3, 4 represent the four vertices of the bounding box, and the order is [upper left, upper right, lower right, lower left]; then the four vertices of the scene text bounding box in the image are connected at the pixel level to obtain the image processed by the visual cue; (1-2) Text prompts; In the text prompts, the discriminative entity reasoning DEI module is first used to infer the scene text words in the text; the DEI module is composed of a pre-trained BERT, two MLP layers and a softmax layer; the DEI model is trained by calculating the cross entropy loss between the output of the DEI module softmax layer and the true label of the word on the TextCaps-OCR dataset; after the above-mentioned separate training, the trained DEI module is used to infer the input text query to obtain the scene text words in the text; then, the scene text words are combined with the following prompt template to form the final text prompt: "There are scene texts:[ST]inthis image." Here, ST represents the scene context words inferred from the text by DEI.

3. The entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompts according to claim 2 is characterized in that: Step 2: For a given image-text pair, the image encoder of the feature extraction module extracts visual features based on the visual and textual heterogeneous cue reasoning strategy, and the text encoder of the feature extraction module extracts text features. (2-1) Extract visual features through image encoder: In the image encoder, the visual cue strategy in step 2 is first used to process the image to extract global features, and then local features and OCR features are extracted from the original image; Global feature extraction: For the images processed by visual cues, the CLIP pre-trained image encoder with frozen parameters is used to extract global features to obtain a 512-dimensional global feature G; Local feature extraction: For the image processed by visual cues, the pre-trained Faster-RCNN is used to extract local features to obtain local features. Where M is the maximum number of significant objects in the image; OCR feature extraction: For the original image, first use the OCR system to obtain the scene text words and the corresponding bounding box set Where N is the maximum number of OCR detected in the image; according to the set S, the scene text in the image is feature extracted from multiple dimensions to achieve a comprehensive understanding of the scene text. The specific operations are as follows: (i) Use the pre-trained FastText to extract features from scene text words and obtain 300-dimensional word embedding features (ii) Use the pre-trained Faster-RCNN to extract the visual information of the object in the OCR bounding box to obtain 2048-dimensional appearance features (iii) Use the pre-trained PHOC to extract features from scene text words to obtain 604-dimensional PHOC semantic features (iv) Bounding box for scene text The tiling operation is then input as a number to obtain the 4-dimensional bounding box feature Final OCR Features Defined as: Where W1, W2, W3 and W4 are learnable projection matrices, LN(·) is the layer normalization operation, and σ is the LeakyReLU activation function; (2-2) Extract text features through text encoder: In the text encoder, the DEI module in step 2 is first used to infer the original text to obtain the scene text words, thereby obtaining the text prompt; then, the text prompt is extracted using the parameter-frozen CLIP pre-trained text encoder to obtain 512-dimensional text prompt features; for the original text, the gated recurrent unit GRU is used to extract the original text features to obtain the original text features, and finally, the text prompt features and the original text features are fused through the addition operation to obtain the final text features T.

4. The entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompts according to claim 3 is characterized in that: Step 3: Specifically, a set of L learnable slots are used to extract the OCR features X from the OCR features X through the slot attention mechanism. ocr Specifically, the PED module first initializes L learnable attribute slots from the Gaussian distribution N(0,1) Where D is the embedding dimension of each slot; then, through multiple layers of weight-sharing cross-attention layers, more discriminative scene text features are distilled, as follows: in, represents the cross attention layer, and t represents the tth iteration; specifically, the cross attention operation is as follows: first, the L initialized attribute slots E t-1 and X ocr Perform layer normalization operation; then, input OCR feature X ocr Projection as a key in the attention mechanism With value E t-1 Projection as query in attention mechanism D h represents the dimension of the hidden layer state; after this, X ocr and E t-1 The attention matrix between is calculated as follows: Among them, Softmax(·) is calculated along the direction of the slot; based on the relationship matrix A of the tth iteration t , the PED module updates the slot, the formula is as follows: Among them, W0 represents the learnable projection matrix and FFN represents the feedforward neural network; finally, a more robust scene text feature E is obtained: E=Avg(E t ) Among them, Avg(·) represents the average pooling operation; Finally, the global feature G and the local feature The scene text feature E is fused according to the following formula to obtain the final visual representation V on the image side: I obj =GRU(GCN(X obj )) V=(E⊙G+G)⊙I obj +I obj Among them, GCN stands for graph convolutional network, GRU stands for gated recurrent unit, and ⊙ stands for element-wise product.

5. The entity reasoning and distillation cross-modal retrieval method based on scene text heterogeneous prompts according to claim 4 is characterized in that: Step 4: Use OCR-aware triplet loss to train the heterogeneous cue-guided entity reasoning and distillation framework consisting of visual and textual heterogeneous cue reasoning strategies, feature extraction modules, and perceptual entity distillation (PED) modules. In order to align visual features V and text features T, OCR-aware triplet ranking loss is used as the training objective, which is defined as follows: Among them, (V,T) represents the positive sample pairs of images and texts in a batch, and represents a batch of image-text negative sample pairs, α>0 is the gain parameter, [x] + =max(x,0), ε(·) represents the cosine similarity function.

Citation Information

Patent Citations

  • Cross-modal retrieval model and method based on anti-fact reasoning and computer equipment

    CN115146100A

  • Image-text cross-modal vehicle retrieval model training method in vehicle dense scene

    CN118968516A