An end-to-end scene text recognition method based on CLIP

Through the CLIP-based end-to-end scene text recognition method, using language cue generators and visual cue generators, combined with a bimodal similarity matching module, the problems of small sample learning and insufficient generalization ability are solved, and the accuracy and applicability of text recognition are improved. It is suitable for fields such as document analysis and autonomous driving.

CN117058667BActive Publication Date: 2025-09-23HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311154735.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-09-23
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

Existing end-to-end scene text recognition methods have deficiencies in small-sample learning and generalization capabilities, especially in complex scenes. Traditional methods are also sensitive to lighting, viewing angle, and font, resulting in performance degradation.

Method used

An end-to-end scene text recognition method based on CLIP is adopted. By constructing a language cue generator, a visual cue generator and a bimodal similarity matching module, combined with the CLIP pre-trained model, the effective fusion of image and text features is achieved. The cross-attention mechanism is used to transfer visual information to the text instance level, and text segmentation is performed through text instance-language matching alignment.

Benefits of technology

It significantly improves the accuracy of scene text detection and end-to-end text recognition, overcomes the challenges of small sample learning and generalization capabilities, and enables the model to perform well in open scenarios, making it suitable for fields such as document analysis and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058667B_ABST
    Figure CN117058667B_ABST
Patent Text Reader

Abstract

The present invention discloses a CLIP-based end-to-end scene text recognition method: by improving the large-scale visual language pre-training model CLIP, using the image encoder and text encoder pre-trained by CLIP, a language prompt generator, a visual prompt generator, and a text instance and language matching module are introduced. By leveraging the language knowledge in CLIP, FastTCM can effectively assist downstream text detection and end-to-end text recognition tasks, thereby significantly improving the accuracy of existing scene text detectors and end-to-end text recognizers. In addition, it can also enhance performance in small sample learning scenarios and improve the generalization ability of the model. It greatly expands the application areas of scene text detection and end-to-end text recognition, and is expected to play an important role in fields such as image annotation and document analysis. By integrating language and visual information, it provides a new paradigm for end-to-end scene text recognition, and makes a positive contribution to the development of deep learning technology in the field of text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and optical character recognition, and more specifically, relates to an end-to-end scene text recognition method based on CLIP. Background Art

[0002] Text recognition has always been a challenging task in the fields of deep learning and optical character recognition. Traditional text recognition methods typically require multiple discrete steps, such as text detection, text segmentation, and character recognition, which are closely interdependent. Furthermore, these methods are often very sensitive to factors such as lighting, viewing angle, and font, resulting in poor performance in real-world scenarios.

[0003] In recent years, advances in deep learning have brought significant improvements to text recognition. End-to-end text recognition methods achieve higher performance and robustness by integrating text detection, segmentation, and recognition steps into a unified framework. However, these methods still face challenges such as limited sample size, insufficient generalization, and difficulty recognizing in complex scenes.

[0004] Furthermore, with the rise of large-scale visual language models (such as CLIP, Contrastive Language-Image Pre-Training, a pre-training method based on contrastive text-image pairs), the deep learning community has begun exploring how to effectively integrate language knowledge and visual information into text recognition. The CLIP model is well-known for its excellent cross-modal understanding capabilities, but how to apply it to scene text recognition remains an open problem. Summary of the Invention

[0005] In response to the above-mentioned defects or improvement needs of the existing technology, the present invention provides an end-to-end scene text recognition method based on CLIP. Its purpose is to utilize the advantages of large-scale visual language models to solve the problems of small sample learning and insufficient generalization capabilities of end-to-end scene text recognition methods in the field of optical character recognition, thereby providing key technical support for document analysis, autonomous driving and other fields, and further improving the efficiency of smart office, smart education and other fields.

[0006] To achieve the above objectives, the present invention provides an end-to-end scene text recognition method based on CLIP, which includes:

[0007] Step 1: Extract image features and use the ResNet 50 model pre-trained by CLIP as the image encoder. Input the image into the image encoder to obtain the global image embedding features.

[0008] Step 2: Construct text input. First, use the predefined language prompt "Text" as part of the input and encode it into a vector through Word Embedding. Second, construct a learnable language prompt vector as part of the text input. Then, merge the predefined language prompt vector and the learnable language prompt vector to obtain a preliminary text encoder input. Finally, the present invention proposes using a language prompt generator to generate a conditional prompt feature vector. For each image, the conditional prompt feature vector is combined with the preliminary text encoder input as the final text encoder input.

[0009] Step 3: Extract text features and use the CLIP pre-trained Transformer model as a text encoder to extract text embeddings.

[0010] Step 4: Enhance text embedding. This paper proposes to use a bimodal similarity matching (BSM) module to enhance the text embedding using global image embedding features to obtain the final text embedding.

[0011] Step 5: Generate conditional visual cues. This paper proposes using a visual cue generator to adaptively propagate fine-grained semantic information from text features into visual features to obtain conditional visual cues. Specifically, this paper uses the cross-attention mechanism in the Transformer to model the interaction between image features and text embeddings. Then, the global image embedding features and conditional visual cues are fused to obtain text-aware local image features.

[0012] Step 6: Perform text instance-language matching on the final text embedding and text-aware local image features, and use a dot product followed by a Sigmoid operation to obtain a text segmentation map.

[0013] Step 7: During the training phase, on the one hand, the corresponding loss of the downstream task head (text detection head or end-to-end text recognition head) is used; on the other hand, the cross-entropy optimized text segmentation map is used as an auxiliary loss, and the two are jointly optimized to optimize the model parameters;

[0014] Step 8: In the inference phase, use the output of the corresponding task head as the final output of the model.

[0015] In one embodiment of the present invention, the step 1 specifically includes: first inputting the image The global image embedding feature I is obtained by inputting the image encoder, i.e., I = ImageEncoder(I′). The default input size of the image is 1024×1024. is the output image feature. C is the number of channels of the image feature, the default is 1024. s is the image downsampling factor, the default is 32.

[0016] In one embodiment of the present invention, the step 2 specifically includes: given a predefined prompt word "Text", a predefined language prompt vector t' with a dimension of D is obtained through the Word Embedding layer in ,Right now The default setting of D is 512. The learnable language hint vector is realized by learnable parameters, denoted as {c1,…,c n}, the length of the learnable parameter n is set to 4 by default. The initial text encoder input can be obtained by concatenating the learnable language prompt vector with the predefined language prompt vector in dimension 0, that is,

[0017] Although combining predefined language prompts and learnable language prompts and feeding them into the text encoder is effective for guiding the CLIP model, in open scenarios, when the test text instances are not in the same distribution range as the training images, they may be affected by limited small sample learning or generalization capabilities. To this end, the present invention proposes the use of a language prompt generator to generate a conditional prompt feature vector cc. Specifically, the present invention introduces a meta query (MQ) and then generates cc through a two-layer feedforward network for feature transformation. This mechanism can decouple the text encoder during the inference process, realize offline calculation, and speed up the inference speed, that is, Among them, MQ is initialized as a learnable parameter with a dimension size of C. is a learnable parameter. LN is the Layer Normalization operation. σ is the activation function ReLu. The meta-query serves as an implicit image condition, guiding the generation of subsequent language cues, thereby mining the pre-trained knowledge from the text encoder. It is important to note that once training is complete, the meta-query remains unchanged. This enables the CLIP text encoder to extract offline computations during inference, thereby shortening inference time and making the present invention more suitable for practical applications.

[0018] The final text encoder input can be expressed as Among them, cc will be broadcast to and t in Same dimension size.

[0019] In one embodiment of the present invention, the step three specifically includes: extracting Features, get the text embedding t of dimension C out ,Right now Where C is the feature dimension of text embedding, which defaults to 1024.

[0020] In one embodiment of the present invention, the step 4 specifically includes: a bimodal similarity matching (BMS) module is used to control the amount of visual modality information that should be used to compensate for the text modality embedding. This method of dynamically enriching text embedding enhances text embedding through visual information, which helps to improve the overall performance of the model. Specifically, first, the global image embedding feature I is globally averaged to obtain the global image average embedding Given a text embedding t out and global image average embedding First calculate t out and The cosine similarity between We use sim as the correlation threshold of the output gate, which controls the amount of visual modality information used to compensate the text modality embedding. Next, we use the correlation threshold sim to adjust t as follows: out and Perform a weighted sum: in This is the final output of the text encoder.

[0021] In one embodiment of the present invention, the step 5 specifically includes: the present invention uses a cross attention mechanism to transfer visual information from the image level to the text instance level, thereby enabling it to have the ability to perceive text areas. Specifically, the interaction between image features and text embeddings results in conditional visual cues. This process can be defined as Where TDec represents the Transformer decoder. In practice, it consists of 6 dual Transformer decoder layers, each layer contains 4 attention heads to ensure sufficient interaction between image features and text embeddings; the width of the Transformer is set to 256 and the feedforward hidden dimension is set to 1024. The final text-aware local image features That is

[0022] In one embodiment of the present invention, the step six specifically includes: the text segmentation map is obtained based on the matching relationship between text instances and languages: τ is the temperature coefficient and is set to 0.07 by default.

[0023] In one embodiment of the present invention, step seven specifically includes: using the label of the text region as supervision of the text segmentation map, and the auxiliary loss is defined as:

[0024]

[0025] Among them, y ij and Pij Represents the label and predicted probability of the text instance at the pixel (i, j) position. The text segmentation map P will be combined with the local image features The input of the downstream detection head and the end-to-end text recognition head are integrated to explicitly incorporate language priors into the text detection and recognition tasks. The loss corresponding to the downstream task head is recorded as The final optimization goal is

[0026] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:

[0027] This paper improves the large-scale visual-language pre-training model CLIP, introducing key components such as a language cue generator, a visual cue generator, and a text instance and language matching module to significantly improve the accuracy of scene text detection and end-to-end text recognition. This innovative framework not only overcomes the challenges of small-sample learning and generalization, but also performs well in open-ended scenarios, providing important technical support for fields such as document analysis and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flow chart of an end-to-end scene text recognition method based on CLIP of the present invention;

[0029] Figure 2 This is a schematic diagram of a language prompt generator;

[0030] Figure 3 Schematic diagram of the visual cue generator. DETAILED DESCRIPTION

[0031] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0032] This invention aims to address the challenges of traditional text detection and end-to-end text recognition methods by providing a CLIP-based, end-to-end scene text recognition method. By effectively integrating linguistic knowledge and visual information, it improves the performance, generalization, and applicability of text recognition. This technological innovation is expected to have a significant impact in areas such as image annotation, document analysis, and autonomous driving.

[0033] like Figure 1 As shown in FIG, the CLIP-based end-to-end scene text recognition method proposed in the present invention includes the following steps:

[0034] Step 1: Extract image features and use the ResNet 50 model pre-trained by CLIP as the image encoder. Input the image into the image encoder to obtain the global image embedding features.

[0035] Specifically, first input the image The global image embedding feature I is obtained from the input image encoder, that is, I=ImageEncoder(I ′ ). The default input size of the image is 1024×1024. is the output image feature. C is the number of channels of the image feature, the default is 1024. s is the image downsampling factor, the default is 32.

[0036] Step 2: Construct text input. First, use the predefined language prompt "Text" as part of the input and encode it into a vector through Word Embedding. Second, construct a learnable language prompt vector as part of the text input. Then, merge the predefined language prompt vector and the learnable language prompt vector to obtain a preliminary text encoder input. Finally, the present invention proposes to use a language prompt generator, such as Figure 2 As shown, a conditional hint feature vector is generated. For each image, the conditional hint feature vector is combined with the input of the preliminary text encoder as the final text encoder input;

[0037] Specifically, given a predefined prompt word "Text", the predefined language prompt vector t with dimension D is obtained through the Word Embedding layer. ′ in ,Right now The default setting of D is 512. The learnable language hint vector is realized by learnable parameters, denoted as {c1,…,c n}, the length of the learnable parameter n is set to 4 by default. The initial text encoder input can be obtained by concatenating the learnable language prompt vector with the predefined language prompt vector in dimension 0, that is,

[0038] Although combining predefined language prompts and learnable language prompts and feeding them into the text encoder is effective for guiding the CLIP model, in open scenarios, when the test text instances are not in the same distribution range as the training images, they may be affected by limited small sample learning or generalization capabilities. To this end, the present invention proposes the use of a language prompt generator to generate a conditional prompt feature vector cc. Specifically, the present invention introduces a meta query (MQ) and then generates cc through a two-layer feedforward network for feature transformation. This mechanism can decouple the text encoder during the inference process, realize offline calculation, and speed up the inference speed, that is, Among them, MQ is initialized as a learnable parameter with a dimension size of C. is a learnable parameter. LN is the Layer Normalization operation. σ is the activation function ReLu. The meta-query acts as an implicit image condition to guide the generation of subsequent language prompts, thereby mining the pre-trained knowledge from the text encoder. It should be noted that once the training is completed, the meta-query will remain unchanged. This enables the CLIP text encoder to extract offline calculations during inference, thereby shortening the inference time and making the present invention more suitable for practical applications. The final text encoder input can be expressed as Among them, cc will be broadcast to and t in Same dimension size.

[0039] Step 3: Extract text features and use the CLIP pre-trained Transformer model as a text encoder to extract text embeddings.

[0040] Specifically, the text encoder is used to extract Features, get the text embedding t of dimension C out ,Right now Where C is the feature dimension of text embedding, which defaults to 1024.

[0041] Step 4: Enhance text embedding. This paper proposes to use a bimodal similarity matching (BSM) module to enhance the text embedding using global image embedding features to obtain the final text embedding.

[0042] Specifically, the Bimodal Similarity Matching (BMS) module is used to control the amount of visual modality information that should be used to compensate for the text modality embedding. This method of dynamically enriching text embeddings, which enhances text embeddings with visual information, helps improve the overall performance of the model. Specifically, the global image embedding feature I is first globally averaged to obtain the global image average embedding Given a text embedding tout and global image average embedding First calculate t out and The cosine similarity between We use sim as the correlation threshold of the output gate, which controls the amount of visual modality information used to compensate the text modality embedding. Next, we use the correlation threshold sim to adjust t as follows: out and Perform a weighted sum: in This is the final output of the text encoder.

[0043] Step 5: Generate conditional visual cues. The present invention proposes to use a visual cue generator, such as Figure 3 As shown in the figure, the fine-grained semantic information in the text features is propagated into the visual features in an adaptive manner to obtain conditional visual cues. Specifically, the present invention uses the cross-attention mechanism in the Transformer to model the interaction between image features and text embeddings; then, the global image embedding features and conditional visual cues are fused to obtain text-aware local image features.

[0044] Specifically, the present invention uses a cross-attention mechanism to transfer visual information from the image level to the text instance level, thereby enabling it to perceive text regions. Specifically, the interaction between image features and text embeddings results in conditional visual cues. This process can be defined as Where TDec represents the Transformer decoder. In practice, it consists of 6 dual Transformer decoder layers, each layer contains 4 attention heads to ensure sufficient interaction between image features and text embeddings; the width of the Transformer is set to 256 and the feedforward hidden dimension is set to 1024. The final text-aware local image features That is

[0045] Step 6: Perform text instance-language matching on the final text embedding and text-aware local image features, and use a dot product followed by a Sigmoid operation to obtain a text segmentation map.

[0046] Specifically, this segmentation map is obtained based on the matching relationship between text instances and languages: τ is the temperature coefficient and is set to 0.07 by default.

[0047] Step 7: During the training phase, on the one hand, the corresponding loss of the downstream task head (text detection head or end-to-end text recognition head) is used; on the other hand, the cross-entropy optimized text segmentation map is used as an auxiliary loss, and the two are jointly optimized to optimize the model parameters;

[0048] Specifically, the labels of text regions are used as supervision for text segmentation maps, and the auxiliary loss is defined as:

[0049]

[0050] Among them, y ij and P ij Represents the label and predicted probability of the text instance at the pixel (i, j) position. The text segmentation map P will be combined with the local image features The input of the downstream detection head and the end-to-end text recognition head are integrated to explicitly incorporate language priors into the text detection and recognition tasks. The loss corresponding to the downstream task head is recorded as The final optimization goal is

[0051] Step 8: In the inference phase, use the output of the corresponding task head as the final output of the model.

[0052] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A CLIP-based end-to-end scene text recognition method, characterized in that: The steps include: Step 1: Extract image features and use the ResNet 50 model pre-trained by CLIP as the image encoder. Input the image into the image encoder to obtain the global image embedding features. Step 2: Construct text input. First, use the predefined language prompt "Text" as part of the input and encode it into a vector using Word Embedding. Second, construct a learnable language prompt vector as part of the text input. Then, merge the predefined language prompt vector and the learnable language prompt vector to obtain the preliminary text encoder input. Finally, use the language prompt generator to generate a conditional prompt feature vector. For each image, the conditional prompt feature vector is combined with the preliminary text encoder input as the final text encoder input. Step 3: Extract text features and use the CLIP pre-trained Transformer model as a text encoder to extract text embeddings. Step 4: Enhance text embedding, using the bimodal similarity matching module, and using the global image embedding feature to enhance the text embedding to obtain the final text embedding; the step 4 specifically includes: the bimodal similarity matching module is used to control the amount of visual modality information that should be used to compensate for the text modality embedding; first, the global image embedding feature I is globally averaged to obtain the global image average embedding Given a text embedding t out and global image average embedding Calculate t out and The cosine similarity between sim is used as the correlation threshold of the output gate, which controls the amount of visual modality information used to compensate the text modality embedding; Next, using the correlation threshold sim, t is adjusted as follows out and Perform a weighted sum: in This is the final output of the text encoder; Step 5: Generate conditional visual cues. Use the visual cue generator to adaptively propagate fine-grained semantic information from text features to visual features to obtain conditional visual cues. Use the cross-attention mechanism in the Transformer to model the interaction between image embedding features and text embeddings. Then, fuse the global image embedding features and conditional visual cues to obtain text-aware local image features. Step 6: Perform text instance-language matching on the final text embedding and text-aware local image features, and use a dot product followed by a Sigmoid operation to obtain a text segmentation map. Step 7: During the training phase, on the one hand, the downstream task head, i.e., the text detection head or the end-to-end text recognition head, is used to obtain the corresponding loss. On the other hand, the cross-entropy optimization of the text segmentation map is used as an auxiliary loss. The two are jointly optimized to optimize the model parameters. Step 8: In the inference phase, use the output of the corresponding task head as the final output of the model.

2. The CLIP-based end-to-end scene text recognition method according to claim 1, characterized in that: The step 1 specifically includes: First, the input image The global image embedding feature I is obtained from the input image encoder, that is, I=ImageEncoder(I ′ ), where the image input size defaults to 1024×1024, is the output image feature, where C is the number of channels of image features, and s is the image downsampling ratio.

3. The CLIP-based end-to-end scene text recognition method according to claim 2, characterized in that: The second step specifically includes: Given a predefined prompt word "Text", the predefined language prompt vector t with dimension D will be obtained through the Word Embedding layer ′ in ,Right now The learnable language hint vector is realized by learnable parameters, denoted as {c1,…,c n }, the length of the learnable parameter is n; the initial text encoder input can be obtained by concatenating the learnable language prompt vector with the predefined language prompt vector in dimension 0, that is, Use the language prompt generator to generate the conditional prompt feature vector cc; introduce the meta-query and then perform feature transformation through a two-layer feedforward network to generate cc, that is, Among them, MQ is initialized as a learnable parameter with a dimension size of C. is a learnable parameter, LN is the Layer Normalization operation, and σ is the activation function ReLu; the meta-query acts as an implicit image condition to guide the generation of subsequent language prompts, thereby mining the pre-trained knowledge from the text encoder; The final text encoder input is represented as Among them, cc will be broadcast to and t in Same dimension size.

4. The CLIP-based end-to-end scene text recognition method according to claim 3, characterized in that: The step three specifically includes: Extraction using text encoder Features, get the text embedding t of dimension C out ,Right now Where C is the feature dimension of text embedding.

5. The CLIP-based end-to-end scene text recognition method according to claim 4, characterized in that: The step five specifically includes: Using the cross-attention mechanism, visual information is transferred from the image level to the text instance level, so that it has the ability to perceive text areas; the interaction between image embedding features and text embeddings gives conditional visual cues This process is defined as Where TDec represents the Transformer decoder; TDec consists of 6 dual Transformer decoder layers, each layer contains 4 attention heads to ensure sufficient interaction between image features and text embeddings; the width of the Transformer is set to 256, and the feedforward hidden dimension is set to 1024; the final text-aware local image features That is 6. The CLIP-based end-to-end scene text recognition method according to claim 5, characterized in that: The step six specifically includes: The text segmentation map is obtained based on the matching relationship between text instances and languages: τ is the temperature coefficient.

7. The CLIP-based end-to-end scene text recognition method according to claim 6, characterized in that: The step seven specifically includes: Using the labels of text regions as supervision for the text segmentation map, the auxiliary loss is defined as: Among them, y ij and P ij Represents the label and predicted probability of the text instance at the pixel (i, j) position; the text segmentation map P and local image features The input of the downstream detection head or end-to-end text recognition head is integrated to explicitly incorporate language priors into the text detection and recognition tasks. The loss corresponding to the downstream task head is recorded as The final optimization goal is

Citation Information

Patent Citations

  • Knowledge graph representation learning method friendly to long text

    CN113761224A

  • Fine adjustment method of visual language pre-training model and image-text retrieval method

    CN115391588A