A method for cross-lingual description generation of multilingual text images

By constructing a multilingual text-image cross-language description generation network and utilizing a multimodal Transformer module to map different modal information in different languages ​​and perform text error correction, the problem of text recognition error in multilingual environments is solved, and high-quality multilingual text-image description generation is achieved.

CN119516548BActive Publication Date: 2025-10-28HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411631533.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-10-28
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Existing text image description generation methods fail to fully utilize the correlation between text regions and visual targets at different locations within an image in multilingual environments, resulting in erroneous text recognition results that affect high-quality description generation.

Method used

A multilingual text-image cross-language description generation network is constructed, including a multilingual image-text detection and recognition module, a multilingual text information encoding module, a multilingual visual information encoding module, a multimodal Transformer module, a knowledge extraction module, and a language embedding module. The multimodal Transformer module maps different language and modal information to a common space, and the text recognition results are corrected by a network pre-trained through a multimodal text error correction task.

Benefits of technology

It generates higher-quality multilingual text image descriptions, overcomes the problem of text recognition error accumulation, improves the accuracy of multilingual text image description generation, and can generate description sentences in a specified language in a multilingual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516548B_ABST
    Figure CN119516548B_ABST
Patent Text Reader

Abstract

This invention discloses a method for cross-lingual description generation of multilingual text images. The steps include: 1. acquiring multilingual text images and annotating them with descriptive sentences; 2. constructing a cross-lingual description generation network for multilingual text images; 3. constructing a dataset for a multimodal text correction task and pre-training some modules in the description generation network; 4. training all modules of the network based on the multilingual text image description generation dataset; 5. using the trained cross-lingual description generation network to generate descriptive sentences in a specified language for any input multilingual text image. This invention can perform deep understanding of input multilingual natural scene text images in multilingual scenarios and output descriptive sentences in a specified language for the multilingual text images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to information processing problems related to multilingual natural scene text images, specifically to a method for generating cross-language descriptions for multilingual text images. Background Technology

[0002] Current research on text-to-image caption generation largely focuses on generating English image captions from English text images. However, given the rapid development of cross-border trade and tourism, research on text-to-image caption generation in multilingual environments is necessary. Text-to-image caption generation methods rely on upstream models, and erroneous text recognition results can negatively impact subsequent modeling steps. For text images in natural scenes, there are relationships between text regions at different locations within the image, as well as between visual objects and text content. Existing text recognition models do not fully utilize these relationships when recognizing text. Incorrect text recognition results affect the generation of high-quality image captions. Summary of the Invention

[0003] To address the shortcomings of the existing technologies, this invention proposes a cross-lingual description generation method for multilingual text images. The method aims to correct the text recognition results based on the correlation between information from multiple perspectives within the text image, and then use the corrected results for description generation, thereby generating higher-quality multilingual text image descriptions.

[0004] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0005] The present invention provides a method for generating cross-language descriptions of multilingual text images, characterized by the following steps:

[0006] Step 1: Obtain multilingual text images and annotate the descriptive sentences to obtain the training set. ,in, Representing multilingual text images, express The corresponding image description, and , express The first in One character, express The number of characters in the text; express Corresponding languages express The corresponding linear representation of structured knowledge, and , express The first in One character, express The number of characters in the text; express The number of multilingual text images;

[0007] Step 2: Construct a cross-lingual description generation network for multilingual text images It includes: a multilingual image and text detection and recognition module, a multilingual text information encoding module, a multilingual visual information encoding module, a multimodal Transformer module, a knowledge extraction module, a language embedding module, and a Transformer decoding module; and will Each sample input in The process is performed to obtain the decoded character probability sequence. And used to construct the total loss function. ;

[0008] Step 3: Construct a dataset for the multimodal text error correction task and used for The multilingual text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module are pre-trained to obtain the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, and the pre-trained multimodal Transformer module.

[0009] Step 4: Based on the total loss function A multilingual text-image cross-language description generative network was developed using the backpropagation algorithm. The multilingual image and text detection and recognition module, the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, the pre-trained multimodal Transformer module, the knowledge extraction module, the language embedding module, and the Transformer decoding module are trained to update the network parameters, thereby obtaining the trained multilingual text and image cross-language description generation model.

[0010] Step 5: Use the trained multilingual text-image cross-language description generation model to generate descriptions for the specified language categories. Any input multilingual text image Make predictions to obtain the results in the specified language category. Down The descriptive statement.

[0011] The cross-lingual description generation method for multilingual text images described in this invention is also characterized in that step 2 includes the following steps:

[0012] Step 2.1: The multilingual image and text detection and recognition module first detects... The text area is output, and its position coordinates are output. Therefore, based on right Cropping is performed to obtain the text region image. and identify The character sequence in the text is used to obtain the text recognition result. ;

[0013] Step 2.2: The multilingual text information encoding module uses the mT5 model encoder to... The text is processed to output a text feature representation matrix. ,in, express The Middle The text representation vector of each position, for The number of semantic vectors in the middle;

[0014] Step 2.3: The multilingual visual information encoding module uses the Vit model to... Visual features are extracted to obtain Visual feature representation matrix ,in, express The Middle Visual representation vectors of image region features for The number of image regions;

[0015] Step 2.4: The multilingual visual information encoding module uses the Vit model to... Visual features are extracted to obtain representations. Visual feature representation matrix ,in, express The visual representation vector of the m-th image region. for The number of image regions;

[0016] Step 2.5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , The inputs are respectively processed by the multimodal Transformer module, and after passing through multiple stacked multi-head attention mechanisms, feedforward operations and residual connection processing, modality fusion and semantic enhancement, the corresponding outputs are obtained. Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding ;

[0017] Step 2.6: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , The input is used in the knowledge extraction module for prediction to obtain the character probability sequence of the knowledge representation sequence. ,in, The knowledge representation sequence represents the first... A probability vector of K characters, where K represents the number of characters in the sequence.

[0018] Step 2.7: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ;

[0019] Step 2.8: The knowledge extraction module... Transformation is performed, and in The character with the highest probability value is selected at each position, thus forming the prediction result of the knowledge representation sequence composed of characters at all positions. ,in, express The first in One character, express The number of characters in the text;

[0020] The knowledge extraction module performs embedding representation on the knowledge representation sequence to obtain the embedding representation of the knowledge representation sequence. ;

[0021] Step 2.9: The language embedding module... Perform embedding representation to obtain Representation vector ;

[0022] Step 2.10: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would , , , , The input is subjected to a stacked multi-head attention operation in the Transformer decoding module to obtain the decoded character probability sequence. ;in, express The Middle A character probability vector at each position. express The number of characters in the text;

[0023] Step 2.11: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ;

[0024] Step 2.12: Calculate the total loss function ,in, express The corresponding weights.

[0025] Furthermore, step 3 includes the following steps:

[0026] Step 3.1: Construct a dataset for the multimodal text correction task ,in, express The corresponding identification label containing the error characters, express The correct identification label;

[0027] Step 3.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is sequentially input into the multilingual text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module for processing to obtain... The corresponding semantically enhanced feature vector sequence ,in, express The corresponding number 1 eigenvector for The number of characters in the text;

[0028] Step 3.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is fed into a character classifier for prediction, and the output is the probability of the character category corresponding to each position. ;in, Indicates the output of the first The probability vector of the location;

[0029] Step 3.4: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Convert to one-hot vector and with Comparison to construct the loss function It is used to perform backpropagation on the language text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module to update the module parameters, thereby obtaining the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, and the pre-trained multimodal Transformer module.

[0030] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the cross-language description generation method, and the processor is configured to execute the program stored in the memory.

[0031] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the cross-language description generation method.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] 1. The multilingual text image cross-language description generation network of this invention can understand multilingual text images and generate descriptive sentences in a specified language from multilingual natural scene text images. It is not limited to generating English or Chinese descriptive sentences and can be better applied to multilingual environments. This invention uses a multilingual image and text detection and recognition module and a multilingual text information encoding module to process information from different languages. It also uses a multimodal Transformer module to map information from different languages ​​and modalities to a common space, overcoming the semantic gap between multilingual and multimodal languages. This invention uses a language embedding module to obtain language representation vectors and inputs them into the decoder, thereby guiding the decoder to generate descriptive sentences in a specified language.

[0034] 2. This invention corrects the recognition results of text images based on the correlation between information from different perspectives within multilingual text images, and uses the corrected results for description generation. This reduces the impact of text misrecognition on the generation of descriptive sentences, thereby generating higher-quality multilingual text image descriptive sentences. This invention uses a multimodal text error correction task to pre-train multiple parts of the cross-language description generation network, overcoming the problem of text recognition error accumulation in existing technologies, thus improving the accuracy of multilingual text image description generation.

[0035] 3. The network framework designed in this invention integrates the information extraction task and description generation task of multilingual text images, enabling the two tasks to complement each other. The inherent relationship between the two tasks allows them to promote each other, overcoming the shortcomings of existing technologies in insufficiently mining the relationships between different modal information within the image when generating descriptive sentences, thereby generating higher quality text image descriptive sentences. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating the usage of the multilingual text-image cross-language description generation method of the present invention;

[0037] Figure 2 This is a network structure diagram of the multilingual text image cross-language description generation method of the present invention. Detailed Implementation

[0038] In this embodiment, as Figure 1 As shown, a structured information extraction method for text images in multilingual natural scenes includes the following steps:

[0039] Step 1: Obtain multilingual text images and annotate the descriptive sentences to obtain the training set. ,in, Representing multilingual text images, express The corresponding image description, and , express The first in One character, express The number of characters in the text. express Corresponding languages express The corresponding linear representation of structured knowledge, and , express The first in One character, express The number of characters in the text. express The number of multilingual text images. For the first... Sample During training, and As input to the network, and As a supervisory signal for the network.

[0040] Step 2: As Figure 2 As shown, a cross-lingual description generation network for multilingual text images is constructed. It includes: a multilingual image and text detection and recognition module (for detecting and recognizing multilingual text in images), a multilingual text information encoding module (for constructing embedded representations of multilingual text modal data), a multilingual visual information encoding module (for constructing embedded representations of visual modal data), a multimodal Transformer module (for mapping multilingual multimodal data to a common space), a knowledge extraction module (for extracting structured knowledge representations from multilingual text images), a language embedding module (for constructing embedded representations for language categories), and a Transformer decoding module (for decoding the character sequence of image description sentences); and will... Each sample input in The process is performed to obtain the decoded character probability sequence. .

[0041] Step 2.1: The multilingual image and text detection and recognition module first detects... The text area is output, and its position coordinates are output. Therefore, based on right Cropping is performed to obtain the text region image. and identify The character sequence in the text is used to obtain the text recognition result. ;

[0042] Step 2.2: The multilingual text information encoding module uses the encoder of the mT5 model to... The text is processed to output a text feature representation matrix. ,in, express The Middle The text representation vector of each position, for The number of semantic vectors in the middle;

[0043] Step 2.3: The multilingual visual information encoding module uses the Vit model to... Visual features are extracted to obtain Visual feature representation matrix ,in, express The Middle Visual representation vectors of image region features for The number of regions in the image.

[0044] Step 2.4: The multilingual visual information encoding module uses the Vit model to... Visual features are extracted to obtain representations. Visual feature representation matrix ,in, express The Middle Visual representation vectors of image regions for The number of image regions;

[0045] Step 2.5: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , The inputs are respectively fed into the multimodal Transformer module, and sequentially pass through multiple stacked multi-head attention mechanisms, feedforward operations, residual connection processing, modality fusion, and semantic enhancement. The outputs at the corresponding positions of the multimodal Transformer module represent... Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding The multimodal Transformer module maps information from different languages ​​and modalities into a common space, overcoming the semantic gap between multiple languages ​​and multiple modalities.

[0046] Step 2.6: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , The input is used in the knowledge extraction module for prediction, resulting in a character probability sequence of the knowledge representation sequence. ,in, The knowledge representation sequence represents the first... A probability vector of K characters, where K represents the number of characters in the sequence.

[0047] Step 2.7: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ; It was used to train the network through an auxiliary task: structured knowledge extraction, giving the network the ability to extract structured knowledge and enabling the network to use structured knowledge to improve the quality of description generation.

[0048] Step 2.8: The knowledge extraction module... Transformation is carried out, in In this method, only the character with the highest probability is selected at each position. The characters at all positions together form the predicted sequence. ,in, express The first in One character, express The number of characters in the sequence; the knowledge extraction module performs embedding representation on the knowledge representation sequence to obtain the embedding representation of the knowledge representation sequence. ;

[0049] Step 2.9: Language Embedding Module Perform embedding representation to obtain Representation vector ; The system contains the language category of the user-specified description, making it easier for the network to receive language information and generate descriptions in the corresponding language.

[0050] Step 2.10: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would , , , , The input is processed through stacked multi-head attention operations in the Transformer decoding module to obtain the decoded character probability sequence. ;in, express The Middle A character probability vector at each position. express The number of characters in the text.

[0051] Step 2.11: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ; It was used to train the network through the main task: cross-lingual description generation of text images.

[0052] Step 2.12: Calculate the total loss function ,in, express The corresponding weights; during training, different weights can be tried. The optimal value is determined through experimental results. Values.

[0053] Step 3: Construct a dataset for the multimodal text error correction task ,right The network is pre-trained using a multilingual text information encoding module, a multilingual visual information encoding module, and a multimodal Transformer module. During training, a multimodal text error correction task is first used to pre-train a portion of the network's structure, endowing the network with text error correction capabilities and mitigating the problem of text recognition error accumulation. Other training tasks are then performed after pre-training is complete.

[0054] Step 3.1: Construct the dataset for the text correction task ,in, express The corresponding identification label containing the error characters, express The correct identification label;

[0055] Step 3.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is sequentially input into the multilingual text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module for processing to obtain... The corresponding semantically enhanced feature vector sequence ,in, express The corresponding number 1 eigenvector for The number of characters in the text;

[0056] Step 3.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is fed into a character classifier for prediction, and the output is the probability of the character category corresponding to each position. ;in, Indicates the output of the first The probability vector of the location.

[0057] Step 3.4: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Convert to one-hot vector ,and Comparison to construct the loss function It is used to backpropagate the language text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module to update the module parameters, thereby obtaining the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, and the pre-trained multimodal Transformer module.

[0058] Step 4: Based on the total loss function A multilingual text-image cross-language description generative network was developed using the backpropagation algorithm. The multilingual image and text detection and recognition module, the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, the pre-trained pre-trained multimodal Transformer module, the knowledge extraction module, the language embedding module, and the Transformer decoding module are trained to update the network parameters, thereby obtaining the trained multilingual text and image cross-language description and generation model. Step 4 uses the total loss function Simultaneously, two training tasks are performed: a knowledge extraction task and a description generation task, to update the parameters in the network. The two tasks can complement each other, thereby improving the performance of description generation.

[0059] Step 5: Use the trained multilingual text-image cross-language description generation model to generate descriptions for the specified language categories. Any input multilingual text image Make predictions to obtain the results in the specified language category. Down The descriptive statement. Specifically, let the input multilingual text image be denoted as... The specified language category is Model for and The process involves predicting each token in the description statement one by one to obtain a probability sequence. Take the token with the highest probability value at each position in the probability sequence, form a description statement, and return it to the user.

[0060] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0061] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

Claims

1. A method for generating cross-language descriptions for multilingual text images, characterized in that, Includes the following steps: Step 1: Obtain multilingual text images and annotate the descriptive sentences to obtain the training set. ,in, Representing multilingual text images, express The corresponding image description, and , express The first in One character, express The number of characters in the text; express Corresponding languages express The corresponding linear representation of structured knowledge, and , express The first in One character, express The number of characters in the text; express The number of multilingual text images; Step 2: Construct a cross-lingual description generation network for multilingual text images It includes: a multilingual image and text detection and recognition module, a multilingual text information encoding module, a multilingual visual information encoding module, a multimodal Transformer module, a knowledge extraction module, a language embedding module, and a Transformer decoding module; and will Each sample input in Processing is carried out, including: Cropping is performed to obtain the text region image. and identify The character sequence in the text is used to obtain the text recognition result. ; then The text is processed to output a text feature representation matrix. At the same time, for Visual features are extracted to obtain Visual feature representation matrix ;right Visual features are extracted to obtain Visual feature representation matrix Then, , , The inputs are respectively entered into the multimodal Transformer module, and the corresponding outputs are... Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding , Corresponding semantically enhanced feature encoding Finally, , , The input is used in the knowledge extraction module for prediction to obtain the decoded character probability sequence. And used to construct the total loss function. ; Step 3: Construct a dataset for the multimodal text error correction task and used for The multilingual text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module are pre-trained to obtain the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, and the pre-trained multimodal Transformer module. Step 4: Based on the total loss function A multilingual text-image cross-language description generative network was developed using the backpropagation algorithm. The multilingual image and text detection and recognition module, the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, the pre-trained multimodal Transformer module, the knowledge extraction module, the language embedding module, and the Transformer decoding module are trained to update the network parameters, thereby obtaining the trained multilingual text and image cross-language description generation model. Step 5: Use the trained multilingual text-image cross-language description generation model to generate descriptions for the specified language categories. Any input multilingual text image Make predictions to obtain the results in the specified language category. Down The descriptive statement.

2. The method for generating cross-language descriptions of multilingual text images according to claim 1, characterized in that, Step 2 also includes the following steps: Step 2.6: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , The input is used in the knowledge extraction module for prediction to obtain the character probability sequence of the knowledge representation sequence. ,in, The knowledge representation sequence represents the first... A probability vector of characters. K Knowledge represents the number of characters in a sequence; Step 2.7: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ; Step 2.8: The knowledge extraction module... Transformation is performed, and in The character with the highest probability value is selected at each position, thus forming the prediction result of the knowledge representation sequence composed of characters at all positions. ,in, express The first in One character, express The number of characters in the text; The knowledge extraction module performs embedding representation on the knowledge representation sequence to obtain the embedding representation of the knowledge representation sequence. ; Step 2.9: The language embedding module... Perform embedding representation to obtain Representation vector ; Step 2.10: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would , , , , The input is subjected to a stacked multi-head attention operation in the Transformer decoding module to obtain the decoded character probability sequence. ;in, express The Middle A character probability vector at each position. express The number of characters in the text; Step 2.11: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Transform into a one-hot vector sequence and... To construct the loss function by making comparisons. ; Step 2.12: Calculate the total loss function ,in, express The corresponding weights.

3. The method for generating cross-language descriptions of multilingual text images according to claim 2, characterized in that, Step 3 includes the following steps: Step 3.1: Construct a dataset for the multimodal text correction task ,in, express The corresponding identification label containing the error characters, express The correct identification label; Step 3.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is sequentially input into the multilingual text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module for processing to obtain... The corresponding semantically enhanced feature vector sequence ,in, express The corresponding number 1 eigenvector for The number of characters in the text; Step 3.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require The data is fed into a character classifier for prediction, and the output is the probability of the character category corresponding to each position. ;in, Indicates the output of the first The probability vector of the location; Step 3.4: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Convert to one-hot vector and with Comparison to construct the loss function It is used to perform backpropagation on the language text information encoding module, the multilingual visual information encoding module, and the multimodal Transformer module to update the module parameters, thereby obtaining the pre-trained language text information encoding module, the pre-trained multilingual visual information encoding module, and the pre-trained multimodal Transformer module.

4. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the cross-language description generation methods of claims 1-3, the processor being configured to execute the program stored in the memory.

5. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by a processor, performs the steps of the cross-language description generation method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Audio information synthesis method and device, computer readable medium and electronic equipment

    CN112767910A

  • Image-multi-language subtitle conversion method based on interactive Transform

    CN114707523A