A Structured Information Extraction Method for Multilingual Natural Scene Text Images

By building a multi-modal information extraction network, combining text branches, visual branches and multi-modal information extraction machines, the problem of structured information extraction in multi-lingual natural scene text images is solved, and deep understanding and structured information extraction in multi-lingual multi-modal environments are achieved.

CN119516563BActive Publication Date: 2025-07-01HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411631527.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-07-01
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle structured information extraction in multilingual natural scene text images, especially in complex visual scenes and multimodal environments.

Method used

A structured information extraction method for multilingual natural scene text images is proposed. By constructing a multi-modal information extraction network, combining text branches, visual branches and multi-modal information extraction devices, it realizes multi-modal information fusion and efficient extraction of structured knowledge of multi-lingual text images.

Benefits of technology

It realizes deep understanding and structured information extraction of natural scene text images in a multilingual multimodal environment, overcomes the problem of poor information extraction effect in the existing technology, and improves the effect of structured knowledge extraction of multilingual text images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516563B_ABST
    Figure CN119516563B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting structured information from multi - language natural scene text images, and its steps include: 1. Constructing a dataset for extracting information from multi - language natural scene text images; 2. Constructing a multi - language and multi - modal information extraction network for natural scene text images; 3. Pre - training the text branch of the multi - modal information extraction network on the multi - language text information extraction dataset; 4. Training the multi - language and multi - modal information extraction network for natural scene text images; 5. Using the trained multi - modal information extraction network to extract information from any input multi - language text image, and obtaining a structured knowledge representation of the visual and language information in the text image. The present invention can, in a multi - language scenario, extract information from the input multi - language natural scene text images, deeply understand the information of different languages and different modalities in the text images, and output a structured knowledge representation of the text images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problems related to the information processing of multi - language natural scene text images, and particularly relates to a method for extracting structured information from multi - language natural scene text images. Background Art

[0002] The existing research on text image information extraction mainly focuses on document images, and realizes the information extraction of document images based on the layout information and text semantic information in the document images. Compared with document images, the background of natural scene text images is more complex, contains specific visual targets, and there is meaningful visual scene information. At the same time, the text content and the form of text presentation in natural scene text images are more flexible and diverse, which brings challenges to the research of structured knowledge extraction. For natural scene text images, the number of words in the text areas contained therein is usually small. The context of the text type in natural scene text images provides relatively weaker support for information extraction. Summary of the Invention

[0003] In order to solve the deficiencies of the above - mentioned existing technologies, the present invention proposes a method for extracting structured information from multi - language natural scene text images, aiming to simultaneously understand three different modalities of information, namely, visual targets in text images, text content in text images, and description statements corresponding to text images in a multi - language environment, so as to achieve efficient extraction of structured knowledge.

[0004] To achieve the above - mentioned invention purpose, the present invention adopts the following technical solutions:

[0005] A method for extracting structured information from multi - language natural scene text images according to the present invention is characterized by including the following steps:

[0006] Step 1: Obtain a multi - language text information extraction data set , where represents the i - th multi - language text, represents the language of the structured knowledge, represents the corresponding structured knowledge, represents the number of multi - language texts in

[0007] Obtain a multi - language natural scene text image set with annotations , where represents the j - th multi - language natural scene text image, represents the language of the structured knowledge annotation, represents the structured knowledge annotation of represents The number of multi - language natural scene text images;

[0008] Step 2: Construct a structured information extraction network for multi - language natural scene text images , including: a text branch, a visual branch, an image descriptor, and a multi - modal information extractor;

[0009] The text branch includes: 1 multi - language text information encoding module, 1 Transformer module, and 1 text information extraction module;

[0010] The visual branch includes: 1 multi - language text and image detection and recognition module, 1 multi - language visual information encoding module, 1 multi - language text information encoding module, and 1 pre - trained multi - modal Transformer module;

[0011] The multi - modal information extractor contains: 1 multi - modal information fusion module and 1 decoding module;

[0012] Step 3: Input into the text branch of the structured information extraction network for pre - training to obtain the pre - trained text branch;

[0013] Step 4: Input into the structured information extraction network for training to obtain the trained structured information extraction model;

[0014] Step 5: Use the trained structured information extraction model to perform information extraction on any input multi - language text image to obtain the predicted structured knowledge representation and output it as the information extraction result.

[0015] Another feature of the structured information extraction method for multi - language natural scene text images according to the present invention is that step 3 includes the following steps:

[0016] Step 3.1: Input into the multi - language text information encoding module and use the encoder of mT5 to process to obtain the embedding representation matrix at each position in , where represents the embedding representation vector at the -th position in , represents the number of characters in , and represents the embedding dimension; represents the embedding dimension;

[0017] Step 3.2: Input It is input into the input Transformer layer and processed through multiple stacked multi-head attention mechanisms, feed-forward operations, and residual connections to obtain the semantic feature matrix of , where represents the semantic feature at the th position in

[0018] Step 3.3: Input into the text information extractor for prediction to obtain the linearized knowledge representation , where represents the th character in the corresponding linearized knowledge representation, and

[0019] represents the number of characters in the linearized knowledge representation; is transformed into the linearized knowledge representation through the tree structure-based rules, and then compared with the predicted linearized knowledge representation to construct a loss function for backpropagation of the text branch to update the network parameters in the text branch, thereby obtaining the pre-trained text branch.

[0020] Furthermore, the step 4 includes the following steps:

[0021] Step 4.1: Transfer the weights of each node in the trained text branch network in step 3 to the nodes at the corresponding positions in the multi-lingual multi-modal information extraction network;

[0022] Step 4.2: Input into the multi-lingual text and image detection and recognition module in the visual branch for text detection and recognition, and respectively obtain the text region position coordinates , the text region cropped image and the recognition result ;

[0023] Step 4.3: Input and into the multi-lingual visual information encoding module for processing to obtain the visual embedding representation of the overall text image and the visual embedding representation of the text region cropped image, where represents the th feature vector in represents the th feature vector in denote the number of vectors in denote the number of vectors in;

[0024] Step 4.4: Based on and the positions of each part in calculate and the corresponding positions of each part in in and ;

[0025] Concatenate with at the corresponding positions to obtain the visual feature encoding of the overall text image;

[0026] Concatenate with at the corresponding positions to obtain the visual feature encoding of the cropped image of the text region;

[0027] Step 4.5: Input into the multi - language text information encoding module for processing to obtain the text embedding representation , where denote the th representation vector in denote the number of embedding vectors;

[0028] Based on the relative positions of each part in calculate to obtain the position encoding corresponding to each part in;

[0029] Concatenate with to obtain the text feature encoding ;

[0030] Step 4.6: Input , , into the pre - trained multi - modal Transformer layer for processing, and output the semantic - enhanced visual feature encoding corresponding to the multi - language text image , the semantic - enhanced visual feature encoding corresponding to the cropped image of the text region , and the recognition result The corresponding semantically enhanced text feature encoding ;

[0031] Step 4.7: Input into the image description module for processing to obtain image description statement ;

[0032] Step 4.8: Input into the pre-trained text branch. After being processed by the multilingual text information encoding module and the Transformer layer in sequence, output the corresponding semantically enhanced text feature encoding , where represents the th encoding vector in , and

[0033] , , , and into the multimodal information extractor for extracting multilingual text image information to obtain the prediction result of the linearized knowledge representation of the text image;

[0034] Step 4.10: Convert to the i-th linearized knowledge representation through tree-based rules, and compare it with the i-th predicted linearized knowledge representation to construct a loss function for backpropagation of the structured information extraction network to update the network parameters, thereby obtaining the trained structured information extraction model.

[0035] Furthermore, step 4.9 includes the following steps:

[0036] Step 4.9.1: The multimodal information fusion module fuses , , , and to obtain the i-th fused feature encoding , where represents the th fused feature vector of

[0037] Step 4.9.2: Input into the decoding module for prediction to obtain the i-th predicted linearized knowledge representation , wherein represents the th character in the corresponding linearized knowledge representation;

[0038] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the structured information extraction method, and the processor is configured to execute the program stored in the memory.

[0039] A computer-readable storage medium according to the present invention, characterized in that a computer program stored on the computer-readable storage medium executes the steps of the structured information extraction method when run by a processor.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] 1. The present invention can realize the information extraction of multi-language natural scene text images. In the case where the image contains complex visual scenes, the structured information extraction network for multi-language natural scene text images can be used to simultaneously understand the multi-language text information and visual scene information in the text image, overcoming the deficiency of the prior art that only uses the text information and text layout information in the image for information extraction, and can be better applied to multi-language and multi-modal environments.

[0042] 2. The present invention transfers the knowledge learned by the information extraction model on the massive text modal information extraction data set to the multi-language multi-modal information extraction model, and alleviates the problem of insufficient sample size of the multi-language text image information extraction data set through the use of external knowledge. Through the transfer of knowledge, the effect of structured knowledge extraction of multi-language text images is improved.

[0043] 3. The network framework designed by the present invention integrates information from different sources, different modalities, and different granularities in multi-language text images, enabling various information to complement each other and supporting the model to deeply understand multi-language natural scene text images from different angles, thereby achieving better structured information extraction performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a usage flow chart of the structured information extraction method for multi-language natural scene text images according to the present invention;

[0045] Figure 2 is a network structure diagram of the structured information extraction method for multi-language natural scene text images according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] In this embodiment, asFigure 1 As shown in Figure 1 , a method for extracting structured information from multi - language natural scene text images includes the following steps:

[0047] Step 1: Obtain a multi - language text information extraction dataset , where represents the i - th multi - language text, represents the language of the structured knowledge, represents the corresponding structured knowledge, represents the number of multi - language texts in For the i - th sample , during training, use and as inputs, as the supervision signal during training.

[0048] Obtain a multi - language natural scene text image set with annotations , where represents the j - th multi - language natural scene text image, represents the language of the structured knowledge annotation, represents the structured knowledge annotation of represents the number of multi - language natural scene text images in For the j - th sample , during training, use and as inputs, as the supervision signal during training.

[0049] Step 2: Construct a structured information extraction network for multi - language natural scene text images , including: a text branch (for extracting knowledge from text - modality inputs), a visual branch (for constructing features for multi - modal information extraction from visual - modality inputs), an image descriptor (for constructing descriptive statements of the input multi - language text images), and a multi - modal information extractor (for information extraction based on multi - modal features);

[0050] The text branch includes: 1 multi - language text information encoding module, 1 Transformer module, and 1 text information extraction module;

[0051] The visual branch includes: 1 multi - language text and image detection and recognition module, 1 multi - language visual information encoding module, 1 multi - language text information encoding module, and 1 pre - trained multi - modal Transformer module;

[0052] The multimodal information extractor includes: 1 multimodal information fusion module and 1 decoding module.

[0053] Step 3: Input to structured information extraction network The text branch in is pre-trained to obtain the pre-trained text branch. The pre-training aims to solve the problem of insufficient training samples for multilingual natural scene text image information extraction, so that the model can learn external knowledge from massive text data, thereby improving the performance of the multilingual text image structured knowledge extraction network.

[0054] Step 3.1: Input the multilingual text information into the encoding module and use the mT5 encoder to Process and obtain The embedding representation matrix for each position in ,in, express Middle The embedding representation vector of the position, express The number of characters in represents the embedding dimension;

[0055] Step 3.2: Input into the Transformer layer, and after multiple stacked multi-head attention mechanisms, feedforward operations and residual connections, we get The semantic feature matrix ,in, express Middle Semantic features of each position;

[0056] Step 3.3: Input the text information extractor for prediction and obtain linearized knowledge representation ,in, express The corresponding linearized knowledge representation characters, Indicates the number of characters in the linearized knowledge representation;

[0057] Step 3.4: Transformation into linear knowledge representation based on tree-structured rules , and then with the predicted linear knowledge representation A comparison is performed to construct a loss function, which is used to backpropagate the text branch to update the network parameters in the text branch, thereby obtaining a pre-trained text branch.

[0058] Step 4: Input structured information extraction network is trained to obtain a trained structured information extraction model;

[0059] Step 4.1: Transfer the weights of each node in the text branch network trained in Step 3 to the corresponding nodes in the multi-lingual multi-modal information extraction network. This is used to transfer the knowledge learned by the information extraction model from the massive text modal information extraction dataset to the multi-lingual multi-modal information extraction model.

[0060] Step 4.2: Input the multi-lingual graphic text detection and recognition module in the visual branch for text detection and recognition, respectively obtaining the text region position coordinates , the text region cropped image and the recognition result ;

[0061] Step 4.3: Input the and into the multi-lingual visual information encoding module for processing, obtaining the visual embedding representation of the overall text image and the visual embedding representation of the text region cropped image , where represents the th characteristic vector in , represents the th characteristic vector in , represents the number of vectors in ,

[0062] Step 4.4: Based on the positions of each part in and , calculate the positions corresponding to each part in in and in , representing and ;

[0063] After splicing and at the corresponding positions, the visual feature encoding of the overall text image is obtained;

[0064] After splicing and at the corresponding positions, the visual feature encoding of the text region cropped image is obtained.

[0065] Step 4.5: Input into the multilingual text information encoding module for processing to obtain text embedding representations , where represents the th representation vector in , and

[0066] Based on the relative positions of the parts in , calculate the position encodings corresponding to the parts in ;

[0067] Concatenate with to obtain the text feature encoding ;

[0068] Step 4.6: Input , , into the pre-trained multimodal Transformer layer for processing, and output the semantically enhanced visual feature encoding corresponding to the multilingual text image , the semantically enhanced visual feature encoding corresponding to the text region cropped image , and the semantically enhanced text feature encoding corresponding to the recognition result .

[0069] Step 4.7: Input into the image description module for processing to obtain the image description statement of ;

[0070] Step 4.8: Input into the pre-trained text branch, and after being processed by the multilingual text information encoding module and the Transformer layer in sequence, output the semantically enhanced text feature encoding corresponding to , where represents the th encoding vector in , and

[0071] Step 4.9: Input , , and Perform information extraction on multilingual text images in the input multimodal information extractor to obtain the prediction results of the linearized knowledge representation of the text images ;

[0072] Step 4.9.1: The multimodal information fusion module fuses , , and to obtain the i-th fused feature encoding , where represents 's -th fused feature vector, represents the number of fused feature vectors. Fusing information from different sources, modalities, and granularities within multilingual text images enables various information to complement each other and support the model's in-depth understanding of multilingual natural scene text images from different perspectives, thereby achieving better structured information extraction performance;

[0073] Step 4.9.2: Input into the decoding module for prediction to obtain the i-th predicted linearized knowledge representation , where represents 's -th character in the corresponding linearized knowledge representation; represents the total number of characters;

[0074] Step 4.10: Convert into the i-th linearized knowledge representation through tree-based rules, and compare it with the i-th predicted linearized knowledge representation to construct a loss function for backpropagation of the structured information extraction network to update the network parameters, thereby obtaining the trained structured information extraction model;

[0075] Step 5: Use the trained structured information extraction model to perform information extraction on any input multilingual text image to obtain the predicted structured knowledge representation and output it as the information extraction result. Specifically, denote the input multilingual text image as , process to predict each character in the linearized knowledge representation one by one to obtain a character sequence , and use tree-based rules to convert the character sequence into a structured knowledge representation and return it to the user as the information extraction result.

[0076] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0077] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

Claims

1. A method for extracting structured information from multilingual natural scene text images, characterized in that: The following steps are involved: Step 1: Obtain a multilingual text information extraction dataset ,in, represents the i-th multilingual text, A language for representing structured knowledge, express The corresponding structured knowledge, express The number of multilingual texts in the Obtain annotated multilingual natural scene text image sets ,in, represents the jth multilingual natural scene text image, The language in which the structured knowledge is annotated. express Structured knowledge annotation of express The number of multilingual natural scene text images in ; Step 2: Build a structured information extraction network for multilingual natural scene text images , including: a text branch, a visual branch, an image descriptor, and a multimodal information extractor; The text branch includes: a multilingual text information encoding module, a Transformer module, and a text information extraction module; The visual branch includes: a multilingual image and text detection and recognition module, a multilingual visual information encoding module, a multilingual text information encoding module, and a pre-trained multimodal Transformer module; The multimodal information extractor comprises: a multimodal information fusion module and a decoding module; Step 3: Input to structured information extraction network Pre-train the text branch in to obtain the pre-trained text branch; Step 4: Input structured information extraction network Train in and obtain the trained structured information extraction model; Step 5: Use the trained structured information extraction model to extract any input multilingual text image Information extraction is performed to obtain a predicted structured knowledge representation, which is output as the information extraction result.

2. The method for extracting structured information from multilingual natural scene text images according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1: Input the multilingual text information into the encoding module and use the mT5 encoder to Process and obtain The embedding representation matrix for each position in ,in, express Middle The embedding representation vector of the position, express The number of characters in represents the embedding dimension; Step 3.2: Input into the Transformer layer, and after multiple stacked multi-head attention mechanisms, feedforward operations and residual connections, we get The semantic feature matrix ,in, express Middle Semantic features of each position; Step 3.3: Input the text information extractor for prediction to obtain linearized knowledge representation ,in, express The corresponding linearized knowledge representation characters, Indicates the number of characters in the linearized knowledge representation; Step 3.4: Transformation into linear knowledge representation based on tree-structured rules , and then with the predicted linear knowledge representation A comparison is performed to construct a loss function, which is used to backpropagate the text branch to update the network parameters in the text branch, thereby obtaining a pre-trained text branch.

3. The method for extracting structured information from multilingual natural scene text images according to claim 2, characterized in that: The step 4 comprises the following steps: Step 4.1: Transfer the weights of each node of the text branch network trained in step 3 to the nodes at the corresponding positions of the multilingual multimodal information extraction network; Step 4.2: Input the multilingual image and text detection and recognition module in the visual branch to perform text detection and recognition, and obtain the text area position coordinates respectively , crop the image in the text area and recognition results ; Step 4.3: and Input the multilingual visual information encoding module for processing to obtain the visual embedding representation of the entire text image and visual embedding representation of cropped images of text regions ,in, express The A representation vector, express The A representation vector, express The number of vectors in , express The number of vectors in ; Step 4.4: Based on and The parts in The position in , calculate and The parts in The corresponding position in and ; Will and After splicing at the corresponding positions, the overall visual feature encoding of the text image is obtained. ; Will and After splicing at the corresponding positions, the visual feature encoding of the cropped image of the text area is obtained. ; Step 4.5: Input the multilingual text information into the encoding module for processing to obtain text embedding representation ,in, express The represents a vector, Represents the number of embedded vectors; based on The parts in The relative position in is calculated The position codes of the parts in ; Will and After splicing, we get the text feature encoding ; Step 4.6: , , Input the pre-trained multimodal Transformer layer for processing and output multilingual text images Corresponding semantically enhanced visual feature encoding , crop the image in the text area Corresponding semantically enhanced visual feature encoding , recognition results Corresponding semantically enhanced text feature encoding ; Step 4.7: Input the image descriptor for processing, and obtain Image description sentence ; Step 4.8: The pre-trained text is input into the branch and processed by the multilingual text information encoding module and the Transformer layer in turn, and then output Corresponding semantically enhanced text feature encoding ,in, express Middle The encoding vector, Indicates the number of vectors; Step 4.9: , , and Input the multimodal information extractor to extract information of multilingual text images, and obtain the prediction result of linearized knowledge representation of text images ; Step 4.10: Transformed into the i-th linearized knowledge representation through tree-based rules , and the linearized knowledge representation of the i-th prediction Compare to construct a loss function for structured information extraction network Back propagation is performed to update the network parameters to obtain the trained structured information extraction model.

4. The method for extracting structured information from multilingual natural scene text images according to claim 3, characterized in that: The step 4.9 comprises the following steps: Step 4.9.1: The multimodal information fusion module , , and Fusion is performed to obtain the feature encoding after the i-th fusion ,in, express No. The fused feature vectors, Indicates the number of fused feature vectors; Step 4.9.2: Input into the decoding module for prediction, and obtain the linearized knowledge representation of the i-th prediction ,in, express The corresponding linearized knowledge representation characters; Indicates the total number of characters.

5. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the structured information extraction method according to any one of claims 1 to 4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the structured information extraction method according to any one of claims 1 to 4 are performed.

Citation Information

Patent Citations

  • Multi-modal scene graph knowledge enhanced adversarial multi-modal pre-training method

    CN115331075A

  • Table semantic information extraction method, system and equipment based on cell coordinate optimization and medium

    CN116543404A