Document image understanding method and system for fine particles

By using the BPE word segmenter and patch segmentation technology, combined with text coordinate mapping and visual feature extraction, the problem of insufficient fine-grained feature capture in document image understanding is solved, achieving efficient and accurate document image understanding and alignment.

CN120976946AActive Publication Date: 2025-11-18CHENGDU DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Patent Information

Application Number
CN202511497376.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing document image understanding technologies cannot effectively capture the spatial distribution features of fine particles, resulting in insufficient visual feature accuracy, which affects the accuracy of alignment. Furthermore, pre-trained language models have high computational complexity, making it difficult to meet the needs of real-time processing scenarios.

Method used

The BPE word segmenter is used for token-level word segmentation, combined with patch segmentation and visual feature extraction. The visual encoding of token-level text is realized through a text coordinate mapping table. The intersection of patch-level visual features and token-level linguistic semantic features is calculated, and alignment is performed using a comparison model to enhance the accuracy of feature extraction and alignment.

Benefits of technology

It improves the efficiency of fine-grained visual feature extraction and the accuracy of language semantic feature extraction, reduces computational complexity, meets real-time processing requirements, and improves the accuracy of document image understanding and alignment precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976946A_ABST
    Figure CN120976946A_ABST
Patent Text Reader

Abstract

The invention relates to a fine particle document image understanding method and system, and the method comprises the steps: carrying out the text content recognition of an input document image, obtaining the pixel coordinate of each character in an obtained text sequence, constructing a text coordinate mapping table, carrying out the Token-level word segmentation processing of the text sequence through a BER word segmentation device, and obtaining a text coordinate mapping table; performing visual coding and language semantic feature extraction on all the obtained Token-level texts according to a text coordinate mapping table to obtain all Token-level visual codes and all Token-level language semantic features, performing patch segmentation on a document image, performing visual feature extraction on each obtained image packet, and obtaining a Token-level image; and calculating the intersection of each patch visual feature and all Token-level visual codes, mapping the patch visual features in the intersection and all Token-level language semantic features, and inputting a comparison model for alignment, so as to realize the understanding of the document image according to the obtained alignment result. Therefore, document image understanding of fine particles is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a fine-grained document image understanding method and system. BACKGROUND

[0002] The document image understanding technology refers to converting the image data of paper or electronic documents into structured text semantic information, and the core goal is to establish the association between "image visual information" and "text semantic information" to support information extraction, table parsing, multilingual translation and other downstream tasks, and is widely used in financial report processing, government form verification, medical prescription analysis and other fields.

[0003] The current mainstream document image understanding technology is to extract visual features of the document image through a convolutional neural network, extract language semantic features of the document image through a pre-trained language model, and then align the extracted visual features and language semantic features to realize the understanding of the document image. However, due to the limited field of view of the convolutional neural network, it can only capture spatial correlations within a limited range and cannot capture fine-grained spatial distribution features, resulting in insufficient precision of the obtained visual features, affecting the accuracy of subsequent alignment, and affecting the accuracy of understanding the document image. In addition, the pre-trained language model contains a large number of parameters, and the computational complexity is high when extracting language semantic features, which is difficult to meet the needs of real-time processing scenarios. SUMMARY

[0004] The technical problem to be solved by the present application is that the present application provides a fine-grained document image understanding method and system, which realizes fine-grained visual feature extraction while improving the efficiency of language semantic feature extraction, meets the needs of real-time processing scenarios, and improves the accuracy of alignment and the accuracy of understanding the document image.

[0005] In order to solve the above technical problems, the technical scheme adopted by the present application is: In a first aspect, the present application provides a fine-grained document image understanding method, comprising: obtaining an input document image, performing text content recognition on the document image to obtain a text sequence, obtaining the pixel coordinates of each character in the text sequence, mapping the pixel coordinates and the text sequence to obtain a text coordinate mapping table; performing Token-level segmentation processing on the text sequence through a BPE tokenizer to obtain all Token-level texts, performing visual coding on all Token-level texts according to the text coordinate mapping table to obtain all Token-level visual codes, and simultaneously performing language semantic feature extraction on all Token-level texts to obtain all Token-level language semantic features; perform patch segmentation on the document image to obtain K image patches, perform visual feature extraction on each image patch to obtain all patch-level visual features, calculate the intersection of each patch-level visual feature and all Token-level visual encodings, map the patch visual features in the intersection to all Token-level language semantic features to obtain mapped patch-level visual features and mapped Token-level language semantic features, input the mapped patch-level visual features and the mapped Token-level language semantic features into a comparison model for alignment to obtain an alignment result, and realize understanding of the document image according to the alignment result.

[0006] The present application has the advantages that: the BPE tokenizer is used to perform Token-level segmentation on the text sequence, i.e., the text sequence is finely granulated into Token-level segmentation, and each Token-level segmentation only carries a single semantic unit in the form of a minimum effective unit, without redundancy, which not only improves the accuracy of language semantic feature extraction, but also reduces the computational complexity and improves the extraction efficiency, meeting the needs of real-time processing scenarios. Meanwhile, the document image is segmented into patches, and visual features are extracted from each image patch to finely granulate the visual space, improve the accuracy of visual feature extraction, improve the accuracy of aligning patch-level visual features with Token-level language semantic features, and improve the accuracy of understanding the document image. The text coordinates mapping table is constructed in advance to visually encode the Token-level text, realize visual anchoring of the Token-level text, avoid misalignment between visual regions and text regions, calculate the intersection of each patch-level visual feature and all Token-level visual encodings, i.e., filter out the patch-level visual features related to the Token-level visual encodings, and map the patch-level visual features in the intersection to all Token-level language semantic features to avoid full-amount patch-level visual features from participating in the calculation, save computing resources, and improve mapping efficiency.

[0007] Optionally, the Token-level segmentation of the text sequence by the BPE tokenizer obtains all Token-level texts, including: obtaining a basic vocabulary table of the BPE tokenizer, performing addition operation including mathematical symbols, multilingual characters and document structure markers on the basic vocabulary table to obtain the basic vocabulary table after the addition operation, performing Token-level segmentation on the text sequence according to the basic vocabulary table after the addition operation to obtain all Token-level texts, and all Token-level texts including Token-level texts matched with the basic vocabulary table after the addition operation and Token-level texts not matched with the basic vocabulary table after the addition operation; The Token-level text that does not match the basic vocabulary table after the addition operation is an out-of-vocabulary word, the character length of each out-of-vocabulary word is calculated, the out-of-vocabulary word is optimized according to the character length, it is judged whether the character length is greater than a length threshold, if not, the out-of-vocabulary word is optimized for unknown word marking to obtain an optimized out-of-vocabulary word, and if yes, the out-of-vocabulary word is optimized for character decomposition by character-level coding to obtain an optimized out-of-vocabulary word; All Token-level texts are updated according to the optimized out-of-vocabulary words to obtain updated Token-level texts.

[0008] According to the above description, the basic vocabulary table of the BPE tokenizer is added, while maintaining its basic tokenization ability, overcoming the defects of being unable to process mathematical symbols and multi-language symbols, avoiding the loss or error segmentation of important information, and differentiating the out-of-vocabulary words according to their character lengths, the out-of-vocabulary words greater than the length threshold are decomposed by character-level coding, which can better preserve morphological information and semantic information, reduce semantic loss, and thus improve the accuracy and integrity of the obtained Token-level texts.

[0009] Optionally, the language semantic features of all Token-level texts are extracted at the same time to obtain Token-level language semantic features of all Token-level texts, including: An embedded character level corresponding to each Token-level text is generated by a character-level embedding formula; An embedded subword level corresponding to each Token-level text is generated by a subword-level embedding formula; An embedded semantic level corresponding to each Token-level text is generated by a semantic-level embedding formula; Token-level language semantic features of each Token-level text are dynamically fused by a gating fusion mechanism to generate Token-level language semantic features of each Token-level text, and Token-level language semantic features of all Token-level texts are obtained; The character-level embedding formula is: ; Wherein, represents an embedded character level corresponding to the i-th Token-level text, represents a character set of the i-th Token-level text, represents a character embedding matrix, and c represents a single character; The subword-level embedding formula is: ; Wherein, denotes the embedding subword level corresponding to the i-th Token level text, denotes the subword embedding matrix, denotes the row index of the i-th Token level text; The semantic level embedding formula is: ; wherein, denotes the embedding semantic level corresponding to the i-th Token level text, denotes the multi-layer perception, denotes the part-of-speech tag of the i-th Token level text, denotes the entity tag of the i-th Token level text, denotes the one-hot encoding function, denotes the feature splicing symbol.

[0010] According to the above description, the embedding character level, the embedding subword level and the embedding semantic level of each Token level text are generated, the embedding character level can capture the semantic information and character pattern inside the Token level text, the embedding subword level can capture the semantic similarity and grammatical relationship between different subwords, and the embedding semantic level can understand the semantic information of the Token level text at the syntactic and semantic levels. And through the gating fusion mechanism, the obtained embedding character level, embedding subword level and embedding semantic level are dynamically fused, so as to improve the accuracy and integrity of the obtained Token level language semantic features.

[0011] Optionally, the patch segmentation of the document image to obtain K image patches comprises: obtaining the image width, image height and document type of the document image; inputting the image width and the image height into a text density formula to calculate the text density of the document image, wherein the text density formula is: ; wherein, denotes the text density of the document image I, H denotes the image height, and W denotes the image width, denotes the indication function of whether the pixel (x, y) is a text pixel; According to the optimal scaling ratio of the document image, the document image is scaled according to the optimal scaling ratio to obtain a scaled document image; According to the document type, the patch size is dynamically adjusted to perform adaptive size patch segmentation on the scaled document image to obtain K image patches.

[0012] According to the above description, when the document image is patch segmented, the optimal scaling ratio of the document image is calculated according to the text density of the document image, and then the patch size is adaptively adjusted according to the document type, so as to realize dynamic patch segmentation, get rid of the limitation of traditional fixed size scaling and fixed size segmentation, and drive the document density and the document type, so that the obtained image patch retains the fine-grained detail information.

[0013] Optionally, the visual feature extraction of each image patch to obtain all patch-level visual features comprises: The internal local feature of each image patch is extracted by using a convolutional neural network to obtain all patch-level local features. All patch-level local features are mapped to a fixed-dimensional embedding space through a linear projection layer to obtain all mapped patch-level local features, and the optimization processing of all mapped patch-level local features is performed to obtain all optimized patch-level local features, and the optimization processing includes residual connection and normalization processing.

[0014] According to the above description, the internal spatial correlation of a single image patch is captured by a convolutional neural network, so that the obtained patch-level local features retain their own detail features, which are mapped to a fixed-dimensional embedding space through a linear projection, eliminating the difference in spatial dimension, realizing dimension standardization, and performing residual connection to retain the original patch-level local features, avoiding feature degradation, and stabilizing the stable distribution of patch-level local features in the embedding space through normalization processing.

[0015] Optionally, the intersection of each patch-level visual feature and all Token-level visual encoding is calculated, and the patch visual feature in the intersection is mapped with all Token-level language semantic features, comprising: The row index and the column index of each image patch in the document image are obtained, and the row index and the column index are input into an absolute position formula for absolute position encoding calculation to obtain the absolute position encoding of each image patch, and the row index and the column index of each two image patches are input into a relative position formula for relative position encoding calculation to obtain the relative position encoding of each image patch. The absolute position encoding of each image patch is fused with the corresponding relative position encoding to obtain a two-dimensional hybrid position encoding of each image patch, and the two-dimensional hybrid position encoding is added to the patch-level visual feature corresponding to the image patch to implement feature enhancement on the patch-level visual feature, thereby obtaining all enhanced patch-level visual features; Sequence position information and a document structure type to which each Token-level language semantic feature corresponds in the text sequence are obtained. Meanwhile, local context information of each Token-level language semantic feature is calculated, and local context encoding is performed within a preset correlation window with each Token-level language semantic feature as the center, thereby obtaining local context information of each Token-level language semantic feature. The local context information, sequence position information and document structure type of each Token-level language semantic feature are fused to obtain context position encoding of each Token-level language semantic feature, and the context position encoding is added to the corresponding Token-level language semantic feature to implement feature enhancement on the Token-level language semantic feature, thereby obtaining all enhanced Token-level language semantic features. An intersection of each enhanced patch-level visual feature and all Token-level visual encodings is calculated, and all enhanced patch-level visual features in the intersection are mapped with all enhanced Token-level language semantic features.

[0016] According to the above description, the two-dimensional hybrid position encoding composed of absolute position encoding and relative position encoding is added to the patch-level visual feature, the position information of the image patch itself is established, and the associated position information between different image patches is also established, thereby solving the spatial disorder of the image patch and ensuring the spatial correlation of the patch-level visual feature. The context position encoding composed of local context information, sequence position information and document structure type is added to the Token-level language semantic feature, thereby eliminating semantic ambiguity, avoiding Token-level language semantic feature sequence disorder, and establishing the logical correlation of the Token-level language semantic feature. Since the corresponding position information is added to the enhanced patch-level visual feature and the enhanced Token-level language semantic feature, cross-modal mapping can be realized, multi-scene requirements can be met, and adaptability and flexibility can be improved.

[0017] Optionally, the mapping of all enhanced patch-level visual features in the intersection with all enhanced Token-level language semantic features comprises: An effective attention range of each enhanced patch-level visual feature is calculated through the sparse attention mechanism, and feature association is performed on the enhanced patch-level visual feature within the effective attention range through the multi-head attention mechanism to obtain all feature-associated patch-level visual features; All feature-associated patch-level visual features are subjected to multi-scale interpolation fusion to obtain multi-scale interpolation fused patch-level visual features, and different pooling strategies are triggered according to the spatial layout of all Token-level visual encodings to perform pooling processing on the multi-scale interpolation fused patch-level visual features, so as to realize mapping of the multi-scale interpolation fused patch-level visual features and all enhanced Token-level language semantic features.

[0018] According to the above description, the multi-head attention mechanism is used to perform feature association on the enhanced patch-level visual features, solving the problem of feature isolation and lack of context of a single enhanced patch-level visual feature. The sparse attention is used to calculate the effective attention range, and the feature association is performed within the effective attention range, avoiding the interference of invalid enhanced patch-level visual features, reducing the parameter size, improving the inference speed of feature association, realizing lightweight, and performing multi-scale interpolation fusion and adaptive pooling processing on all feature-associated patch-level visual features, further improving the detail distinction of the patch-level visual features, overcoming the problem of important information loss or noise introduction existing in the traditional fixed pooling processing, and ensuring the accuracy of the mapping of the multi-scale interpolation fused patch-level visual features and all enhanced Token-level language semantic features.

[0019] Optionally, the mapped patch-level visual features and the mapped Token-level language semantic features are input into a comparison model for alignment to obtain an alignment result, and the understanding of the document image is realized according to the alignment result, including: Token-level similarity, phrase-level similarity and sentence-level similarity between each mapped patch-level visual feature and each mapped Token-level language semantic feature are calculated; The mapped patch-level visual features and the mapped Token-level language semantic features are subjected to multi-level alignment according to the Token-level similarity, the phrase-level similarity and the sentence-level similarity to obtain a multi-level alignment result; The multi-level alignment result is subjected to alignment fusion to obtain a final alignment result, and the understanding of the document image is realized according to the final alignment result.

[0020] According to the above description, through the multi-level alignment of the Token level similarity, the phrase level similarity and the sentence level similarity, progressive alignment from fine-grained, medium-grained and coarse-grained is realized, and the multi-level alignment results are fused, so that the final alignment result takes into account the multi-granularity requirement, improves the alignment accuracy, and improves the accuracy of document image understanding.

[0021] Optionally, further comprising: The alignment quality degree of the alignment result is calculated through an alignment quality formula, and the alignment consistency degree of the alignment result is calculated through an alignment consistency detection formula; The alignment quality degree and the alignment consistency degree are input into an alignment scoring formula for scoring to obtain a scoring result, and it is judged whether the scoring result is lower than a scoring threshold, if yes, the alignment model is optimized according to the scoring result to obtain an optimized alignment model; The alignment quality formula is: ; Wherein, The alignment quality degree is represented by N, the total number of the mapped patch level visual features and the mapped Token level language semantic features, The i-th mapped patch level visual feature is represented by The i-th mapped Token level language semantic feature is represented by The final similarity of the i-th mapped patch level visual feature and the i-th mapped Token level language semantic feature is represented by The consistency detection formula is: ; Wherein, The alignment consistency degree is represented by Var, the variance, The final similarity of the i-th mapped patch level visual feature and the i-th mapped Token level language semantic feature is represented by The maximum variance is represented by N, the total number of the mapped patch level visual features and the mapped Token level language semantic features; The alignment scoring formula is: ; Wherein, The scoring result is represented by The quality weight coefficient of the alignment quality degree is represented by The consistency degree weight coefficient of the alignment consistency degree is represented by The alignment quality degree is represented by The alignment consistency degree is represented by

[0022] According to the above description, the alignment quality degree and the alignment consistency degree of the alignment result are calculated, the alignment result is scored according to the alignment quality degree and the alignment consistency degree, the alignment model is optimized according to the scoring result, and the alignment accuracy of the alignment model is improved.

[0023] In a second aspect, the present application provides a fine-grained document image understanding system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the fine-grained document image understanding method of the first aspect when executing the computer program.

[0024] The technical effects of the fine-grained document image understanding system of the second aspect are the same as those of the fine-grained document image understanding method of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 A flowchart of the fine-grained document image understanding method provided by the present embodiment; Figure 2 A schematic diagram of the overall flow of the fine-grained document image understanding method provided by the present embodiment; Figure 3 A schematic diagram of the process of mapping the patch visual features in the intersection to all Token-level language semantic features involved in the present embodiment; Figure 4 A schematic diagram of the structure of the fine-grained document image understanding system provided by the present embodiment.

[0026] REFERENCE NUMERALS 1. A fine-grained document image understanding system; 2. A processor; 3. A memory. DETAILED DESCRIPTION

[0027] In order to better understand the above technical solutions, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer, more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0028] Embodiment One

[0029] Please refer to Figures 1 to 3 The present application provides a fine-grained document image understanding method, comprising the steps of: S1, obtain an input document image, perform text content recognition on the document image to obtain a text sequence, obtain pixel coordinates of each character in the text sequence, map the pixel coordinates and the text sequence to obtain a text coordinate mapping table; In the embodiment, as shown in the figure, the input document image is subjected to text content recognition to obtain a text sequence. At this time, a lightweight OCR model can be used for text content recognition. Pixel coordinates of each character in the text sequence are obtained to realize mapping of the pixel coordinates and the text sequence and to construct and obtain a text coordinate mapping table. Figure 2 In the embodiment, as shown in the figure, the input document image is subjected to text content recognition to obtain a text sequence. At this time, a lightweight OCR model can be used for text content recognition. Pixel coordinates of each character in the text sequence are obtained to realize mapping of the pixel coordinates and the text sequence and to construct and obtain a text coordinate mapping table.

[0030] S2, perform Token-level segmentation processing on the text sequence by using a BPE segmenter to obtain all Token-level texts, perform visual coding on all Token-level texts according to the text coordinate mapping table to obtain all Token-level visual codes, and perform language semantic feature extraction on all Token-level texts to obtain all Token-level language semantic features; In the embodiment, Token-level segmentation processing is performed on the text sequence obtained in step S1 by using a BPE segmenter to obtain all Token-level texts. Visual coding is performed on all Token-level texts according to the text coordinate mapping table. At this time, pixel coordinates of all characters constituting each Token-level text are extracted from the text coordinate mapping table, a minimum enclosing matrix of the pixel coordinates is calculated, a binary matrix with consistent dimensions of the document image is constructed, pixels in the minimum enclosing matrix are marked as 1 and the rest are marked as 0, and visual coding is performed on each Token-level text to obtain all Token-level visual codes. The Token-level visual codes at this time only contain pixel range information and do not contain language semantic information, which are used for positioning patch visual features of the image patch. Language semantic feature extraction is performed on all Token-level texts to obtain all Token-level language semantic features.

[0031] At this time, the step of performing Token-level segmentation processing on the text sequence by using a BPE segmenter to obtain all Token-level texts includes: S21, obtain a basic vocabulary table of the BPE segmenter, perform addition operation including mathematical symbols, multilingual characters and document structure markers on the basic vocabulary table to obtain a basic vocabulary table after addition operation, perform Token-level segmentation processing on the text sequence according to the basic vocabulary table after addition operation to obtain all Token-level texts, and all Token-level texts include Token-level texts matched with the basic vocabulary table after addition operation and Token-level texts not matched with the basic vocabulary table after addition operation; S22, the Token-level text that does not match the basic vocabulary table after the addition operation is taken as an out-of-vocabulary word, the character length of each out-of-vocabulary word is calculated, the out-of-vocabulary word is optimized according to the character length, it is judged whether the character length is greater than a length threshold, if not, the out-of-vocabulary word is optimized for unknown word marking, and the optimized out-of-vocabulary word is obtained, if yes, the out-of-vocabulary word is optimized for character decomposition by character-level coding, and the optimized out-of-vocabulary word is obtained; S23, all Token-level texts are updated according to the optimized out-of-vocabulary word, and all updated Token-level texts are obtained.

[0032] In the embodiment, as shown in the figure, Figure 2 the BPE tokenizer originally based on the basic vocabulary table is added with mathematical symbols, multi-language characters and document structure marks, the tokenization capability of the BPE tokenizer is expanded, the Token-level text sequence is tokenized according to the basic vocabulary table after the addition operation, and all Token-level texts are obtained, including Token-level texts matching the basic vocabulary table after the addition operation and Token-level texts not matching the basic vocabulary table after the addition operation. For Token-level texts not matching the basic vocabulary table after the addition operation, the Token-level texts are taken as out-of-vocabulary words, and the out-of-vocabulary words are optimized differently. It is judged whether the character length of the out-of-vocabulary word is greater than a length threshold. Since it is considered that the character length not greater than 3 is usually an abbreviation, a symbol or a functional word, the semantic information is relatively simple and does not cause serious information loss, and the character length greater than 3 usually contains rich semantic information, therefore, the length threshold is set to 3, that is, the out-of-vocabulary word with the character length not greater than 3 is directly optimized for unknown word marking, and the out-of-vocabulary word with the character length greater than 3 is directly decomposed into a character sequence, so as to obtain the optimized out-of-vocabulary word. The optimized out-of-vocabulary word is replaced by the original Token-level text, that is, the Token-level text is updated, and all updated Token-level texts are obtained.

[0033] At this time, the language semantic features of all Token-level texts are extracted at the same time in step S2, and all Token-level language semantic features are obtained, including: S24, an embedding character level corresponding to each Token-level text is generated by a character-level embedding formula; S25, an embedding subword level corresponding to each Token-level text is generated by a subword-level embedding formula; S26, an embedding semantic level corresponding to each Token-level text is generated by a semantic-level embedding formula; S27, dynamically fusing the embedded character level, the embedded subword level and the embedded semantic level of each Token-level text through a gated fusion mechanism to generate Token-level language semantic features of each Token-level text, and obtaining all Token-level language semantic features; The character level embedding formula is: ; wherein, represents the embedded character level corresponding to the i-th Token-level text, represents the character set of the i-th Token-level text, represents a character embedding matrix, and c represents a single character; The subword level embedding formula is: ; wherein, represents the embedded subword level corresponding to the i-th Token-level text, represents a subword embedding matrix, represents the row index of the i-th Token-level text; The semantic level embedding formula is: ; wherein, represents the embedded semantic level corresponding to the i-th Token-level text, represents a multi-layer perception, represents the part-of-speech tag of the i-th Token-level text, represents the entity tag of the i-th Token-level text, represents a one-hot encoding function, represents a feature splicing symbol.

[0034] In this embodiment, as shown in Figure 2 , the corresponding embedded character level, embedded word level and embedded semantic level are generated for each Token-level text through the character level embedding formula, the subword level embedding formula and the semantic level embedding formula, and the embedded character level, the embedded subword level and the embedded semantic level of each Token-level text are dynamically fused through a gated fusion mechanism. When dynamically fusing each Token-level text, the gated fusion mechanism splices the embedded character level, the embedded subword level and the embedded semantic level corresponding to each Token-level text along the feature dimension and inputs them into a gated weight calculation network for calculation to generate dynamic gated weights in three dimensions, i.e., the weight of the embedded character level, the weight of the embedded subword level, and the weight of the embedded semantic level. The specific formula is as follows: ; wherein, a weight of an embedding character level of an i-th Token level text, a weight of an embedding subword level of an i-th Token level text, a weight of an embedding semantic level of an i-th Token level text, a weight of an embedding character level of an i-th Token level text, a function, a bias quantity with a dimension of d, a gating weight matrix, and a dimension of the gating weight matrix is d x 3d.

[0035] The embedding character level, the embedding subword level, and the embedding semantic level are further weighted and fused by the dynamic gating weight of three dimensions to generate a Token level language semantic feature of each Token level text, and a specific formula is as follows: ; wherein, a Token level language semantic feature.

[0036] S3, performing patch segmentation on the document image to obtain K image patches, performing visual feature extraction on each image patch to obtain all patch level visual features, calculating an intersection of each patch level visual feature and all Token level visual encodings, mapping patch visual features in the intersection to all Token level language semantic features to obtain mapped patch level visual features and mapped Token level language semantic features, inputting the mapped patch level visual features and the mapped Token level language semantic features into a comparison model for alignment to obtain an alignment result, and realizing understanding of the document image according to the alignment result.

[0037] At this time, the performing patch segmentation on the document image in step S3 to obtain K image patches comprises: S31, obtaining an image width, an image height, and a document type of the document image; S32, inputting the image width and the image height into a text density formula to calculate a text density of the document image, and the text density formula is: ; wherein, denotes a text density of a document image I, H denotes an image height, W denotes an image width, denotes an indicator function of whether a pixel (x, y) is a text pixel; S33, calculating an optimal scaling ratio of the document image according to the text density, and scaling the document image according to the optimal scaling ratio to obtain a scaled document image. S34. Dynamically adjust the patch size according to the document type to perform adaptive patch segmentation on the scaled document image, resulting in K image patches.

[0038] In this embodiment, as Figure 2 As shown, the document image's width, height, and document type are obtained. The text density of the document image is calculated using the formula for image height, width, and text density. During the optimal scaling calculation, the document image is preprocessed according to dimensions from a preset size set, resulting in a preprocessed document image. The text density is then recalculated based on the preprocessed image's width and height, yielding the preprocessed document image's document density. The optimal scaling ratio is then calculated based on all text densities. The document image is scaled according to this optimal scaling ratio, resulting in a scaled document image. The patch size is dynamically adjusted based on the document type to adaptively segment the scaled document image. Document types include: natural scene text, tables and charts, dense documents, and GIU interfaces. Natural scene text uses 16×16 pixels, tables and charts use 14×14 pixels, dense documents use 12×12 pixels, and GIU interfaces use 18×18 pixels. The preset size set is {0.5, 0.75, 1.0, 1.25, 1.5}, and the formula for the optimal scaling ratio is as follows: ; in, This indicates the optimal scaling ratio. This represents the text density of a document image after preprocessing it according to size s from a preset set of sizes. This represents the optimal text density threshold, s represents the current size, and S represents the preset size set. express function.

[0039] At this point, the visual feature extraction performed on each image patch in step S3 yields all patch-level visual features, including: S35. Use a convolutional neural network to extract the internal local features of each image patch to obtain all patch-level local features; S36. All patch-level local features are mapped to a fixed-dimensional embedding space through a linear projection layer to obtain all mapped patch-level local features. All mapped patch-level local features are then optimized to obtain all optimized patch-level local features. The optimization process includes residual connection and normalization.

[0040] In the embodiment, a 3x3 convolutional neural network is used to extract internal features of each image patch to obtain all patch-level local features. First, a flattening operation is performed on all patch-level local features to convert them into patch-level local features with the same size as the image patch. Then, a linear projection layer is used to map all patch-level local features to a fixed-dimensional embedding space, where the fixed dimension is d. Finally, optimization processing including residual connection and normalization processing is performed on all mapped patch-level local features.

[0041] At this time, when the patch visual features in the intersection are mapped with all Token-level language semantic features in step S3, both the patch visual features and the Token-level language semantic features are enhanced. The specific steps are as follows: Obtain the row index and column index of each image patch in the document image, input the row index and column index into the absolute position formula to calculate the absolute position encoding, and obtain the absolute position encoding of each image patch. At the same time, input the row index and column index of each two image patches into the relative position formula to calculate the relative position encoding, and obtain the relative position encoding of each image patch. Weighted fusion of the absolute position encoding and the corresponding relative position encoding of each image patch is performed to obtain the two-dimensional mixed position encoding of each image patch. The two-dimensional mixed position encoding is added to the patch-level visual feature corresponding to the image patch to enhance the patch-level visual feature, and all enhanced patch-level visual features are obtained. Obtain the sequence position information and the document structure type of the Token-level text corresponding to each Token-level language semantic feature in the text sequence. At the same time, local context information calculation is performed on each Token-level language semantic feature. The local context encoding is performed within a preset association window with each Token-level language semantic feature as the center to obtain the local context information of each Token-level language semantic feature. Weighted fusion of the local context information, sequence position information and document structure type of each Token-level language semantic feature is performed to obtain the context position encoding of each Token-level language semantic feature. The context position encoding is added to the corresponding Token-level language semantic feature to enhance the Token-level language semantic feature, and all enhanced Token-level language semantic features are obtained. The intersection of each enhanced patch-level visual feature and all Token-level visual encodings is calculated, and all enhanced patch-level visual features in the intersection are mapped with all enhanced Token-level language semantic features.

[0042] In this embodiment, as shown in FIG. 6, the row index and column index of each image patch in the document image are input into the absolute position formula for absolute position encoding calculation, and the row index and column index of each two image patches are input into the relative position formula for relative position encoding calculation to obtain the absolute position encoding and the relative position encoding of each image patch, and the absolute position encoding and the relative position encoding are weighted and fused to obtain a two-dimensional hybrid position encoding of an image patch. Figure 3 ​​​​​​​​​​​​​​​​​The two-dimensional mixed position coding is added to the patch-level visual feature corresponding to the image patch to realize feature enhancement on the patch-level visual feature, and all enhanced patch-level visual features are obtained. The sequence position of the Token-level text corresponding to each Token-level language semantic feature in the text sequence and the document structure type are obtained, and local context information calculation is performed on each Token-level language semantic feature. The local context information calculation is respectively performed with each Token-level language semantic feature as the center, and local context coding is performed within a preset correlation window to obtain the local context information of each Token-level language semantic feature. At this time, the preset correlation window is 5. The local context information, sequence position information and document structure type of each Token-level language semantic feature are weighted and fused to obtain the context position coding of each Token-level language semantic feature. The context position coding is added to the corresponding Token-level language semantic feature to realize feature enhancement on the Token-level language semantic feature, and all enhanced Token-level language semantic features are obtained. The intersection of each enhanced patch-level visual feature and all Token-level visual encodings is calculated, and all enhanced patch-level visual features in the intersection are mapped with all enhanced Token-level language semantic features.

[0043] At this time, when all enhanced patch-level visual features in the intersection are mapped with all enhanced Token-level language semantic features, a sparse attention mechanism and a multi-head attention mechanism are used, and the specific steps are as follows: The effective attention range of each enhanced patch-level visual feature is calculated by the sparse attention mechanism, and the enhanced patch-level visual feature is associated by the multi-head attention mechanism within the effective attention range, and all feature-associated patch-level visual features are obtained. All feature-associated patch-level visual features are subjected to multi-scale interpolation fusion to obtain multi-scale interpolation fused patch-level visual features. Different pooling strategies are triggered according to the spatial layout of all Token-level visual encodings to pool the multi-scale interpolation fused patch-level visual features, so as to realize the mapping of the multi-scale interpolation fused patch-level visual features and all enhanced Token-level language semantic features.

[0044] In this embodiment, as Figure 3As shown, the effective attention range of each enhanced patch-level visual feature is calculated by the sparse attention mechanism, wherein the effective attention and range include attention within a preset local window and attention corresponding to a globally key Token text, which is filtered from all Token-level texts according to a preset key rule, and the enhanced patch-level visual feature is associated with features within the effective attention range by the multi-head attention mechanism, to obtain all feature-associated patch-level visual features. The all feature-associated patch-level visual features are subjected to multi-scale interpolation fusion, and the multi-scale is 0.5, 1.0 and 2.0 at this time. Different pooling strategies are triggered according to the spatial layout of all Token-level visual encodings to pool the multi-scale interpolation fused patch-level visual features, such as: when the spatial layout of all Token-level visual encodings is: compact distribution, the average pooling strategy is used, when the spatial layout of all Token-level visual encodings is: complex distribution, the attention pooling strategy is used, and when the spatial layout of all Token-level visual encodings is: text-intensive, the maximum pooling strategy is used.

[0045] At this time, the mapped patch-level visual features and the mapped Token-level language semantic features are input into the alignment model for alignment in step S3 to obtain an alignment result, and understanding of the document image is realized according to the alignment result, including: S37, calculating Token-level similarity, phrase-level similarity and sentence-level similarity between each mapped patch-level visual feature and each mapped Token-level language semantic feature; S38, performing multi-level alignment of the mapped patch-level visual features and the mapped Token-level language semantic features according to the Token-level similarity, the phrase-level similarity and the sentence-level similarity to obtain a multi-level alignment result; S39, performing alignment fusion on the multi-level alignment result to obtain a final alignment result, and understanding of the document image is realized according to the final alignment result.

[0046] In the embodiment, the Token-level similarity, the phrase-level similarity and the sentence-level similarity between each mapped patch-level visual feature and each mapped Token-level language semantic feature are calculated. When calculating the Token-level similarity, the Token-level complexity of the current Token-level language semantic feature is calculated first, and then the similarity formula is selected adaptively according to the Token-level complexity. Similarly, when calculating the phrase-level similarity, the phrase-level complexity of the current Token-level language semantic feature is calculated first, and then the similarity formula is selected adaptively according to the phrase-level complexity. Similarly, when calculating the sentence-level similarity, the sentence-level complexity of the current Token-level language semantic feature is calculated first, and then the similarity formula is selected adaptively according to the sentence-level complexity. When selecting the similarity formula, the following principles are followed: when the calculated complexity is less than a first complexity threshold, the cosine similarity calculation formula is used; when the calculated complexity is between the first complexity threshold and a second complexity threshold, the bilinear similarity calculation formula is used; and when the calculated complexity is greater than or equal to the second threshold, the MLP similarity calculation formula is used.

[0047] According to the Token-level similarity, the phrase-level similarity and the sentence-level similarity, the mapped patch-level visual features and the mapped Token-level language semantic features are aligned at multiple levels to obtain a multi-level alignment result. The multi-level alignment result is fused to obtain a final alignment result, and the final alignment result is used to realize understanding of the document image.

[0048] In the embodiment, the final alignment result is also evaluated to realize optimization of the alignment model in the reverse direction. The specific steps are as follows: The alignment quality degree of the alignment result is calculated by an alignment quality formula, and the alignment consistency degree of the alignment result is calculated by an alignment consistency detection formula; The alignment quality degree and the alignment consistency degree are input into an alignment scoring formula to obtain a scoring result. It is judged whether the scoring result is lower than a scoring threshold. If yes, the alignment model is optimized according to the scoring result to obtain an optimized alignment model; The alignment quality formula is as follows: ; Wherein, represents the alignment quality degree, N represents the total number of the mapped patch-level visual features and the mapped Token-level language semantic features, represents the i-th mapped patch-level visual feature, represents the i-th mapped Token-level language semantic feature, the final similarity between the i-th mapped patch-level visual feature and the i-th mapped Token-level language semantic feature; The consistency detection formula is: ; wherein, represents the alignment consistency degree, and Var represents the variance, the final similarity between the i-th mapped patch-level visual feature and the i-th mapped Token-level language semantic feature, represents the maximum variance, and N represents the total number of the mapped patch-level visual features and the mapped Token-level language semantic features; The alignment score formula is: ; wherein, represents the score result, represents the quality weight coefficient of the alignment quality degree, represents the consistency weight coefficient of the alignment consistency degree, represents the alignment quality degree, represents the alignment consistency degree.

[0049] In the embodiment, as shown in Figure 2 , the alignment quality degree of the alignment result is calculated by the alignment formula, and the alignment consistency degree of the alignment result is calculated by the alignment consistency detection formula, the alignment result is scored according to the obtained alignment quality degree, alignment consistency degree and alignment score formula, if the score result is lower than the score threshold, the comparison model is optimized according to the score result, the comparison model will calculate the comparison loss of the Token-level similarity, the comparison loss of the phrase-level similarity and the comparison loss of the sentence-level similarity in the process of optimization, and the losses are dynamically fused by the loss dynamic fusion mechanism to realize joint optimization, and the pre-trained visual feature extraction model used when visual feature extraction is performed and the pre-trained language semantic feature extraction model used when language semantic feature extraction is performed are traced back to be jointly optimized, so as to obtain the optimized comparison model, the optimized visual feature extraction model and the optimized language semantic feature extraction model, wherein it should be clear that the pre-trained visual feature extraction model and the pre-trained language semantic feature extraction model are used respectively when the visual feature extraction and the language semantic feature extraction are performed in the above, and the two are jointly trained in the early stage, and the intelligent negative sample mining strategy is adopted to dynamically construct high-quality negative samples for the construction of training samples.

[0050] Embodiment Two

[0051] Please refer to Figure 4The application provides a fine-grained document image understanding system 1, comprising a memory 3, a processor 2, and a computer program stored in the memory 3 and capable of running on the processor 2, wherein the processor 2 implements the steps in the first embodiment when the computer program is executed.

[0052] The specific structure and variations of the system / device used for the method of the embodiments of the application can be understood by those skilled in the art based on the method of the embodiments of the application, and thus will not be described here again. The system / device used for the method of the embodiments of the application belongs to the scope of the application.

[0053] Those skilled in the art should understand that the embodiments of the application can be provided as a method, a system or a computer program product. Therefore, the application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0054] The application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions.

[0055] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps not listed in the claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In the claims, the word 'comprising' does not exclude the presence of other elements or steps than those listed in the claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In the claims, the word 'first','second', 'third', etc. does not imply any order. The terms 'first','second', 'third', etc. are to be understood as names of elements.

[0056] Moreover, it is to be understood that the description of the present application set forth herein is illustrative of the present application and is not intended to limit the scope of the present application as defined in the following claims. Various modifications of the preferred embodiment as described herein, will be apparent to those with ordinary skill in the art and can be made without departing from the spirit and scope of the application.

[0057] Although the preferred embodiment of the application has been described, those skilled in the art will be able to make modifications and alterations to this embodiment without departing from the spirit and scope of the application. Accordingly, it is intended that the scope of the application be governed by the following claims and their equivalents.

[0058] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A fine-grained document image understanding method, characterized in that, include: The input document image is obtained, the text content of the document image is recognized to obtain a text sequence, the pixel coordinates of each character in the text sequence are obtained, and the pixel coordinates are mapped to the text sequence to obtain a text coordinate mapping table. The text sequence is segmented into tokens using a BPE word segmenter to obtain all token-level texts. All token-level texts are then visually encoded according to the text coordinate mapping table to obtain all token-level visual encodings. Simultaneously, all token-level texts are extracted for linguistic semantic features to obtain all token-level linguistic semantic features. The document image is segmented into K image patches. Visual features are extracted from each image patch to obtain all patch-level visual features. The intersection of each patch-level visual feature and all token-level visual codes is calculated. The patch visual features in the intersection are mapped to all token-level linguistic semantic features to obtain mapped patch-level visual features and mapped token-level linguistic semantic features. The mapped patch-level visual features and mapped token-level linguistic semantic features are input into a comparison model for alignment to obtain the alignment result. The document image is understood based on the alignment result.

2. The fine-grained document image understanding method as described in claim 1, characterized in that, The step of performing token-level segmentation on the text sequence using the BPE tokenizer to obtain all token-level text includes: Obtain the basic vocabulary of the BPE word segmenter, and perform addition operations on the basic vocabulary including mathematical symbols, multilingual characters and document structure tags to obtain the basic vocabulary after the addition operations. Perform token-level word segmentation on the text sequence based on the basic vocabulary after the addition operations to obtain all token-level texts. All token-level texts include token-level texts that match the basic vocabulary after the addition operations and token-level texts that do not match the basic vocabulary after the addition operations. Token-level texts that do not match the basic vocabulary after the addition operation are taken as out-of-vocabulary words. The character length of each out-of-vocabulary word is calculated. The out-of-vocabulary words are optimized based on the character length. It is determined whether the character length is greater than the length threshold. If not, the out-of-vocabulary words are optimized by unknown word tagging to obtain optimized out-of-vocabulary words. If so, character-level encoding is used to optimize character decomposition of the out-of-vocabulary words to obtain optimized out-of-vocabulary words. Update all token-level texts based on the optimized out-of-vocabulary terms to obtain the updated token-level texts.

3. The fine-grained document image understanding method as described in claim 1, characterized in that, Simultaneously, linguistic semantic features are extracted from all token-level texts, resulting in all token-level linguistic semantic features including: The embedded character level is generated for each token-level text using a character-level embedding formula; The embedded sub-words corresponding to each token-level text are generated using sub-word-level embedding formulas. The semantic level of embedding is generated for each token-level text using a semantic-level embedding formula. By using a gating fusion mechanism, the embedded character level, embedded sub-word level, and embedded semantic level of each token-level text are dynamically fused to generate token-level linguistic semantic features for each token-level text, thus obtaining all token-level linguistic semantic features. The character-level embedding formula is: ; in, This represents the embedded character level corresponding to the i-th token-level text. This represents the character set of the i-th token-level text. This represents a character embedding matrix, where c represents a single character; The sub-word-level embedding formula is as follows: ; in, This represents the embedded sub-word level corresponding to the i-th token-level text. Represents the subword embedding matrix, Represents the line index of the i-th token-level text; The semantic-level embedding formula is: ; in, This represents the embedding semantic level corresponding to the i-th token-level text. This represents a multilayer perceptron. This represents the part-of-speech tag of the i-th token-level text. This represents the entity label of the i-th token-level text. This represents the one-hot encoding function. Symbols indicating feature splicing.

4. The fine-grained document image understanding method as described in claim 1, characterized in that, The step of performing patch segmentation on the document image to obtain K image patches includes: Obtain the image width, image height, and document type of the document image; The text density of the document image is calculated by inputting the image width and image height into the text density formula, which is: ; in, H represents the text density of document image I, H represents the image height, and W represents the image width. An indicator function that indicates whether a pixel (x, y) is a text pixel; The optimal scaling ratio of the document image is calculated based on the text density, and the document image is scaled according to the optimal scaling ratio to obtain the scaled document image. The patch size is dynamically adjusted according to the document type to perform adaptive patch segmentation on the scaled document image, resulting in K image patches.

5. The fine-grained document image understanding method as described in claim 1, characterized in that, The step of extracting visual features for each image patch to obtain all patch-level visual features includes: A convolutional neural network is used to extract internal local features from each image patch to obtain all patch-level local features; All patch-level local features are mapped to a fixed-dimensional embedding space through a linear projection layer to obtain all mapped patch-level local features. All mapped patch-level local features are then optimized to obtain all optimized patch-level local features. The optimization process includes residual connection and normalization.

6. The fine-grained document image understanding method as described in claim 1, characterized in that, The step of calculating the intersection of each patch-level visual feature and all token-level visual codes, and mapping the patch-level visual features in the intersection to all token-level linguistic semantic features, includes: Obtain the row index and column index of each image patch in the document image, input the row index and column index into the absolute position formula to calculate the absolute position encoding, so as to obtain the absolute position encoding of each image patch. At the same time, input the row index and column index of every two image patches into the relative position formula to calculate the relative position encoding, so as to obtain the relative position encoding of each image patch. The absolute position code of each image patch is weighted and fused with the corresponding relative position code to obtain a two-dimensional hybrid position code for each image patch. The two-dimensional hybrid position code is added to the patch-level visual feature corresponding to the image patch to enhance the patch-level visual feature, and all enhanced patch-level visual features are obtained. Obtain the sequence position information and document structure type of the token-level text corresponding to each token-level linguistic semantic feature in the text sequence; Simultaneously, local context information is calculated for each Token-level linguistic semantic feature. Local context encoding is performed within a preset association window, centered on each Token-level linguistic semantic feature, to obtain the local context information of each Token-level linguistic semantic feature. The local context information, sequence position information and document structure type of each token-level linguistic semantic feature are weighted and fused to obtain the context position code of each token-level linguistic semantic feature. The context position code is added to the corresponding token-level linguistic semantic feature to enhance the token-level linguistic semantic feature, and all enhanced token-level linguistic semantic features are obtained. Calculate the intersection of each enhanced patch-level visual feature with all token-level visual codes, and map all enhanced patch-level visual features in the intersection to all enhanced token-level linguistic semantic features.

7. The fine-grained document image understanding method as described in claim 6, characterized in that, The mapping of all enhanced patch-level visual features and all enhanced token-level linguistic semantic features in the intersection includes: The effective attention range of each enhanced patch-level visual feature is calculated by a sparse attention mechanism, and the enhanced patch-level visual features are associated within the effective attention range by a multi-head attention mechanism to obtain the patch-level visual features after all features are associated. The patch-level visual features, after associating all features, are subjected to multi-scale interpolation fusion to obtain multi-scale interpolated patch-level visual features. Different pooling strategies are triggered based on the spatial layout of all token-level visual codes to perform pooling processing on the multi-scale interpolated patch-level visual features, so as to realize the mapping of the multi-scale interpolated patch-level visual features with all enhanced token-level linguistic semantic features.

8. The fine-grained document image understanding method as described in claim 1, characterized in that, The step of aligning the mapped patch-level visual features with the mapped token-level linguistic semantic features in a comparison model to obtain an alignment result, and then using the alignment result to understand the document image, includes: Calculate the token-level similarity, phrase-level similarity, and sentence-level similarity between each mapped patch-level visual feature and each mapped token-level linguistic semantic feature; Based on the token-level similarity, phrase-level similarity, and sentence-level similarity, the mapped patch-level visual features and the mapped token-level linguistic semantic features are aligned at multiple levels to obtain multi-level alignment results. The multi-level alignment results are aligned and merged to obtain the final alignment result, and the document image is understood based on the final alignment result.

9. The fine-grained document image understanding method as described in claim 8, characterized in that, Also includes: The alignment quality score of the alignment result is calculated using the alignment quality formula, and the alignment consistency score of the alignment result is calculated using the alignment consistency detection formula. The alignment quality score and the alignment consistency score are input into the alignment scoring formula for scoring to obtain a scoring result. It is then determined whether the scoring result is lower than the scoring threshold. If so, the alignment model is optimized based on the scoring result to obtain an optimized alignment model. The alignment quality formula is: ; in, This represents the alignment quality score, where N represents the total number of mapped patch-level visual features and mapped token-level linguistic semantic features. This represents the patch-level visual feature after the i-th mapping. This represents the token-level linguistic semantic features after the i-th mapping. This represents the final similarity between the patch-level visual features after the i-th mapping and the token-level linguistic semantic features after the i-th mapping; The consistency detection formula is as follows: ; in, Indicates alignment consistency; Var represents variance. This represents the final similarity between the patch-level visual features after the i-th mapping and the token-level linguistic semantic features after the i-th mapping. denoted by , where represents the maximum variance, and N represents the total number of mapped patch-level visual features and mapped token-level linguistic semantic features; The alignment scoring formula is as follows: ; in, Indicates the scoring result. The quality weighting coefficient represents the alignment quality score. The consistency weighting coefficient represents the degree of alignment consistency. Indicates alignment quality. Indicates the degree of alignment consistency.

10. A fine-grained document image understanding system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cross-modal information retrieval method based on pre-training model

    CN116796047A

  • Long text similarity calculation method and device based on semantic progressive fusion

    CN117113094A

  • Code search system and method based on pre-training model

    CN117992572A

  • Zero sample image classification method and device

    CN119600643A

  • Lao image text recognition method and device fused with language error correction model

    CN119785367A

Cited By

  • Text image and formula image unified identification method and system, storage medium and equipment

    CN121366423A