Table understanding system and method based on large-scale visual language model

By employing cross-modal attention computation and optical character recognition technology, irrelevant image regions are erased, solving the problem of visual attention deviation in large-scale visual language models during table comprehension and improving the accuracy and robustness of table reasoning.

CN120913233APending Publication Date: 2025-11-07NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511041292.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Large-scale visual language models are easily affected by visual redundancy when understanding tables, causing visual attention to deviate from key information and affecting the performance of table reasoning tasks.

Method used

A table understanding system based on a large-scale visual language model is adopted. By combining cross-modal attention calculation, observational contribution calculation, counterfactual contribution calculation and causal inference analysis with optical character recognition technology, irrelevant image areas are erased, thereby improving the model's attention to and accuracy of table content.

Benefits of technology

It significantly reduces the interference of irrelevant image regions on the model's attention, improving the model's accuracy and robustness in the table semantic understanding task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913233A_ABST
    Figure CN120913233A_ABST
Patent Text Reader

Abstract

The invention provides a table understanding system and method based on a large-scale visual language model, and relates to the technical field of natural language processing. The method comprises the steps of firstly obtaining a target table image and a corresponding question text, and performing preprocessing to obtain an image block and a text mark; inputting into a large-scale visual language model to obtain an attention matrix; estimating a cross-modal attention score of the problem text to the image based on the attention matrix; obtaining an interference problem text, and inputting the original table image and the interfered problem text into the large-scale visual language model again; calculating a causal effect score of the question on image attention, and calculating a corresponding question correlation degree for each text element by adopting a space weighted average method; erasing image areas corresponding to the text elements of which the problem relevancy is lower than a specific threshold value; and inputting the processed table image and the original problem into a large-scale visual language model to obtain an optimized reasoning result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, and particularly relates to a table understanding system and method based on a large-scale vision language model. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, visual understanding has shown broad prospects in various practical applications. As an important branch of visual document understanding, visual table understanding aims to extract and understand the content of image-forming tables. With the emergence of large-scale vision language models (LVLMs) such as GPT-4V, researchers have begun to explore how to apply the powerful cross-modal reasoning capabilities of these models to structured document understanding. These models are usually pre-trained based on image-text pairs and have the ability to answer questions, fill in the blanks, and reconstruct table content, providing a new path for semantic understanding and question answering of complex tables. Although existing methods have made some progress, when dealing with complex table pictures, the model is easily disturbed by redundant content within the table, causing the visual attention to deviate from the key information, resulting in poor performance of table reasoning tasks. SUMMARY

[0003] In view of the deficiencies of the prior art, the present application provides a table understanding system and method based on a large-scale vision language model, which provides a table picture processing system and method to solve the problem of visual attention deviation in understanding tables by a large-scale vision language model.

[0004] In one aspect, a table understanding system based on a large-scale vision language model includes the following modules: an image and question input module, a cross-modal attention calculation module, an observational contribution calculation module, an counterfactual contribution calculation module, a causal inference analysis module, an image processing module, and a reasoning module.

[0005] The image and question input module receives an input target table image and a question text corresponding to the target table image and performs a preprocessing operation. The preprocessing specifically includes dividing the target table image into a plurality of fixed-size image blocks and decomposing the question text into a series of discrete text tokens using a text tokenizer.

[0006] The cross-modal attention calculation module specifically inputs the image blocks and text tokens obtained by preprocessing into a large-scale vision language model LVLM to generate output text tokens in a self-recurrent manner, and calculates cross-modal attention scores of the text to the image during the generation process.

[0007] The calculation of the cross-modal attention scores analyzes the attention degree of the question tokens to the visual tokens by aggregating the cross-modal attention weights in multiple decoder layers. Let the cross-modal attention weight of the i-th decoder layer be denoted as The cross-modal attention weight in the layer decoder is wherein respectively represent the number of question text tags, represent the number of visual image blocks, the attention weight is averaged on all decoder layers and all question tag dimensions, and the calculation method is as follows:

[0008] ;

[0009] wherein, represents the cross-modal attention distribution attention weight corresponding to the i-th question tag in the j-th layer, is the total number of layers of the model decoder, represents the number of question text tags, represents the cross-modal attention score from the question text to the image , which is used to reflect the attention degree of the model to each image region in the generation process; The observation contribution calculation module: in the factual scenario, based on the original input target table image v and question text t, the cross-modal attention calculation module is called to obtain the cross-modal attention score of the question to the image , denoted as the observation attention contribution.

[0010] The counterfactual contribution calculation module: in the counterfactual scenario, the question text is replaced with a space tag

[0011] , the cross-modal attention calculation module is called to obtain the cross-modal attention score under the interference text and the original table image input condition , denoted as the counterfactual attention contribution; The causal inference analysis module: the factual scenario and the counterfactual scenario are introduced, the observation attention contribution and the counterfactual attention contribution are obtained, and then the causal effect of the question to the cross-modal attention score is calculated; the causal effect is defined as follows:

[0012] ;

[0013] ;

[0014] wherein, represents the causal effect score of the question to the image attention, which is obtained by subtracting the counterfactual attention contribution from the observation attention contribution;

[0015] The image processing module erases the image region irrelevant to the question and answer according to the causal effect score;

[0016] ​Specifically, the image blocks are refined into masks at the text element level to indicate table areas relevant to the question; optical character recognition (OCR) technology is used to locate text elements in the table image, obtaining a set of text elements. Then, the causal effect score for each text element region is calculated.

[0017] To calculate the causal effect score for each OCR text element, a spatially weighted average method is used to aggregate the causal effects of multiple image patches. The aggregation weight is proportional to the overlapping area of ​​the image patches within the text region. The final causal effect score is defined as follows:

[0018] ;

[0019] in, and Represents text elements and image block regions for optical character recognition, with i and j being the corresponding indices; This is a function for calculating the area of ​​an image region. Represents image blocks Text element r covering OCR detection i The area percentage ranges from [0, 1]. For the set of all image patches;

[0020] Then, the aggregated causal effect score is below the threshold. text elements Perform an erase operation, which involves filling the corresponding area with the background color, and setting a threshold. The causal effect score was obtained by averaging the scores of all image patches.

[0021] ;

[0022] in The number of all image blocks. Score the causal effect for the j-th image patch. Finally, based on the threshold... Remove irrelevant content from the table image to obtain the table image. .

[0023] The reasoning module will display the table image. With the original question text Inputting these into a large-scale visual language model generates more accurate answers and improves reasoning accuracy.

[0024] On the other hand, a table understanding method based on a large-scale visual language model, implemented based on the aforementioned table understanding system based on a large-scale visual language model, includes the following steps:

[0025] Step 1: Obtain the target table image and the corresponding question text;

[0026] Step 2: Preprocess the target table image and the corresponding question text to obtain image blocks and text labels;

[0027] Step 3: Input the preprocessed image blocks and text labels into large-scale visual language models (LVLMs) to obtain attention matrices;

[0028] Step 4: Estimate the aforementioned question text based on the attention matrices The cross-modal attention score of the image ; ;

[0029] Step 5: Record the cross-modal attention score as an observed attention contribution;

[0030] Step 6: Obtain the interference question text, input the original table image and the interference question text again into the large-scale visual language model, calculate the cross-modal attention score , and record it as a counterfactual attention contribution;

[0031] Step 7: Calculate the difference between the observed attention contribution and the counterfactual attention contribution as the causal effect score of the question on the image attention;

[0032] Step 8: Use optical character recognition technology to extract text elements in the table;

[0033] Step 9: According to the causal effect score and the optical character recognition text elements, use a spatially weighted average method to calculate the corresponding question relevance of each text element;

[0034] Step 10: Erase the image area corresponding to the text element whose question relevance is lower than a certain threshold;

[0035] Step 11: Input the processed table image and the original question into the large-scale visual language model to obtain the optimized reasoning result.

[0036] The beneficial effects produced by the above technical solutions are:

[0037] The present application provides a table understanding system and method based on large-scale visual language models. The present application can significantly reduce the interference of irrelevant image areas on model attention, thereby improving the accuracy and robustness of the model in table semantic understanding tasks. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The present application provides a table understanding system and method based on large-scale visual language models. The present application can significantly reduce the interference of irrelevant image areas on model attention, thereby improving the accuracy and robustness of the model in table semantic understanding tasks.

[0039] Figure 2 Causal relationship diagram for the present application;

[0040] wherein (a) - observed attention score, (b) - counterfactual attention score, (c) - causal effect score;

[0041] Figure 3 An apparatus structure diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0042] The specific embodiments of the present application are described in further detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application, but are not used to limit the scope of the present application.

[0043] In one aspect, a table understanding system based on a large-scale visual language model, as shown in Figure 1 includes the following modules: image and question input module, cross-modal attention calculation module, observed contribution calculation module, counterfactual contribution calculation module, causal inference analysis module, image processing module, reasoning module.

[0044] The image and question input module receives the input target table image and the question text corresponding to the target table image, and performs a preprocessing operation; the preprocessing is specifically: dividing the target table image into a plurality of fixed-size image blocks (image patches); for example, for a table image with a resolution of 224x224 pixels, it can be divided into 16x16 image blocks, thereby obtaining 196 image patches. At the same time, the question text is decomposed by a text tokenizer to convert it into a series of discrete text tokens; for example, for the question "Which expenditure is the most in 2022 in this table?", the tokenizer may decompose it into: ["this", "table", "in", "2022", "year", "which", "expenditure", "is", "the", "most", "?"] and further mapped to the ID in the word table.

[0045] The cross-modal attention calculation module is specifically for inputting the image blocks and text tokens obtained by preprocessing into the large-scale visual language model LVLMs to generate output text tokens in a self-recurrent manner, and calculating the cross-modal attention score of the text to the image during the generation process.

[0046] In a typical architecture of LVLMs, the model is composed of multiple stacked Transformer modules. The self-attention mechanism (Self-Attention) contained in each module internally generates query (Query), key (Key) and value (Value) vectors by performing linear transformation on the input embedding. For any query vector, the attention mechanism calculates the similarity scores between it and all key vectors, and uses these scores as weights to weight-sum the corresponding value vectors, thereby generating attention output; the specific calculation method is as follows: Wherein Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Wherein, L represents the sequence length, d represents the dimension of the attention head; the obtained attention weight matrix, i.e. the attention score describes the attention allocation relationship between positions in the sequence; The function is defined as , which is used to convert the unnormalized attention score into a probability distribution.

[0047] Based on the above attention mechanism, the module further proposes a cross-modal attention score calculation method for measuring the attention degree from the question label to each image block. Specifically, the calculation of the cross-modal attention score globally analyzes the attention degree of the question label to the visual label by aggregating the cross-modal attention weights in multiple decoder layers; let the cross-modal attention weight in the i-th decoder layer be Wherein Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all L represents the sequence length,

[0048] ;

[0049] Wherein, Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all Q, K and V represent the query matrix, the key matrix and the value matrix respectively, and the dimensions are all

[0050] The observation contribution calculation module: in the factual scenario, based on the original input target table image v and question text t, calling the cross-modal attention calculation module to obtain the cross-modal attention score of the question to the image , denoted as observation attention contribution.

[0051] The counterfactual contribution calculation module: in the counterfactual scenario, replacing the question text with a space token , destroying the semantics of the question text , calling the cross-modal attention calculation module to obtain the cross-modal attention score under the condition of the interference text and the original table image input , denoted as counterfactual attention contribution; this score reflects the cross-modal attention prior bias of the model, causing the model to deviate from the table content in reasoning.

[0052] The causal inference analysis module: introducing the factual scenario and the counterfactual scenario, obtaining the observation attention contribution and the counterfactual attention contribution, and then calculating the causal effect of the question to the cross-modal attention score to estimate the real attention of the question text to the image block. The causal effect is defined as follows:

[0053] ;

[0054] Wherein, represents the causal effect score of the question to the image attention, which is obtained by subtracting the observation attention contribution from the counterfactual attention contribution; the causal effect score can effectively alleviate the cross-modal prior bias and provide a more accurate estimate of the correlation between the image block and the question.

[0055] The image processing module erases the image area irrelevant to the question and answer according to the causal effect score;

[0056] Specifically, a single image block may span multiple rows and columns, potentially overlapping multiple table cells. To solve this problem, the image block is refined into a more detailed mask with text elements as the granularity, which is used to indicate the table area related to the question; considering the structured layout of the table, we use the optical character recognition technology OCR to locate the text elements (including Chinese characters, letters, numbers and symbol composed character sequence) in the table image to obtain the text element set , and then calculate the causal effect score of each text element area.

[0057] In order to calculate the causal effect score of each OCR text element, the spatial weighted average method is used to aggregate the causal effects of multiple image blocks, and the aggregation weight is proportional to the overlapping area of the image block in the text area; the final causal effect score is defined as follows:

[0058] ;

[0059] wherein, and denote the text elements and image block regions of the optical character recognition, i and j are the corresponding indices; is a function to calculate the area of an image region; denotes the image block covering the text element r i of the OCR detection, whose area proportion is in the range [0, 1]; is the set of all image blocks. This weighting scheme ensures that the image blocks with a larger proportion in the text element have a greater contribution to the final causal effect score .

[0060] Then, for the text elements with an aggregated causal effect score below a threshold , an erasing operation is performed, i.e., filling the corresponding region with a background color, effectively removing the irrelevant content to the question. This erasing operation realizes the control of the content at the cell level in the table, which is based on the aforementioned obtained causal effect score. The threshold is obtained by averaging the causal effect scores of all image blocks:

[0061] ;

[0062] wherein is the number of all image blocks, is the causal effect score of the jth image block. Finally, according to the threshold , the irrelevant content is erased from the table image, obtaining the table image .

[0063] The inference module inputs the table image and the original question text to an existing large-scale visual language model (such as Qwen2-VL) together, thereby generating more accurate answers and improving the inference accuracy.

[0064] On the other hand, a table understanding method based on a large-scale visual language model, based on the aforementioned table understanding system based on a large-scale visual language model, comprises the following steps:

[0065] Step 1: Obtain the target table image and the corresponding question text;

[0066] Step 2: Preprocess the target table image and the corresponding question text to obtain image blocks and text labels;

[0067] Step 3: Input the preprocessed image patches and text tags into large-scale visual language models (LVLMs) to obtain the attention matrix;

[0068] Step 4: Estimate the aforementioned problem text based on the attention matrix. For images Cross-modal attention score ;

[0069] Step 5: Record the cross-modal attention score as the contribution of observational attention;

[0070] Step 6: Obtain the text containing interfering information and extract the original table image. Problem text after interference Input the large-scale visual language model again and calculate the cross-modal attention score. And record it as a contribution to counterfactual attention;

[0071] Step 7: Calculate the difference between observational attention contribution and counterfactual attention contribution as the causal effect score of the question on image attention;

[0072] Step 8: Use optical character recognition technology to extract text elements from the table;

[0073] Step 9: Based on the causal effect score and the text elements after optical character recognition, calculate the question relevance of each text element using a spatially weighted average method; for example... Figure 2 The diagram shown is a schematic diagram of causal relationships in this embodiment; Figure (a) is the observational attention score, Figure (b) is the counterfactual attention score, and Figure (c) is the causal effect score.

[0074] Step 10: Erasure the image regions corresponding to text elements whose relevance to the question is below a certain threshold;

[0075] Step 11: Input the processed table image and the original question into a large-scale visual language model to obtain the optimized inference results.

[0076] An embodiment of the apparatus for processing table images provided by this invention can be applied to any electronic device with data processing capabilities, such as a computer, server, or embedded device. The apparatus can be implemented in software, or through hardware circuitry, or a combination of hardware and software.

[0077] In the case of software implementation, the apparatus can be regarded as logical function modules whose functions are executed by the processor in the device storing the computer program instructions in the non-volatile memory and realized by loading the instructions into the memory and running. The processor realizes the processing of the table image and the question text, the analysis of the attention matrix, the causal effect evaluation, the OCR recognition and the image region erasing and other operations by executing the program, so as to complete the task flow of the whole table semantic understanding and the enhanced reasoning.

[0078] From the perspective of hardware structure, as shown in Figure 3 The apparatus disclosed by the present application can be deployed in a device with data processing capability including a processor, a memory, a network interface and a non-volatile memory. In addition, according to different application scenarios, the device can further include other hardware components, such as an input module, an output module, an image acquisition module, an image acceleration unit, etc., which will not be described here.

[0079] The above description is only the preferred embodiment of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions.

Claims

1. A table understanding system based on a large-scale visual language model, characterized by, The method comprises the following modules: an image and question input module, a cross-modal attention calculation module, an observational contribution calculation module, an counterfactual contribution calculation module, a causal inference analysis module, an image processing module, and an inference module. The image and question input module receives an input target table image and a question text corresponding to the target table image, and performs a preprocessing operation. The cross-modal attention calculation module specifically inputs the image blocks and text tokens obtained through preprocessing into large-scale visual language models (LVLMs) to generate output text tokens in a self-recurrent manner, and calculates cross-modal attention scores during the generation process. The observation contribution calculation module: in the fact scenario, based on the original input target table image v and question text t, call the cross-modal attention calculation module to obtain the cross-modal attention score of the question to the image , recorded as observation attention contribution; The counter-factual contribution calculation module: in the counter-factual scene, replace the question text with a space mark , call the cross-modal attention calculation module to obtain the cross-modal attention score under the interference text and the original table image input condition , denoted as counter-factual attention contribution; The causal inference analysis module introduces factual and counterfactual scenarios, obtains observational and counterfactual attention contributions, and then calculates the causal effect of the question on the cross-modal attention score. The image processing module erases image regions irrelevant to the question and answer according to the causal effect score. The reasoning module inputs the table image together with the original question text into a large-scale visual language model, thereby generating more accurate answers and improving reasoning accuracy.

2. The table understanding system based on a large-scale visual language model according to claim 1, wherein, The calculation of the cross-modal attention score analyzes the attention of the question tokens to the visual tokens through the aggregation of cross-modal attention weights in multiple decoder layers.

3. The table understanding system based on a large-scale visual language model according to claim 2, wherein, Set the Cross-modal attention weights in the layer decoder are where Np and Nv represent the number of question text tokens and the number of visual image patches, respectively, The attention weights are averaged over all decoder layers and all question token dimensions, computed as follows: ; wherein, represents the layer corresponding to the cross-modal attention distribution attention weight of the total number of layers of the model decoder, represents the number of question text tokens, represents the cross-modal attention score from the question text to the image , used to reflect the attention degree of the model to each image region in the generation process.

4. The table understanding system based on a large-scale visual language model according to claim 1, wherein, The causal effect is defined as follows: ; wherein, denotes the causal effect score of the question on the image attention, obtained by differencing the observed attention contribution and the counterfactual attention contribution.

5. The table understanding system based on a large-scale visual language model according to claim 1, wherein, The image processing module specifically refines the image block into a mask with text elements as granularity, which is used to indicate the table area related to the question; specifically, an optical character recognition technology (OCR) is adopted to locate the text elements in the table image, and a text element set is obtained Then, a causal effect score of each text element area is calculated.

6. The table understanding system based on a large-scale visual language model according to claim 5, wherein, To calculate the causal effect score of each OCR text element, the spatial weighted average method is used to aggregate the causal effects of multiple image blocks, and the aggregation weight is proportional to the overlapping area of the image block in the text region. ; wherein, and denote the text elements and image block regions of the optical character recognition, i and j are the corresponding indices; is a function to calculate the area of an image region; denote the image block coverage of the text elements r i detected by OCR, ranging in [0, 1]; is the set of all image blocks; Then, the text elements whose aggregated causal effect scores are below a threshold are removed An erasing operation is performed, i.e. the corresponding area is filled with the background color, the threshold is obtained by averaging the causal effect scores of all picture tiles: ; wherein is the number of all picture blocks, is the causal effect score of the j-th image block; finally, according to a threshold the table picture is erased of irrelevant content, resulting in a table image .

7. A table understanding method based on a large-scale visual language model, implemented by the table understanding system based on the large-scale visual language model in claim 1. The method comprises the following steps: Step 1: Obtain a target table image and a corresponding question text. Step 2: Preprocess the target table image and the corresponding question text to obtain image blocks and text tokens. Step 3: Input the preprocessed image blocks and text tokens into large-scale visual language models (LVLMs) to obtain an attention matrix. Step 4: estimating the preceding question text based on the attention matrix to the image cross-modal attention score ; Step 5: Record the cross-modal attention score as the observational attention contribution. Step 6: Obtain the interference problem text, and the original table image and the interference problem text Input the large-scale visual language model again, calculate the cross-modal attention score and record it as the counterfactual attention contribution; Step 7: Calculate the difference between the observational attention contribution and the counterfactual attention contribution as the causal effect score of the question on the image attention. Step 8: Use optical character recognition technology to extract text elements from the table. Step 9: Calculate the question relevance of each text element using the spatial weighted average method based on the causal effect score and the optical character recognition text elements. Step 10: Erase the image regions corresponding to the text elements with a question relevance lower than a certain threshold. Step 11: Input the processed table image and the original question into the large-scale visual language model to obtain an optimized inference result.